Skip to main content

Generate your anomaly test with Elementary AI

Let our Slack chatbot create the anomaly test you need.
elementary.dimension_anomalies The test counts rows grouped by given dimensions (columns/expressions). This test practically monitors the frequency of values in the configured dimension over time, and alerts on unexpected changes in the distribution. It is best to configure it on low-cardinality fields. If timestamp_column is configured, the distribution is collected per time_bucket. If not, it counts the total rows per dimension. Each value is compared against its own past, with its own expected range. So a value that drops to zero gets caught even when the table’s total row count barely changes.

Understanding dimension anomalies

What happens when you list more than one column, how to think about detection_delay, and why a long detection period can hide problems.

Test configuration

Required configuration: dimensions
data_tests:
  — elementary.dimension_anomalies:
    arguments:
      dimensions: sql expression
      timestamp_column: column name
      where_expression: sql expression
      anomaly_sensitivity: int
      anomaly_direction: [both | spike | drop]
      detection_period:
        period: [hour | day | week | month]
        count: int
      training_period:
        period: [hour | day | week | month]
        count: int
      time_bucket:
        period: [hour | day | week | month]
        count: int
      seasonality: day_of_week
      detection_delay:
        period: [hour | day | week | month]
        count: int
      ignore_small_changes:
        spike_failure_percent_threshold: int
        drop_failure_percent_threshold: int
      anomaly_exclude_metrics: [SQL expression]
      exclude_final_results: [SQL expression]