Skip to content

Anomalies

The Anomalies report lists unusual changes in service quality and invalid measurement values. Use it to check whether a problem affects one client, several clients, or a measured service.

You can reach the Anomalies report from Reports > Anomalies in the sidebar.

Note

This entry only appears when anomaly detection is enabled for your organization.

How anomaly detection works

A background process runs every 15 minutes. It builds a baseline from a rolling window of recent measurements, seven days by default, and compares current values with that baseline.

Each client has a separate baseline for every service and target host. A client on a slow connection is flagged when it becomes slower than usual, not because it is slower than another client.

Detection uses two complementary passes:

  • Spike uses a short window, around one hour depending on measurement frequency, to find sudden changes.
  • Shift uses a longer window of around six hours to find sustained drift.

A separate data quality pass finds values outside the valid range for a metric, such as a MOS score above 5 or a negative latency. These usually indicate a measurement problem.

Probe data and Player SDK data

The detector handles probe and Player SDK data. Probe measurements come from devices running Surfmeter Automator. Player SDK measurements come from end-user video players and are pooled by domain. Player SDK anomalies are therefore reported as fleet-wide. Per-probe and cross-probe detection apply only to probe data.

Comparing a probe to itself, and to its peers

Per-probe detection compares each probe with its own history. It finds changes local to that probe.

Cross-probe detection compares that change with other probes measuring the same target. A probe that differs from its peers likely has a local problem.

When a service degrades for every probe, no single probe stands out and no cross-probe anomaly is raised. Cross-probe detection needs at least three probes measuring the same target.

Diagnosis

Each probe anomaly has a diagnosis based on the per-probe and cross-probe results:

  • Probe-local – one probe is affected on one target. Suggests an issue specific to that probe reaching that target, such as its routing or DNS for that one service.
  • Probe infrastructure – one probe is affected across several targets. Suggests a problem with that probe's own connection or hardware rather than any one service.
  • Service-wide – several probes are affected on the same target. Suggests a problem with that service, or the network that delivers it.
  • Network-wide – several probes are affected across several targets at once. Suggests a broad upstream or internet-provider problem.

Episodes

An episode tracks one anomaly from its first detection until resolution. Later 15-minute checks extend the same episode instead of creating duplicates. The episode closes after the anomaly stops appearing.

Email notifications

Under Settings > Notifications, enable Anomaly events and choose immediate or digest delivery. Immediate delivery sends new episodes after the detection cycle. Digest delivery batches episodes on the schedule you select.

To stop one client from producing new anomaly events, open its detail page and select Notifications > Mute anomalies. The client still contributes to fleet baselines and cross-probe comparison. See Notification preferences.

Anomaly list

The list page shows all detected anomaly episodes within the selected time range.

Time range and interval

Use the time range picker and interval selector at the top to adjust the reporting period. The list shows all episodes that overlap the selected window, including long-running episodes that started before the window but are still active within it.

Filters

Use the filters to narrow the list by status, detection details, measurement subject, client, or location.

Overview:

  • Status – Active (new) or closed (resolved) episodes, or both. Defaults to new so current issues show first.
  • Severity – Warning or critical. Critical events are those that escalated by persisting, or that started with an extreme deviation.
  • Diagnosis – Where the problem most likely sits: Probe-local, Probe infrastructure, Service-wide, or Network-wide (see Diagnosis).
  • Source – Whether the anomaly comes from our own probes or from Player SDK data (real end-users).

Detection (how it was found):

  • Detection Type – Per-probe, Cross-probe, Fleet-wide, or Data quality.
  • Detection Pass – Spike, shift, or data quality.
  • Detection Method – The statistical method behind the flag (MAD, IQR, percentile, proportion, range check).
  • Confidence – High, medium, or low, reflecting how much the signal can be trusted given the number of measurements and how many probes were available to compare against.
  • Deviation Direction – Whether the value went above or below normal (or above/below the physical bounds for data quality events).

Subject (what was measured):

  • Measurement Type – Video, Web, Network, Speedtest, or Conferencing.
  • Subject – The service (e.g. "netflix", "youtube"). For network measurements this is the technology.
  • Statistic Name – The specific metric (e.g. p1203_overall_mos, download).
  • Domain – Target domain.
  • Hostname – Target hostname.

Client and location:

  • Client Label – Client device name.
  • ISP – Internet provider name.
  • Country / City – Where the probe is located.

You can hover over any filter badge to see a plain-language explanation of what the value means and how it relates to detection.

Note

Client and location filters (ISP, country, city) only match per-probe anomalies. Fleet-wide (Player SDK) anomalies are aggregated across a whole domain and do not carry these fields.

Timeline chart

A stacked bar chart shows the number of new anomaly episodes over time, broken down by a grouping dimension you can choose from a dropdown. For example, group by subject to see which services generated the most anomalies, by severity to see the ratio of warnings to critical events, or by diagnosis to see at a glance whether issues are mostly probe-side or service-side.

The chart buckets episodes by their start time (first_seen_at), so a long-running episode contributes one bar at its onset rather than appearing in every interval.

Events table

Below the chart, a paginated table lists individual anomaly episodes. By default, episodes are sorted by last activity (most recently active first). Each row shows:

  • Status and severity badges
  • The diagnosis (Probe-local, Service-wide, and so on; data-quality and fleet-wide events show their detection type instead)
  • When the episode started and was last seen (or resolved)
  • How many times the anomaly re-fired (occurrence count)
  • The measurement type, service subject or domain, and client
  • The affected metric and metric vs. baseline values
  • A brief explanation of the anomaly

Click any row to open the detail page for that episode.

Resolving anomalies

You can manually mark anomaly episodes as resolved directly from the list. Select one or more episodes using the row checkboxes, then click the Resolve selected button in the toolbar above the table. A confirmation dialog asks you to confirm the action. Once resolved, the episodes move out of the default "new" filter view.

This is useful for known false positives, expected maintenance windows, or other cases where the automatic cleanup has not yet closed the episode.

Note

Only users with the admin, editor, or organization admin role can resolve anomalies.

Anomaly detail page

Click an anomaly row to open its detail page.

Info card

The top section shows:

  • Severity and status badges
  • A plain-language explanation of what was detected
  • The measurement type, subject, and hostname/domain
  • Timeline: when the episode started, when it was last seen or resolved, and how many times it re-fired
  • Detection metadata: detection type, pass, method, source, confidence, the diagnosis (for probe anomalies), and sample count. Hover over each badge for a brief explanation of what the value means.

Origin card

Shows the client that triggered the anomaly (linked to the client's detail page), along with ISP and location information when available. Fleet-wide anomalies display a "Fleet-wide" badge instead of a client link.

Metric vs. baseline card

A visual comparison of three key values:

  • Metric value – The measured value that triggered the anomaly, highlighted in red with a directional arrow.
  • Baseline value – The expected value learned from recent history. Hover for an explanation of what the baseline represents for this detection type.
  • Threshold – The cutoff that was crossed. For per-probe and fleet-wide anomalies, this is the baseline plus a multiple of the recent spread; for data quality anomalies, it is an absolute physical bound. Cross-probe anomalies are judged against the other probes rather than a fixed cutoff, so they show no threshold.

KPI context chart

A time-series chart of the affected metric, centered on the episode window. The chart shows the same metric for the same client and service, spanning from before the episode started to after it was last seen. This lets you see the degradation in context: the normal behavior before the episode, the anomaly itself, and the recovery (if any).

Affected measurements table

Below the chart, a table lists the individual measurements that fell within the episode window and scope. This mirrors what you would see in the Measurements Explorer if you filtered to the same client, service, and time range.

Actions

From the detail page you can:

  • View in Explorer – Jump to the Measurements Explorer pre-filtered to the same client, service, and time window.
  • Resolve – Mark this single episode as resolved (same confirmation flow as the list page).
  • Copy JSON – Copy the raw anomaly event data to your clipboard for further analysis.

Severity levels

Anomalies are classified as either warning or critical:

  • Warning is the initial severity for any anomaly that crosses the detection threshold.
  • Critical means the anomaly has either shown an extreme initial deviation or has persisted across several checks (roughly 2.5 hours for spike events, 3 hours for shift events), indicating a sustained issue rather than a passing blip. Once an episode reaches critical severity, it stays critical even if the deviation eases, preventing flapping between severity levels.

Episode lifecycle

Each anomaly is tracked as an episode with the following lifecycle:

  • New – The episode is active. The anomaly has been detected and is still appearing in later checks.
  • Resolved – The episode is closed. This happens automatically when the anomaly stops appearing for a configurable timeout, 60 minutes by default. You can also resolve episodes from the list or detail page.