Get started Dashboard
Alert triage ·
Part 1 of Getting off call: alerts and incidents for small teams

How to group noisy alerts into one incident

Explore with AI

Decide which labels mean the same problem, such as service, cluster and environment, and group on them in your alert router so related firings join one notification. In Prometheus Alertmanager that means group_by on each route, sensible group_wait, group_interval and repeat_interval timers, and inhibition rules that mute downstream symptoms. Then key the incident record on the alert group's key, so every new firing updates the incident that is already open.

On this page

To group noisy alerts into one incident, first decide which labels mean “the same problem”. Then group on those labels in your alert router, and key the incident on that group so every new firing updates the record that is already open. In Prometheus Alertmanager this comes down to three things: group_by on each route, three timers that decide when the grouped notification goes out, and inhibition rules that mute the symptoms of a failure you already know about.

Grouping only works if alerts carry the right labels when they fire, so most of the work sits in your alerting rules. The steps below take one service from a page per instance to one incident per real problem. They cover the config, the commands to test it, and the settings that quietly undo it.

What grouping alerts into one incident means

Five mechanisms get lumped together under “noise reduction”. Each does a different job, and you usually need more than one.

  • Deduplication collapses repeat firings of the same alert. Alertmanager deduplicates the alerts it receives, and its webhook payload gives each alert a fingerprint that identifies it. So an alert sent again on every evaluation stays one alert.
  • Grouping batches different alerts that share values for chosen labels into one notification. The Alertmanager docs give the example of alerts for cluster=A and alertname=LatencyHigh batched into a single group.
  • Inhibition mutes some alerts while another alert is firing. For example, it can mute every service alert in a cluster while the alert saying the whole cluster is unreachable is active.
  • Silences mute alerts that match a set of matchers for a fixed time, for planned work.
  • Correlation links alerts with different labels that share a cause: a dependency, a deploy, a time window. Label grouping can’t do this alone. You either encode the relationship in labels or use a tool that knows your topology.

An incident is the record a person works. The goal is one incident per underlying problem, with every alert that belongs to it attached as evidence.

Why one fault fires dozens of alerts

Most alert storms come from rules written per instance. The Alertmanager docs describe a network partition where half of a service’s instances can no longer reach the database. A rule that alerts per instance then sends hundreds of alerts for one fault. In a large outage, hundreds to thousands of alerts can be firing at once.

The second source is the cascade. A struggling database host raises CPU, disk I/O and latency alerts on itself, and the services that depend on it raise error rate alerts of their own. Each alert is correct. Each one names a different metric, so exact-match deduplication leaves them all standing.

The third is flapping. A metric that hovers around its threshold resolves and fires again, and each new firing can look like a new problem.

Choose the labels that mean “the same problem”

The group key is the set of label values you group on, and every alert with the same values lands in the same group. Choosing it is a decision about your failure domains, so write that decision down before you touch the config.

  • Too narrow: grouping on alertname and instance gives one group per instance, which is the storm you started with.
  • Too wide: grouping on alertname alone merges high CPU in production with high CPU in staging, or two clusters failing for different reasons.
  • A sensible default: service, cluster (or region) and environment. One group then means one service having trouble in one place, whatever the individual alerts are called.

Leave instance out of group_by. The instances stay listed in the notification, and they all fold into one group.

None of this works unless every rule uses the same label names. I built and led the Workers observability team at Cloudflare, and the habit I’d push hardest is agreeing on those names once and enforcing them in review. A rule that writes svc where the others write service quietly forms a group of its own.

Grouping alerts in Prometheus Alertmanager, step by step

  1. Label every alerting rule at the source. Add the grouping labels, a severity and an owning team. Use for so short blips stay pending, and keep_firing_for so a brief dip below the threshold doesn’t resolve the alert and fire it again.
# rules/checkout.yml
groups:
- name: checkout
  labels:
    service: checkout
    team: payments
  rules:
  - alert: CheckoutDBUnreachable
    expr: checkout_db_up == 0
    for: 5m
    keep_firing_for: 5m
    labels:
      severity: critical
      database: orders-db
    annotations:
      summary: "{{ $labels.instance }} cannot reach orders-db"
      runbook: https://runbooks.example.com/checkout-db
  - alert: CheckoutHighErrorRate
    expr: job:checkout_errors:ratio5m > 0.05
    for: 10m
    labels:
      severity: critical
      database: orders-db

Labels set at group level apply to every rule in the group. With for: 5m, Prometheus checks that the condition holds on every evaluation for 5 minutes before the alert fires. Until then the alert stays pending.

  1. Set group_by on the routes. The root route catches everything. Child routes inherit its settings and can override them.
# alertmanager.yml
route:
  receiver: default
  group_by: [alertname, cluster, service]
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h
  routes:
    - matchers:
        - team="payments"
      receiver: payments-pager
      group_by: [cluster, service]

The payments route drops alertname from the key. CheckoutDBUnreachable and CheckoutHighErrorRate in the same cluster then arrive as one notification: one incident per service per cluster.

  1. Add inhibition for the cascades you know about. The inhibition section below has the rules and the labels they depend on.

  2. Check the config and test the routing without firing anything.

amtool check-config alertmanager.yml
amtool config routes test --config.file=alertmanager.yml --tree \
  --verify.receivers=payments-pager team=payments service=checkout cluster=eu-1

A valid file prints SUCCESS and lists what it found:

Checking 'alertmanager.yml'  SUCCESS
Found:
 - global config
 - route
 - 2 inhibit rules
 - 2 receivers
 - 0 templates

With --verify.receivers, amtool returns error code 1 when the label set reaches a different receiver, so the routing test can run in CI.

  1. Reload Alertmanager. Send SIGHUP to the process or POST to the reload endpoint. If the new file is malformed, Alertmanager logs an error and keeps running the old config.
curl -X POST http://localhost:9093/-/reload
  1. Watch the groups form. During the next incident, query the live alerts by your grouping labels:
amtool -o extended alert query service=checkout cluster=eu-1

When the grouped notification actually goes out

Three timers on each route decide the shape of what on-call receives. The defaults are 30 seconds for group_wait, 5 minutes for group_interval and 4 hours for repeat_interval. Child routes inherit them unless they set their own.

  • group_wait is how long Alertmanager waits after a new group appears before it sends the first notification. That gives the rest of the storm, and any inhibiting alerts, time to arrive. If an alert resolves before group_wait passes, no notification goes out for it at all, which quietly absorbs flapping. Set it too short and the first page is incomplete. Set it too long and real pages arrive late.
  • group_interval is a recurring timer that starts once group_wait has passed. At each tick, Alertmanager sends an update only if alerts in the group have fired or resolved since the last tick. A new instance joining the group arrives in the group’s next update, and does not start a separate group.
  • repeat_interval sends an unchanged group again as a reminder. Alertmanager checks it after each group_interval, so make it a multiple of group_interval. If it isn’t one, Alertmanager rounds it up to the next multiple.

The example config in the Alertmanager docs sets group_wait: 10s on a database route. A shorter value like that suits a route where you want the first page sooner.

One detail bites slow receivers. group_interval also sets the timeout for each send, so a notification that takes longer than group_interval is cancelled.

Muting downstream symptoms with inhibition

Grouping folds together alerts that share labels. Inhibition handles the cascade, where the symptoms carry different labels from the cause. An inhibition rule mutes target alerts while a source alert is firing, provided both have the same values for the labels listed in equal.

inhibit_rules:
  - name: cluster-down-mutes-services
    source_matchers:
      - alertname="ClusterUnreachable"
    target_matchers:
      - alertname!="ClusterUnreachable"
    equal: [cluster]
  - name: database-down-mutes-dependents
    source_matchers:
      - alertname="OrdersDBDown"
    target_matchers:
      - alertname!="OrdersDBDown"
    equal: [database]

The second rule is why the checkout rules above carry database: orders-db. With the dependency in a label, a plain inhibition rule can say “while the database is down, the checkout error alerts belong to the same incident”. The source alert still pages, so the incident has one headline alert that points at the cause.

Write target matchers that the source alert can’t match. Alertmanager has special handling for alerts that match both sides of a rule, and its docs advise choosing matchers that avoid that case, because it is hard to reason about.

Turning one alert group into one incident record

Alertmanager groups notifications, and your incident record needs the same identity so updates land on it. The webhook receiver sends a groupKey that identifies the group, plus a fingerprint for each alert inside it. Key the incident on groupKey and every update for that group touches the same record.

receivers:
  - name: payments-pager
    webhook_configs:
      - url: http://incident-bridge:8080/alertmanager
        send_resolved: true
# incident_bridge.py
from flask import Flask, request

app = Flask(__name__)
incidents = {}  # groupKey -> incident

@app.post("/alertmanager")
def receive():
    payload = request.get_json()
    incident = incidents.setdefault(payload["groupKey"], {
        "labels": payload["groupLabels"],
        "alerts": {},
        "status": "open",
    })
    for alert in payload["alerts"]:
        incident["alerts"][alert["fingerprint"]] = alert["status"]
    incident["status"] = "resolved" if payload["status"] == "resolved" else "open"
    return "", 204

If your incident tool accepts a deduplication key on the events it receives, send it the group key and you get the same behaviour without the bridge. Also decide what happens when a resolved group fires again. Reopening the old incident within a short window keeps the history in one place. Opening a new incident after that window keeps unrelated recurrences apart.

Grouping by cause when the labels differ

Labels group alerts that look alike. Some incidents are made of alerts that look nothing alike: a deploy that raises latency on one service, errors on another and queue depth on a third. Folding those together takes context that labels alone don’t carry.

  • Dependencies. Encode what each service depends on as labels, as with database above, or use a service map so an upstream alert can absorb the alerts from its dependents.
  • Changes. Record deploys and config changes with a timestamp and the service they touched. Alerts that start just after a change to the same service are strong candidates for one incident.
  • Time, with care. Alerts that start within a short window and share some context can belong together. Time alone is weak evidence, because two unrelated failures can start in the same minute. Prefer causal context over timestamps.

Keep the raw alerts searchable after you fold them. Grouping changes who gets paged and how often, and the individual alerts stay the evidence someone reads during the investigation and the review.

Where alert grouping goes wrong

  • group_by: ['...'] turns grouping off. That special value groups by every label, so each alert passes through on its own. Use it only when the system downstream does its own grouping.
  • The group is so wide it hides a second incident. A key of alertname alone can bundle a minor problem with a critical one. Fix: put service and cluster or environment in the key, and give critical and warning alerts separate routes.
  • An inhibition rule mutes everything. Alertmanager treats a missing label and an empty label as the same thing. So if every label in equal is missing from both the source and the target alerts, the rule still applies. Fix: use only equal labels that every relevant rule sets, and test the rule with real label sets.
  • Rules use different label names. svc, service and app form three groups for one service. Fix: one label schema, checked in review.
  • Prometheus sends to a load balancer. In high-availability mode, Alertmanager expects every alert to reach every instance. Point Prometheus at the full list of Alertmanagers, with no load balancer in between.
  • group_wait is set long to calm things down. The storm goes quiet, and so does the real page. Keep group_wait short, and cut noise with for, keep_firing_for and inhibition.
  • Silences outlive the work. A silence with no comment and a long expiry hides the next real failure. Fix: set comment_required: true in the amtool config, and silence with the narrowest matchers you can, for example amtool silence add service=checkout cluster=eu-1 --comment="orders-db failover".

A checklist before you ship the change

  • Every alerting rule sets service, team, severity and a cluster or environment label.
  • Rules use for, plus keep_firing_for where a metric hovers around its threshold.
  • Each route’s group_by matches a failure domain you wrote down, and leaves instance out.
  • The timers are set explicitly on the root route, and overridden only where a team needs different values.
  • Every known cascade has an inhibition rule whose equal labels exist on both sides.
  • amtool check-config and amtool config routes test --verify.receivers run in CI.
  • The incident record is keyed on the group key, and you have decided when a recurrence reopens it.

Folding repeat alerts into one issue with Polylane

Polylane records each of its own detections, and each alert forwarded from a connected tool, as an issue with a fingerprint. For its own detections the fingerprint comes from the resource and the check, and for forwarded alerts from the integration and the alert id. A firing that matches an open issue raises that issue’s occurrence count, and a finding that belongs to an open or recently resolved issue joins it. If the same problem comes back within 24 hours, the existing issue reopens, so one agent keeps working it in one thread.

See how Polylane's alert intelligence works.

Common questions.

What is the difference between alert deduplication and alert grouping?

Deduplication collapses repeat firings of the same alert, so an alert sent on every evaluation stays one alert. Grouping batches different alerts that share the labels in group_by, such as every instance of one service in one cluster, into one notification. Most setups need both, plus inhibition for cascades where the symptoms carry different labels from the cause.

Which labels should I group alerts by?

Start with service, cluster or region, and environment, and leave instance out. That gives one group per service having trouble in one place. Grouping on alertname alone is too wide, and alertname plus instance is too narrow to reduce anything.

What are the default group_wait, group_interval and repeat_interval values in Alertmanager?

group_wait defaults to 30 seconds, group_interval to 5 minutes and repeat_interval to 4 hours. Child routes inherit these from their parent unless they set their own. repeat_interval should be a multiple of group_interval, and Alertmanager rounds it up to the next multiple if it isn't one.

Why did an alert fire but never send a notification?

If the alert resolved before group_wait passed, Alertmanager sends nothing for it. Next, run amtool silence query to look for an active silence, and check whether an inhibition rule muted it. amtool config routes test with the alert's labels shows which receiver it would reach.

How do I stop dependent services paging when a database goes down?

Add an inhibit rule with the database alert as the source, the dependent alerts as targets, and equal set to a label both carry, such as database. Then add that label to the dependent services' alerting rules. The database alert still pages, so the incident has one alert that points at the cause.

How do I test alert grouping and routing without firing real alerts?

Run amtool check-config on the file, then amtool config routes test with a label set and --verify.receivers. The test returns error code 1 if the alert would reach a different receiver, so you can run it in CI on every change to the Alertmanager config.

How do I get one incident per alert group in my incident tool?

Send notifications through a webhook receiver. The payload includes a groupKey that identifies the group, and a fingerprint for each alert in it. Key the incident on groupKey. If your incident tool accepts a deduplication key, pass the group key as that key.

Does alert grouping still work with several Alertmanager instances?

Yes, as long as every alert reaches every instance. Point Prometheus at the full list of Alertmanagers in its alerting config, and don't put a load balancer between Prometheus and the Alertmanagers.

Sources

  1. Alertmanager concepts (Prometheus documentation)
  2. Alertmanager configuration (Prometheus documentation)
  3. Alerting rules (Prometheus documentation)
  4. prometheus/alertmanager README, including amtool
  5. Alert Grouping and Deduplication (ADHDecode)
  6. Reducing alert noise without hiding real production failures (NHI Mgmt Group)
  7. Polylane documentation

About the author

Boris Tane

Founder of Polylane

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

Related

Nobody should be on-call. Polylane watches your infra, finds what broke, and writes the fix.

Get started for free