Skip to content

Lint Rules

The prometheus lexicon’s rules use the PROM prefix. PROM0xx rules read your TypeScript source during chant lint. PROM101-PROM107 and PROM211-PROM219 run after a build over every rule file in the output, the other PROM2xx over every alertmanager.yml, and PROM3xx join the rule files with the rest of the build. Documents are recognized by shape, so a rule file or Alertmanager config that reaches the output another way is checked too.

IdSeverityCatches
PROM001errora literal credential in a Receiver or AlertmanagerSettings
PROM002errora literal expr in a RuleGroup that is not valid PromQL
PROM003erroran Slo whose literal objective is not strictly between 0 and 1, whose window is not a positive duration, or whose SLI expression lacks {{window}} or is not PromQL
PROM101errortwo groups in one rule file with the same name
PROM102warningtwo rules with the same name and label set
PROM103errora for, keep_firing_for, interval or query_offset that is not a Prometheus duration
PROM104erroran expr that is not valid PromQL
PROM105errora rule with both or neither of record and alert, an empty name, a recording rule with alert-only fields, or a group with no name
PROM106warningan alerting rule with no severity label on it or its group
PROM107warningan alerting rule with no summary or description annotation
PROM201errora route naming a receiver that is not declared
PROM202warningan alert severity no route below the root matches
PROM203errortwo receivers, or two time intervals, with the same name
PROM204errora route naming a time interval that is not declared
PROM205errorno root route, a root route without a receiver, or a root route with matchers
PROM206errora route or inhibit-rule matcher that does not parse
PROM207warninga receiver no route sends to
PROM208erroran Alertmanager duration that does not parse: group_wait, group_interval, repeat_interval, resolve_timeout and Jira reopen_duration as Prometheus durations, and the webhook, Slack, PagerDuty and incident.io timeout and Pushover retry, expire and ttl as Go durations
PROM209errora receiver integration missing a destination, credential or required field that Alertmanager asks for, after its global default
PROM210erroran integration or global setting Alertmanager rejects: a value and its *_file both set, settings that exclude each other, or a value outside the allowed set
PROM211warningan alerting rule with no for, or for: 0s
PROM212warningan alerting rule with no runbook_url annotation (opt-in)
PROM213warningan alert expression with no comparison, so it fires for every series it returns
PROM214warningan alert template reading $labels.x where the expression’s by or without drops x
PROM215warningrate, irate or increase over a name not ending in _total, _count, _sum or _bucket
PROM216warninghistogram_quantile over a series without _bucket in its name, or an aggregation that drops le
PROM217warninga recording rule not named level:metric:operations
PROM218warninga =~ or !~ matcher with no regex metacharacters, or with ^ or $ anchors
PROM219warningan alerting rule, or its group, setting the alertname label
PROM220warningtls_config.insecure_skip_verify: true on a receiver or in global
PROM221warningan email receiver sending SMTP credentials with require_tls: false
PROM222warningan integration sending a credential to an http:// URL
PROM223warninga route whose repeat_interval is shorter than its group_interval
PROM224warningan inhibit rule with no equal whose source and target matchers can match the same alert
PROM225errora labeldrop or labelkeep relabel step with source_labels, separator, target_label, modulus or replacement
PROM301warninga rule reading a spanmetrics, servicegraph or GenAI metric no collector config in the build emits, or grouping one by a label that is not a declared dimension

Alertmanager does not expand ${ENV} in its config, so the only way to keep a secret out of the file is its *_file field, pointing at a mounted secret. Replace routing_key: "..." with routing_key_file: "/etc/alertmanager/secrets/pagerduty-key". A Slack api_url of https://slack.com/api/chat.postMessage is not flagged. That is the bot endpoint, the bot token goes in http_config.authorization, and update_message needs that URL.

Both parse with @prometheus-io/lezer-promql, the grammar the Prometheus project publishes and its web UI uses, pinned as PROMQL_GRAMMAR. PROM002 reports literals in the editor; PROM104 reports every rule in the built file, including expressions assembled at runtime. The message gives the offset of the first error. The grammar checks syntax, not types: rate(up) parses. promtool check rules catches the rest, see Checking with promtool and amtool.

Slo() throws on these when the build runs it; PROM003 reports them in the editor, at the literal. Write the objective as a fraction (0.995, not 99.5), the window as a Prometheus duration (28d, not 4 weeks), and each SLI expression with {{window}} where the range goes, e.g. sum(rate(requests_total[{{window}}])). See SLOs.

Follows promtool’s duplicate-rules lint. The key is the rule’s kind, name and full label set (group labels merged under the rule’s), across the whole file. An alert at severity: "page" and the same alert at severity: "ticket" is not a duplicate.

For each severity value an alert in the build’s rule files carries, some route below the root must match it. Only matchers on severity are read, since other labels come from the series and aren’t known at build time; a route counts when it or a route above it matches on severity and every severity matcher on its path accepts the value. An unmatched severity falls through to the root’s default receiver, which is usually not where a page should go.

One build root only. PROM202 joins the rule files with the Alertmanager config in the same chant build output (chant #1939). When the rules and the Alertmanager config are built in different build roots, it has nothing to join and says nothing, which is indistinguishable from “all routed”. Declare them in the same build root, as the alerting example does.

Alertmanager requires exactly one root route with a receiver and no matchers. In chant, the root is the Route entity no other route nests; nest every other route under it.

Alertmanager reads its durations two ways. The route timers, global.resolve_timeout and Jira’s reopen_duration are Prometheus durations: 30s, 4h, 1d, no fractions. The timeout of webhook, Slack, PagerDuty and incident.io, and Pushover’s retry, expire and ttl, are Go durations (time.ParseDuration): 1.5s and 500ms are fine, 1d is not. Write 24h where you would write 1d.

PROM209 and PROM210: receiver integrations

Section titled “PROM209 and PROM210: receiver integrations”

Both follow the checks Alertmanager v0.34.1 makes when it loads the file, so amtool check-config rejects the same configs; the checks catch them at build time and name the receiver and entry. Alertmanager stops at the first error, these report every one. The sources are each integration’s UnmarshalYAML in config/notifiers.go and notify/<name>/config.go, and Config.UnmarshalYAML in config/config.go for the global fallbacks.

PROM209, something missing:

IntegrationNeedsFalls back to
discord_configs, msteams_configs, msteamsv2_configswebhook_url or webhook_url_file
email_configsto, smarthost, fromglobal.smtp_smarthost, global.smtp_from
incidentio_configsurl or url_file; when http_config is set, alert_source_token(_file) or http_config.authorization
jira_configsproject, issue_type, api_urlglobal.jira_api_url (no default)
mattermost_configswebhook_url or webhook_url_fileglobal.mattermost_webhook_url(_file)
opsgenie_configsapi_key or api_key_file; each responder an id, username or name, and a typeglobal.opsgenie_api_key(_file)
pagerduty_configsrouting_key(_file) or service_key(_file)
pushover_configsuser_key(_file) and token(_file)
rocketchat_configstoken_id(_file) and token(_file)global.rocketchat_token_id(_file), global.rocketchat_token(_file)
slack_configsapi_url(_file) or app_token(_file)global.slack_api_url(_file), or global.slack_app_token(_file) when the entry has no http_config.authorization
sns_configsone of topic_arn, phone_number, target_arn
telegram_configschat_id or chat_id_file, and bot_token(_file)global.telegram_bot_token(_file)
victorops_configsrouting_key, and api_key(_file)global.victorops_api_key(_file)
webex_configsroom_id, and http_config.authorization (or bearer_token(_file)) on the entry itself
webhook_configsurl or url_file
wechat_configsapi_secret(_file) and corp_idglobal.wechat_api_secret(_file), global.wechat_api_corp_id

Webex is checked before the global.http_config is applied, so a global authorization does not count. incident.io asks for a credential only when the entry has an http_config.

PROM210, a setting Alertmanager rejects:

  • A value and its *_file sibling both set, on any integration (webhook_url, url, api_key, routing_key, service_key, user_key, token, token_id, bot_token, chat_id, api_secret, alert_source_token, api_url, app_token) or in global (slack_api_url, slack_app_token, opsgenie_api_key, victorops_api_key, telegram_bot_token, smtp_auth_password, smtp_auth_secret, rocketchat_token, rocketchat_token_id, wechat_api_secret, mattermost_webhook_url).
  • Slack: api_url(_file) with app_token(_file); app_token(_file) with http_config.authorization; update_message without api_url: https://slack.com/api/chat.postMessage (the bot token then goes in http_config.authorization); a field without title and value; an action without type, text, and url or name; a confirmation without text. In global, slack_app_token(_file) with a slack_api_url other than the slack_app_url.
  • incident.io: alert_source_token(_file) with http_config.authorization.
  • Pushover: html and monospace together.
  • SNS: two of topic_arn, phone_number, target_arn. Alertmanager’s check is a chained XOR, so all three pass it; PROM210 does the same.
  • Email: a smarthost (or global.smtp_smarthost) that is not host:port; the same header twice in different case; with threading.enabled, a References or In-Reply-To header, or a thread_by_date other than none or daily.
  • Values outside the allowed set: WeChat message_type (text, markdown), Telegram parse_mode (Markdown, MarkdownV2, HTML), Jira api_type (auto, cloud, datacenter), an Opsgenie responder type (team, teams, user, escalation, schedule, or a template), a VictorOps custom_fields key the integration reserves (routing_key, message_type, state_message, entity_display_name, monitoring_tool, entity_id, entity_state).
  • Mattermost: a field, on the message or an attachment, without title and value.

The http_config block’s own rules (one of basic_auth, oauth2 and authorization, and so on) come from prometheus/common and are not checked here; amtool check-config catches them.

An alert with no for fires on the first evaluation that matches, so one bad scrape can page. Set for to how long the condition must hold. Three kinds of alert are left alone. An expression that reads no series, such as vector(1), is a deliberate always-firing alert. An expression that reads every series through a *_over_time function already spans a window, as GenAiRules’ budget alerts do. Two conditions joined by and are the multi-window burn-rate form Slo builds, where the short window does what for would.

PROM212 is off by default: the lexicon’s recommended lint preset, the one chant build uses, leaves it out. Turn it on in chant.config.ts with the all preset, or by naming it, which keeps it whatever the preset:

import type { ChantConfig } from "@intentius/chant";
export default {
lexicons: ["prometheus"],
lint: {
presets: { prometheus: "all" },
// or: rules: { PROM212: "warning" },
},
} satisfies ChantConfig;

Slo and GenAiRules take annotations on each alert for the runbook_url. validateRunbookUrls(file) returns the same findings for a parsed file.

PROM213 and PROM214: what an alert’s expression returns

Section titled “PROM213 and PROM214: what an alert’s expression returns”

Every series an alert’s expression returns is an alert. PROM213 reports an expression with nothing that filters it: no comparison (a comparison with bool returns every series too), no absent() and no unless. With or, both sides need a condition. An expression that reads no series is left alone, as for PROM211.

PROM214 follows the labels through the expression. sum by (job) keeps only job, sum without (instance) drops instance, histogram_quantile drops le, and one-to-one matching with on (...) keeps only the matching labels. A template that reads a dropped label renders it empty. A label the rule or its group sets is not reported. Prometheus fills $labels from the series alone (rules/alerting.go), so read such a label without $labels.

PROM215 and PROM216: counters and histograms by name

Section titled “PROM215 and PROM216: counters and histograms by name”

Both go by the metric’s name, since the type is only known to a running Prometheus. PROM215 reads rate, irate and increase taken straight from a selector. PROM216 assumes classic histograms. A native histogram has no _bucket series, so histogram_quantile over one is reported; leave PROM216 off with lint.rules: { PROM216: "off" } if you use native histograms.

Prometheus sets alertname to the rule’s name after it applies the rule’s labels (rules/alerting.go), so an alertname label is overwritten. promtool check rules accepts it.

amtool check-config accepts all three. PROM220 reports insecure_skip_verify: true in any tls_config of a receiver, or of global. PROM221 reports an email receiver with SMTP auth (auth_password, auth_secret or their *_file, its own or global’s) whose require_tls is false, its own or global.smtp_require_tls; force_implicit_tls is left alone. PROM222 reports an integration whose url, api_url or webhook_url, or the global URL it falls back to, is http:// while the request carries a credential: user info in the URL, http_config auth (its own, else global’s), or a key or token field.

PROM223 and PROM224: notification timing and inhibition

Section titled “PROM223 and PROM224: notification timing and inhibition”

Alertmanager resends a group’s notification only when it flushes the group, every group_interval, so a shorter repeat_interval never takes effect. PROM223 follows both timers down the routing tree from Alertmanager’s defaults (5m and 4h) and reports the route that sets one of them.

An inhibit rule with no equal mutes every target alert while any source alert fires. PROM224 reports one whose source and target matchers can both match the same alert, so alerts of that kind mute each other. Matchers that name values (=, or a regex of plain alternatives such as page|ticket) are compared; any other regex is taken to overlap.

labeldrop and labelkeep match regex against label names, so a step that also sets source_labels, separator, target_label, modulus or replacement has fields with nothing to act on, and Prometheus refuses to load the file. PROM225 reads every prometheus.yml in the build, in relabel_configs, metric_relabel_configs, alert_relabel_configs and write_relabel_configs, and names the step.

PROM301: metrics the build’s collectors emit

Section titled “PROM301: metrics the build’s collectors emit”

PROM301 joins the rule files with the collector configs in the same build. chant build gives each lexicon’s checks only its own output, but every entity, so the otel lexicon’s config is rebuilt from the build’s otel entities. A config found in the output, or in a Kubernetes ConfigMap, is read too. The names follow the otel lexicon’s spanMetricsNames(), serviceGraphNames() and the GenAI preset’s sum and signaltometrics connectors, with each config’s prometheus exporter namespace and add_metric_suffixes. A connector under a renamed type (span_metrics, service_graph, signal_to_metrics) counts as the built-in it names.

Each connector owns a namespace: a spanmetrics namespace, traces_service_graph, and the common prefix of a sum or signaltometrics connector’s names (genai.tokens, gen_ai.client). The spanmetrics default, traces_span_metrics, and traces_service_graph are owned whenever the build has a collector config, so a rule reading traces_span_metrics_calls_total from a build whose connector is named shop is reported. For a selector under an owned namespace, PROM301 reports a name no config emits. For an aggregation over such metrics, it reports a by (...) label none of them carries. The allowed labels are the declared dimensions, job, instance, the otel_scope_* labels, and le on _bucket series. An aggregation over label_replace or label_join is not checked. An exporter with resource_to_telemetry_conversion turns the label check off for its config, since any resource attribute can become a label.

PROM301 says nothing when the build has no collector config, and nothing about names outside the owned namespaces. The grafana lexicon’s GRAF118 runs the same check, collectorMetricIssues in collector-metrics.ts, over dashboard queries.

validateRuleFile(file), validateAlertmanagerConfig(config) and validateSeverityRouting(ruleFiles, config) return the same findings for any parsed file. The k8s lexicon’s tests use them on the groups a PrometheusRule renders.