Skip to content

Getting Started

This walks through one alert for an HTTP API, from declaration to a rule file Prometheus loads and an alertmanager.yml that routes the alert.

To start from a scaffold instead, chant init --lexicon prometheus <dir> writes the rule group and routing this page builds. --template picks one of three others:

TemplateWhat it writes
rulesThe rule group only, for a setup whose Alertmanager config lives elsewhere
slo-styleAn error ratio recorded over two windows and a burn-rate alert that needs both, written out as plain rules
sloAn Slo, which builds the error ratios, the error budget and the page and ticket burn-rate alerts, and the routing for both severities, with a page muting the same SLO’s ticket. See SLOs

Each one builds clean and passes the PROM checks.

Terminal window
npm install --save-dev @intentius/chant @intentius/chant-lexicon-prometheus
chant.config.ts
import type { ChantConfig } from "@intentius/chant";
export default { lexicons: ["prometheus"] } satisfies ChantConfig;

A RuleGroup takes the rule file’s own keys. Rules are plain objects: record and expr for a recording rule, alert and expr for an alerting rule.

src/rules.ts
import { RuleGroup, type Rule } from "@intentius/chant-lexicon-prometheus";
const rules: Rule[] = [
{
record: "job:http_errors:ratio5m",
expr: 'sum by (job) (rate(http_requests_total{code=~"5.."}[5m])) / sum by (job) (rate(http_requests_total[5m]))',
},
{
alert: "ApiErrorRatioHigh",
expr: "job:http_errors:ratio5m > 0.05",
for: "10m",
labels: { severity: "page" },
annotations: { summary: "{{ $labels.job }} is failing over 5% of requests" },
},
];
const api = new RuleGroup({ name: "api", interval: "30s", rules });
export { api };
src/alertmanager.ts
import { Receiver, Route, type RouteProps, type WebhookConfig } from "@intentius/chant-lexicon-prometheus";
const hook: WebhookConfig[] = [{ url: "http://alert-sink.monitoring:8080/" }];
const oncall = new Receiver({ name: "oncall", webhook_configs: hook });
const fallback = new Receiver({ name: "default" });
const children: RouteProps[] = [{ matchers: ['severity="page"'], receiver: oncall }];
const root = new Route({ receiver: fallback, group_by: ["alertname", "job"], routes: children });
export { oncall, fallback, root };

The root route is the Route no other route nests. Its receiver takes every alert no child route matches.

Terminal window
chant build src --lexicon prometheus -o dist/rules.yml

dist/rules.yml holds the group, and dist/alertmanager.yml is written beside it:

dist/rules.yml
groups:
- name: api
interval: 30s
rules:
- record: job:http_errors:ratio5m
expr: sum by (job) (rate(http_requests_total{code=~"5.."}[5m])) / sum by (job) (rate(http_requests_total[5m]))
- alert: ApiErrorRatioHigh
expr: job:http_errors:ratio5m > 0.05
for: 10m
labels:
severity: page
annotations:
summary: "{{ $labels.job }} is failing over 5% of requests"

Every expr was parsed as PromQL (PROM104), durations were checked (PROM103), the route’s receiver exists (PROM201), and the page severity has a route (PROM202). Change severity: "page" to severity: "critical" and rebuild: PROM202 reports that no route matches severity="critical", so the alert would go to the default receiver.

With promtool and amtool installed:

Terminal window
promtool check rules dist/rules.yml
amtool check-config dist/alertmanager.yml

The same checks run from TypeScript with promtoolCheckRules and amtoolCheckConfig.