Skip to content

GenAI Rules

GenAiRules builds one RuleGroup for workloads that send OpenTelemetry GenAI spans through the otel lexicon’s GenAI preset: request rate, error ratio and latency per provider, model and operation, token rates, spend from a price table you declare, and per-tool rates and errors. It reads every metric and label name from the GenAiMetrics the collector was built with, so a collector with another namespace, or with the conventions’ client metrics switched on, moves the expressions with it.

import { genAiMetrics } from "@intentius/chant-lexicon-otel";
import { GenAiRules } from "@intentius/chant-lexicon-prometheus";
// The same options the collector's genAiPipeline() was built with.
const metrics = genAiMetrics({ clientMetrics: "derive" });
export const genai = GenAiRules({
genAi: metrics,
prices: [
{ provider: "anthropic", model: "claude-x", inputPerMTok: 3, outputPerMTok: 15, currency: "USD", source: "https://example.com/pricing", asOf: "2026-09-29" },
],
alerts: { errorRatio: true, budgets: [{ amount: 50, currency: "USD", per: "day" }] },
});

GenAiRules is a composite, called without new. It returns { rules }, a RuleGroup named genai unless name says otherwise. Pass genai.rules to a k8s PrometheusRule to render the same rules as a CRD.

FieldTypeMeaning
genAiGenAiMetrics or { metrics }genAiMetrics(options) with the collector’s options, or genAiComponents(options).
source"client" or "spans"Which metrics the model rules read. Default client when the metrics include the conventions’ client metrics, otherwise spans.
pricesGenAiPrice[]Prices per provider and model. A model missing here gets no cost series.
alertsGenAiAlertingOpt-in alerts. None is built unless it is named here.
namestringThe rule group’s name. Default genai.
prefixstringThe first part of every recorded series name. Default gen_ai. Two GenAiRules in one Prometheus need different prefixes.
rateWindowdurationThe range every rate() reads. Default 5m.
groupBystring[]More Prometheus labels every rule keeps, such as job or service_name.
labelsRecord<string, string>Added to every rule as group labels.
intervaldurationThe group’s evaluation interval.

The model rules read one of two sets of metrics.

With source: "client", they read the conventions’ gen_ai.client.operation.duration and gen_ai.client.token.usage, which the collector derives from spans with clientMetrics: "derive" or passes through from the SDK with "passthrough". Requests are the duration histogram’s _count, errors are requests with an error.type, and token rates are the token histogram’s _sum with gen_ai.token.type input or output. These carry the provider, so every model series is split by provider, model and operation.

With source: "spans", they read the preset’s own genai.calls, genai.duration, genai.tokens.input and genai.tokens.output. Errors are spans with status.code STATUS_CODE_ERROR. These metrics carry the provider only when the collector sets providerDimensions: true, and the token sums never do, so the cost rules put the provider from the price table on their series. For the same reason a span-metric build refuses a price table that prices one model under two providers.

The tool rules read genai.calls and genai.duration whichever source the model rules use, because the conventions’ client metrics carry no tool name.

Spans without a request model (tool calls, in-process agent steps) are left out of the model series.

With the defaults (prefix gen_ai, rateWindow 5m):

SeriesLabelsValue
gen_ai:requests:rate5mprovider, model, operationrequests per second
gen_ai:errors:rate5mthe same and error_typefailed requests per second
gen_ai:error_ratio:rate5mprovider, model, operationfailed over all requests, 0 when none failed
gen_ai:error_ratio_by_type:rate5mthe same and error_typefailures of each type over all requests
gen_ai:operation_duration_seconds:quantile_rate5mthe same and quantile (0.5, 0.95, 0.99)operation latency in seconds
gen_ai:tokens:rate5mprovider, model and gen_ai_token_typetokens per second
gen_ai:cost:rate5mprovider, model, gen_ai_token_type and currencyspend per second in the price’s currency
gen_ai:tool_calls:rate5mgen_ai_tool_nametool calls per second
gen_ai:tool_errors:rate5mgen_ai_tool_namefailed tool calls per second
gen_ai:tool_error_ratio:rate5mgen_ai_tool_namefailed over all tool calls
gen_ai:tool_duration_seconds:quantile_rate5mgen_ai_tool_name, quantile="0.95"tool call latency in seconds

Provider is absent from the request, error and latency series of a span-metric build without providerDimensions. The labels are the attribute keys as the prometheus exporter writes them: gen_ai_provider_name, gen_ai_request_model, gen_ai_operation_name, error_type.

Recording rules come first in the group and the alerts after, so each evaluation’s alerts and ratios read the rates it has just recorded.

The semantic conventions define no cost metric, and the lexicon ships no prices, since providers change them without notice. Each entry in prices becomes one recording rule that multiplies the model’s recorded token rates by its prices:

- record: gen_ai:cost:rate5m
expr: |-
gen_ai:tokens:rate5m{gen_ai_provider_name="anthropic", gen_ai_request_model="claude-x", gen_ai_token_type="input"} * 3 / 1000000
or
gen_ai:tokens:rate5m{gen_ai_provider_name="anthropic", gen_ai_request_model="claude-x", gen_ai_token_type="output"} * 15 / 1000000
labels:
gen_ai_provider_name: anthropic
gen_ai_request_model: claude-x
currency: USD
Price fieldMeaning
providerThe gen_ai.provider.name value.
modelThe gen_ai.request.model value, exactly as spans report it. A provider that answers with a dated model name still matches on the name that was requested.
inputPerMTok, outputPerMTokPrice per million input and output tokens, at or above 0.
currencyRequired. chant converts no currency, and every cost series carries it as a label.
sourceRequired. Where the prices come from, such as the provider’s pricing page.
asOfThe date the prices were read, YYYY-MM-DD.

currency and source are the same two fields a workspace run’s cost record carries (#3033), so spend measured here and spend recorded for a run can be compared.

A model the table does not price gets no cost series at all, never a cost of zero, so a sum of gen_ai:cost:rate5m covers the priced models only. At the conventions’ pin, input tokens include cached input tokens, so a model with prompt caching is costed as if every input token were billed at the full input price. A cache-read price waits on the cache token metrics of #3044.

No alert is built unless alerts names it. Each takes true for its defaults, or an object that may set for (default 10m), severity (default warning), labels and annotations besides its threshold.

FieldAlertFires whenDefault threshold
errorRatioGenAiErrorRatioHighgen_ai:error_ratio:rate5m is over threshold0.05
latencyGenAiLatencyHighp95 operation latency is over thresholdSeconds30
toolErrorRatioGenAiToolErrorRatioHigha tool’s error ratio is over threshold0.1
budgetsGenAiSpendOverBudgetspend over the last hour or day, in one currency, is over amountnone; each budget sets amount, currency and per

Spend over a window is the average of gen_ai:cost:rate5m over the window, times its length in seconds:

- alert: GenAiSpendOverBudget
expr: sum by (currency) (avg_over_time(gen_ai:cost:rate5m{currency="USD"}[1d])) * 86400 > 50
labels:
budget: day
currency: USD
severity: warning

A budget fires as soon as it is over, without a for, unless it sets one. Its currency must be the currency of some price, and there can be one budget per window and currency.

GenAiRules() throws when genAi is not the otel lexicon’s GenAI metrics, when source is client and the metrics have no client metrics, when a metric lacks an attribute a rule is split by, when a price is missing currency or source or has a negative price, when a model is priced twice in a way the token series can’t tell apart, when a ratio threshold is not between 0 and 1, or when a duration, prefix or label name is not valid.

The lexicon’s tests build the rules from both genAiMetrics() and genAiMetrics({ clientMetrics: "derive" }) and evaluate them over fixed series with the evaluator in src/rule-eval.ts: request rates, error ratios, latency quantiles, token rates and cost per token type, a model without a price getting no cost series, and each alert firing above its threshold and not below it. With promtool on the path, promtool check rules checks the built file as well.

genAiRuleMetrics(rules) returns what the rules record, so a dashboard reads series names from the declaration instead of repeating them. It takes the GenAiRules(...) result, its rules group, or the props. The grafana lexicon’s AgentDashboard builds a dashboard from it with AgentDashboard({ rules, datasource }); build the rules with groupBy: ["service_name"] or ["job"] if the dashboard should have a service picker.

import { genAiRuleMetrics } from "@intentius/chant-lexicon-prometheus";
const m = genAiRuleMetrics(genai);
m.requests; // "gen_ai:requests:rate5m"
m.latency; // { record: "gen_ai:operation_duration_seconds:quantile_rate5m", quantiles: ["0.5", "0.95", "0.99"] }
m.cost; // "gen_ai:cost:rate5m"
m.labels.model; // "gen_ai_request_model"
m.currencies; // ["USD"]
FieldMeaning
group, source, rateWindowThe group name, the source the model rules read, and the rate window.
labelsLabel names: provider (absent when the source has none), model, operation, errorType, tokenType, tool (absent without a tool dimension), currency, quantile.
modelLabels, tokenLabelsThe labels the model series and the token and cost series are split by, groupBy last.
requests, errors, errorRatio, errorRatioByType, tokensSeries names.
latency{ record, quantiles }.
costThe cost series name. Absent without prices.
tool{ calls, errors, errorRatio, latency }. Absent when the span metrics carry no tool name.
pricesThe price table without the prices: provider, model, currency, source, asOf.
currenciesEvery currency in the price table.
alertsEach alert built: alert, kind, severity, threshold, and currency and per for a budget.

Agent rules (run rate, error ratio and run duration per agent), the average number of tool and model calls per agent run, and the cache hit ratio need the agent and cache metrics of #3044, and are left to it.