Agent Observability on k3d
An agent that misreads a quantity still returns 200, so seeing its failures takes more than error logs. You need GenAI spans, a collector that samples without dropping the failed runs, and an SLO over those runs with dashboards on top. This tutorial walks the agent-observability example, which declares that whole stack in one chant project and runs it on k3d.
None. The cluster is k3d on your machine and every backend is an open-source image (the OpenTelemetry Collector, the Prometheus and Grafana projects’ own images), so nothing asks for an account or a key.
How the lexicons fit together
Section titled “How the lexicons fit together”Each layer is declared with the lexicon that owns it, and each one reads the declarations it depends on instead of repeating them.
| Layer | Lexicon | Declared with |
|---|---|---|
| Cluster | k3d | Cluster, brought up by k3dUp |
| Collector configs | otel | components, Pipeline, SpanMetricsConnector, TailSamplingProcessor, genAiComponents() |
| Collectors on Kubernetes | k8s | OtelCollector (agent DaemonSet), OtelCollectorGateway, gatewayExporter() |
| Rules and routing | prometheus | Slo, Route, Receiver, InhibitRule |
| Dashboards | grafana | Datasource, RedDashboard, SloDashboard, AgentDashboard |
The joins are where chant earns its keep:
- The agent’s exporters come from
gatewayExporter(gateway, { loadBalance: true }), so renaming the gateway or moving its namespace moves the agents with it, and theloadbalancingexporter sends every span of a trace to one gateway replica. - The gateway runs tail sampling after the
spanmetricsconnector and the GenAI preset’s metrics branch, so metrics count every span and only traces are thinned. - The SLO’s SLI reads the
spanmetricsmetric names, and the SLO dashboard reads the SLO’s recorded series fromsloMetrics(). - The RED and agent dashboards read their metric names from the connector and the GenAI preset, so renaming a namespace at the source moves the panels.
- Prometheus, Alertmanager and Grafana mount ConfigMaps holding exactly what the prometheus and grafana builds write, through
ruleFileYaml,alertmanagerYamland the grafana lexicon’sGrafanaConfigMapsandgrafanaVolumes.
Build and lint
Section titled “Build and lint”cd examples/agent-observabilitynpm run buildnpm run lintbuild writes one output per lexicon under dist/. The k8s manifests carry the prometheus and grafana outputs in ConfigMaps, so the cluster runs exactly the files those builds wrote.
Run it
Section titled “Run it”npm run imagek3d cluster create --config dist/k3d.yamlk3d image import agent-observability-demo:0.1.0 -c agent-observabilityexport KUBECONFIG=$(k3d kubeconfig write agent-observability)kubectl apply -f dist/k8s.yamlkubectl -n observability port-forward svc/grafana 3000:80Open http://localhost:3000. The demo agent fails one run in five against a 99% objective, on purpose. Within a few minutes the SLO’s page alert fires and Alertmanager routes it to oncall, while the ticket alert for the same SLO is inhibited by the page. Tempo holds every failed and every slow run but only some of the ordinary ones, and Prometheus counts them all.
The example’s on-demand e2e does all of this for you and checks each step:
npx vitest run --project e2e examples/agent-observabilityWhere to go next
Section titled “Where to go next”Each reference page covers its part in depth. Connectors, sampling and the GenAI preset are in the OTel lexicon docs, and the agent and gateway composites with their placement checks are in OTel collectors on Kubernetes. For Slo and routing see the Prometheus lexicon; for dashboards and provisioning, the Grafana lexicon.