Skip to content

Agent Observability on k3d

An agent that misreads a quantity still returns 200, so seeing its failures takes more than error logs. You need GenAI spans, a collector that samples without dropping the failed runs, and an SLO over those runs with dashboards on top. This tutorial walks the agent-observability example, which declares that whole stack in one chant project and runs it on k3d.

None. The cluster is k3d on your machine and every backend is an open-source image (the OpenTelemetry Collector, the Prometheus and Grafana projects’ own images), so nothing asks for an account or a key.

Each layer is declared with the lexicon that owns it, and each one reads the declarations it depends on instead of repeating them.

LayerLexiconDeclared with
Clusterk3dCluster, brought up by k3dUp
Collector configsotelcomponents, Pipeline, SpanMetricsConnector, TailSamplingProcessor, genAiComponents()
Collectors on Kubernetesk8sOtelCollector (agent DaemonSet), OtelCollectorGateway, gatewayExporter()
Rules and routingprometheusSlo, Route, Receiver, InhibitRule
DashboardsgrafanaDatasource, RedDashboard, SloDashboard, AgentDashboard

The joins are where chant earns its keep:

  • The agent’s exporters come from gatewayExporter(gateway, { loadBalance: true }), so renaming the gateway or moving its namespace moves the agents with it, and the loadbalancing exporter sends every span of a trace to one gateway replica.
  • The gateway runs tail sampling after the spanmetrics connector and the GenAI preset’s metrics branch, so metrics count every span and only traces are thinned.
  • The SLO’s SLI reads the spanmetrics metric names, and the SLO dashboard reads the SLO’s recorded series from sloMetrics().
  • The RED and agent dashboards read their metric names from the connector and the GenAI preset, so renaming a namespace at the source moves the panels.
  • Prometheus, Alertmanager and Grafana mount ConfigMaps holding exactly what the prometheus and grafana builds write, through ruleFileYaml, alertmanagerYaml and the grafana lexicon’s GrafanaConfigMaps and grafanaVolumes.
Terminal window
cd examples/agent-observability
npm run build
npm run lint

build writes one output per lexicon under dist/. The k8s manifests carry the prometheus and grafana outputs in ConfigMaps, so the cluster runs exactly the files those builds wrote.

Terminal window
npm run image
k3d cluster create --config dist/k3d.yaml
k3d image import agent-observability-demo:0.1.0 -c agent-observability
export KUBECONFIG=$(k3d kubeconfig write agent-observability)
kubectl apply -f dist/k8s.yaml
kubectl -n observability port-forward svc/grafana 3000:80

Open http://localhost:3000. The demo agent fails one run in five against a 99% objective, on purpose. Within a few minutes the SLO’s page alert fires and Alertmanager routes it to oncall, while the ticket alert for the same SLO is inhibited by the page. Tempo holds every failed and every slow run but only some of the ordinary ones, and Prometheus counts them all.

The example’s on-demand e2e does all of this for you and checks each step:

Terminal window
npx vitest run --project e2e examples/agent-observability

Each reference page covers its part in depth. Connectors, sampling and the GenAI preset are in the OTel lexicon docs, and the agent and gateway composites with their placement checks are in OTel collectors on Kubernetes. For Slo and routing see the Prometheus lexicon; for dashboards and provisioning, the Grafana lexicon.