Skip to content

Traces and metrics

llms.txtlists every page for an agent
Optional: hand this page to your coding agentThe steps work by hand too.
Show the whole prompt
Read https://intentius.io/terragucci/reference/observability/.
Add the OTLP variables and `dashboards: true` to terragucci.yml as the page says, run `npx terragucci init`, and open a pull request. List the secrets I must add; do not create them.
Never apply, approve (a pull request review or `terragucci approve`), override a policy denial (`terragucci override`), use `--mode apply`, or merge; never touch `.chant/allowed_signers` or `chant/lifecycle`.

Nothing is sent until an endpoint is set; any backend with an OTLP endpoint works.

terragucci reads the standard OpenTelemetry variables. Set the endpoint for every job in terragucci.yml:

env:
OTEL_EXPORTER_OTLP_ENDPOINT: https://otel-collector.example.com:4318
Variable Effect
OTEL_EXPORTER_OTLP_ENDPOINT the collector’s OTLP/HTTP base URL; traces go to /v1/traces and metrics to /v1/metrics
OTEL_EXPORTER_OTLP_TRACES_ENDPOINT, OTEL_EXPORTER_OTLP_METRICS_ENDPOINT one signal’s full URL, used as given
OTEL_EXPORTER_OTLP_HEADERS key=value pairs sent with every request, such as an API key
OTEL_TRACES_EXPORTER=none, OTEL_METRICS_EXPORTER=none turn one signal off
OTEL_SDK_DISABLED=true turn both off
OTEL_SERVICE_NAME, OTEL_RESOURCE_ATTRIBUTES the service name (default terragucci) and extra resource attributes
TRACEPARENT the stage’s trace joins this one, when your CI starts a trace of its own
Setting Value
protocol OTLP/JSON over HTTP, so the collector’s port 4318, not gRPC
keys a CI secret named by telemetry: { headers_secret: OTLP_HEADERS }; the generated workflows set OTEL_EXPORTER_OTLP_HEADERS from it on the plan, apply and drift jobs
a fork’s pull request gets no secrets
GitLab mark the variable masked
an unreachable collector never fails a stage
terragucci tf-plan project, commit, binary, change set, totals
wave 2 set digest, approval
root envs/prod/payments plan digest, changes by action, status
tofu init
tofu plan the binary's own spans, under this one
tofu show

Each binary run is a span, passed as TRACEPARENT. OpenTofu and choudoufu nest theirs in it; Terraform starts its own, tagged terragucci.trace_id, terragucci.span_id and terragucci.root.

One tf-plan run's trace in Grafana's Explore, read from Tempo: the terragucci tf-plan stage span, the root span for envs/dev/orders, the tofu init, plan and show spans under it, and OpenTofu's own provider and resource spans inside the planOne tf-plan run's trace in Grafana's Explore, read from Tempo: the terragucci tf-plan stage span, the root span for envs/dev/orders, the tofu init, plan and show spans under it, and OpenTofu's own provider and resource spans inside the plan

The binary exports to a loopback receiver, OTLP over HTTP only, which forwards batches to your collector.

With binary: choudoufu, the OpenTofu fork from the team behind terragucci (set it up), the report also times each provider call, sums timings by resource type on a large estate, and times every state lock wait.

Each plan, drift and wave report lists where the time went, even without an endpoint:

  • roots by wall time, slowest first, with each one’s plan time and, in a wave, its apply time;
  • the slowest resource instances in each root, with their refresh time;
  • the time of each provider call, with the resource it was for;
  • provider start-up;
  • waits for a state lock, from choudoufu on a backend that locks.

Past a span budget (CHOUDOUFU_TRACE_DETAIL, CHOUDOUFU_TRACE_SPAN_BUDGET) choudoufu sums timings by resource type or provider method, and the report says so.

choudoufu adds two more:

Timing Sent by Spans In report.json
provider calls choudoufu tfplugin5.Provider/<method> and tfplugin6.Provider/<method>, with rpc.method roots[].timings.provider_calls
summed timings choudoufu, once a walk passes its span budget (2000 by default) Aggregate: <span name> roots[].timings.aggregates
Binary What the report gets
Terraform no traces; the report says so
OpenTofu a span for each resource instance it plans or applies, so the report lists the slowest resources
choudoufu what OpenTofu gives, plus provider calls, summed timings past the span budget, and state lock waits
Terragrunt a unit’s time from Terragrunt’s run report; a TG_TF_PATH wrapper points each plan at the receiver

On a root whose backend locks (S3 with use_lockfile, say), choudoufu (set it up) records each wait as a State lock wait span with its attempts. It shows in the report and in terragucci_lock_wait_seconds. An estate with a record store takes no lock, so it shows no waits (locking with choudoufu).

Run State lock wait
a Terragrunt plan up to five minutes (-lock-timeout=5m)
a tf-apply wave’s plan and apply up to five minutes (-lock-timeout=5m)

TF_CLI_ARGS, TF_CLI_ARGS_plan or TF_CLI_ARGS_apply set a different wait. A wave writes its report to terragucci-report/ and the reports bucket when set.

The HTML report shows them under “Where the time went”, and report.json carries timings and roots[].timings; see Report JSON schema.

A stage pushes its metrics once, when it ends; point Prometheus exporter or remote write on your collector at them. A tf-apply wave sends terragucci_root_apply_seconds, terragucci_provider_init_seconds, terragucci_lock_wait_seconds and the terragucci_wave_* gauges. All are gauges with project and stage labels, plus those below.

Metric Labels Answers Series per project and stage
terragucci_stage_duration_seconds result how long each stage takes, and whether it failed one per stage and result
terragucci_roots_planned how many roots a change touches one
terragucci_root_plan_seconds root the slowest roots one per root
terragucci_root_apply_seconds root, wave the slowest root applies; sent by a tf-apply wave one per root and wave
terragucci_plan_changes action creates, updates, replaces and deletes, over time one per action
terragucci_plan_groups how far a change folds one
terragucci_binary_version binary, version the binary versions in use across repos one per binary version in use
terragucci_last_run_seconds when each stage last ran, in Unix seconds one
terragucci_roots_changed pull_request on a plan how many roots a pull request changes one per pull request, which grows with every pull request
terragucci_tips rule tips by rule one per rule
terragucci_module_pin root, module, version which version of each module each root pins one per root, module and version, which grows with every version a root has pinned
terragucci_provider_init_seconds root, provider provider start-up one per root and provider
terragucci_lock_wait_seconds root time spent waiting for a state lock one per root that waited
terragucci_resource_seconds root, address the run’s slowest resources up to the ten slowest resources of a run, so the addresses change from run to run
terragucci_drift_roots roots a drift run found drifted one
terragucci_drift_since_seconds, terragucci_drift_clear_seconds when the open drift was first found (its drift issue opened), and when a drift run last found none one each
terragucci_wave_roots wave roots in a tf-apply wave one per wave
terragucci_wave_waiting_since_seconds, terragucci_wave_settled_seconds wave when a wave started waiting for its approval, and when it last stopped waiting one per wave
terragucci_dora_deployments_per_week, terragucci_dora_change_failure_ratio, terragucci_dora_restore_seconds project only, * for the estate; sent by terragucci estate the delivery metrics one each per project
terragucci_dora_lead_time_seconds project only, and segment; sent by terragucci estate the median lead time and its split four per project

pull_request, address and version are unbounded, so a long-retention Prometheus grows with them. The Change review dashboard needs pull_request and the Runs dashboard needs address; dropping either empties those panels.

The stage span’s terragucci.result is success or failure for a plan or drift run, and one of applied, nothing, waiting, refused and failed for a wave.

Metrics carry only terragucci.project and service.version as resource attributes, and spans also carry vcs.ref.head.revision and cicd.pipeline.run.url.full.

report.json records the run’s trace id as run.trace_id, and the HTML report shows it. Set telemetry.trace_url to a link with {trace_id} in it and the report links the trace instead:

telemetry:
trace_url: "https://grafana.example.com/explore?left=%7B%22datasource%22:%22tempo%22,%22queries%22:%5B%7B%22query%22:%22{trace_id}%22%7D%5D%7D"

The stage span carries the run’s report.html address as terragucci.report.url when reports.url is set.

With dashboards: true, init and reconcile write Grafana dashboards and alert rules; init leaves files it did not write alone.

observability/terragucci/
grafana/dashboards/<uid>.json one per dashboard
grafana/provisioning/dashboards/terragucci.yaml the provider that loads them
grafana/provisioning/alerting/terragucci.yaml the SLO burn-rate alerts, as Grafana-managed rules
prometheus/terragucci.rules.yml the SLO recording rules and alerts, and the pipeline alerts

init writes nine dashboards, the six below and one for each of the three SLOs.

Dashboard Shows
Pipeline health runs per hour, errors and duration for each stage and project, and runs by result
Change review roots changed and groups for each pull request, and creates, updates, replaces and destroys over time
Rollouts and waves waves waiting for an approval and for how long, wave runs by result, and how many roots pin each module version
Drift drifted roots by project, how old the open drift is, and roots corrected
Estate roots per project, the binary and terragucci versions each runs, the delivery metrics terragucci estate sends, module pins and tips by rule
Runs the slowest roots, root applies and resources, provider start-up, lock waits, stage durations, and the trace of each run from Tempo
One per SLO the SLI against its objective, the error budget left and the burn rates
The Estate dashboard's first row for the example: 15 roots in the project, run with tofu by terragucci 0.4.4The Estate dashboard's first row for the example: 15 roots in the project, run with tofu by terragucci 0.4.4

The SLOs, over 28 days: plans finish within ten minutes 95% of the time, wave applies succeed 99%, drift is corrected within a day 90%.

The plan time SLO dashboard's tiles for the example: the SLI over 28 days at 100% against a 95% objective, all of the error budget left, and no burn-rate alert firingThe plan time SLO dashboard's tiles for the example: the SLI over 28 days at 100% against a 95% objective, all of the error budget left, and no burn-rate alert firing
Alert Fires when
TerragucciDriftOld a project’s open drift is older than drift_age
TerragucciWaveWaiting a wave has waited for its approval longer than wave_wait
TerragucciApplyFailed an apply failed in the last hour
TerragucciWaveRefused a wave was refused in the last hour because its plans moved
TerragucciDriftStopped a project’s drift run has not run for schedule
Key under dashboards Default Meaning
dir observability/terragucci where the files go
prometheus prometheus the uid of the Grafana datasource reading the Prometheus that holds the metrics
tempo tempo the uid of the Grafana datasource reading Tempo
folder terragucci the Grafana folder for the dashboards and the Grafana-managed rules
path /var/lib/grafana/dashboards/terragucci where Grafana finds the dashboard files
drift_age, wave_wait, schedule 1d, 4h, 2d the alert thresholds

Mount grafana/provisioning/ under /etc/grafana/provisioning/ and grafana/dashboards/ at path. Add prometheus/terragucci.rules.yml to rule_files. Page from the alerting file or the ErrorBudgetBurn alerts, not both.

The dashboards read span metrics from a spanmetrics connector in your collector:

connectors:
spanmetrics:
namespace: terragucci.spans
dimensions:
- name: terragucci.stage
- name: terragucci.project
- name: terragucci.result
resource_metrics_key_attributes: [service.name, terragucci.project]
histogram:
unit: s
explicit:
buckets: [5s, 15s, 30s, 60s, 120s, 300s, 600s, 1200s, 1800s, 3600s]

Only the Runs dashboard’s trace list needs Tempo.

With reports.url set, the Runs dashboard links each trace to traces/<trace id>.html and Estate links each project to its report index.

terragucci

These docs count page views and clicks with PostHog. They set no cookies, store nothing in your browser, and send nothing when your browser asks not to be tracked.