Ops Reference
Reference for the *.op.ts file contract: the built-in step builders and retry profiles, what an Op’s labels are for, what bounds an ApplyOp delete on each target, and what chant build emits. For how to define, gate, and run an Op, see Ops.
Step builders
Section titled “Step builders”Pre-built step builders:
| Builder | What it does |
|---|---|
shell(cmd, opts?) | Run an arbitrary shell command; returns stdout, stderr, exitCode, and with json: true the parsed stdout as json |
build(path) | Run chant build |
kubectlApply(manifest, opts?) | Server-side apply a manifest as chant’s field manager |
helmInstall(name, chart, opts?) | helm upgrade --install |
waitForStack(stackFile, opts?) | Poll until a chant stack output file is ready |
gitlabPipeline(projectId, ref, opts?) | Trigger a GitLab pipeline and wait |
lifecycleSnapshot(env, opts?) | chant lifecycle snapshot |
teardown(path, opts?) | Build + destroy via the project’s own npm run teardown script |
envTeardown(env, opts?) | chant lifecycle teardown <env> --yes in-process: delete the environment’s marker-owned resources. A prod-like name needs confirmProd: true in opts; precede with a gate(...) step for a human approval |
effect(receipt, steps, opts?) | Read-compare-run-write over an effect receipt: read the live receipt, skip the wrapped steps when it matches the expectation (“effect already applied”), otherwise run them and write the receipt only on their success, last. receipt is the typed EffectReceipt declaration — there is no string form. See Effects |
Retry profiles
Section titled “Retry profiles”Each step takes an optional profile that controls the step’s retry and timeout policy. The table is ACTIVITY_PROFILES in core, so every runtime reads the same four:
| Profile | Suitable for |
|---|---|
fastIdempotent (default) | Quick, safe-to-retry steps |
longInfra | Slow infra changes (cluster create, Helm install) |
k8sWait | Polling until K8s resources are ready |
humanGate | Steps that may take hours |
shell() defaults to atMostOnce, twenty minutes and one attempt, because chant can’t know whether the command is safe to repeat.
A step’s own timeout sets how long one attempt may run, in place of its profile’s, and keeps the profile’s retries (#2787). shell("./build.sh", { timeout: "45m" }) is still one attempt, now of 45 minutes. activity(fn, args, { timeout }) and a step literal’s timeout field do the same. The value is a duration such as 45m or 1h30m, more than zero and at most six hours (MAX_STEP_TIMEOUT). A longer wait is waiting on a person, which a gate step is for. OPS012 flags a bad value at chant build, and the run fails the step if one gets that far. op.json carries the value on the step beside its profile.
Labels
Section titled “Labels”An Op’s labels are its discovery keys: free-form name: value pairs that say what kind of Op this is and where it belongs. ConvergeOp sets Converge: "true" alongside an Env. WatchOp sets Watch. chant operator --env <name> filters on labels.Env, and chant run list prints whatever an Op declares.
export default Op({ name: "alb-deploy", overview: "Deploy ALB stack", labels: { Environment: "staging", Region: "us-east-1", }, phases: [ phase("Build", [/* ... */]), phase("Deploy", [/* ... */]), ],});An Op’s labels describe the declaration. What a single run did (down to each step’s status) is a fact about that run and lands on the run ledger instead. See chant run.
Labels on the record
Section titled “Labels on the record”Every run appends one record to the run ledger on the chant/lifecycle branch, at <env>/runs__<name>.jsonl, and the Op’s labels ride along on it. <env> is the Op’s own labels.Env, or local when it declares none. That is what makes a label a filter after the fact: chant operator log --env staging and chant run log <name> both read the same records back.
OpName is set to the Op’s name by default, and a declared label wins over it: set labels: { OpName: "custom" } and custom is what lands.
Outcome attributes
Section titled “Outcome attributes”Activity steps support an optional outcomeAttribute field that captures the activity’s return value onto the run record’s outcomes:
phase("Diff", [ { kind: "activity", fn: "lifecycleDiff", args: { env: "prod", live: true }, // Capture lifecycleDiff's `drifted` field as Drift=true/false on the record outcomeAttribute: { name: "Drift", from: "drifted" }, },]),from is a dot-path into the return value; when omitted, the whole return value is stringified. chant run <name> --json prints the record the run wrote with its outcomes on it, which is the same document chant operator log reads back later.
Progress and gate state
Section titled “Progress and gate state”chant run status <name> reports the state the runtime holds for the latest run, and a picture needs per-phase detail: which step is in which phase, and which gate, if any, the run stopped at.
--progress-json
Section titled “--progress-json”chant run <name> --progress-json streams one NDJSON StepRecord per line to stdout as each declared step settles, fed by whatever runtime is hosting the run and matching chant run <name> --json’s record shape:
interface StepRecord { phase: string; fn: string; status: "ok" | "fail" | "skipped"; durationMs: number; error?: string;}A record is emitted once a step reaches a terminal state — "ok" or "fail". "skipped" only appears once the run itself settles: a declared step that never ran this time round (an onFailure phase on a run that succeeded, or an effect() step’s untaken branch) is reported skipped rather than left silently absent. Purely additive: omitting the flag leaves chant run’s output and exit code unchanged.
The same reconstruction backs the op-status MCP tool’s progress field and the chant://ops/<name>/runs/latest resource, so a point-in-time query and a live stream answer with the same shape.
Gate state
Section titled “Gate state”A run that stopped at an unapproved gate says so on the record it wrote, and the pending fact itself lives on the gate ledger at _gates/<op>.jsonl:
interface GateState { name: string; description?: string; since: string; // ISO timestamp expiresAt: string; // the gate's `timeout`, 48h by default}Read it off chant run status <name>’s Gate line, the op-status MCP tool’s gate field, or chant operator status’s pending-gates list. An Op with no gate steps reports nothing rather than an error.
The GateStep the run stopped at names itself on gate, which is the argument chant approve <op> <gate> takes; the same key names the gate on a component’s kind: "gate" step and inside ApplyOp’s gate config. That key was signalName through 0.58.0, and while a declaration still spelling it that way is read as before, chant build warns once per build and the old key is removed in 0.60.0 (#2202). The terraform composites spell it gateName, since gate there already selects the gate mode.
Delete paths
Section titled “Delete paths”delete controls how apply treats resources no longer declared. The delete rides the target’s own delete path, and what bounds that path differs per target:
| target | delete path | what bounds it |
|---|---|---|
kubectl | a sweep filtered by app.kubernetes.io/managed-by=chant (and by the stack when ownership.stack is set), restricted to the namespaces the apply touched | the ownership marker — an unmarked object is never touched |
arm | pruneArmOrphans, which lists the resource group and deletes only resources carrying chant’s ownership tag | the ownership tag — an untagged resource is never touched |
cloudformation | the stack deletes resources removed from its template | the stack — a resource CloudFormation did not create is not in it |
gcp | gcpApply’s prune, which lists each manifested kind and deletes only resources carrying chant’s ownership label | the ownership label — an unlabeled resource is never touched; a kind the lexicon cannot list is reported not-prunable rather than pruned |
grafana | grafanaApply’s prune, which deletes the project’s dashboards and folders that the build no longer declares, dashboards first | the ownership labels (managed-by, stack and env) — an unlabelled dashboard or folder is never touched; library panels, Grafana 11, and a project with no ownership.stack are reported not-prunable rather than pruned |
clickhouse | clickhouseApply’s prune, which drops the tables, views and databases in the declared databases that the build no longer declares, views first | the ownership marker on the object’s comment (managed-by, stack and env); an object without it is never touched, a database still holding one is kept, and a project with no ownership.stack is reported not-prunable rather than pruned |
postgres | postgresApply’s prune, which drops the objects in the declared schemas that the build no longer declares, views first and schemas last, never with CASCADE | the ownership marker on the object’s COMMENT (managed-by, stack and env); an object without it, or kept by another tool (an ORM’s revision table, a managed provider’s schema), is never touched, an object something else still depends on is kept, and a project with no ownership.stack is reported not-prunable rather than pruned |
fly | flyApply’s prune: machines by chant’s metadata marker; volumes/ips/certs/secrets app-scoped — anything the plan no longer declares under a managed app | the metadata marker for machines (an unmarked machine is never touched); the managed app for the metadata-less types, which carry no marker of their own |
All of them are owned-only by construction — bounded by a marker, tag, label, stack, or (for fly’s metadata-less types) the managed app itself.
An ARM delete needs an apiVersion, and the resource-group listing does not return one per resource. pruneArmOrphans takes it from the template when the template still declares the type, then from the lexicon’s pinned apiVersion registry (the same one the serializer stamps into the template), then from ARM’s provider metadata (GET /subscriptions/{sub}/providers/{namespace}?api-version=2021-04-01, newest stable version, one call per namespace per run). Removing the last resource of a type from source therefore still prunes it (#1472). Only a type none of the three sources know is reported as not-prunable with detail no-api-version.
Effects
Section titled “Effects”The effect(receipt, steps) builder wraps its steps in read-compare-run-write over an effect receipt. On each run the workflow reads the receipt through the receipt store (provided by the receipt row’s lexicon), compares it against the expectation, and skips the wrapped steps on a match. On a mismatch the steps run in authored order, and the receipt is written only when every one of them succeeded — the effect step is the sole writer of a receipt, and the write is always last. A failed step leaves the receipt untouched (stale), so the next run proposes the effect again. A gate(...) authored inside the wrapped steps is reached only when the effect will actually fire.
| Config | What it does |
|---|---|
effects: "gated" on ApplyOp | Insert the approval gate before the apply (beside delete: "gated", same gate shape) so the Plan phase’s effect-will-fire rows are reviewed before any effect runs |
receipts: [...] on WatchOp | Add a read-only Receipts phase reporting absent or differing receipts as findings (a StaleReceipts outcome on the run record) — nothing runs, nothing is written |
Codegen
Section titled “Codegen”chant build writes dist/ops/<name>/op.json for every declared Op: a deterministic, engine-neutral restatement of the step graph, produced by core itself rather than by a lexicon serializer.
| Section | What it holds |
|---|---|
| the step graph | every phase and step as declared, with each step’s profile and each gate’s timeout resolved to its effective value |
activityProfiles | the actual retry and timeout policy behind every profile the Op references, out of ACTIVITY_PROFILES |
activityContracts | args and returns as JSON Schema, for every step whose activity has a registered contract in the registry the build passed |
entities | the declared entity names a step’s contract identifies, so a reader can join a run against the estate |
Nothing has to import chant to read it, which is the point. A dashboard, an agent, or a runtime chant does not ship can answer “what does this Op do” from the file alone. Round-tripping op.json back through the IR and re-serializing produces a byte-identical file.
A hosting lexicon may write its own artifacts beside it under the same dist/ops/<name>/ directory; that location is fixed so the runtime knows where to look.
See Ops for how to define, gate, and run an Op.