Skip to content

Results

Which infrastructure toolchain lets an agent answer questions about an AWS estate for the least money. Pick a tool to see what its answers cost.

Pass rate

Ranked by what 100 correct answers cost: the spend on one question, divided by the share the tool gets right, times a hundred. Being cheap at being wrong does not help, and a hundred is a number worth having rather than four decimal places of cents.

Each row lists every run in the arm's replicate set, and the figure is the middle one. A single run cannot carry this: at three attempts per question these arms move about three trials in 24 with nothing changed between them. Ranking on the newest run put one arm's best and another's worst against each other and called it an order.

Select a row to see what that tool spent, how hard it worked, and the environment its agent was given.

Reproduce any of this

Every number here comes from a run anyone can repeat. It deploys to a local emulator, so it costs nothing and touches no AWS account.

git clone https://github.com/INTENTIUS/chant-bench && cd chant-bench
just setup                 # fetches the benchmark, builds every arm
just run chant             # one arm, about ten minutes
just ingest ../aws-bench   # bring the result into this site

Full instructions · each arm's exact command and briefing are under Agent environment on its panel below.

Pass rate by question

Of 24 trials: 8 questions, 3 attempts each.

describe-ec-instances-cross-regi3/3

Describe my EC2 instances across the three regions.

Graded against 4 / 1 / 1 by region✓ ✓ ✓

Everyone else: No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 2/3, Alchemy 3/3.

ec-instances-without-default-vpc3/3

Which of my EC2 instances don't have a default VPC?

Graded against 5✓ ✓ ✓

Everyone else: No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 1/3, Alchemy 3/3.

find-ec-instances-in-public-subn3/3

Find my EC2 instances that are in a public subnet.

Graded against 5✓ ✓ ✓

Everyone else: No tool (AWS CLI) 2/3, Pulumi 0/3, Terraform 1/3, AWS CDK 0/3, Alchemy v2 (Effect) 2/3, Alchemy 3/3.

list-ec-instances-all-regions3/3

List my account's EC2 instance ids in all regions.

Graded against 6 instances across 3 regions✓ ✓ ✓

Everyone else: No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 3/3, Alchemy v2 (Effect) 3/3, Alchemy 3/3.

list-ec-instances-all-regions-13/3

Which EC2 instances are reachable via SSH from the internet?

Graded against 2 — one only through its launch template✓ ✓ ✓

Everyone else: No tool (AWS CLI) 0/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 3/3, Alchemy 1/3.

list-ec-instances-by-vpc-across3/3

Which EC2 instances are in which VPCs across all regions?

Graded against 6 instances across 4 VPCs✓ ✓ ✓

Everyone else: No tool (AWS CLI) 3/3, Pulumi 2/3, Terraform 3/3, AWS CDK 1/3, Alchemy v2 (Effect) 1/3, Alchemy 3/3.

list-ec-private-ips-all-regions3/3

List all of my EC2 and their private ip in a table.

Graded against 6 instances with private IPs✓ ✓ ✓

Everyone else: No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 3/3, Alchemy v2 (Effect) 3/3, Alchemy 3/3.

list-unused-security-groups-all1/3

Provide me a list of unused Security Groups by all regions.

Graded against 4 attached to nothing✗ ✓ ✗

Everyone else: No tool (AWS CLI) 1/3, Pulumi 0/3, Terraform 0/3, AWS CDK 0/3, Alchemy v2 (Effect) 0/3, Alchemy 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
chant$0.0332
No tool (AWS CLI), runner up$0.0504
field average$0.1000
per question asked
chant$0.0304
No tool (AWS CLI), runner up$0.0378
field average$0.0700
tokens in
chant121,549
No tool (AWS CLI), runner up123,695
field average298,955
tokens out
chant2,088
No tool (AWS CLI), runner up2,871
field average4,162

Work per answer

What the agent had to do to get there.

commands
chant2.83
No tool (AWS CLI), runner up4.21
field average9.53
turns
chant4.83
No tool (AWS CLI), runner up6.04
field average11.69
clock time
chant43s
No tool (AWS CLI), best37s
field average108s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads
chant0
Pulumi, runner up0
field average35.57

Agent environment

Identical for every arm except the briefing, which is the one thing the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
chant-g3
what the run cost
$0.7307 — 24 questions at $0.0304 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/chant
harness
bfa85f8-dirty
briefing
briefing-chant-snapshot.md · 9ce3707f885e

Repeat this run:./benchmarks/agent-env/run-arm.sh chant

The briefing this agent received, in full
# Answer estate questions with `chant search` — the recorded state is the source of truth

This AWS estate was deployed from the chant project mounted at
`/workspace/chant`, and the chant CLI is installed in it. A state snapshot was
recorded at deploy time: it holds every managed resource with its resolved
physical id, the resources the estate depends on but does not declare, and the
edges between them. chant folds that graph into typed answers.

**Query the recorded state rather than enumerating the account resource by
resource.** A raw `aws ec2` sweep returns per-resource facts with no
relationships; the snapshot already holds the topology, and `--explain` reports
the universe it matched against, so you know the denominator.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root. Three read commands, each answering a different
shape of question:

**`chant lifecycle show floci`** — the complete recorded inventory: every
managed resource with its logical name, type, physical id and status, plus the
resources the estate depends on. This is the census, so you know the
denominator before you filter.

**`chant search "<query>" --at latest --env floci [--explain] [--show a,b]`** —
filter and join over that inventory. The main tool for any question narrower
than "list everything".

**`chant graph --format ir --at latest --env floci`** — the whole graph as JSON
on stdout. `nodes` carry `id`, `kind`, `physicalId` and `attrs`; `edges` carry
`from`, `to` and `viaAttr` (the attribute the reference travels through). For a
question about how resources relate rather than about one resource's properties.

Warnings go to stderr, so stdout is already valid JSON — redirect with
`2>/dev/null`, not `2>&1`, or the warnings land in the JSON and break the parse.
Both `search` and `graph` take `--at latest` to read the recording.

The snapshot already includes resources of a kind this estate manages that exist
in the account without being declared or referenced — a default security group,
something left behind. They are in every `--at` answer, marked distinctly; there
is no flag to add.

Every answer states what backed it — `— observed from snapshot <commit> taken
<time> · bound N/M` — so you can see the estate has already been read, and how
completely, without re-reading it yourself.

Values match exactly or by substring — there is no wildcard, so `attr:x=*foo`
matches nothing. When a query returns no matches, the footer names the
attributes the queried kind carries, and for an attribute you did query it lists
the values actually present. A miss is worth reading rather than working around.

Query grammar (space-separated terms, all must match):

- `kind:<substr>` — resource kind, e.g. `kind:EC2::Instance`
- `attr:<name>=<val>` — an attribute equals/contains a value
- `tag:<key>=<val>` — a tag with that key and value
- `!<term>` — prefix any term to require its ABSENCE. `!<-kind:X` selects nodes
  nothing of kind X points at, which is how you ask what is unattached. An edge
  term needs a target: say what would have referenced it.
- `->attr:n=v` / `->kind:X` — this resource has an edge TO one matching the
  right side; `<-` reverses it. This performs the join across the relationship,
  so `kind:EC2::Instance ->attr:MapPublicIpOnLaunch=true` selects instances by a
  property of their subnet.

Terms compose:

    chant search "kind:EC2::Subnet !<-kind:EC2::Instance" --at latest --env floci
    chant search "kind:EC2::Instance" --at latest --env floci --show VpcId,PrivateIpAddress

Each result row is `<logicalId>  <kind>  <physicalId>  <shown attrs>`. `--show`
takes the resource's own property names as the account reports them.
`--explain` adds a footer with the universe count ("N of M Instances matched")
and, for each non-match, the term it failed.

## Derived attributes

Besides the attributes AWS returns directly, chant records two facts about every
resource — `region`, and `providerDefault: true` on the ones AWS created rather
than anyone declaring them (a default VPC and its subnets, a VPC's default
security group, a main route table, AWS-managed keys and policies). Both are
plain attributes: query them with `attr:`, show them with `--show`.

It also folds multi-hop topology onto each instance and exposes the result as an
attribute:

- `internetFacing` — whether the instance's subnet routes to an internet
  gateway, resolved through the route table, including a default VPC's main
  route-table association.
- `effectiveIngress` — ingress rules that reach the instance, resolved across
  both its directly attached security groups and any reached through its launch
  template. Values take the form `<proto>:<port>:<cidr>`.

## Path to estate facts, in order

1. `chant search "<query>" --at latest --env floci --explain` — the default, for
   every question. Add `->`/`<-` when the answer depends on a relationship.
   `chant lifecycle show floci` when a census answers more directly than a
   filter, and `chant graph --format ir --at latest --env floci` when you want
   the raw graph to work over.
2. The typed source under `/workspace/chant/*/src/` — for intent the grammar
   doesn't cover.
3. `aws ec2 …` — for runtime values the recorded state does not carry (instance
   states, allocated addresses).

Pass rate by question

Of 24 trials: 8 questions, 3 attempts each.

describe-ec-instances-cross-regi3/3

Describe my EC2 instances across the three regions.

Graded against 4 / 1 / 1 by region✓ ✓ ✓

Everyone else: chant 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 2/3, Alchemy 3/3.

ec-instances-without-default-vpc3/3

Which of my EC2 instances don't have a default VPC?

Graded against 5✓ ✓ ✓

Everyone else: chant 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 1/3, Alchemy 3/3.

find-ec-instances-in-public-subn2/3

Find my EC2 instances that are in a public subnet.

Graded against 5✗ ✓ ✓

Everyone else: chant 3/3, Pulumi 0/3, Terraform 1/3, AWS CDK 0/3, Alchemy v2 (Effect) 2/3, Alchemy 3/3.

list-ec-instances-all-regions3/3

List my account's EC2 instance ids in all regions.

Graded against 6 instances across 3 regions✓ ✓ ✓

Everyone else: chant 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 3/3, Alchemy v2 (Effect) 3/3, Alchemy 3/3.

list-ec-instances-all-regions-10/3

Which EC2 instances are reachable via SSH from the internet?

Graded against 2 — one only through its launch template✗ ✗ ✗

Everyone else: chant 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 3/3, Alchemy 1/3.

list-ec-instances-by-vpc-across3/3

Which EC2 instances are in which VPCs across all regions?

Graded against 6 instances across 4 VPCs✓ ✓ ✓

Everyone else: chant 3/3, Pulumi 2/3, Terraform 3/3, AWS CDK 1/3, Alchemy v2 (Effect) 1/3, Alchemy 3/3.

list-ec-private-ips-all-regions3/3

List all of my EC2 and their private ip in a table.

Graded against 6 instances with private IPs✓ ✓ ✓

Everyone else: chant 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 3/3, Alchemy v2 (Effect) 3/3, Alchemy 3/3.

list-unused-security-groups-all1/3

Provide me a list of unused Security Groups by all regions.

Graded against 4 attached to nothing✓ ✗ ✗

Everyone else: chant 1/3, Pulumi 0/3, Terraform 0/3, AWS CDK 0/3, Alchemy v2 (Effect) 0/3, Alchemy 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
No tool (AWS CLI)$0.0504
chant, best$0.0332
field average$0.1000
per question asked
No tool (AWS CLI)$0.0378
chant, best$0.0304
field average$0.0700
tokens in
No tool (AWS CLI)123,695
chant, best121,549
field average298,955
tokens out
No tool (AWS CLI)2,871
chant, best2,088
field average4,162

Work per answer

What the agent had to do to get there.

commands
No tool (AWS CLI)4.21
chant, best2.83
field average9.53
turns
No tool (AWS CLI)6.04
chant, best4.83
field average11.69
clock time
No tool (AWS CLI)37s
chant, runner up43s
field average108s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads by design
No tool (AWS CLI)83
Pulumi, best0
field average35.57

Agent environment

Identical for every arm except the briefing, which is the one thing the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
bare-g3
what the run cost
$0.9068 — 24 questions at $0.0378 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/bare
harness
bfa85f8-dirty
briefing
briefing-bare.md · 166c7534c252

Repeat this run:./benchmarks/agent-env/run-arm.sh bare

The briefing this agent received, in full
# Answer estate questions from the AWS API

There is no infrastructure toolchain here — no state file, no synthesized
template, no recorded snapshot. The AWS CLI is installed and configured against
the account, and that is the whole surface.

**Every answer has to be assembled from API calls.** `describe-instances`,
`describe-security-groups`, `describe-subnets`, `describe-route-tables` and
friends each return one slice; a question that spans resources means calling
several and joining the results yourself.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

The account spans **us-east-1**, **us-west-1** and **us-west-2**. Most EC2 calls
are regional, so a question about "all regions" means asking each one — pass
`--region` explicitly rather than relying on the default.

`--output json` piped through `jq` is usually easier to join than the table
output. `--query` filters server-side if you would rather narrow before it
reaches you.

Path to estate facts, in order:

1. `aws ec2 …`, `aws iam …` — the default, for every question. Join across calls
   when the answer spans resources.
2. `aws cloudformation describe-stack-resources` / `describe-stacks` — if the
   estate was deployed from a stack, this maps logical ids to physical ones.

Pass rate by question

Of 24 trials: 8 questions, 3 attempts each.

describe-ec-instances-cross-regi3/3

Describe my EC2 instances across the three regions.

Graded against 4 / 1 / 1 by region✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 2/3, Alchemy 3/3.

ec-instances-without-default-vpc3/3

Which of my EC2 instances don't have a default VPC?

Graded against 5✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 1/3, Alchemy 3/3.

find-ec-instances-in-public-subn0/3

Find my EC2 instances that are in a public subnet.

Graded against 5✗ ✗ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, Terraform 1/3, AWS CDK 0/3, Alchemy v2 (Effect) 2/3, Alchemy 3/3.

list-ec-instances-all-regions3/3

List my account's EC2 instance ids in all regions.

Graded against 6 instances across 3 regions✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Terraform 3/3, AWS CDK 3/3, Alchemy v2 (Effect) 3/3, Alchemy 3/3.

list-ec-instances-all-regions-13/3

Which EC2 instances are reachable via SSH from the internet?

Graded against 2 — one only through its launch template✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 0/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 3/3, Alchemy 1/3.

list-ec-instances-by-vpc-across2/3

Which EC2 instances are in which VPCs across all regions?

Graded against 6 instances across 4 VPCs✗ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Terraform 3/3, AWS CDK 1/3, Alchemy v2 (Effect) 1/3, Alchemy 3/3.

list-ec-private-ips-all-regions3/3

List all of my EC2 and their private ip in a table.

Graded against 6 instances with private IPs✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Terraform 3/3, AWS CDK 3/3, Alchemy v2 (Effect) 3/3, Alchemy 3/3.

list-unused-security-groups-all0/3

Provide me a list of unused Security Groups by all regions.

Graded against 4 attached to nothing✗ ✗ ✗

Everyone else: chant 1/3, No tool (AWS CLI) 1/3, Terraform 0/3, AWS CDK 0/3, Alchemy v2 (Effect) 0/3, Alchemy 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
Pulumi$0.0891
chant, best$0.0332
field average$0.1000
per question asked
Pulumi$0.0631
chant, best$0.0304
field average$0.0700
tokens in
Pulumi257,291
chant, best121,549
field average298,955
tokens out
Pulumi4,001
chant, best2,088
field average4,162

Work per answer

What the agent had to do to get there.

commands
Pulumi7.62
chant, best2.83
field average9.53
turns
Pulumi9.29
chant, best4.83
field average11.69
clock time
Pulumi47s
No tool (AWS CLI), best37s
field average108s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads
Pulumi0
Terraform, runner up0
field average35.57

Agent environment

Identical for every arm except the briefing, which is the one thing the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
pulumi-g3
what the run cost
$1.5147 — 24 questions at $0.0631 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/pulumi
harness
bfa85f8-dirty
briefing
briefing-pulumi.md · a06c6b73c0eb

Repeat this run:./benchmarks/agent-env/run-arm.sh pulumi

The briefing this agent received, in full
# Answer estate questions from the Pulumi state — it is the source of truth

This AWS estate was deployed from the Pulumi program mounted read-only at
`/workspace/pulumi`, already applied. The exported state records every resource
with its resolved live ids, its inputs and outputs, and the dependency edges
between resources.

**Query the state rather than enumerating the account resource by resource.** A
raw `aws ec2` sweep returns per-resource facts with no relationships; the state
export already holds the graph, and it is the complete set of managed resources,
so you know the denominator.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root:

- `cd /workspace/pulumi && ./pulumi-export` — the whole applied state as JSON.
  Each entry under `.deployment.resources[]` has:
  - `type` — the resource type, e.g. `aws:ec2/instance:Instance`
  - `urn` — its unique name
  - `inputs` — what was declared
  - `outputs` — the resolved attributes, including physical ids
  - `parent` and `dependencies` — the edges to other resources

  `jq` over `.deployment.resources[]` answers relationship questions without
  hand-joining CLI output — filter by `type`, then follow `dependencies` or an
  output id into the resources that reference it.

Path to estate facts, in order:

1. `./pulumi-export` piped through `jq` — the default, for every question. Use
   `dependencies`/`parent` and output ids when the answer spans resources.
2. The `index.ts` source under `/workspace/pulumi` — for intent and
   configuration the export doesn't surface directly.
3. `aws ec2 …` — for runtime values the state does not carry (instance states,
   allocated addresses).

Pass rate by question

Of 24 trials: 8 questions, 3 attempts each.

describe-ec-instances-cross-regi3/3

Describe my EC2 instances across the three regions.

Graded against 4 / 1 / 1 by region✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 2/3, Alchemy 3/3.

ec-instances-without-default-vpc3/3

Which of my EC2 instances don't have a default VPC?

Graded against 5✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 1/3, Alchemy 3/3.

find-ec-instances-in-public-subn1/3

Find my EC2 instances that are in a public subnet.

Graded against 5✗ ✗ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, Pulumi 0/3, AWS CDK 0/3, Alchemy v2 (Effect) 2/3, Alchemy 3/3.

list-ec-instances-all-regions3/3

List my account's EC2 instance ids in all regions.

Graded against 6 instances across 3 regions✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, AWS CDK 3/3, Alchemy v2 (Effect) 3/3, Alchemy 3/3.

list-ec-instances-all-regions-13/3

Which EC2 instances are reachable via SSH from the internet?

Graded against 2 — one only through its launch template✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 0/3, Pulumi 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 3/3, Alchemy 1/3.

list-ec-instances-by-vpc-across3/3

Which EC2 instances are in which VPCs across all regions?

Graded against 6 instances across 4 VPCs✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 2/3, AWS CDK 1/3, Alchemy v2 (Effect) 1/3, Alchemy 3/3.

list-ec-private-ips-all-regions3/3

List all of my EC2 and their private ip in a table.

Graded against 6 instances with private IPs✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, AWS CDK 3/3, Alchemy v2 (Effect) 3/3, Alchemy 3/3.

list-unused-security-groups-all0/3

Provide me a list of unused Security Groups by all regions.

Graded against 4 attached to nothing✗ ✗ ✗

Everyone else: chant 1/3, No tool (AWS CLI) 1/3, Pulumi 0/3, AWS CDK 0/3, Alchemy v2 (Effect) 0/3, Alchemy 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
Terraform$0.0890
chant, best$0.0332
field average$0.1000
per question asked
Terraform$0.0705
chant, best$0.0304
field average$0.0700
tokens in
Terraform310,938
chant, best121,549
field average298,955
tokens out
Terraform3,658
chant, best2,088
field average4,162

Work per answer

What the agent had to do to get there.

commands
Terraform8.88
chant, best2.83
field average9.53
turns
Terraform11.08
chant, best4.83
field average11.69
clock time
Terraform61s
No tool (AWS CLI), best37s
field average108s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads
Terraform0
Pulumi, runner up0
field average35.57

Agent environment

Identical for every arm except the briefing, which is the one thing the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
terraform-g3
what the run cost
$1.6929 — 24 questions at $0.0705 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/terraform
harness
bfa85f8-dirty
briefing
briefing-terraform.md · 7822d55ca7ca

Repeat this run:./benchmarks/agent-env/run-arm.sh terraform

The briefing this agent received, in full
# Answer estate questions from Terraform state — it is the source of truth

This AWS estate was deployed from the Terraform configuration mounted read-only
at `/workspace/terraform`, already applied, and the Terraform CLI is vendored in
the workspace. The applied state records every managed resource with its
resolved live ids, its attributes, and the references between resources.

**Query the state rather than enumerating the account resource by resource.** A
raw `aws ec2` sweep returns per-resource facts with no relationships; the state
already holds how resources reference one another, and `state list` gives you
the complete set under management, so you know the denominator.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root (use the vendored binary, `./terraform`):

- `cd /workspace/terraform && ./terraform state list` — every resource address
  under management, one per line. This is the full inventory.
- `cd /workspace/terraform && ./terraform state show <address>` — one resource
  with all of its resolved attributes.
- `cd /workspace/terraform && ./terraform show -json` — the whole applied state
  as JSON. Resources live under `.values.root_module` (recurse
  `child_modules`); each has `type`, `address`, and a `values` object with the
  resolved attributes. `jq` over this answers relationship questions without
  hand-joining CLI output.
- `cd /workspace/terraform && ./terraform output -json` — the declared outputs.

Path to estate facts, in order:

1. `./terraform show -json` or `state show` — the default, for every question.
   Follow attribute references (subnet ids, security-group ids, launch-template
   ids) between resources to answer questions that span them.
2. The `.tf` source under `/workspace/terraform` — for intent and configuration
   the state doesn't surface directly.
3. `aws ec2 …` — for runtime values the state does not carry (instance states,
   allocated addresses).

Pass rate by question

Of 24 trials: 8 questions, 3 attempts each.

describe-ec-instances-cross-regi2/3

Describe my EC2 instances across the three regions.

Graded against 4 / 1 / 1 by region✓ ✓ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, Alchemy v2 (Effect) 2/3, Alchemy 3/3.

ec-instances-without-default-vpc2/3

Which of my EC2 instances don't have a default VPC?

Graded against 5✗ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, Alchemy v2 (Effect) 1/3, Alchemy 3/3.

find-ec-instances-in-public-subn0/3

Find my EC2 instances that are in a public subnet.

Graded against 5✗ ✗ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, Pulumi 0/3, Terraform 1/3, Alchemy v2 (Effect) 2/3, Alchemy 3/3.

list-ec-instances-all-regions3/3

List my account's EC2 instance ids in all regions.

Graded against 6 instances across 3 regions✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, Alchemy v2 (Effect) 3/3, Alchemy 3/3.

list-ec-instances-all-regions-12/3

Which EC2 instances are reachable via SSH from the internet?

Graded against 2 — one only through its launch template✗ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 0/3, Pulumi 3/3, Terraform 3/3, Alchemy v2 (Effect) 3/3, Alchemy 1/3.

list-ec-instances-by-vpc-across1/3

Which EC2 instances are in which VPCs across all regions?

Graded against 6 instances across 4 VPCs✗ ✓ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 2/3, Terraform 3/3, Alchemy v2 (Effect) 1/3, Alchemy 3/3.

list-ec-private-ips-all-regions3/3

List all of my EC2 and their private ip in a table.

Graded against 6 instances with private IPs✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, Alchemy v2 (Effect) 3/3, Alchemy 3/3.

list-unused-security-groups-all0/3

Provide me a list of unused Security Groups by all regions.

Graded against 4 attached to nothing✗ ✗ ✗

Everyone else: chant 1/3, No tool (AWS CLI) 1/3, Pulumi 0/3, Terraform 0/3, Alchemy v2 (Effect) 0/3, Alchemy 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
AWS CDK$0.1599
chant, best$0.0332
field average$0.1000
per question asked
AWS CDK$0.0866
chant, best$0.0304
field average$0.0700
tokens in
AWS CDK352,189
chant, best121,549
field average298,955
tokens out
AWS CDK5,552
chant, best2,088
field average4,162

Work per answer

What the agent had to do to get there.

commands
AWS CDK11.62
chant, best2.83
field average9.53
turns
AWS CDK13.92
chant, best4.83
field average11.69
clock time
AWS CDK142s
No tool (AWS CLI), best37s
field average108s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads by design
AWS CDK109
Pulumi, best0
field average35.57

Agent environment

Identical for every arm except the briefing, which is the one thing the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
cdk-g3
what the run cost
$2.0780 — 24 questions at $0.0866 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/cdk
harness
bfa85f8-dirty
briefing
briefing-cdk.md · f4b4c7082924

Repeat this run:./benchmarks/agent-env/run-arm.sh cdk

The briefing this agent received, in full
# Answer estate questions from the CDK app and its stacks — they are the source of truth

This AWS estate was deployed from the AWS CDK application mounted read-only at
`/workspace/cdk_app`, and the CDK CLI is installed in it. CDK's deployed state
is CloudFormation: the synthesized templates hold the complete declared shape,
and the CloudFormation API maps each logical id to the physical id it deployed
to.

**Query the templates and the stacks rather than enumerating the account
resource by resource.** A raw `aws ec2` sweep returns per-resource facts with no
relationships; a synthesized template holds every resource, its properties, and
its `Ref`/`Fn::GetAtt` references to other resources — including the resources
L2 constructs generate that the source never names, so it is the complete
inventory and tells you the denominator.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root:

- `cd /workspace/cdk_app && npx cdk ls` — every stack the app defines.
- `cd /workspace/cdk_app && npx cdk synth <stack> --json` — the synthesized
  CloudFormation template: all resources with their properties, logical ids, and
  the `Ref`/`Fn::GetAtt` edges between them. `jq` over this answers relationship
  questions without hand-joining CLI output.

    `synth` prints **YAML** unless you pass `--json`, so piping it straight into
    `jq` fails with `Invalid numeric literal`. Warnings go to stderr, so redirect
    with `2>/dev/null`, not `2>&1`. The same templates are written as JSON to
    `cdk.out/*.template.json` if you would rather read them from there.
- `aws cloudformation describe-stack-resources --stack-name <stack> --region <region>`
  — the deployed logical id → physical id mapping for that stack.
- `aws cloudformation describe-stacks --stack-name <stack> --region <region>` —
  the stack's outputs and status.

Path to estate facts, in order:

1. `npx cdk synth --json` (or the templates in `cdk.out/`) for the declared shape and
   the relationships, joined to `describe-stack-resources` for the physical ids
   — the default, for every question. The app spans several stacks and regions;
   cover each.
2. `lib/`, `stacks/` and `environment.ts` under `/workspace/cdk_app` — for
   intent the template doesn't make obvious.
3. `aws ec2 …` — for runtime values the templates do not carry (instance states,
   allocated addresses).

Pass rate by question

Of 24 trials: 8 questions, 3 attempts each.

describe-ec-instances-cross-regi2/3

Describe my EC2 instances across the three regions.

Graded against 4 / 1 / 1 by region✓ ✓ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy 3/3.

ec-instances-without-default-vpc1/3

Which of my EC2 instances don't have a default VPC?

Graded against 5✓ ✗ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy 3/3.

find-ec-instances-in-public-subn2/3

Find my EC2 instances that are in a public subnet.

Graded against 5✗ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, Pulumi 0/3, Terraform 1/3, AWS CDK 0/3, Alchemy 3/3.

list-ec-instances-all-regions3/3

List my account's EC2 instance ids in all regions.

Graded against 6 instances across 3 regions✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 3/3, Alchemy 3/3.

list-ec-instances-all-regions-13/3

Which EC2 instances are reachable via SSH from the internet?

Graded against 2 — one only through its launch template✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 0/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy 1/3.

list-ec-instances-by-vpc-across1/3

Which EC2 instances are in which VPCs across all regions?

Graded against 6 instances across 4 VPCs✗ ✓ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 2/3, Terraform 3/3, AWS CDK 1/3, Alchemy 3/3.

list-ec-private-ips-all-regions3/3

List all of my EC2 and their private ip in a table.

Graded against 6 instances with private IPs✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 3/3, Alchemy 3/3.

list-unused-security-groups-all0/3

Provide me a list of unused Security Groups by all regions.

Graded against 4 attached to nothing✗ ✗ ✗

Everyone else: chant 1/3, No tool (AWS CLI) 1/3, Pulumi 0/3, Terraform 0/3, AWS CDK 0/3, Alchemy 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
Alchemy v2 (Effect)$0.1392
chant, best$0.0332
field average$0.1000
per question asked
Alchemy v2 (Effect)$0.0870
chant, best$0.0304
field average$0.0700
tokens in
Alchemy v2 (Effect)395,613
chant, best121,549
field average298,955
tokens out
Alchemy v2 (Effect)5,323
chant, best2,088
field average4,162

Work per answer

What the agent had to do to get there.

commands
Alchemy v2 (Effect)14.62
chant, best2.83
field average9.53
turns
Alchemy v2 (Effect)17.29
chant, best4.83
field average11.69
clock time
Alchemy v2 (Effect)317s
No tool (AWS CLI), best37s
field average108s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads
Alchemy v2 (Effect)6
Pulumi, best0
field average35.57

Agent environment

Identical for every arm except the briefing, which is the one thing the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
alchemy-effect-g2
what the run cost
$2.0886 — 24 questions at $0.0870 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/alchemy
harness
bfa85f8-dirty
briefing
briefing-alchemy-effect.md · 6e6e55fd2fc6

Repeat this run:./benchmarks/agent-env/run-arm.sh alchemy-effect

The briefing this agent received, in full
# Answer estate questions from the Alchemy state — it is the source of truth

This AWS estate was deployed from the Alchemy program mounted read-only at
`/workspace/alchemy`, already applied, and the Alchemy CLI is installed in it.
The applied state records every resource with its resolved live ids and
attributes. This estate is deployed as one stack per region, with an entrypoint
each: `us-east-1.run.ts`, `us-west-1.run.ts`, `us-west-2.run.ts`.

**Query the state rather than enumerating the account resource by resource.** A
raw `aws ec2` sweep returns per-resource facts with no relationships; the state
already holds each resource's resolved attributes and the ids it references, and
`state resources` is the complete set per stack, so you know the denominator.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root with `--local`, which reads the on-disk store under
`.alchemy/state`. That store holds all three regions, and one entrypoint reaches
every stack in it — `--stack` is what selects the region, not the entrypoint. Use
`us-west-1.run.ts` as the handle throughout:

- `alchemy state tree us-west-1.run.ts --local` — every stack and stage with the
  resources under it.
- `alchemy state stacks us-west-1.run.ts --local` and
  `alchemy state stages us-west-1.run.ts --local` — the stacks and stages present.
- `alchemy state resources --stack <stack> --stage <stage> us-west-1.run.ts --local`
  — the fully-qualified name of every resource there. This is the full
  inventory for that stack.
- `alchemy state get --stack <stack> --stage <stage> --fqn <fqn> us-west-1.run.ts --local`
  — one resource with its resolved attributes, including physical ids and the
  subnet and security-group ids it references.

`alchemy state stacks` lists all three region stacks whichever entrypoint you
name, so one command per question covers the estate. The same records are on disk
under `/workspace/alchemy/.alchemy/state/*/bench/*.json` — one stack directory
per region, one JSON file per resource, each with a `resourceType`, a `props`
object holding the declared configuration and an `attr` object holding the
resolved attributes — if you would rather `jq` or grep the files directly.

Path to estate facts, in order:

1. `alchemy state resources` / `alchemy state get` per entrypoint — the default,
   for every question. Follow referenced ids between records when the answer
   spans resources.
2. The `*.run.ts` stacks and `src/` under `/workspace/alchemy` — for intent the
   state doesn't surface directly.
3. `aws ec2 …` — Alchemy treats cloud state as authoritative, so use it for
   runtime values the state does not carry (instance states, allocated
   addresses).

Pass rate by question

Of 24 trials: 8 questions, 3 attempts each.

describe-ec-instances-cross-regi3/3

Describe my EC2 instances across the three regions.

Graded against 4 / 1 / 1 by region✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 2/3.

ec-instances-without-default-vpc3/3

Which of my EC2 instances don't have a default VPC?

Graded against 5✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 1/3.

find-ec-instances-in-public-subn3/3

Find my EC2 instances that are in a public subnet.

Graded against 5✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, Pulumi 0/3, Terraform 1/3, AWS CDK 0/3, Alchemy v2 (Effect) 2/3.

list-ec-instances-all-regions3/3

List my account's EC2 instance ids in all regions.

Graded against 6 instances across 3 regions✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 3/3, Alchemy v2 (Effect) 3/3.

list-ec-instances-all-regions-11/3

Which EC2 instances are reachable via SSH from the internet?

Graded against 2 — one only through its launch template✓ ✗ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 0/3, Pulumi 3/3, Terraform 3/3, AWS CDK 2/3, Alchemy v2 (Effect) 3/3.

list-ec-instances-by-vpc-across3/3

Which EC2 instances are in which VPCs across all regions?

Graded against 6 instances across 4 VPCs✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 2/3, Terraform 3/3, AWS CDK 1/3, Alchemy v2 (Effect) 1/3.

list-ec-private-ips-all-regions3/3

List all of my EC2 and their private ip in a table.

Graded against 6 instances with private IPs✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 3/3, Pulumi 3/3, Terraform 3/3, AWS CDK 3/3, Alchemy v2 (Effect) 3/3.

list-unused-security-groups-all0/3

Provide me a list of unused Security Groups by all regions.

Graded against 4 attached to nothing✗ ✗ ✗

Everyone else: chant 1/3, No tool (AWS CLI) 1/3, Pulumi 0/3, Terraform 0/3, AWS CDK 0/3, Alchemy v2 (Effect) 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
Alchemy$0.1383
chant, best$0.0332
field average$0.1000
per question asked
Alchemy$0.1095
chant, best$0.0304
field average$0.0700
tokens in
Alchemy531,410
chant, best121,549
field average298,955
tokens out
Alchemy5,642
chant, best2,088
field average4,162

Work per answer

What the agent had to do to get there.

commands
Alchemy16.92
chant, best2.83
field average9.53
turns
Alchemy19.38
chant, best4.83
field average11.69
clock time
Alchemy106s
No tool (AWS CLI), best37s
field average108s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads
Alchemy51
Pulumi, best0
field average35.57

Agent environment

Identical for every arm except the briefing, which is the one thing the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
alchemy-g3
what the run cost
$2.6281 — 24 questions at $0.1095 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/alchemy
harness
bfa85f8-dirty
briefing
briefing-alchemy.md · 596be04902b9

Repeat this run:./benchmarks/agent-env/run-arm.sh alchemy

The briefing this agent received, in full
# Answer estate questions from the Alchemy state — it is the source of truth

This AWS estate was deployed from the Alchemy program mounted read-only at
`/workspace/alchemy`, already applied, and the Alchemy CLI is installed in it.
The applied state records every resource with its resolved live ids and
attributes.

**Query the state rather than enumerating the account resource by resource.** A
raw `aws ec2` sweep returns per-resource facts with no relationships; the state
already holds each resource's resolved outputs and the ids it references, and
`state list` is the complete set of managed resources, so you know the
denominator.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root:

- `cd /workspace/alchemy && alchemy state tree` — every stack and stage with the
  resources under it.
- `cd /workspace/alchemy && alchemy state list` — the fully-qualified name of
  every resource, one per line. This is the full inventory.
- `cd /workspace/alchemy && alchemy state get <fqn>` — one resource as JSON:
  `kind` is the resource type (e.g. `aws::Instance`, `aws::SecurityGroupRule`)
  and `output` holds the resolved attributes — physical ids, IPs, and the subnet
  and security-group ids it references. Following those ids into other records
  answers questions that span resources.

Fully-qualified names look like `<app>/<stage>/<resource-id>`, so
`alchemy state list` then `alchemy state get` over the names walks the estate.
The same records are on disk under
`/workspace/alchemy/.alchemy/alchemy-ec2-multiregion/bench/*.json` if you would
rather `jq` or grep the files directly.

Path to estate facts, in order:

1. `alchemy state list` / `alchemy state get` — the default, for every question.
   Follow referenced ids between records when the answer spans resources.
2. `alchemy.run.ts` and `src/` under `/workspace/alchemy` — for intent the state
   doesn't surface directly.
3. `aws ec2 …` — Alchemy treats cloud state as authoritative, so use it for
   runtime values the state does not carry (instance states, allocated
   addresses).
What this is measuring

Every tool here can reach these answers. The agent keeps calling the API until it does. What differs is who does the work.

For most arms the model is the query engine. It sweeps, joins, and reasons over results, holding the estate in its context. That is what the token counts are buying.

chant moves the join into the tool. The model writes one query, the tool answers it. Same answer, a third of the tokens, and the answer comes back with the query that produced it. You can read it, re-run it, put it in CI.

That part is not an efficiency gain. It is the difference between an agent read the account and thinks four groups are unused and a line that can be checked.

Read every arm against No tool, which is upstream aws-bench's own experiment: an agent with the AWS CLI and nothing else. A tool that does not get there more cheaply is not earning its place.

Figures in these panels are per question, over that arm's latest valid run — one run, so they will not always agree with the ranking outside, which is the middle of the three. Where the two differ, the arm's runs disagree with each other, and the row above says by how much. Cost is the agent's own billed total, not tokens times a rate card. Bars are scaled against the highest value any arm recorded, so a short amber bar is the good one.

Reading account reads

A tool that answers from state it already holds is worth more than one that re-reads the cloud. CDK is the honest exception. It keeps no state of its own, so its reads are its sanctioned path, not a fallback.