Skip to content

ec2-multiregion-negatives — results

A second scenario, scored separately from the board. Same estate, two extra questions, every arm.

Six trials, not twenty-four

These questions are ours, not aws-bench's, and they are scored over 2 questions at k=3 — six trials, against the board's 24. The figures are laid out the same way so they are readable the same way; they are not comparable with the board's, and nothing here belongs on it.

Six trials also move further than 24 do: one build of chant returned 3, 4 and 6 of 6 with nothing changed between the runs. Read the replicate set beside the figure, not after it.

Why these two

list-unused-security-groups-all-regions is the most interesting result on the board and the least representative. Every arm that keeps a state file is at zero on it, and an agent with no infrastructure tooling beats all of them. The answer is a negative about things a state file does not contain, and reading your own state cannot find what nothing points at.

That is one question out of eight, which is an anecdote. These two share the property that makes it hard: the account's default VPCs and their subnets were created by no deployment, so an arm reading only its own state sees a subset and cannot know what it is missing.

They are easier than the question they are modelled on. The no-tool baseline gets them with a sweep of two API calls, where the security-group question needs every network interface cross-referenced and even account-reading agents manage only 28%. What they test is the same structure, not the same difficulty.

Ground truth, computed live off the deployed estate at run time: 8 empty subnets of 13 across three regions, and 2 empty VPCs of 6.

Select a row to see what that tool spent, how hard it worked, and the environment its agent was given.

Reproduce any of this

These score an estate that is already deployed, so they run in about three minutes on top of an arm's board run.

just run chant                                    # deploy + board
./benchmarks/agent-env/run-negatives.sh chant     # then these two

Full instructions · each arm's exact command and briefing are under Agent environment on its panel below.

Pass rate by question

Of 6 trials: 2 questions, 3 attempts each.

subnets-with-no-network-interfac3/3

Which of my subnets have no network interfaces in them?

Graded against 8 of 13, across three regions✓ ✓ ✓

Everyone else: No tool (AWS CLI) 2/3, AWS CDK 2/3, Alchemy 2/3, Pulumi 0/3, Alchemy v2 (Effect) 0/3, Terraform 0/3.

vpcs-with-no-running-instances3/3

Which of my VPCs have no running instances?

Graded against 2 of 6✓ ✓ ✓

Everyone else: No tool (AWS CLI) 2/3, AWS CDK 3/3, Alchemy 2/3, Pulumi 0/3, Alchemy v2 (Effect) 1/3, Terraform 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
chant$0.0249
No tool (AWS CLI), runner up$0.0351
field average$0.1700
per question asked
chant$0.0249
No tool (AWS CLI), best$0.0234
field average$0.0800
tokens in
chant97,228
No tool (AWS CLI), best85,040
field average360,870
tokens out
chant1,707
No tool (AWS CLI), runner up1,780
field average4,585

Work per answer

What the agent had to do to get there.

commands
chant2
No tool (AWS CLI), runner up3.5
field average11.93
turns
chant4
No tool (AWS CLI), runner up5.17
field average14.1
clock time
chant26s
No tool (AWS CLI), best23s
field average81s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads
chant0
Pulumi, runner up0
field average15

Agent environment

Identical for every arm except the tool and its briefing, which are what the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
chant-neg-i3
tool under test
@intentius/chant 0.41.0
what the run cost
$0.1495 — 6 questions at $0.0249 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/chant
harness
e8c259c
briefing
briefing-chant-snapshot.md · 9ce3707f885e

Repeat this run:./benchmarks/agent-env/run-negatives.sh chant

The briefing this agent received, in full
# Answer estate questions with `chant search` — the recorded state is the source of truth

This AWS estate was deployed from the chant project mounted at
`/workspace/chant`, and the chant CLI is installed in it. A state snapshot was
recorded at deploy time: it holds every managed resource with its resolved
physical id, the resources the estate depends on but does not declare, and the
edges between them. chant folds that graph into typed answers.

**Query the recorded state rather than enumerating the account resource by
resource.** A raw `aws ec2` sweep returns per-resource facts with no
relationships; the snapshot already holds the topology, and `--explain` reports
the universe it matched against, so you know the denominator.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root. Three read commands, each answering a different
shape of question:

**`chant lifecycle show floci`** — the complete recorded inventory: every
managed resource with its logical name, type, physical id and status, plus the
resources the estate depends on. This is the census, so you know the
denominator before you filter.

**`chant search "<query>" --at latest --env floci [--explain] [--show a,b]`** —
filter and join over that inventory. The main tool for any question narrower
than "list everything".

**`chant graph --format ir --at latest --env floci`** — the whole graph as JSON
on stdout. `nodes` carry `id`, `kind`, `physicalId` and `attrs`; `edges` carry
`from`, `to` and `viaAttr` (the attribute the reference travels through). For a
question about how resources relate rather than about one resource's properties.

Warnings go to stderr, so stdout is already valid JSON — redirect with
`2>/dev/null`, not `2>&1`, or the warnings land in the JSON and break the parse.
Both `search` and `graph` take `--at latest` to read the recording.

The snapshot already includes resources of a kind this estate manages that exist
in the account without being declared or referenced — a default security group,
something left behind. They are in every `--at` answer, marked distinctly; there
is no flag to add.

Every answer states what backed it — `— observed from snapshot <commit> taken
<time> · bound N/M` — so you can see the estate has already been read, and how
completely, without re-reading it yourself.

Values match exactly or by substring — there is no wildcard, so `attr:x=*foo`
matches nothing. When a query returns no matches, the footer names the
attributes the queried kind carries, and for an attribute you did query it lists
the values actually present. A miss is worth reading rather than working around.

Query grammar (space-separated terms, all must match):

- `kind:<substr>` — resource kind, e.g. `kind:EC2::Instance`
- `attr:<name>=<val>` — an attribute equals/contains a value
- `tag:<key>=<val>` — a tag with that key and value
- `!<term>` — prefix any term to require its ABSENCE. `!<-kind:X` selects nodes
  nothing of kind X points at, which is how you ask what is unattached. An edge
  term needs a target: say what would have referenced it.
- `->attr:n=v` / `->kind:X` — this resource has an edge TO one matching the
  right side; `<-` reverses it. This performs the join across the relationship,
  so `kind:EC2::Instance ->attr:MapPublicIpOnLaunch=true` selects instances by a
  property of their subnet.

Terms compose:

    chant search "kind:EC2::Subnet !<-kind:EC2::Instance" --at latest --env floci
    chant search "kind:EC2::Instance" --at latest --env floci --show VpcId,PrivateIpAddress

Each result row is `<logicalId>  <kind>  <physicalId>  <shown attrs>`. `--show`
takes the resource's own property names as the account reports them.
`--explain` adds a footer with the universe count ("N of M Instances matched")
and, for each non-match, the term it failed.

## Derived attributes

Besides the attributes AWS returns directly, chant records two facts about every
resource — `region`, and `providerDefault: true` on the ones AWS created rather
than anyone declaring them (a default VPC and its subnets, a VPC's default
security group, a main route table, AWS-managed keys and policies). Both are
plain attributes: query them with `attr:`, show them with `--show`.

It also folds multi-hop topology onto each instance and exposes the result as an
attribute:

- `internetFacing` — whether the instance's subnet routes to an internet
  gateway, resolved through the route table, including a default VPC's main
  route-table association.
- `effectiveIngress` — ingress rules that reach the instance, resolved across
  both its directly attached security groups and any reached through its launch
  template. Values take the form `<proto>:<port>:<cidr>`.

## Path to estate facts, in order

1. `chant search "<query>" --at latest --env floci --explain` — the default, for
   every question. Add `->`/`<-` when the answer depends on a relationship.
   `chant lifecycle show floci` when a census answers more directly than a
   filter, and `chant graph --format ir --at latest --env floci` when you want
   the raw graph to work over.
2. The typed source under `/workspace/chant/*/src/` — for intent the grammar
   doesn't cover.
3. `aws ec2 …` — for runtime values the recorded state does not carry (instance
   states, allocated addresses).

Pass rate by question

Of 6 trials: 2 questions, 3 attempts each.

subnets-with-no-network-interfac2/3

Which of my subnets have no network interfaces in them?

Graded against 8 of 13, across three regions✗ ✓ ✓

Everyone else: chant 3/3, AWS CDK 2/3, Alchemy 2/3, Pulumi 0/3, Alchemy v2 (Effect) 0/3, Terraform 0/3.

vpcs-with-no-running-instances2/3

Which of my VPCs have no running instances?

Graded against 2 of 6✓ ✗ ✓

Everyone else: chant 3/3, AWS CDK 3/3, Alchemy 2/3, Pulumi 0/3, Alchemy v2 (Effect) 1/3, Terraform 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
No tool (AWS CLI)$0.0351
chant, best$0.0249
field average$0.1700
per question asked
No tool (AWS CLI)$0.0234
chant, runner up$0.0249
field average$0.0800
tokens in
No tool (AWS CLI)85,040
chant, runner up97,228
field average360,870
tokens out
No tool (AWS CLI)1,780
chant, best1,707
field average4,585

Work per answer

What the agent had to do to get there.

commands
No tool (AWS CLI)3.5
chant, best2
field average11.93
turns
No tool (AWS CLI)5.17
chant, best4
field average14.1
clock time
No tool (AWS CLI)23s
chant, runner up26s
field average81s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads by design
No tool (AWS CLI)17
Pulumi, best0
field average15

Agent environment

Identical for every arm except the tool and its briefing, which are what the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
bare-neg-i3
tool under test
aws-cli 2.36.14
what the run cost
$0.1405 — 6 questions at $0.0234 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/bare
harness
e8c259c
briefing
briefing-bare.md · 166c7534c252

Repeat this run:./benchmarks/agent-env/run-negatives.sh bare

The briefing this agent received, in full
# Answer estate questions from the AWS API

There is no infrastructure toolchain here — no state file, no synthesized
template, no recorded snapshot. The AWS CLI is installed and configured against
the account, and that is the whole surface.

**Every answer has to be assembled from API calls.** `describe-instances`,
`describe-security-groups`, `describe-subnets`, `describe-route-tables` and
friends each return one slice; a question that spans resources means calling
several and joining the results yourself.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

The account spans **us-east-1**, **us-west-1** and **us-west-2**. Most EC2 calls
are regional, so a question about "all regions" means asking each one — pass
`--region` explicitly rather than relying on the default.

`--output json` piped through `jq` is usually easier to join than the table
output. `--query` filters server-side if you would rather narrow before it
reaches you.

Path to estate facts, in order:

1. `aws ec2 …`, `aws iam …` — the default, for every question. Join across calls
   when the answer spans resources.
2. `aws cloudformation describe-stack-resources` / `describe-stacks` — if the
   estate was deployed from a stack, this maps logical ids to physical ones.

Pass rate by question

Of 6 trials: 2 questions, 3 attempts each.

subnets-with-no-network-interfac2/3

Which of my subnets have no network interfaces in them?

Graded against 8 of 13, across three regions✗ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, Alchemy 2/3, Pulumi 0/3, Alchemy v2 (Effect) 0/3, Terraform 0/3.

vpcs-with-no-running-instances3/3

Which of my VPCs have no running instances?

Graded against 2 of 6✓ ✓ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, Alchemy 2/3, Pulumi 0/3, Alchemy v2 (Effect) 1/3, Terraform 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
AWS CDK$0.1092
chant, best$0.0249
field average$0.1700
per question asked
AWS CDK$0.0910
No tool (AWS CLI), best$0.0234
field average$0.0800
tokens in
AWS CDK418,551
No tool (AWS CLI), best85,040
field average360,870
tokens out
AWS CDK6,015
chant, best1,707
field average4,585

Work per answer

What the agent had to do to get there.

commands
AWS CDK17.67
chant, best2
field average11.93
turns
AWS CDK19.33
chant, best4
field average14.1
clock time
AWS CDK140s
No tool (AWS CLI), best23s
field average81s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads by design
AWS CDK64
Pulumi, best0
field average15

Agent environment

Identical for every arm except the tool and its briefing, which are what the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
cdk-neg-i1
tool under test
aws-cdk 2.1131.0
what the run cost
$0.5458 — 6 questions at $0.0910 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/cdk
harness
e8c259c
briefing
briefing-cdk.md · f4b4c7082924

Repeat this run:./benchmarks/agent-env/run-negatives.sh cdk

The briefing this agent received, in full
# Answer estate questions from the CDK app and its stacks — they are the source of truth

This AWS estate was deployed from the AWS CDK application mounted read-only at
`/workspace/cdk_app`, and the CDK CLI is installed in it. CDK's deployed state
is CloudFormation: the synthesized templates hold the complete declared shape,
and the CloudFormation API maps each logical id to the physical id it deployed
to.

**Query the templates and the stacks rather than enumerating the account
resource by resource.** A raw `aws ec2` sweep returns per-resource facts with no
relationships; a synthesized template holds every resource, its properties, and
its `Ref`/`Fn::GetAtt` references to other resources — including the resources
L2 constructs generate that the source never names, so it is the complete
inventory and tells you the denominator.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root:

- `cd /workspace/cdk_app && npx cdk ls` — every stack the app defines.
- `cd /workspace/cdk_app && npx cdk synth <stack> --json` — the synthesized
  CloudFormation template: all resources with their properties, logical ids, and
  the `Ref`/`Fn::GetAtt` edges between them. `jq` over this answers relationship
  questions without hand-joining CLI output.

    `synth` prints **YAML** unless you pass `--json`, so piping it straight into
    `jq` fails with `Invalid numeric literal`. Warnings go to stderr, so redirect
    with `2>/dev/null`, not `2>&1`. The same templates are written as JSON to
    `cdk.out/*.template.json` if you would rather read them from there.
- `aws cloudformation describe-stack-resources --stack-name <stack> --region <region>`
  — the deployed logical id → physical id mapping for that stack.
- `aws cloudformation describe-stacks --stack-name <stack> --region <region>` —
  the stack's outputs and status.

Path to estate facts, in order:

1. `npx cdk synth --json` (or the templates in `cdk.out/`) for the declared shape and
   the relationships, joined to `describe-stack-resources` for the physical ids
   — the default, for every question. The app spans several stacks and regions;
   cover each.
2. `lib/`, `stacks/` and `environment.ts` under `/workspace/cdk_app` — for
   intent the template doesn't make obvious.
3. `aws ec2 …` — for runtime values the templates do not carry (instance states,
   allocated addresses).

Pass rate by question

Of 6 trials: 2 questions, 3 attempts each.

subnets-with-no-network-interfac2/3

Which of my subnets have no network interfaces in them?

Graded against 8 of 13, across three regions✓ ✗ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, AWS CDK 2/3, Pulumi 0/3, Alchemy v2 (Effect) 0/3, Terraform 0/3.

vpcs-with-no-running-instances2/3

Which of my VPCs have no running instances?

Graded against 2 of 6✓ ✗ ✓

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, AWS CDK 3/3, Pulumi 0/3, Alchemy v2 (Effect) 1/3, Terraform 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
Alchemy$0.1939
chant, best$0.0249
field average$0.1700
per question asked
Alchemy$0.1293
No tool (AWS CLI), best$0.0234
field average$0.0800
tokens in
Alchemy671,242
No tool (AWS CLI), best85,040
field average360,870
tokens out
Alchemy6,690
chant, best1,707
field average4,585

Work per answer

What the agent had to do to get there.

commands
Alchemy22.67
chant, best2
field average11.93
turns
Alchemy24.83
chant, best4
field average14.1
clock time
Alchemy141s
No tool (AWS CLI), best23s
field average81s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads
Alchemy22
Pulumi, best0
field average15

Agent environment

Identical for every arm except the tool and its briefing, which are what the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
alchemy-neg-i3
tool under test
alchemy 0.93.12
what the run cost
$0.7761 — 6 questions at $0.1293 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/alchemy
harness
e8c259c
briefing
briefing-alchemy.md · 596be04902b9

Repeat this run:./benchmarks/agent-env/run-negatives.sh alchemy

The briefing this agent received, in full
# Answer estate questions from the Alchemy state — it is the source of truth

This AWS estate was deployed from the Alchemy program mounted read-only at
`/workspace/alchemy`, already applied, and the Alchemy CLI is installed in it.
The applied state records every resource with its resolved live ids and
attributes.

**Query the state rather than enumerating the account resource by resource.** A
raw `aws ec2` sweep returns per-resource facts with no relationships; the state
already holds each resource's resolved outputs and the ids it references, and
`state list` is the complete set of managed resources, so you know the
denominator.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root:

- `cd /workspace/alchemy && alchemy state tree` — every stack and stage with the
  resources under it.
- `cd /workspace/alchemy && alchemy state list` — the fully-qualified name of
  every resource, one per line. This is the full inventory.
- `cd /workspace/alchemy && alchemy state get <fqn>` — one resource as JSON:
  `kind` is the resource type (e.g. `aws::Instance`, `aws::SecurityGroupRule`)
  and `output` holds the resolved attributes — physical ids, IPs, and the subnet
  and security-group ids it references. Following those ids into other records
  answers questions that span resources.

Fully-qualified names look like `<app>/<stage>/<resource-id>`, so
`alchemy state list` then `alchemy state get` over the names walks the estate.
The same records are on disk under
`/workspace/alchemy/.alchemy/alchemy-ec2-multiregion/bench/*.json` if you would
rather `jq` or grep the files directly.

Path to estate facts, in order:

1. `alchemy state list` / `alchemy state get` — the default, for every question.
   Follow referenced ids between records when the answer spans resources.
2. `alchemy.run.ts` and `src/` under `/workspace/alchemy` — for intent the state
   doesn't surface directly.
3. `aws ec2 …` — Alchemy treats cloud state as authoritative, so use it for
   runtime values the state does not carry (instance states, allocated
   addresses).

Pass rate by question

Of 6 trials: 2 questions, 3 attempts each.

subnets-with-no-network-interfac0/3

Which of my subnets have no network interfaces in them?

Graded against 8 of 13, across three regions✗ ✗ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, AWS CDK 2/3, Alchemy 2/3, Alchemy v2 (Effect) 0/3, Terraform 0/3.

vpcs-with-no-running-instances0/3

Which of my VPCs have no running instances?

Graded against 2 of 6✗ ✗ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, AWS CDK 3/3, Alchemy 2/3, Alchemy v2 (Effect) 1/3, Terraform 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
Pulumi—
chant, best$0.0249
field average$0.1700
per question asked
Pulumi$0.0671
No tool (AWS CLI), best$0.0234
field average$0.0800
tokens in
Pulumi298,344
No tool (AWS CLI), best85,040
field average360,870
tokens out
Pulumi4,184
chant, best1,707
field average4,585

Work per answer

What the agent had to do to get there.

commands
Pulumi8.83
chant, best2
field average11.93
turns
Pulumi11
chant, best4
field average14.1
clock time
Pulumi51s
No tool (AWS CLI), best23s
field average81s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads
Pulumi0
Terraform, runner up0
field average15

Agent environment

Identical for every arm except the tool and its briefing, which are what the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
pulumi-neg-i3
tool under test
pulumi 3.255.0
what the run cost
$0.4025 — 6 questions at $0.0671 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/pulumi
harness
e8c259c
briefing
briefing-pulumi.md · a06c6b73c0eb

Repeat this run:./benchmarks/agent-env/run-negatives.sh pulumi

The briefing this agent received, in full
# Answer estate questions from the Pulumi state — it is the source of truth

This AWS estate was deployed from the Pulumi program mounted read-only at
`/workspace/pulumi`, already applied. The exported state records every resource
with its resolved live ids, its inputs and outputs, and the dependency edges
between resources.

**Query the state rather than enumerating the account resource by resource.** A
raw `aws ec2` sweep returns per-resource facts with no relationships; the state
export already holds the graph, and it is the complete set of managed resources,
so you know the denominator.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root:

- `cd /workspace/pulumi && ./pulumi-export` — the whole applied state as JSON.
  Each entry under `.deployment.resources[]` has:
  - `type` — the resource type, e.g. `aws:ec2/instance:Instance`
  - `urn` — its unique name
  - `inputs` — what was declared
  - `outputs` — the resolved attributes, including physical ids
  - `parent` and `dependencies` — the edges to other resources

  `jq` over `.deployment.resources[]` answers relationship questions without
  hand-joining CLI output — filter by `type`, then follow `dependencies` or an
  output id into the resources that reference it.

Path to estate facts, in order:

1. `./pulumi-export` piped through `jq` — the default, for every question. Use
   `dependencies`/`parent` and output ids when the answer spans resources.
2. The `index.ts` source under `/workspace/pulumi` — for intent and
   configuration the export doesn't surface directly.
3. `aws ec2 …` — for runtime values the state does not carry (instance states,
   allocated addresses).

Pass rate by question

Of 6 trials: 2 questions, 3 attempts each.

subnets-with-no-network-interfac0/3

Which of my subnets have no network interfaces in them?

Graded against 8 of 13, across three regions✗ ✗ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, AWS CDK 2/3, Alchemy 2/3, Pulumi 0/3, Terraform 0/3.

vpcs-with-no-running-instances1/3

Which of my VPCs have no running instances?

Graded against 2 of 6✗ ✓ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, AWS CDK 3/3, Alchemy 2/3, Pulumi 0/3, Terraform 0/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
Alchemy v2 (Effect)$0.5021
chant, best$0.0249
field average$0.1700
per question asked
Alchemy v2 (Effect)$0.0837
No tool (AWS CLI), best$0.0234
field average$0.0800
tokens in
Alchemy v2 (Effect)404,892
No tool (AWS CLI), best85,040
field average360,870
tokens out
Alchemy v2 (Effect)4,564
chant, best1,707
field average4,585

Work per answer

What the agent had to do to get there.

commands
Alchemy v2 (Effect)11.67
chant, best2
field average11.93
turns
Alchemy v2 (Effect)14.17
chant, best4
field average14.1
clock time
Alchemy v2 (Effect)76s
No tool (AWS CLI), best23s
field average81s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads
Alchemy v2 (Effect)2
Pulumi, best0
field average15

Agent environment

Identical for every arm except the tool and its briefing, which are what the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
alchemy-effect-neg-i3
tool under test
alchemy 2.0.0-beta.70
what the run cost
$0.5022 — 6 questions at $0.0837 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/alchemy
harness
e8c259c
briefing
briefing-alchemy-effect.md · fddba9c087d1

Repeat this run:./benchmarks/agent-env/run-negatives.sh alchemy-effect

The briefing this agent received, in full
# Answer estate questions from the Alchemy state — it is the source of truth

This AWS estate was deployed from the Alchemy program mounted read-only at
`/workspace/alchemy`, already applied, and the Alchemy CLI is installed in it.
The applied state records every resource with its resolved live ids and
attributes. This estate is deployed as one stack per region, with an entrypoint
each: `us-east-1.run.ts`, `us-west-1.run.ts`, `us-west-2.run.ts`.

**Query the state rather than enumerating the account resource by resource.** A
raw `aws ec2` sweep returns per-resource facts with no relationships; the state
already holds each resource's resolved attributes and the ids it references, and
`state export` returns every record in the store as one JSON document, so you
know the denominator from a single call.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root with `--local`, which reads the on-disk store under
`.alchemy/state`. That store holds all three regions, and one entrypoint reaches
every stack in it — `--stack` is what selects the region, not the entrypoint. Use
`us-west-1.run.ts` as the handle throughout:

- `alchemy state export us-west-1.run.ts --local` — **every resource in every
  stack as one JSON document**: a flat `resources` array of
  `{stack, stage, fqn, state}`, where `state` is the same record `state get`
  prints. One call covers all three regions; filter it with `jq`. This answers
  most questions by itself.
- `alchemy state export --stack <stack> us-west-1.run.ts --local` — the same,
  narrowed to one region's stack.
- `alchemy state tree us-west-1.run.ts --local` — every stack and stage with the
  resources under it, when you want the census without the records.
- `alchemy state stacks us-west-1.run.ts --local` and
  `alchemy state stages us-west-1.run.ts --local` — the stacks and stages present.
- `alchemy state get --stack <stack> --stage <stage> --fqn <fqn> us-west-1.run.ts --local`
  — one resource with its resolved attributes, when you already know its name.

`alchemy state stacks` lists all three region stacks whichever entrypoint you
name, so one command per question covers the estate. The same records are on disk
under `/workspace/alchemy/.alchemy/state/*/bench/*.json` — one stack directory
per region, one JSON file per resource, each with a `resourceType`, a `props`
object holding the declared configuration and an `attr` object holding the
resolved attributes — if you would rather `jq` or grep the files directly.

Path to estate facts, in order:

1. `alchemy state export … --local` piped through `jq` — the default, for every
   question. The whole estate is in one document, so relationship questions are
   a join over the array rather than a walk between commands.
2. The `*.run.ts` stacks and `src/` under `/workspace/alchemy` — for intent the
   state doesn't surface directly.
3. `aws ec2 …` — Alchemy treats cloud state as authoritative, so use it for
   runtime values the state does not carry (instance states, allocated
   addresses).

Pass rate by question

Of 6 trials: 2 questions, 3 attempts each.

subnets-with-no-network-interfac0/3

Which of my subnets have no network interfaces in them?

Graded against 8 of 13, across three regions✗ ✗ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, AWS CDK 2/3, Alchemy 2/3, Pulumi 0/3, Alchemy v2 (Effect) 0/3.

vpcs-with-no-running-instances0/3

Which of my VPCs have no running instances?

Graded against 2 of 6✗ ✗ ✗

Everyone else: chant 3/3, No tool (AWS CLI) 2/3, AWS CDK 3/3, Alchemy 2/3, Pulumi 0/3, Alchemy v2 (Effect) 1/3.

What one answer cost

The agent's own billed total, not tokens times a rate card. Per correct answer is that divided by the share the tool gets right — the expected spend before an answer arrives that holds up.

per correct answer
Terraform—
chant, best$0.0249
field average$0.1700
per question asked
Terraform$0.1240
No tool (AWS CLI), best$0.0234
field average$0.0800
tokens in
Terraform550,789
No tool (AWS CLI), best85,040
field average360,870
tokens out
Terraform7,153
chant, best1,707
field average4,585

Work per answer

What the agent had to do to get there.

commands
Terraform17.17
chant, best2
field average11.93
turns
Terraform20.17
chant, best4
field average14.1
clock time
Terraform108s
No tool (AWS CLI), best23s
field average81s

Independence

Reads of the live account while answering. CDK and the baseline keep no state, so theirs are the sanctioned path.

account reads
Terraform0
Pulumi, runner up0
field average15

Agent environment

Identical for every arm except the tool and its briefing, which are what the comparison is about. A run only compares with another that shares the harness commit and the briefing hash.

run
terraform-neg-i3
tool under test
terraform 1.15.8
what the run cost
$0.7439 — 6 questions at $0.1240 each
agent
claude-code
model
claude-haiku-4-5-20251001
attempts per question
k=3
substrate
floci emulator, no AWS account and no spend
workdir
/workspace/terraform
harness
e8c259c
briefing
briefing-terraform.md · 7822d55ca7ca

Repeat this run:./benchmarks/agent-env/run-negatives.sh terraform

The briefing this agent received, in full
# Answer estate questions from Terraform state — it is the source of truth

This AWS estate was deployed from the Terraform configuration mounted read-only
at `/workspace/terraform`, already applied, and the Terraform CLI is vendored in
the workspace. The applied state records every managed resource with its
resolved live ids, its attributes, and the references between resources.

**Query the state rather than enumerating the account resource by resource.** A
raw `aws ec2` sweep returns per-resource facts with no relationships; the state
already holds how resources reference one another, and `state list` gives you
the complete set under management, so you know the denominator.

A security group can reach an instance indirectly: a launch template can carry
security-group ids that the instance's own record never lists. Anything you
conclude about what reaches an instance has to account for both the groups
attached to it directly and any it picks up from a template it was launched
from.

Run from the project root (use the vendored binary, `./terraform`):

- `cd /workspace/terraform && ./terraform state list` — every resource address
  under management, one per line. This is the full inventory.
- `cd /workspace/terraform && ./terraform state show <address>` — one resource
  with all of its resolved attributes.
- `cd /workspace/terraform && ./terraform show -json` — the whole applied state
  as JSON. Resources live under `.values.root_module` (recurse
  `child_modules`); each has `type`, `address`, and a `values` object with the
  resolved attributes. `jq` over this answers relationship questions without
  hand-joining CLI output.
- `cd /workspace/terraform && ./terraform output -json` — the declared outputs.

Path to estate facts, in order:

1. `./terraform show -json` or `state show` — the default, for every question.
   Follow attribute references (subnet ids, security-group ids, launch-template
   ids) between resources to answer questions that span them.
2. The `.tf` source under `/workspace/terraform` — for intent and configuration
   the state doesn't surface directly.
3. `aws ec2 …` — for runtime values the state does not carry (instance states,
   allocated addresses).
What this is measuring

Whether a tool can find what nothing points at. Both questions ask for resources that no deployment created and no state file records: subnets holding no network interface, VPCs holding no instance.

The arms split on one axis, and it is not whether they have tooling. AWS CDK scores as well as anything here, because it keeps no state of its own and has to read the account to answer at all — the same route the no-tool baseline takes. The arms that hold only a record of what they deployed cannot see the rest of the estate, and score zero.

So the question a row answers is not does this tool help, it is what can this tool see.

Reading account reads

Here the column cuts both ways. Reads are how the account-reading arms reach these answers at all, so a zero beside a high score is the interesting cell: it means the arm answered from state it already held.