Evidence · How close AWS is

reference-k8s-platform-app

hand-written reference shape kept in this repository, the kubernetes lane’s two-estate pair (#1883, epic #1885): network (a Namespace, a NetworkPolicy, a Service, two ConfigMaps and two root outputs) and app (a Namespace, three ConfigMaps, a ServiceAccount, a Service, a Deployment, a HorizontalPodAutoscaler v2 and a kubernetes_manifest ConfigMap in network’s namespace) - 14 objects over seven kinds, both estates on record_store “kubernetes”, app reading network’s outputs through reads_outputs_of and data terraform_estate_outputs where stock reads terraform_remote_state. Measures what only smoke scenarios covered: records in the cluster, a declared cross-estate read refused and then granted, live-mv -from-estate on a kubernetes_manifest killed between its two requests and re-run, the cache-served unchanged plan, and teardown app first. Under #1107’s fallback rule: no published Kubernetes-only root splits one team’s objects from another’s and reads across the split

Set: growing. Lane: kubernetes.

Clear. Every headline stage passes.

StageVerdictDurationDetail
Cold deploypass1m5splain terraform against kind v1.37.0: network’s 5 objects (Namespace, NetworkPolicy, Service, 2 ConfigMaps) and two root outputs, then app’s 9 (Namespace, 3 ConfigMaps, Service, Deployment, HorizontalPodAutoscaler v2, ServiceAccount, and a kubernetes_manifest ConfigMap in network’s namespace) reading network’s outputs through data terraform_remote_state; two terraform.tfstate files, zero tofu-estate labels read back with kubectl; the same two estates cold-deployed by stock on a second cluster as every later stage’s oracle
Migratepass43snetwork: 5 resource(s) newly stamped, 0 already stamped, 0 newly recorded, 0 re-recorded for sensitivity only, 0 already recorded, 0 failed, 0 skipped; its first apply recorded both root outputs as Secrets in tofu-records-network (Apply complete! Resources: 0 added, 0 changed, 0 destroyed.). app: 9 resource(s) newly stamped, 0 already stamped, 0 newly recorded, 0 re-recorded for sensitivity only, 0 already recorded, 0 failed, 0 skipped, its live-import reading network’s outputs through data terraform_estate_outputs (#1862). Every object carries its estate’s label, read back with kubectl (network 5 of 5, app 9 of 9, the handoff manifest in network’s namespace carrying tofu-estate=app); records in the cluster, record_store “kubernetes”: no .tofu-records directory and no terraform.tfstate in either live root
Replan from nothingpass35sboth plans with no state file are empty as the cluster admin (network 5 objects, app 9). app’s plan under app-planner - the estate fence for app, get/list/watch cluster-wide on every listable kind but Secrets, every verb on app’s declared kinds and its own records in tofu-records-app, nothing in tofu-records-network - refused with “This estate may not read another estate’s outputs” naming estate network and tofu-records-network (exit 1); with one Role granting get on network’s 2 output Secret(s) by name, and still no list there, the same plan is empty and says the values are as of network’s last apply. Nothing in tofu-records-network changed across either plan (names and resourceVersions). BREAK_READ=1 grants the read first and the refusal correctly does not appear
No-op applypass34sno-op apply in both estates (0 added, 0 changed, 0 destroyed each); objects carrying tofu-estate unchanged at network 5 and app 9, counted with kubectl outside the records namespaces
Drift and reconvergepass29sapp-config tampered with kubectl patch; choudoufu proposed exactly kubernetes_config_map.app (0 add, 1 change, 0 destroy), matching stock’s own plan on the oracle cluster; apply changed 1 and the gateway value - built from network’s recorded outputs - reads back as gateway.platform.svc.cluster.local. BREAK=1 tampers a second object and the single-object assertion correctly fails
Renamepass30smoved block kubernetes_service_account.app -> .team: no add and no destroy, one in-place change confined to the address annotation rewrite (0 add, 1 change, 0 destroy); the ServiceAccount untouched and still labelled, read with kubectl; stock’s plan for the same moved block on the oracle cluster is zero churn, since stock never writes the annotation. BREAK=1 renames metadata.name instead, a real identity change, and the in-place assertion correctly fails
Remove a blockpass40sdeleting kubernetes_service_account.team’s block proposed exactly one destroy at the sweep’s orphan address kubernetes_service_account_v1.orphan_shop_app, applied cleanly, the ServiceAccount gone (kubectl) and the next plan empty; stock’s plan for the same removal on the oracle cluster is also exactly one destroy. BREAK_REMOVE=1 keeps the block and no destroy is proposed
Change countpass1m7sscaling kubernetes_config_map.shard 2 -> 1 destroyed exactly shard-1 at the sweep’s orphan address kubernetes_config_map.orphan_shop_shard-1 (shard-0 untouched, kubectl); back to 2 created exactly shard[1]; the next plan is empty; stock’s plans for the same two changes on the oracle cluster have the identical shape. BREAK_COUNT=1 asserts the lower index was destroyed and correctly fails
Replace with create_before_destroypass1m16sa create_before_destroy ConfigMap whose content-hashed name changes (cfg-a -> cfg-b) plans as stock’s replace: ‘kubernetes_config_map.hashed must be replaced’, +/- create replacement and then destroy, 1 add and 1 destroy, with no orphan destroy beside it, because cfg-a carries the block’s address annotation and the sweep binds it (#1640). At -parallelism=1 the apply log shows cfg-b’s creation complete (line 66) before cfg-a’s deposed destroy starts (line 67), the same order stock’s apply shows on the oracle cluster; kubectl confirms cfg-b alone remains, carrying the annotation, and the next plan is empty (#1541). The block is removed from both roots afterwards. BREAK_REPLACE=1 recreates cfg-a carrying the block’s annotation and the next plan correctly proposes destroying it
Crash mid-applypass5m27sAn apply creating two objects was killed by the engine’s own SIGTERM the instant kubernetes_secret.crash_first’s create committed (exit 1, -parallelism=1, crash_second reads crash_first’s name); the next plan proposed exactly the remainder (Plan: 1 to add, 0 to change, 0 to destroy) and nothing for crash-first, matching stock’s plan from the same position on the oracle cluster, and one apply finished it. The record the interrupted apply wrote is a Secret in tofu-records-app (envelopes 13 -> 14) carrying residue wait_for_service_account_token: taking that Secret out turns the plan into Plan: 1 to add, 1 to change, 0 to destroy, putting it back restores the remainder. Oracle on B: terraform state rm in app and terraform import into network; stock’s import of a kubernetes_manifest leaves manifest unset in state, so stock needed an apply after the import (0 added, 1 changed, 0 destroyed, the live ConfigMap’s data, labels and annotations unchanged) before both its plans were empty. live-mv -from-estate=app of kubernetes_manifest.handoff (a ConfigMap in platform) into network was killed by SIGKILL between its two requests (TOFU_E2E_LIVE_MV_INTERRUPT on the e2e build, exit 137): kubectl read tofu-estate=network and the new address on the object with the markers still held by an Update entry; the plain rerun finished the hand-off (no Update entry holds a marker), network’s and app’s plans are empty and network’s apply goes through, the same end state as stock’s state rm and import on the oracle cluster. The create_before_destroy rename window is interrupted too (#1768): a kubernetes_config_map renamed under create_before_destroy was killed by the engine’s own hook the instant the new object’s create committed, at -parallelism=1, leaving both objects carrying the block’s address annotation and the record holding the old one as the address’s deposed object. The verdict is the end state stock’s replace leaves, not the plan’s wording: after one more apply exactly the new object remains, the old one is gone, the record’s deposed entry is cleared and the replan is empty. Both of the rerun’s paths reached it - orphan leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘orphan_shop_crash-rename-a’; deposed leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘(deposed object’; the same configuration through the orphan destroy, and the name read from another block’s attribute (#1539’s shape) through the deposed record (#1683), where stock’s plan reads the same deposed-object destroy. BREAK_CRASH=1 asserts nothing is proposed after the apply interrupts and correctly fails; BREAK_MV=1 asserts the hand-off finished without the rerun and correctly fails
Teardownpass1m3sapp’s apply -destroy removed exactly its 9 objects while network’s 6 stood untouched and app’s records left tofu-records-app with them (0 instance record Secrets; the store’s own provisioning sentinel stays by design); then network’s removed its 6 (the handoff manifest it took over by live-mv included), both namespaces are gone, no object of either estate carries tofu-estate (kubectl, every namespace but the records ones), and tofu-records-network holds 0 record and 0 output Secrets - so app’s read of network’s outputs would now be told they are not recorded. Guided-discovery hints left in the records namespaces: 2. Stock’s destroys of the same two estates on the oracle cluster, app then network, both completed
Plan, review, applypass40splan -out wrote one update (app-config gains reviewed=yes); a stray label on shard-0 (kubectl) moved the world and apply of the saved plan refused with “The approved plan no longer matches the live system” at exit 3, nothing applied; with the label removed the identical file applied, 0 added, 1 changed, 0 destroyed, and reviewed=yes reads back; stock’s own planfile applied on the oracle cluster. BREAK_APPROVAL=1 expects success after the move and correctly fails
Greenfield applypass1m26snetwork then app applied fresh with live blocks and record_store “kubernetes” (5 and 9 added), no terraform.tfstate and no .tofu-records directory in either root; every object labelled with its estate (kubectl); network’s two outputs recorded as Secrets in tofu-records-network and app’s 8 record envelope(s) read back from tofu-records-app; both replanned empty with and without the cache. The two estates’ inventory - the NetworkPolicy, both Services, the five ConfigMaps (the handoff manifest’s included), the Deployment and its GATEWAY_HOST built from network’s outputs, the HPA and the ServiceAccount - matches stock’s cold deploy on the same cluster object by object, labels never compared. BREAK=1 drops the HPA from the expected inventory and the match correctly fails
Strict profile (not a headline stage)pass45severy strict toggle on (secrets = refuse, no_source_create = refuse, marker_repair = never with a markers “record” selection naming kubernetes_config_map) against a scratch estate carrying random_password.db: exactly one refusal, Logical resource is not admitted under strict { secrets = “refuse” }. BREAK_STRICT=1 turns secrets back to “store” and the refusal disappears
Plan with no local state (not a headline stage)pass50sapp’s records live in the cluster, so the only local state is the state cache; with it deleted - a fresh clone - the plan found every declared object by its label and namespace and name: nothing created, destroyed or replaced, 0 in-place update(s). The unchanged plan before the deletion was served from the cache (written by a no-op apply, since a plan never persists it): 9 instance(s) answered by a cache hit (#1864’s vouch), 81 requests against reads = “full”’s 104 with 0 hits, both plans empty. plan_calls_no_local_state=104 plan_calls_cache_serving=81 ratio=1.28x (all plan -refresh=false, Kubernetes API requests counted on the wire through live/smoke/k8sproxy.py); stock in this position has no plan at all, only one import block per object. BREAK_NO_LOCAL_STATE=1 strips the web Service’s label and the check correctly fails

Last run at commit 4bfb459f95 on 2026-10-05T02:52:31Z, exit code 0, against substrate image kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5. Total run time 17m11.9s. Oracle: stock terraform 1.15.8, stock tofu 1.12.5. Stale: the current pin is terraform 1.16.1, tofu 1.13.0. Engine: OpenTofu base 1.13.0 (matches the current base).

Reproduce it

go run ./tools/gauntlet run reference-k8s-platform-app

Needs Docker (the emulator is pulled at the pinned digest), the AWS CLI, and a stock terraform or tofu binary on PATH for the cold deploy. The script is live/e2e/reference-k8s-platform-app/run.sh; BREAK=1 corrupts its assertions to show they are load-bearing.