Evidence · How close AWS is

reference-k8s-stateful

hand-written reference shape kept in this repository, the kubernetes lane’s stateful estate (#1175, surface 2 of #1107): a Namespace, ServiceAccount, Secret, two ConfigMaps and a two-instance count ConfigMap, headless Services and StatefulSets for postgres:17-alpine (1 replica) and redis:7-alpine (2 replicas) whose volume_claim_template labels disagree with its selector on purpose, a ClusterIP Service and Deployment reading both, and a PodDisruptionBudget - 14 objects over eight kinds, typed _v1 resources only, no StorageClass so kind’s own standard class does the WaitForFirstConsumer binding. #1107’s second research comment fetched and ran the published field for this surface and nothing cleared the bar, so this is that issue’s explicit hand-written fallback

Set: growing. Lane: kubernetes.

Clear. Every headline stage passes.

StageVerdictDurationDetail
Cold deploypass1m34s14 objects over eight kinds (namespace, ServiceAccount, Secret, 4 ConfigMaps, 2 headless Services and a ClusterIP one, 2 StatefulSets, a Deployment, a PodDisruptionBudget) from plain terraform against kind v1.37.0, a real terraform.tfstate with 14 instances, zero tofu-estate labels read back with kubectl. The two volume_claim_templates produced 3 PVCs nobody declared, all Bound on kind’s own standard class (rancher.io/local-path, WaitForFirstConsumer, reclaim Delete) with no StorageClass in the root, and stock’s state holds none of them; the identical shape cold-deployed by stock on a second cluster as every later stage’s oracle
Migratepass2s14 of 14 stamped, 0 failed, 0 skipped (14 resource(s) newly stamped, 0 already stamped, 0 newly recorded, 0 re-recorded for sensitivity only, 0 already recorded, 0 failed, 0 skipped); every object carries tofu-estate=reference-k8s-stateful, read back with kubectl over eight kinds, and none of the three controller-created PVCs does
Replan from nothingpass11sthe plan with no state file is empty; all 14 identities (NAMESPACE/NAME) confirmed present with kubectl across eight kinds
No-op applypass12sno-op apply (0 added, 0 changed, 0 destroyed); objects carrying tofu-estate=reference-k8s-stateful unchanged at 14 across the estate’s eight kinds, and still zero of the three PVCs, counted with kubectl
Drift and reconvergepass23sone ConfigMap tampered with kubectl patch; choudoufu proposed exactly kubernetes_config_map_v1.api (0 add, 1 change, 0 destroy), matching stock’s own plan on the oracle cluster for the same tamper; apply changed 1 and the value reads back as configured. Neither the shard ConfigMaps nor any of the three undeclared PVCs was proposed. BREAK=1 tampers a second object and the single-object assertion correctly fails
Renamepass23smoved block: kubernetes_service_account_v1.app -> .team, no add and no destroy, one in-place change confined to the address annotation rewrite (0 add, 1 change, 0 destroy) - the marker rewritten in place, the same shape the AWS lanes assert for a rename, not literal zero churn; the live object untouched and still labelled, read with kubectl; stock’s plan for the same moved block on the oracle cluster is zero churn, since stock never writes this annotation. The moved-block half only: live-mv also has a Kubernetes leg since #1639, not exercised by this stage. A bare block rename without a moved block plans the same one in-place change, since the block name is not part of the object’s identity; BREAK=1 renames the object’s own metadata.name instead, which is a genuine identity change and plans a destroy and a create
Remove a blockpass1m29sdeleting kubernetes_stateful_set_v1.redis’s block proposed exactly one destroy (0 add, 0 change, 1 destroy) at the sweep’s synthetic orphan address kubernetes_stateful_set_v1.orphan_refk8sst_redis (“Owned and undeclared: 1 live resource will be destroyed”), applied cleanly, the object gone from the cluster (kubectl get statefulset redis: NotFound) and the next plan empty; stock’s plan for the same removal on the oracle cluster is also exactly one destroy. The three controller-created PVCs survive the removal unchanged and Bound, as they do under stock (3 left there too), carrying labels the controller wrote - app=redis from the SELECTOR merged over the claim template’s app=redis-data, plus the template’s tier=storage and role=cache-volume - the pvc-protection finalizer, and no ownerReferences, since the default persistentVolumeClaimRetentionPolicy is Retain and they are built to outlive the StatefulSet. Their managedFields name kube-scheduler and kube-controller-manager, and kube-controller-manager alone wrote their spec, which is the signal kubesweep.ControllerMade judges after #1179 (the owner-reference signal its package comment credited is absent here, and kubectl get hides managedFields without –show-managed-fields, which is how #1179 came to record them as absent too). With the estate’s label put on one of them by kubectl - a harsher state than a claim template carrying the marker, since kubectl also registers itself as a field manager (data-redis-0 labels=[app=redis,role=cache-volume,tier=storage,tofu-estate=reference-k8s-stateful] ownerReferences=absent managers=[kube-controller-manager,kube-scheduler,kubectl-label] specWrittenBy=[kube-controller-manager] finalizers=kubernetes.io/pvc-protection storageClass=standard phase=Bound) - the sweep still proposed nothing, so the controller-copy exclusion held. The two-sided arm then labelled that PVC again and declared one with kubectl (spec written by kubectl-create, nobody’s control plane) in the same namespace, and one plan told them apart: Plan: 0 to add, 0 to change, 1 to destroy, naming kubernetes_persistent_volume_claim_v1.orphan_refk8sst_declared-orphan and not kubernetes_persistent_volume_claim_v1.orphan_refk8sst_data-redis-0, so the exclusion is not the sweep being blind to the kind. Both were cleaned up and the plan is empty again. BREAK_REMOVE=1 keeps the block and no destroy is proposed; BREAK_PVC=1 runs the probe unlabelled and correctly fails
Change countpass57sscaling kubernetes_config_map_v1.shard from 2 to 1 destroyed exactly shard-1, planned at the sweep’s orphan address kubernetes_config_map_v1.orphan_refk8sst_shard-1 since the label carries no index (shard-0 untouched, both read with kubectl); back to 2 created exactly kubernetes_config_map_v1.shard[1] under the same name; the next plan is empty; stock’s plans for the same two changes on the oracle cluster have the identical shape. BREAK_COUNT=1 asserts the lower index was destroyed and correctly fails
Replace with create_before_destroypass58sa create_before_destroy ConfigMap whose content-hashed name changes (cfg-a -> cfg-b) plans as stock’s replace: ‘kubernetes_config_map.hashed must be replaced’, +/- create replacement and then destroy, 1 add and 1 destroy, with no orphan destroy beside it, because cfg-a carries the block’s address annotation and the sweep binds it (#1640). At -parallelism=1 the apply log shows cfg-b’s creation complete (line 64) before cfg-a’s deposed destroy starts (line 65), the same order stock’s apply shows on the oracle cluster; kubectl confirms cfg-b alone remains, carrying the annotation, and the next plan is empty (#1541). The block is removed from both roots afterwards. BREAK_REPLACE=1 recreates cfg-a carrying the block’s annotation and the next plan correctly proposes destroying it
Crash mid-applypass3m35san apply creating two objects was interrupted by a real SIGTERM (exit 1), delivered by the engine itself inside the -parallelism=1 graph walker the instant kubernetes_config_map_v1.crash_first’s create committed (internal/command/apply_e2etesting_crash.go); crash_second reads crash_first’s name, so the walker cannot have reached it - kubectl confirms crash-first exists carrying tofu-estate=reference-k8s-stateful and crash-second does not. The next plan proposed exactly the remainder (Plan: 1 to add, 0 to change, 0 to destroy, kubernetes_config_map_v1.crash_second created) and proposed nothing at all for crash-first, which it bound by its label and its namespace and name - not a second create the API server would refuse, not an orphan sweep - matching stock’s own plan from the same position on the oracle cluster; the recovery apply added exactly one object, both read back with kubectl, the plan after it is empty and no PVC carries the estate’s label. The record store’s contribution is read, not counted: the interrupted apply wrote exactly one record (files 15 -> 16) for kubernetes_secret_v1.crash_first carrying residue wait_for_service_account_token, and taking that one file out of the store and replanning from the identical position turns the recovery plan from Plan: 1 to add, 0 to change, 0 to destroy into Plan: 1 to add, 1 to change, 0 to destroy, proposing wait_for_service_account_token back on the object the crash left behind; putting it back restores the exact-remainder plan. The crash pair’s first object is a Secret and not a ConfigMap for that reason (#1235): a kubernetes_config_map(_v1) has neither a ratified identity row nor a config-only argument, records nothing, and made this line the 12 -> 12 count #1188 could not read anything out of. BREAK_CRASH=1 asserts nothing is proposed and correctly fails; BREAK_CRASH_UNBOUND=1 strips the label off crash-first and the same recovery check correctly fails The create_before_destroy rename window is interrupted too (#1768): a kubernetes_config_map renamed under create_before_destroy was killed by the engine’s own hook the instant the new object’s create committed, at -parallelism=1, leaving both objects carrying the block’s address annotation and the record holding the old one as the address’s deposed object. The verdict is the end state stock’s replace leaves, not the plan’s wording: after one more apply exactly the new object remains, the old one is gone, the record’s deposed entry is cleared and the replan is empty. Both of the rerun’s paths reached it - orphan leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘orphan_refk8sst_crash-rename-a’; deposed leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘(deposed object’; the same configuration through the orphan destroy, and the name read from another block’s attribute (#1539’s shape) through the deposed record (#1683), where stock’s plan reads the same deposed-object destroy.
Teardownpass1m57sapply -destroy removed exactly the 15 remaining objects in one apply, naming no PVC; the namespace is gone and no object of any of the estate’s eight kinds carries tofu-estate=reference-k8s-stateful (kubectl, every namespace). The 3 PVCs nobody declared went with it, to zero cluster-wide, because the Namespace is in the root and its deletion cascades - which is exactly why the day2_remove residue above has never shown up in a teardown; stock’s destroy of the same estate on the oracle cluster also removed exactly the 15 its state held
Plan, review, applypass34splan -out wrote one update (api-config gains reviewed=yes); the world then moved out of band (a stray label on shard-0, kubectl, never choudoufu) and apply of the saved plan refused with “The approved plan no longer matches the live system” at exit 3, nothing applied (kubectl reads no reviewed key); with the label removed the identical file applied, 0 added, 1 changed, 0 destroyed, and reviewed=yes reads back; stock’s own planfile applied on the oracle cluster in the unchanged case. BREAK_APPROVAL=1 expects success after the move and correctly fails
Greenfield applypass1m24s14 objects applied fresh with a live block and no terraform.tfstate, every one labelled tofu-estate=reference-k8s-stateful (kubectl, eight kinds) and none of the three PVCs; the record store held 10 record envelope(s) - envelopes counted by their own address field, not files, so guided discovery’s hint at tofu-hints/reference-k8s-stateful and any root output in the same store are not mistaken for records (#1291); replanned empty with and without the cache. Deleting the whole record store and the cache and replanning - the reading six AWS estates take here and no Kubernetes estate took - proposed Plan: 0 to add, 6 to change, 0 to destroy: nothing created, nothing destroyed, nothing swept as an orphan, every object still bound by its label and its namespace and name, and 6 in-place update(s) putting back the residue the store held (wait_for_load_balancer wait_for_rollout wait_for_service_account_token); one apply reconverged and the plan after it is empty, so a lost store costs an apply here and not an object (#1188, #1235). The cluster’s inventory matches stock’s cold deploy on the same cluster object by object, labels never compared - ConfigMap and Secret data, each Service’s headless-ness, ports and selector, both StatefulSets’ replicas, serviceName, containers and volume_claim_templates (claim-template labels included), the Deployment, the PodDisruptionBudget, and the three controller-created PVCs with their merged labels, absent ownerReferences and pvc-protection finalizer. BREAK=1 drops the redis StatefulSet from the expected inventory and the match correctly fails
Strict profile (not a headline stage)pass1m4severy strict toggle on (secrets = refuse, no_source_create = refuse, marker_repair = never with a markers “record” selection naming kubernetes_config_map_v1) against a scratch estate carrying random_password.db: exactly one refusal, Logical resource is not admitted under strict { secrets = “refuse” }; the other two toggles are on and silent. BREAK_STRICT=1 turns secrets back to “store” and the refusal disappears
Plan with no local state (not a headline stage)not run

Last run at commit 4bfb459f95 on 2026-10-05T02:47:27Z, exit code 0, against substrate image kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5. Total run time 14m43.7s. Oracle: stock terraform 1.15.8, stock tofu 1.12.5. Stale: the current pin is terraform 1.16.1, tofu 1.13.0. Engine: OpenTofu base 1.13.0 (matches the current base).

Reproduce it

go run ./tools/gauntlet run reference-k8s-stateful

Needs Docker (the emulator is pulled at the pinned digest), the AWS CLI, and a stock terraform or tofu binary on PATH for the cold deploy. The script is live/e2e/reference-k8s-stateful/run.sh; BREAK=1 corrupts its assertions to show they are load-bearing.