Evidence · How close AWS is
reference-k8s-workloads
hand-written reference shape kept in this repository, the kubernetes lane’s workload-breadth estate (#1884, epic #1885): 22 objects in two namespaces declaring every typed workload kind no other estate did - kubernetes_job_v1 (completes), kubernetes_cron_job_v1 (suspended, never fires), kubernetes_daemon_set_v1, kubernetes_ingress_v1, kubernetes_network_policy_v1, kubernetes_horizontal_pod_autoscaler_v2 over a two-replica Deployment, a directly declared kubernetes_persistent_volume_claim_v1, kubernetes_role_v1 and kubernetes_role_binding_v1, kubernetes_limit_range_v1 and kubernetes_resource_quota_v1 - plus the deprecated aliases kubernetes_daemonset, kubernetes_role and kubernetes_network_policy on APIs the cluster still serves. Small multi-arch images, no StorageClass. #1107’s fallback rule: the published roots with these kinds use Helm or a cloud provider, or are two to seven objects
Set: growing. Lane: kubernetes.
Clear. Every headline stage passes.
| Stage | Verdict | Duration | Detail |
|---|---|---|---|
| Cold deploy | pass | 1m18s | 22 objects in two namespaces from plain terraform against kind v1.37.0: Job, CronJob (suspended), two DaemonSets, Ingress, two NetworkPolicies, HPA v2 over a two-replica Deployment, a directly declared PVC, two Roles and a RoleBinding, a LimitRange and a count ResourceQuota, plus a Namespace pair, a ServiceAccount, a Service and three ConfigMaps - 19 on the _v1/_v2 types and 3 on the deprecated aliases kubernetes_daemonset, kubernetes_role and kubernetes_network_policy. A real terraform.tfstate with 22 instances, stock’s own replan empty, zero tofu-estate labels read back with kubectl over 17 kinds. The Job completed (succeeded=1, its Pod migrate-s5794), the suspended CronJob owns 0 Jobs, the declared PVC is Bound by its first consumer on kind’s standard class, the DaemonSets read 1/1 ready on 1 node(s). Stock’s deprecation warnings: Warning: Deprecated Resource (+8 similar elsewhere);Warning: Deprecated Resource @ kubernetes_daemonset.legacy_agent. The identical shape cold-deployed by stock on a second cluster as every later stage’s oracle |
| Migrate | pass | 2s | 22 of 22 stamped, 0 failed, 0 skipped (22 resource(s) newly stamped, 0 already stamped, 0 newly recorded, 0 re-recorded for sensitivity only, 0 already recorded, 0 failed, 0 skipped); every object carries tofu-estate=reference-k8s-workloads, read back with kubectl over 17 kinds, the three deprecated-alias objects included, and no Pod does - not the Job’s, not the Deployment’s, not either DaemonSet’s |
| Replan from nothing | pass | 23s | the plan with no state file is empty; all 22 identities (KIND NAMESPACE/NAME) confirmed present with kubectl across 17 kinds and two namespaces, the deprecated-alias objects bound by the same label-and-name join as their _v1 kinds (kubernetes_daemonset’s included, which joined to a kind the server does not list until #1884). The deprecation warnings choudoufu printed are exactly stock’s for the same configuration on the oracle cluster, by warning and by address: Warning: Deprecated Resource @ kubernetes_daemonset.legacy_agent;Warning: Deprecated Resource @ kubernetes_network_policy.legacy_deny;Warning: Deprecated Resource @ kubernetes_role.legacy_reader, and no error |
| No-op apply | pass | 13s | no-op apply (0 added, 0 changed, 0 destroyed); objects carrying tofu-estate=reference-k8s-workloads unchanged at 22 across the estate’s 17 kinds, and still zero Pods, counted with kubectl. The apply’s deprecation warnings: Warning: Deprecated Resource (+5 similar elsewhere);Warning: Deprecated Resource @ kubernetes_daemonset.legacy_agent |
| Drift and reconverge | pass | 35s | the HPA’s minReplicas raised 2 -> 3 with kubectl patch; the HPA controller then rewrote the Deployment’s spec.replicas 2 -> 3 on its own (waited for, bounded). choudoufu proposed exactly kubernetes_horizontal_pod_autoscaler_v2.web (min_replicas 3 -> 2) and kubernetes_deployment_v1.web (replicas 3 -> 2), 0 add, 2 change, 0 destroy, matching stock’s own plan on the oracle cluster for the same drift; no ReplicaSet, Pod or shard ConfigMap was named. The apply changed 2, the HPA’s floor first (the Deployment depends on it, so the controller is never left below a raised floor), both read back as configured, and the next plan is empty. BREAK=1 tampers a third object and the exact-two assertion correctly fails |
| Rename | pass | 23s | moved block: kubernetes_service_account_v1.web -> .frontend, no add and no destroy, one in-place change confined to the address annotation rewrite (0 add, 1 change, 0 destroy) - the marker rewritten in place; the live object untouched and still labelled, and the RoleBinding naming it by its unchanged name untouched, read with kubectl; stock’s plan for the same moved block on the oracle cluster is zero churn, since stock never writes this annotation. The moved-block half only: live-mv’s Kubernetes leg (#1639) is not exercised here. BREAK=1 renames the object’s own metadata.name instead, which plans a destroy and a create |
| Remove a block | pass | 1m8s | deleting kubernetes_cron_job_v1.nightly’s block proposed exactly one destroy (0 add, 0 change, 1 destroy) at the sweep’s orphan address kubernetes_cron_job_v1.orphan_refk8swl-batch_nightly (“Owned and undeclared: 1 live resource will be destroyed”), under the versioned type since no block declares a CronJob any more; applied cleanly, kubectl reads the CronJob gone, it owned no Job at any point (suspended), and the next plan is empty; stock’s plan for the same removal on the oracle cluster is also exactly one destroy. The controller-copy exclusion, on the Job’s own Pod migrate-s5794 (ownerReference to the Job): with the estate’s label put on it by kubectl the sweep proposed nothing. The two-sided arm then declared a bare Pod with kubectl in the same namespace, labelled the same way, and one plan told them apart: Plan: 0 to add, 0 to change, 1 to destroy, naming kubernetes_pod_v1.orphan_refk8swl-batch_declared-orphan and not kubernetes_pod_v1.orphan_refk8swl-batch_migrate-s5794. Both were cleaned up, the plan is empty again and no Pod carries the label. BREAK_REMOVE=1 keeps the block and no destroy is proposed; BREAK_POD=1 runs the probe unlabelled and correctly fails |
| Change count | pass | 57s | scaling kubernetes_config_map_v1.shard from 2 to 1 destroyed exactly shard-1, planned at the sweep’s orphan address kubernetes_config_map_v1.orphan_refk8swl_shard-1 since the label carries no index (shard-0 untouched, both read with kubectl); back to 2 created exactly kubernetes_config_map_v1.shard[1] under the same name; the next plan is empty; stock’s plans for the same two changes on the oracle cluster have the identical shape. BREAK_COUNT=1 asserts the lower index was destroyed and correctly fails |
| Replace with create_before_destroy | pass | 42s | two legs. The shared create_before_destroy rename (gauntlet_kind_day2_replace) passed first: cfg-a -> cfg-b planned as stock’s replace, create first, the new object’s create complete before the old one’s deposed destroy at -parallelism=1, kubectl reading cfg-b alone with the block’s annotation, next plan empty, the block removed again. Then a Job, immutable once created: changing kubernetes_job_v1.migrate’s pod template planned ‘# kubernetes_job_v1.migrate must be replaced’, 1 add and 1 destroy, no orphan beside it, the same shape as stock’s plan on the oracle cluster; the apply replaced it (uid uid=b8ebe858-bf87-435f-bda9-b7d869113c96 -> uid=4e63692e-dd2a-4039-b210-2bdc6dfaa323), the new Job completed with the rev-2 template and carries the estate’s label, it is the only Job in refk8swl-batch, no Pod carries the label, and the next plan is empty. BREAK_REPLACE=1 runs the shared leg’s Break line and skips the Job leg |
| Crash mid-apply | pass | 3m35s | an apply creating two objects was interrupted by a real SIGTERM (exit 1), delivered by the engine itself inside the -parallelism=1 walker the instant kubernetes_secret_v1.crash_first’s create committed; crash_second reads crash_first’s name, so the walker cannot have reached it - kubectl confirms crash-first exists carrying tofu-estate=reference-k8s-workloads and crash-second does not. The next plan proposed exactly the remainder (Plan: 1 to add, 0 to change, 0 to destroy) and nothing for crash-first, bound by its label and its namespace and name, matching stock’s own plan from the same position on the oracle cluster; the recovery apply added exactly one object and the plan after it is empty. The record’s contribution is read, not counted: the interrupted apply wrote exactly one record (envelopes 16 -> 17) carrying residue wait_for_service_account_token, and taking it out turns the recovery plan into Plan: 1 to add, 1 to change, 0 to destroy, proposing wait_for_service_account_token back; putting it back restores the remainder. BREAK_CRASH=1 and BREAK_CRASH_UNBOUND=1 correctly fail The create_before_destroy rename window is interrupted too (#1768): a kubernetes_config_map renamed under create_before_destroy was killed by the engine’s own hook the instant the new object’s create committed, at -parallelism=1, leaving both objects carrying the block’s address annotation and the record holding the old one as the address’s deposed object. The verdict is the end state stock’s replace leaves, not the plan’s wording: after one more apply exactly the new object remains, the old one is gone, the record’s deposed entry is cleared and the replan is empty. Both of the rerun’s paths reached it - orphan leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘orphan_refk8swl_crash-rename-a’; deposed leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘(deposed object’; the same configuration through the orphan destroy, and the name read from another block’s attribute (#1539’s shape) through the deposed record (#1683), where stock’s plan reads the same deposed-object destroy. |
| Teardown | pass | 33s | apply -destroy removed exactly the 23 remaining labelled objects in one apply, naming no Pod or ReplicaSet; both namespaces are gone (waited for, bounded) and no object of any of the estate’s 17 kinds carries tofu-estate=reference-k8s-workloads (kubectl, every namespace); the Job’s Pod, the Deployment’s ReplicaSet and Pods and both DaemonSets’ Pods went with their owners and namespaces; stock’s destroy of the same estate on the oracle cluster also removed exactly the 23 its state held |
| Plan, review, apply | pass | 34s | plan -out wrote one update (web-config gains reviewed=yes); the world then moved out of band (a stray label on shard-0, kubectl, never choudoufu) and apply of the saved plan refused with “The approved plan no longer matches the live system” at exit 3, nothing applied (kubectl reads no reviewed key); with the label removed the identical file applied, 0 added, 1 changed, 0 destroyed, and reviewed=yes reads back; stock’s own planfile applied on the oracle cluster in the unchanged case. BREAK_APPROVAL=1 expects success after the move and correctly fails |
| Greenfield apply | pass | 1m17s | 22 objects applied fresh with a live block and no terraform.tfstate, every one labelled tofu-estate=reference-k8s-workloads (kubectl, 17 kinds) and no Pod; the apply’s deprecation warnings are exactly stock’s cold-deploy apply’s; the record store held 11 record envelope(s); replanned empty with and without the cache. Deleting the whole record store and the cache and replanning proposed Plan: 0 to add, 5 to change, 0 to destroy: nothing created, destroyed or swept, every object still bound by its label and its namespace and name, and 5 in-place update(s) putting back the residue the store held (create match_labels update wait_for_load_balancer wait_for_rollout); one apply reconverged and the plan after it is empty. The cluster’s inventory matches stock’s cold deploy on the same cluster object by object - every object’s spec, data, rules, role reference and subjects with metadata and status left out and server-assigned cluster IPs, bound volume and controller-uid selectors scrubbed - and the Job completed again. BREAK=1 drops the legacy DaemonSet from the expected inventory and the match correctly fails |
| Strict profile (not a headline stage) | pass | 19s | every strict toggle on (secrets = refuse, no_source_create = refuse, marker_repair = never with a markers “record” selection naming kubernetes_config_map_v1) against a scratch estate carrying random_password.db: exactly one refusal, Logical resource is not admitted under strict { secrets = “refuse” }; the other two toggles are on and silent. BREAK_STRICT=1 turns secrets back to “store” and the refusal disappears |
| Plan with no local state (not a headline stage) | not run |
Last run at commit 4bfb459f95 on 2026-10-05T02:35:04Z, exit code 0, against substrate image kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5. Total run time 12m58.2s.
Oracle: stock terraform 1.15.8, stock tofu 1.12.5. Stale: the current pin is terraform 1.16.1, tofu 1.13.0.
Engine: OpenTofu base 1.13.0 (matches the current base).
Reproduce it
go run ./tools/gauntlet run reference-k8s-workloads
Needs Docker (the emulator is pulled at the pinned digest), the AWS CLI, and a
stock terraform or tofu binary on PATH for the cold deploy. The script is
live/e2e/reference-k8s-workloads/run.sh; BREAK=1 corrupts its assertions to show they are load-bearing.