Evidence · How close AWS is
corpus-k8s-metrics-server
kubernetes/k8s.io infra/aws/terraform/kops-infra-ci/metrics-server.tf (commit 4046258e7d), the Kubernetes project’s own metrics-server installation for its kops CI cluster, cut out of its root: two ClusterRoles, two ClusterRoleBindings, a RoleBinding, a ServiceAccount, a Deployment and a Service in kube-system, and the aggregated APIService v1beta1.metrics.k8s.io, all on the provider’s non-_v1 type names (#1880, epic #1885). Measures an aggregated API (kubernetes_api_service) through the sweep and through a rollout that leaves it unavailable, and a teardown that leaves kube-system, a namespace the estate writes into and never declared, as it found it
Source: https://github.com/kubernetes/k8s.io.git at 4046258e7d54fdc156e2f39ad8da51b57be88f55.
Set: growing. Lane: kubernetes.
Clear. Every headline stage passes.
| Stage | Verdict | Duration | Detail |
|---|---|---|---|
| Cold deploy | pass | 1m17s | 9 objects (two ClusterRoles, two ClusterRoleBindings, a RoleBinding, a ServiceAccount, a Deployment and a Service in kube-system, and the APIService v1beta1.metrics.k8s.io) from plain terraform on k8s.io’s own metrics-server.tf at 4046258e7d54fdc156e2f39ad8da51b57be88f55 plus its deltas (cut out of kops-infra-ci with a kind-pointed provider block, –kubelet-insecure-tls), against kind v1.37.0; a real terraform.tfstate with 9 instances, zero tofu-estate labels; the APIService reports Available=True and /apis/metrics.k8s.io/v1beta1 answers through the API server; kube-system’s own objects read beforehand as the teardown’s baseline; the identical root cold-deployed by stock on a second cluster as every later stage’s oracle |
| Migrate | pass | 3s | 9 of 9 stamped, 0 skipped, from the stock state file; every object carries tofu-estate=corpus-k8s-metrics-server, read back with kubectl across the estate’s kinds, the cluster-scoped APIService v1beta1.metrics.k8s.io among them |
| Replan from nothing | pass | 11s | the plan with no state file is empty; all nine identities confirmed present by name with kubectl - the five cluster-scoped ones (the APIService, both ClusterRoles, both ClusterRoleBindings) by NAME, the four in kube-system by NAMESPACE/NAME |
| No-op apply | pass | 11s | no-op apply (0 added, 0 changed, 0 destroyed); objects carrying tofu-estate=corpus-k8s-metrics-server unchanged at 9 across the estate’s kinds, counted with kubectl |
| Drift and reconverge | pass | 46s | metrics-server’s –secure-port moved 10250 -> 10251 with kubectl patch, a rollout after which the APIService v1beta1.metrics.k8s.io reported Available=False (FailedDiscoveryCheck) and /apis/metrics.k8s.io/v1beta1 stopped answering; planned in that state, choudoufu proposed exactly kubernetes_deployment.metrics (0 add, 1 change, 0 destroy) with no sweep-unavailable warning, the failed discovery group costing nothing, matching stock’s own plan on the oracle cluster for the same tamper; apply changed 1, the argument reads back 10250, the APIService reported Available again and the aggregated API answers, the next plan empty. BREAK=1 tampers the Service’s declared label too and the single-object assertion correctly fails |
| Rename | pass | 23s | moved block: kubernetes_service_account.metrics -> .collector, with the three bindings’ subjects and the Deployment’s service_account_name following, no add and no destroy, one in-place change confined to the address annotation rewrite (0 add, 1 change, 0 destroy) - the marker rewritten in place, the shape every lane asserts for a rename; the live ServiceAccount in kube-system untouched and still labelled; stock’s plan for the same moved block on the oracle cluster is zero churn, since stock never writes this annotation. The moved-block half only. BREAK=1 renames the object’s own metadata.name and the marker-rewritten-in-place assertion correctly fails |
| Remove a block | pass | 33s | deleting kubernetes_api_service.metrics’s block proposed exactly one destroy (0 add, 0 change, 1 destroy) at the sweep’s orphan address kubernetes_api_service_v1.orphan_v1beta1_metrics_k8s_io - the cluster-scoped aggregated APIService found by its label alone under apiregistration.k8s.io, a group client-go’s scheme does not register - applied cleanly, the object gone (kubectl get apiservice: NotFound), the Service behind it untouched, the next plan empty; stock’s plan for the same removal on the oracle cluster is also exactly one destroy; the Deployment’s ReplicaSet and Pod, which carry no estate label, were never proposed. BREAK_REMOVE=1 keeps the block and no destroy is proposed |
| Change count | pass | 1m9s | a two-instance count ConfigMap added in kube-system beside the published file (its own shape has no count block): scaling 2 to 1 destroyed exactly ms-shard-1, planned at the sweep’s orphan address kubernetes_config_map_v1.orphan_kube-system_ms-shard-1 since the label carries no index (ms-shard-0 untouched, both read with kubectl); back to 2 created exactly kubernetes_config_map_v1.shard[1] under the same name; the next plan is empty; stock’s plans for the same two changes on the oracle cluster have the identical shape. BREAK_COUNT=1 asserts the lower index was destroyed and correctly fails |
| Replace with create_before_destroy | pass | 56s | a create_before_destroy ConfigMap whose content-hashed name changes (cfg-a -> cfg-b) plans as stock’s replace: ‘kubernetes_config_map.hashed must be replaced’, +/- create replacement and then destroy, 1 add and 1 destroy, with no orphan destroy beside it, because cfg-a carries the block’s address annotation and the sweep binds it (#1640). At -parallelism=1 the apply log shows cfg-b’s creation complete (line 64) before cfg-a’s deposed destroy starts (line 65), the same order stock’s apply shows on the oracle cluster; kubectl confirms cfg-b alone remains, carrying the annotation, and the next plan is empty (#1541). The block is removed from both roots afterwards. BREAK_REPLACE=1 recreates cfg-a carrying the block’s annotation and the next plan correctly proposes destroying it |
| Crash mid-apply | pass | 3m35s | a Secret/ConfigMap pair added in kube-system beside the published file: the apply creating both was interrupted by a real SIGTERM (exit 1), delivered by the engine itself inside the -parallelism=1 graph walker the instant kubernetes_secret_v1.crash_first’s create committed; crash_second reads crash_first’s name, so the walker cannot have reached it - kubectl confirms ms-crash-first exists carrying tofu-estate=corpus-k8s-metrics-server and ms-crash-second does not. The next plan proposed exactly the remainder (Plan: 1 to add, 0 to change, 0 to destroy) and nothing for crash_first, which it bound by its label and its namespace and name, matching stock’s own plan from the same position on the oracle cluster; the recovery apply added exactly one object and the plan after it is empty. The record store’s contribution is read, not counted: the interrupted apply wrote exactly one record (files 11 -> 12) carrying residue wait_for_service_account_token, and taking it out and replanning turns Plan: 1 to add, 0 to change, 0 to destroy into Plan: 1 to add, 1 to change, 0 to destroy, proposing wait_for_service_account_token back; putting it back restores the exact-remainder plan. BREAK_CRASH=1 and BREAK_CRASH_UNBOUND=1 correctly fail The create_before_destroy rename window is interrupted too (#1768): a kubernetes_config_map renamed under create_before_destroy was killed by the engine’s own hook the instant the new object’s create committed, at -parallelism=1, leaving both objects carrying the block’s address annotation and the record holding the old one as the address’s deposed object. The verdict is the end state stock’s replace leaves, not the plan’s wording: after one more apply exactly the new object remains, the old one is gone, the record’s deposed entry is cleared and the replan is empty. Both of the rerun’s paths reached it - orphan leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘orphan_kube-system_crash-rename-a’; deposed leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘(deposed object’; the same configuration through the orphan destroy, and the name read from another block’s attribute (#1539’s shape) through the deposed record (#1683), where stock’s plan reads the same deposed-object destroy. |
| Teardown | pass | 14s | apply -destroy removed exactly the 12 remaining objects in one apply (the four cluster-scoped RBAC objects among them), no object of any of the estate’s kinds carrying tofu-estate=corpus-k8s-metrics-server afterwards (kubectl, every namespace); kube-system, which the estate writes into and never declared, is still Active, and every one of its 236 objects of the compared kinds (ConfigMaps, ServiceAccounts, Services, Deployments, DaemonSets, Roles, RoleBindings, plus every ClusterRole, ClusterRoleBinding and APIService) read before cold_deploy is there and nothing else is; stock’s destroy on the oracle cluster removed exactly the 12 its state held and passes the same comparison. BREAK_TEARDOWN=1 deletes one of kube-system’s own ConfigMaps and the comparison correctly fails |
| Plan, review, apply | pass | 33s | plan -out wrote one update (the cluster-scoped ClusterRole system:metrics-server gains reviewed=yes); the world then moved out of band (a stray label on the metrics-server Service in kube-system, kubectl, never choudoufu) and apply of the saved plan refused with “The approved plan no longer matches the live system” at exit 3, nothing applied; with the label removed the identical file applied, 0 added, 1 changed, 0 destroyed, and reviewed=yes reads back; stock’s own planfile applied on the oracle cluster in the unchanged case. BREAK_APPROVAL=1 expects success after the move and correctly fails |
| Greenfield apply | pass | 1m11s | the cut-out root applied fresh with a live block and no terraform.tfstate: 9 objects, every one labelled tofu-estate=corpus-k8s-metrics-server (kubectl), the APIService reporting Available=True again; the record store held 6 record envelope(s), counted by their own address field; replanned empty with and without the cache. Deleting the whole record store and the cache and replanning proposed Plan: 0 to add, 2 to change, 0 to destroy: nothing created, destroyed or swept, every object still bound by its label and its name, and 2 in-place update(s) putting back the residue the store held (wait_for_load_balancer wait_for_rollout); one apply reconverged and the plan after it is empty. The cluster’s inventory (both ClusterRoles’ rules and aggregation labels, the three bindings’ roles and subjects, the ServiceAccount, the APIService’s group, version, priorities, TLS setting and backing Service, the Deployment’s containers, args, ports, service account and priority class, the Service’s ports and selector) matches stock’s cold deploy on the same cluster object by object, labels never compared. BREAK=1 drops the APIService from the expected inventory and the match correctly fails |
| Strict profile (not a headline stage) | pass | 13s | every strict toggle on (secrets = refuse, no_source_create = refuse, marker_repair = never with a markers “record” selection naming kubernetes_config_map_v1) against a scratch estate carrying random_password.db: exactly one refusal, Logical resource is not admitted under strict { secrets = “refuse” }; the other two toggles are on and silent. BREAK_STRICT=1 turns secrets back to “store” and the refusal disappears |
| Plan with no local state (not a headline stage) | not run |
Last run at commit 4bfb459f95 on 2026-10-05T02:58:57Z, exit code 0, against substrate image kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5. Total run time 11m15.2s.
Oracle: stock terraform 1.15.8, stock tofu 1.12.5. Stale: the current pin is terraform 1.16.1, tofu 1.13.0.
Engine: OpenTofu base 1.13.0 (matches the current base).
Reproduce it
go run ./tools/gauntlet run corpus-k8s-metrics-server
Needs Docker (the emulator is pulled at the pinned digest), the AWS CLI, and a
stock terraform or tofu binary on PATH for the cold deploy. The script is
live/e2e/corpus-k8s-metrics-server/run.sh; BREAK=1 corrupts its assertions to show they are load-bearing.