Evidence · How close AWS is
corpus-cloud-platform-components
ministryofjustice/cloud-platform-infrastructure terraform/aws-accounts/cloud-platform-aws/vpc/eks/core/components (commit 6e1eca7be0, MIT), the Ministry of Justice Cloud Platform’s own cluster-components root, top-level resources only: three StorageClasses on the unversioned kubernetes_storage_class, three PriorityClasses, two ClusterRoleBindings, a ServiceAccount in kube-system and kubernetes_manifest.vpa, five VerticalPodAutoscalers keyed by for_each over a static map (the manager workspace’s, selected in the configuration); plus the VPA project’s own two CRDs (converted by convert.sh) and the three namespaces the VPAs live in, pre-applied - 19 objects over seven kinds. Pruned: the ~15 ministryofjustice module calls (not in .corpus/_modules), the backend, remote states and aws data sources, and storage.tf’s kubectl_manifest gp2 default flip, which stock applies on the side from a root of its own
Source: https://github.com/ministryofjustice/cloud-platform-infrastructure.git at 6e1eca7be0d4224718a993e82508df8b7b19857d.
Set: growing. Lane: kubernetes.
Clear. Every headline stage passes.
| Stage | Verdict | Duration | Detail |
|---|---|---|---|
| Cold deploy | pass | 1m10s | 19 objects from plain terraform against kind v1.37.0: the pinned components root’s three StorageClasses (the unversioned kubernetes_storage_class), three PriorityClasses, two ClusterRoleBindings, a ServiceAccount in kube-system and kubernetes_manifest.vpa’s five VerticalPodAutoscalers keyed by for_each over a static map (each key read back from stock’s own state), plus the VPA project’s two CRDs and the three namespaces the VPAs live in; a real terraform.tfstate with 19 instances and zero tofu-estate labels read back with kubectl; the identical root cold-deployed by stock on a second cluster as every later stage’s oracle. Two applies, and the first is declared: declared pre-apply, #1173: 5 address(es) declared at pre_apply in live/gauntlet/estates.json, applied with -target on every side (estate,oracle) before the main apply - the full list is in that file and in the GAUNTLET pre_apply= line this run printed, which the runner checks address by address; forced by: manager-vpas.tf declares kubernetes_manifest objects of kind VerticalPodAutoscaler (autoscaling.k8s.io/v1), and kubernetes_manifest builds an object’s schema at PLAN time, so the root cannot be planned at all against a cluster that does not serve that kind - depends_on does not help and re-running fails identically (#1173). On cloud-platform’s own clusters the CRD and the three namespaces the VPAs live in come from module calls this crossing prunes, so the VPA project’s two CRDs (root/vpa-crd.tf) and those namespaces (root/vpa-namespaces.tf) go up first, with -target, on every side; the namespaces are in the list because kubernetes_manifest dry-runs a namespaced object at plan time and a dry run into a missing namespace is refused. The five VPAs and the nine typed objects are what the main apply creates. Control, run first on this same cluster: the un-targeted one-pass plan exits 1 before creating anything - “Error: API did not recognize GroupVersionKind from manifest (CRD may not be installed)” - so the pre-apply is load-bearing. storage.tf line 56’s kubectl_manifest (the gp2 default flip) was applied by stock on both clusters from a root of its own (delta 4) and reads is-default-class=false, outside the estate. BREAK_PREAPPLY=1 requires the one-pass plan to succeed and correctly fails |
| Migrate | pass | 6s | 19 of 19 stamped, 0 skipped, from the stock state file; every object carries tofu-estate=corpus-cloud-platform-components, counted back with kubectl across the estate’s seven kinds; the five for_each instances of kubernetes_manifest.vpa are each named in the stamped report by their static-map key (" |
| Replan from nothing | pass | 14s | the plan with no state file is empty; all 19 objects bind by namespace and name, and twelve identities compared by value with kubectl: a StorageClass of each provisioner (gp3, io1-expand), the global-default PriorityClass, a ClusterRoleBinding, the kube-system ServiceAccount (kube-system/concourse-build-environments), the VPA CRD, a namespace, and all five for_each instances of kubernetes_manifest.vpa, each at the NAMESPACE/NAME its static-map key spells and pointing at the target kind the key names |
| No-op apply | pass | 14s | no-op apply (0 added, 0 changed, 0 destroyed); objects carrying tofu-estate=corpus-cloud-platform-components unchanged at 19 across the estate’s seven kinds, counted with kubectl |
| Drift and reconverge | pass | 27s | one for_each instance deleted out of band with kubectl (the thanos-compactor VPA in monitoring); choudoufu proposed putting back exactly kubernetes_manifest.vpa[“monitoring/Deployment/thanos-compactor”] (1 add, 0 change, 0 destroy) - the instance found missing by its static-map key, its four siblings untouched - matching stock’s own plan on the oracle cluster for the same delete; apply created 1, the targetRef reads back as configured, and the recreated object carries the estate label again (19 labelled). A delete rather than a patch because kubernetes_manifest applies server-side and a kubectl patch would measure field-manager ownership, not drift. BREAK=1 deletes a second VPA and the single-object assertion correctly fails |
| Rename | pass | 28s | moved block over a cluster-scoped, unversioned type: kubernetes_cluster_role_binding.webops -> .webops_admin, no add and no destroy, one in-place change confined to the address annotation rewrite (0 add, 1 change, 0 destroy) - the marker rewritten in place, not literal zero churn; the live ClusterRoleBinding untouched and still labelled; stock’s plan for the same moved block on the oracle cluster is zero churn, since stock never writes this annotation. The moved-block half only: live-mv also has a Kubernetes leg since #1639, not exercised by this stage. BREAK=1 renames the object’s own metadata.name and the assertion correctly fails |
| Remove a block | pass | 40s | deleting kubernetes_priority_class.node_critical’s block proposed exactly one destroy (0 add, 0 change, 1 destroy) at the sweep’s orphan address kubernetes_priority_class.orphan_node-critical - a cluster-scoped object found by its label, which carries no address - applied cleanly, the object gone (kubectl get priorityclass: NotFound), the next plan empty, 18 objects still labelled; kind’s own system-cluster-critical and system-node-critical, which carry no estate label, were never proposed; stock’s plan for the same removal on the oracle cluster is also exactly one destroy. BREAK_REMOVE=1 keeps the block and no destroy is proposed |
| Change count | pass | 1m41s | a two-instance counted VerticalPodAutoscaler (kubernetes_manifest.vpa_shard, name vpa-shard-${count.index} inside the manifest object) added beside the published root, whose own for_each is a static map: the replan right after creating both is empty, scaling 2 to 1 destroyed exactly vpa-shard-1, planned at the sweep’s orphan address kubernetes_manifest.orphan_verticalpodautoscaler_monitoring_vpa-shard-1 since the label carries no index (vpa-shard-0 untouched, both read with kubectl); back to 2 created exactly kubernetes_manifest.vpa_shard[1] under the same name; the next plan is empty; stock’s plans for the same two changes on the oracle cluster have the identical shape. BREAK_COUNT=1 asserts the lower index was destroyed and correctly fails |
| Replace with create_before_destroy | pass | 2m5s | Immutable fields first: io1-expand’s parameters.iopsPerGB (26 -> 50) and cluster-critical’s value (999999000 -> 999998000), both ForceNew in the provider and refused as updates by the API server, planned as kubernetes_storage_class.io1 and kubernetes_priority_class.cluster_critical ‘must be replaced’, -/+ destroy and then create replacement, 2 add and 2 destroy with no orphan beside them - the same plan stock makes on the oracle cluster; applied 2 added, 2 destroyed, the new values and the estate label read back with kubectl, the replan empty, 18 objects still labelled. Then the shared create_before_destroy rename: a create_before_destroy ConfigMap whose content-hashed name changes (cfg-a -> cfg-b) plans as stock’s replace: ‘kubernetes_config_map.hashed must be replaced’, +/- create replacement and then destroy, 1 add and 1 destroy, with no orphan destroy beside it, because cfg-a carries the block’s address annotation and the sweep binds it (#1640). At -parallelism=1 the apply log shows cfg-b’s creation complete (line 64) before cfg-a’s deposed destroy starts (line 65), the same order stock’s apply shows on the oracle cluster; kubectl confirms cfg-b alone remains, carrying the annotation, and the next plan is empty (#1541). The block is removed from both roots afterwards. BREAK_REPLACE=1 recreates cfg-a carrying the block’s annotation and the next plan correctly proposes destroying it |
| Crash mid-apply | pass | 3m53s | a Secret and a ConfigMap added beside the published root in the concourse namespace, the second reading the first’s name: the apply creating both was interrupted by a real SIGTERM (exit 1) delivered by the engine inside the -parallelism=1 walker the instant kubernetes_secret_v1.crash_first’s create committed; kubectl confirms crash-first exists carrying tofu-estate=corpus-cloud-platform-components and crash-second does not; the interrupted apply wrote exactly one record (envelopes 26 -> 27) carrying residue wait_for_service_account_token. The next plan proposed exactly the remainder (Plan: 1 to add, 0 to change, 0 to destroy, crash_second created) and nothing for crash-first, bound by its label and its namespace and name, matching stock’s plan from the same position on the oracle cluster; the recovery apply added one and the replan is empty. BREAK_CRASH=1 and BREAK_CRASH_UNBOUND=1 correctly fail. The create_before_destroy rename window is interrupted too (#1768): a kubernetes_config_map renamed under create_before_destroy was killed by the engine’s own hook the instant the new object’s create committed, at -parallelism=1, leaving both objects carrying the block’s address annotation and the record holding the old one as the address’s deposed object. The verdict is the end state stock’s replace leaves, not the plan’s wording: after one more apply exactly the new object remains, the old one is gone, the record’s deposed entry is cleared and the replan is empty. Both of the rerun’s paths reached it - orphan leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘orphan_concourse_crash-rename-a’; deposed leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘(deposed object’; the same configuration through the orphan destroy, and the name read from another block’s attribute (#1539’s shape) through the deposed record (#1683), where stock’s plan reads the same deposed-object destroy. |
| Teardown | pass | 26s | apply -destroy removed exactly the 20 remaining objects in one apply, in an order the API server accepted - the three namespaces gone, both VPA CRDs gone, and no object of any of the estate’s kinds carrying tofu-estate=corpus-cloud-platform-components (kubectl, every namespace); the stock-side gp2 StorageClass (delta 4, never the estate’s) is still there, not default and unlabelled; stock’s destroy of the same estate on the oracle cluster also removed exactly the 20 its state held |
| Plan, review, apply | pass | 39s | plan -out wrote one in-place update (the global-default PriorityClass’s description, a mutable field); the world then moved out of band (a stray label on the gp3 StorageClass, kubectl, never choudoufu) and apply of the saved plan refused with “The approved plan no longer matches the live system” at exit 3, nothing applied; with the label removed the identical file applied, 0 added, 1 changed, 0 destroyed, and the new description reads back; stock’s own planfile applied on the oracle cluster. BREAK_APPROVAL=1 expects success after the move and correctly fails |
| Greenfield apply | pass | 1m54s | the root applied fresh with a live block and no terraform.tfstate, in the same two applies the cold deploy took. Before the pre-apply, against the empty cluster, the un-targeted plan exits 1 with exactly one refusal, “Kubernetes kind not served by the cluster”, naming kubernetes_manifest.vpa, VerticalPodAutoscaler and autoscaling.k8s.io/v1 - one for the block, not one per for_each instance - and nothing else; choudoufu then performs the declared pre-apply itself (5 addresses, read from live/gauntlet/estates.json), and once the kind is served the same plan is clean, 14 to add. 19 objects, every one labelled tofu-estate=corpus-cloud-platform-components (kubectl, seven kinds); the record store held 19 record envelope(s); replanned empty with and without the cache. With the whole record store and the cache deleted the plan proposed No changes. Your infrastructure matches the configuration: nothing created, destroyed or replaced, 0 in-place update(s) putting back the residue the store held (none); one apply reconverged and the plan after it is empty (#1235). The inventory - every StorageClass’s provisioner, parameters, reclaim policy, expansion and default-class annotation, every PriorityClass’s value, default flag and description, both ClusterRoleBindings’ role and subjects, the ServiceAccount, the three namespaces, both CRDs’ group, kind, scope and versions, and all five VPAs’ targetRef and updatePolicy - matches stock’s cold deploy on the same cluster object by object, labels never compared. BREAK=1 drops io1-expand from the expected inventory and the match correctly fails |
| Strict profile (not a headline stage) | pass | 19s | every strict toggle on (secrets = refuse, no_source_create = refuse, marker_repair = never with a markers “record” selection naming kubernetes_priority_class) against a scratch estate carrying random_password.db: exactly one refusal, Logical resource is not admitted under strict { secrets = “refuse” }; the other two toggles are on and silent. BREAK_STRICT=1 turns secrets back to “store” and the refusal disappears |
| Plan with no local state (not a headline stage) | not run |
Last run at commit 4bfb459f95 on 2026-10-05T03:06:49Z, exit code 0, against substrate image kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5. Total run time 14m16.7s.
Oracle: stock terraform 1.15.8, stock tofu 1.12.5. Stale: the current pin is terraform 1.16.1, tofu 1.13.0.
Engine: OpenTofu base 1.13.0 (matches the current base).
Reproduce it
go run ./tools/gauntlet run corpus-cloud-platform-components
Needs Docker (the emulator is pulled at the pinned digest), the AWS CLI, and a
stock terraform or tofu binary on PATH for the cold deploy. The script is
live/e2e/corpus-cloud-platform-components/run.sh; BREAK=1 corrupts its assertions to show they are load-bearing.