Evidence · How close AWS is
corpus-govuk-cluster-access
alphagov/govuk-infrastructure terraform/deployments/cluster-access (commit c02504fa4abb439234669e9ea9b662bdb35209e9, MIT), GOV.UK’s platform team’s own cluster access root: three Namespaces, then module “access-entry” called eight times, each a ClusterRole with dynamic rule blocks and its ClusterRoleBinding, plus a Role and RoleBinding per namespace through for_each for four of them - 31 objects, the kubernetes lane’s first typed RBAC beyond one binding and its first module fan-out; the EKS access entries are dropped on kind
Source: https://github.com/alphagov/govuk-infrastructure.git at c02504fa4abb439234669e9ea9b662bdb35209e9.
Set: growing. Lane: kubernetes.
Clear. Every headline stage passes.
| Stage | Verdict | Duration | Detail |
|---|---|---|---|
| Cold deploy | pass | 1m3s | 31 objects (3 Namespaces, then module “access-entry” called eight times: 8 ClusterRoles with dynamic rule blocks, 8 ClusterRoleBindings, 6 Roles and 6 RoleBindings through for_each over namespaces) from plain terraform on the published root plus its deltas (Terraform Cloud block, EKS access entries and the AWS and tfe data sources dropped, provider pointed at the run’s cluster, GOV.UK’s own integration var files), against kind v1.37.0; a real terraform.tfstate with 31 instances, 28 of them under module calls, zero tofu-estate labels, 31 objects inventoried with kubectl; the identical root cold-deployed by stock on a second cluster as every later stage’s oracle |
| Migrate | pass | 2s | 31 of 31 stamped, 0 skipped, from the stock state file, 28 of them addressed through the eight module instances and 12 through for_each keys inside them; every object carries tofu-estate=corpus-govuk-cluster-access, read back with kubectl across the estate’s five kinds; every write was judged by the estate boundary (live/kubernetes/estate-boundary.yaml), installed on the cluster first and observed by the API server |
| Replan from nothing | pass | 11s | the plan with no state file is empty; all 31 identities compared by value with kubectl - each object found by NAMESPACE/NAME (NAME for the 3 Namespaces, 8 ClusterRoles and 8 ClusterRoleBindings), carrying tofu-estate=corpus-govuk-cluster-access and the escaped instance address in choudoufu.intentius.io/tofu-address, module. |
| No-op apply | pass | 12s | no-op apply (0 added, 0 changed, 0 destroyed); objects carrying tofu-estate=corpus-govuk-cluster-access unchanged at 31 across five kinds, counted with kubectl |
| Drift and reconverge | pass | 34s | the ithctester ClusterRole’s dynamic rules replaced out of band with one rule (kubectl patch); choudoufu proposed exactly module.ithctester.kubernetes_cluster_role_v1.cluster_role (0 add, 1 change, 0 destroy), matching stock’s own plan on the oracle cluster for the same tamper; apply changed 1, the 6 declared rules (as stock rendered them on B) read back, the next plan is empty. The write was judged by the estate boundary first: the identical patch by system:serviceaccount:default:gauntlet-rbac-editor, holding RBAC patch and escalate on clusterroles but not the estate, was refused by ValidatingAdmissionPolicy ‘choudoufu-estate-boundary’ (“is not bound to that estate”) with the rules unchanged, and admitted once live/kubernetes/estate-grant.yaml granted it corpus-govuk-cluster-access - that admitted write is the drift. BREAK=1 tampers the readonly ClusterRole too and the single-object assertion correctly fails |
| Rename | pass | 34s | a module call renamed through a moved block: module.readonly -> module.viewer moves 6 instances at once - its ClusterRole, ClusterRoleBinding, and the Role and RoleBinding its for_each keys put in apps and licensify - with no add and no destroy, 6 in-place changes confined to the address annotation rewrite (0 add, 6 change, 0 destroy, every one module.readonly.* -> module.viewer.*, the for_each keys carried across), the marker rewritten in place; all 31 identities then compared by value with kubectl at their new addresses, each object’s metadata.name untouched, and the next plan empty; stock’s plan for the same moved block on the oracle cluster is zero churn, since stock never writes this annotation. The moved-block half only: live-mv also has a Kubernetes leg since #1639, not exercised by this stage. BREAK=1 changes the module call’s name argument - the objects’ own names - and the zero-churn assertion correctly fails |
| Remove a block | pass | 35s | deleting the licensinguser module block - one of eight instances of module “access-entry” - proposed exactly four destroys (0 add, 0 change, 4 destroy), one of each RBAC kind and all of them that instance’s: kubernetes_cluster_role_binding_v1.orphan_licensinguser-binding kubernetes_cluster_role_v1.orphan_licensinguser kubernetes_role_binding_v1.orphan_licensify_licensinguser-binding kubernetes_role_v1.orphan_licensify_licensinguser; applied cleanly, all four gone (kubectl: the ClusterRole, the ClusterRoleBinding, and the Role and RoleBinding in licensify NotFound), while the 27 identities that remain - developer’s and viewer’s Role and RoleBinding in the same licensify namespace among them - still match by value, and the next plan is empty; stock’s plan for the same removal on the oracle cluster is also exactly four destroys. BREAK_REMOVE=1 keeps the block and no destroy is proposed |
| Change count | pass | 1m8s | a two-instance count ConfigMap added beside the published root in its apps namespace (the estate’s own shape has no count block): scaling 2 to 1 destroyed exactly shard-1, planned at the sweep’s orphan address kubernetes_config_map_v1.orphan_apps_shard-1 since the label carries no index (shard-0 untouched, both read with kubectl); back to 2 created exactly kubernetes_config_map_v1.shard[1] under the same name; the next plan is empty; stock’s plans for the same two changes on the oracle cluster have the identical shape. BREAK_COUNT=1 asserts the lower index was destroyed and correctly fails |
| Replace with create_before_destroy | pass | 58s | a create_before_destroy ConfigMap whose content-hashed name changes (cfg-a -> cfg-b) plans as stock’s replace: ‘kubernetes_config_map.hashed must be replaced’, +/- create replacement and then destroy, 1 add and 1 destroy, with no orphan destroy beside it, because cfg-a carries the block’s address annotation and the sweep binds it (#1640). At -parallelism=1 the apply log shows cfg-b’s creation complete (line 64) before cfg-a’s deposed destroy starts (line 65), the same order stock’s apply shows on the oracle cluster; kubectl confirms cfg-b alone remains, carrying the annotation, and the next plan is empty (#1541). The block is removed from both roots afterwards. BREAK_REPLACE=1 recreates cfg-a carrying the block’s annotation and the next plan correctly proposes destroying it |
| Crash mid-apply | pass | 3m36s | a Secret and a ConfigMap added beside the published root in its apps namespace (the estate’s own shape has no two objects with an edge between them to be caught halfway through): the apply creating both was interrupted by a real SIGTERM (exit 1), delivered by the engine itself inside the -parallelism=1 graph walker the instant kubernetes_secret_v1.crash_first’s create committed (internal/command/apply_e2etesting_crash.go); crash_second reads crash_first’s name, so the walker cannot have reached it - kubectl confirms crash-first exists carrying tofu-estate=corpus-govuk-cluster-access and crash-second does not. The next plan proposed exactly the remainder (Plan: 1 to add, 0 to change, 0 to destroy, kubernetes_config_map_v1.crash_second created) and nothing at all for crash-first, which it bound by its label and its namespace and name - not a second create the API server would refuse, not an orphan sweep - matching stock’s own plan from the same position on the oracle cluster; the recovery apply added exactly one object, both read back with kubectl, and the plan after it is empty. The record store’s contribution is read, not counted: the interrupted apply wrote exactly one record (files 7 -> 8) for kubernetes_secret_v1.crash_first carrying residue wait_for_service_account_token, and taking that one file out of the store and replanning from the identical position turns the recovery plan from Plan: 1 to add, 0 to change, 0 to destroy into Plan: 1 to add, 1 to change, 0 to destroy, proposing wait_for_service_account_token back on the object the crash left behind; putting it back restores the exact-remainder plan. BREAK_CRASH=1 asserts nothing is proposed and correctly fails; BREAK_CRASH_UNBOUND=1 strips the label off crash-first and the same recovery check correctly fails The create_before_destroy rename window is interrupted too (#1768): a kubernetes_config_map renamed under create_before_destroy was killed by the engine’s own hook the instant the new object’s create committed, at -parallelism=1, leaving both objects carrying the block’s address annotation and the record holding the old one as the address’s deposed object. The verdict is the end state stock’s replace leaves, not the plan’s wording: after one more apply exactly the new object remains, the old one is gone, the record’s deposed entry is cleared and the replan is empty. Both of the rerun’s paths reached it - orphan leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘orphan_apps_crash-rename-a’; deposed leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘(deposed object’; the same configuration through the orphan destroy, and the name read from another block’s attribute (#1539’s shape) through the deposed record (#1683), where stock’s plan reads the same deposed-object destroy. |
| Teardown | pass | 24s | apply -destroy removed exactly the 31 remaining objects in one apply - the module instances’ cluster-scoped and namespaced RBAC objects, the script’s own ConfigMaps and Secret, and the namespaces - in an order the API server accepted, the three namespaces gone and no object of any of the estate’s five kinds carrying tofu-estate=corpus-govuk-cluster-access (kubectl, every namespace); stock’s destroy of the same estate on the oracle cluster also removed exactly the 31 its state held |
| Plan, review, apply | pass | 34s | plan -out wrote one update (the apps namespace gains reviewed=yes); the world then moved out of band (a stray label on the fulladmin-binding ClusterRoleBinding, kubectl, never choudoufu) and apply of the saved plan refused with “The approved plan no longer matches the live system” at exit 3, nothing applied; with the label removed the identical file applied, 0 added, 1 changed, 0 destroyed, and reviewed=yes reads back; stock’s own planfile applied on the oracle cluster in the unchanged case. BREAK_APPROVAL=1 expects success after the move and correctly fails |
| Greenfield apply | pass | 1m8s | the published root applied fresh with a live block and no terraform.tfstate: 31 objects, every one labelled tofu-estate=corpus-govuk-cluster-access and every identity, module instance and for_each key included, matching by value (kubectl, five kinds); the record store held 3 record envelope(s), counted by their own address field (#1291); replanned empty with and without the cache. Deleting the whole record store and the cache and replanning proposed No changes. Your infrastructure matches the configuration: nothing created, destroyed or swept, every object still bound by its label and its namespace and name, and 0 in-place update(s) putting back the residue the store held (none); one apply reconverged and the plan after it is empty (#1188, #1235). The cluster’s inventory (each Namespace’s labels and annotations, every ClusterRole’s and Role’s rules as the dynamic blocks rendered them, every binding’s role reference and subjects) matches stock’s cold deploy on the same cluster object by object, the marker label and this fork’s annotations normalised out. BREAK=1 drops the ithctester ClusterRole from the expected inventory and the match correctly fails |
| Strict profile (not a headline stage) | pass | 18s | every strict toggle on (secrets = refuse, no_source_create = refuse, marker_repair = never with a markers “record” selection naming kubernetes_config_map_v1) against a scratch estate carrying random_password.db: exactly one refusal, Logical resource is not admitted under strict { secrets = “refuse” }; the other two toggles are on and silent. BREAK_STRICT=1 turns secrets back to “store” and the refusal disappears |
| Plan with no local state (not a headline stage) | not run |
Last run at commit 4bfb459f95 on 2026-10-05T03:04:07Z, exit code 0, against substrate image kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5. Total run time 11m18.6s.
Oracle: stock terraform 1.15.8, stock tofu 1.12.5. Stale: the current pin is terraform 1.16.1, tofu 1.13.0.
Engine: OpenTofu base 1.13.0 (matches the current base).
Reproduce it
go run ./tools/gauntlet run corpus-govuk-cluster-access
Needs Docker (the emulator is pulled at the pinned digest), the AWS CLI, and a
stock terraform or tofu binary on PATH for the cold deploy. The script is
live/e2e/corpus-govuk-cluster-access/run.sh; BREAK=1 corrupts its assertions to show they are load-bearing.