Evidence · How close AWS is
corpus-govuk-cluster-services
alphagov/govuk-infrastructure terraform/deployments/cluster-services (commit c02504fa4abb439234669e9ea9b662bdb35209e9), GOV.UK’s own in-cluster platform services root for its EKS clusters, split into a live slice (namespaces, dex client Secrets, kubernetes_labels and kubernetes_annotations with force) and a recorded stock half of Helm, kubectl and AWS blocks; the first published root driving the field-granular types
Source: https://github.com/alphagov/govuk-infrastructure.git at c02504fa4abb439234669e9ea9b662bdb35209e9.
Set: growing. Lane: kubernetes.
Clear. Every headline stage passes.
| Stage | Verdict | Duration | Detail |
|---|---|---|---|
| Cold deploy | pass | 1m6s | 38 instances from plain terraform on the published root’s live slice (2 namespaces, one in the local gatekeeper module; 15 dex-client Secrets over 3 namespaces; kubernetes_labels.argocd_secret x3 on three of those Secrets; kubernetes_annotations.rm_default_storageclass with force = true; 17 random_* values), split mechanically by split.py from the stock half (data aws_caller_identity x2, data aws_eks_cluster_auth x1, data aws_iam_policy_document x5, data aws_region x1, data tfe_outputs x2, module constraint_templates x1, module constraints x1, module renovate_irsa x1, module secure_s3_bucket_tempo x1, module tag_image_iam_role x1, module tempo_iam_role x1, provider aws x1, provider helm x1, provider kubectl x1, provider kubernetes x1, resource aws_iam_policy x3, resource aws_iam_role_policy_attachment x1, resource aws_s3_bucket x1, resource aws_s3_bucket_cors_configuration x1, resource aws_s3_bucket_lifecycle_configuration x1, resource aws_s3_bucket_logging x1, resource aws_s3_bucket_object_lock_configuration x1, resource aws_s3_bucket_ownership_controls x1, resource aws_s3_bucket_policy x1, resource aws_s3_bucket_public_access_block x1, resource aws_s3_bucket_replication_configuration x1, resource aws_s3_bucket_server_side_encryption_configuration x1, resource aws_s3_bucket_versioning x1, resource helm_release x14, resource kubectl_manifest x3, resource terraform_data x1, resource time_sleep x2, terraform x1 - recorded beside the root at terraform/deployments/cluster-services-platform, not applied: its charts need IRSA, an AWS load balancer controller and public DNS), against kind v1.37.0; a real terraform.tfstate with 38 instances and zero tofu-estate labels. Stock’s force = true took storageclass.kubernetes.io/is-default-class on kind’s standard StorageClass from the platform’s own manager [kubectl-client-side-apply] (true -> false, now [Terraform]); app.kubernetes.io/part-of=argocd on the three dex-client-argocd Secrets owned by Terraform; the identical root cold-deployed by stock on a second cluster as every later stage’s oracle |
| Migrate | pass | 2s | 21 resource(s) newly stamped, 0 already stamped, 17 newly recorded, 0 re-recorded for sensitivity only, 0 already recorded, 0 failed, 0 skipped: 17 namespaces and Secrets stamped tofu-estate=corpus-govuk-cluster-services (read back with kubectl), the 4 field-granular instances handed over by field manager rather than labelled (1 kubernetes_annotations, 3 kubernetes_labels, each reported as Handed … from “Terraform”), the 17 random_* values seeded into the record store; field manager choudoufu:corpus-govuk-cluster-services owns the StorageClass’s storageclass.kubernetes.io/is-default-class annotation (false) and app.kubernetes.io/part-of on all three dex-client-argocd Secrets, read off managedFields with no Terraform entry left; the annotations instance’s record carries handover_from = Terraform, the migration evidence #1869’s later plans read |
| Replan from nothing | pass | 12s | the plan with no state file is empty; both namespaces, six of the fifteen Secrets (argocd and grafana in each namespace) and the patched StorageClass confirmed present by name with kubectl; field manager choudoufu:corpus-govuk-cluster-services owns the StorageClass’s storageclass.kubernetes.io/is-default-class annotation (false) and app.kubernetes.io/part-of on all three dex-client-argocd Secrets, read off managedFields with no Terraform entry left |
| No-op apply | pass | 12s | no-op apply (0 added, 0 changed, 0 destroyed); objects carrying tofu-estate=corpus-govuk-cluster-services unchanged at 17 across namespaces, Secrets and ConfigMaps, counted with kubectl; field manager choudoufu:corpus-govuk-cluster-services owns the StorageClass’s storageclass.kubernetes.io/is-default-class annotation (false) and app.kubernetes.io/part-of on all three dex-client-argocd Secrets, read off managedFields with no Terraform entry left |
| Drift and reconverge | pass | 37s | the monitoring/dex-client-grafana Secret’s clientID tampered with kubectl patch; choudoufu proposed exactly kubernetes_secret_v1.dex_client[“monitoring-grafana”] (0 add, 1 change, 0 destroy), matching stock’s own plan on the oracle cluster for the same tamper; apply changed 1, the value reads back as the record store’s random_bytes value and the replan is empty; field manager choudoufu:corpus-govuk-cluster-services owns the StorageClass’s storageclass.kubernetes.io/is-default-class annotation (false) and app.kubernetes.io/part-of on all three dex-client-argocd Secrets, read off managedFields with no Terraform entry left, untouched by the Secret’s own whole-object write. BREAK=1 tampers dex-client-prometheus too and the single-object assertion correctly fails |
| Rename | pass | 25s | moved block: kubernetes_namespace_v1.monitoring -> .monitoring_namespace, with dex_client’s depends_on following, no add and no destroy, one in-place change confined to the address annotation rewrite (0 add, 1 change, 0 destroy) - the marker rewritten in place, the same shape the other kind estates assert; the namespace and the five Secrets in it untouched and still labelled; stock’s plan for the same moved block on the oracle cluster is zero churn, since stock never writes this annotation; field manager choudoufu:corpus-govuk-cluster-services owns the StorageClass’s storageclass.kubernetes.io/is-default-class annotation (false) and app.kubernetes.io/part-of on all three dex-client-argocd Secrets, read off managedFields with no Terraform entry left. BREAK=1 renames the namespace’s own metadata.name and the assertion correctly fails |
| Remove a block | pass | 36s | deleting kubernetes_annotations.rm_default_storageclass’s block proposed exactly one destroy (0 add, 0 change, 1 destroy) at the sweep’s field-manager orphan address kubernetes_annotations.orphan_storageclass_standard - an object the estate never owned, found by the fields choudoufu:corpus-govuk-cluster-services held on it, not by a label - applied cleanly; the standard StorageClass still exists, choudoufu:corpus-govuk-cluster-services no longer owns storageclass.kubernetes.io/is-default-class (owners now []), and the annotation reads ‘’, the same end state stock’s destroy of the same block left on the oracle cluster; the next plan is empty and the labelled count unchanged at 17; field manager choudoufu:corpus-govuk-cluster-services owns app.kubernetes.io/part-of on all three dex-client-argocd Secrets, read off managedFields with no Terraform entry left. BREAK_REMOVE=1 keeps the block and no destroy is proposed |
| Change count | pass | 1m12s | a two-instance count ConfigMap added beside the published root in the monitoring namespace (the live slice’s only count block is eph_account’s, 0 outside an ephemeral env): scaling 2 to 1 destroyed exactly shard-1, planned at kubernetes_config_map_v1.orphan_monitoring_shard-1 (shard-0 untouched, both read with kubectl); back to 2 created exactly kubernetes_config_map_v1.shard[1] under the same name; the next plan is empty; stock’s plans for the same two changes on the oracle cluster have the identical shape; field manager choudoufu:corpus-govuk-cluster-services owns app.kubernetes.io/part-of on all three dex-client-argocd Secrets, read off managedFields with no Terraform entry left. BREAK_COUNT=1 asserts the lower index was destroyed and correctly fails |
| Replace with create_before_destroy | pass | 1m1s | a create_before_destroy ConfigMap whose content-hashed name changes (cfg-a -> cfg-b) plans as stock’s replace: ‘kubernetes_config_map.hashed must be replaced’, +/- create replacement and then destroy, 1 add and 1 destroy, with no orphan destroy beside it, because cfg-a carries the block’s address annotation and the sweep binds it (#1640). At -parallelism=1 the apply log shows cfg-b’s creation complete (line 64) before cfg-a’s deposed destroy starts (line 65), the same order stock’s apply shows on the oracle cluster; kubectl confirms cfg-b alone remains, carrying the annotation, and the next plan is empty (#1541). The block is removed from both roots afterwards. BREAK_REPLACE=1 recreates cfg-a carrying the block’s annotation and the next plan correctly proposes destroying it |
| Crash mid-apply | pass | 3m47s | a Secret and a ConfigMap added beside the published root in the monitoring namespace, the second reading the first’s name: the apply creating both was interrupted by a real SIGTERM (exit 1), delivered by the engine itself inside the -parallelism=1 graph walker the instant kubernetes_secret_v1.crash_first’s create committed; kubectl confirms crash-first exists carrying tofu-estate=corpus-govuk-cluster-services and crash-second does not. The next plan proposed exactly the remainder (Plan: 1 to add, 0 to change, 0 to destroy, kubernetes_config_map_v1.crash_second created) and nothing for crash-first, which it bound by its label and its namespace and name, matching stock’s own plan from the same position on the oracle cluster; the recovery apply added exactly one object and the plan after it is empty. The record’s contribution is read, not counted: the interrupted apply wrote exactly one record (envelopes 41 -> 42) carrying residue wait_for_service_account_token, and taking it out of the store turns the recovery plan from Plan: 1 to add, 0 to change, 0 to destroy into Plan: 1 to add, 1 to change, 0 to destroy, proposing wait_for_service_account_token back; putting it back restores the exact remainder; field manager choudoufu:corpus-govuk-cluster-services owns app.kubernetes.io/part-of on all three dex-client-argocd Secrets, read off managedFields with no Terraform entry left. BREAK_CRASH=1 asserts nothing is proposed and correctly fails; BREAK_CRASH_UNBOUND=1 strips the label off crash-first and the same recovery check correctly fails The create_before_destroy rename window is interrupted too (#1768): a kubernetes_config_map renamed under create_before_destroy was killed by the engine’s own hook the instant the new object’s create committed, at -parallelism=1, leaving both objects carrying the block’s address annotation and the record holding the old one as the address’s deposed object. The verdict is the end state stock’s replace leaves, not the plan’s wording: after one more apply exactly the new object remains, the old one is gone, the record’s deposed entry is cleared and the replan is empty. Both of the rerun’s paths reached it - orphan leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘orphan_monitoring_crash-rename-a’; deposed leg: Plan: 0 to add, 0 to change, 1 to destroy, the destroy at ‘(deposed object’; the same configuration through the orphan destroy, and the name read from another block’s attribute (#1539’s shape) through the deposed record (#1683), where stock’s plan reads the same deposed-object destroy. |
| Teardown | pass | 25s | apply -destroy removed exactly the 41 remaining instances in one apply (21 labelled objects, the three kubernetes_labels writes and 17 random values), the monitoring and gatekeeper-system namespaces gone, no dex-client Secret left in the stand-in namespaces, no object of the estate’s kinds carrying tofu-estate=corpus-govuk-cluster-services (kubectl, every namespace), and kind’s standard StorageClass, never the estate’s, still there; stock’s destroy of the same estate on the oracle cluster also removed exactly the 41 its state held |
| Plan, review, apply | pass | 35s | plan -out wrote one update (the monitoring namespace gains reviewed=yes); the world then moved out of band (a stray label on the gatekeeper-system namespace, kubectl, never choudoufu) and apply of the saved plan refused with “The approved plan no longer matches the live system” at exit 3, nothing applied; with the label removed the identical file applied, 0 added, 1 changed, 0 destroyed, and reviewed=yes reads back; stock’s own planfile applied on the oracle cluster in the unchanged case; field manager choudoufu:corpus-govuk-cluster-services owns the StorageClass’s storageclass.kubernetes.io/is-default-class annotation (false) and app.kubernetes.io/part-of on all three dex-client-argocd Secrets, read off managedFields with no Terraform entry left. BREAK_APPROVAL=1 expects success after the move and correctly fails |
| Greenfield apply | pass | 37s | the split root applied fresh with a live block and no terraform.tfstate: 38 instances, the 17 namespaces and Secrets labelled tofu-estate=corpus-govuk-cluster-services (kubectl), the record store holding 37 record envelope(s); force = true took storageclass.kubernetes.io/is-default-class on kind’s standard StorageClass from the platform’s own manager [kubectl-annotate] with no refusal, since that manager is not an estate’s; field manager choudoufu:corpus-govuk-cluster-services owns the StorageClass’s storageclass.kubernetes.io/is-default-class annotation (false) and app.kubernetes.io/part-of on all three dex-client-argocd Secrets, read off managedFields with no Terraform entry left; replanned empty with and without the cache. The cluster’s inventory (both namespaces and the gatekeeper labels, every dex-client Secret’s keys and type in all three namespaces, app.kubernetes.io/part-of on the three argocd Secrets, the StorageClass’s default-class annotation) matches stock’s cold deploy on the same cluster object by object, the estate label never compared. BREAK=1 drops the StorageClass from the expected inventory and the match correctly fails |
| Strict profile (not a headline stage) | pass | 19s | every strict toggle on (secrets = refuse, no_source_create = refuse, marker_repair = never with a markers “record” selection naming kubernetes_config_map_v1) against a scratch estate carrying the root’s own random_password shape (dex_secret’s): exactly one refusal, Logical resource is not admitted under strict { secrets = “refuse” }; the other two toggles are on and silent. BREAK_STRICT=1 turns secrets back to “store” and the refusal disappears |
| Plan with no local state (not a headline stage) | not run |
Last run at commit 4bfb459f95 on 2026-10-05T02:32:28Z, exit code 0, against substrate image kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5. Total run time 11m7.5s.
Oracle: stock terraform 1.15.8, stock tofu 1.12.5. Stale: the current pin is terraform 1.16.1, tofu 1.13.0.
Engine: OpenTofu base 1.13.0 (matches the current base).
Reproduce it
go run ./tools/gauntlet run corpus-govuk-cluster-services
Needs Docker (the emulator is pulled at the pinned digest), the AWS CLI, and a
stock terraform or tofu binary on PATH for the cold deploy. The script is
live/e2e/corpus-govuk-cluster-services/run.sh; BREAK=1 corrupts its assertions to show they are load-bearing.