Terraform workflows break down at scale because every team builds its own pipeline, and nothing forces those pipelines to agree.
The pattern looks the same everywhere. A repository per team, a slightly different GitHub Actions workflow in each one, more state backends than anybody can account for, and a terraform apply that somebody ran from a laptop during an incident. Some teams run a plan step but no policy check. One team has long-lived AWS access keys in repository secrets that nobody has rotated since the engineer who created them left.
Now try to answer a few questions about that setup. Who changed the shared VPC? Which accounts still allow publicly accessible RDS instances? Does the running infrastructure still match what is in main?
If answering requires opening a terminal, the problem is not Terraform. The problem is that four separate control planes are missing, and each one needs a different fix.
Centralizing Terraform means putting state and credentials, execution, policy, and visibility under shared ownership. Most teams centralize execution first because it is the visible pain, then wonder why governance still feels loose. Work from the bottom up instead.
terraform plan -detailed-exitcode. Buy it as drift detection with automatic reconciliation.Start by splitting state per team and per environment, then removing every static credential.
State boundaries are blast radius boundaries. One monolithic state file means a typo in a staging module locks the production pipeline, and every plan refreshes 900 resources it has no reason to touch.
If you still run a DynamoDB table for state locking, delete it. Terraform 1.10 added native S3 locking, and 1.11 promoted it to generally available while deprecating the DynamoDB argument.
terraform {
backend "s3" {
bucket = "acme-tfstate-prod"
key = "platform/networking/terraform.tfstate"
region = "us-east-1"
encrypt = true
use_lockfile = true
}
}
Next, replace long-lived access keys with OIDC federation. Each workflow run then receives a short-lived role session scoped to the repository and trigger that requested it.
permissions:
id-token: write
contents: read
jobs:
plan:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: arn:aws:iam::111122223333:role/terraform-plan
aws-region: us-east-1
The IAM trust policy is where governance actually lives. Scope the subject claim to a specific repository and trigger, because a wildcard such as repo:acme/* is functionally the same as sharing the credentials with every repository you own.
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::111122223333:oidc-provider/token.actions.githubusercontent.com"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com"
},
"StringLike": {
"token.actions.githubusercontent.com:sub": "repo:acme/platform-networking:pull_request"
}
}
}
Use two roles, not one. The plan role above is the one a pull request can assume, and it is read-only in the account apart from s3:GetObject, s3:PutObject, and s3:DeleteObject on the lock file, which use_lockfile needs in order to take a lock. The apply role gets write permissions and a subject claim of repo:acme/platform-networking:ref:refs/heads/main.
Check the subject format before you write either policy. Repositories created, renamed, or transferred after 15 July 2026 carry immutable numeric IDs in the claim, as in repo:acme@123456/platform-networking@789012:pull_request, and a condition matching only the older shape fails closed with no useful error.
Three options, in increasing order of cost and capability.
Publish one workflow_call workflow in a central repository and have every team reference it by tag. This is the cheapest real fix and it eliminates pipeline drift immediately. It does not close the gap between plan and apply. The plan a reviewer approved ran against a state that a concurrent merge has since changed, and nothing prevents two applies from racing. Secrets and role ARNs also stay scattered across repositories.
Atlantis is an open source, pull request driven Terraform runner, and it is the standard answer for teams of four to ten. The governance lives in the server-side configuration, not the repository file, because an atlantis.yaml that any team can edit enforces nothing.
# repos.yaml, on the Atlantis server
repos:
- id: /github.com/acme/.*/
apply_requirements: [approved, mergeable, undiverged]
allowed_overrides: [workflow]
allow_custom_workflows: false
workflow: default
workflows:
default:
plan:
steps:
- init
- plan
- run: terraform show -json $PLANFILE > $SHOWFILE
- run: conftest test --policy /policies $SHOWFILE
undiverged is the requirement most teams miss. It rejects an apply when the branch is behind its base, which is exactly the race condition that reusable CI workflows leave open. It only takes effect under the merge checkout strategy, so start the server with --checkout-strategy=merge or the requirement silently does nothing.
Atlantis does not provide role-based access control beyond what your version control system offers, dependencies across repositories or outputs passed between stacks, or anything on a schedule. depends_on orders projects within a single pull request and stops there. It also means operating a server that holds apply credentials for every account you manage.
HCP Terraform and Spacelift both run Terraform on their own workers rather than in your CI, and that architectural choice is what buys you the capabilities the first two options cannot reach. You get runs on a schedule, dependency graphs that pass outputs between stacks, policy evaluated server side where nobody can skip it, and access control that is independent of version control permissions. The tradeoff is that a hosted worker needs credentials for your accounts. Both support self-hosted workers running in your own network if that tradeoff is unacceptable, at the cost of operating something again.
Enforce governance against the plan, not the HCL, because HCL does not tell you what is about to be destroyed.
Static scanners still earn their place for the cheap mistakes. Run Checkov or Trivy in pre-commit so engineers get feedback before CI does. For everything else, convert the plan to JSON and evaluate it with Open Policy Agent.
terraform plan -out=tfplan.binary
terraform show -json tfplan.binary > plan.json
conftest test --policy ./policies plan.json
package main
import rego.v1
mutating := {"create", "update"}
deny contains msg if {
rc := input.resource_changes[_]
rc.type == "aws_db_instance"
rc.change.actions[_] in mutating
rc.change.after.publicly_accessible
msg := sprintf("%s sets publicly_accessible = true", [rc.address])
}
destroys contains rc.address if {
rc := input.resource_changes[_]
"delete" in rc.change.actions
}
warn contains msg if {
count(destroys) > 3
msg := sprintf("plan destroys %d resources, needs a second reviewer", [count(destroys)])
}
The reason to write governance in Rego rather than bash is portability. The rule logic survives the move from Conftest on a laptop to a server-side policy in HCP Terraform or Spacelift, even though the input path does not. The plan sits at input.resource_changes under Conftest, input.plan in HCP Terraform, and input.terraform in Spacelift, so you write the rule once and change one line per destination. Sentinel is the alternative if you have standardized on HCP Terraform, with the tradeoff that Sentinel runs nowhere outside HashiCorp's own products.
Detect drift on a schedule, because CI cannot. This is a structural limit rather than a missing feature. CI runs on merge, and drift happens when somebody edits a security group in the console to unblock a customer.
Building detection yourself is straightforward for a handful of stacks. Exit code 2 means the plan found changes, and you have to capture that code yourself because a failed step exposes no exit status to later steps.
# .github/workflows/drift.yml
on:
schedule:
- cron: "0 6 * * *"
jobs:
detect:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: terraform init -input=false
- id: plan
run: |
set +e
terraform plan -detailed-exitcode -lock=false -no-color > plan.txt
echo "exitcode=$?" >> "$GITHUB_OUTPUT"
- if: steps.plan.outputs.exitcode == 2
env:
GH_TOKEN: ${{ github.token }}
run: gh issue create --title "Drift in $STACK" --body-file plan.txt
Multiply that across 30 stacks and you are maintaining a fleet of cron jobs that nobody owns. The scheduling, fan-out, and reconciliation are the part worth paying for.
Know which half you are buying. HCP Terraform covers detection through health assessments, which are scheduled refresh-only runs that report drifted resources without touching them. Spacelift schedules detection per stack and can fire a reconciliation run when drift appears, declared as a resource alongside the stack so your orchestration configuration stays reviewed Terraform like everything else.
Reconciliation is not the same as automatic repair. A reconciliation run is an ordinary tracked run and obeys the stack's own rules, so it applies unattended only if that stack has autodeploy enabled, and it still passes through your plan policies on the way. Enable both where the code is authoritative, and leave autodeploy off where an on-call engineer might legitimately need to hold a manual change.
Pin module versions and publish shared modules to a private registry while you are here. Teams copying a vpc module between repositories is the reason your standards diverged in the first place.
Answer three questions about your own setup.
Migrate one layer at a time, and never all stacks at once.
Step four is the one teams skip, and it decides whether engineers treat your policies as guardrails or as an obstacle to route around. Warn mode surfaces the rules that are wrong while a wrong rule is still harmless, rather than in the middle of a change that has to ship.
Terraform is centralized when you can answer who changed what, when, under which policy, and whether the live infrastructure still matches the code, without opening a terminal. Every layer above exists to make that one question answerable.