Why Terraform Workflows Break Down Across Multiple Teams
Terraform workflows break down at scale because every team builds its own pipeline, and nothing forc 2026-9-27 19:0:0 Author: hackernoon.com(查看原文) 阅读量:1 收藏

Terraform workflows break down at scale because every team builds its own pipeline, and nothing forces those pipelines to agree.

The pattern looks the same everywhere. A repository per team, a slightly different GitHub Actions workflow in each one, more state backends than anybody can account for, and a terraform apply that somebody ran from a laptop during an incident. Some teams run a plan step but no policy check. One team has long-lived AWS access keys in repository secrets that nobody has rotated since the engineer who created them left.

Now try to answer a few questions about that setup. Who changed the shared VPC? Which accounts still allow publicly accessible RDS instances? Does the running infrastructure still match what is in main?

If answering requires opening a terminal, the problem is not Terraform. The problem is that four separate control planes are missing, and each one needs a different fix.

What does centralizing Terraform actually mean?

Centralizing Terraform means putting state and credentials, execution, policy, and visibility under shared ownership. Most teams centralize execution first because it is the visible pain, then wonder why governance still feels loose. Work from the bottom up instead.

  • State and credentials: Without a shared standard, you get corrupted state files, standing cloud credentials, and no clear blast radius. Build it with a remote state backend and OIDC federation. Managed platforms solve it by issuing short-lived credentials per run.
  • Execution: Without a shared runner, engineers apply from laptops, concurrent applies race each other, and the plan a reviewer approved is not the plan that runs. Build it with a reusable CI workflow or self-hosted Atlantis. Buy it as HCP Terraform or Spacelift.
  • Policy: Without policy as code, each team enforces its own rules, or nobody enforces any. Build it with Conftest, Checkov, or Trivy. Managed platforms run the same checks server side, where developers cannot skip them.
  • Visibility: Without scheduled reconciliation, manual console changes go undetected until production breaks. Build it with a cron job running terraform plan -detailed-exitcode. Buy it as drift detection with automatic reconciliation.

How do you centralize Terraform state and credentials?

Start by splitting state per team and per environment, then removing every static credential.

State boundaries are blast radius boundaries. One monolithic state file means a typo in a staging module locks the production pipeline, and every plan refreshes 900 resources it has no reason to touch.

If you still run a DynamoDB table for state locking, delete it. Terraform 1.10 added native S3 locking, and 1.11 promoted it to generally available while deprecating the DynamoDB argument.

terraform {
  backend "s3" {
    bucket       = "acme-tfstate-prod"
    key          = "platform/networking/terraform.tfstate"
    region       = "us-east-1"
    encrypt      = true
    use_lockfile = true
  }
}

Next, replace long-lived access keys with OIDC federation. Each workflow run then receives a short-lived role session scoped to the repository and trigger that requested it.

permissions:
  id-token: write
  contents: read

jobs:
  plan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::111122223333:role/terraform-plan
          aws-region: us-east-1

The IAM trust policy is where governance actually lives. Scope the subject claim to a specific repository and trigger, because a wildcard such as repo:acme/* is functionally the same as sharing the credentials with every repository you own.

{
  "Effect": "Allow",
  "Principal": {
    "Federated": "arn:aws:iam::111122223333:oidc-provider/token.actions.githubusercontent.com"
  },
  "Action": "sts:AssumeRoleWithWebIdentity",
  "Condition": {
    "StringEquals": {
      "token.actions.githubusercontent.com:aud": "sts.amazonaws.com"
    },
    "StringLike": {
      "token.actions.githubusercontent.com:sub": "repo:acme/platform-networking:pull_request"
    }
  }
}

Use two roles, not one. The plan role above is the one a pull request can assume, and it is read-only in the account apart from s3:GetObject, s3:PutObject, and s3:DeleteObject on the lock file, which use_lockfile needs in order to take a lock. The apply role gets write permissions and a subject claim of repo:acme/platform-networking:ref:refs/heads/main.

Check the subject format before you write either policy. Repositories created, renamed, or transferred after 15 July 2026 carry immutable numeric IDs in the claim, as in repo:acme@123456/platform-networking@789012:pull_request, and a condition matching only the older shape fails closed with no useful error.

How do you centralize Terraform execution?

Three options, in increasing order of cost and capability.

Option 1. A reusable CI workflow

Publish one workflow_call workflow in a central repository and have every team reference it by tag. This is the cheapest real fix and it eliminates pipeline drift immediately. It does not close the gap between plan and apply. The plan a reviewer approved ran against a state that a concurrent merge has since changed, and nothing prevents two applies from racing. Secrets and role ARNs also stay scattered across repositories.

Option 2. Atlantis

Atlantis is an open source, pull request driven Terraform runner, and it is the standard answer for teams of four to ten. The governance lives in the server-side configuration, not the repository file, because an atlantis.yaml that any team can edit enforces nothing.

# repos.yaml, on the Atlantis server
repos:
  - id: /github.com/acme/.*/
    apply_requirements: [approved, mergeable, undiverged]
    allowed_overrides: [workflow]
    allow_custom_workflows: false
    workflow: default

workflows:
  default:
    plan:
      steps:
        - init
        - plan
        - run: terraform show -json $PLANFILE > $SHOWFILE
        - run: conftest test --policy /policies $SHOWFILE

undiverged is the requirement most teams miss. It rejects an apply when the branch is behind its base, which is exactly the race condition that reusable CI workflows leave open. It only takes effect under the merge checkout strategy, so start the server with --checkout-strategy=merge or the requirement silently does nothing.

Atlantis does not provide role-based access control beyond what your version control system offers, dependencies across repositories or outputs passed between stacks, or anything on a schedule. depends_on orders projects within a single pull request and stops there. It also means operating a server that holds apply credentials for every account you manage.

Option 3. A managed orchestrator

HCP Terraform and Spacelift both run Terraform on their own workers rather than in your CI, and that architectural choice is what buys you the capabilities the first two options cannot reach. You get runs on a schedule, dependency graphs that pass outputs between stacks, policy evaluated server side where nobody can skip it, and access control that is independent of version control permissions. The tradeoff is that a hosted worker needs credentials for your accounts. Both support self-hosted workers running in your own network if that tradeoff is unacceptable, at the cost of operating something again.

How do you enforce Terraform governance with policy as code?

Enforce governance against the plan, not the HCL, because HCL does not tell you what is about to be destroyed.

Static scanners still earn their place for the cheap mistakes. Run Checkov or Trivy in pre-commit so engineers get feedback before CI does. For everything else, convert the plan to JSON and evaluate it with Open Policy Agent.

terraform plan -out=tfplan.binary
terraform show -json tfplan.binary > plan.json
conftest test --policy ./policies plan.json
package main

import rego.v1

mutating := {"create", "update"}

deny contains msg if {
	rc := input.resource_changes[_]
	rc.type == "aws_db_instance"
	rc.change.actions[_] in mutating
	rc.change.after.publicly_accessible
	msg := sprintf("%s sets publicly_accessible = true", [rc.address])
}

destroys contains rc.address if {
	rc := input.resource_changes[_]
	"delete" in rc.change.actions
}

warn contains msg if {
	count(destroys) > 3
	msg := sprintf("plan destroys %d resources, needs a second reviewer", [count(destroys)])
}

The reason to write governance in Rego rather than bash is portability. The rule logic survives the move from Conftest on a laptop to a server-side policy in HCP Terraform or Spacelift, even though the input path does not. The plan sits at input.resource_changes under Conftest, input.plan in HCP Terraform, and input.terraform in Spacelift, so you write the rule once and change one line per destination. Sentinel is the alternative if you have standardized on HCP Terraform, with the tradeoff that Sentinel runs nowhere outside HashiCorp's own products.

How do you detect Terraform drift across teams?

Detect drift on a schedule, because CI cannot. This is a structural limit rather than a missing feature. CI runs on merge, and drift happens when somebody edits a security group in the console to unblock a customer.

Building detection yourself is straightforward for a handful of stacks. Exit code 2 means the plan found changes, and you have to capture that code yourself because a failed step exposes no exit status to later steps.

# .github/workflows/drift.yml
on:
  schedule:
    - cron: "0 6 * * *"

jobs:
  detect:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: terraform init -input=false
      - id: plan
        run: |
          set +e
          terraform plan -detailed-exitcode -lock=false -no-color > plan.txt
          echo "exitcode=$?" >> "$GITHUB_OUTPUT"
      - if: steps.plan.outputs.exitcode == 2
        env:
          GH_TOKEN: ${{ github.token }}
        run: gh issue create --title "Drift in $STACK" --body-file plan.txt

Multiply that across 30 stacks and you are maintaining a fleet of cron jobs that nobody owns. The scheduling, fan-out, and reconciliation are the part worth paying for.

Know which half you are buying. HCP Terraform covers detection through health assessments, which are scheduled refresh-only runs that report drifted resources without touching them. Spacelift schedules detection per stack and can fire a reconciliation run when drift appears, declared as a resource alongside the stack so your orchestration configuration stays reviewed Terraform like everything else.

Reconciliation is not the same as automatic repair. A reconciliation run is an ordinary tracked run and obeys the stack's own rules, so it applies unattended only if that stack has autodeploy enabled, and it still passes through your plan policies on the way. Enable both where the code is authoritative, and leave autodeploy off where an on-call engineer might legitimately need to hold a manual change.

Pin module versions and publish shared modules to a private registry while you are here. Teams copying a vpc module between repositories is the reason your standards diverged in the first place.

Which approach should you choose?

Answer three questions about your own setup.

  • One to three teams, isolated accounts, low blast radius: Use a reusable CI workflow, OIDC, and Conftest. Buy nothing. Evaluating platforms will cost you more effort than a platform saves you.
  • Four to ten teams sharing networking or accounts: You need queueing and enforced policy. Choose Atlantis if branch protection satisfies your governance requirements and you are willing to operate the server. Choose a managed orchestrator if running a credential-holding server is not a job anybody on your team wants.
  • Compliance requirements, dozens of stacks, or infrastructure beyond Terraform: Once you need audit trails, granular access control, cross-stack dependencies, and one workflow covering Pulumi, CloudFormation, Ansible, and Kubernetes manifests, you are building an internal product. Buy one instead.

How do you migrate without freezing deployments?

Migrate one layer at a time, and never all stacks at once.

  1. Move state to a consistent backend with native locking, one stack at a time. Nothing else changes yet.
  2. Replace static credentials with OIDC. Delete the old keys and watch what breaks. Something will.
  3. Move one low-risk stack to the new execution model and run it in parallel with the old pipeline until both agree on a real change.
  4. Add policies in warn mode and leave them there long enough to read what they catch.
  5. Switch to deny, then migrate the remaining stacks.

Step four is the one teams skip, and it decides whether engineers treat your policies as guardrails or as an obstacle to route around. Warn mode surfaces the rules that are wrong while a wrong rule is still harmless, rather than in the middle of a change that has to ship.

How do you know it worked?

Terraform is centralized when you can answer who changed what, when, under which policy, and whether the live infrastructure still matches the code, without opening a terminal. Every layer above exists to make that one question answerable.


文章来源: https://hackernoon.com/why-terraform-workflows-break-down-across-multiple-teams?source=rss
如有侵权请联系:admin#unsafe.sh