A cloud landing zone is the pre-built foundation that every workload lands on: the account hierarchy, the identity model, the network backbone, the logging pipeline and the guardrails, all created by code before the first application team arrives. It is not a product you buy. AWS Control Tower, the Azure landing zone reference implementations and Google's enterprise foundations blueprint are accelerators for building one, but the decisions they encode are yours to own.
The reason to get it right early is that these decisions are expensive to change later. Moving a running workload between accounts, renumbering overlapping IP ranges or retrofitting audit logging that never existed are multi-month projects. This article explains each layer from first principles, using AWS names with Azure and Google Cloud equivalents. It covers how guardrails actually evaluate, how an account vending pipeline works, a worked request traced end to end, the failure modes we see most often, and a checklist to start from. For the identity primitives underneath, read cloud IAM first.
The unit of isolation
Everything starts from one choice: what is your unit of isolation? On AWS the strongest boundary that is cheap to create is the account. IAM permissions, service quotas, most resource limits and the billing line all stop at the account edge, so a compromised role or a runaway script in one account cannot touch another unless you explicitly allow it. Azure's equivalent unit is the subscription, grouped under management groups; Google Cloud uses projects grouped into folders under an organization node.
The common pattern is one account per workload per environment: payments-prod, payments-dev and so on. That sounds like a lot of accounts, and it is. A few hundred is normal for a mid-sized company. That is only manageable because the landing zone makes accounts cheap, uniform and automatically governed. The cost of many accounts is cross-account plumbing; the cost of few accounts is blast radius and tangled permissions, which is worse.
| Concept | AWS | Azure | Google Cloud |
|---|---|---|---|
| Isolation unit | Account | Subscription | Project |
| Grouping for policy | Organizational unit (OU) | Management group | Folder |
| Preventive policy | SCPs, RCPs, declarative policies | Azure Policy (deny effect) | Organization Policy constraints |
| Workforce access | IAM Identity Center | Entra ID + RBAC | Cloud Identity + IAM |
| Reference accelerator | Control Tower, LZA | Azure landing zones (CAF) | Enterprise foundations blueprint |
Organizational units and the core accounts
Organizational units exist to attach policy, not to mirror the org chart. Design them around how differently you want accounts governed, because a team reorganisation should not force an account move. A durable starting shape is shown below: a Security OU holding a log archive account and an audit or security-tooling account, an Infrastructure OU for the network hub and shared services, a Workloads OU split into production and non-production, a Sandbox OU with loose controls but hard budgets, and a Suspended OU that denies everything and holds accounts on their way out.
The management account deserves its own rule: run nothing in it. Service control policies do not apply to the management account, so anything deployed there runs outside your guardrails. Keep its access limited to a small group with hardware MFA, and delegate administration of security services such as GuardDuty and Security Hub to the audit account so day-to-day security work never needs management-account access.
Guardrails: how preventive policy really evaluates
Guardrails come in three kinds. Preventive controls stop an API call before it happens. Detective controls, such as AWS Config rules, find drift after the fact. Proactive controls, such as policy-as-code checks in the deployment pipeline, reject a template before it is applied. You want all three, but preventive controls are the ones that make a mistake impossible rather than merely visible.
To write SCPs correctly you have to know how they evaluate. An SCP never grants anything. It sets the maximum permissions available to principals in the accounts below it, and a request succeeds only if every SCP from the root down to the account allows it and the principal's own IAM policy allows it. An explicit deny anywhere wins. SCPs do not affect service-linked roles, and they do not apply to the management account. Resource control policies (RCPs), added in late 2024, are the mirror image: they cap what can be done to resources in your accounts, for the services that support them, regardless of who is calling. RCPs are how you stop an S3 bucket from being opened to principals outside your organization.
Here is a typical pair of deny statements. The first restricts usage to approved regions while exempting global services and the break-glass and pipeline roles. The second makes the logging baseline tamper-proof for everyone except the pipeline that manages it.
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "DenyOutsideApprovedRegions",
"Effect": "Deny",
"NotAction": ["iam:*", "organizations:*", "sts:*", "support:*",
"cloudfront:*", "route53:*", "budgets:*", "waf:*"],
"Resource": "*",
"Condition": {
"StringNotEquals": {"aws:RequestedRegion": ["eu-west-1", "eu-central-1"]},
"ArnNotLike": {"aws:PrincipalARN": [
"arn:aws:iam::*:role/lz-break-glass",
"arn:aws:iam::*:role/lz-pipeline"]}
}
},
{
"Sid": "ProtectLoggingBaseline",
"Effect": "Deny",
"Action": ["cloudtrail:StopLogging", "cloudtrail:DeleteTrail",
"config:StopConfigurationRecorder",
"config:DeleteConfigurationRecorder"],
"Resource": "*",
"Condition": {"ArnNotLike": {"aws:PrincipalARN": "arn:aws:iam::*:role/lz-pipeline"}}
}
]
}Two practical limits shape SCP design: a policy document is capped at 5,120 characters and an OU or account can have at most five SCPs attached directly. That pushes you towards a small number of well-commented policies attached high in the tree. Test every change on a Sandbox OU first; a bad region-deny can break global services you forgot to exempt.
Identity: federated, short-lived, reviewable
No human should have a long-lived IAM user in a workload account. Federate your workforce identity provider (Entra ID, Okta, Google Workspace) into IAM Identity Center, define permission sets such as ReadOnly, Developer and WorkloadAdmin, and assign identity-provider groups to permission sets per account. A person's access then becomes a group membership you can review, and every console or CLI session is short-lived and traceable to a named user in CloudTrail.
Keep two exceptions on paper. Break-glass roles exist for when the identity provider is down; they live in a small number of accounts, require hardware MFA, and page the security team whenever they are assumed. Machine identities for CI/CD should use OIDC federation from the CI provider into a narrowly scoped role, never stored access keys. The cloud security baseline article goes deeper on the controls that sit on top of this identity layer.
Network: address planning and hub-and-spoke
Networking is the layer most often regretted. The two decisions that matter most are address planning and traffic topology. For addresses, allocate CIDR ranges from a central IP address manager (AWS VPC IPAM or an equivalent) so no two VPCs that might ever need to talk overlap. Reserve a large block per region and per environment, and hand each account a fixed-size slice, for example a /22, from the right pool.
For topology, the common choice is hub-and-spoke. Each workload VPC attaches to a Transit Gateway owned by the network account. Separate route tables keep production and non-production spokes from reaching each other. Internet egress goes through a central egress VPC with NAT and inspection, so outbound traffic has one place to log and filter. DNS resolution for on-premises and private zones is centralised in the hub too. The trade-off is cost and latency: every packet that crosses the Transit Gateway pays a per-GB processing charge, and centralised NAT concentrates data-transfer charges in one account. High-volume east-west paths sometimes justify direct peering or PrivateLink instead. Cloud networking covers these building blocks individually.
Logging and security tooling
Turn on an organization trail so CloudTrail records API activity in every current and future account, delivered to a bucket in the log archive account. Enable AWS Config recording in every account and region you use, aggregated in the audit account. Send VPC flow logs and DNS query logs to the same archive. Then make the archive hard to tamper with. Use a separate KMS key owned by the log archive account, S3 Object Lock in compliance mode for the retention period your regulators require, and a bucket policy that denies deletes from every principal. The deny statement in the SCP above stops a workload admin from switching the recorders off.
Security services follow a delegated-administrator pattern. GuardDuty, Security Hub, Inspector and IAM Access Analyzer are enabled organization-wide and administered from the audit account, which becomes the single pane for findings. Route high-severity findings to an on-call queue, not just a dashboard.
The account vending pipeline
Account vending is the pipeline that turns a request into a fully governed account. Control Tower's Account Factory, Account Factory for Terraform (AFT) and the Landing Zone Accelerator are ready-made versions. Whichever you use, the steps are the same, and each step must be idempotent so a failed run can simply be retried:
def vend_account(req):
validate(req) # owner, team, env, cost_center, data_class
acct = org.find_by_email(req.email) or org.create_account(req.name, req.email)
org.wait_until_active(acct)
org.move_account(acct, target_ou(req.env, req.data_class)) # guardrails apply here
cidr = ipam.allocate(pool=f"{req.region}-{req.env}", netmask=22)
baseline.apply(acct, vars={"cidr": cidr, "tags": req.tags}) # VPC, roles, log config
network.attach_to_tgw(acct, cidr, route_table=req.env)
sso.assign(acct, group=f"{req.team}-{req.env}-admins", permission_set="WorkloadAdmin")
budgets.create(acct, monthly_usd=req.budget, notify=[req.owner])
smoke_test(acct) # e.g. a call in a denied region must fail
registry.record(acct, req) # account inventory = source of truth
return acctThe smoke test is the step people skip and regret. It proves the guardrails actually bind, for example by attempting an action in a disallowed region with a test role and asserting AccessDenied, before the account is handed over. The baseline itself should be versioned. When it changes, the pipeline re-applies it to every account in waves, starting with sandbox, so drift between old and new accounts never accumulates. Config drift reconciliation describes how to keep those accounts converged afterwards.
Worked example: vending a regulated production account
Trace one request. The payments team asks for a production account in eu-west-1 for a card-processing service handling regulated data. They fill in a form that commits to a pull request in the account-request repository: owner, team, environment prod, data class pci, cost centre and a monthly budget of 8,000 USD.
Review approves the PR. The pipeline creates payments-prod, moves it to the Prod/PCI OU (which inherits the region and logging SCPs from Workloads, plus a stricter SCP that denies unencrypted storage), allocates 10.40.12.0/22 from the eu-west-1 production pool, applies baseline v14, attaches the VPC to the Transit Gateway's production route table, and assigns the payments-prod-admins group. The smoke test tries ec2:RunInstances in us-east-1 and gets the expected denial. The whole run takes about the time it takes the account to become active, typically minutes, not days of ticket queues.
A week later an engineer tries to stop the CloudTrail trail while debugging noisy logs. The call fails with an explicit deny from the SCP, the attempt itself is recorded in the archive, and a Config rule confirms the recorder is still on. Their spend appears in the cost report under the payments cost centre because the tags were applied at vending time. Nobody had to remember any of this. See cloud FinOps for how that allocation feeds chargeback.
Failure modes and trade-offs
These are the failure modes that cost the most:
- Workloads in the management account. They escape every SCP. Move them out early, while it is still cheap.
- OUs that mirror the org chart. Reorganisations then force account moves and policy churn. Group by governance need instead.
- Overlapping CIDRs. Two VPCs on 10.0.0.0/16 can never be routed together without NAT tricks. Allocate from IPAM from day one.
- Click-ops exceptions. A hand-edited SCP or route table drifts from code and is overwritten on the next pipeline run, or worse, is not. Every exception goes through the repository, with an expiry date.
- Automated destruction. Auto-remediation that deletes non-compliant resources, such as an untagged bucket, will one day delete data that mattered. Quarantine or block, notify, and let a human delete.
- Guardrails nobody tested. A region-deny that also blocks a global service can break sign-in or DNS across the whole OU. Roll policy changes through Sandbox, then Non-prod, then Prod.
The central trade-off is control against autonomy. Tight preventive guardrails reduce incidents but generate exception requests; loose ones shift effort to detection and clean-up. A good rule is to prevent what is irreversible or regulated, such as deleting logs, using unapproved regions or exposing data publicly, and to detect and coach on everything else.
What to do next
- Write down your isolation unit and account naming rule, and confirm nothing runs in the management account.
- Draw your OU tree around governance differences: Security, Infrastructure, Workloads (Prod, Non-prod), Sandbox, Suspended.
- Enable an organization CloudTrail and Config aggregation into a locked-down log archive account with Object Lock.
- Federate your identity provider into IAM Identity Center and remove IAM users with console passwords.
- Reserve address space in IPAM per region and environment before the next VPC is created.
- Put your region-deny and logging-protection SCPs in code, and test them on a sandbox OU first.
- Build or adopt a vending pipeline with idempotent steps and a guardrail smoke test, and route all new accounts through it.