Microsoft's AI Landing Zones project publishes a design checklist of 44 recommendations across 10 areas. It is a good checklist and a hard one to act on: it is organized by topic, not by when a decision becomes expensive to reverse.
This is the same checklist, restated in plain language and sorted into three phases: what to decide first, what to put in before the first agent ships, and what a second team or an auditor will test. The recommendations are Microsoft's. The order is ours.
What's Inside
- All 44 of Microsoft's recommendations, each tagged with its original checklist ID
- Three phases: decide first, before the first agent ships, before a second team or production
- A region check against the project's list of full-deployment regions
- A printable version with evidence prompts, a score sheet, a decision log and platform-team questions
Read this first
- The project is in preview. Its README says so and notes it may use preview services (1). As of October 9, 2026 it has no tagged releases, so pin the exact commit you deploy.
- The checklist still says "Azure AI Foundry". That was the product's earlier name. Search for it when you cross-reference (2).
- Regions. The README lists the regions that support a full deployment and says partial deployments elsewhere are not supported. No Canadian region is on that list as of October 9, 2026. If Canadian residency matters to you, settle this early: check the project's roadmap, ask your Microsoft account team, and consider whether an adapted build following the same checklist fits better. We make no statement about residency or compliance for any region; re-read the README before you decide (1).
How to use it
Gather your platform lead, a security or risk owner and whoever owns the budget. Mark a check "yes" only if you can show the artifact. Score each applicable check (yes = 2, partly = 1, no = 0), leave out the ones that don't apply, and start with the lowest-scoring phase. If Phase A is low, stop and fix it before anyone builds more.
A. Decide first
These choices set your network shape, who and what can authenticate, where data lives, and how cost is carved up. Changing them after agents are running means rebuilding, so decide them before anyone builds.
- No PaaS service or AI model endpoint is reachable from the public internet. Each one is reached through a private endpoint. (ALZ N-R3, Networking)
- Outbound traffic is restricted by default to approved services or domain names, and someone maintains the list of trusted destinations. (ALZ N-R9, Networking)
- Managed identities with least-privilege access are used on every Azure service that supports them. (ALZ I-R1, Identity)
- Static API keys are replaced by Microsoft Entra ID wherever possible, and anyone who still holds a key has been audited. (ALZ I-R3, Identity)
- Key-based access to AI model endpoints is disabled, so clients must authenticate with Microsoft Entra ID. (ALZ I-R6, Identity)
- Regions were chosen after confirming that every AI service you need is available there with the features you need. (ALZ R-R1, Resource Organization)
- You reviewed the quota needed to deploy these resources in your chosen region, before deploying. (ALZ R-R2, Resource Organization)
- Your resource organization fits Azure's subscription and regional quota limits, so growth doesn't cause surprise outages. (ALZ R-R3, Resource Organization)
- You scale through multiple Foundry resources and projects, with one Foundry resource per billing boundary so cost can be allocated by team. (ALZ R-R4, Resource Organization)
- Foundry Agent Service uses the standard setup, so threads, messages and uploaded files are stored in Azure resources you own. (ALZ D-R1, Data)
- Thread storage, file storage and vector search are separated by project, not shared across every project. (ALZ D-R2, Data)
- Compute is standardized across models, orchestrators, self-hosted agents and application parts, favoring PaaS options such as Azure Container Apps, App Service or AKS. (ALZ C-R1, Compute)
B. Before the first agent ships
The controls that stop the first agent from becoming the template for everything that follows: the gateway, the firewall, the policy, the baseline alerts and the content filter.
- Cost was estimated before building, using the pricing and billing model of Foundry and the services it uses. (ALZ CO-R1, Cost)
- Public-facing workloads have Azure DDoS Protection enabled, or the platform landing zone's central DDoS service covers them. (ALZ N-R1, Networking)
- Developer access goes through a jumpbox reached via Azure Bastion, or the platform landing zone's central equivalent. (ALZ N-R2, Networking)
- Network security groups are applied on every virtual network in the landing zone. (ALZ N-R4, Networking)
- A public front end sits behind Application Gateway or Azure Front Door with a Web Application Firewall. (ALZ N-R5, Networking)
- Azure API Management sits between front ends and AI endpoints as a generative AI gateway, including load balancing across endpoints. (ALZ N-R6, Networking)
- Outbound traffic passes through Azure Firewall (or a third-party firewall) using user-defined routes, preferably the platform landing zone's. (ALZ N-R7, Networking)
- Private DNS zones resolve your private endpoints, preferably the platform landing zone's, with policy that enforces private endpoints and DNS. (ALZ N-R8, Networking)
- MFA is enforced, and sensitive accounts use secondary admin accounts or just-in-time access through Privileged Identity Management. (ALZ I-R2, Identity)
- Conditional Access policies respond to unusual sign-ins, require MFA for critical AI resources, and can limit access by location or device compliance. (ALZ I-R4, Identity)
- Azure RBAC grants only the access each role needs, and assignments are reviewed regularly. (ALZ I-R5, Identity)
- Built-in Azure Policy for AI resources is applied automatically, at the management-group level. (ALZ G-R1, Governance)
- A baseline content filter from Azure AI Content Safety is defined for your approved models. (ALZ G-R4, Governance)
- Which models teams may deploy is controlled by an Azure Policy allowlist, started in audit mode and moved to deny only once you understand what teams need. (ALZ G-R5, Governance)
- Model outputs are inspected regularly, and Prompt Shields scan inputs for prompt-attack attempts. (ALZ S-R5, Security)
- Models, AI resources, data and applications are monitored so they stay aligned with workload KPIs. (ALZ M-R1, Monitoring)
- The recommended Azure Monitor baseline alert rules for AI are enabled. (ALZ M-R2, Monitoring)
- Diagnostic settings send logs and metrics from every deployed resource to a Log Analytics workspace, including audit logs for key services. (ALZ M-R4, Monitoring)
C. Before a second team or production
What a second team, an auditor or a bad day will test: risk standards, drift, disaster recovery and the cost levers. Some of these can wait until the first agent works, but not past the second team.
- Industry standards such as the NIST AI Risk Management Framework were reviewed, and the matching regulatory-compliance policy initiatives are applied. (ALZ G-R2, Governance)
- Responsible-AI standards are adopted, with reporting on model outputs. (ALZ G-R3, Governance)
- Microsoft Defender for Cloud recommendations are reviewed, including discovery of generative AI workloads and AI security posture management. (ALZ S-R1, Security)
- The Microsoft cloud security baseline and the Azure service guides are followed. (ALZ S-R2, Security)
- Microsoft Purview, for example Insider Risk Management, is used to assess data risks such as leaks and oversharing in AI workflows. (ALZ S-R3, Security)
- MITRE ATLAS and the OWASP generative AI risks are used to identify risks across the architecture. (ALZ S-R4, Security)
- Generative AI quality is monitored with Foundry's built-in and manual evaluation, alongside latency and vector-search accuracy, with tracing enabled. (ALZ M-R3, Monitoring)
- Model and data drift are tracked continuously, with custom alerts on performance thresholds. (ALZ M-R5, Monitoring)
- Azure Monitor Network Insights and Network Watcher are in place for network troubleshooting. (ALZ M-R6, Monitoring)
- A business-continuity and disaster-recovery policy covers your AI endpoints and AI data, with at least two regions considered. (ALZ R-R1, Reliability)
- Predictable workloads use provisioned throughput (PTU) on the primary endpoint, with a consumption-based endpoint for spillover. (ALZ CO-R2, Cost)
- Deployment types were compared, including global deployment, which offers lower cost per token on certain OpenAI models. (ALZ CO-R3, Cost)
- Non-production compute, such as VMs and compute instances in Foundry and Azure Machine Learning, shuts down automatically under policy. (ALZ CO-R4, Cost)
- If you use Microsoft Fabric, its data reaches Foundry through the Fabric data agent. (ALZ D-R3, Data)
Get the printable version
The checks above are free to read and share. The printable PDF adds what doesn't fit in a web page: an "ask for" line under each check naming the evidence that proves it, the score sheet, a region-check worksheet, a decision log for the choices the checklist leaves to you, and eight questions for your platform team.
Sources
- AI Landing Zones: project README (Azure/AI-Landing-Zones on GitHub), MIT licence, not archived, no tagged releases; README last changed 2026-09-23; states the project is in preview
- AI Landing Zones: design checklist (44 recommendations, 10 design areas), file last changed 2026-06-30 (commit 998a53e2)
- What is an Azure landing zone? (Cloud Adoption Framework), page updated 2025-07-24
- AI strategy (Cloud Adoption Framework): the six phases of AI adoption, page updated 2026-06-26
This guide restates Microsoft's recommendations in our own words and adds our own ordering. It is general information, not legal, security or compliance advice. The upstream project is in preview, so verify against the current source.
Where to go next
- Landing zone sorted, agents next? Score the controls that sit on top of it with our free Agent Platform Production Checklist.
- Want the design checklist explained area by area? See Azure AI Landing Zones: Fast Track to Enterprise AI.
- Not sure where your organization stands? Take the free Agentic AI Readiness Scorecard. It takes a few minutes and needs no email.
Kevin Evans
Founder, Code To Cloud Inc.
Kevin Evans leads agentic DevOps and application modernization engagements for enterprise and mid-market teams, and offers fractional CTO advisory to growing businesses. He spent nearly five years at Microsoft, rising to Senior Solutions Engineer, and is a CNCF Ambassador and Calgary chapter lead, based in Calgary, Alberta. More about Kevin
