Agent projects usually stall when a second team wants in and nobody can say who owns what, what gets tested, what it costs, or who can switch it off.
This is the checklist we use to ask those questions. It is 33 yes-or-no checks across four guardrails and a foundation. Answer each one only if you can produce the evidence behind it, and treat "we think so" as a no.
What's Inside
- Four guardrails: an evaluation gate, identity on every model call, spend tied to what shipped, and a person where it matters
- A foundation check: can a second team reuse what the first one built?
- Every check traced to Microsoft documentation, with page links
- A printable version with evidence prompts, a score sheet, an ownership map, ten questions for any partner, and a one-page executive brief
How to use it
Gather an engineering lead, someone from security or risk, and whoever owns the budget. Mark a check "yes" only if you can show the artifact. Score each guardrail (yes = 2, partly = 1, no = 0), and start with whichever scores lowest. If several are low, start with identity: every other control is easier to enforce once each call is attributable to a team.
A source number or link after a check points to Microsoft documentation. "Our recommendation" marks a check that is Code To Cloud's own view, not a Microsoft requirement. Several Foundry features named here are in preview, so check each one's current status before you depend on it. Sources were checked on October 9, 2026.
1. An evaluation gate before production
Without a gate, the only test a new agent version gets is whoever happens to try it first. The gate is an automated check every version must pass before it reaches production, so a second team can rely on the result.
- Every agent version is evaluated automatically before it reaches production, as a step in the pipeline. (4, 5)
- A written pass/fail threshold exists, and missing it fails the job. (5)
- The test dataset is yours, versioned, and includes cases where an agent went wrong before. (4)
- Quality and safety evaluation also runs on a sample of live production traffic (continuous evaluation). (4)
- A scheduled evaluation against the same test data catches drift between releases. (4)
- Adversarial testing (red teaming) runs before first release and again after major model or architecture changes. (1, 2, 4)
- A new version is validated as a candidate while production stays pinned to the current one, and is promoted only after the checks pass and a person approves. (5)
2. Identity on every model call
If you can't say which team and which purpose a call belongs to, you can't govern it, bill it or revoke it.
- Every agent has its own identity. No two agents share an account. (1)
- Agents authenticate with managed identity. No API keys sit in code, configuration or chat. (1, 3)
- Every model and tool call goes through one gateway. Nothing calls a model endpoint directly. (2, 3)
- Each call is attributable to a team and a purpose in the logs. (3)
- Agents and tools have least-privilege access: tools enforce the user's permissions or use narrowly scoped service accounts. (1)
- External tools and MCP servers are limited to an approved list. (1, 3)
- You can pause or revoke one agent or tool quickly. (1, 2)
3. Spend tied to what shipped
Token counts show activity. They don't show whether an agent paid for itself, so someone has to compare what each agent cost with what it delivered.
- You can see token use for each team and each agent, as well as the monthly total. (1, 3)
- Token limits or quotas stop one team from using up a shared allocation. (3)
- Resources are tagged by agent or use case, and budget alerts are set. (1)
- A named person reads cost against outcomes delivered (value per token), not tokens alone, and owns that number. (our recommendation)
- Model choice is reviewed: routine work goes to cheaper models, and routing (for example Foundry's model router) is tested against your own baseline. (2, 6)
- Consumption is reviewed monthly, and quotas follow business priority. (2)
- A quarterly audit retires agents nobody uses. (2)
4. A person where it matters
Agents do the routine work. Decide in advance which steps need a person to approve them, and write that into the pipeline so it doesn't depend on who is on shift.
- Every agent has a named, accountable human owner. (1)
- One registry lists every agent with its owner, purpose, platform and access scope. (1)
- Actions that change the world (send, write, delete, pay) are classified, and the ones that need human approval are written down and enforced in the pipeline. (our recommendation)
- A person approves each production deployment and monitors ongoing compliance. (1)
- A rehearsed plan exists to disable an agent fast, preserve its logs and inform stakeholders. (1)
- People are told when they are dealing with an AI agent. (1)
- Agent alerts reach your security operations team. (1, 2)
Foundation: can a second team reuse this?
The pilot team's pipeline is the platform's first draft. These checks tell you whether the next team inherits the controls or rebuilds them.
- A governed landing zone (identity, network, policy) exists for AI workloads. (1)
- Environments are built from code (Bicep or Terraform), so they can be reproduced. (2, 5)
- Development, test and production are separate environments. (7)
- Pipelines, prompts and integration patterns are reusable templates, so a new team inherits the controls. (2)
- Agents that touch internal data are kept apart from public-facing ones, for example in separate subscriptions. (1)
Get the printable version
The checklist above is free to read and share. The printable PDF adds what doesn't fit in a web page: an "ask for" line under each check naming the evidence that proves it, a score sheet, an ownership map, ten questions to ask any partner who builds your agents (including us), a one-page executive brief, and a page on what an engagement with us covers.
Sources
- Govern and secure AI agents across the organization (Cloud Adoption Framework), updated 2026-10-01
- Manage AI agents across your organization (Cloud Adoption Framework), updated 2026-06-26
- AI gateway capabilities in Azure API Management, updated 2026-06-25
- Observability in generative AI (Microsoft Foundry), updated 2026-08-26
- Set up CI/CD for hosted agents with the Azure Developer CLI (Microsoft Foundry), updated 2026-09-30
- Model router for Microsoft Foundry, updated 2026-09-02
- Agentic AI adoption maturity model (Microsoft), updated 2026-09-30
This is general information, not legal, security or compliance advice. Vendor products change quickly, so verify against current documentation.
Where to go next
- Not sure where your organization stands? Take the free Agentic AI Readiness Scorecard. It takes a few minutes and needs no email.
- Want to run the same questions against your own platform? The Agent Platform Readiness Review is a two-week look at evaluation gates, identity and gateway, cost controls and your landing zone. Book a discovery call below and we'll scope it together.
Kevin Evans
Founder, Code To Cloud Inc.
Kevin Evans leads agentic DevOps and application modernization engagements for enterprise and mid-market teams, and offers fractional CTO advisory to growing businesses. He spent nearly five years at Microsoft, rising to Senior Solutions Engineer, and is a CNCF Ambassador and Calgary chapter lead, based in Calgary, Alberta. More about Kevin

