On this page
AI Agent Segregation of Duties: A Control Blueprint
tl;dr
AI agent segregation of duties is an urgent architecture problem, not a future policy exercise. Gartner projects 40% of enterprise applications will include task-specific AI agents by 2026, yet only 13% of organizations report having adequate agent governance. Agents can combine cross-system permissions at machine speed, creating unapproved privileges that traditional human-centric controls cannot catch.
By the end of 2026, Gartner expects 40% of enterprise applications to include task-specific AI agents, which makes AI agent segregation of duties an architecture problem rather than a future policy exercise.
That projection describes application integration, not widespread production deployment. Gartner data cited by NeuralCoreTech puts current organizational deployment at 17% and says only 13% believe they have the right agent governance in place. The gap matters because agents don’t wait for quarterly reviews. They read data, choose actions, call systems, and produce evidence continuously.
Why does AI Agent Segregation of Duties collapse?
Segregation of duties—SoD for short—keeps incompatible functions, such as initiating a payment and approving that payment, with different actors. Human-centric controls work because each action is relatively bounded, a person uses an attributable identity, and an auditor can reconstruct who did what.
Agents stretch those assumptions. A control-collapse analysis from Singulr argues that least privilege depends on predictable scope, bounded action, and accountable execution. An agent can cross several systems at machine speed, combine permissions from each environment, and create new effective privileges without a manager ever seeing the combination in one role.
Consider a finance agent that reads a purchase request, recommends a supplier, creates the purchase order, approves the payment, and writes the supporting log. Each function may look acceptable in isolation. Together, they give one machine identity the ability to control an entire transaction. The problem isn’t that the agent is “acting like an employee.” It’s that the employee is holding every seat in the control chain.
The underlying control frameworks care about roles and functions, not biological identity. A SoD guide from Hernan Huwyler maps the problematic pattern to COSO Principle 10, ISO/IEC 27001 Annex A control 5.3, and NIST SP 800-53 AC-5. Those controls don’t become optional because the actor is a model connected to an API.
What actions should an agent never control alone?
The practical boundary is simple: an agent may recommend, classify, reconcile, or draft, but it shouldn’t initiate, approve, execute, and record the same consequential action.
That boundary doesn’t require every agent action to stop for human review. A reconciliation assistant can classify a match without asking an operator. The point of intervention is where actions become irreversible or materially alter financial position. A payment approval, journal entry, cash movement, or privileged access change belongs on the other side of that line.
Huwyler’s analysis also explains why the default “human in the loop” answer breaks under high-volume pipelines. The relevant control boundary sits between inference steps, not merely between people. A person can approve the overall workflow while separate agent steps retain incompatible authority between prompts and tool calls.
The better pattern assigns different enforcement points to each function:
- Recommendation agent: reads approved inputs and proposes an action.
- Policy engine: checks the proposal against explicit business rules.
- Execution agent: acts only with a short-lived, purpose-scoped credential.
- Approval service: requires a named human for designated high-risk actions.
- Independent recorder: writes the decision and execution evidence outside the acting identity.
del.ai proposes a stricter substrate-level model for ERP finance agents. Its four-class permission design places journal entries, cash movements, and payment approvals in a MONEY-MOVE category that requires a named human in the approval chain. That’s a design document from a pre-revenue company, not validated customer evidence, but the architectural principle is sound: enforce the restriction below the model, where a prompt or role change can’t bypass it.
Can existing IAM and GRC controls handle agents?
Existing identity and access management can help, but only when teams extend it from role-based entitlement review to the agent’s actual behavior. A clean-looking role may be harmless in isolation while becoming dangerous beside a credential, delegated workflow, or API permission in another system.
That’s why activity-based SoD review is stronger than role-only attestation. NHI Mgmt Group’s guidance recommends correlating what an identity actually did with what it was allowed to do, drawing from transaction logs, API calls, workflow traces, and connected applications. The output is a more accurate view of “toxic combinations”—permissions that become dangerous only when assembled.
Visibility is already the weak point. SecurEnds research cited by NHI Mgmt Group says only 5.7% of organizations have full visibility into their service accounts and 97% of non-human identities carry excessive privileges. An agent inventory built only from IAM directories will miss credentials held by orchestration platforms, integration services, developer tools, and shadow deployments.
This is where governance, risk, and compliance tools—GRC platforms—can retain value, but their role changes. They can translate policies, collect evidence, route exceptions, and coordinate reviews. They shouldn’t be the only preventive control over a live agent action. A Compliance Council analysis makes the same distinction: a shared service account can leave an agent unattributable even when deployment governance was formally approved.
The evidence also supports retrofitting in the right places. An AMI Praha implementation story using MidPoint describes a financial institution using a central identity governance platform to detect cross-system SoD conflicts before access is granted. The open lesson is architectural: centralize policy evaluation where you can, then enforce it at provisioning and runtime. Don’t force every system to discover conflicts independently.
How much will agent controls cost?
Control spending is not limited to agent-orchestration software. It includes identity governance, access review, compliance evidence, runtime controls, and the labor required to route exceptions.
Reported pricing spans a wide range. AI Cyber Check’s 2026 pricing guide places mid-market agentic compliance platforms at $25K per year to $120K per year, while enterprise deployments run from $150K per year to $500K+ per year.
These options solve different problems, so the comparison starts with control scope rather than feature count.
| Tool or approach | Pricing in research | Primary control role | Best fit |
|---|---|---|---|
| Microsoft Agent 365 | $15/user/month | Agent inventory, shadow-agent detection, Entra identity integration, and conditional access | Organizations already centered on the Microsoft 365 ecosystem |
| OPAG access-review workflow | — | Prepares source-linked ERP review packets, highlights SoD conflicts, routes owners, and records human decisions | ERP access review involving IT, finance, security, and audit |
| del.ai ERP finance design | — | Proposes substrate-enforced permission classes and named-human approval for MONEY-MOVE actions | Teams evaluating architectural controls for ERP finance agents |
Per-seat pricing can look predictable while scaling poorly across the entire workforce. The only scenario supported by the research puts a 50-user Microsoft Agent 365 deployment at $9,000 per year: 50 × $15 × 12. Don’t extrapolate that figure to other team sizes or assume it covers the execution controls for every agent. The cited review says its method is research-based rather than hands-on testing.
The larger financial argument is about exposure. RSA cites IBM research finding that incidents involving shadow AI cost $670,000 more on average than standard incidents. A governance platform that merely creates another dashboard may be cheap next to an agent that bypasses the control system entirely.
What should an agent control architecture include?
Start with identity, because an agent without a distinct identity can’t be constrained or attributed cleanly. Give each deployment its own non-human identity, named owner, business purpose, permitted systems, credential lifetime, and decommission condition.
Discovery has to extend beyond the identity provider. Dataiku’s September 2026 launch coverage notes that fewer than one in five organizations keep a complete, current inventory of their AI systems. An inventory that covers only agents deployed through one cloud or one builder leaves the most important gaps open.
Next, enforce purpose at execution. Retrieval, planning, approval, and action should receive separate capabilities where the risk warrants it. A tool that can read a payment file shouldn’t automatically receive permission to initiate a payment. A credential should be issued for the current task and revoked when that task ends.
Runtime policy then closes the path between authorization and use. Archer’s agentic compliance launch describes policy-as-code guardrails applied before a model responds. Proofpoint’s agent-security approach translates business rules into runtime controls across systems. While neither product is a complete SoD architecture, both illustrate the right location for enforcement: before the consequential action, not weeks later during review.
Finally, record activity independently of the agent. A bank-focused control analysis calls for evidence covering successful, blocked, and failed actions, with enough context to replay a decision. It also recommends separating the system acting from the system recording the action. Otherwise, the agent can approve an action and author its own clean-looking evidence.
How should you roll out agent SoD controls?
Roll out one consequential workflow at a time, starting with a process where actions are bounded and reversible. A purchase-order assistant is safer than a payment-release agent because mistakes are easier to identify and correct. You need evidence from normal traffic, conflicts, denied actions, and attempted privilege expansion before expanding the mandate.
Use this implementation sequence:
- Map the transaction lifecycle. List initiation, authorization, execution, custody, and recording functions across every connected system.
- Inventory identities and behavior. Find credentials, delegated access, API calls, actual transactions, and shadow agents—not just assigned roles.
- Define prohibited combinations. Translate each toxic combination into a preventive rule with an owner, exception path, and expiration condition.
- Enforce at runtime. Use scoped credentials, purpose-bound tools, independent approvals, and pre-action policy checks.
- Test the control path. Attempt to create, approve, execute, and record a restricted action. Verify that at least one step fails for the intended reason.
- Reconcile continuously. Compare intended activity with actual activity and update rules when connected systems or agent behavior changes.
Human approval is appropriate for irreversible actions, but it shouldn’t be the universal answer. In multi-agent pipelines, the deeper concern is inherited authority: a downstream agent may receive a credential broad enough to ignore the boundary the upstream approval established. Our guide to agent delegation patterns that survive production reality covers scope narrowing and traceable handoffs in more detail.
Budget for controls that operate during execution, not just afterward. Seat-based visibility won’t show how many inference loops, tool calls, or background jobs an agent creates, a cost problem covered in AI agent budget controls. Security controls face the same scaling trap, especially when a coding agent can reach repositories, deployment systems, and cloud infrastructure through a reusable credential. Our dependency-security control guide explains why source, sandbox, permission, and financial boundaries need one connected enforcement path.
For a payment or journal-entry workflow, start by naming every identity and action, separating initiation from approval and execution from recording, then test whether a runtime credential can halt the agent before an irreversible action. If the system can complete that test, you have a control architecture. If it only has a policy page and a quarterly report, you have documentation.
Recommended Reading
-
Adversarial Testing for AI Agents: A Buyer’s Guide
Adversarial testing for AI agents is an infrastructure decision, not an optional security experiment. Commercial automated red-teaming platforms cost $20,000 to $300,000 annually, so teams must select tools aligned with their visibility, ownership, and budget constraints.
-
AI Agent Dependency Security: A Practical Control Guide
Effective AI agent dependency security requires an end-to-end control path from source code through sandbox execution, with enforceable financial and permission limits, not just standalone inventory tools. Autonomous agents expand attack surfaces beyond traditional CVE scanners, with documented incidents including 2,090 malicious RubyGems published in hours and unconstrained recursive loops incurring 50,000 USD in cloud costs in under an hour.
-
How to Build Multi-Agent Systems Without Losing Control
Most organizations deploy AI agents without governance, creating risk and cost overruns. Build the control plane first, tune the harness instead of the model, and pick frameworks by fit. Open-source stacks offer cost transparency and IP control.