# How Should Teams Secure Delegation Across Multi-Agent AI Workflows?

Colton Ramsey · October 2, 2026

> What Agent Delegation Security Actually Protects Agent delegation security is the set of technical, operational, and governance controls used when one...

## What Agent Delegation Security Actually Protects

Agent delegation security is the set of technical, operational, and governance controls used when one AI agent assigns work, permissions, data, or authority to another agent. In a multi-agent workflow, delegation changes who can act on a user’s behalf: a planner may ask a researcher to search internal documents, a coding agent may request deployment rights, or a purchasing agent may authorize payment below a fixed limit. The security problem is not simply whether each individual action is valid, but whether that action remains appropriate after authority passes through one or more agents.

**Also worth reading:** [How do enterprises secure agentic AI workflows against data leakage and autonomous errors?](https://tryinterlock.com/knowledge/how_do_enterprises_secure_agentic_ai_workflows_against_data_leakage_and_autonomous_errors.php) · [How Should You Implement OpenTelemetry Agent Tracing for Production AI Workflows?](https://tryinterlock.com/knowledge/how_should_you_implement_opentelemetry_agent_tracing_for_production_ai_workflows.php) · [How Should Enterprises Control Agent Identity Security Without Slowing AI Workflows?](https://tryinterlock.com/knowledge/how_should_enterprises_control_agent_identity_security_without_slowing_ai_workflows.php)

A direct prompt from a user to a model differs from a chain in which the user delegates to Agent A, Agent A delegates to Agent B, and Agent B calls a database through a narrowly scoped credential. Each hop can expand the meaning of the original request or accidentally expose more authority than intended. Research on the “delegation problem” and reports about vague tasks combined with total tool access describe the same basic risk: an underspecified instruction can receive a broadly privileged implementation. Delegation security therefore needs to treat agents as partially trusted actors rather than transparent extensions of the person who launched them.

The goal is controlled autonomy: agents should complete useful work without allowing unverified instructions, confused-deputy attacks, overlong sessions, or compromised downstream tools to become system-wide incidents. This does not mean removing all agent-to-agent access. It means attaching identity, purpose, scope, expiration, auditability, and revocation rules to every delegated capability. The correct baseline is bounded authority, explicit trust boundaries, and evidence that can be reviewed after the fact.

## Why Delegation Creates More Risk Than Ordinary API Calls

Traditional API authorization usually begins with a known application identity requesting a defined operation. Agentic systems complicate that model because instructions are probabilistic, tasks can be decomposed dynamically, and natural-language goals may not translate cleanly into machine-enforceable constraints. An agent can understand the general objective—“prepare the quarterly report”—without being able to determine exactly which files, records, tools, and dollar amounts are required. A convenient implementation may therefore give it broad access and ask it to exercise restraint, even though restraint is not a security boundary.

Authority can also accumulate unintentionally. Agent A receives read access to a customer record, Agent B is allowed to summarize it, and Agent C can send that summary to an external service. If credentials are copied at each stage, the final service may receive an identity with more power than the original workflow required. OAuth 2.0 Token Exchange, standardized in RFC 8693, offers a useful mechanism for exchanging one subject token for another when authorization crosses domains. It does not, by itself, define whether the requested agent or task should receive the resulting token; policy still has to decide that.

Agent delegation also introduces prompt injection as an authorization problem. Untrusted text inside an email, document, web page, or tool response may tell an agent to ignore its task and disclose data or invoke another tool. Model-level safeguards cannot provide the deterministic guarantees expected of an access-control layer. The agent may be instructed never to reveal secrets, but a reliable policy boundary should deny access to those secrets in the first place. This is why a secure workflow combines model instructions, runtime policy, scoped credentials, filtered context, and human approval gates.

## The Security Model for Controlled Agent-to-Agent Work

A practical delegation model should represent the original user, the delegating agent, the receiving agent, the target resource, the requested action, and the delegated task. A token or policy decision should answer concrete questions: Is Agent B allowed to act for this user? Is this action limited to the stated customer and report period? Can it write but not delete? Is external transmission permitted? Does authorization expire after 15 minutes or 500 tool calls? Who can revoke it? Without these attributes, the system often falls back to a reusable service credential that all agents share.

Least privilege is the central control, but “least” must be calculated from the task rather than copied from a generic role. If a summarization agent only needs three report files, the credential should refer to those files or an equivalent data-selection rule. If a research agent may query approved sources but not post the result, network and tool permissions should differ. Cedar, the open-source authorization language discussed by AWS for enforcing least-privilege policy, can express contextual rules such as resource ownership, action type, principal identity, and delegation depth. It is useful for policy evaluation, although a mature deployment still needs identity issuance, enforcement points, logging, and tested operational procedures.

Zero trust is often treated as a slogan, but it has a precise application here. Every delegation should be authenticated, authorized independently, and scoped to a narrow purpose. Agents should not be trusted merely because they run in the same platform or were selected by another agent. Long-lived shared credentials should be replaced where possible with short-lived tokens, audience-restricted claims, workload identity, or separately issued credentials. A receiving agent must also reject a delegation that is expired, intended for another audience, missing a task binding, or inconsistent with local policy. Trust is therefore continuous rather than inherited forever.

## How to Secure a Multi-Agent Workflow in Practice

Begin by inventorying agents, tools, identities, data classes, and delegation paths. A modest three-agent workflow may contain five distinct authority transitions, while a dynamic system can create decisions that are difficult to reproduce. Record which agent can call each tool, which identities are available at each step, whether tokens are copied or exchanged, and where untrusted content enters the context. A useful initial threshold is to block any delegation that cannot be tied to an identifiable principal, purpose, resource set, and expiration policy.

Next, define policies in machine-readable form and test them before production use. Start with deny-by-default permissions, then add only the actions required by a documented workflow. Separate read, write, delete, financial, administrative, and external-communication capabilities rather than grouping them into one “editor” or “manager” role. Set hard limits for actions per task, data volume, spend, destination domains, and agent-hop depth. A 30-minute token is already a major improvement over a credential that remains valid for a year, but a 5- or 10-minute token may be better for a short tool operation if identity infrastructure supports it reliably.

Human approval should be reserved for decisions where mistakes are difficult to reverse or outside ordinary policy. Typical examples include sending regulated information to an external recipient, transferring more than $500, changing production access, deleting customer data, or making a public commitment. Approval must occur at the point of action and display the exact recipient, resource, scope, and intended effect; approving a vague description earlier does not necessarily authorize a later consequential step. After rollout, monitor denied requests, unusual delegation chains, token reuse, repeated retries, and activity outside normal operating hours. NIST and ISO AI-governance guidance can support this process, but neither framework makes an unsafe workflow safe merely because a control has a formal name.

## Delegation Security Controls Compared

Organizations can secure agent handoffs through several complementary layers. No single option covers identity, intent, runtime enforcement, and audit evidence equally well, so the strongest design usually combines policies, isolated credentials, isolated execution, monitoring, and human checkpoints. The following comparison is about control fit rather than a ranking of vendors.

| Feature | Central policy and scoped authorization | Separate identities and isolated runtimes | Human approval gates | Prompt and output safeguards |
| --- | --- | --- | --- | --- |
| Primary purpose | Decide whether an action is allowed | Limit what a compromised agent can reach | Prevent high-impact mistakes | Reduce unsafe or malicious behavior at model level |
| Enforcement strength | High when checked at every sensitive call | High blast-radius control | High for covered actions | Variable and probabilistic |
| Best for | Cross-agent and cross-service policy | High-value data and production tools | Irreversible or regulated decisions | Defense in depth |
| Operational cost | Policy design and evaluation work | Provisioning, patching, observability | User or operator review time | Testing and model-output monitoring |
| Common weakness | Policy gaps or token over-scoping | Fragmented administration | Approval fatigue or vague prompts | Injection can still cross weak boundaries |

A central policy engine is preferable when many agents need consistent decisions across tools. Separate credentials and runtimes become more important when one agent could cause broad damage through direct filesystem, network, or administrative access. Human review is effective for consequential exceptions but is expensive, so organizations should approve categories rather than every routine read. Model safeguards add useful behavioral controls, yet they should never be the only barrier between an untrusted document and a production credential.
A layered example illustrates the point. A coordinator delegates a vendor comparison task to three research agents for no more than 20 minutes. Each agent receives access only to approved source domains, cannot view internal contracts, and cannot contact suppliers directly. The aggregation agent can write to a private workspace but cannot publish results. A policy service checks the user, task identifier, permitted domains, and delegation depth on each tool call. If one research agent encounters instructions embedded in a webpage, it still has no credential for contracts or publishing, while the security team can revoke the associated workload identity and inspect the chain. The agents remain productive because unnecessary human checkpoints are removed.

## Common Mistakes and Trade-Offs in Agent Security

The most frequent mistake is treating an agent’s system prompt as an authorization policy. Statements such as “only access necessary records” help behavior but do not enforce access. Another common error is sharing one service account across agents because it is easier to operate. That practice destroys attribution and allows a compromised research agent to inherit the purchasing or deployment authority needed by a different workflow. Credentials should be isolated at least by agent role, customer, environment, and tool class, even when two agents serve the same application.

Teams also over-enforce human approval. If users must approve every search or internal read, autonomy collapses and alerts become routine noise. The better threshold is reversibility: low-impact, bounded reads can proceed automatically, while external publication, privilege changes, material spend, deletion, and regulated-data transfer usually deserve an explicit checkpoint. Approvals should expire quickly and be bound to the reviewed action, rather than becoming broad permissions that can be reused later. A policy covering “up to 1,000 records” may be unsuitable when one record is a medical result and another is a public event listing.

Other failures include copying user permissions directly into a downstream token, granting agents access to secrets merely so they can decide whether secrets are relevant, and logging complete prompts that contain sensitive data. Full payloads may be needed for investigation, but access, retention, and redaction rules should be explicit. Teams should also avoid unlimited loops: one agent can repeatedly delegate work unless token exchange, hop limits, budget ceilings, and termination conditions are enforced. Finally, authorization tests should include negative cases such as wrong audience, expired task, modified resource, excessive cost, and an attempt to delegate beyond the original grant. A policy that passes only happy-path tests has not been adequately evaluated.

## Alternatives, Standards, and Their Limits

There is no single standard that solves enterprise agent delegation. OAuth 2.0 and RFC 8693 address interoperable token-based authorization, including delegated and exchanged credentials, but they do not reason about whether an AI agent’s dynamic plan is safe. Cedar provides a flexible way to express fine-grained authorization policies and can fit centralized decision points. It still requires reliable identities, correct inputs, and enforcement at every resource that matters. NIST and ISO frameworks help organizations structure governance, risk assessment, and accountability; they are not substitutes for runtime technical controls.

Open authorization proposals and agent-security projects can accelerate experimentation, but draft status matters. A specification submitted to an IETF working group may evolve, may not become a standard, and may not solve all identity or policy problems. Organizations should review implementation support, threat assumptions, interoperability tests, licensing, and maintenance before making a production dependency. WebID-based approaches such as Solid OIDC and WebID-TLS delegation are relevant where decentralized identity and user-controlled data are central, although they introduce their own governance and key-management requirements.

Some teams choose a no-delegation architecture in which a single agent performs research and tool use under one controlled runtime. That can reduce handoff risk for small workflows, but it does not eliminate prompt injection, excessive tool access, or compromised data sources. Others use deterministic workflow engines for routing and allow models only at selected reasoning nodes. This is often the best option for payments, compliance decisions, and production changes because the orchestration layer can enforce exact transitions. Commercial orchestration platforms may offer policy, tracing, and approval features, but pricing and capability vary widely; no responsible generic price can be assigned without knowing agent volume, tool calls, retention requirements, identity integrations, and deployment model.

## When to Act and What Security Should Cost

Act before a multi-agent system handles production data, customer actions, money, privileged infrastructure, or regulated information. A practical trigger is the first workflow that crosses a trust boundary—such as an external website, another organization, or a privileged internal API. Another trigger is the first agent that delegates to a dynamically selected recipient rather than a fixed agent registry. At that point, ordinary API security review may no longer cover the decisions and failure modes created by probabilistic planning.

Early risk can be reduced with an intermediate deployment pattern: fixed agent identities, explicit routing, temporary credentials, read-only data by default, destination allowlists, and human approval for writes. A useful pilot limit is 3 agents, 5 tools, 1,000 tool calls per day, and a 30-day evaluation window, but these are starting numbers rather than universal standards. Organizations should set thresholds based on asset value and recoverability. If exposure is low, short-lived sandboxes may offer better protection than an elaborate approval system. If exposure is high, a $10 monthly token-control tool cannot compensate for a production administrator credential that never expires.

Cost has three components. Identity and policy services may include paid plans, while open-source runtimes and policy languages can reduce license fees but still require engineering and maintenance. Observability rises with every tool call and retained trace. Human review becomes the dominant cost when approval volume approaches thousands of decisions per day, which is a strong reason to reserve checkpoints for high-impact actions. Security testing should include authorization bypass attempts, token replay, prompt injection, credential theft, delegation-chain depth, and emergency revocation. The investment is justified when it reduces the chance and blast radius of unauthorized action, not because every agent workflow needs the largest available governance suite.

## A Defensive Implementation Standard

The definitive answer is to treat every agent handoff as a temporary security-sensitive transaction. The delegating principal, receiving principal, task, target resource, permitted action, expiration, and revocation path should be explicit. Each agent should receive only the identities and tools necessary for its assigned step, with sensitive resources protected by deterministic authorization and isolated credentials. Untrusted content must never receive the same authority as a trusted instruction, and model instructions should be viewed as behavioral guidance rather than the main security control.

A sound architecture combines short-lived credentials, audience-restricted token exchange where appropriate, centralized least-privilege policy, destination and resource allowlists, hop and budget limits, runtime isolation, traceable decision logs, and targeted human approval. Testing must challenge both the policy layer and the orchestration logic, including expired grants, modified resources, recursive delegation, token replay, prompt injection, and attempts to exceed cost or scope. As of October 2026, the market and standards surrounding agent authorization remain active, so teams should prefer well-supported controls over assumptions that a draft protocol or vendor label has already solved delegation security.

Delegation is valuable because it allows agents to divide complex work, but distributed authority creates distributed risk. The right objective is not zero delegation; it is bounded delegation with independent checks. Teams that measure that objective in permissions, token lifetimes, delegation depth, approval thresholds, incident-response time, and verified audit records will be better prepared than teams that merely add warnings to a prompt. The mature pattern is controlled autonomy: agents act quickly within clear limits, escalate narrowly when needed, and leave evidence that a human can evaluate.

## Quick answers

### What is the safest way for one AI agent to delegate work to another?

Use a separate, short-lived identity for each agent and bind its permissions to a specific task, audience, resource set, and expiration time. The receiving agent should evaluate the delegation locally rather than inheriting unrestricted authority from the sender. Avoid copying a broad user or service credential across the handoff.

### Does OAuth 2.0 Token Exchange make agent delegation secure?

RFC 8693 provides a standardized way to exchange tokens when authorization crosses security domains or involves a different subject. It does not decide whether the requested task is appropriate, so teams still need least-privilege policy, short lifetimes, audience restrictions, and revocation.

### How many hops should a multi-agent workflow allow?

There is no universal limit, but a fixed maximum such as three or four hops is a useful starting point for experimentation. Recursive or dynamically unbounded delegation should be blocked unless the system can enforce cost, identity, task, and policy limits at every transition.

### Should humans approve every agent tool call?

No. Routine, reversible, low-impact reads can usually proceed automatically when their scope is tightly bounded. Human approval is more appropriate for external transmission, deletion, financial commitments, privilege changes, and other actions that are difficult to reverse.

### Can prompt injection be prevented by agent delegation policies?

Delegation policies can greatly reduce impact by ensuring that injected instructions do not expose unrelated credentials or sensitive resources. They cannot reliably detect every malicious instruction, so model safeguards, content isolation, least-privilege credentials, and runtime controls are still needed.

Canonical: https://tryinterlock.com/knowledge/how_should_teams_secure_delegation_across_multi-agent_ai_workflows.php
Markdown: https://tryinterlock.com/knowledge/how_should_teams_secure_delegation_across_multi-agent_ai_workflows.php/index.md
