# Which Three Agent Security Architectures Still Leave Security Unresolved in 2026?

Colton Ramsey · September 28, 2026

> Direct Answer: Three Incomplete Architectures As of 28 September 2026, three commonly proposed agent security architectures still leave important risks...

## Direct Answer: Three Incomplete Architectures

As of 28 September 2026, three commonly proposed agent security architectures still leave important risks unresolved: a centralized agent control plane, a sandboxed execution architecture, and an identity-centric zero-trust architecture. Each model addresses a real part of the problem, but none provides a complete boundary for autonomous tools, delegated authority, changing plans, cross-agent messages, and sensitive data. The control plane centralizes policy and visibility but creates a valuable target and does not automatically stop a confused or compromised agent. Sandboxing restricts what an agent can touch but says little about whether its current objective should be permitted. Identity-centric controls record who launched a task and which service an agent represents, but long-lived credentials can still be misused after authorization.

**Also worth reading:** [How Does Multi-Agent Workflow Automation Actually Function Within Modern Enterprise Architectures?](https://tryinterlock.com/knowledge/how_does_multi-agent_workflow_automation_actually_function_within_modern_enterprise_architectures.php) · [What Are MCP Gateway Security Controls And How Do They Protect Multi-Agent Systems?](https://tryinterlock.com/knowledge/what_are_mcp_gateway_security_controls_and_how_do_they_protect_multi-agent_systems.php) · [How Do Enterprise Security Teams Build an AI Agent Governance Framework Checklist in 2026?](https://tryinterlock.com/knowledge/how_do_enterprise_security_teams_build_an_ai_agent_governance_framework_checklist_in_2026.php)

These architectures are not mutually exclusive. A production system may use all three while still failing because the handoff points between them remain undefined. For example, an authenticated agent can receive a legitimate identity, pass through a correctly configured sandbox, and then invoke a permitted API with a harmful sequence of actions. Conversely, a tightly sandboxed process can be safe at runtime while receiving instructions through an untrusted prompt. The unresolved question is therefore not simply where to put an agent, but how policy should follow identity, intent, context, data, tools, and state across every execution path.

For multi-agent workflow orchestration, the practical answer is a layered architecture with an independent policy decision point, short-lived workload identity, isolated execution, explicit tool permissions, message provenance, state limits, and human approval gates. That is a direction, not a claim that one product or framework solves agent security. The remaining gap is measurable assurance: organizations need evidence that all routes obey the same rules even when agents, vendors, frameworks, and model providers change.

## Architecture 1: A Centralized Agent Control Plane

A centralized control plane is the first common architecture. It maintains inventories of agents, assigns policies, routes tasks, records tool calls, and may mediate every action through a gateway or orchestration service. Its appeal is straightforward: administrators can see which agents exist, determine which model each uses, apply organization-wide rules, and investigate activity from one place. This makes it easier to enforce requirements such as denying access to production systems, restricting particular data sources, requiring approval for external email, or retaining an audit trail. The model also fits emerging shared-security initiatives, including the Blueprint Alliance discussed in 2026, because vendors can agree on common control concepts without requiring every customer to assemble a separate security stack.

The unresolved issue is concentration of authority. If every tool call passes through the control plane, compromise or design failure there can affect many agents at once. The problem is greater when agents can modify plans, create new subtasks, select tools, or issue new instructions to peer agents. A policy engine may approve an individual call as technically valid without recognizing that thousands of individually approved calls collectively create an attack path. Rate limits and budgets help, but they must distinguish among users, tenants, agents, tools, destinations, and action types; a global limit can either block legitimate work or fail to contain one malicious workflow.

A centralized plane also faces latency, availability, and governance costs. An additional network hop and policy evaluation can add milliseconds to fast workflows, while dependency on the control plane can stop all delegated work during an outage. Teams may resist routing highly sensitive execution through a shared service, particularly where data residency or sovereign compute rules apply. The architecture should therefore define failure behavior explicitly: fail closed for privileged tools, but not necessarily for read-only work that can safely continue under cached, short-lived policy. Most importantly, the orchestration service should not be the only security boundary; agents still need independent runtime isolation and least-privilege credentials.

## Architecture 2: Sandboxed Browser, Terminal, and Code Execution

The second architecture isolates agents in controlled browsers, terminals, containers, virtual machines, or microVMs. Browser agents may run in a remote session rather than on an operator’s laptop, while coding agents receive temporary filesystems, restricted shell commands, and access to selected repositories. This approach directly reduces damage from prompt injection, malicious generated code, accidental deletion, and unsafe network access. It supports a zero-install experience and makes it easier to reproduce or destroy an execution environment after a task. Research and product activity around browser-based multi-agent terminals, coding-agent policy enforcement, and kernel-level sentinels all reflect the attractiveness of moving security controls closer to execution.

Sandboxing alone nevertheless leaves several questions unresolved. A process can be technically isolated but still allowed to read production credentials, query an internal metadata endpoint, or send an attacker-controlled payload to the public internet. Strong isolation therefore requires deny-by-default egress, separate secrets, restricted file mounts, CPU and memory quotas, process limits, and tool-specific network destinations. Even then, the host can become weak if an agent can request broader permissions mid-task. A sandbox that starts with access to ten repositories and later gains access to a customer database has crossed an approval boundary regardless of whether both operations were implemented in the same runtime.

The second unresolved issue is semantic safety. Operating-system controls determine what an agent can do, not whether its chosen sequence of allowed actions makes sense. An agent might be permitted to read invoices and send email, yet should not use those permissions to export records to an unrelated address. Sandboxing is also not a complete defense against prompt injection because untrusted text remains capable of influencing a model that holds legitimate tools. Effective systems combine execution isolation with instruction-data separation, constrained outputs, approval for irreversible operations, and independent checks outside the model. In short, the sandbox is necessary for containing consequences, but it cannot be treated as the authority deciding the purpose of a task.

## Architecture 3: Identity-Centric, Zero-Trust Agent Access

The third architecture treats every agent as a distinct workload identity rather than borrowing a human user’s session. Each agent receives short-lived credentials, explicit scopes, a narrow role, and an auditable chain of responsibility. Policy is based on device or workload posture, task context, target resource, data classification, and action risk instead of network location. This is attractive because agents are software actors, not employees, and conventional user access models often fail to represent delegated decisions, tool-to-tool calls, or machine-generated service requests. An identity-first design can distinguish an agent launched by a support user from one launched by a scheduled workflow even if both use the same model.

Identity resolves attribution more readily than intent. Once a credential is issued, an attacker or faulty agent may use it within the permitted scope. OAuth access tokens, API keys, browser sessions, and delegated service accounts can also be copied, cached too long, or exchanged for broader tokens. Strong implementations should bind credentials to the intended workload, audience, session, and cryptographic holder, while preventing one agent from impersonating another during peer messages. Just-in-time issuance and automatic expiration reduce the value of stolen credentials, but they do not establish whether a particular action belongs to the user’s original goal. A signed message can prove origin without proving that its contents are safe or truthful.

The zero-trust model also needs revocation and recovery procedures. If a planner agent detects that a research worker is compromised, it must be possible to stop downstream calls, invalidate caches, revoke delegated tokens, and identify every state object or external side effect that may already have been affected. This becomes difficult when agents communicate through third-party models, external APIs, email, issue trackers, or browser sessions. Cost is another constraint: identity brokers, policy engines, token services, audit storage, and key management add platform expense and operational work. Identity-centric security is therefore a necessary control layer, not a standalone answer. It provides strong attribution and least privilege, but runtime isolation and centralized decisioning still address different failure modes.

## Side-by-Side Comparison of the Three Architectures

The three architectures solve different problems, so choosing only one creates predictable blind spots. The table below compares their primary strengths, unresolved risks, and appropriate role in an agent security architecture. It also identifies where each model should be mandatory.

| Feature | Centralized control plane | Sandboxed execution | Identity-centric zero trust |
| --- | --- | --- | --- |
| Primary control goal | Observe, coordinate, and enforce cross-agent policy | Contain actions and technical compromise | Attribute and restrict each workload |
| Strongest use case | Fleet-wide governance and audit | Untrusted code, browsers, and terminals | API access and delegated service calls |
| Main unresolved risk | Single control-plane target and collective action risk | Allowed actions may still be unsafe or irrelevant | Valid identity may still perform an inappropriate task |
| Typical failure mode | Policy approves each step but misses the sequence | Agent escapes intent boundaries inside a permitted sandbox | Stolen or over-scoped credential is misused |
| Best deployment role | Independent decision and enforcement layer | Mandatory local execution boundary | Mandatory identity for every agent and tool |
| Human approval fit | Central approval for high-risk workflows | Prevent unattended execution beyond a trust boundary | Reauthorize elevated actions for a specific workload |

A mature system uses the control plane to decide, the sandbox to contain, and identity to attribute. It should not make the model itself responsible for all three functions. This division matters because a compromised model or orchestration library can bypass controls designed to live only inside the same process. Independence can be achieved without deploying three separate products: an internal policy service, an isolated runtime service, and a workload identity provider may provide separate trust zones and audit streams.

## Practical Implementation Steps for an Orchestration Platform

Begin with an inventory that records every agent, model, tool, account, data source, destination, and human owner. A reasonable pilot may contain fewer than 10 agents and 20 tool integrations; the exact number depends on the workflow, but small scopes make review easier. Assign stable identities, remove shared credentials, and map each permission to a concrete operation rather than a broad administrative role. Require at least four categories of independent evidence: who initiated the task, which instructions were accepted, which tools were invoked, and which external side effects occurred. Logging model output alone is insufficient because useful evidence also includes prompts, policy decisions, tool arguments, returned data classifications, token audiences, and workflow state transitions.

Next, put untrusted browser and code actions inside disposable runtimes. Use read-only mounts by default, separate development from production credentials, restrict outbound network access, and cap execution time, memory, storage, subprocesses, and monetary usage. Define approval thresholds before deployment: irreversible production changes, external messages to more than a defined recipient group, access to regulated data, new credential issuance, and spending above an agreed dollar or token limit should normally require confirmation. For higher-risk tasks, ask for approval on the proposed action and parameters rather than showing a generic “Allow this agent?” prompt, because users cannot meaningfully approve an action they do not understand.

Finally, test the architecture against ordinary failures as well as attacks. Revoke an agent token, interrupt a tool call, rotate a secret during a task, disconnect the policy service, return malformed tool output, and simulate prompt injection in a web page. Measure median approval and policy-evaluation latency, percentage of calls denied, mean time to revoke access, and the number of external side effects completed after revocation. Security review should occur before any production launch and after every material change to agents, tools, identity providers, or data paths. The platform should remain neutral about models and vendors while enforcing stable rules across them.

## Common Mistakes and Cost Trade-offs

A frequent mistake is calling prompt filtering “agent security.” Content filters can reduce obvious manipulation, but they do not replace authorization, isolation, or audit. Another error is assuming that a more autonomous planner creates better coordination. Additional autonomy can increase tool-call volume, state duration, and the number of trust boundaries, so a five-agent workflow may require more controls than a single agent even if all five use the same model. Teams also underestimate prompt injection through browser pages, documents, repository files, email, and tool results. Any content retrieved for context should be treated as untrusted data rather than an instruction with the same authority as the system task.

Shared credentials remain a serious weakness because they erase attribution and make least privilege difficult. Long-lived API keys, broad cloud roles, and inherited human sessions should be replaced where possible with short-lived, audience-bound credentials. Centralized logging can also become a data-security problem if traces contain secrets or regulated records, so logs need encryption, retention limits, role-based access, and redaction. High availability matters because a security service that fails open exposes every connected agent, while a service that fails closed may interrupt all workflows. Policy classification and risk tiers are more useful than binary trust labels because most actions are neither completely safe nor universally forbidden.

Pricing varies by architecture and is rarely limited to license fees. A small open-source or developer pilot might cost from $0 to several hundred dollars per month, excluding labor, while commercial orchestration, identity, sandbox, and policy services can range from hundreds to tens of thousands of dollars per month. Enterprise deployments may add private networking, dedicated compute, audit retention, compliance review, and incident response, pushing total cost higher. Token usage is only one variable: isolated runtimes consume compute, policy evaluation adds service time, and human approvals can dominate operating cost. Organizations should compare the cost of the workflow with expected loss from credential theft, data exposure, production changes, and external communication. Cheaper agents are not economical if every incident requires manual investigation.

## When Organizations Should Act

Action is justified when an agent can modify data, send communications, spend money, access production systems, execute generated code, or represent an organization to another party. Read-only assistants with no credentials, no external network, and no persistent state form a lower-risk category, but they can still expose supplied documents or leak conversation data. A staged rollout should normally begin with sandboxed research, progress to internal read-only integrations, and only then permit controlled write operations. There is no universal waiting period because exposure depends on permissions and data rather than the novelty of agents. A threshold based on one destructive tool call is more meaningful than a calendar-based promise to review the design later.

Regulated or internet-facing deployments deserve earlier review because they combine sensitive data with external inputs and explicit accountability requirements. Teams should establish ownership before launch: one party approves business purpose, another reviews agent permissions, and a security owner manages the technical boundary. Avoid claiming compliance merely because a tool includes policy enforcement software. Documentation, access reviews, incident exercises, vendor assurance, and evidence retention determine whether controls work in practice. By 2026, interest in shared agent-security architecture is growing, but standards and cross-vendor implementations remain less settled than conventional cloud identity or application security.

The correct decision is not whether to build only a control plane, only a sandbox, or only an identity system. It is whether those controls are independent enough to survive model error and agent compromise, yet connected enough to share evidence and enforce one policy vocabulary. Organizations with powerful tools should begin that integration immediately; teams experimenting with low-impact agents can use the same pilot to define thresholds before expanding authority. Waiting until an agent performs a harmful action confuses risk management with incident response.

## Security Requirements for the Next-Generation Architecture

The next architecture must account for actions rather than merely conversations. It should bind each task to an initiating user, a declared objective, a bounded budget, and an expiration time. Every message between agents needs provenance and type information, while tool results need data labels and instruction-isolation rules. Policy decisions should include the action, resource, identity, context, confidence or risk score, and reason for denial, with deterministic controls for critical boundaries. Models may propose actions, but a separate enforcement component should decide whether a tool accepts them. This separation reduces the chance that prompt injection in the model context also disables the policy logic.

An effective shared standard would also define lifecycle events: agent registration, credential issuance, permission changes, task handoffs, revocation, incident containment, and retirement. It should specify minimum audit fields and retention behavior without forcing every platform to use the same vendor. Interoperability should cover policy exchange, identity claims, approval receipts, and evidence export. Cost and latency budgets matter because safety controls that make ordinary workflows unusable will be bypassed. Targets might include sub-second policy checks for low-risk calls, sub-second revocation propagation for high-risk tools, and zero production credentials in general-purpose sandboxes, although exact values must be set from the organization’s risk and performance requirements.

The three existing architectures will continue to coexist. The central plane provides coordination, the sandbox supplies containment, and identity supplies attribution. What remains unresolved is their assurance under changing agents, generated plans, indirect prompt injection, credential theft, and multi-step actions. The most credible answer is therefore a defense-in-depth architecture in which no single component can authorize, execute, and audit its own unrestricted behavior. That approach does not eliminate uncertainty, but it makes uncertainty observable, limits consequences, and gives human owners meaningful control before agent security becomes an incident rather than a design requirement.

## Quick answers

### What is the safest architecture for enterprise AI agents?

No single architecture is sufficient. The strongest practical design combines an independent policy and orchestration control plane, sandboxed browser or code execution, and short-lived identity for every agent and tool. Human approval should be added for irreversible, regulated, financial, or externally visible actions.

### Can sandboxing alone secure AI agents?

No. Sandboxing can contain generated code, file access, and network activity, but it cannot determine whether a permitted sequence of actions is appropriate. It works best with least-privilege credentials, deny-by-default egress, data classification, policy checks, and approval for high-impact operations.

### How should multi-agent systems handle tool permissions?

Each agent should receive a distinct identity and narrowly scoped, short-lived permission for every tool. Permissions should follow the user, task, resource, and action context rather than being inherited from a shared administrator account. Revocation must also stop downstream agents and invalidate cached or delegated tokens.

### What security risks are unique to multi-agent workflows?

The main added risks are privilege chaining, indirect prompt injection, confused deputies, excessive handoffs, and coordinated actions that are individually permitted but collectively harmful. Independent policy checks, message provenance, cumulative budgets, and cross-agent revocation are therefore more important than treating each isolated agent as a separate project.

### When should a company require human approval for agent actions?

Approval is appropriate when an action is irreversible, externally visible, financial, privileged, or capable of changing production data. The prompt should show the target, intended change, relevant parameters, and expected cost rather than merely asking whether the agent may continue. Read-only research can usually use lower-friction controls when its data and tools are restricted.

Canonical: https://tryinterlock.com/knowledge/which_three_agent_security_architectures_still_leave_security_unresolved_in_2026.php
Markdown: https://tryinterlock.com/knowledge/which_three_agent_security_architectures_still_leave_security_unresolved_in_2026.php/index.md
