Context Graph vs Agent Sandbox
Execution Isolation Is Not Decision Governance
AI agent sandboxes are becoming standard infrastructure. Managed agents run in controlled environments with isolated files, browser access, tool execution, credentials, and network rules.
That is progress. It is also easy to misread. A sandbox protects the runtime from agent behavior. A context graph protects the business from invalid agent decisions.
The distinction is simple: a sandbox answers where an agent can run. A context graph answers whether the proposed action is valid before it reaches the real world.

The Core Difference
A sandbox is an execution boundary. It constrains the process. It can prevent tenant escape, unsafe filesystem access, uncontrolled network calls, and host-level damage.
A context graph is a decision boundary. It evaluates the proposed action against structured context: scope, policy, temporal validity, applicability, provenance, and prior decision traces.
| Layer | Question | Protects | Does not solve | Artifact |
|---|---|---|---|---|
| Agent sandbox | Can this agent run safely here? | Machine, tenant, filesystem, network, process boundary | Business validity, applicability, decision scope, policy version | Execution log |
| Context graph | Is this proposed action valid now? | Customer, contract, policy, data scope, regulated workflow | Host-level isolation and process containment | Causal decision trace |
Why Sandboxes Still Fail Business Logic
Refund approval
A support agent can run in a sandbox and call only the approved refund API. It can still approve a refund outside the return window, after entitlement changed, or under a regional exception that does not apply to the customer.
KYC screening
A financial agent can access sanctioned-party data and beneficial ownership records safely. The harder question is which jurisdiction, policy version, threshold, and source authority govern this decision right now.
CRM account changes
A revenue agent can update Salesforce from a constrained runtime. It can still assign the wrong owner, expose a scoped field, trigger renewal before legal review, or overwrite a superseded contract state.
The Four-Layer Agent Stack
Production agent architecture now has four distinct layers. Most teams are investing in transport, execution, and reasoning. The missing layer is decision governance.
| Layer | Function | Example primitive |
|---|---|---|
| Transport | Connect agents to tools and data | MCP, APIs, connectors |
| Execution | Run agents in controlled environments | Sandbox, worktree, browser, remote runtime |
| Decision | Validate the proposed action before execution | Decision context graph |
| Reasoning | Propose the next action | LLM, planner, agent harness |
The decision layer is where pre-execution enforcement lives. The agent proposes an action. The context graph checks whether the action is applicable, scoped, current, policy-compliant, and traceable. Only then does execution proceed.
When You Need Both
Sandboxes and context graphs are complementary. A sandbox should contain the agent runtime. A context graph should govern the action that runtime is trying to take.
If the agent only drafts text for human review, a sandbox and audit log may be enough. If the agent mutates systems of record, sends external messages, approves transactions, changes customer state, or touches regulated workflows, the sandbox is not the control point. The decision boundary has to sit before execution.
That is why an accountable agent is not just an autonomous agent in a safer runtime. It is an agent whose actions are validated before they become real.
Research Validation
In June 2026, a systems-research team introduced ActPlane, an OS-level policy-enforcement layer for agent harnesses that intercepts an agent’s system actions before they execute, using eBPF in the kernel. Its motivation validates the distinction on this page: the authors show that tool-call interception alone cannot observe the indirect execution paths an agent can take, so enforcement has to sit below the harness, at the operating system.
The paper then draws the same line this page draws. It locates the policy context with the agent closest to the task, while — in its words — “enforcement must happen at the OS to cover all execution paths.” Enforcement and decision are two different jobs. ActPlane is an enforcement mechanism — a highly capable sandbox that reaches every syscall — but it still has to be told which actions are legitimate. Whether an action is applicable, in scope, current, and policy-compliant for this actor and record is upstream context, not a kernel hook.
That is the division of labor this page argues for. A sandbox — even one as thorough as syscall-complete OS-level enforcement — contains where and how an agent runs. A decision context graph computes whether the proposed action is valid before it runs, and hands that judgment to whatever layer enforces it. The two are complementary: stronger execution boundaries make the decision boundary more valuable, not less.
Source: Zheng et al., “ActPlane: Programmable OS-Level Policy Enforcement for Agent Harnesses,” arXiv:2606.25189, June 2026.
Related memo
A Sandbox Is Not a Decision Boundary
The memo that introduced this frame for managed agents, Codex enterprise deployments, and agent runtime sandboxes.
Read the memo →FAQ
Is an agent sandbox a governance layer?
No. It is an execution isolation layer. Governance requires validating whether a proposed action is authorized, applicable, current, and traceable before execution.
Can a context graph replace a sandbox?
No. A context graph does not isolate processes or protect the host environment. Production systems need both: sandbox for execution isolation, context graph for decision governance.
What is the shortest way to remember the distinction?
A sandbox protects the machine. A decision context graph protects the business.