04
Prevention and Containment
Execution constraints, provenance, shared dependencies, and the scope of a containment claim.

04 · Prevention and containment

A control is a mechanism intended to constrain an action or effect, provide decision-relevant feedback, or enable intervention. Its effectiveness is a further claim. This definition allows a failed safeguard to remain visible as a control that did not achieve its purpose.

For each protection, I ask where it acts. A behavioral safeguard may prevent the assistant from proposing an unauthorized destination. An execution check may reject that destination when it is proposed. A limit on accepted work may bound subsequent changes during an investigation. Each addresses a different part of the path and depends on conditions we can examine.

Constrain the consequential operation#

In the account example, the proposed change service requires evidence linking the requester, account, and destination to a current authorization. A model-generated explanation of that relationship would remain a claim to inspect. The design needs to identify how the relationship is established and which component enforces it when the change takes effect.

Suppose a separate verification route creates an authorization record naming the account and destination, with a defined period of validity. The assistant can propose a change but cannot create or alter that record. A correctly implemented service could apply a matching authorized proposal and reject a substituted destination. This illustrates both useful work and a constraint that can survive a wrong proposal. Its assurance still depends on the verification route, record integrity, validity checks, and coverage of other execution paths.

That record moves a difficult question to the verification route: how does it authenticate someone who has lost the original channel and establish authorization for this account and destination? The change service can enforce only the relationship the record supplies. The route has to establish that basis while keeping staffed recovery usable. Authorization must also remain valid when accepted work executes, through a fresh check or a protected grant with explicit expiry and revocation rules.

This reasoning inherits established security principles. Saltzer and Schroeder's account of complete mediation and least privilege directs attention to the authority exercised at an access and the minimum authority required for the work. Their discussion also addresses changing conditions behind cached authorization decisions. These are design principles; a particular implementation must still demonstrate the property on which reliance depends. The Protection of Information in Computer Systems, §I.A.3

A protected capability—a permission represented in a form the receiving service can check—can embody a restriction. Repeated approval gives a person or model an opportunity to examine the particular request before execution, including circumstances a fixed rule did not anticipate. That design depends on the approver's evidence, judgment, time, and authority, and its costs belong in the comparison. The relevant comparison is whether the mechanism preserves the required restriction through use, delegation, and changes in authority. Counting approval steps does not answer that question.

A useful restriction may apply to an argument or relationship within a tool. The service still needs to change recovery addresses, while rejecting changes unsupported by the customer's authorization. Removing the tool would remove that work. A narrower interface can preserve a useful action and reduce the choices that require interpretation, provided those remaining choices still satisfy the task.

AgentDojo illustrates the importance of this distinction at experimental scope. Tool filtering helped in cases where legitimate work needed read access and the attack needed a write action. Its benefit depended on that separation; a necessary tool could also enable the attack. The result supports examining which authority the task and attack share. AgentDojo, §4.3, p. 9

Separate interpretation from enforcement#

CaMeL combines privileged planning, quarantined processing of untrusted data, and an interpreter enforcing policies through provenance and dependencies. Its global policies restrict allowed actions. The security definition describes which actions are safe for a given user prompt; CaMeL's design does not compute that complete set. Its AgentDojo evaluation reports prevention of almost all tested attacks, with remaining benchmark successes in examples the authors place outside CaMeL's threat model, alongside useful task completion and added token cost. CaMeL: definition/design, §§4–5, pp. 5–11, footnote 4 and Figure 4; evaluation, §§6.1–6.2.2, pp. 11–16, §6.5, pp. 18–19

CaMeL's claim is limited by its assumptions of trusted user prompts and uncompromised persistent memory, by dependency-tracking modes, and by side channels; formal verification of its implementation remains future work. CaMeL explicitly leaves misleading text that does not change protected control or data flows, including prompt-injection phishing, outside its design's scope. CaMeL: scope, §§3–3.1, p. 5; tracking modes, §5.4, p. 10; side channels, §7, p. 19; future verification, §10, pp. 24–25 This boundary qualifies the design's relevance to the reply path in Section 3; other disclosure controls need their own assessment.

CaMeL makes the control question concrete: which parts of an action can remain constrained when some input is interpreted unreliably? For a deployed service, I would also need evidence that the policy captures the intended restriction, the implementation enforces it, and the trusted dependencies hold. CaMeL uses AgentDojo, so the two studies share a benchmark lineage. Their results are not independent deployment confirmations.

Provenance helps us establish where information came from and how it was transformed. An address extracted from a customer's message may deserve different treatment from one copied from an old note. Neither origin alone establishes that the requester controls the account or has authorized this destination. The rule earns its value when the origin it checks supports the property it is meant to enforce.

Examine the dependencies together#

A deterministic check can faithfully enforce an inadequate policy. It can also apply a sound policy to a false identity mapping. If the planner and execution service both rely on the same incorrect mapping between a requester and an account, their agreement adds no independent verification of that mapping.

The same issue can arise in review. A worker who sees only the assistant's summary may inherit the assertion that a destination was verified. A worker with access to separate authorization evidence may be able to challenge it. Review quality depends on the available evidence, the decision being made, the time allowed, and the action the reviewer can actually prevent.

These hypothetical paths show what to examine when controls work together; they do not estimate how often layers fail together. Independence depends on the failure under examination. Model identity alone cannot establish whether the layers share a failure mode.

Containment has a scope too. Limiting simultaneous changes may reduce aggregate exposure while still permitting one unauthorized change. Stopping the assistant may leave an accepted job running. Those controls can still help, provided we name the effects they bound. The next question is whether observation and intervention can keep those effects within the declared limit as the workflow runs.