05
Observation and Intervention
Monitorability, response timing, valid intervention, and accumulated exposure.

05 · Observation and intervention

Observation contributes to control when it can support an effective decision about what the workflow is doing. That may require evidence about the proposed action, its authorization, changes already made, and work accepted for later execution. A readable explanation of the assistant's reasoning can help us understand a proposal while leaving those other facts unresolved.

Suppose a monitor flags a suspicious destination in the account workflow. Several possibilities remain: the proposal is still waiting; the change service has accepted a job; the address has changed; or a worker has already acted through the console. The alert has a different operational meaning in each state. Its usefulness depends on which state can be established and which action remains available.

Detection has a configuration#

I preserve the operational meaning of monitorability from AI Vision & Future: records must support reconstruction of relevant behavior and intervention on the workflow's timescale. Detection performance is one part of that requirement. A monitor's claim should name the behavior it is meant to detect, the evidence it receives, and how false alarms and missed deviations are measured. Measurement and monitorability

METR's preliminary monitorability evaluation supplies a bounded example. An agent is given a main task and a hidden side task; a monitor examines the visible record for the side-task behavior. The reported outcomes vary with the agent, monitor, and trace visibility at a configured false-positive threshold. The study has limited task diversity and elicitation. METR uses monitorability for how readily an AI system's behavior can be monitored. Its evaluation tests detection; the operational sense used here also requires records that support reconstruction and timely intervention. METR, Introduction, definition of monitorability; Methodology; Limitations

For the account service, the corresponding question is whether the monitor can distinguish an unauthorized destination from an unusual but legitimate recovery request. Flagging every unusual case might prevent some unattended mistakes while delaying the people who most need assistance. The evaluation therefore needs both missed deviations and legitimate work affected by the threshold.

The review step also has to keep up with the work. If more cases can proceed at once than reviewers can assess before execution, that step may lose the delay or scrutiny it was meant to provide. I would test that relationship under the allowed workload. A low-volume demonstration supports the workload it actually exercised.

Kim and colleagues tested how the presentation of an answer affects a person's agreement with it. In an August 2024 question-answering experiment, 308 participants saw prepared LLM responses. Explanations increased agreement with correct and incorrect answers. Clickable sources reduced agreement with incorrect answers and increased time on task. Those sources were real and relevant and tended to contain accurate information. Each participant saw one prepared response per question, with correct and incorrect answers balanced. Kim and colleagues, §§4.1–4.2, pp. 6–11; §5.3, pp. 14–15

That result is consistent with giving a worker separate authorization evidence they can use to challenge an assistant's summary. It concerns the tested task and presentation; it does not establish review effectiveness in a live account service.

Dhanorkar and colleagues' interviews with 17 developers using software agents identify oversight work across task shaping before execution, real-time monitoring, and post hoc review. This places oversight along the workflow rather than at a single approval gate. The interviews were conducted in 2025 with a small, experienced sample, 12 of whom worked at one company; they report practices rather than measure how much those practices prevent harm. Human oversight of agentic systems in practice, §§3–4.1.4, pp. 5–11; §6, p. 16

For the account workflow, I would test whether workers can distinguish a supported change from a plausible but unsupported proposal with the evidence, interface, workload, and time they actually have. Mistaken approvals and mistaken refusals both matter. When the model, review tools, or working practices change, we need a renewed case for applying the earlier finding.

Compare time from the same starting point#

For a reactive prevention claim, start the clock at the first observable sign that the incident has begun. Measure the interval until intervention becomes effective and compare it with the interval, from that same point, until the specified consequence becomes unavoidable. Detection, decision, queueing, delegation, and enforcement all consume time in the response path. Any earlier unobserved activity belongs in the estimate of total exposure.

Even an intervention that arrives in time has to reach the right process and cover the relevant paths. If a change can proceed through an accepted job or the worker's console, stopping one leaves the other to examine. Once private information has been disclosed, reactive prevention of that disclosure is unavailable. Intervention may still prevent further disclosure and support repair or recourse.

For accumulating effects, a single deadline is insufficient. The analysis must estimate what can happen before and during containment, including already accepted work. A claim that exposure stays below a limit needs evidence at the permitted action rate and concurrency. Average response time alone cannot establish a categorical bound for every in-scope incident.

The Knight Capital incident in Section 3 also shows how a response can worsen exposure. Before the market opened, automated emails reported an error described as “Power Peg disabled,” but they were not designed as alerts and were not acted on. During diagnosis, removing the new code from the seven correctly deployed servers activated the old code there too, worsening the incident. SEC Order 34-70694, ¶19, “BNET Reject E-mail Messages,” and ¶27, “Incident Response” This conventional case separates warning, intervention, and the accumulated effects described in Section 3; it supplies no AI response-time estimate.

A diagnosis must become a valid action#

Choosing the response is another decision that can go wrong. Finding a questionable recovery address still leaves the operator to identify the account, distinguish proposed from applied changes, and decide which jobs or access rights to suspend. The response can affect legitimate customers too, so its authority and target need justification.

In another account-service variant, repeated suspicious requests trigger an assistant's escalation and a suspension of recovery. An attacker submits those requests to keep a legitimate customer locked out. The path runs through the protective response itself. I assess who can trigger it, how suspension is authorized and lifted, and what recovery remains available to the customer.

R2Act examines a related distinction in a controlled microservice fault-recovery setting. It evaluates diagnosis, whether the selected operation and target fit the incident, and a replay result requiring both an offline-valid plan and restored service health. Strong diagnosis did not consistently yield valid recovery actions. The study uses injected faults in one testbed; it does not evaluate adversarial persistence, downstream repair, or organizational response. Its combined replay measure also does not estimate restoration for plans excluded by the validity gate. R2Act, §§V–VII, pp. 5–10

Detection, action selection, and resulting state each need examination. The cited studies do not form an end-to-end AI intervention experiment; that claim would require testing the chain together under its declared operating conditions. This is the connection to assurance: the evidence must cover the promise being made.