Framing
How operators should think about AI systems#
Before choosing a model or architecture, describe the work you are prepared to entrust to the system. That description gives the team something concrete to build, evaluate, and govern.
Framing determines where responsibility lives. An operator needs to know what success means, which errors matter, and who can change the system when its behavior falls outside those expectations.
I begin with a practical question: what can this system be trusted to do repeatedly under the conditions in which people will use it?
The unit is the system#
Apply the paper's system-level accountability to the workflow in front of you: identify the components, people, and downstream actions covered by its operating commitment.
That system includes models, prompts, tools, data access, state, evaluation, deployment paths, user interfaces, monitoring, governance, and organizational incentives. Its reliability depends on how those parts behave together. A model error can be the proximate cause of an incident; tools, permissions, and review determine how far its consequences travel.
Operators learn to ask:
- What outputs are allowed to cross trust boundaries?
- Where does uncertainty get surfaced or suppressed?
- Which feedback loops are fast, and which are slow?
- What happens when the system is wrong in a way that looks right?
Start with the workflow and follow its consequences through the people and systems that depend on it.
Why this framing matters in practice#
The operating boundary states the commitment: the work and authority entrusted to the system, under stated acceptance criteria and error limits.
Write that commitment in terms an operator can inspect:
- Work and authority: which cases the system handles, which actions it may take, and which decisions remain with people.
- Acceptance criteria: the outcome quality, timeliness, and constraints required for this work.
- Error limits: tolerable failure frequency and severity for a named task mix, exposure, and observation period, including prohibited actions that require immediate containment.
- Owner and intervention: who answers for the workflow, who can interrupt it, and what conditions require review, contraction, or shutdown.
For a support assistant, the commitment might cover classification and draft routing recommendations while a person approves every ticket creation. Improving the explanation prompt can fit inside that boundary. Allowing the system to create tickets independently changes the authority entrusted to it and requires a new decision.
The agency envelope specifies how delegated action is controlled: permissions, horizon, state, delegation/concurrency, reversibility/containment, and oversight capacity. Operations & Governance turns those dimensions into an operating checklist. Controls can change while the commitment remains fixed; record which kind of change you are making.
From experimentation to operation#
Use the paper's four-level evidence ladder to locate what you have observed and what the next decision requires.
- Demonstrated capability: preserve the inputs, configuration, assistance, and successful result. Use them to investigate whether the task repeats beyond the constructed or favorable conditions of the demonstration.
- Repeatable evaluation: record the configuration, task distribution, repeated results, failures, and coverage gaps. Decide whether this evidence and the available controls support a bounded live trial.
- Bounded deployment: record live outcomes, interventions, costs, and exposure under the named tools, permissions, review, and error limits. Use that account to decide whether to maintain or change the commitment.
- Scaled operation: assess performance and costs across the greater volume, duration, variation, and organizational use actually observed. For the next increment, identify the exposure and control requirements that remain untested.
The levels describe the scope of inference. Study design determines confidence within each level. An observational deployment can contain serious selection bias, and a rigorous evaluation can leave live operating conditions untested. Approval requires an accountable judgment about both the evidence and the controls for the proposed work.
Realized value means benefits remain after the full costs of integration, supervision, infrastructure, governance, and failure. Name the level at which you assess it: task, person, organization, market, or society. A bounded deployment can already yield value. Time saved on a task supports a task-level finding; organizational value depends on what happens to that time and the rest of the workflow.
Before a material change, assemble the decision record in Execution. It connects the proposed work, the evidence, the expected value, and the person authorized to decide.
A note on posture#
This handbook takes a working-model posture. Its guidance applies to systems operated under real constraints and imperfect information.
I treat expansion as a choice that needs a reason and evidence. Holding a useful boundary, narrowing authority, or retiring a workflow can also be a sound result. The operator's task is to make that choice explainable and to keep the means of acting on it available.
The sections that follow develop that practice: learning within a boundary, testing additional delegation, and sustaining the controls and economics on which dependable operation rests.