A3
Scenario C: Failure, Recovery, and Trust
A fictional scenario showing containment, downstream repair, and evidence for partial re-entry.

Scenario C: Failure, Recovery, and Trust

The failure happened on an ordinary afternoon.

The team operated a compliance-triage system that classified potentially risky transactions and routed review cases. People retained authority over transaction approval and disposition. The system's grant covered creating and prioritizing cases from supported input schemas.

The workflow owner had defined a stop condition: an unsupported or ambiguous schema could enter a quarantine queue for human review, and any automatic classification from it required pausing the affected path. The on-call operator had authority to do so. A separate sampled review checked classification quality within supported inputs.

A downstream reviewer noticed several cases whose priority conflicted with the underlying records. The reviewer escalated the discrepancy. The automatic schema check had accepted an upstream change that preserved field names while changing their meaning.

The operator paused case creation for that input path and routed incoming records to the quarantine queue. Recommendations on the unaffected source continued. The response included queued and in-flight work: the team stopped pending writes and reconciled uncertain tool results against case identifiers before retrying anything.

Containment reduced new exposure. It left a second task: find and address the classifications already produced.

The linked ingestion, configuration, action, and case records established when the changed source first arrived and which cases it had affected. The team recorded the confirmed misprioritizations separately from cases whose downstream effects were still under review. Transaction decisions had remained with people, but delayed review and displaced attention still needed to be assessed.

The model had interpreted the altered fields consistently with the context it received. The incident account included that behavior, the upstream semantic change, and the input control that had failed to detect it.

Restoring the old model configuration would have left the changed input in place. The team instead restored the previous validated input contract for the unaffected source and kept the changed source under human review. Responders corrected priorities, reconciled duplicate or incomplete cases, and notified the affected teams of the exposure and remaining uncertainty.

Later that day, the owner assigned a control change: semantic contract checks using representative records, including the failure case. Separate cases tested whether those checks could distinguish supported inputs from ambiguous ones. A rehearsal verified that a failed check reached quarantine and that pausing the path stopped further writes.

The re-entry decision was deliberately narrower than the original service.

Automatic triage resumed only on the validated source after evaluation met its existing acceptance criteria and the response owner confirmed downstream reconciliation. The changed source remained in quarantine pending agreement on its meaning and fresh evaluation. The business owner included the additional human review in the cost of continuing service.

The incident record linked the breached limit, source and task mix, affected case count, severity assessment, timeline, containment action, repair work, and re-entry evidence. It named the remaining action owner and the observation that would permit reconsidering the restricted path.

A later review found that the new contract check intercepted another incompatible input, distinct from the original failure case. That established the control's response on one further case. The earlier separate-case tests and this interception left coverage questions open: which semantic changes could still pass, and how often would supported inputs be quarantined unnecessarily? The owner continued labeled sample reviews of accepted and quarantined inputs. Broader authority remained subject to its own reliability and value assessment.

Users had a concrete account of what happened, which work had been repaired, and which limits still applied. Continued reliance could be based on that evidence and the controls available to them.

Recovery had restored useful operation at a narrower boundary. That was the commitment the team could support.