03 · Flywheel
Overview#
The Flywheel is a reinforcing system: we learn from operating it, carry that learning into changes, and test whether later operation improves. Retained learning supplies the momentum.
Valid evaluation turns observations into evidence. The loop closes when that evidence informs a retained decision or system change and its effects are tested in later operation.
I use the Flywheel to describe improvement within a current operating boundary: the work and authority entrusted to the system, under stated acceptance criteria and error limits. Retrieval, validation, or tool implementation can change within that commitment. Delegating additional work or authority changes the commitment and creates a candidate Helix step.
The operating loop#
Each rotation has six parts:
- Bounded operation:
- A model-assisted workflow acts under defined tools, permissions, review, and acceptance criteria.
- Real-world outcomes:
- Use produces results, errors, interventions, costs, and downstream effects.
- Observation and telemetry:
- The system records relevant inputs, actions, state, outputs, and outcomes.
- Evaluation and attribution:
- Those observations are tested against the intended objective and connected to system behavior with enough confidence to act.
- Retained decisions:
- Evidence informs changes to prompts, tools, retrieval, evaluations, models, or human process, or a documented decision to retain the current configuration.
- Tested performance:
- Later operation tests whether the decision improved or sustained reliability, cost, safety, or usefulness within the same boundary.
Consider a patch-drafting system that repeatedly misses a known compatibility requirement. The team adds validation and tests whether proposals improve. People still approve and execute changes. Successful validation improves the system within the same authority: a Flywheel improvement.
Measurement as the closure mechanism#
Measurement belongs in the product from the beginning. Its records need to help someone understand what happened and decide what to change.
To close the loop, we generally need:
- A stable, decision-relevant definition of success.
- Observations that represent the operating conditions, including difficult and unsuccessful cases.
- Attribution sufficient to distinguish a model error, tool failure, policy failure, interface problem, or human-process failure.
- Evaluations that remain predictive outside the test set and are checked against live outcomes.
- A path from observed failure to a change that can be deployed, compared, and reversed.
Validity means that a measurement supports the decision being made. Adoption, activity, preference, task speed, and outcome quality answer different questions, even when measured precisely. Pre-deployment evaluation can justify a bounded trial; live reliability requires live evidence.
Monitorability requires records sufficient to reconstruct relevant behavior and support intervention on the workflow's timescale. A record that explains an incident afterward may still arrive too late to help someone stop it.
What creates momentum#
Operating assets carry learning into later iterations:
- Evaluation suites preserve known success and failure cases.
- Stable tool and policy interfaces make controls reusable.
- Provenance and incident records preserve why a change was made.
- Operator knowledge becomes explicit, reusable process.
These assets carry the momentum: each iteration builds on prior learning. The claim is that retaining them improves outcomes or reduces the cost of comparable later operation. If valid learning repeatedly yields neither sustained improvement nor lower cost within the current boundary, the local compounding claim weakens. Helix tests transfer across an expanded boundary.
When the Flywheel becomes agentic#
Agentic operation can accelerate the Flywheel; closure still depends on valid measurement and retained learning.
When a system selects and executes actions over time:
- the interval between observation, decision, and action can shrink,
- mistakes can alter the environment and become the context for later steps,
- and action volume can outpace evaluation, monitoring, or intervention.
A faster loop with weak measurement can compound error more efficiently. Section 04 therefore treats agentic scope as an explicit agency envelope whose dimensions can vary separately while their effects interact.
Failure modes and limits#
Common failures that break or misdirect the loop:
- Metric mismatch: the proxy improves while the outcome does not.
- Selection effects: the observed users, tasks, or interventions cease to represent the operating population.
- Feedback contamination: system-influenced outputs are reused as evidence without independent checks.
- Evaluation drift: tests are overfit, stale, or insensitive to new failure modes.
- Tool-interface drift: changing schemas, permissions, or data shapes invalidate earlier results.
- Hidden human correction: untracked review creates apparent reliability that the system cannot reproduce.
- Monitorability gaps: activity is recorded but cannot be reconstructed or reviewed on the required timescale.
Some limits are structural. Learning may stall when correctness is expensive or delayed, feedback cannot be used, errors are irreversible, or the environment changes faster than evaluation can follow.