Appendix — References
Overview#
This appendix curates a small set of papers and projects across foundational areas that repeatedly show up in AGI-adjacent system design: world modeling, memory, reasoning & planning, embodiment, autonomy, and safety & alignment.
How to use this appendix#
The appendix grounds claims about autonomy, feedback, and system reliability in existing technical work. The main paper states its requirements and hypotheses in Sections 00–06; this appendix maps the surrounding research for readers who want to explore it.
Evidence behind this revision#
I use these selected studies and research syntheses to assess capability, operating performance, measurement, and economics. Four of the nine sources are METR studies of technical tasks and measurement; broader inference requires evidence across institutions and operating settings. Sources were checked on September 6, 2026; the dates below identify the publications, updates, and relevant observation periods.
METR, Task-Completion Time Horizons#
Updated May 8, 2026. Measured task horizons have advanced on selected software, machine-learning, and cybersecurity tasks. Horizons express task difficulty in human completion time at a specified success rate. The tasks are self-contained and well specified. METR regards estimates above 16 human-hours as unreliable with this suite; transfer to ordinary work requires separate testing.
Brynjolfsson, Li, and Raymond, Generative AI at Work#
Manuscript revised November 6, 2024; journal issue May 2025; rollout mainly autumn 2020–winter 2021. A field study estimates increased customer-support resolutions per hour on average, with substantial differences across workers. It concerns one firm's assisted workflow, with people retaining final discretion. The staggered rollout requires causal identification assumptions; the estimate is bounded by that design and setting. Full-cost returns and broader delegation require additional evidence.
METR, We are Changing our Developer Productivity Experiment Design#
February 24, 2026; follow-up experiment began August 2025. METR's early-2025 randomized experiment found that AI access slowed experienced open-source developers working in familiar repositories. The later update reports increasing nonparticipation and withheld tasks because developers might be assigned to work without AI. The authors believe later tools likely accelerated work; selection effects and concurrent-agent time accounting left the gain's size uncertain. Both results are bounded by their participants, tasks, and tool generations.
METR, Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity#
May 11, 2026; survey fielded February–April. Reported speed gains exceeded reported value gains. Changes in task mix help explain why these measures can differ. The 349-worker convenience sample and low response rates create selection risk. The findings describe participants' perceptions; organizational value requires outcome measurement and cost accounting.
U.S. Census Bureau, Large Firms With at Least 20 Employees Biggest AI Users#
May 26, 2026; data December 14, 2025–May 3, 2026. Reported AI use among U.S. employer businesses was 17–20% in the observed period, with differences by firm size and sector. The measure records use in any business function during the previous two weeks. Wording broadened in November 2025; comparisons with earlier estimates must account for that change.
International AI Safety Report 2026#
February 3, 2026. The report synthesizes evidence of a gap between controlled evaluations and live performance, alongside progress and limits in safeguards. Its focus is general-purpose AI safety. Applying its findings to a workflow requires direct evidence about that system's behavior and operating conditions.
METR, Early work on monitorability evaluations#
January 22, 2026. Stronger agents improved at completing deliberately assigned side tasks and evading detection, while stronger monitors detected more. Access to reasoning traces improved detection in some configurations. The prototype uses a small, specialized task set, with limited testing of strategies available to agents and monitors. Production oversight requires validation in its own operating conditions.
Stanford HAI, 2025 AI Index Report#
Published 2025; comparison November 2022–October 2024. This edition reports a large decline in model inference price at a fixed GPT-3.5-level performance threshold. Dependable-outcome cost and net returns require full-system accounting.
IEA, Key Questions on Energy and AI — World Energy Outlook Special Report#
Published April 16, 2026; historical demand comparison covers 2025. The IEA reports improving energy efficiency per AI task alongside rising data-centre electricity demand in 2025, including AI-focused facilities. The comparison distinguishes unit efficiency from aggregate demand. Whether cheaper inference induced additional use requires causal evidence; dependable-outcome cost and net returns require full-system accounting.
I read these sources as evidence of measured task progress, setting-specific productivity effects, and distinct unit-cost and aggregate-demand trends. Dependable economics across deployments and reliable agency at scale remain open questions in this assessment. Helix requires its own test of whether retained operational learning supports repeated, economical expansion of entrusted work and authority.
Chapters#
| Chapter | Scope | Route |
|---|---|---|
| World models | Predictive representations, latent dynamics, model-based control. | Open |
| Memory | Retrieval, persistence, and memory as an agent-system interface. | Open |
| Reasoning & planning | Decomposition, tool use, search over thoughts/programs, planning loops. | Open |
| Embodiment | Vision-language-action policies, imitation, and robotics benchmarks. | Open |
| Autonomy | Agents in web/software environments, exploration, orchestration, guardrails. | Open |
| Safety & alignment | Preference learning, evaluation, interpretability, monitoring. | Open |