Tokens Up 124×, Progress Up 10%: New Analysis Challenges AI Intelligence Explosion Theory
Internal lab data shows AI research gains that are real but nowhere near self-sustaining.

The AI industry has spent the recent two months loudly declaring the arrival of recursive self-improvement — AI systems building better successors that could trigger a self-sustaining cascade of capability growth. OpenAI has made AI-driven research its top organizational priority. Anthropic reports that more than 90 percent of its research tasks now involve AI collaboration or AI-led execution. Alibaba claimed at its cloud summit that Qwen3.8-Max ran 33 automated training jobs over roughly five days, moving its Artificial Analysis benchmark score from 40 to 45.
Against this backdrop, Ramez Naam — a computer scientist who spent 13 years at Microsoft and formulated Naam's Law on solar cost curves in 2011 — published an 8,500-word quantitative challenge on September 27 in Noah Smith's Noahpinion Substack. His conclusion, drawn from internal data released by both OpenAI and Anthropic: the AI self-improvement loop is genuine, but it is operating at roughly one-fifth of the efficiency required to sustain itself — and the gap between current performance and the intelligence explosion threshold is five to ten times, not five to ten percent.
Autonomy Is Not the Same as Acceleration
The first move in Naam's argument is definitional. The industry tends to conflate five distinct levels of recursive self-improvement as if they formed a single runway. Understanding which level has actually been reached — and which has not — is essential to evaluating whether an intelligence explosion is near.
Level one is AI-assisted researchers: humans still direct all the work, but AI tools raise individual productivity. This is widespread today. Level two is AI improving weaker models: AI systems autonomously train or refine models less capable than themselves. This has also occurred, with multiple examples from major labs. Level three — the critical contested frontier — is AI autonomously iterating on peer-tier models, though a human-plus-AI combination still outperforms pure AI. Alibaba's Qwen claims sit somewhere in this territory, though Naam notes they remain unverified by independent parties.
Levels four and five — full AI substitution for human researchers, then fast takeoff where each improvement cycle is cheaper and faster than the last, compounding without limit — have not been demonstrated or approached.
The crucial insight is that the question is not whether AI can act autonomously in a research context. It is whether returns on AI research effort are rising or falling. An AI system that operates without human input but still faces exponentially harder problems at each step is on a treadmill, not approaching a singularity.
The Macro Ledger: Exponential Investment, Linear Output
Before examining the feedback loop's mechanics, Naam assembles a macroeconomic scorecard. Since early 2023, AI chip capacity — measured in Nvidia H100 equivalents by Epoch AI — has grown approximately 127-fold. Roughly 3 percent of US GDP now flows toward AI infrastructure, with capital investment well outpacing what cloud hyperscalers can recoup from operating profits.
Against this backdrop, Epoch AI's Epoch Capabilities Index (ECI) — a composite benchmark aggregating performance across multiple test suites to resist saturation — shows frontier AI advancing at approximately 14 to 16 ECI points per year since reasoning models arrived in late 2024. That is a genuinely fast clip by the standards of any prior technology. But it is a straight line, not an accelerating curve.
OpenAI's own research acceleration report makes the mismatch visible in a single comparison. Across its internal research infrastructure, token consumption per researcher rose 124 times over the measurement period. Code output per engineer grew roughly sevenfold. But the number of experiments per researcher — the metric Naam argues most closely tracks useful research throughput — increased only 1.6 times, reaching an August 2026 peak relative to the 2025 average.
This discrepancy is not evidence that AI tools are useless. It is evidence of process inefficiency: the same underlying research problem requires progressively more tokens, iterations, and compute to make incremental headway. Naam applies a standard power-law decay exponent of 0.2 — derived from studies of research productivity in established scientific fields — to translate the 60 percent experiment increase into an estimated 10 percent improvement in actual research speed.
Anthropic's internal data reaches a compatible conclusion. A survey of 130 technical staff found that Claude Mythos Preview delivered a geometric mean productivity multiplier of approximately 4× — a meaningful number. But Anthropic estimated that doubling its rate of model progress would require productivity gains roughly an order of magnitude larger: approximately 40×. That "4× to 40×" gap is the technical resistance wall.
Anthropic's own phrasing in the Claude Fable 5.1 and Mythos 5.1 system card confirms the picture: internal AI usage had been "a key factor in maintaining the current rate of progress" — not in accelerating it.
Read more: AI Is Already Accelerating AI: Where the RSI Loop Stands Today
The Cunningham Threshold: How Strong Does the Loop Need to Be?
The most rigorous quantitative framing of what an intelligence explosion requires comes not from Naam but from Tom Cunningham, Lukas Althoff, and seven co-authors at the Elasticity Institute, whose paper "The Economics of Recursive Self-Improvement" appeared on arXiv on September 13.
Cunningham's framework models the RSI feedback loop as two linked elasticities. The first: how much does a one-unit improvement in AI capability (one ECI point) translate into higher AI research productivity? The second: how much does higher research productivity translate back into new AI capability gains? For the loop to become self-sustaining — each cycle feeding enough into the next to maintain or increase momentum — the product of these two elasticities must clear a critical threshold.
Using historical data from Stockfish — the open-source chess engine whose decades-long improvement record provides one of the few clean empirical examples of software improving software — the paper calibrates the threshold at approximately 15 percent: a one-ECI-point improvement in AI capability must yield at least a 15 percent increase in AI research productivity for the loop to sustain itself. Naam's own recalibration using the Stockfish data pushes the estimate toward 19 percent, though he acknowledges the difference carries significant uncertainty.
The paper's own estimate, based on Anthropic's 130-person productivity survey, placed current RSI loop efficiency at approximately 9 percent — below the threshold, though close enough to remain in the conversation.
Naam then re-ran the calculation using OpenAI's logged experiment data, which covers a far larger sample than Anthropic's opt-in survey and is directly measured rather than self-reported. If a 9 percent efficiency were accurate, a 16-point ECI gain should have produced roughly a 4× productivity increase. The observed increase was 1.6×. Working backward through the elasticity model, Naam estimated the true current RSI efficiency at 2 to 3 percent per ECI point — roughly one-fifth to one-tenth of the self-sustaining threshold.
Why Narrow Superintelligence Doesn't Generalize
Before examining what the lab data says directly about AI research capability, Naam confronts the strongest evidence on the other side: OpenAI's September 8 announcement that a swarm of approximately 10,000 concurrent AI agents had generated a proof addressing one of the Clay Mathematics Institute's seven Millennium Prize Problems — the existence and smoothness question in the Navier–Stokes equations. The 88-hour effort consumed roughly 130 billion output tokens and represented a genuine scientific milestone.
Naam calls this "narrow superintelligence": AI performing tasks far beyond human capacity within a highly constrained, fully verifiable domain. Mathematics is tractable for AI because every proof step can be verified mechanically, synthetic training problems can be generated without limit, and a Lean proof checker either accepts or rejects a result with no ambiguity. AI architecture research has none of these properties. When a lab wants to know whether a new training technique works, it must run expensive training runs, wait weeks for results, and make judgment calls about whether an improvement generalizes. There is no automated verifier, and no penalty structure for creative failure.
Anthropic made this limitation explicit in the Opus 5.5 system card released on September 22. Claude on open-ended research tasks "mostly tests incremental ideas and prefers less ambitious hypotheses," the company wrote, explaining that training data is dominated by examples of gradual improvement rather than conceptual leaps. On Anthropic's own CoBench benchmark — designed to test AI performance on historical AI research debugging tasks — Opus 5.5 scored 55.8 percent, versus the 85 percent threshold Anthropic has internally defined as necessary for a model to fully substitute for human research staff. The improvement from Claude Opus 5 to Opus 5.5 on this metric was just 2.6 percentage points — within statistical error bounds.
OpenAI's internal data provides the sharpest direct measurement. When measuring the fraction of AI research tasks that an AI agent could complete independently at an 80 percent success threshold, the company found that agents reliably completed only tasks requiring less than 15 minutes of human time. METR's independent task-horizon evaluations, by contrast, show AI agents succeeding at coding tasks that take humans up to 14 hours at a 50 percent success rate, with that horizon doubling roughly every four months. That 16-to-44-times discrepancy reflects the fundamental difference between well-specified, verifiable tasks and the open-ended judgment calls that drive frontier AI research.
Three Acceleration Shortcuts, Three Dead Ends
With the feedback loop running too weakly to sustain itself and AI research capabilities far below what benchmarks suggest, the industry has leaned on three tactical approaches to force faster progress. Naam examines each and finds the same obstacle.
Test-time compute scaling — letting models reason longer before answering — offers genuine but logarithmically diminishing returns. Each doubling of inference budget costs twice as much while delivering a progressively smaller accuracy increment. The approach raises the ceiling on performance for any given model, but the gains slow as the budget grows.
Agent swarms — running hundreds or thousands of AI instances in parallel — improve throughput but face two independent bottlenecks. Oxford philosopher Toby Ord's September 21 analysis of swarm scaling showed empirically that a swarm of N agents achieves N^λ speedup rather than N speedup, where λ falls between 0.48 and 0.68 depending on task type. A tenfold increase in agents delivers a three-to-fivefold throughput improvement at tenfold the cost. The second bottleneck is idea diversity: a study comparing AI output with 467 human participants found that the first ten AI responses collectively produce roughly eight to ten distinct ideas, but each additional AI response beyond that delivers diminishing novel contributions — eventually requiring two AI outputs to match the incremental value of a single human's additional suggestion.
Scaling model size with reinforcement learning remains the most powerful approach and the one Naam respects most. OpenAI's Navier–Stokes proof was built on a base of large-scale reinforcement learning post-training layered on pretraining. But this path also hits diminishing returns: as dense models grow larger, the compute cost per token rises faster than linear, narrowing the economic margin between model improvement and inference efficiency.
The Three-Loop Counter-Argument
The most sophisticated challenge to Naam's single-loop analysis comes from economists Tom Davidson, Basil Halperin, Thomas Houlden, and Anton Korinek, whose NBER working paper constructs a three-loop model of explosive growth. Their framework links three mutually reinforcing loops: a software loop (AI improves training algorithms), a hardware loop (AI accelerates chip design), and an economic loop (AI productivity growth generates capital for both). When all three strengthen together, the aggregate trajectory could clear the explosion threshold even if no single loop does independently.
Naam calls this the most compelling pro-explosion argument he has examined. But he identifies two structural gaps: the software loop calibration may inherit the same overestimate as the survey-based productivity figures, and the hardware loop is constrained by physical manufacturing timelines — better chip architectures still require 18 to 24 months to move through fabrication and reach deployment.
Neither objection is decisive. The three-loop model remains a legitimate description of how explosive growth could eventually occur. What it does not provide is evidence that it is occurring now.
What Would Need to Change
The RSI feedback loop is real. It is not strong enough to sustain itself. Naam identifies three paths that could change the calculation.
The first is external fuel: continued growth in compute capacity, data center construction, and capital investment can push frontier capability higher through raw force even without a self-sustaining software loop. This path maintains the current linear trajectory — but if hardware scaling eventually slows, the AI research tools that labs have built become proportionally more critical to sustaining any progress at all.
The second is qualitative change in the loop itself: better training data, architectural breakthroughs of the scale the Transformer represented in 2017, or AI systems that genuinely develop research intuition rather than optimizing for incremental improvement. Naam explicitly holds this open. Current models are predominantly trained on examples of gradual progress, which may explain the systematic bias toward incremental hypotheses that Anthropic documented in Opus 5.5 — a structural constraint that a fundamentally different training regime could potentially correct.
The third is multi-loop reinforcement: if the software, hardware, and economic loops strengthen each other faster than individual diminishing returns drain them, the composite trajectory could clear the threshold even if no single loop does independently. This is Davidson et al.'s argument; Naam's disagreement is about calibration, not structure.
Naam's contribution is not to rule out the intelligence explosion — he explicitly declines to do that. It is to provide a quantitative estimate of how far current evidence puts the system from a self-sustaining state. That distance is five to ten times the current loop strength, not five to ten percent. Closing it would require changes larger than anything the last several years of frontier AI development have delivered — and the data from both OpenAI and Anthropic confirms that the treadmill is running. It is not yet capable of running on its own.