MIT Researchers Prove Attention Concentration Drives the 1/3 Neural Scaling Law
Attention heads — not the language-modeling output layer — are the training bottleneck in LLMs

For years, the empirical observation that large language model training loss decays roughly as a one-third power of training time — L ~ τ^(−1/3) — has shaped billions of dollars in compute investments without a satisfying theoretical explanation. A new preprint from researchers at MIT argues that the exponent is not a coincidence or a statistical feature of training data, but a structural consequence of how softmax attention heads learn to specialize. Crucially, the team identifies which part of a standard Transformer is responsible: not the final language-modeling output layer, as prior theory implied, but the attention heads deeper in the network.
The work, posted to arXiv on September 26, 2026, as arXiv:2609.32100, is the fourth paper in a series by Yizhou Liu, Sara Kangaslahti, and MIT physics professor Jeff Gore that derives neural scaling laws from architecture rather than from properties of the training data.
Why the 1/3 Exponent Has Resisted Explanation
The scaling law community has known since Kaplan et al. (2020) and the Chinchilla paper (Hoffmann et al., 2022) that training loss in large language models scales approximately as a power law with compute, model size, and dataset size, with exponents near one-third. Chinchilla measured roughly 0.28 for dataset-size scaling and 0.34 for parameter-size scaling — close to one-third, but never derived from first principles.
The dominant explanation pointed to the training data: natural language follows Zipfian frequency distributions, and power-law inputs, the argument went, produce power-law learning curves. The trouble is that this account makes the exponent an accident of dataset statistics rather than a predictable architectural consequence — and it offers no guidance on how to change the exponent, or whether it is fixed by the model design.
The Liu et al. series takes a different approach, deriving scaling laws from the geometry of the training objective and the network's nonlinear operations, without assuming anything specific about the data distribution.
How a Softmax Head Generates a 1/3 Power Law
The core mathematical insight applies to any softmax function trained via cross-entropy to produce a peaked output distribution — one where most of the probability mass concentrates on one or a few options.
In a Transformer attention head, this is the normal end-state of learning. A head that has specialized to detect, say, syntactic dependencies between specific tokens must route nearly all its attention weight to those tokens for the relevant inputs. Producing such a peaked distribution requires the pre-softmax logits — the raw dot-product scores between query and key vectors — to become large in magnitude. Counterintuitively, a softmax distribution cannot be sharply peaked unless the underlying logits are large, because the exponential normalization in softmax suppresses differences between moderately-sized logits.
The researchers show mathematically that the interplay between softmax's exponential normalization and cross-entropy's logarithmic gradient causes logit magnitudes to grow as τ^(1/3) during training, where τ is training time. A quantity they call logit coldness — measuring the sharpness of the attention distribution — grows along the same trajectory. The loss associated with this head therefore decays as the reciprocal: L ~ τ^(−1/3). This result was established for a single softmax head in a prior paper in the series, accepted at ICML 2026. The new September 2026 paper addresses what happens in a real Transformer, which contains many softmax operations: dozens to hundreds of attention heads across multiple layers, plus the final language-modeling head that maps hidden states to vocabulary probabilities.
The Binding Bottleneck: Why Attention Heads, Not the LM Head
In a multi-softmax system, the overall training loss is dominated by whichever softmax component converges most slowly — the binding bottleneck. If all components converge at the same τ^(−1/3) rate, the overall loss also follows τ^(−1/3). But if different components converge at different rates, the slowest one determines the exponent that appears in the empirical training curve.
The question the new paper addresses is: in a real LLM, which softmax is the bottleneck — the attention heads or the final language-modeling head?
The intuitive answer might favor the LM head, which must produce correct probability distributions over tens of thousands of vocabulary tokens. But the researchers argue, and show empirically using the Pythia family of open-source LLMs from EleutherAI, that this intuition is wrong. The LM head's output distribution tends not to become as sharply peaked as an attention head's distribution during training — or its logit magnitudes grow at a different rate — making the attention heads the slower-converging, dominant bottleneck.
The team measures logit coldness across head types at many training checkpoints, available through Pythia's publicly released intermediate states. The empirical fits confirm the theoretical prediction: attention-head logit coldness grows as τ^(1/3) and dominates the convergence rate of the full model.
A Series That Explains Width, Depth, and Time From First Principles
This paper is the fourth in a systematic research program that may be the most ambitious recent achievement in the theory of deep learning. Paper 1, which received Best Paper Runner-up at NeurIPS 2025, derived width scaling (loss ~ m^(−1)) from representation superposition — overparameterized networks packing multiple features into shared activation directions. Paper 2 (February 2026) derived depth scaling (loss ~ ℓ^(−1)) from ensemble averaging across layers. Paper 3, accepted at ICML 2026, proved τ^(−1/3) time scaling for a single softmax system.
Together, the four papers argue that Chinchilla's scaling exponents — near one-third across all three axes — are not independent empirical coincidences but different signatures of the same architectural mechanics, requiring no assumptions about data statistics.
What This Implies for Compute Strategy and Architecture Design
The architectural attribution carries a specific implication: to change the training-time scaling exponent, one would need to change how attention heads develop concentrated distributions. Interventions that have been explored in attention efficiency research — such as entropy regularization to encourage flat attention distributions, logit capping to prevent runaway logit magnitudes, or replacing softmax with alternatives such as sigmoid or linear attention — might in principle alter the exponent, though the paper does not include intervention experiments.
The researchers also note that mixture-of-experts routing gates fall in the same theoretical class as attention heads if their gating networks produce peaked routing distributions — that is, if most tokens are routed to a small subset of experts. If so, MoE training curves would exhibit the same τ^(−1/3) time scaling, a prediction that remains to be tested empirically on large-scale MoE systems.
For practitioners using Chinchilla compute-optimal training strategies, the paper provides theoretical backing for a practical number that previously rested on curve-fitting alone. The one-third exponent is now derivable from the architecture; if your model contains standard softmax attention heads that specialize during training, that exponent should be expected.
Limitations and What Independent Verification Will Require
The paper is a preprint as of this writing and has not yet completed formal peer review, though the prior three papers in the series cleared NeurIPS and ICML. The empirical validation relies on the Pythia suite, a well-regarded open benchmark, but Pythia covers a specific range of model sizes and one training data distribution (The Pile). Whether the bottleneck identification holds across other major LLM families — GPT-style decoder-only models with grouped-query attention, Llama-style architectures with RoPE positional encodings, or models with heterogeneous head configurations — has not yet been independently confirmed.
The theoretical framework also makes a clean assumption that the slowest-converging head determines the overall exponent, which holds if head convergence rates are sufficiently separated. In architectures where multiple heads compete to be the bottleneck, the result could be a weighted or composite exponent rather than a clean one-third.
Independent replication of the logit coldness measurements on other model families would substantially strengthen the paper's claims. Given the series' track record, such replication seems likely to appear in the months following publication.
If the framework is confirmed, the finding would mark a conceptual boundary: neural scaling laws in Transformers are not a reflection of the world's statistical structure, but a signature of the architecture solving its own optimization problem — a mathematical inevitability written into the softmax function and the attention mechanism from the first forward pass.