GPTS24AI 情报
News · September 30, 2026 · Research · 27 sections · ~1,347 words · 30 views

MIT Researchers Prove Attention Concentration Drives the 1/3 Neural Scaling Law

Attention heads — not the language-modeling output layer — are the training bottleneck in LLMs


Cover image generated by AI