Karpathy: AI Comprehension Is the Next Bottleneck, and Four Formats Can Help
The OpenAI co-founder recommends ASD-STE100, diagrams, HTML pages, and video to cut LLM verbosity

Andrej Karpathy, co-founder of OpenAI and now a researcher at Anthropic, posted four concrete techniques on X on October 2 for making large language model output genuinely readable — framing the problem not as a model deficiency but as a coming bottleneck in how humans consume AI-generated knowledge.
The post opens with a precise warning: "We'll be spending a lot more time trying to understand the outputs of language models." As models grow more capable and generate more text more quickly, the throughput of AI-generated content is increasing faster than human reading capacity. Karpathy's argument is that the primary skill of working with AI is shifting from prompting — getting models to produce — toward comprehension and verification: parsing, validating, and acting on what the model already produced. The four techniques he outlines are practical answers to that shift, and they represent a progression from structured text all the way to passively consumed video.
Why LLMs Are Verbose by Design
The verbosity problem Karpathy is addressing is not incidental. It is a structural consequence of how frontier models are trained. Reinforcement learning from human feedback, the dominant post-training alignment method, builds models by having human raters compare output pairs and indicate which they prefer. Raters consistently favor longer, more thorough answers over shorter ones, even when brevity would better serve the user. The reward model that emerges from that process equates comprehensiveness with quality. The result is output that piles qualifications, background, and caveats onto answers that could be a single sentence — not because the model lacks confidence in the short answer, but because the training signal taught it that longer is better.
Practical mitigations already exist — system prompts instructing brevity, API length parameters, fine-tuned models trained on shorter responses — but each trades something: system prompts constrain globally; length caps cut answers mid-thought; brevity fine-tuning sacrifices nuance. Karpathy's techniques are different: they redirect output format rather than truncate content, using the model's full capability while changing how it is expressed.
ASD-STE100: The Aerospace Standard That Disciplines AI Prose
The first and most technically specific technique is asking an LLM to explain something using ASD-STE100 — a controlled-language specification maintained by the AeroSpace and Defence Industries Association of Europe. The project began in 1979 and the working group was formally established in 1983, when the European aerospace industry set out to standardize aviation maintenance documentation across manufacturers, languages, and training backgrounds. STE100 enforces strict structural rules: sentences are limited to 20 words maximum for procedural steps, each sentence covers one topic only, active voice is preferred over passive, and imperative mood is used for instructions rather than constructions like "it is recommended that the technician should..."
The practical effect on AI output is visible. A prose explanation of how to restart a network router — delivered in a typical LLM style with clauses nested inside clauses — can run to sixty words in a single sentence. The same information structured by STE principles produces five sentences averaging fewer than ten words each. The instructions can be followed without parsing; they read as a sequence of actions rather than a paragraph of reasoning.
When used as a prompt constraint, STE100 rules pressure the model into stripping away the hedging, nested qualification, and synonym-laden elaborations that make LLM prose hard to follow. Karpathy notes he sometimes asks for "80% of STE100," acknowledging that full compliance is impractical. The standard also restricts vocabulary to approximately 900 approved words with a single permitted meaning each — a list the model cannot enforce without explicit access to it. ASD's own FAQ cautions that no tool can replace the standard itself, and notes that AI-generated text can appear structurally consistent with Simplified Technical English while violating the vocabulary requirements. The technique produces noticeably cleaner output. It does not produce certified aerospace documentation.
Beyond Text: Diagrams and Interactive HTML Artifacts
Karpathy's second technique is asking for a diagram rather than a text explanation. The choice reflects a basic cognitive difference: visual parsing — identifying relationships, comparing magnitudes, following spatial flows — operates differently from verbal decoding of dense prose. For topics where structure matters more than sequential argument, a diagram can convey in one pass what a paragraph achieves across multiple readings. Modern LLMs can generate both static and animated graphics with enough reliability to make this practical on demand.
The third technique applies the same logic to document structure rather than visuals. Asking for output "in HTML" shifts the model from prose composition to web document generation. The model draws on its training in web development to produce navigable documents with section headers, collapsible panels, interactive elements, and visual hierarchy. This is technically using the model's code-generation capabilities rather than its essay-writing capabilities — a form of modality switching that trades linear reading for spatial navigation. Karpathy described this category of output in a May 2026 post as prompting models to structure responses as HTML and view the generated file in a browser — documents cheap enough to generate purpose-built for a single question and then discard, rather than consulting general documentation written for a different audience. That earlier post drew significant engagement from practitioners who found the framing useful; the October 2 post extends it into a four-technique system.
The Video Frontier: 3Blue1Brown Style and ElevenLabs Narration
The technique Karpathy describes as most promising is generating a custom explainer video. His specific recommendation is to ask for a video in the style of 3Blue1Brown — the YouTube channel by mathematician Grant Sanderson, known for animations that transform abstract mathematical and computational concepts into spatial, visual processes — with narration produced by ElevenLabs. He also noted that users without an ElevenLabs API key can ask the model to find free local alternatives.
The appeal is cognitive-load reduction taken to its logical endpoint. Reading requires active construction of meaning; watching a well-designed animation with synchronized narration moves closer to passive reception. For concepts involving transformation or multi-step processes, animation can make visible what prose can only describe. Karpathy's characterization is "most bullish" rather than "production-ready," signaling this as a near-term direction more than a fully reliable technique today.
The current capability is real but uneven. Animation accuracy is not guaranteed, and very short videos leave limited room for nuance on complex topics. The pipeline likely involves the model generating animation code and narration separately, with ElevenLabs handling voice synthesis — adding friction compared to a direct text answer. As RuntimeWire noted, Karpathy's post reports no comparative results or measured comprehension improvement; the recommendation is practitioner advice, not a controlled study.
The Cognitive-Load Ladder and What It Signals for AI Interface Design
Karpathy's four techniques map onto a progression of decreasing cognitive effort: constrained text (structured for efficiency, still read linearly) → diagrams (visual parsing, non-linear) → interactive HTML (navigational, user-directed) → video (passively consumed, no decoding required). Each rung trades some precision and density for lower processing cost on the reader's side.
The broader implication is that AI interface design is entering a phase where output format becomes as important as output quality. Most current AI tooling focuses on the prompting side — how to elicit better model outputs. Karpathy is pointing at the consumption side: given that models already produce more text than users can easily process, the next design problem is not generation but comprehension. Answers to that problem at the user level look like his four techniques. Answers at the model and platform level would require changing the training signal so that brevity and structural clarity are rewarded as consistently as comprehensiveness currently is — a more fundamental change that no major lab has announced. Until that changes, the bottleneck will stay where Karpathy identified it.