OpenAI Launches GPT-6.1 Sol Ultrafast at 6x the Cost and 8x Codex Speed
API access is open to all developers, but ChatGPT requires a $500/month Pro 500 plan; Cerebras may be out

OpenAI on Thursday rolled out GPT-6.1 Sol Ultrafast across its API, Codex, and ChatGPT Work — a new speed tier that the company says delivers near-flagship AI performance at roughly 300 tokens per second. In Codex, that represents eight times the standard Sol throughput rate; through the API, developers gain access at six times the standard per-token price. The launch arrived as Day 4 of Codex lead Tibo Sottiaux's 28-day quality-of-life improvement pledge, and it settled the biggest open question left over from DevDay: when, exactly, would the Ultrafast tier that OpenAI had teased on September 29 actually ship. The answer turned out to be ten days later — and the rollout came with a pricing structure that has drawn praise from API developers and criticism from ChatGPT subscribers who found themselves locked out.
Read more: GPT-6.1 Sol's benchmark positioning against Astra
What "Eight Times Faster" Actually Means — and Where the Claim Falls Short
The 8x figure in OpenAI's announcement describes token generation throughput in Codex: how quickly the model produces output tokens once it starts responding. Standard GPT-6.1 Sol runs at roughly 30 to 50 tokens per second across subscription products, a rate that OpenAI raised from about 30 to about 50 tokens per second on Day 1 of the 28-day campaign, four days before Ultrafast launched. Ultrafast brings that number to approximately 300 tokens per second — at least in OpenAI's controlled testing environment.
The 8x number does not account for several things that affect the actual user experience: time to first token, connection overhead, tool execution within agentic sessions, and application-side processing. A small independent pilot by Kingy AI, running 16 attempts on a coding fix and a research brief, measured actual task completion at roughly 3.7 to 3.9 times faster than Standard — a meaningful speedup, but one that is less than half the headline figure. OpenAI's recommendation to use WebSocket connections rather than standard HTTP for Ultrafast agentic sessions reflects this reality: persistent connections reduce per-turn overhead, allowing more of the speed gain to reach the end user. Developers who enable Ultrafast via the API's service_tier: "ultrafast" parameter and still use repeated HTTP cycles may find the overhead absorbs a significant fraction of the throughput advantage.
That said, even a 3.7× real-world speedup is material for the use cases OpenAI is targeting. In incident response, where an on-call engineer is watching a system fail in real time, or in a computer-use agent navigating a browser to complete a task, or in a voice interface waiting for a conversational reply, seconds matter. The business case for Ultrafast is strongest anywhere a user is watching the model work and the cost of their time or attention exceeds the token cost difference. OpenAI's developer community announcement contains full technical specifications for the tier.
The Pricing Arithmetic: Sol Ultrafast as the Stealth Astra Upgrade
API pricing for Sol Ultrafast sits at $12 per million input tokens and $60 per million output tokens — exactly six times the Standard Sol rate of $2 and $10. The seemingly punishing multiplier becomes more interesting when compared against GPT-6 Astra's standard API price: $10 input and $50 output per million tokens. Sol Ultrafast costs approximately 1.2 times what Astra costs at standard speed, while offering performance that OpenAI says is comparable to Astra and speed that far exceeds it.
OpenAI Developer Relations Engineer Dominik Kundel framed it directly: "Intelligence close to Astra level, 8 times the speed of Sol, and the cost is only 1.2 times that of Astra."
The comparison flips when looking at Astra's own Ultrafast tier. Astra Ultrafast is priced at $60 input and $300 output per million tokens — five times the price of Sol Ultrafast. For teams that need both near-flagship intelligence and near-maximum speed, Sol Ultrafast eliminates most of the argument for the $60/$300 Astra Ultrafast tier, except in cases where the marginal quality gap between Astra and Sol matters more than the 5x price difference.
To anchor this in practical terms: a request with 20,000 input tokens and 2,000 output tokens costs about $0.06 on Sol Standard and about $0.36 on Sol Ultrafast. Heavy agentic workloads with large contexts cost proportionally more. Long-context output above 272,000 tokens incurs a further surcharge, with output priced at $90 per million tokens rather than $60. Full pricing details are available in OpenAI's official pricing documentation.
The Infrastructure Question: Nvidia, Not Cerebras
When OpenAI launched its first Ultrafast tier in August 2026 — a preview for GPT-5.6 Sol — the infrastructure story was central to the announcement. Cerebras powered it, using the WSE-3 wafer-scale chip: a processor the size of an entire silicon wafer, with 44 gigabytes of on-chip SRAM and 21 petabytes per second of aggregate on-wafer memory bandwidth. The design eliminates the "memory wall" that constrains GPU inference, where model weights must move from off-chip HBM memory to compute units for each token. On Cerebras hardware, weights reside permanently next to the cores that use them. The resulting speed for GPT-5.6 Sol Ultrafast was 750 tokens per second — a claimed 14 times the standard rate.
GPT-6.1 Sol Ultrafast appears to run on different hardware. Infrastructure research firm SemiAnalysis reported, the same day as the DevDay announcement on September 29, that the new tier runs at low batch size on Nvidia GPUs rather than Cerebras chips. OpenAI has not confirmed or denied the claim. Nvidia's official account replied to SemiAnalysis's post on X with an eyes emoji, which is not a statement. Cerebras had not commented publicly at the time.
The significance of a hardware shift goes beyond inference mechanics. Cerebras, which went public earlier this year, is understood to count OpenAI as a significant fraction of its inference backlog. If GPT-6.1 Sol Ultrafast — a substantially more capable model than GPT-5.6 Sol — does not run on Cerebras, the dependency questions for Cerebras's business become more acute. On the technical side, GPT-6.1 Sol's architecture details are undisclosed; it is possible the model's size or design does not map cleanly to Cerebras's layered partitioning approach, which GPT-5.6 Sol used by distributing layer groups across multiple CS-3 systems. Cerebras capacity scaling and OpenAI's launch timing may also play a role — using Nvidia GPU overcapacity allows a faster rollout while specialized inference capacity ramps.
The slower headline multiplier — 8× rather than 14× — is consistent with low-batch-size Nvidia inference rather than WSE-3 wafer-scale speed. Cerebras's published figures for unrelated open-source models on WSE-3 include 3,000 tokens per second for GPT OSS 120B and 1,850 for Gemma 4 31B, orders of magnitude above what large proprietary models achieve on conventional hardware.
What Developers Get — and What They Are Locked Out Of
Access to Ultrafast follows a split-track structure. Through the API, any developer can enable the tier by setting service_tier to "ultrafast" in a Responses API call, paying the per-token rate on a usage basis. Rate limits default to 1 million tokens per minute at the Build tier, scaling to 4 million at Launch and 40 million at Grow.
The picture inside ChatGPT Work and Codex is different. There, Ultrafast is available only to Pro 500 subscribers ($500 per month, which includes 25 times the Plus usage allowance), eligible usage-based Enterprise customers with the tier enabled by a workspace administrator, and credit-based Edu plan users. ChatGPT Plus ($20 per month) and Pro ($200 per month) subscribers cannot access Ultrafast inside ChatGPT or Codex regardless of how much they pay — not even in exchange for burning through their allowance faster.
That restriction is the source of developer friction. When Ultrafast runs inside ChatGPT, it depletes included usage at eight times the Standard rate, making the Pro 500 allowance equivalent to about three times the Plus allowance at Standard speed. Some developers who tested the tier under XHigh effort reasoning settings reported burning approximately one percent of their Pro 500 allowance in minutes. The burn rate is a real constraint, not just marketing fine print.
The counterargument — raised by several developers in OpenAI's developer community forums — is that OpenAI should let any paid subscriber elect to burn their allowance faster in exchange for higher speed. If the usage pool is finite regardless, access gating by tier rather than by user choice reads as an artificial restriction rather than a capacity limitation. Since the API has no such gate, the argument has some force: the technical constraint is on per-token cost, which the 6x API premium already addresses.
Read more: OpenAI's 28-day improvement pledge for Codex
Where Sol Ultrafast Sits in the Inference Speed Landscape
The competitive frame for Ultrafast is not primarily other AI providers' current models — it is OpenAI's own product tiers. For teams already using Sol Standard who need faster responses, the upgrade path is clear and comparably cheap. For teams considering Astra Standard, Sol Ultrafast offers comparable speed at a lower price unless the intelligence gap — which GPTS24 previously reported as 1.3 to 3 benchmark points depending on task — materially affects their use case.
Compared to other frontier models, independently measured throughput data suggests Gemini 3.5 Flash runs at roughly 167 to 180 tokens per second and Gemini 2.5 Flash at around 204 tokens per second — fast models, but below Sol Ultrafast's claimed 300 tokens per second. Anthropic has not announced an equivalent ultra-low-latency tier for Claude Sonnet 5.5, which otherwise matches Sol Standard's per-token pricing at $2 input and $10 output per million. The absence of a Sonnet Ultrafast tier leaves a gap that Sol Ultrafast occupies. These comparisons are imperfect across providers because latency depends on model size, hardware allocation, request shape, and geographic routing, and the figures for competing models were not collected under equivalent conditions.
The comparison that matters most for OpenAI's developer ecosystem is between Sol Ultrafast and the prior GPT-5.6 Sol Ultrafast preview. The earlier tier claimed 750 tokens per second on Cerebras hardware, nearly 2.5 times faster than what Sol Ultrafast delivers. The raw speed step-back is real. The argument that this generation's Ultrafast is more valuable despite being slower — because GPT-6.1 Sol is a substantially more capable model than GPT-5.6 Sol — is plausible but rests on capability improvements that are themselves partially self-reported. Independent benchmark data from BenchLM places Sol at 71.9% against Astra's 73.2%, a gap of 1.3 points — context for evaluating the trade-off between speed and intelligence at different price levels.
The 28-Day Campaign and What Comes Next
Sol Ultrafast's release as Day 4 of the 28-day improvement sprint frames it as one milestone in a longer sequence. The first three days delivered a baseline speed increase for all models, API infrastructure simplifications, the Decisions API (a lightweight classification and routing layer built on the Luna model), and GPT-6's first appearance in the main ChatGPT interface. Days 5 through 28 remain unannounced.
What OpenAI has yet to address publicly is the infrastructure question SemiAnalysis raised. Whether GPT-6.1 Sol Ultrafast will eventually migrate to Cerebras or another specialized accelerator, whether the current Nvidia-based deployment is intentional or transitional, and whether a true hardware-architecture speed ceiling exists for Sol at low batch size — none of this has been clarified. The 14× Cerebras benchmark from August remains the reference point for what Ultrafast architecture can achieve; Sol Ultrafast's 8× on what appears to be conventional GPU hardware leaves open whether a genuine speed upgrade arrives in the remaining 24 days or later this year. For developers building applications where every 50 milliseconds of latency changes user behavior, that question is not academic.