SpaceXAI Launches Grok 4.7 at Grok 4.6 Prices as Independent Coding Tests Trail Its Own Claims
Musk ranks the AI coding model third behind Anthropic and OpenAI, but independent testers flag heavy token use

SpaceXAI released Grok 4.7 on September 21, 2026, calling it the company's most capable model for coding and knowledge work and pricing it exactly like its predecessor: $2 per million input tokens and $6 per million output tokens. The company says the new model works longer on difficult tasks, checks its own work more carefully, and ships with its best-calibrated safeguards to date, while being served at the same price and speed as Grok 4.6.
Elon Musk moved quickly to frame the launch. In a post on X, he wrote that "Grok 4.7 places @SpaceXAI as third, after Anthropic & OpenAI, for agentic coding," adding that the model's speed and lower cost make it a strong everyday option. Independent evaluations published the same day support part of that claim, but they also show that Grok 4.7's strongest coding results depend heavily on SpaceXAI's own agent software, and that its low per-token price does not automatically translate into a low bill.
Grok 4.7 is available in Cursor, the AI coding company SpaceXAI now owns, and in Grok Build, SpaceXAI's own coding agent. It is also offered through the Grok API under the model ID grok-4.7, through third-party coding harnesses, and through model routers and cloud platforms. GitHub said Grok 4.7 is rolling out gradually in GitHub Copilot for Copilot Pro, Pro+, Max, Business, and Enterprise plans, with support in Visual Studio Code, Visual Studio, Copilot CLI, the Copilot cloud agent, JetBrains, Xcode, and Eclipse.
SpaceXAI also serves a Fast variant that it says delivers twice the output speed at twice the price. According to VentureBeat, Cursor lists Grok 4.7 Fast at $4 and $12 per million input and output tokens and makes Fast the default for Pro and higher plans, though neither SpaceXAI nor Cursor has published a specific tokens-per-second figure for it.
Read more: Grok Bot gains native access to X accounts
What Changed Under the Hood
According to SpaceXAI, Grok 4.7 is built on a new, larger base model than Grok 4.6. It was then trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. The company says the result is a model that is better at verifying its own work and at managing long context.
SpaceXAI also says it trained Grok 4.7 to natively understand the harness behind Grok Bot, its agent product for professional work, which the company says improves conversational tasks and general knowledge work. SpaceXAI's launch post does not state a parameter count for Grok 4.7. Some coverage, including Android Headlines, has reported a base model of about 2.1 trillion parameters, compared with 1.5 trillion for its predecessor, a figure that circulated before launch but does not appear in the company's announcement.
The context window is 500,000 tokens, which Artificial Analysis notes is unchanged from Grok 4.6. Cached input is discounted to $0.50 per million tokens, also matching the previous generation.
The launch came later than Musk had signaled. He had pointed to a release around September 12, then said the model needed more time after the team found that reinforcement learning had penalized response length too heavily, causing the model to give up on hard problems too early and skip careful review of its own answers. The official release arrived nine days after that initial target.
SpaceXAI's Own Benchmarks
The company's launch table compares Grok 4.7 at its xhigh reasoning setting with Grok 4.6, OpenAI's GPT-5.6 Sol Max, and Anthropic's Claude Fable 5.1 Max. All of these figures are self-reported by SpaceXAI.
On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 scored 46.3%, up from 40.4% for Grok 4.6 and ahead of GPT-5.6 Sol Max at 41.7%, but behind Fable 5.1 Max at 51.8%. On DeepSWE v1.1, Grok 4.7 reached 71.0% at high effort, compared with 72.7% for GPT-5.6 Sol Max and 70.0% for Fable 5.1 Max.
The model led the comparison on two specialized tests. It scored 64.0% on EEBench, an electrical engineering benchmark, against 56.4% for Fable 5.1 Max, and 19.6% on the Harvey Legal Agent Benchmark, compared with 6.7% for Fable 5.1 Max and 2.5% for GPT-5.6 Sol Max.
It trailed on others. SpaceXAI lists Grok 4.7 at 38.0% on Terminal-Bench 4.0 for multi-hour terminal work, far behind Fable 5.1 Max at 57.9%, and at 56.7% on HealthBench Professional for clinical reasoning, below GPT-5.6 Sol Max at 60.5% and Fable 5.1 Max at 62.1%.
Independent Tests Tell a More Mixed Story
Artificial Analysis, which runs its own standardized evaluations, gave Grok 4.7 a score of 46 on its Intelligence Index, just two points above Grok 4.6. The index's leaders, Claude Fable 5.1 and OpenAI's GPT-6 Astra, sit at 53. Artificial Analysis said the result places SpaceXAI among the top four AI labs.
The coding picture depends on how the model is tested. When Artificial Analysis paired Grok 4.7 with Grok Build, the model scored 56 on the firm's Coding Agent Index, up nine points from Grok 4.6. Within that index, Terminal-Bench 4.0 rose from 18% to 33%, DeepSWE v1.1 from 65% to 73%, and SWE-Atlas-QnA from 58% to 63%. Among models tested in their native harnesses, Grok 4.7 with Grok Build ranks fourth, behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5.
That ranking explains Musk's "third" claim: two of the three models ahead of Grok 4.7 come from Anthropic, so SpaceXAI ranks third when counted by lab rather than by model.
Artificial Analysis stresses that its Grok Build results are separate from its Intelligence Index, which uses a standardized evaluation harness across models. Under that standardized setup, VentureBeat reported, Artificial Analysis measured Grok 4.7 at roughly 26% on Terminal-Bench 4.0, compared with 59.6% for GPT-6 Astra and about 49% for Claude Opus 5 at maximum effort.
Vals AI, another independent evaluator, evaluated Grok 4.7 across its benchmark suite and reached a less favorable result. In a post on X, Vals said Grok 4.7 ranks 24th on the Vals Index at 54.2%, down 5 points from Grok 4.6, which ranked 14th at 59.2%. Vals attributed much of the decline to finance and research-heavy tasks and noted that, even at the xhigh setting, Grok 4.7 used less than half the reasoning tokens of Grok 4.6 on its suite. Vals also said Grok 4.7 is its best-performing Grok model to date on several legal and medical benchmarks.
Read more: How GPT-6 Astra reset the frontier benchmark race
Knowledge Work Is the Clearest Gain
The strongest independent evidence of improvement comes from long-horizon office work. On AA-Briefcase, Artificial Analysis's private benchmark for multi-hour professional tasks, Grok 4.7 scored 1,657 Elo, a gain of 111 over Grok 4.6 at high effort, placing it just behind Claude Opus 5 and Claude Fable 5.1. Artificial Analysis said the improvement came mainly from analytical quality, which rose from 1,690 to 1,994 Elo, while presentation quality stayed roughly flat.
On GDPval-AA, which asks models to produce practical work products such as documents, spreadsheets, and slides, Grok 4.7 scored 1,695 Elo, up 90 from Grok 4.6. SpaceXAI's own GDPval chart places that result between Fable 5.1 at 1,735 and GPT-6 Astra at 1,542.
Artificial Analysis also found a lower hallucination rate on its AA-Omniscience test, at 29% compared with 34% for Grok 4.6, with accuracy broadly unchanged. Outside agentic knowledge work, however, the firm said Grok 4.7 broadly matches its predecessor, with small gains on some tests and regressions on AA-LCR and AutomationBench-AA.
Cheap Tokens, Expensive Tasks
On list price, Grok 4.7 undercuts its closest rivals by a wide margin. Claude Fable 5.1 and GPT-6 Astra both list at $10 per million input tokens and $50 per million output tokens, making Grok 4.7 five times cheaper on input and more than eight times cheaper on output.
Token consumption narrows that gap. Artificial Analysis found that Grok 4.7 at xhigh used about 81,000 output tokens per Intelligence Index task, compared with 36,000 for Grok 4.6 at high effort and 27,000 for GPT-6 Astra at maximum effort. That is 125% more than Grok 4.6 and 196% more than GPT-6 Astra.
Artificial Analysis puts the average cost per Intelligence Index task at about $3.74 for Grok 4.7 at xhigh and $2.73 at high. VentureBeat noted that the same comparison puts GPT-5.6 Sol Max at about $1.99 per task, despite that model's much higher $4 and $20 list pricing. Vals AI similarly reported that Grok 4.7 costs slightly more per Vals Index test than Grok 4.6, at $4.78 versus $4.34.
Supporters point to a different metric. In SpaceXAI's CursorBench 4.0 cost comparison, highlighted by AI commentator @haider1 and cited by VentureBeat, Grok 4.7 at xhigh scored 46.3% at about $6.01 per task, compared with 46.8% at $7.05 for Fable 5.1 at medium effort and 41.7% at $8.23 for GPT-5.6 Sol Max. The takeaway for engineering teams is that cost per completed task, not the rate card, determines the real bill.
Speed remains a strength. Artificial Analysis measured Grok 4.7's answer output speed at approximately 188 tokens per second on long prompts, though the model averaged about 7.1 minutes per Intelligence Index task because of its heavy token use.
A New Safeguard Stack
SpaceXAI says Grok 4.7 was built with an entirely new safeguard stack and calls it the strongest model it has tested on refusals and jailbreak resistance. The company reports that the model topped LatchBio's biosafety benchmark at 62.4% and that, on HackerBench v0.3, SpaceXAI's internal benchmark for risky and malicious cyber tasks, it allowed only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work.
These are company-reported results and have not yet been independently replicated. SpaceXAI also said it has begun giving select cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defense research.
Harness and Model Are Now Sold Together
The gap between Grok 4.7's Grok Build score and its standardized results points to a broader shift in how frontier AI coding systems compete. SpaceXAI trained Grok 4.7 specifically for its Grok Bot harness and markets it first through Grok Build and Cursor, two products it controls. Anthropic and OpenAI pursue similar strategies by pairing their models with their own coding agents.
For developers, that makes the testing setup part of any benchmark claim. A leaderboard score measured inside a vendor's own agent can differ sharply from one measured in a neutral harness, as Grok 4.7's Terminal-Bench results show. VentureBeat's advice to engineering leads is to benchmark the model on representative internal jobs and track cost per successful completion, token consumption, latency, retries, and required human intervention alongside the API rate.
Grok 4.8 Is Already in the Pipeline
Musk has already moved on to the next release. In posts on X on September 13, before Grok 4.7 shipped, he said Grok 4.8 is a 2.5-trillion-parameter model trained on SpaceXAI's new C++ software stack and that it would finish its initial training that week before moving into reinforcement learning. He has also described Grok 4.9 as likely reaching the class of OpenAI's Astra and Anthropic's Fable models, and has framed Grok 5 as the company's bid for frontier leadership.
SpaceXAI has published no announcement, model card, pricing, or release date for Grok 4.8, and its specifications remain claims from Musk rather than confirmed product details. For now, Grok 4.7 positions SpaceXAI just below the frontier, competing on price and speed while independent testers continue to measure how much of its performance holds up outside its own tools.