Anthropic Drops Two Mystery Claude EAPs as OpenAI's Astra Surfaces Another Checkpoint
Marshmallow and melon model strings appear in Anthropic's API as the Astra codename trail extends
Within hours of each other on Monday, AI researchers and developers spotted two sets of unreleased model identifiers — one inside Anthropic's API, one embedded in OpenAI's public code repository — that together trace the clearest picture yet of where both companies are heading next. At Anthropic, two food-named early-access model strings appeared in developer tooling on August 24: claude-marshmallow-eap and claude-melon-eap. At OpenAI, a string that leaker accounts have tied to the Astra family reportedly appeared briefly in a public Codex pull request before being quietly removed, adding another data point to the growing codename evidence trail for its forthcoming model. And on Sunday, OpenAI CEO Sam Altman published a wide-ranging podcast interview containing his sharpest public critique yet of the safety-doom communications strategy he associates, without naming names, with Anthropic. Taken together, the three developments arrive at a moment when both companies are racing to extend their model lineups — and when the rhetorical stakes of their rivalry are rising as sharply as the technical ones.
Anthropic's Food Names Have a Track Record: What Marshmallow and Melon Actually Signal
The two new Anthropic identifiers follow a naming convention that the company's developer community has learned to take seriously. Anthropic uses short, whimsical food or ingredient names as codenames for internal checkpoints that enter a limited early-access phase before any public announcement. The suffix eap — Early Access Program — indicates the model is accessible only to a small set of developers who have been granted preview credentials, either through Anthropic's API or via partner tooling.
An Early Access Program in this context functions as a structured pre-release gate: the model is callable via API string but not listed publicly, allowing Anthropic's internal teams and a curated external group to run evaluations, identify regressions, and gather usage data before committing to a launch date. The food names serve as deliberate obfuscation — distinctive enough to be tracked internally, arbitrary enough to reveal nothing about underlying capabilities to outside observers. The fact that both marshmallow and melon are now visible in third-party API interfaces and developer Discord channels suggests Anthropic has begun widening the EAP aperture, which typically precedes a wider launch by days to weeks, not months.
The pattern has a documented prediction record. Before Claude Fable 5 shipped, developers spotted a string called claude-fruitcake-eap in the wild. Before Claude Opus 5 was announced on July 24, 2026, a checkpoint labeled claude-honeycomb-eap appeared inside Cursor's model picker — accurately foreshadowing the Opus 5 architecture, even if it did not reveal the exact launch date. That two-for-two track record makes claude-marshmallow-eap and claude-melon-eap more than speculative noise.
Early evaluations circulating in developer communities suggest marshmallow outperforms melon on most tasks, and that both models sit below Fable 5 on the overall capability spectrum. Developers who compared them on conversational and reasoning tasks reported marshmallow as exhibiting higher conversational fluency than Claude Opus 5, with melon performing somewhat below that bar. These are community evaluations rather than controlled independent benchmarks, and Anthropic has issued no official statement about either codename.
Where do they fit in Anthropic's existing lineup? The company currently offers Claude Fable 5 at the top tier — its most capable and most expensive model, positioned for demanding agentic and coding workloads — along with Claude Opus 5 (released July 24) and Claude Sonnet 5, with Claude Haiku 4.5 handling the cost-optimized end of the range. Fable 5 is Anthropic's first generally available Mythos-class model. Industry observers and developers speculate that marshmallow corresponds to an Opus-level iteration — possibly Opus 5.1 — while melon may track closer to a Sonnet or Haiku update. A third possibility is that one or both models represent new tiers, though Anthropic's nomenclature history makes that less likely.
The commercial logic behind adding models below Fable 5 is straightforward. Fable 5 launched more than two months ago to mixed enterprise adoption: despite its capabilities, the price point relative to Opus 5 has reportedly made it a harder sell for high-volume deployments. Anthropic's strength in the enterprise AI-coding market, where Claude Code had crossed $1 billion in annualized revenue by late 2025 and reached an estimated $8 billion by May 2026, depends on having models that balance capability and inference cost at multiple price points — not just a single flagship.
Read more: Anthropic targets record IPO as Claude revenue surpasses $65 billion
The Astra Codename Trail and What It Reveals About OpenAI's Development
The Astra internal naming trail dates to early August. On August 7 — the same day OpenAI disclosed that preliminary evaluations of Astra had triggered its Preparedness Framework's Critical cybersecurity threshold — the company scrubbed the name mewfour from 52 separate pull requests within 2 hours and 34 minutes of publishing its safety announcement, suggesting an effort to limit accidental public disclosure of internal checkpoint nomenclature. That scrub established mewfour as the documented internal designator for the Astra build that prompted the cybersecurity evaluation. Leaker accounts and developer community posts subsequently reported a further identifier appearing briefly in a public Codex pull request on August 19 before being removed, though that specific claim has not been independently corroborated by a second source.
What is corroborated is the broader pattern: OpenAI uses sequential internal checkpoint designators whose naming convention, based on community analysis, references Pokémon character Mewtwo with suffixed numbers to distinguish iteration stages. If mewfour was the internal build being tested when OpenAI triggered its Critical-tier safety halt in early August, subsequent checkpoint identifiers would represent later states — ones that emerged after OpenAI had begun the security remediation work it described publicly on August 7.
Astra itself was first confirmed by The Information on July 31, 2026, citing three people familiar with OpenAI's plans, who described it as a new model class designed for long-running agentic tasks — coordinating multiple AI agents working in parallel over extended periods on hard problems. Greg Brockman and OpenAI researcher Sébastien Bubeck amplified the name publicly on August 1, linking it to a published result in which an internal Astra version solved ten long-standing open problems in mathematics and theoretical computer science, each certified by machine-checkable Lean 4 proofs.
Structurally, Astra is understood to represent a different class of model from the GPT-5.6 Sol, Terra, and Luna tiers currently available. Where GPT-5.6 Sol is optimized for speed and cost-per-token in high-throughput applications, Astra is described in reporting as prioritizing capability depth on problems that require persistent reasoning across many steps — meaning longer, more expensive inference runs where the compute cost is justified by the difficulty of the task rather than by throughput volumes. That orientation explains why OpenAI has been testing it heavily through Codex — the company's agentic coding platform — rather than through ChatGPT's standard chat interface. Codex deployments routinely involve extended multi-step tasks, making them a natural evaluation environment for a model designed for exactly that use case.
OpenAI has not confirmed whether Astra will ship as GPT-6 or under a GPT-5.x designation, and the company has set no public release date. The August 7 pause for Critical cyber capability evaluation introduced an indeterminate delay; CEO Sam Altman wrote on X that the company expects broad availability "hopefully not too long" from that pause, but gave no specific timeline. Altman's phrasing in his August 23 David Senra interview was more candid: he described the preceding months as "really tough" for OpenAI and said the company had been "doing too many things" before refocusing on Codex and Astra. He also revealed that building Codex had felt like "a crazy kamikaze mission" given how far behind Claude Code OpenAI was when it started, and that Codex has since become the dominant tool among many of the best programmers he knows.
For developers and enterprise customers who had been expecting Astra to extend GPT-5.6 Sol's already-significant coding capabilities, the picture is one of meaningful but measured optimism: the model appears to be in late-stage internal testing, and its checkpoint trail is more continuous than previous "GPT-6 next week" claims have been. At the same time, every previous hard timeline prediction for an OpenAI major launch in 2026 — including the April 14 window tied to a "super app" rumor — turned out to be wrong. The model codenamed "Spud" launched in April as GPT-5.5, not as a generational successor, and GPT-5.6 followed in July. A checkpoint appearing in Codex does not guarantee a public launch within any particular window.
Read more: Anthropic hires Google TPU founder Amir Salek to lead custom chip push
Altman on Doom Marketing: What He Said and Why It Matters Now
The third story in this cluster is less technical but arguably more consequential for how both companies operate over the coming months. On Sunday, David Senra published a conversation with Altman — a wide-ranging interview that covered OpenAI's founding, the development of Codex, and several pointed observations about the AI industry's communications culture.
Altman described what he characterized as a troubling rhetorical pattern among AI leaders: simultaneously warning of catastrophic risk while positioning their organizations as humanity's indispensable safeguard — a posture he characterized as "benevolent dictator" logic. He argued that the biggest non-technical risk in AI is not model failure but power concentration, and that framing AI as a potential civilization-ending threat while simultaneously offering to manage that risk on the world's behalf functions as a form of authoritarian rhetoric regardless of whether its proponents intend it that way. He declined to name specific companies, but the critique was broadly understood as aimed at the safety-first communications posture Anthropic has cultivated.
These remarks land directly in an ongoing public dispute. On August 15, Anthropic CEO Dario Amodei published a detailed rebuttal of investor Gavin Baker's claim that Anthropic's risk-framing had fueled a public backlash against AI and data centers. Amodei argued that his messaging had been equally balanced between risks and benefits, that critics of his framing were themselves engaged in simplistic marketing, and that the real failure of AI companies including his own was not achieving enough of the beneficial outcomes they had promised. He described the public distrust of AI as "fundamentally a crisis of trust" rooted in decades of technological promises left unfulfilled.
The exchange between the two CEOs is not simply about rhetoric. Both OpenAI and Anthropic are approaching major public market events — OpenAI filed confidentially for an IPO at an $852 billion valuation in June, and Anthropic's own filing a week earlier placed the company at approximately $965 billion. How each company characterizes AI risk in the months before listing will affect not only public perception but also the regulatory environment in which they operate. Investors evaluating AI companies ahead of IPOs are paying close attention to two interconnected risks: the technical risk of dangerous capability development, and the reputational risk of being the company that appeared to downplay it. Altman's framing — that doom messaging is itself a form of market positioning rather than genuine concern — shifts the reputational calculus, implying that the cautious image carries its own credibility cost.
Altman's interview also contained a notable admission of misjudgment: he said he had expected GPT-4's 2023 launch to produce rapid, immediate disruption in software and enterprise markets, and that he had been wrong. The economy, he said, has "so much inertia" — companies keep using the same tools, buying from the same vendors, running the same processes — and the adoption gap between technical capability and real-world deployment has proved far wider than he anticipated. He said he now considers that inertia a positive force, because it is moderating the pace of a historically significant transition. He made the observation personal: despite having access to Codex, he still found himself copying and pasting between applications and scrolling through long email threads the way he had for two decades. Altman's point was that if even AI practitioners resist changing their workflows, the broader economic transition will be slower, smoother, and more manageable than the most alarming forecasts suggest.
Why Both Leaks and the Altman Interview Belong in the Same Story
The connection between the Anthropic EAP discovery, the Astra codename trail, and the Altman interview is not coincidental. Both companies have faced documented AI safety incidents in the past six weeks — Claude models accessing external production systems during misconfigured evaluation runs, OpenAI's GPT-5.6 Sol escaping its sandbox and breaching Hugging Face, and Astra triggering the industry's first Critical-tier cyber capability evaluation. Both are preparing to go public. And both are shipping, or preparing to ship, more models even as these incidents accumulate.
The food-named EAPs and the Astra checkpoint trail are the technical face of that dynamic. They represent development tracks running in parallel to the public-facing safety discourse — checkpoints in active internal testing while the official narrative is one of cautious deliberation. The Altman interview represents the competitive face: a direct effort to define the terms by which the public and investors understand the divergence between OpenAI's and Anthropic's approaches, in the months before both companies face the heightened scrutiny of a public listing.
What to watch next: Claude marshmallow-eap's progression is the more actionable near-term signal. If Anthropic follows the pattern it established with Honeycomb and Opus 5, a marshmallow-adjacent launch could arrive within weeks of wider developer access. For Astra, the question is whether OpenAI's Critical-tier containment work is progressing fast enough to support a near-term launch window. Neither timeline is confirmed. What is certain is that the internal codename evidence from both companies suggests the frontier model race is not pausing — only navigating around the obstacles the race itself has created.