Google Launches Gemini 4 Argon With 1M Output Tokens and Autonomous Vulnerability Patching
Argon leads 12 of 18 benchmarks and uncovers a hospital software flaw, but access stays gated

Google DeepMind launched Gemini 4 Argon on Tuesday, its first new flagship model generation since Gemini 3, introducing a 1-million output token ceiling, an autonomous three-stage vulnerability patching pipeline, and a gated rollout to cybersecurity organizations — while stopping short of announcing any broad public release date. The launch arrives alongside a Bloomberg report that some Google employees question whether Argon's benchmark coding scores reflect real-world everyday performance, a tension that defines the moment for one of the most consequential AI releases of 2026.
A 15.6× Jump in Output Tokens Changes What Agents Can Generate in a Single Call
The most architecturally significant number in Tuesday's launch is the output token ceiling: 1,000,000 tokens, up from 64,000 in predecessor Gemini models — a 15.6× increase that crosses a threshold making a new category of agentic workflow viable.
At 64K tokens, rewriting large codebases, generating exhaustive security reports, or producing full Rust migration diffs requires splitting tasks into chunks and stitching outputs across multiple model calls, each junction a potential point of inconsistency. At 1 million output tokens, many of those workflows can run inside a single inference pass. Google has demonstrated this in practice: Argon agents rewrote the libgav1 open-source video decoder — 32,000 lines of C++ SIMD code — into Rust through multiple rounds of profile-guided experiments, achieving a 2.7× performance improvement over the prior Rust port while producing identical video output. On the Fuchsia operating system's Zircon kernel, Google has shown Argon agents handling migrations of 800,000-plus lines of C++ code.
These are company-claimed figures on company-owned projects, but they are verifiable in principle through Google's open-source repositories, which sets them apart from purely internal benchmark tables. The architecture enabling 1M output tokens is not publicly documented, and no technical paper for Argon has been released. Google stated that when the model has headroom to generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one continuous pass.
Autonomous Vulnerability Patching: Three Stages, One Real Hospital Software Flaw Already Found
The second major capability is more structurally novel: Argon can autonomously find, validate, and patch software vulnerabilities without requiring a human to specify what to look for. The pipeline operates in three stages — generating candidate vulnerability hypotheses, constructing an exploit to validate whether the vulnerability is real and exploitable, then proposing a remediation patch. On CWE-bench v1, which measures a model's ability to remediate security vulnerabilities, Argon ties for first place with a score of 68%.
Before the public launch, Google deployed this capability through the Wiz Scan for Good partnership. Wiz used pre-release Argon to scan hospital software and identified a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide — a finding that earlier frontier models had missed and that was responsibly disclosed before public announcement. This is a documented case of the pipeline operating on real production software with a real security outcome.
The capability is being distributed through the Fairwind Program, a gated initiative giving 650 or more trusted cybersecurity organizations early access. The naming carries its own signal: favorable winds, controlled navigation. A model capable of generating working exploits and patches at speed is a dual-use capability, and Google's gating is an implicit acknowledgment that broad availability without trust controls would create genuine risk for the software infrastructure these same organizations are meant to defend. Trusted Fairwind partners and internal Google teams receive Argon without cyber guardrails so they can leverage its full frontier-level capabilities.
Benchmark Leads, Acknowledged Gaps, and Internal Skepticism
Google's launch materials claim Argon leads on 12 of 18 published benchmarks compared with GPT-6 Astra and Claude Fable 5.1. Leading dimensions include DeepSWE v1.1 at 77.9% (real-world long-horizon software engineering), the Vals Index (knowledge work, GDP-weighted), AutomationBench by Zapier at 51.3% (business task execution), LVBench at 91.7% (long video understanding), and Gray Swan IPI (indirect prompt injection robustness).
All 18 results are company-reported, and no independent reproduction of the full set had been announced at launch. Benchmark comparisons in competitive AI announcements require interpretation: methodology, prompting strategy, and hardware configurations can shift meaningfully between organizations even when benchmark names match. The 12-of-18 figure is a claim to be evaluated against independent testing over time.
Two benchmarks warrant specific attention where competitors hold the lead. On FrontierSWE v2 — the harder of two agentic software engineering suites — GPT-6 Astra leads with 65.5% against Argon's 55.0%, a ten-point gap. On Terminal-Bench 4.0, which measures autonomous command-line task completion, Claude Opus 5.5 leads at 66.4% against Argon's 57.4%. FrontierSWE v2 measures long-horizon agentic software engineering — the kind of sustained development work most relevant to teams deploying coding agents on real projects. The terminal-task gap matters equally for any organization whose engineers run AI agents at the command line.
Bloomberg reported the same day as the launch, citing people with direct access to the model, that some Google employees who have worked with Argon internally say its coding performance on front-end tasks and everyday development does not match the structured benchmark results. The sources are anonymous. Google pushed back directly, telling Bloomberg it would be inaccurate to characterize the model as underperforming in coding, and one employee told Bloomberg there is "large consensus" internally that Gemini 4 is at the frontier. The pattern Bloomberg described — strong benchmark performance, variable real-world generalization — is a recurring finding across frontier AI coding models and cannot be confirmed until independent API access is available.
Read more: Claude Sonnet 5.5 tops agentic coding tests at a lower price as Opus 5.5 leads Terminal-Bench 4.0
How Argon Sits in the Competitive Landscape
Gemini 4 Argon enters a market where Google, OpenAI, and Anthropic have all released frontier-class models in the past year. On the dimensions where Argon leads — cybersecurity benchmarks, knowledge-work performance, business automation, long-video understanding — the advantages are meaningful rather than marginal. On FrontierSWE v2 and Terminal-Bench 4.0, GPT-6 Astra and Claude Opus 5.5 respectively hold durable leads that matter for teams whose primary use case is sustained, agent-driven software development.
The 1M output token ceiling is currently the largest announced output limit in the competitive set. For any organization working on large-scale code migration, long document generation, or single-pass agentic tasks, Argon's output ceiling alone changes the cost-benefit calculation of which model to use. API pricing at launch is $2 per million input tokens and $10 per million output tokens during the introductory period, with a 95% discount on cached input tokens — a structure designed for agentic workflows where the same long context is reused across generation steps. Post-introductory pricing doubles to $4 and $20 per million tokens respectively.
The Rust Migration Capability and Its Infrastructure Consequences
C and C++ power the majority of foundational software infrastructure — operating system kernels, device drivers, embedded systems, network stacks — and memory unsafety in these codebases accounts for the majority of exploited vulnerabilities in production software. The NSA and CISA formally recommended migrating critical infrastructure to memory-safe languages, with Rust as a primary target, in a joint guidance document. Manual migration at scale is prohibitively expensive; a codebase of 800,000 lines would require years of human engineering effort.
Argon's demonstrated ability to handle migrations at that scale — with the libgav1 result showing not just functional equivalence but a 2.7× performance improvement — describes a practical acceleration path for an infrastructure modernization effort the industry has largely treated as aspirational. Google noted that even its own large-scale C++ to Rust migrations are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production, given the criticality of the systems involved. What remains to be validated by external teams is whether migration quality holds on third-party codebases that vary in documentation quality, legacy dependencies, and build complexity.
What Comes Next: Independent Validation and the Broad Release Date
Google stated it is actively engaged in the U.S. government's voluntary process for pre-release model access. The Fairwind gating and the absence of a public availability date together describe a release strategy shaped by the dual-use nature of what Argon can do. Google is not treating this as a conventional product launch with a general availability date; broader access will start with paid API customers and Google AI Ultra subscribers, without a timeline announced.
The next meaningful milestones are independent benchmark reproduction — whether external AI researchers and enterprises produce results consistent with Google's 12-of-18 claim, particularly on FrontierSWE v2 where GPT-6 Astra already leads by ten points — and the first large-scale third-party assessment of the Fairwind vulnerability pipeline on real infrastructure. Whether OpenAI or Anthropic respond to the 1M output token ceiling before Argon reaches general availability may determine whether that ceiling becomes a durable competitive advantage or a feature the field matches within months.