Gemini 4 Argon: 1M Token Output Limits and Pricing
Google announced Gemini 4 Argon with a 1-million-token output cap via Long Decode Continuation, introductory API pricing at $2 per million input tokens, and top engineering benchmark scores.
Gemini 4 Argon delivers a 1-million-token output limit and matches the top score on the Artificial Analysis Intelligence Index at 53. Announced by Google on September 30, 2026, the model marks the debut of the Gemini 4 generation, raising output limits from the previous 64K tokens to handle sustained agentic coding, security remediation, and context-heavy execution.12

Rollout strategy and the Fairwind Program
Google initiated access for Gemini 4 Argon by releasing the model without cyber guardrails to trusted defenders through the Fairwind Program. This deployment allows verified security teams to use the model's full frontier capabilities for vulnerability analysis and defense.1
Subsequent rollout expands access in phases. Priority is allocated to paid API customers and Google AI Ultra subscribers before access broadens to general developers, enterprise tiers, and consumers.3
External cybersecurity teams have already applied the model in production. Wiz used Gemini 4 Argon through its Scan for Good initiative to locate a critical vulnerability exposing personal information in healthcare software deployed across hospitals globally.1
How Long Decode Continuation achieves 1 million output tokens
The increase from the prior 64K-token cap to 1 million output tokens relies on an API mechanism called Long Decode Continuation. Rather than attempting an uninterrupted stream that risks timeout failures or degraded attention, Long Decode Continuation pauses generation across sustained tasks and resumes responses across subsequent API calls.12
Independent evaluation harness limits vary depending on implementation. While Artificial Analysis achieved 1 million tokens using the continuation feature, independent benchmarking platform Vals lists the maximum output limit at 262K tokens for standard evaluation setups.2
API pricing tiers and context caching economics
Argon enters the market with a 50% introductory discount, pricing input tokens at $2 per million and output tokens at $10 per million. Standard post-introductory rates rise to $4 per million input tokens and $20 per million output tokens, though Google has not announced an end date for the introductory phase.12
For workflows reusing large system instructions, context windows, or codebases, context caching shifts the operational math. Cached input tokens receive a 95% discount off the base input rate, reducing recurring input overhead to a fraction of standard cost. Teams comparing infrastructure options across proprietary endpoints and self-hosted runtimes can evaluate deployment costs in our self-hosted AI agents guide.1
Argon vs GPT-6 Astra and Claude Opus 5.5 benchmarks
On standard software engineering evaluations, Gemini 4 Argon holds leads in agentic coding and automation benchmarks, while trailing frontier models on specialized scientific terminal tasks.124
On DeepSWE v1.1, Argon sets a state of the art at 77.9%, outperforming Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%. On Zapier's AutomationBench measuring core enterprise execution, Argon ranks first with 51.3%, alongside a 77.5% score on AutomationBench-AA.12
Long-context and multimodal benchmarks reflect similar strengths. Argon achieved 91.7% on the LVBench long-video evaluation and 84.2% on the GraphWalks 256K–1M long-context benchmark. It tied for first place on CWE-bench v1 with a 68% vulnerability remediation score, reached 70% on CyberBench proof-of-concept tasks, and recorded 100% on IOI 2024–2026.123
Public arena tracking places Argon first in Text Arena with an Elo rating of 1525, while it sits at eighth in Code Arena WebDev with an Elo of 1679. In interactive web development, it built 30 Vibe Code Bench applications perfectly, compared to 25 for Opus 5 and 24 for Astra. On PostTrainBench, Argon scored 45.3%, improving over Gemini 3.1 Pro's 21.99%.2
However, GPT-6 Astra retains an advantage on difficult execution environments. On FrontierSWE v2, Astra scored 65.5% against Argon's 55.0%. Astra also leads on Terminal-Bench Science, scoring 68.1% compared to Argon's 57.6% (which improved from 19.0% on Terminal-Bench 4.0).24
Task expenses and real-world agentic execution
Evaluation expenses vary wildly across benchmark types. On Artificial Analysis standard tasks, Argon costs an average of $1.99 per task under introductory pricing, compared to $3.26 for GPT-6 Astra. Argon averages 62K output tokens per task on these evaluations, more than double Astra's 27K.2
On the independent Vals Index covering 41 models, Argon ranks first overall with a 68.9% score, averaging $15.68 per task while consuming about a quarter of Claude Sonnet 5.5's output tokens. However, on deep agentic execution such as CUA-bench evaluations, Argon's operational cost reached $193.78 per task. Practitioners assessing return on investment across multi-step execution loops can review pricing architectures in our AI agents for business pricing breakdown.245
Argon also changes hallucination trade-offs. On the AA-Omniscience test, Argon recorded a 15% hallucination rate compared to 51% for GPT-6 Astra. The trade-off is lower absolute accuracy: Argon answered 50% correctly compared to Astra's 63%, choosing refusal or conservative replies over ungrounded speculation.2
Google's internal infrastructure engineering demonstrates how these long-horizon agentic workflows operate in production. Working on Google's open-source libgav1 video decoder, Argon agents replaced 32K lines of SIMD code by running rounds of profile-guided experiments, generating safe Rust code that compiled into an implementation running 2.7x faster than the earlier Rust port. Teams can contrast this agentic workflow with developer tools covered in our GitHub Copilot Computer Use report.1
Argon agents are also migrating C/C++ codebases to Rust across Google at scale, progressing from libraries like re2 up to 800K+ lines for the Fuchsia Zircon kernel. In data centers, agent teams analyzed profiling telemetry to autonomously free over 300 TiB of memory, with expected savings between 500 TiB and 1 PiB. In quantum computing, Argon optimized subroutine spacetime resources, beating published baselines by 40% in minutes.1
Alongside frontier releases like Argon, Google maintains its open-weight Gemma family. Our release tracking data shows google gemma-4-12B launched on 2026-05-23 with 12B parameters under the apache-2.0 licence, providing local alternatives for constrained environments.6
For production engineering teams, choosing Gemini 4 Argon depends on output volume requirements and context reuse. If your pipeline requires continuous code refactoring, video analysis, or cached document contexts where a 95% input discount applies, Argon offers superior unit economics under introductory rates. If your workload centers on frontier terminal science or single-attempt accuracy without conservative refusals, GPT-6 Astra remains the stronger baseline.124
Questions
Is Gemini 4 Argon out?
Yes. Google announced Gemini 4 Argon on September 30, 2026. It is currently available to trusted cyber defenders via the Fairwind Program, with access expanding to paid API customers and Google AI Ultra subscribers.
Is Gemini 4 Argon free?
No. Gemini 4 Argon is a commercial frontier model with introductory API pricing set at $2 per million input tokens and $10 per million output tokens, moving to standard rates of $4 input and $20 output.
Where can I use Gemini 4 Argon?
Gemini 4 Argon is accessed through Google's developer API for paid accounts and through the Fairwind Program for vetted cybersecurity defenders, with rollout planned for Google AI Ultra subscribers.
Who can use Gemini Argon?
Initial access is granted to trusted cyber defense partners without cyber guardrails under the Fairwind Program. Wider access prioritizes paid API users and Google AI Ultra subscribers ahead of enterprise and consumer availability.