Evidence at a glance

64K 1MEvidence
$2Input/$10Evidence
Input$0.10/ , 95%Evidence
$4Input/$20Evidence
128KEvidence
18 12 , 1Evidence

From 64K to One Million, the Task Shape Changes

Google DeepMind has announced Gemini 4 Argon, the first frontier model in the Gemini 4 generation. It targets long-horizon software engineering, enterprise knowledge work in areas such as legal and finance, and cybersecurity defense. Its most distinctive specification is not a conventional parameter claim or context-window headline, but a maximum single-response output of one million tokens. Earlier Gemini models topped out at 64K, while Claude Opus 5.5, Claude Fable 5.1, and GPT-6 Astra allow 128K.

The difference changes the shape of a task, not merely the length of an answer. A large code refactor, a long legal report, or a multi-step execution workflow could in principle require fewer interruptions, state transfers, and manual reconstructions. The model has room to analyze, generate, revise, and deliver within one continuous trajectory. But one million tokens is an output ceiling, not proof that the model can read one million input tokens. Google has not disclosed Argon’s input context window, so the specification should not be treated as unlimited memory.

The Benchmarks Show Strength in Long Tasks, Not Universal Dominance

In Google’s comparison across 18 benchmarks, Argon leads outright on 12 and ties for first on one. The results that best support its positioning come from long-horizon work. On DeepSWE v1.1, Argon scores 77.9%, ahead of Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%. It ranks first on the Vals Index at 68.9% and on AutomationBench for end-to-end business execution at 51.3%, well above Opus 5.5’s 42.5%. Argon also reaches 19.6% on the Harvey Legal Agent Benchmark, versus 5.4% for GPT-6 Astra, and sets a new result on LVBench at 91.7%.

Those numbers cannot be reduced to a claim that Argon beats every rival everywhere. On FrontierSWE v2, Argon scores 55.0%, behind GPT-6 Astra’s 65.5%. On Terminal-Bench 4.0, its 57.4% trails Claude Opus 5.5 at 66.4%, while on OSWorld-2.0 computer use it scores 69.2%, below GPT-6 Astra’s 72.6%. The more useful reading is a division of labor: Argon appears particularly suited to tasks that must preserve a goal over a long trajectory, produce substantial intermediate work, and close a business workflow. That does not make it universally more reliable in terminal work or computer operation.

Pricing Turns Long Trajectories into a Calculable Bet

Argon’s pricing reinforces this positioning. The introductory rates are $2 per million input tokens and $10 per million output tokens, with cached input tokens receiving a 95% discount to $0.10 per million. After the introductory period, the rates rise to $4 for input and $20 for output. At the current discounted rate, a full one-million-token response costs $10, rising to $20 later. This does not make every call cheap. It makes a long task trajectory easier to price against a workflow that would otherwise require multiple models, turns, and human handoffs.

Artificial Analysis reported that Argon matches GPT-6 Astra on its Intelligence Index while costing roughly 60% as much per task at discounted prices. Such comparisons still depend on task definitions, output length, and the agent harness, so they do not directly establish enterprise total cost of ownership. For technical leaders, the key question is whether a longer trajectory actually reduces retries, state synchronization, code review, and human intervention. If one million tokens only produces more unverified material, a lower unit price simply creates a larger waste stream.

For Cyber Defense, the Ability to Stop Matters as Much as the Ability to Act

Argon has been trained to autonomously find, validate, and patch critical software vulnerabilities. It ties for first on CWE-bench v1 at 68%, although the competing models on that leaderboard run inside their own agent harnesses, so the score cannot be separated from the execution environment. A more concrete example comes from Wiz’s Scan for Good initiative: Argon identified a critical vulnerability in healthcare software used by hospitals worldwide, which Google says previous frontier models had missed. The case suggests that the value of a long trajectory may lie less in producing more prose and more in sustaining a complete find-validate-patch loop.

Cybersecurity is also the setting where capability ceilings cannot be evaluated without control ceilings. Google says it is strengthening defenses against cyber and CBRN misuse, monitoring indirect prompt injection, watching for misalignment between reasoning and actions, and preserving the ability to stop execution. It is also using sealed, isolated sandboxes for high-risk training and evaluation. Argon is being released in phases through a U.S. government voluntary pre-release access process, with unguarded cyber capabilities available to trusted defenders and internal Google teams. Ordinary developers should not read that restricted deployment as a default, guardrail-free security agent.