Google introduced Gemini 4 Argon on September 30 as its new frontier model for software engineering, professional work, and cyber defense. It did not arrive with a button for everyone to press. Access has started with a group of selected defenders in the Fairwind Program while Google expands its safety testing.

That choice says as much about the product as its benchmarks do. Google attributes to Argon the ability to sustain long-running tasks, work across modalities, and generate up to one million output tokens in a single trajectory, up from the 64,000 limit of the previous model cited by the company. This is an output limit, not a promise that extremely long responses will always be useful, correct, or economical.

Coding, documents, and long trajectories

In Google's own published evaluations, Argon scored 77.9% on DeepSWE v1.1, which targets long-horizon software engineering, and 51.3% on AutomationBench, which measures end-to-end execution across business functions. It also ranked first on the Vals Index and scored 91.7% on LVBench for long-video understanding.

The company offered more tangible internal examples. Argon agents reportedly helped migrate C and C++ codebases to Rust, including work on Fuchsia's Zircon kernel, and identify memory optimizations across data centers. In another case, an existing Rust port of the libgav1 video decoder became 2.7 times faster after model-guided experimentation. These figures come from Google and its internal environments; they do not establish that outside teams will reproduce the same results with different codebases, tools, and acceptance criteria.

Argon's release is more than a benchmark contest. It tests how much autonomous work a company is willing to expose before it trusts its own control systems.

Cyber capability explains the gate

Google says Argon can find, validate, and patch critical vulnerabilities. The model tied for the top score on CWE-bench v1 at 68%. The company plans to give trusted defenders a version without some cyber guardrails so they can retain the capabilities needed for legitimate defense work.

That design also increases misuse risk. Before a broader rollout, Google says it is strengthening refusals for cyber and chemical, biological, radiological, and nuclear attacks, testing resistance to prompt injection, monitoring for goal misalignment, and isolating execution environments. These are controls described by the company. Their effectiveness still needs to be tested through external and adversarial use.

Argon will carry an introductory API price of $2 per million input tokens and $10 per million output tokens, with cached input priced at a 95% discount. After the introductory period, Google says those rates will rise to $4 and $20. It has not stated how long the lower price will last.

Paid API customers and Google AI Ultra subscribers are next in line, followed by enterprises and consumers, but Google has not provided a public date. That makes an immediate procurement decision based on token pricing premature. Technical leaders will still need latency, operational limits, administrative controls, and cost-per-completed-task data. One million output tokens creates room for longer trajectories; it also creates room for expensive mistakes that take longer to notice.

Source: Google's official Gemini 4 Argon announcement.