Google has introduced Gemini 4 Argon, initially rolling it out to selected cybersecurity defenders ahead of broader availability.
Results in Google’s published benchmark table put Argon ahead of OpenAI’s GPT-6 Astra on several tests and give enterprise buyers an early measure of how Google’s latest model compares with one of its biggest rivals.
Argon also enters with pricing that could make frontier AI more affordable to run at scale.
A model that keeps working through much longer jobs
Google says Argon supports up to 1 million output tokens in one trajectory, compared with its previous 64,000-token limit. A larger ceiling lets the model continue through far longer coding or agent tasks before a run has to be split or restarted. As enterprise AI takes on larger projects, fewer forced stops could become important.
Internal testing included data-center work that freed more than 300 TiB of memory. Engineers also used the model to convert C and C++ software to Rust, including work involving more than 800,000 lines in the Fuchsia Zircon kernel.
Another project replaced 32,000 lines of specialized code in a Rust video decoder and produced a 2.7-times speed gain over an earlier Rust version. Projects at that scale put the million-token ceiling next to the kind of AI infrastructure and software work large organizations may encounter.
Cyber defenders get Argon first
Selected cybersecurity teams are first in line through Google DeepMind’s Fairwind Program. Participants can receive an Argon configuration with fewer restrictions than the standard release, subject to vetting and controls over who can access it.
Wiz has also tested the model through Scan for Good. Argon found a critical vulnerability in health care software that earlier frontier models had missed during that work.
Paid API customers and Google AI Ultra subscribers are next in the rollout, followed by additional developers and businesses using AI in enterprise applications. Early access gives the tech giant time to collect feedback and test safeguards before the new AI model reaches a larger audience.
What eWeek found: Argon topped Astra in most tests at a lower cost
eWeek counted 19 rows in Google’s published Gemini benchmark table where both Argon and GPT-6 Astra report results.Google’s model finished ahead in 14, while OpenAI’s led in four, with one tie.
Price differences are substantial, with Argon’s introductory input and output rates 80% below Astra’s listed prices and its later rates 60% lower.
Comparison | Gemini 4 Argon | GPT-6 Astra |
| Benchmark rows led | 14 of 19 | 4 of 19 |
| Input price per 1M tokens | $2 introductory | $10 |
| Output price per 1M tokens | $10 introductory | $50 |
| Later input price per 1M tokens | $4 | $10 |
| Later output price per 1M tokens | $20 | $50 |
| 100M input + 20M output example | $400 introductory, $800 later | $2,000 |
Price becomes more consequential once AI goes from pilots to daily use. Coding agents and research workflows can burn through millions of tokens over repeated steps, so lower rates can stretch the same budget across more runs, more employees, or longer jobs.
Argon could also change how companies divide work between models. If real-world use shows it can maintain similar performance on a company’s own workloads, teams could assign more high-volume jobs to Argon and reserve Astra for tasks where it performs better.
Enterprise teams should test identical jobs on both models and compare the cost of each successful completion. Total token use, retries, and tool calls should all count toward that comparison. Lower list prices only translate into real savings if the model can finish the work efficiently.
Read more: OpenAI is giving enterprise and developer workflows another upgrade with GPT-6.1 Sol’s stronger coding and computer-use performance.


