
On September 30, Google unveiled Gemini 4 Argon – a flagship AI model for programming, enterprise tasks, and cybersecurity. The model is being tested by a limited group of security specialists as part of the Fairwind program. Google is participating in a voluntary advanced model review program organized by the U.S. government.
A key change is the increase in the maximum output volume from 64,000 to 1 million tokens. Starting price: $2 for 1 million input and $10 for 1 million output tokens; cached input – $0.10. After the introductory period, rates will rise to $4 and $20, with no timeline provided by Google. On DeepSWE v1.1, the model scored 77.9%. On FrontierSWE v2 – 55% compared to 65.5% for GPT-6 Astra and 62.3% for Claude Opus 5.5. In Vals Finance Agent v2 – 65.4%, in AutomationBench by Zapier – 51.3%, in LVBench – 91.7%.
The model is already being used by thousands of Google employees. Agents analyzed data center operations and found ways to free up over 300 TiB of memory; potential savings range from 500 TiB to 1 PiB. Others assisted in porting code from C/C++ to Rust, including the Zircon kernel of the Fuchsia system with over 800,000 lines. In libgav1, Argon reworked 32,000 lines of SIMD code – the Rust version runs 2.7 times faster. In one experiment, the model improved quantum algorithm optimization by 40%. In cybersecurity, Argon was trained to autonomously identify and fix vulnerabilities. In CWE-bench v1 – 68% pass@1, matching GPT-6 Astra and Grok 4.7. Wiz is testing the model in the Scan for Good program. Argon identified a critical vulnerability in medical software that threatened patient personal data. Google is enhancing protection against prompt injection and deviations from user intent. According to Artificial Analysis, Argon has a 15% hallucination rate – the lowest among leading models.





