Google announced Gemini 4 Argon on 30 September, presenting a coding benchmark score above those it gave for rival models, 9to5Google reported. The announcement came as employees with direct access to Gemini 4 questioned how well it handles certain coding work, according to Bloomberg. Google has made Argon available to trusted testers and cybersecurity defenders, with broader access described as coming soon.
Key points
- Google gives Gemini 4 Argon a DeepSWE v1.1 coding score of 77.9%.
- Employees with access say it struggles with some coding tasks, Bloomberg reported.
- Google says Argon can produce up to 1 million output tokens and is working on internal code migrations.
- Trusted testers and cyber defenders have access ahead of AI Ultra subscribers and paid API customers.
DeepSWE v1.1 meets coding doubts inside Google
Google puts Argon’s score on DeepSWE v1.1 at 77.9%, alongside 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra. Those are Google’s reported figures for a named coding evaluation. They describe how the models performed on that test, rather than how each would handle every software project brought to it.
The distinction is immediate for Argon. People with direct access to the effort told Bloomberg that Gemini 4 performed well on industry benchmarks but fared less well when employees put it to work; they said it struggled with certain coding tasks. The people requested anonymity because they were discussing an internal matter.
That account is narrower than a verdict on all of Argon’s coding ability. It identifies difficulties with some tasks, while Google’s published score concerns DeepSWE v1.1. The two accounts can sit together: a set evaluation gives a repeatable score, whereas software work can demand changes to a particular codebase and checks that the changed program still behaves correctly.
Google DeepMind had said Gemini 4 was in post-training ahead of release, as AI Affairs reported in its earlier coverage. Post-training is the stage after a model’s main development in which it is adjusted using additional feedback. Google had shifted resources towards smaller Flash models rather than continuing with Gemini 3.5 Pro, TokenPost reported on 27 September.
Fairwind access begins with cyber defenders
Google has made Argon available to trusted testers and cybersecurity defenders through its Fairwind Program. Google says it plans to make the model available without cyber guardrails to trusted defenders and internal Google teams. Google describes AI Ultra subscribers and paid API customers as the intended recipients of broader access.
For vulnerability repair, Google reports a score of 68% on CWE-bench v1, an evaluation of a model’s ability to remedy security flaws, and says Argon tied for first place. That is a separate measure from DeepSWE v1.1: repairing a flaw tests a different task from the broader coding work at issue in the employees’ account.
Google also says it is working on defences against misuse and prompt-injection attacks, in which hostile instructions are placed in material a model encounters while carrying out another request. Its described measures include monitoring the model’s internal activity, isolating sandboxed environments before high-risk work, and stopping execution when monitoring detects behaviour of concern. These provisions matter for developers seeking to give a model access to software systems: generating an answer and allowing that answer to act on a system carry different risks.
One million output tokens and Google’s Rust migrations
Google says Argon’s output limit is 1 million tokens, up from 64K. Tokens are pieces of text a model processes or produces; an output limit governs how much it can produce in a response, rather than how much material it can take in. A higher limit gives Argon room to write a longer result in one run, which Google links to more extended coding and reasoning work.
Inside Google, Argon agents are working on migrations from C/C++ to Rust, including core libraries and the Fuchsia OS Zircon kernel. Google says those large rewrites undergo automated and manual audits, emulation tests and review before production use. The checking is part of the work even when a model can generate a substantial stretch of replacement code.
Updating an old program could be attempted in longer stretches if more of a proposed rewrite arrived in one response, but the replacement would still need the testing and review Google describes before it went into use.
Google offers a more specific account of its work on libgav1, software for decoding video. It says Argon agents replaced 32K lines of SIMD code in an existing Rust port through repeated experiments guided by performance profiles, examining compiler output as they worked. Google reports that the resulting decoder produced identical video output and ran 2.7x faster than that Rust port.