Gemini 4 Argon delivers frontier achievement in complex workflows throughout real-world application engineering, endeavor cognition activity akin lawful and finance, and cybersecurity defense.
Listen to article
[[duration]] minutes
This satisfied is generated by Google AI. Generative AI is experimental
Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to prolong profound reasoning throughout complex, long-horizon workflows, Argon is basically changing the way we activity and build at Google. It delivers frontier achievement in complex workflows throughout real-world application engineering, endeavor cognition activity akin lawful and finance, and cybersecurity defense.
Safely releasing frontier capabilities at this flat requires a phased approach. We are actively busy in the U.S. government’s voluntary procedure for pre-release example admission during we gradually develop access. We’ll continue to collect feedback from first testers as we iterate on guardrails before making Argon accessible to developers, enterprises, and consumers as shortly as possible.
Argon volition initiate at an introductory price 1 of $2 per myriad input tokens and $10 per myriad output tokens, alongside cached input tokens priced at 95% off input token price.
Changing how we activity and build at Google
Gemini 4 Argon is already powering our inner workflows, alongside thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and penning quality. It’s assisting teams build faster and shove the boundaries of engineering efficiency and accelerating breakthroughs:
- Quantum algorithmic optimization: Argon is assisting our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck crucial applications. In one example, it attack the published baseline by 40% in a matter of minutes.
- Memory efficiency: A squad of Argon agents analyzed fleet-wide profiling telemetry to autonomously acknowledge and use recollection optimizations throughout Google’s data centers, freeing up complete 300 TiB of recollection formerly rolled out, alongside an estimated 500 TiB to 1 PiB in total savings.
- Large Scale Codebase Migrations and Optimizations: Argon agents are operating on migrating C/C++ codebases to Rust throughout Google—scaling from tens of thousands of lines in center libraries akin re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel. Given the criticality of many of these systems, specified large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and assessment before rolling out to production.
For libgav1, Google's open origin application for decoding video, Argon agents took an existing Rust harbor and replaced 32K lines of SIMD code by operating many rounds of profile-guided experiments, studying the compiler's output, producing harmless Rust so the compiler would vectorize it automatically. The end outcome is a memory-safe video decoder that runs 2.7x faster than the Rust port, alongside identical video output, bringing it nearer to the optimized C++.
Working harder on your most complex problems
To assistance Gemini 4 Argon’s capabilities throughout longer, additional complex use cases, we are considerably expanding the model’s output token bounds to an industry-leading 1M tokens, up from the former 64K tokens. When the example has the headroom to think deeply and create hundreds of thousands of tokens in a sole trajectory, it adds a new flat of degree in reasoning to resolve durable problems in one go.
Enabling coding and endeavor workflows throughout domains
Gemini 4 Argon’s capabilities throughout coding, reasoning, and multimodality and its capability to prolong long, multi-step tasks allow it to excel throughout a range of endeavor workflows.
Google engineers have been using Argon for their regular tasks, from mundane debugging to large-scale codebase migrations and algorithm designs. It sets a new province of the art on DeepSWE v1.1 (77.9%), which measures a model’s achievement in real-world long-horizon application engineering tasks.
Beyond coding, Argon is the foremost example on the Vals Index, which measures financial effect throughout finance, coding, legal, and tax work, alongside all field valued by its contribution to U.S. GDP. We see likewise foremost achievement throughout another domain particular evaluations, akin Vals Finance Agent v2 (multi-step financial research) and Harvey’s Legal Agent Benchmark (legal investigation and drafting). On AutomationBench, Zapier’s benchmark measuring end-to-end implementation throughout center endeavor functions, Argon ranks #1 alongside a mark of 51.3%.
Argon is additionally uniquely powerful whenever cognition activity requires ocular understanding. It’s capable to run expert diagram analysis, acknowledge particulars from lengthy videos, and obtain act according to a sequence of documents. For example, on LVBench, which measures lengthy video understanding, Argon is province of the art alongside a mark of 91.7%.
Leading in protective cybersecurity
To improved equip cyber defenders for the new era of cyberattacks, we trained Gemini 4 Argon to be extremely capable at cybersecurity defense. Argon can autonomously find, validate, and place crucial application vulnerabilities. For trusted defenders and our own inner teams at Google, we’ll be releasing Argon without cyber guardrails so they can leverage its complete frontier-level cybersecurity defence capabilities.
Wiz is already using Argon for cybersecurity defence through its Scan for Good undertaking – a program dedicated to protecting crucial community infrastructure for liberated by finding and remediating high-risk exposures. In an first display of its impact, the example uncovered a crucial exposure exposing delicate individual data throughout healthcare application used by hospitals worldwide, identifying a serious hazard that former frontier models had missed.
On CWE-bench v1, which evaluates the model’s capability to remediate safety vulnerabilities, Argon ties for archetypal location alongside a top mark of 68%, construction on 3.8 Flash Cyber’s frontier achievement on CWE-bench v0.
Gemini 4 Argon demonstrates notable leaps in exposure finding complete 3.8 Flash Cyber. For example:
- On Google’s inner thorough exposure benchmark, Argon uncovered a broad range of exposures throughout complex codebases spanning 20 programming languages.
- On Wiz’s inner black-box intrusion evaluation benchmark, which tests a model’s capability to analyze live web systems without origin code, Argon outperforms 3.8 Flash Cyber in discovering the assault surface, identifying vulnerabilities, and producing proof-of-concept evidence to validate them.
Strengthening frontier safeguards before broad availability
Before rolling out Gemini 4 Argon broadly, we’re continuing to fortify crucial frontier safeguards throughout four chief areas:
Defending against misuse: To forestall bad actors from using Argon for cyber or chemical, biological, radiological, and nuclear (CBRN) attacks, the example is designed to refuse harmful requests during preserving legitimate, dual-use specialized research, as per our Frontier Safety Framework. We are strengthening the robustness of our safeguards for this launch, including improving our techniques to detect the model’s internal activations to place misuse. These safeguards underwent robustness evaluation by inner and external red teams using a blend of manual and automated assault methods.
Defending against immediate injection attacks: Argon is additionally our most resilient example yet against indirect immediate injections, anywhere malicious instructions or environment are used to hijack a model’s behavior. These are complex attacks that necessitate changeless vigilance and multiple layers of defense. Through automated red teaming and adversarial training, Gemini 4 Argon is foremost in immediate injection robustness on the Gray Swan’s Indirect Prompt Injection (IPI) benchmark.
Monitoring for misalignment: In command to forestall Argon from strolling out of limits to try to accomplish a project in a way that goes beyond the user’s intentions, we are deploying misalignment mitigations that detect Argon’s chain-of-thought and actions and halt implementation whenever necessary.
We used a akin scheme to detect our training runs and dispatch alerts to a dedicated event reply team, taking careful precautions against feeding the findings rear into training so as to not hazard shaping Argon’s reasoning to evade our monitoring. We strongly advance the remainder of the industry to maintain reasoning transparency in these pivotal moments of risen capabilities during navigating alignment risks, so that example thoughts remain helpful in identifying and diagnosing misalignment.
Hardening systems: As frontier models develop increasingly capable, safely evaluation them requires safe environments that can keep up alongside the systems themselves. In row alongside our agent authority roadmap, we are hardening our sandboxed environments by isolating and sealing them before high-risk training or evaluations begin. We’re committed to sharing these delegate safety finest practices alongside our partners to enhance safety throughout the ecosystem.
Rolling out soon
We built Gemini 4 Argon alongside frontier-level capabilities on coding, cognition work, cybersecurity defense, and imaginative penning to be a partner for developers, professionals, and enterprises during they tackle the most difficult problems. We’re grateful for the first cohort of cyber defenders and trusted testers whose real-world evaluations and feedback volition assistance us fortify our systems before we publish to developers, enterprises, and consumers, starting alongside paid API customers and Google AI Ultra subscribers.