Gemini 4 Argon: what Google announced and what it means
Gemini 4 Argon is Google's new frontier model, announced on 30 September 2026 at 20:00 UTC (22:00 in Poland). The headline change is answer length: up to 1M output tokens, up from the previous 64K. In the API it will launch at $2 per million input tokens and $10 per million output tokens, then move to $4 and $20. Outside Google, only the US government, selected cyber defenders in the Fairwind Program and trusted testers have it for now. Google has given no date for a wider release.
I have not tested Argon. It is not generally available yet. Google's selected partners and testers have it. This post is a map of the announcement, built from Google's pages, posts on X and a few checked outside sources. The benchmark numbers come from the vendor itself. The benchmarks, the API price and availability each get their own article.
The 30-second version
1M
output tokens in a single answer, up from 64K
$2/$10
per 1M input and output tokens at the introductory price
$4/$20
per 1M tokens once the introductory period ends
95%
off cached input, so $0.10 at launch
BENCHMARKS
Gemini 4 vs GPT-6 Astra vs Claude
Where Argon wins Google's table, and where it loses
API PRICING
API pricing and the 1M-token answer
What one long answer costs, and when caching helps
RELEASE
Who gets Argon first, and when
Fairwind, the paid API, Google AI Ultra and the EU
In short: Google stretched the model's answer to 1M tokens. At launch it priced it the same as GPT-6.1 Sol. For now only the US government, selected cyber defenders and trusted testers have it.
Timeline (Polish time)
The times come from the metadata of Google's post and from timestamps on X. I give them in Polish time (CEST, UTC+2); subtract two hours for UTC.
- 22:00 Google's blog publishes the post by Koray Kavukcuoglu of Google DeepMind: "Gemini 4 Argon: our next era of frontier intelligence".
- 22:02 Sundar Pichai introduces Argon on X. In a second post he writes that the model is with the US government and going to trusted defenders in the Fairwind Program.
- 22:03 Logan Kilpatrick gives the launch price: $2 in and $10 out. Google DeepMind posts about the 1M-token output limit.
- 22:04 Vals AI, which runs the Vals Index, calls Argon its new number one.
- 22:05 The Google AI account repeats that broader availability will come "as soon as possible", with no date.
- 22:30 Arena reports that Argon took first place on its text leaderboard with 1525 points.
- 22:40 Google edits the blog post. Against the Internet Archive copy from 22:06, the text and figures stay the same. The charts become animated and an author bio appears.
- 23:46 A second edit. In the text, "Fuchsia OS Zircon kernel" becomes "Fuchsia Zircon kernel".
I quote the version of the post fetched at 00:04 on 1 October, after both edits.
The day before, at DevDay 2026, OpenAI showed GPT-6.1 Sol among other things. That model costs the same in the API as Argon at launch: $2 in and $10 out. I put together what else OpenAI announced at DevDay, with the prices of its new models.
What Gemini 4 Argon is
Google calls Argon its new frontier model, built for long, multi-step work. The post names three areas: real-world software engineering, enterprise knowledge work (legal and finance, for example) and cybersecurity defense. The model also reads charts, long videos and series of documents. According to Google, thousands of its employees already use it, and its engineers use it for daily tasks.
The official name is Gemini 4 Argon. Unofficial pre-launch leaks on X and tech sites talked about "Gemini 4 Pro", but none of Google's launch pages uses that name. A Google spokesperson told Reuters that Argon is larger than the previous line of Pro models. Google has announced no other Gemini 4 models.
Google has not published an API model ID, a context window or a knowledge cutoff. In Google's materials, inputs of up to 1M tokens appear only in the GraphWalks test, and a test setting is not an official context window.
1M output tokens: what fits, and who checks it
For me this is the number that matters most. Google raises the answer limit to 1M tokens, "up from the previous 64K tokens". The post does not say which model had 64K. The docs for current Gemini 3.x models, such as 3.8 Flash and 3.1 Pro, list 65,536 tokens, which is exactly 64K. GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 stop an answer at 128K tokens (Opus 5.5 reaches 300K only in the Batch API beta). Argon can write almost eight times as much in one answer.
One number does not match. Vals AI, which tested Argon before launch, lists a 262K output limit on its model page and used that setting by default in its tests. Google says 1M, and the reason for the gap is not known yet.
Google DeepMind explains it this way: with that limit Argon reasons more deeply and solves long problems in one go. A whole migrated module or a full report fits in a single answer. Today a result like that takes several calls, or more than a dozen, and each one sends its input again.
There is one more caveat. In current Gemini models, thinking tokens count toward that limit and cost the same as answer text. Google has not confirmed this for Argon yet. If the rule holds, part of the million goes to reasoning alone.
The cost is easy to work out. A million answer tokens is $10 at the introductory price and $20 after it, plus input. Claude Opus 5.5 charges the same $20, but it has to split that text into at least eight answers. I worked out separately what a single million-token answer costs and when caching cuts the bill.
In my view the bottleneck moves from writing to reading. The model writes a million tokens in one pass. Nobody reads that carefully in an afternoon. Even Google does not send such code straight to production. Its large migrations first go through automated and manual auditing, emulation testing and review. Before you pay for long answers, decide who will check them and with what.
Benchmarks in brief
Google published a table with 19 rows and 18 tests, because GraphWalks takes two rows. Argon sits next to GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. I count by row. Argon wins 13, ties one and loses five.
The tie is CWE-bench v1, a test of fixing vulnerabilities in code, where Argon and GPT-6 Astra both score 68.0%. GPT-6 Astra has the top score on FrontierSWE v2, Terminal-Bench Science 0.1 and OSWorld-2.0. Claude Opus 5.5 tops Terminal-bench 4.0 and PostTrainBench. It is strongest in knowledge work and long context. On GraphWalks with inputs from 256K to 1M tokens it scores 84.2%, against 71.8% for GPT-6 Astra. On DeepSWE v1.1, a test of long software engineering tasks, it scores 77.9%.
This is the vendor's table, so I have three caveats. First, Google chose the rivals. On the Vibe Code Bench leaderboard Claude Sonnet 5.5 is first, and it is not in the table. On Google's own CWE-bench v1 chart, Grok 4.7 also scores 68%. Second, on five rows Google computed Argon's score itself and took the rivals' scores from leaderboards or vendors. On OSWorld-2.0 Google took Argon's best of three runs. Third, a Google spokesperson was more careful with Reuters, saying Google sees Argon as "comparable to" Astra and Opus "on key coding and cyber benchmarks".
According to Arena's official account, Argon leads its text leaderboard, which is built on user votes. It has a preliminary score of 1525 points, with style control on. On the Code Arena: WebDev leaderboard it is eighth, behind all three rivals from the table.
I took the table apart in a separate post. There you will find all 19 rows, who built each test and who computed which score.
Designing or deploying agents for clients? In our collective of AI consultants you prepare for AI vendor partner certifications and learn with other consultants. When a suitable client brief arrives, we may invite you to a project.
Examples from Google itself
Google described five examples from its own work. These are Google's internal reports, with no published outside check.
Quantum computing. Argon helps researchers cut the resources of subroutines, measured as qubits times gates. In one example it beat the published baseline by 40% within minutes. The post does not say which subroutine.
Data-center memory. A team of AI agents built on Argon went through profiling data from the whole server fleet and applied memory optimizations on its own. The post reports over 300 TiB of memory freed "once rolled out". Whether that rollout is finished is not clear. Google estimates the total saving at 500 TiB to 1 PiB.
Migrating C and C++ to Rust. AI agents built on Argon are rewriting code across Google. The smallest targets are libraries of tens of thousands of lines, such as re2 and libgav1. The largest is the Zircon kernel of the Fuchsia OS, at over 800K lines. The work is ongoing, and the code goes through auditing and review before production. The post does not say how much of Zircon has been migrated.
libgav1, Google's video decoder. Argon replaced 32K lines of SIMD code in an existing Rust port with safe code. The decoder runs 2.7x faster than that port and produces identical video. Against the original C++, Google says only that the new code is "closer". That suggests it is still slower.
Wiz. The company uses Argon in its Scan for Good program, which protects critical public infrastructure for free. The model found a critical vulnerability in healthcare software used by hospitals worldwide. The flaw exposed sensitive personal data. The post names neither the vendor nor a CVE number. Wiz has been part of Google since March 2026.
Every company with old code faces a similar job. In my view, connecting new code to a system that is still running is harder than the rewrite itself. I wrote separately about why AI projects so often run aground on old systems, and how to work around them without rewriting everything at once.
Safety and the two versions of the model
In effect, Argon will come in two versions. Trusted defenders and Google's internal teams will get it "without cyber guardrails", meaning without the blocks on cybersecurity tasks. The post does not define exactly what that removes. Gemini 3.8 Flash Cyber from September works in a similar way: it has looser safeguards and goes only to trusted defenders. The version for everyone will get guardrails that Google is still working on.
Even the unrestricted version is bound by the Fairwind Program's rules. Only defensive and research tasks are allowed, and creating malware is banned. Access is limited to the partner's internal security teams. The partner may not resell it.
The model is also with the US government. Google writes that it is actively engaged in the voluntary process through which the US government gets models before release. It does not say which body runs that process or how long it takes.
Before the broad release, Google is strengthening safeguards in four areas. The model is meant to refuse help with cyberattacks and CBRN weapons, and Google monitors its internal activations. Separate systems watch the model's chain of thought and actions and stop it when necessary. High-risk training and tests run in sealed sandboxes. There is no model card and no Frontier Safety Framework report for Argon yet.
The fourth area matters most to companies: indirect prompt injection. Google calls Argon its most resilient model yet against this attack. On the Gray Swan test the attack succeeded 0.7% of the time over 15 attempts, the lowest of the 13 models on Google's chart. Claude Opus 5.5 and Claude Fable 5.1 are at 1.0% each, GPT-6 Astra at 8.5%. A low rate does not excuse you from designing the AI agent carefully. In a separate post I take apart the attack in which an instruction tucked into a fetched page or file takes over an AI agent.
The Agentic Architect
Practical patterns, case studies and AI news. Zero spam, once a week.
You'll get one email to confirm. Unsubscribe with one click. Privacy policy
What it means for companies in the EU
The announcement says nothing about Poland or the EU. The channels Argon is due to arrive through already work in Poland. The Gemini API and Google AI Studio support Poland and all 27 EU member states. Google AI Ultra costs PLN 469.99 a month in Poland with limits 5x higher than AI Pro, or PLN 979.99 with limits 20x higher.
There are two precedents. Gemini 3 Pro and Gemini 3.1 Pro launched globally in the Gemini app on the same day. Some agent features in the Ultra plan skip Europe, though. Gemini Spark works everywhere except the European Economic Area, Switzerland, the United Kingdom and Nigeria.
Check where the model processes data, too. On Google Cloud, Gemini 3.1 Pro Preview runs only on the global endpoint, and Gemini 3.8 Flash Cyber globally and in the US multi-region. If your company must keep data in the EU, settle that before the pilot.
Fairwind is a global program with over 650 partners, but only some of them get Argon. Apart from Wiz, which Google owns, the announcement does not say which. Governments, critical infrastructure operators and core technology platforms come first. I collected the details in a separate post. There you can check who gets Argon in what order, and what is known about a release in the EU.
What should you do this week? Collect a few dozen tasks from your own company and write down what a good result looks like. Budget on $4 and $20, because the launch price has an end date that Google has not published yet. Name the person who will check an answer of several hundred thousand tokens. When Argon reaches the paid API, you run those tasks on day one.
Sources and fact-check date
- Google: Gemini 4 Argon: our next era of frontier intelligence
- Internet Archive: Google's post as of 22:06
- Google DeepMind: Gemini 4 Argon evaluation methodology (PDF)
- Google DeepMind: Fairwind Program
- Gemini API: Gemini 3.8 Flash model limits
- Gemini API: countries where the API and AI Studio are available
- Google AI Ultra: prices in Poland
- Sundar Pichai on X about the US government and the Fairwind Program
- Arena on X: Argon's text leaderboard score
- Reuters: what the Google spokesperson said
- Vals AI: Vibe Code Bench leaderboard
- Vals AI: Gemini 4 Argon model page and test settings
As of 1 October 2026.
Frequently asked questions
What is Gemini 4 Argon?
Google's new frontier model, announced on 30 September 2026. Google built it for long, multi-step work: software engineering, legal and finance document work, and cybersecurity defense. Google says a single answer from the model can run to 1M tokens.
When will Gemini 4 Argon be available?
Google has not given a date. Selected defenders in the Fairwind Program get the model first. Paid API customers and Google AI Ultra subscribers come next, as soon as possible.
How much does Gemini 4 Argon cost in the API?
At launch, $2 per million input tokens and $10 per million output tokens. Cached input is 95% cheaper, so it costs $0.10. After the introductory period the price rises to $4 and $20. Google has not said how long the launch price will last.
Is Gemini 4 Argon the same as Gemini 4 Pro?
None of Google's launch pages uses the name Gemini 4 Pro. It comes from unofficial pre-launch leaks. The official name is Gemini 4 Argon. A Google spokesperson told Reuters that Argon is larger than the previous line of Pro models.
Will Gemini 4 Argon be available in the EU?
Google's announcement says nothing about the EU. The Gemini API, Google AI Studio and Google AI Ultra already work in EU countries such as Poland, and earlier Gemini models launched globally on the same day. That is only a precedent. Google has promised nothing here.