GEMINI 4 ARGON · COMPARISON • 11 min read •

Gemini 4 vs GPT-6 Astra vs Claude: what Google's benchmark table says

Gemini 4 Argon has the best score in 13 of the 19 rows of the table Google published at launch. It ties GPT-6 Astra in one. In five a rival has the top score: GPT-6 Astra in three, Claude Opus 5.5 in two. In two of those rows Argon comes last of the four models. These are Google's numbers, and Argon is not generally available yet.

Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 in Google's benchmark table: Argon best in 13 of 19 rows, 1 tie, 5 losses

I compare four models: Gemini 4 Argon from Google, GPT-6 Astra from OpenAI, and Claude Fable 5.1 and Claude Opus 5.5 from Anthropic. From here on I use the short names Argon, Astra, Fable 5.1 and Opus 5.5. Argon is not generally available yet, and I have not tested it, so I read Google's table and check it against the benchmark owners' own leaderboards.

For the record: I passed the Claude Certified Architect Foundations exam, and my company holds OpenAI Select Partner status. That is why every number below says who reported it and what it measures.

The rest of the announcement, from availability to the output limit and Google's own examples, is in my overview of what Google announced on 30 September.

Google's full table: 19 rows, four models

The table has 19 rows but 18 benchmarks, because GraphWalks takes two rows: up to 128K tokens, and 256K to 1M. I count rows. A win is the highest score in a row, and a tie is a score equal to the best rival. Count benchmarks instead and you get 12 wins, 1 tie and 5 losses. VentureBeat counts the same way.

Google's benchmark table from the Gemini 4 Argon launch, scores in percent. Best score in each row highlighted (Google's data, 30 September 2026)
BenchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5Best in row
Knowledge work
Vals Index68.9%63.1%65.8%67.0%Argon
AutomationBench51.3%41.4%31.4%42.5%Argon
Vals Finance Agent v265.4%53.5%58.9%58.6%Argon
Harvey's Legal Agent Benchmark19.6%5.4%6.7%3.8%Argon
Agentic coding
DeepSWE v1.177.9%74.1%67.4%74.2%Argon
FrontierSWE v255.0%65.5%56.3%62.3%GPT-6 Astra
Vibe Code Bench91.9%89.6%90.3%90.3%Argon
Terminal-bench 4.057.4%58.2%57.9%66.4%Claude Opus 5.5
ML engineering
PostTrainBench45.3%44.3%40.2%49.3%Claude Opus 5.5
Science and math
Terminal-Bench Science 0.157.6%68.1%52.6%63.3%GPT-6 Astra
LABBench 288.8%85.4%68.6%73.1%Argon
RiemannBench76.0%72.0%65.6%69.6%Argon
Long context
GraphWalks, up to 128K (BFS, F1)99.7%98.7%91.4%90.6%Argon
GraphWalks, 256K to 1M (BFS, F1)84.2%71.8%65.0%66.8%Argon
Computer use
Agent's Last Exam (pass rate)39.5%34.2%n/a38.2%Argon
OSWorld-2.0 (offline subset, partial score)69.2%72.6%n/an/aGPT-6 Astra
Multimodal understanding
Chartography71.6%71.0%46.2%66.3%Argon
LVBench91.7%87.5%79.7%83.7%Argon
Cybersecurity
CWE-bench v168.0%68.0%58.0%67.0%Tie: Argon and GPT-6 Astra

"n/a" means Google gives no score. That is why only Argon and Astra remain in OSWorld-2.0.

Where Gemini 4 Argon wins

It is strongest in knowledge work, where it wins all four rows. The Vals Index combines finance, coding, legal and tax tasks, and weights each sector by its share of US GDP. Argon scores 68.9% there, Opus 5.5 67.0%. On Zapier's AutomationBench it leads by 8.8 points, and on Vals Finance Agent v2 by 6.5.

The second strength is long context. In GraphWalks the model gets a graph written as an edge list and has to traverse it. With contexts from 256K to 1M tokens, Argon beats Astra by 12.4 points. With shorter contexts the lead shrinks to one point.

Five of the wins are slim, though. On the Vals Index, Vibe Code Bench, the shorter GraphWalks, Agent's Last Exam and Chartography the lead is under 2 points.

FOR AI ARCHITECTS AND CONSULTANTS

Designing or deploying agents for clients? In our collective of AI consultants you prepare for AI vendor partner certifications and learn with other consultants. When a suitable client brief arrives, we may invite you to a project.

Where Gemini 4 Argon loses

In agentic coding Argon wins two rows out of four. In the other two it comes last. On FrontierSWE v2 it trails Astra by 10.5 points, and on Terminal-bench 4.0 it trails Opus 5.5 by 9.

On Terminal-Bench Science 0.1 Argon also trails Astra by 10.5 points, but finishes third. The other two losses are smaller. In PostTrainBench an AI agent spends 10 hours post-training a language model on a single GPU, and Opus 5.5 wins it by 4 points. On OSWorld-2.0, a computer-use test, Astra is ahead by 3.4.

The only tie is CWE-bench v1, a test of patching vulnerabilities in code: Argon and Astra both score 68.0%. A Google spokesperson was, in fact, more cautious with Reuters than the blog was. In the spokesperson's words, Argon is comparable to Astra and Opus on key coding and cyber benchmarks.

Who built the benchmarks and who ran the numbers

Who wrote the tests? Mostly someone outside Google: Vals AI, Zapier, Surge AI or Harvey. GraphWalks even comes from OpenAI. Google, however, decided which 18 tests made the table.

What matters more is who ran the numbers, and Google spells that out in its methodology. In 9 rows, Google says, every score comes from public leaderboards. Argon has 7 wins, 1 tie and 1 loss there. Five rows Google computed itself, for every model, and Argon wins four of them. In the last five Google ran only Argon and took the rivals' scores from other leaderboards or from the vendors. Interestingly, this is where Argon loses most often: three times out of five.

Four benchmark owners, namely Vals AI, Zapier, Proximal and Collinear AI, already list Argon on their leaderboards. So they must have had it before launch. I consider these the best numbers we have today. Some of them also show something Google's table lacks: the spread of a result. On FrontierSWE v2 the worst and best of five runs lie 18 to 20 points apart.

Conditions are not always equal either. On Terminal-Bench Science 0.1 Google gave Argon six times the verifier timeout. On OSWorld-2.0 Google reports Argon's best of three runs, and takes Astra's score from OpenAI's blog. And on LVBench Argon watched video at 1 frame per second, while the rivals got between 300 and 800 frames, as their APIs allowed.

I also found mismatches with the sources. Google says it took the rivals' Terminal-bench 4.0 scores from the official leaderboard. Opus 5.5 is not on that leaderboard, though, and its 66.4% matches the figure Anthropic reported itself. Meanwhile Google lists Fable 5.1 at 52.6% on Terminal-Bench Science, while the official leaderboard shows 40.0%. That substitution works in the rival's favour. And Surge AI's public leaderboards, which Google cites for RiemannBench and Chartography, do not list Argon yet.

What is missing from the table? Models outside the four. On CWE-bench v1 the owner shows a three-way tie, and Grok 4.7 ranks ahead of Argon once the tie is broken on pass@4. Harvey's benchmark is solved more often by four Muse Spark models from Meta. Vibe Code Bench is led by Claude Sonnet 5.5.

Settings can move a score by dozens of points. Anthropic reports Chartography with tools, where Opus 5.5 scores 89.0% and Fable 5.1 88.4%. Google's table gives the no-tools results. The same models then score 66.3% and 46.2%.

Then there is the harness: the agent loop, its tools and the way it manages context. Even on CWE-bench each model runs in a different one: Argon in Antigravity, Astra in Codex, Opus 5.5 in Claude Code. So the score measures the model together with its wrapper. I wrote separately about what the harness adds to a model's score and what you rewrite when you swap the model.

Arena: first place at 1525

Arena builds its ranking from user votes. On 30 September its official account announced that Argon (High) leads the Text Arena at 1525. The leaderboard page marks that score "Preliminary", with a ±9 margin and 4,942 votes. The ranking is computed with style control on. A screenshot showing 1533 went around online, but the official figure is 1525.

Opus 5.5 (High) is fourth at 1504, and Astra is not in the top 25. In Code Arena: WebDev the picture flips. There Argon ranks eighth (1679) and trails all three rivals: Opus 5.5 is first (1818), Astra second (1789) and Fable 5.1 fourth (1751).

Newsletter

The Agentic Architect

Practical patterns, case studies and AI news. Zero spam, once a week.

You'll get one email to confirm. Unsubscribe with one click. Privacy policy

Price per capability: rates, cost per task and answer length

Google published two prices, an introductory one and a standard one. How long will the introductory price last? The announcement does not say. Nor does it say whether cached input stays 95% cheaper afterwards.

API prices in USD per million tokens, standard rates as of 30 September 2026, prompts up to 272K tokens
ModelInputCached inputOutputMax output (tokens)
Gemini 4 Argon, introductory20.10 (computed: 95% off)101M (per Google)
Gemini 4 Argon, after the introductory period4not stated201M (per Google)
GPT-6 Astra10150128K
Claude Fable 5.1100.2550128K
Claude Opus 5.540.2020128K

VentureBeat wrote that Argon costs one-fifth of Astra's price. That holds, but only at the introductory price. At the standard price it is 40%, and Argon's standard price equals Opus 5.5's. On the other hand, OpenAI's GPT-6.1 Sol already costs as much as Argon on promotion. Astra also gets pricier on prompts above 272K tokens: input then costs twice as much, output one and a half times as much.

The token price alone does not tell you what a task costs. Vals AI gives a cost per Vals Index task and prices Argon at its standard rate. That comes to $15.68 for Argon, $18.46 for Astra, $28.71 for Fable 5.1 and $32.14 for Opus 5.5. On FrontierSWE v2 the order looks different. There a single trial costs, per Proximal, $98.87 on Opus 5.5, $129.36 on Argon and $1,029.65 on Astra. Proximal adds that Argon ran in a non-production setting with a much lower cache hit rate.

The second difference is the length of a single answer. According to Google, Argon is to return up to 1M tokens. Google describes this as up from 64K, and current Gemini 3.x models list a limit of 65,536 in the docs. Vals AI, which has already evaluated Argon, lists a 262K output limit and uses it as the default setting in its tests. The rivals stop at 128K. The exception is Opus 5.5 on the Batch API, which reaches 300K in beta. At a 128K limit, a million tokens takes at least eight calls. And somebody then has to read every such answer.

The full bill, caching included, is in my calculation of what one million-token Argon answer costs.

How to choose a model for a client

Google's table shows the tasks of benchmark authors. Your client has their own. In my view, choosing a model starts with your own test set. Someone else's table can at most suggest which models to let into it.

You freeze the set before the first run. A new model stays only if it improves the score on the same tasks. I described this rule as a yardstick you do not touch between one model and the next.

SIX STEPS BEFORE YOU PICK A MODEL

  1. 1. Collect 30 to 50 real client tasks with correct answers.
  2. 2. Freeze the tasks and the scoring before you run the first model.
  3. 3. Run each model in the harness you will use in production.
  4. 4. Repeat each task several times and record the spread of results.
  5. 5. Count the cost per correctly completed task, at the standard price.
  6. 6. Check that the model is available in the API and in the client's region.

You can add Argon to the test once Google opens API access. Build the task set now.

Sources and check date

As of 1 October 2026.

Frequently asked questions

Is Gemini 4 better than GPT-6?

In Google's table Gemini 4 Argon beats GPT-6 Astra in 14 of 19 rows, ties in one and loses in four. These are Google's own results, and Argon is not generally available yet. In Code Arena: WebDev GPT-6 Astra ranks higher.

Gemini 4 vs Claude: which should you choose?

It depends on the tasks. In Google's table Argon beats Claude Opus 5.5 in 14 rows and loses in four, including Terminal-bench 4.0 and PostTrainBench. Opus 5.5 is available now and costs $4 and $20 per million tokens, the same as Argon after the promotion. Test both models on your own tasks.

Can I use Gemini 4 Argon yet?

Not yet. Google is giving it first to selected partners in the Fairwind Program, which it runs for cyber defenders. Paid API customers and Google AI Ultra subscribers come next. The announcement gives no date.

SP

Szymon Paluch

ex-CTO · AI Strategy

Advising clients on which model to pick?

The Certified AI Consultant Program is a collective of AI consultants. Together we prepare for AI vendor partner certifications. Today those are the Claude certifications, through our partner organization. The program does not grant a Google certification. You learn alongside other consultants, and after a separate profile assessment you may be invited to a client project. Applying is free; you pay only after acceptance.

See the program for AI consultants
Related posts
What Is a Forward Deployed Engineer? Role, Pay and Who Hires
Forward Deployed AI Engineer: The Job, the Work and Who Hires
How to Become a Forward Deployed Engineer: 5 Steps