GEMINI 4 ARGON · API PRICING • 11 min read •

Gemini 4 Argon API pricing and the 1M-token answer

Gemini 4 Argon API pricing starts at $2 per million input tokens and $10 per million output tokens. Cached input tokens are 95% off, so they cost $0.10 per million. After an introductory period the rates rise to $4 and $20, and Google has not said when. One Argon answer can run to a million tokens. The output tokens of such an answer alone cost $10 today and $20 after the promotion.

Gemini 4 Argon API pricing: $2 and $10 per million tokens at launch, $4 and $20 after the introductory period, 95% off cached input and a 1M-token answer limit

I price AI agent projects for clients, so from Google's announcement I take two numbers: what an answer that uses the whole new limit costs, and who checks it afterwards. I have not tested Argon. It is not generally available yet, so I am working from the price list.

Benchmarks, use cases and the timeline are in my overview of the announcement and what it means for companies. Here I stay with the money.

How much does Gemini 4 Argon cost in the API?

Google gave two prices. The introductory one is $2 per million input tokens, $10 per million output tokens and 95% off cached input, which comes to $0.10. The standard one is $4 and $20. The footnote with that second price gives no date and says nothing about caching. The Gemini 3.8 Flash promotion has an end date on the price list, 31 December 2026, and Argon has none.

API prices, USD per million tokens (as of 1 October 2026)
ModelInputCached inputOutputMax output (tokens)
Gemini 4 Argon, introductory price 2 0.10 10 1M
Gemini 4 Argon, after the introductory period 4 0.20, if the 95% discount stays 20 1M
Gemini 3.1 Pro Preview (prompts up to 200K tokens) 2 0.20 plus storage 12 65,536
Gemini 3.8 Flash (promotion until 31 Dec 2026) 0.75 0.075 plus storage 3.75 65,536
Claude Opus 5.5 4 0.20 (write 5) 20 128K (300K in the Batch API, beta)
Claude Fable 5.1 10 0.25 (write 12.50) 50 128K
GPT-6 Astra (prompts up to 272K tokens) 10 1 (write 12.50) 50 128K

After the promotion, the closest match is Claude Opus 5.5, which costs exactly the same: $4 and $20, with cached input at $0.20. GPT-6 Astra and Claude Fable 5.1 cost $10 and $50. VentureBeat notes that at the introductory price Argon costs one fifth of GPT-6 Astra's price. After the promotion it is 40%.

In OpenAI's lineup, GPT-6.1 Sol has Argon's launch price: $2 and $10 in the API. After the promotion Argon will cost twice as much. I broke down the whole OpenAI price list after DevDay, from GPT-6 Luna to the ChatGPT Pro plans, in a separate piece.

What does one 1M-token answer cost?

Google raised the output token limit to one million, from the previous 64K. Current Gemini 3.x text models list a limit of 65,536 tokens in the docs. GPT-6 Astra, Claude Opus 5.5 and Claude Fable 5.1 stop at 128K tokens. Google explains the change with reasoning. A model with that much headroom can generate hundreds of thousands of tokens in one trajectory, think more deeply and solve a hard problem in one go.

The arithmetic is simple. One million output tokens times $10 per million is $10 at the introductory price, and after the promotion 1M times $20, which is $20. Then add the input tokens. With a 200K-token prompt you add 0.2 times $2, or $0.40, and $0.80 after the promotion. A full answer with that prompt therefore costs $10.40 today and $20.80 later.

One caveat. On current Gemini models, thinking tokens count toward the limit and Google bills them as output tokens. Google has not yet published that rule for Argon. If it keeps it, the million tokens cover both the reasoning and the answer. You get a shorter text and pay for the whole thing.

On other models the same million output tokens take many calls. Claude Opus 5.5 charges $20 but needs at least eight calls, because 1,000,000 divided by 128,000 is 7.8. In the Batch API beta you get down to four, and Batch halves the price. GPT-6 Astra and Claude Fable 5.1 charge $50 each, also over at least eight calls. Gemini 3.1 Pro Preview charges $12, but with a 65,536-token limit it needs at least sixteen. Every call sends the prompt again, so you pay for it several times.

After the promotion, Argon and Opus 5.5 cost the same per token. Argon's gain is that the whole result comes out of one run and you do not have to stitch pieces together.

What fits in a million tokens of output?

Google does not convert tokens into pages or lines of code, so I work from an approximation. An average line of code has a few dozen characters, and a token is usually a few characters. I assume roughly 10 tokens per line as a working figure. It is my yardstick, Google gives none, and your language and tokenizer will give a different result.

On that yardstick a million tokens is on the order of 100K lines of code. A whole migrated module fits in one answer, tests included. Google writes that AI agents built on Argon are moving libraries of tens of thousands of lines to Rust, and a library like that fits. The Zircon kernel, which Google puts at over 800K lines, does not.

Text works the same way. A few characters per token gives a few million characters, which is over a thousand pages of 1,800 characters, minus whatever the thinking takes. A full report with appendices fits with room to spare.

One caveat on that number. Vals AI tested Argon before launch. Its model page lists a 262K-token output limit, and that is the setting it ran its tests with. Google's announcement says a million. I do not know where the difference comes from. Before you plan a task around the full million, check the limit in the API docs once Argon appears there.

FOR AI ARCHITECTS AND CONSULTANTS

Designing or deploying agents for clients? In our collective of AI consultants you prepare for AI vendor partner certifications and learn with other consultants. When a suitable client brief arrives, we may invite you to a project.

Caching at 95% off: when does it pay?

With long prompts, caching moves the bill more than the answer limit does. Argon takes 95% off cached input. Current Gemini models take 90% off, for example $0.20 instead of $2 on Gemini 3.1 Pro Preview.

Take a company knowledge base: 200K tokens of policies and documentation placed at the start of every prompt. A hundred users each ask one question, so the model reads 100 times 0.2M, or 20M tokens.

Without caching you pay 20 times $2, or $40 at the introductory price. With caching the first prompt costs 0.2 times $2, or $0.40. The other 99 read 19.8M tokens from the cache at $0.10, or $1.98. In total you pay $2.38, which is 94% less. After the promotion it comes to $80 without caching and $4.76 with it, if the discount stays.

What does this bill include? Input tokens only, because you count the answers separately. I assume every one of the 99 following prompts hits the cache. I leave out a fee for keeping data in the cache, because Google has not given one for Argon. Gemini 3.1 Pro Preview charges $4.50 per million tokens per hour, which for 200K tokens comes to $0.90 an hour.

A 200K-token prompt sent 100 times, input only, USD per the price lists
ModelWithout cachingWith caching
Gemini 4 Argon, introductory price 40.00 2.38
Gemini 4 Argon, after the promotion (if the 95% discount stays) 80.00 4.76
Gemini 3.1 Pro Preview 40.00 4.36 plus storage
Gemini 3.8 Flash (promotion until 31 Dec 2026) 15.00 1.64 plus storage
Claude Opus 5.5 (first write to a 5-minute cache) 80.00 4.96
Claude Fable 5.1 200.00 7.45
GPT-6 Astra 200.00 22.30

After the promotion, Argon and Opus 5.5 come out almost level here: $4.76 and $4.96. Caching pays when many prompts start with the same long block. That is what a knowledge base for a hundred users looks like, and so does an AI agent's fixed system prompt or the same tool definitions at every step. Keep the fixed part at the start and the user's question at the end. When every prompt carries a different document, the discount does little.

The introductory price trap: budget on $4 and $20

Building a product on Argon? Price it at $4 and $20. The promotion will end, and Google has not said when. So you do not know how many months of your client contract it will cover. In my view it is safer to assume none.

After the promotion, input and output tokens cost exactly twice as much. Your model cost doubles too, provided the cache keeps its 95% discount. If it dropped to 90%, as on current Gemini models, cached input would cost $0.40. The knowledge base example would then cost $8.72 instead of $4.76. That is my scenario. Google has not announced any such change.

The price list also does not tell you how many tokens a model burns per task. Vals AI, which tested Argon before launch, measured that on its Vals Index at $4 and $20. The average cost per test came to $15.68 for Argon and $32.14 for Claude Opus 5.5, although both models have the same price per token. Vals measured on its own tasks, and yours may come out differently.

Newsletter

The Agentic Architect

Practical patterns, case studies and AI news. Zero spam, once a week.

You'll get one email to confirm. Unsubscribe with one click. Privacy policy

Who will read a million tokens?

After the promotion, a million output tokens cost $20. An engineer who has to read 100K lines of generated code will spend many days on them, and each of those days costs more than the model's whole answer. In my view that is where the real cost of a long answer sits. A million tokens nobody checked is a risk inside the client's system.

Google does not skip this either. The Rust migrations in the announcement go through automated and manual auditing, emulation testing and review before they reach production. Google does not say they are there yet.

The libgav1 example, Google's video decoder, gives a good pattern. AI agents built on Argon took an existing Rust port and replaced 32K lines of SIMD code. Google describes the result with two facts you can measure: the video output is identical, and the decoder runs 2.7x faster than the earlier Rust port. The baseline is that port, because according to Google the decoder only got closer to the C++ version. You check neither fact by reading the code line by line. That is what an acceptance condition for a long answer looks like.

Set one before you ask for a million tokens. It can be tests, matching results against the old system or a timing measurement. A second model that reads the result and approves it proves little. I wrote up separately a real bug that agreement between three AI agents would not have caught, and the check that rejects it.

Across a whole project budget the proportion looks similar. I showed it in a cost breakdown where checking and maintaining an AI agent after launch can outgrow the build cost within a year.

What is not yet known about Argon's price?

The announcement leaves several gaps, and each of them can change your bill.

  • How long the introductory period lasts. The footnote says only that it will expire.
  • Whether cached input keeps its 95% discount at the $4 price. The footnote names only input and output tokens. I calculate $0.20 on the assumption that the discount stays. VentureBeat makes the same assumption and flags it too.
  • What Argon will cost in the Batch API. Other paid Gemini models get 50% off there, and Google has not given that price for Argon.
  • Whether a prompt above 200K tokens gets a higher rate. Gemini 3.1 Pro Preview then charges $4 per million input tokens and $18 per million output tokens.
  • What Google will charge for keeping data in the cache.
  • How Google will bill thinking tokens. On current Gemini models they are output tokens, counted toward the limit.
  • What the model will be called in the API. Argon is not in the Gemini API docs yet.
  • Whether "1M" means 1,000,000 or 1,048,576 tokens.

It is also unknown when paid API customers will get the model. Google says only that it will happen as soon as possible. I collected separately who gets Argon first and what is known about its availability in Poland and the EU.

What to calculate before Argon reaches the API

  • Recalculate the project budget at $4 and $20. Count cached input at $0.20, and at $0.40 in the cautious variant.
  • Check in your logs how much of the prompt repeats between calls. That decides how much caching saves you.
  • Decide how many output tokens you allow per task. On current Gemini models the max_output_tokens limit includes thinking tokens too.
  • Write down the acceptance condition for a long answer before you ask for one.

Once you have access, run your set of real tasks on Argon. Divide the whole bill by the number of tasks completed correctly and compare it with the model you use today.

Sources and fact-check date

As of 1 October 2026.

SP

Szymon Paluch

ex-CTO · AI Strategy

Advising clients on which model to build AI agents on?

The Certified AI Consultant Program is a collective of AI consultants. Together we prepare for AI vendor partner certifications. Today those are the Claude certifications, through our partner organization. The program does not grant a Google certification. You learn alongside other consultants, and after a separate profile assessment you may be invited to a client project. Applying is free; you pay only after acceptance.

See the program for AI consultants
Related posts
What Is a Forward Deployed Engineer? Role, Pay and Who Hires
Forward Deployed AI Engineer: The Job, the Work and Who Hires
How to Become a Forward Deployed Engineer: 5 Steps