9 min read B2C power user

Gemini 4 Argon API Pricing: A 50% Launch Discount That Doubles

Gemini 4 Argon launches at $2/$10 per million tokens, then doubles to $4/$20. Here is what the intro pricing means for a heavy AI bill at scale.

Gemini 4 Argon API Pricing: A 50% Launch Discount That Doubles

Search interest in gemini api pricing runs around 5,400 queries a month, with a cost per click near $3.13 and low competition, and it is about to climb. On September 30, 2026, Google DeepMind revealed Gemini 4 Argon, its first proprietary model above the Flash class in more than seven months. The announcement is a benchmark story first and a pricing story second, but the pricing is the part that will land on your bill. Argon launches at an introductory rate of $2 per million input tokens and $10 per million output tokens, with cached input at 95 percent off, then the rate doubles to $4 per million input and $20 per million output once the promotion ends. Google has not confirmed the end date.

That structure is worth pausing on. This is not a price cut. It is a discount with a cliff, and the cliff is invisible on the launch page. If you route serious work onto Argon without modeling the post-promotion rate, you will absorb a surprise increase that no rate card warned you about.

What Gemini 4 Argon actually ships

Argon is a frontier model built for long, multi-step work. Google reports a new state of the art on DeepSWE v1.1 at 77.9 percent for real-world software engineering tasks, first place on Zapier’s AutomationBench with 51.3 percent, and leadership on the Vals Index, which weights finance, coding, legal, and tax work by each sector’s share of U.S. GDP. On LVBench, which measures long video understanding, it scores 91.7 percent.

The output limit is the structural change. Argon raises the output token ceiling to an industry-leading 1 million tokens, up from 64,000. A new Gemini API feature called Long Decode Continuation pauses long responses and resumes them across follow-up calls, so reasoning can run to 1M output tokens without a request timeout. In practice, that means an agent can produce a genuinely large artifact, a migration plan, a long document, a wide code change, in a single trajectory instead of a chain of truncated retries.

For heavy users, the interesting number is not the benchmark ceiling. It is the hallucination rate. On AA-Omniscience, Argon posts a 15 percent hallucination rate, the lowest of any model scoring 45 or above on the Artificial Analysis Intelligence Index, against 51 percent for GPT-6 Astra at max and 54 percent for GPT-6.1 Sol at max. A model that admits it does not know is cheaper to run, because a hallucinated answer you act on is the most expensive output there is.

Gemini pricing: the intro rate and the doubling

Lay the numbers out plainly, because the shape is the whole story.

At launch, Argon costs $2 per million input tokens and $10 per million output tokens, with cached input discounted 95 percent to $0.10 per million. After the introductory period expires, the price returns to $4 per million input and $20 per million output, with the same 95 percent cache discount.

Artificial Analysis measured the practical consequence. At the discounted rate, Argon costs $1.99 per Intelligence Index task, roughly 60 percent of what GPT-6 Astra at max costs for a comparable intelligence score. When the discount ends, that climbs to $3.98 per task, about 1.2 times GPT-6 Astra at max. The intelligence score does not move. Only the price does.

There is a second detail that compounds the first. Argon’s cost efficiency at launch comes from its token price, not from token thrift. It averages about 62,000 output tokens per task, against 27,000 for GPT-6 Astra at max. A model that thinks longer is fine when the output token is cheap. At $20 per million output tokens and 62k tokens per task, the post-promotion math is where heavy users need to pay attention.

Google Gemini pricing in context

TokenKarma does not cover Google in isolation, and Argon only makes sense against the field it is trying to beat. On the Artificial Analysis Intelligence Index, Argon at high reasoning scores 53, matching GPT-6 Astra at max and sitting one point ahead of GPT-6.1 Sol at max. That is 23 points above Google’s previous non-Flash model, Gemini 3.1 Pro Preview, which scored 30.

The agentic picture is more mixed and more useful. Argon ranks first on AutomationBench-AA at 78 percent, seven points ahead of Claude Sonnet 5.5 at max. On Terminal-Bench 4, it reaches 57 percent, behind Claude Sonnet 5.5 at 64, Claude Opus 5.5 at 60, and GPT-6 Astra at 59. So Argon’s strength is broad end-to-end business execution and long-context understanding, not necessarily the hardest terminal-style coding loops where Anthropic still leads.

The cost story mirrors that split. Google’s new model promises frontier intelligence at a discount, and for a window of at least a month, that is materially true. But OpenAI and Anthropic have been pushing in the opposite direction: doing less work per task rather than charging less per token. Cursor’s own harness post this week is the clearest example, with a 7 percent cost reduction for users delivered entirely through prompt and tool-context trimming. Argon takes the other road, a lower rate for a model that spends more tokens.

Both roads are legitimate. They are not interchangeable, and a heavy user should know which one the discount is actually buying.

Gemini cost: what to do before the price doubles

If Argon is on your roadmap, the actions are concrete and time-boxed.

Model the post-promotion rate before you migrate anything. Budget the work you plan to move onto Argon at $4 and $20, not $2 and $10. If the discounted number is what makes the case, the case does not survive the cliff. For 62k-output-token tasks, the difference is roughly a twofold increase in output spend overnight.

Protect the cache-read ratio. Argon’s 95 percent cache discount drops cached input to $0.10 per million at the promotional rate. Cached reads are the single largest lever on any long agent run. Keep the stable prefix of your prompts genuinely stable so the cache keeps hitting, because a cache miss at $4 per million input erases most of the launch discount.

Use Long Decode Continuation deliberately. The 1M output ceiling is real, but output tokens are the expensive half of the meter. Spend the long trajectory on work that genuinely needs it, like large migrations or long-form generation, and stop short on tasks that a shorter answer resolves. A bigger ceiling is a bigger bill if you treat it as free.

Measure cost per solved task, not cost per token. This is the same discipline the Claude and OpenAI releases forced this month. Argon’s lower hallucination rate and higher rubric pass rate should mean fewer wasted retries, and that efficiency is real. The open question is whether it is large enough to offset the doubling. Run one representative workflow at both price points before you commit a production pipeline to it.

Watch the promo end date. Google said the discount lasts at least a month and has not fixed an end date. Treat the promotional rate as a trial period, not a plan. When the rate changes, your cost per task changes with it, and no migration will soften that.

Argon earns its place at the top of the leaderboard, and for the duration of the launch discount it is the most cost-efficient frontier model on the market. The discipline is to adopt it for the right tasks at the right rate, and to know exactly what your bill does the day the discount ends.