Google's third Flash model in six weeks burns 30% more tokens
Three Flash generations in six weeks means three shifts in citation behaviour. Brands treating Google AI search as a stable channel are already behind.
Key takeaways
- Gemini 3.8 Flash burns ~30% more output tokens per task than its predecessor, raising effective costs despite identical token pricing.
- Benchmark price-per-token figures are unreliable for reasoning-heavy workloads; enterprises should measure cost-per-completed-workflow.
- Google's Flash acceleration is an infrastructure and API play, not evidence of improving AI search quality.
- Three model releases in six weeks means citation patterns in AI Overviews may have shifted three times without any change on a brand's part.
- Perplexity and ChatGPT, running newer frontier reasoning, may gain ground on complex multi-source citation queries Google is deprioritising.
Three Flash models in six weeks. Google is not iterating; it is flooding the zone.
The Decoder reports that Gemini 3.8 Flash, released in late September 2026, matches Anthropic's Claude Opus 5 on select agentic coding benchmarks while carrying a lower headline token rate than its rival. On paper, that is a significant claim: a budget-tier model competitive with a frontier one on a task class that enterprise buyers care about. In practice, the picture is messier, and the messiness matters more than the benchmark number.
The catch is in the token burn. Gemini 3.8 Flash's "working harder" reasoning mode generates approximately 30% more output tokens per task than its predecessor. Token pricing is identical to the previous Flash generation, so the cost per completed task rises regardless of the lower per-token rate. Buyers who benchmark on price-per-million-tokens will be surprised when their invoice arrives.
The gap between listed price and effective cost
This is not a quirk; it is a pattern that enterprise procurement teams have been slow to absorb. AI models are priced like hotel rooms: the rack rate tells you little about what checkout costs. Reasoning-heavy tasks, agentic loops, and multi-step workflows all multiply output tokens in ways that flat pricing sheets obscure. A financial services firm running compliance document review, or a multilateral organisation processing policy submissions through an AI pipeline, will discover that a 30% token-burn increase translates directly into a budget overrun at scale.
For B2B brands building AI-mediated workflows rather than buying AI search visibility, the implication is straightforward: evaluate models on cost-per-completed-workflow, not cost-per-token. The two numbers diverge sharply when reasoning is involved.
The visibility implication runs in a different direction. As Google accelerates Flash model releases, the models that power AI Overviews, Gemini-assisted search, and third-party AI applications built on Google's API are changing faster than most brands' content and technical teams can track. Each model update brings different citation preferences, different summarisation behaviour, and different threshold sensitivity to source authority signals. Three major model iterations in six weeks means three opportunities for a brand's citation frequency to shift without any change on the brand's part.
What Google's frontier absence signals
The Decoder flags what the Flash cadence obscures: Google's frontier models remain absent. No Gemini Ultra successor has materialised while the Flash family has had three generations. That gap suggests Google is prioritising deployable, cost-efficient models over benchmark leadership at the frontier. The commercial logic is defensible. Enterprises and API customers run Flash; they rarely need frontier capability for production workloads.
But for AI search specifically, frontier model quality determines the depth and accuracy of synthesis in long-form answers. If Google is deprioritising frontier development, the long-answer quality of AI Overviews may plateau relative to competitors running newer frontier reasoning. That matters to brands whose authority depends on being cited in complex, multi-source answers, the kind that policy institutions and industrial groups tend to generate queries about. Perplexity and ChatGPT, both running more recent frontier reasoning, may increasingly win those high-stakes citation slots.