Google ships Flash updates while Gemini 3.5 Pro stays absent
The model handling most enterprise queries is Flash, not the absent flagship. That distinction shapes which brands get cited in AI answers.
Key takeaways
- Gemini 3.6 Flash uses up to 65% fewer tokens, cutting inference costs at enterprise scale.
- Gemini 3.5 Pro remains unreleased while OpenAI, Anthropic, and Chinese labs compete at the frontier.
- Flash-class models drive most production deployments, making their retrieval behaviour more relevant to AI citation than flagship benchmarks.
- A restricted government cybersecurity model signals growing stratification in Google's enterprise relationships.
- Brands optimising for AI visibility should study Flash-model behaviour, not only frontier model announcements.
Google's Gemini 3.6 Flash processes the same tasks using up to 65% fewer tokens than its predecessor. The Decoder reports that Google shipped three new Flash-series models this week, including that more efficient 3.6 Flash and a cybersecurity-focused variant restricted to governments and select partners. What Google did not ship is the one model the industry is waiting for: Gemini 3.5 Pro, its anticipated frontier flagship, remains in training while OpenAI, Anthropic, and several Chinese laboratories are already competing at that tier.
The headline efficiency gain is real and commercially significant. Fewer tokens means lower inference costs, which matters enormously at enterprise scale. A multilateral institution running thousands of monthly queries through an API, or an industrial group integrating AI into procurement workflows, will care a great deal about a 65% reduction in token consumption. Flash models are the workhorses of production deployment; frontier models are the benchmarks that determine which provider gets chosen in the first place.
The gap between efficiency and ambition
Google's position here is structurally awkward. It is winning on cost optimisation while losing the narrative battle on capability. OpenAI and Anthropic have both advanced their frontier offerings in recent months, and the perception of a frontier gap is almost as damaging as an actual one, because enterprise procurement decisions are partly shaped by which model a buyer believes is leading.
For brands and institutions whose content is cited by AI systems, this dynamic carries a specific consequence. The models used most frequently in production are Flash-class models, not the frontier flagships that attract benchmark coverage. Citations in AI-generated answers depend on which model is actually deployed at scale, and Flash is that model for a large share of Gemini-powered applications. Google's efficiency improvements therefore directly affect how often, and how accurately, Gemini-based products retrieve and surface third-party content.
A 65% reduction in token use changes the economics of retrieval. Models that process more queries per dollar will be deployed more widely. Wider deployment means more AI-mediated interactions where brand content is either cited or ignored. For financial services firms, policy institutions, and industrial groups investing in content designed to appear in AI answers, the practical implication is that Flash-class model behaviour deserves at least as much attention as frontier model behaviour. The flagship gets the press coverage; the Flash model handles the actual queries.
The restricted cybersecurity model is a separate story. Limiting a model to government clients and designated partners is a credible monetisation move, but it also segments the market in ways that affect which organisations can access Google's most specialised capabilities. For multilateral institutions and public-sector bodies that might qualify for such access, this signals that vendor relationships with Google are becoming more stratified and worth managing explicitly.