Google’s Gemini 3.8 Flash boosts performance — but could raise token bills
Google’s Gemini 3.8 Flash delivers stronger reasoning and engineering performance but may increase token consumption and per-task costs despite unchanged per-token rates.

Google has introduced Gemini 3.8 Flash, a new version of its high-performance language model that prioritizes extra reasoning and iterative tool use — a trade-off that could increase the token costs of running the model.
The company says the model “works harder” than its Gemini 3.7 Flash predecessor by taking more reasoning steps on complex tasks and calling external tools more frequently. While Google set the same introductory per-token prices as the previous build—$0.75 per million input tokens and $3.75 per million output tokens—the firm warns the model may consume more tokens to maximize performance, especially at higher effort levels.
Performance gains and cost implications
Google highlights substantial accuracy and capability improvements for software engineering tasks and autonomous agents. In internal and benchmark testing, Gemini 3.8 Flash outperformed Gemini 3.7 Flash on the DeepSWE v1.1 software engineering benchmark and also scored ahead of competing frontier models on the Vals Finance Agent V2 and Harvey’s Legal Agent evaluations.
Early third-party analysis suggests those capability gains can translate into higher per-task token use. Artificial Analysis reported that, despite unchanged per-token pricing, costs per task could rise by roughly 40% due to a 30% increase in output tokens and more agent turns. Industry users have praised the model’s speed and coding quality compared with rivals, while noting the changed cost dynamics.
Security controls and limited cyber offering
Alongside the main release, Google launched Gemini 3.8 Flash Cyber under its new Fairwind Program, which is restricted to governments and trusted partners. The Fairwind roster includes about 650 members such as CrowdStrike and the Center for Internet Security, providing access to the cyber-focused model plus Google’s CodeMender agent designed to autonomously find and fix vulnerabilities.
Google says the general 3.8 Flash model is shipped with safeguards to reduce misuse in sensitive domains including chemical, biological, radiological, and nuclear (CBRN) topics and cyber offense, and access to the Cyber variant is limited to vetted organizations.
Gemini 3.8 Flash is now available to consumers who subscribe to Google AI Pro or Ultra, and to developers and enterprise customers. Developers who wish to limit token spend can continue using Gemini 3.7 Flash to minimize usage.
