The everyday AI models just got cheaper — and that matters more than a flagship
Google's new Flash models use fewer tokens and cost less to run, which means AI-powered products are now cheaper to build and operate.

- 17%
- fewer output tokens for Gemini 3.6 Flash
- $7.50
- per million output tokens (down from $9)
- $0.30 / $2.50
- Flash-Lite input / output price per million tokens
- 12%
- faster in production, reported by Harvey AI
Most news about AI focuses on the big flagship models — the ones that win benchmarks and dominate headlines. But most real products run on smaller, faster models called Flash, which are priced for high-volume, everyday use. On July 21, Google updated three of them.
Gemini 3.6 Flash is the headline update. Google says it uses 17% fewer output tokens to complete the same work as its predecessor. Tokens are the small chunks of text (and code and data) that AI charges you for — fewer tokens means a lower bill. The output price dropped too, from $9 to $7.50 per million tokens.
The cheapest option, Gemini 3.5 Flash-Lite, costs just $0.30 per million input tokens and $2.50 per million output tokens. Harvey AI — a legal technology company that uses Gemini in production — reported that its workflows ran 12% faster on average with the new model. That kind of real-world result matters more than benchmark scores.
Google's flagship Gemini 3.5 Pro is still only in limited testing, while competitors like Anthropic and OpenAI have shipped major updates. The competition at the top stays open. But for most builders who depend on Flash-tier models, a 17% token cut adds up fast at production scale — and now is the right moment to revisit your AI cost estimates.
- Google Blog — Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
- TechCrunch — Google releases three new Gemini models — but no 3.5 Pro
- SiliconANGLE — Google expands Gemini with cheaper models and a bug-hunter it keeps on a leash
- Unite.AI — Google Ships Three Gemini Flash Models as Its Flagship Slips