raoul.studio Blog
AI in Practice · August 15, 2026

A better AI for coding arrived at half price. Test it before January.

Google's new Gemini 3.7 Flash is dramatically better at coding tasks and available at half its normal price until the end of 2026 — but it still fails on most complex, multi-step work.

Key facts
65.3%
score on the DeepSWE software engineering benchmark
$0.75 / 1M tokens
input price through December 31, 2026
340 tokens/sec
output speed — fastest of 186 models tested
23 days
time between Gemini 3.6 Flash and 3.7 Flash

Google just made one of the biggest single-version jumps in AI coding performance. On August 13, the company released Gemini 3.7 Flash — a model that jumped dramatically on every coding test compared to its predecessor, released just 23 days earlier. On DeepSWE, a benchmark that simulates real software engineering work, the score climbed from 49% to 65.3%. Another test, AutomationBench, nearly doubled from 17% to 30.4%.

The price is what makes this news matter right now for anyone building software. Until December 31, 2026, Google charges $0.75 per million tokens — tokens are the little chunks of text an AI reads and writes, roughly three-quarters of a word each. That is half the normal price until January, when the rate doubles to $1.50. The next five months are the best window to build with this model and measure how much it helps your real work.

There is an important limit. Even at this better level, Gemini 3.7 Flash still fails on most complex, multi-step tasks — the kind where you ask an AI agent to handle a whole project from start to finish. AutomationBench nearly doubled, which sounds big, but seven out of ten complex attempts still fail. And for now, users in Europe cannot access the consumer version at all, with no timeline given.

In practice, Gemini 3.7 Flash works well for single-step coding jobs: writing a function, reviewing a pull request, generating tests, or explaining code you have not seen before. The right time to test is now, while the introductory price holds. Run it on your actual work — the gap between AI coding tools shifts every few weeks, and hands-on tests tell you more than any leaderboard.

Sources
The AI Experiment Is Over When a Powerful AI Costs Nothing Per Use
Digital product studioA better AI for coding arrived at half price. Test it before January.raoul.studio