Gemini 3.8 Flash keeps 3.7's $0.75/$3.75 price until 31 Dec 2026, then doubles. 90.8% on Terminal-Bench 2.1, more thinking tokens per task. Migrate or stay?
Gemini 3.8 Flash is generally available as gemini-3.8-flash since 2 September 2026 at $0.75 per million input tokens and $3.75 per million output — exactly 3.7 Flash’s introductory price — and both numbers double to $1.50 / $7.50 on 1 January 2027. It beats 3.7 Flash on every benchmark Google published, most sharply on Terminal-Bench 2.1 (90.8% vs 81.6%), and beats GPT-5.6 Terra (87.4%) and Claude Sonnet 5 (80.4%) on the same test. The catch is in the model card: 3.8 is built on 3.7 rather than a new base, it “works harder” by spending more thinking tokens, and Google tells efficiency-first workloads to stay on 3.7.
That last sentence is unusual for a launch, and it is the reason this guide is a decision table rather than a migration checklist. Sources: Google’s model page and the DataCamp and eesel breakdowns published on launch day.
Pricing, and the 1 January cliff
| Period | Input / MTok | Output / MTok (includes thinking) |
|---|---|---|
| Now to 31 December 2026 | $0.75 | $3.75 |
| From 1 January 2027 | $1.50 | $7.50 |
Batch and Flex are half those rates; Priority is 1.8×. Multimodal input — text, image, audio, video, PDF — sits inside the same 1M-token window with 64k output. Three things follow from the table. Output pricing includes thinking tokens, and 3.8 thinks more than 3.7 at the same effort setting, so the per-request cost is higher even at identical list prices. Any budget you approve on September pricing is wrong by 2× in four months. And Gemini remains an order of magnitude below the frontier tier: Claude Fable 5.1 and GPT-6 Astra are $10 / $50. Thirteen Gemini 3.8 Flash requests cost roughly one Astra request at the same token counts — before Astra’s cache discount and before Gemini’s extra thinking, which is why you should model both in the AI model cost calculator rather than eyeballing it.
Benchmarks: a coding release, not a reasoning release
| Benchmark | Gemini 3.8 Flash | Gemini 3.7 Flash | Delta |
|---|---|---|---|
| Terminal-Bench 2.1 | 90.8% | 81.6% | +9.2 |
| SWE-Bench Pro | 61.6% | 60.4% | +1.2 |
| SWE-Atlas | 51.9% | 48.0% | +3.9 |
| τ³-bench Banking (tool use) | 38.1% | 30.9% | +7.2 |
| CharXiv (multimodal reasoning) | 86.2% | 84.5% | +1.7 |
| Humanity’s Last Exam | 45.4% | 45.7% | −0.3 |
| HLE-Verified | 54.9% | — | — |
The shape is unmistakable. Terminal work and multi-step tool use jump; general reasoning is flat (HLE actually dips a tenth of a point). Google also says 3.8 Flash outperforms “most larger frontier models” on DeepSWE v1.1 for long-horizon coding without publishing the percentage, and that it beats Claude Opus 5 on three benchmarks. Against the competition on Terminal-Bench 2.1, 3.8 Flash’s 90.8% sits above GPT-5.6 Terra at 87.4% and Claude Sonnet 5 at 80.4%.
If you have read our Gemini 3.5 Flash guide from May, note the velocity: 3.5 Flash scored 76.2% on Terminal-Bench 2.1 four months ago. The Flash line has gained 14.6 points on that benchmark in one summer, across three releases (3.6 on 21 July, 3.7 on 13 August, 3.8 on 2 September).
Comments · 0
Beta: comments are stored locally on your device and not visible to other readers.
No comments yet. Be the first to share your thoughts.