All articles

July 21, 2026

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite: Halving Time per Task

Google has released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Both halve time per task relative to their predecessors and increase token efficiency, Gemini 3.5 Flash-Lite improves by 11 Intelligence Index points while Gemini 3.6 Flash does not improve in intelligence over 3.5 Flash

Google DeepMind has released the latest updates to the Gemini model family with two new models. We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash-Lite ahead of release across Intelligence, Time per Task, and Cost per Task

Key takeaways for Gemini 3.6 Flash (high reasoning):

Maintains the same Intelligence as Gemini 3.5 Flash: Gemini 3.6 Flash scores 50 on the Artificial Analysis Intelligence Index, matching Gemini 3.5 Flash, and just below recently released models Muse Spark 1.1 (xhigh, 51) and GPT-5.6 Luna (max, 51). Compared to Gemini 3.5 Flash, Gemini 3.6 Flash maintains similar scores across the Index, with an improvement in GDPval-AA v2 (1421, +72) and a slight regression in HLE (38%, -3 points)

➤ Half the Time per Task: Gemini 3.6 Flash records an average time per task of 1.3 minutes, a more than 50% reduction compared to Gemini 3.5 Flash (2.7). This is driven by increased token efficiency and faster output, with speeds measured at 304 output tokens per second in our pre-launch testing

Slightly lower Cost per Task: Gemini 3.6 Flash’s cost per task decreases ~18%, from $0.59 to $0.50. This is driven by lower output token use and new pricing of $1.50/$7.50 per 1M input/output tokens, down from $1.50/$9.00 for Gemini 3.5 Flash

Key takeaways for Gemini 3.5 Flash-Lite (high reasoning):

➤ Significant Intelligence improvements over Gemini 3.1 Flash Lite: Gemini 3.5 Flash-Lite scores 36 on the Artificial Analysis Intelligence Index, up 11 points from Gemini 3.1 Flash-Lite (25). This places it behind models such as Nemotron 3 Ultra (38) and DeepSeek V4 Flash (max, 40), and above Mistral Medium 3.5 (30). The biggest intelligence gains compared to Gemini 3.1 Flash-Lite are in agentic evaluations, with improvements in GDPval-AA v2 (1140, +498), TerminalBench v2.1 (53.6, +22.5 points) and Tau3-Banking (16.5%, +7.8 points)

➤ Nearly half the Time per Task: Gemini 3.5 Flash-Lite records an average time per task of 0.6 minutes, nearly a 50% reduction compared to Gemini 3.1 Flash-Lite (1.0). This is driven by increased token efficiency and fast output speed, measured at 350 output tokens per second in our pre-launch testing

More expensive with 2x Cost per Task: Gemini 3.5 Flash-Lite’s average cost per task increases from $0.04 to $0.09, driven by new pricing of $0.30/$2.50 per 1M input/output tokens, up from $0.25/$1.50 for Gemini 3.1 Flash-Lite. This cost increase comes despite using fewer output tokens, falling from 20k to 13k average output tokens per task

Key model details:

Context window: Both models retain the same 1M context window as their predecessors

Multimodality: Both models have text, image, video, and speech input with text output only

Pricing: Gemini 3.6 Flash is priced at $1.50/$7.50 per million input/output tokens, down from Gemini 3.5 Flash at $1.50/$9.00. Gemini 3.5 Flash-Lite is priced at $0.30/$2.50 per million input/output tokens, with the same input pricing across all input modalities. This is an increase from Gemini 3.1 Flash-Lite, which is priced at $0.25/$1.50 per million input/output tokens, with input audio tokens at $0.50. Both models retain the same 90% discount for cached input tokens