Google details Gemini 3.5 Flash-Lite and Pro roadmap

Google details Gemini 3.5 Flash-Lite and Pro roadmap

As of August 2, 2026, Google DeepMind’s model catalog lists Gemini 3.5 Flash-Lite live and a 3.5 Pro tier “coming soon,” signaling a staggered refresh of the Gemini family built, in DeepMind’s words, for “frontier intelligence with action.” The page also spells out capability themes, price bands, and benchmark deltas across the lineup, offering a rare look at how Google is segmenting cost and performance inside its agent push (DeepMind).

What Gemini 3.5 Flash-Lite is meant to do

Google frames Gemini 3.5 Flash-Lite as “best for high-volume tasks that need efficiency and intelligence,” placing it alongside a faster 3.6 Flash model and above earlier 3.1 tiers (DeepMind). The site groups the series around four jobs: agentic coding, advanced multimodal understanding, long-horizon execution, and multi-step problem solving. That reads like a blueprint for production agents that juggle tools, documents, and long-running workflows rather than a one-off chat.

Flash-Lite slots in as the throughput option. It’s positioned for workloads where tokens dominate budgets and responses must stay competent under heavy load. Think bulk code review hints, content triage, transcript summarization, or pre-processing for retrieval. Google routes developers to the consumer-facing Gemini app and the builder side in Google AI Studio, which makes it simple to test prompts, wire up APIs, and watch token counts as you go.

Pricing and benchmarks for the Gemini 3.5 tier

DeepMind publishes headline pricing and test scores across the current stack. According to the model page, Gemini 3.5 Flash is listed at $1.50 per 1 million input tokens and $9.00 per 1 million output tokens, while 3.6 Flash keeps the same input price and drops output to $7.50. Gemini 3.1 Pro carries $2.00 in and $12.00 out. All figures come from Google’s table; real bills depend on your usage mix and any caching or batching you set up.

On the agentic coding front, the page cites SWE-Bench Pro and DeepSWE v1.1 to show long-horizon software performance. DeepMind’s table reports 3.6 Flash at 58.7% on SWE-Bench Pro (public), 3.5 Flash at 55.1%, and 3.1 Pro at 54.2%. The table also includes competitor models for comparison, but Google’s own tiers tell the main story: tighter prices for the fast path, slight score lifts at the top, and a middle band that’s good enough for most agent loops.

If you’re new to planning by tokens, remember that both prompt length and response size matter. A short primer on tokenization helps frame why “input” and “output” are billed separately and how chunking or compression can lower cost without hurting accuracy.

Why Google is staggering the 3.5 lineup

The sequencing is the tell. Google is shipping 3.6 Flash for maximum speed and efficiency today, offering Gemini 3.5 Flash-Lite as the budget workhorse, and keeping 3.1 Pro in play for complex creative tasks while a 3.5 Pro tier is still pending (DeepMind). That split maps to how agent systems behave in production: most cycles are simple tool calls or retrieval passes, with fewer hops that truly need deeper reasoning.

In plain terms, this lineup gives teams a throttle. Use the 3.6 Flash tier when latency and price per response rule the day. Bring in 3.1 Pro or the upcoming 3.5 Pro only when the chain hits ambiguity and needs more deliberation. Park Gemini 3.5 Flash-Lite where you batch volume—classification, extraction, or templated edits—so you don’t pay Pro rates for work a lighter tier can do cleanly.

DeepMind’s capability labels reinforce that intent. “Agentic coding” suggests sustained tool use and repo-scale context. “Long horizon tasks” points to workflows that span many steps and minutes, even hours. Those are the places where costs can spiral without careful tiering. Publishing per‑million token prices up front, next to benchmark deltas, nudges architects to plan routing strategies before deployment, not after overruns arrive.

What’s next for Gemini 3.5 Pro

The open question is timing and scope. The catalog flags “Gemini 3.5 Pro coming soon” but keeps 3.1 Pro as the complex-task option in the meantime (DeepMind). Expect the Pro slot to aim at the cases where tool use and retrieval stop short—hard synthesis, error analysis, or gnarly multimodal reads—while keeping an eye on cost so it can sit inside agent routing without blowing budgets.

Two watchpoints stand out. First, how much of the 3.6 Flash efficiency trickles down to Pro-grade outputs when 3.5 Pro arrives. Second, whether Google raises the ceiling on “long horizon” reliability, the thing that separates toy agents from serious ones. The site’s “frontier intelligence with action” framing hints at the latter focus; agents that keep state, call tools, and recover from missteps will dominate practical gains over raw single-shot scores.

For teams building today, the playbook is straightforward. Prototype in Google AI Studio, measure token flow, and route simple hops to 3.6 Flash or Gemini 3.5 Flash-Lite. Reserve the higher tier for thorny branches. That mix matches how Google has arranged the shelf and, if DeepMind’s own tables hold up under load, it’s the cheapest way to get reliable agent loops at scale.

As Google rounds out the series, keep an eye on the “Hands-on” and “Showcase” sections for concrete workflows and guardrails, and on the benchmark table for any shifts tied to updated training data or tool use. The way the company is packaging tiers—speed first, volume second, depth reserved—makes sense for production agents. The open item is whether Gemini 3.5 Flash-Lite and its siblings stay predictable on long runs as 3.5 Pro arrives.

Related reading: CopilotOpenAIProductivity & AI