Gemini 2.5: Updates to our family of thinking models

JUNE 17, 2025

Shrestha Basu Mallick Product Google DeepMind

Logan Kilpatrick Group Product Manager

Today we are excited to share updates across the board to our Gemini 2.5 model family:

Gemini 2.5 Pro is generally available and stable (no changes from the 06-05 preview)

Gemini 2.5 Flash is generally available and stable (no changes from the 05-20 preview, see pricing updates below)

Gemini 2.5 Flash-Lite is now available in preview

Gemini 2.5 models are thinking models, capable of reasoning through their thoughts before responding, resulting in enhanced performance and improved accuracy. Each model has control over the thinking budget, giving developers the ability to choose when and how much the model “thinks” before generating a response.

Overview of our family of Gemini 2.5 thinking models

Introducing Gemini 2.5 Flash-Lite

Today, we’re introducing 2.5 Flash-Lite in preview with the lowest latency and cost in the 2.5 model family. It’s designed as a cost-effective upgrade from our previous 1.5 and 2.0 Flash models. It also offers better performance across most evals, and lower time to first token while also achieving higher tokens per second decode. This model is great for high throughput tasks like classification or summarization at scale.

Gemini 2.5 Flash-Lite is a reasoning model, which allows for dynamic control of the thinking budget with an API parameter. Because Flash-Lite is optimized for cost and speed, “thinking” is off by default, unlike our other models. 2.5 Flash-Lite also supports all of our native tools like Grounding with Google Search, Code Execution, and URL Context in addition to function calling.

Benchmarks for Gemini 2.5 Flash-Lite

Updates to Gemini 2.5 Flash and pricing

Over the last year, our research teams have continued to push the pareto frontier with our Flash model series. When 2.5 Flash was initially announced, we had not yet finalized the capabilities for 2.5 Flash-Lite. We also launched with a “thinking” and “non-thinking price”, which led to developer confusion.

With the stable version of Gemini 2.5 Flash rolling out (which is the same 05-20 model preview we made available at Google I/O), and the incredible performance of 2.5 Flash, we are updating the pricing for 2.5 Flash:

$0.30 / 1M input tokens (*up from $0.15 input)

$2.50 / 1M output tokens (*down from $3.50 output)

We removed the thinking vs. non-thinking price difference

We kept a single price tier regardless of input token size

While we strive to maintain consistent pricing between preview and stable releases to minimize disruption, this is a specific adjustment reflecting Flash’s exceptional value, still offering the best cost-per-intelligence available.

And with Gemini 2.5 Flash-Lite, we now have an even lower cost option (with or without thinking) for cost and latency sensitive use cases that require less model intelligence.

Pricing updates for our Gemini Flash family

If you are using the Gemini 2.5 Flash Preview 04-17 , the existing preview pricing will remain in effect until its planned deprecation on July 15, 2025, at which point that model endpoint will be turned off. You can transition to the generally available model “gemini-2.5-flash”, or switch to 2.5 Flash-Lite Preview as a lower cost option.

Continued growth of Gemini 2.5 Pro

The growth and demand for Gemini 2.5 Pro continues to be the steepest of any of our models we have ever seen. To allow more customers to build on this model in production, we are making the 06-05 version of the model stable, with the same pareto frontier price point as before.

We expect that cases where you need the highest intelligence and most capabilities are where you will see Pro shine, like coding and agentic tasks. Gemini 2.5 Pro is at the heart of many of the most loved developer tools.

Top developer tools using Gemini 2.5 Pro

If you are using 2.5 Pro Preview 05-06, the model will remain available until June 19, 2025 and then will be turned off. If you are using 2.5 Pro Preview 06-05, you can simply update your model string to “gemini-2.5-pro”.

We can’t wait to see even more domains benefit from the intelligence of 2.5 Pro and look forward to sharing more about scaling beyond Pro in the near future.

AI Cloud Announcements

TorchTPU: Running PyTorch Natively on TPUs at Google Scale

APRIL 7, 2026

Gemini Web AI Tutorials How-To Guides

Turn creative prompts into interactive XR experiences with Gemini

FEB. 19, 2026

Mobile Web Announcements

Bring state-of-the-art agentic skills to the edge with Gemma 4

APRIL 2, 2026

Gemini Google AI Studio AI Events

How we built the Google I/O 2026 Save the Date experience

MARCH 3, 2026