← Retour au blog
tech 9 September 2026

Mercury 2.5: A Major Leap in Diffusion Models

Mercury 2.5 represents a turning point in diffusion language models, delivering increased intelligence and reduced latency, while remaining cost-effective. Discover how this innovation could transform your business.

Article inspired by the original source
Mercury 2.5 ↗ www.inceptionlabs.ai

Introduction

The era of diffusion language models has reached a new milestone with the release of Mercury 2.5 by Inception Labs. This advanced production model promises to deliver superior intelligence, increased execution speed, and all at a reduced cost. While Mercury 2 has already captivated thousands of developers and seen exponential growth in its utilization, Mercury 2.5 goes even further in terms of capabilities and performance.

What's Changed with Mercury 2.5

Mercury 2.5 positions itself as the most advanced diffusion model on the current market. With a 40% increase in intelligence over Mercury 2, it compares to frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. The model processes 1,107 tokens per second on widely available NVIDIA GPUs, with a context capacity of 260,000 tokens.

In terms of cost, Mercury 2.5 is offered at a competitive price of $0.20 per million input and $0.75 per million output, with an impressive 80% launch discount.

Practical Applications of Mercury 2.5

Search Agents and RAG Pipelines

In the realm of search, Mercury 2.5 proves to be an invaluable asset. Each search request can trigger dozens of model calls to plan the search, rewrite queries, rerank results, structure facts, summarize sources, and verify answers. Thanks to its speed, Mercury allows these calls to stay within a single user interaction, which is crucial for search infrastructure companies that have integrated it into production.

Voice Agents and Interactive Applications

In voice applications, latency is a crucial factor. OpenCall, which develops AI phone agents to handle live customer calls, saw its median model response time drop to close to 170 milliseconds with Mercury. According to Oliver Silverstein, co-founder and CEO of OpenCall, switching to Mercury reduced their P99 response time from several minutes to one second, and their P50 from 0.4 seconds to under 0.2 seconds.

Conclusion

Mercury 2.5 is more than just an update; it's a transformation in how diffusion models can be utilized in real-world production environments. With its enhanced capabilities, speed, and affordable cost, it opens new possibilities for businesses looking to integrate advanced AI models into their operations.

Let's discuss your project in 15 minutes.

Mercury 2.5 diffusion models AI latency NVIDIA
Deepthix newsletter · 100% AI · every Monday 8am

An AI agent reads tech for you.

Our AI agent scans ~200 sources per week and ships the best articles to your inbox Monday 8am. Free. One click to unsubscribe.

Visit the newsletter page →

Want to automate your operations?

Let's talk about your project in 15 minutes.

Book a call