← Retour au blog
tech 23 June 2026

VibeThinker: A 3 Billion Parameter Model Surpassing Opus 4.5 in Reasoning

VibeThinker-3B redefines the capabilities of compact models with its novel SFT+GRPO approach, outperforming much larger models.

Introduction

In the world of artificial intelligence, model size has long been considered a direct indicator of power. However, VibeThinker-3B disrupts this notion. With only 3 billion parameters, it surpasses Opus 4.5 through a unique approach combining Supervised Fine-Tuning (SFT) and Guided Reinforcement Policy Optimization (GRPO).

What is VibeThinker-3B?

VibeThinker-3B is a compact language model designed to test the limits of verifiable reasoning within a restricted model framework. It uses the Spectrum-to-Signal post-training paradigm, optimizing the model through an enhanced pipeline that includes curriculum-based supervised fine-tuning, multi-domain reinforcement learning, and offline self-distillation.

Impressive Performance

VibeThinker-3B's performance on verifiable reasoning tasks is remarkable. It scores 94.3 on AIME26, even reaching 97.1 with claim-level real-time scaling. On LiveCodeBench v6, it achieves a Pass@1 of 80.2 and shows strong out-of-distribution generalization with a 96.1% acceptance rate on recent unseen LeetCode contests.

SFT+GRPO: A Revolution in Reasoning

The coupling of Supervised Fine-Tuning with Guided Reinforcement Policy Optimization enables VibeThinker-3B to maximize its reasoning efficiency. This unique combination enhances the model's ability to adapt to new contexts while maintaining high accuracy.

A Compact Yet Powerful Model

Contrary to popular belief, compact models like VibeThinker-3B are not just lighter alternatives but offer a complementary path towards top-tier performance in parameter-dense capability regimes. The model excels in the Parametric Compression-Coverage Hypothesis, confirming that verifiable reasoning can be compressed into compact reasoning cores.

Implications and Applications

Tech companies can leverage models like VibeThinker-3B to enhance their systems without investing in costly infrastructure. Applications range from complex data analysis to decision-making process automation, proving that compactness does not sacrifice competence.

Conclusion

VibeThinker-3B demonstrates that small models have their place in the language model hierarchy, offering a unique balance between performance and efficiency. As we continue to explore the limits of AI, models like this show that size isn't everything.

Let's discuss your project in 15 minutes.

VibeThinker-3B SFT+GRPO AI reasoning compact models language models
Deepthix newsletter · 100% AI · every Monday 8am

An AI agent reads tech for you.

Our AI agent scans ~200 sources per week and ships the best articles to your inbox Monday 8am. Free. One click to unsubscribe.

Visit the newsletter page →

Want to automate your operations?

Let's talk about your project in 15 minutes.

Book a call