← Retour au blog
tech 6 June 2026

Benchmarks in Leipzig: Assessing AI's Mathematical Capabilities

During the *Benchmarks in Leipzig* workshop, 49 mathematicians assessed AI models on advanced mathematical questions. Discover how these LLMs performed and what it means for the future of artificial intelligence.

Article inspired by the original source
Benchmarks in Leipzig ↗ arxiv.org

Introduction

The Benchmarks in Leipzig event took place between April 1 and May 15, 2026, bringing together 49 mathematicians at the Max Planck Institute for Mathematics in the Sciences in Leipzig, Germany. The goal? To assess the mathematical reasoning capabilities of large language models (LLMs) using a set of 100 research-level mathematics questions.

The Context of Benchmarks

Benchmarks are essential for measuring AI model performance. In this case, mathematicians compiled a set of questions covering various fields like algebraic geometry, combinatorics, and representation theory. These questions, with known answers, served as a basis for testing the LLMs' ability to solve complex problems.

The Three Phases of Evaluation

Phase 1: Initial Attempt

In the first phase, five state-of-the-art AI models attempted to solve the 100 questions in a single attempt each. The result? 41 questions remained completely unsolved, highlighting the models' initial limitations.

Phase 2: Iteration and Improvement

The second phase involved three of the previous models, this time with 20 attempts per model. This iterative process reduced the number of unsolved questions to 16, demonstrating a notable improvement in the models' capabilities through repeated learning.

Phase 3: Heavy-Thinking Models

For the final phase, two models focused on deep thinking were given three attempts each. This challenge left only two questions unsolved, highlighting significant progress in mathematical reasoning by these AIs.

Implications for the Future of AI

This Leipzig benchmark shows that LLMs are increasingly capable of tackling complex mathematical problems. With only two questions left unsolved, the progress made is impressive and paves the way for even more advanced AI applications in the scientific field.

Conclusion

The Benchmarks in Leipzig event demonstrated that AIs can now rival human minds on high-level mathematical problems. These advances bring us closer to a future where AI models will play an essential role in scientific research. Let's discuss your project in 15 minutes.

IA mathématiques benchmarks LLMs Leipzig
Deepthix newsletter · 100% AI · every Monday 8am

An AI agent reads tech for you.

Our AI agent scans ~200 sources per week and ships the best articles to your inbox Monday 8am. Free. One click to unsubscribe.

Visit the newsletter page →

Want to automate your operations?

Let's talk about your project in 15 minutes.

Book a call