← Retour au blog
tech 17 July 2026

Kimi K3, and what we can still learn from the pelican benchmark

Kimi K3, Moonshot AI's latest model with 2.8 trillion parameters, is reshaping AI expectations. Explore how it stacks up against market leaders and what the pelican test reveals about its capabilities.

Article inspired by the original source
Kimi K3, and what we can still learn from the pelican benchmark ↗ simonwillison.net

Introduction

The world of artificial intelligence is constantly evolving, with each new advancement pushing the boundaries of what's possible. Recently, the Chinese lab Moonshot AI announced the launch of Kimi K3, their most advanced model to date, boasting 2.8 trillion parameters. This model is not only a technological marvel but also promises to redefine industry standards. So, what can we learn from Kimi K3 and the pelican benchmark test?

Kimi K3: A New Benchmark

Moonshot AI presents Kimi K3 as the first open 3T-class model. With 2.8 trillion parameters, it significantly outperforms its predecessor, Kimi K2.6, in terms of performance. According to internal benchmarks, Kimi K3 primarily beats Claude Opus 4.8 max and GPT-5.5 high, although it still lags behind Claude Fable 5 and GPT-5.6 Sol. One standout feature is its cost per task, which stands at $0.94, almost half the price of Opus 4.8 ($1.80).

Technical Capabilities of Kimi K3

One of the most interesting aspects of Kimi K3 is its token consumption efficiency. According to the Artificial Analysis Intelligence Index, output token usage decreased by 21% compared to Kimi K2.6. This means the model is not only more powerful but also more resource-efficient, which is crucial for large-scale applications.

Performance on the Frontend Code Arena

Kimi K3 is now the leading model on Arena.ai's Frontend Code arena, even surpassing Claude Fable 5. This performance illustrates its capability to adapt to complex frontend development tasks, a domain where optimization and efficiency are essential.

The Pelican Test: An Indicator of Capability

The pelican experiment is a playful yet revealing way to test Kimi K3's image generation capabilities. Using OpenRouter with the llm-openrouter plugin, an SVG image of a pelican riding a bicycle was generated, requiring 95 input tokens and 16,658 output tokens. While this test is simple, it demonstrates the model's ability to handle complex image generation tasks with impressive accuracy.

What We Learn

The pelican test, although unconventional, shows that Kimi K3 can not only generate complex images but also interpret them accurately. This capability is crucial for practical applications such as computer vision and automated visual content generation.

Conclusion

Kimi K3, with its 2.8 trillion parameters, marks a significant milestone in the development of advanced AI. With its impressive performance and efficiency, it sets new expectations for future models. The pelican test is a perfect example of how creative benchmarks can reveal essential underlying capabilities.

Let's discuss your project in 15 minutes.

Kimi K3 AI benchmarks Moonshot AI Pelican test AI model performance
Deepthix newsletter · 100% AI · every Monday 8am

An AI agent reads tech for you.

Our AI agent scans ~200 sources per week and ships the best articles to your inbox Monday 8am. Free. One click to unsubscribe.

Visit the newsletter page →

Want to automate your operations?

Let's talk about your project in 15 minutes.

Book a call