← Retour au blog
tech 22 July 2026

Are AI Labs Pelicanmaxxing?

The pelican on a bicycle test has become a famous benchmark for assessing AI models. But are AI labs tweaking their models to excel in this specific test?

Article inspired by the original source
Are AI Labs Pelicanmaxxing? ↗ dylancastillo.co

Introduction

In recent years, an apparently absurd test has become a benchmark in the AI world: asking a model to generate an SVG image of a pelican riding a bicycle. What started as a joke is now a serious tool for evaluating the creative capabilities of language models. With billions of dollars at stake, could AI labs be tempted to 'pelicanmaxx,' that is, to optimize their models specifically for this test?

The Pelican on a Bicycle Test

Simon Willison, a renowned developer and blogger, popularized this test by using it for every major language model release. The test has become a staple on forums like Hacker News, prompting discussions about its validity and the temptation for labs to improve solely on this criterion.

Methodology of the Experiment

To explore the question of 'pelicanmaxxing,' Dylan Castillo conducted an experiment generating 1,008 SVG images across seven leading models. The models tested included GPT-5.6 Terra, Claude Sonnet 5, and others. Each model received a series of 48 prompts, combining different animals and vehicles, to test the models' flexibility and creativity.

Experiment Results

The results showed that AI models did not perform any better with the pelican on a bicycle than with other animal-vehicle combinations. For example, the success rate for generating a coherent image was similar between a pelican on a bike and a flamingo on a scooter. This suggests that labs are not pelicanmaxxing their models, at least not significantly.

Limitations and Interpretations

It is important to note that the experiment covered only a limited subset of models and scenarios. The results cannot be generalized to the entire AI field. However, they provide valuable insights into model behavior when faced with specific and complex tasks.

Conclusion

So, are AI labs pelicanmaxxing? The results from Dylan Castillo's experiment suggest they are not. The models seem to have generalized capabilities rather than being optimized for specific tasks like the pelican on a bicycle. That said, the allure of easily manipulated benchmarks remains a valid concern in model evaluation.

Let's discuss your project in 15 minutes.

AI pelicanmaxxing benchmarking language models SVG generation
Deepthix newsletter · 100% AI · every Monday 8am

An AI agent reads tech for you.

Our AI agent scans ~200 sources per week and ships the best articles to your inbox Monday 8am. Free. One click to unsubscribe.

Visit the newsletter page →

Want to automate your operations?

Let's talk about your project in 15 minutes.

Book a call