Introduction
In recent years, an apparently absurd test has become a benchmark in the AI world: asking a model to generate an SVG image of a pelican riding a bicycle. What started as a joke is now a serious tool for evaluating the creative capabilities of language models. With billions of dollars at stake, could AI labs be tempted to 'pelicanmaxx,' that is, to optimize their models specifically for this test?
The Pelican on a Bicycle Test
Simon Willison, a renowned developer and blogger, popularized this test by using it for every major language model release. The test has become a staple on forums like Hacker News, prompting discussions about its validity and the temptation for labs to improve solely on this criterion.
Methodology of the Experiment
To explore the question of 'pelicanmaxxing,' Dylan Castillo conducted an experiment generating 1,008 SVG images across seven leading models. The models tested included GPT-5.6 Terra, Claude Sonnet 5, and others. Each model received a series of 48 prompts, combining different animals and vehicles, to test the models' flexibility and creativity.
Experiment Results
The results showed that AI models did not perform any better with the pelican on a bicycle than with other animal-vehicle combinations. For example, the success rate for generating a coherent image was similar between a pelican on a bike and a flamingo on a scooter. This suggests that labs are not pelicanmaxxing their models, at least not significantly.
Limitations and Interpretations
It is important to note that the experiment covered only a limited subset of models and scenarios. The results cannot be generalized to the entire AI field. However, they provide valuable insights into model behavior when faced with specific and complex tasks.
Conclusion
So, are AI labs pelicanmaxxing? The results from Dylan Castillo's experiment suggest they are not. The models seem to have generalized capabilities rather than being optimized for specific tasks like the pelican on a bicycle. That said, the allure of easily manipulated benchmarks remains a valid concern in model evaluation.
Let's discuss your project in 15 minutes.