← Retour au blog
tech 1 September 2026

44% on ARC-AGI-1 in 67 cents

Discover how a small transformer achieved remarkable efficiency on the ARC-AGI-1 benchmark, while being economical and open-source.

Article inspired by the original source
44% on ARC-AGI-1 in 67 cents ↗ mvakde.github.io

Introduction

In the AI world, sample efficiency is one of the most pressing challenges. Reducing costs while maintaining high performance is crucial. This is precisely what Mithil Vakde achieved with his transformer model, reaching 44% on the ARC-AGI-1 benchmark, all for just 67 cents. Let's dive into the details of this technical feat.

The Importance of Sample Efficiency

Sample efficiency refers to a model's ability to learn effectively with a limited number of examples. ARC (Abstraction and Reasoning Corpus) is a meta-learning benchmark characterized by its unsaturated nature and minimal need for synthetic data. This makes it an ideal playground for exploring the limits of sample efficiency.

Methodology and Innovations

Vakde enhanced his previous transformer model with several key innovations. Each input-output pair is converted into sequences of tokens, which the transformer trains on autoregressively. This approach is combined with 3D RoPE embeddings to better handle the 2D grids of the puzzles.

The algorithm also incorporates color and dihedral permutations, increasing data diversity and enabling cross-task learning. Architectural changes such as using SwiGlu instead of GELU and RMSnorm replacing layernorm significantly improved efficiency.

Performance and Costs

Vakde's model not only outperforms many LLMs but does so at a minimal cost. The training took only 1.5 hours and required minimal resources, making this advancement accessible even to researchers with limited budgets.

Future Prospects

The intention behind this work is to push the limits of sample efficiency without increasing costs. Vakde plans to continue his research to further enhance the architecture and algorithm. The goal is to democratize access to performant and affordable models.

Conclusion

This advancement demonstrates that with the right optimizations, it is possible to achieve impressive feats in terms of performance and cost. For those interested in AI and model efficiency, this case study is a source of inspiration.

Let's discuss your project in 15 minutes.

IA transformer efficacité échantillonnale ARC open source
Deepthix newsletter · 100% AI · every Monday 8am

An AI agent reads tech for you.

Our AI agent scans ~200 sources per week and ships the best articles to your inbox Monday 8am. Free. One click to unsubscribe.

Visit the newsletter page →

Want to automate your operations?

Let's talk about your project in 15 minutes.

Book a call