← Retour au blog
tech 20 July 2026

I Burned All My Tokens Researching How to Save Tokens

In a world where optimizing AI agent costs is crucial, learn how to avoid wasting your tokens and master the art of model orchestration.

Article inspired by the original source
I burned all my tokens researching how to save tokens ↗ quesma.com

Introduction

In a world where cost optimization is crucial for tech companies, mismanaging resources can quickly become a financial sinkhole. Bartosz Kotrys from Quesma experienced this firsthand when he burned all his tokens researching how to save tokens. This paradox highlights a common challenge faced by many developers and teams working with AI agents. We will explore how to avoid wasting your tokens and master the art of model orchestration for more effective research.

The Context: What is Model Orchestration?

Model orchestration involves using multiple AI models in tandem to accomplish complex tasks while optimizing costs and performance. Bartosz employed a pipeline of agents that relied on a mix of models like Claude, Codex, and Antigravity. The idea is that each model can be utilized for its specific strengths, sharing memories and thus reducing the resource demands of more expensive models.

The Problem: Burning Through Tokens Quickly

On his first attempt, Bartosz launched his research with his deep research pipeline, but quickly reached the limit of his Claude Max 5x plan in just 30 minutes. This scenario is not uncommon. Many AI users find themselves running out of tokens long before tangible results are achieved, raising a crucial question: how to optimize token usage while ensuring reliable results?

The Solution: Efficient Resource Utilization

To solve this problem, Bartosz took an innovative approach by leveraging all the subscription resources he already had. By integrating a shared memory plugin, he enabled Claude, Codex, and Antigravity to share learned data during sessions. This not only reduced costs but also increased the efficiency of the conducted research.

Cheaper Models as Sub-Agents

One of the key strategies was to use less expensive models as sub-agents for simpler tasks. For instance, Claude Code served as the main model for controlling the research flow, while models like Claude Opus 4.8 and GPT-5.5 were used for specific tasks, thus reducing the need to use expensive models like Claude Fable for every step of the process.

Optimization Tools

Using tools like Terminal-Bench and SWE-bench Pro allowed the evaluation of costs and performances of each model, providing valuable data for decision-making. These benchmarks, although useful, should be used with caution as they measure specific aspects that may not be representative of all scenarios.

Conclusion

Optimizing token usage in an AI environment is essential to maximize return on investment. By judiciously orchestrating different models and using benchmark tools to guide decisions, companies can significantly reduce costs while maintaining high performance. If you want to explore how these strategies can apply to your project, let's discuss it in 15 minutes.

Call to Action

Let's discuss your project in 15 minutes.

AI agents token optimization model orchestration cost management deep research
Deepthix newsletter · 100% AI · every Monday 8am

An AI agent reads tech for you.

Our AI agent scans ~200 sources per week and ships the best articles to your inbox Monday 8am. Free. One click to unsubscribe.

Visit the newsletter page →

Want to automate your operations?

Let's talk about your project in 15 minutes.

Book a call