Introduction
In a world where cost optimization is crucial for tech companies, mismanaging resources can quickly become a financial sinkhole. Bartosz Kotrys from Quesma experienced this firsthand when he burned all his tokens researching how to save tokens. This paradox highlights a common challenge faced by many developers and teams working with AI agents. We will explore how to avoid wasting your tokens and master the art of model orchestration for more effective research.
The Context: What is Model Orchestration?
Model orchestration involves using multiple AI models in tandem to accomplish complex tasks while optimizing costs and performance. Bartosz employed a pipeline of agents that relied on a mix of models like Claude, Codex, and Antigravity. The idea is that each model can be utilized for its specific strengths, sharing memories and thus reducing the resource demands of more expensive models.
The Problem: Burning Through Tokens Quickly
On his first attempt, Bartosz launched his research with his deep research pipeline, but quickly reached the limit of his Claude Max 5x plan in just 30 minutes. This scenario is not uncommon. Many AI users find themselves running out of tokens long before tangible results are achieved, raising a crucial question: how to optimize token usage while ensuring reliable results?
The Solution: Efficient Resource Utilization
To solve this problem, Bartosz took an innovative approach by leveraging all the subscription resources he already had. By integrating a shared memory plugin, he enabled Claude, Codex, and Antigravity to share learned data during sessions. This not only reduced costs but also increased the efficiency of the conducted research.
Cheaper Models as Sub-Agents
One of the key strategies was to use less expensive models as sub-agents for simpler tasks. For instance, Claude Code served as the main model for controlling the research flow, while models like Claude Opus 4.8 and GPT-5.5 were used for specific tasks, thus reducing the need to use expensive models like Claude Fable for every step of the process.
Optimization Tools
Using tools like Terminal-Bench and SWE-bench Pro allowed the evaluation of costs and performances of each model, providing valuable data for decision-making. These benchmarks, although useful, should be used with caution as they measure specific aspects that may not be representative of all scenarios.
Conclusion
Optimizing token usage in an AI environment is essential to maximize return on investment. By judiciously orchestrating different models and using benchmark tools to guide decisions, companies can significantly reduce costs while maintaining high performance. If you want to explore how these strategies can apply to your project, let's discuss it in 15 minutes.
Call to Action
Let's discuss your project in 15 minutes.