← Retour au blog
tech 9 September 2026

Kimi K3 (2.8T): Performance at 1 token/s on MacBook Pro, streamed from four SSDs

Discover how the Kimi K3 (2.8T) model, streamed from four SSDs, achieves 1 token/s performance on a MacBook Pro. A breakthrough for AI developers on Apple Silicon.

Article inspired by the original source
Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs ↗ github.com

Introduction

With the rise of Apple Silicon processors, new possibilities emerge for developers and software engineers. The Kimi K3 (2.8T) model, a mixture of experts (MoE) model, is a recent example demonstrating how these processors can be effectively used for artificial intelligence. This model is capable of processing 1 token per second when streamed from four SSDs on a MacBook Pro.

The Kimi K3 Model and Its Features

The Kimi K3 model is a MoE with 2.8 trillion parameters, making it one of the most advanced models in its category. With 1.45 TB of expert weights, it requires efficient resource management to operate optimally. The decision to use multiple SSDs for streaming data is a strategic one, leveraging increased bandwidth and fast access times.

Why Use Multiple SSDs?

Using four SSDs to stream the Kimi K3 model is no coincidence. Each of these SSDs offers read and write speeds that can exceed 500 MB/s, which is crucial for managing the data flow needed to run such a model. By distributing the load across multiple SSDs, bottlenecks are minimized, ensuring smooth data streaming.

Performance on MacBook Pro

MacBook Pros equipped with the M1 chip have shown impressive performance in various benchmarks. For the Kimi K3, processing 1 token per second is a significant achievement, highlighting these machines' ability to handle intensive workloads. This paves the way for more complex AI applications on portable machines, which was previously the realm of dedicated servers.

Use Cases

Portable AI Development

Imagine developing and testing complex models anywhere, thanks to the power of a MacBook Pro. With the Kimi K3, researchers can experiment with advanced architectures without relying on expensive server infrastructure.

Real-Time Applications

Applications requiring quick responses, such as voice assistants or real-time translation, directly benefit from this advancement. Processing 1 token per second allows for rapid user request responses.

Conclusion

Streaming the Kimi K3 model from four SSDs on a MacBook Pro represents a significant advancement for portable AI and the efficient use of modern hardware resources. This approach opens new possibilities for developers looking to integrate advanced AI capabilities into their projects.

Let's discuss your project in 15 minutes.

Kimi K3 Apple Silicon SSD streaming AI performance MacBook Pro
Deepthix newsletter · 100% AI · every Monday 8am

An AI agent reads tech for you.

Our AI agent scans ~200 sources per week and ships the best articles to your inbox Monday 8am. Free. One click to unsubscribe.

Visit the newsletter page →

Want to automate your operations?

Let's talk about your project in 15 minutes.

Book a call