Running Kimi K3 (2.8T) at 1 token/s on a MacBook Pro via Four SSDs
Engineers managed to run the massive 2.8-trillion parameter Kimi K3 model on a MacBook Pro by streaming data from four SSDs.

The tech community is buzzing over a remarkable hardware feat involving large language models. Engineers successfully demonstrated running the massive Kimi K3 model, featuring 2.8 trillion parameters, locally on a MacBook Pro.
To achieve this breakthrough, the system streamed data concurrently from four high-speed SSDs, reaching a generation speed of 1 token per second. While not practical for heavy daily workloads, it showcases extreme hardware optimization.
Historically, models of this immense scale require server clusters packed with specialized, high-end AI accelerators. Leveraging consumer hardware combined with massive storage throughput represents a creative workaround for memory bandwidth bottlenecks.
For developers and tech enthusiasts worldwide, this experiment highlights the ongoing evolution of local AI inference capabilities. It pushes the boundaries of what is possible with portable workstations and sparks new ideas for distributed memory handling in machine learning.



