Swiftlet fits an 80B-level Qwen giant into Mac: peak memory is only 4.3GB, iPhone 17 can also run 35B natively
The Swiftlet project uses Swift + Metal to enable MoE models like Qwen3-Next to run with minimal size on Mac: only a small dense core is kept in memory, and expert weights are stored on SSD and loaded on-demand. On an M5 Mac, the 4-bit version of Qwen3.6-35B-A3B occupies only 18GB of disk space, with a peak memory of 2.6GB and a decoding speed of 7–11 tok/s.