Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Swiftlet is a Swift and Metal runtime enabling large Qwen models to run on Apple devices with limited RAM by streaming expert weights from disk on demand while keeping only the dense model core in memory. This achieves running an 80B model with 4.3 GB of RAM on a Mac and a 35B model on an iPhone.
How is it possible to run an 80-billion parameter language model on a typical Mac?
Swiftlet streams expert weights from disk on demand while keeping only the small dense core of the model resident in memory (about 2.5 GB). This works because Qwen models use a sparse Mixture-of-Experts architecture where only about 3 billion parameters activate per token, requiring only a small subset of experts to be loaded at any time.
What are the performance differences between running Swiftlet on Mac versus iPhone?
On Apple Silicon Macs, the 35B model runs in 2.6 GB RAM achieving 7-11 tokens/second, while the 80B version uses 4.3 GB RAM at 4.5-5 tokens/second. On iPhone, the 35B model runs in approximately 2.5 GB RAM at about 1 token/second, which is sufficient for interactive conversation despite the lower throughput.
What role does Mixture-of-Experts architecture play in Swiftlet's efficiency?
Swiftlet targets Qwen models that use MoE, where each layer routes tokens to a small subset of experts from a large pool (e.g., 10 of 512 experts for the 80B model). Only a fraction of parameters activate per token, enabling efficient execution on memory-constrained devices without degrading model capabilities significantly.
How does Swiftlet differ from other solutions for running LLMs on Apple devices?
Swiftlet is written in Swift and Metal with runtime shader compilation, requiring no Metal toolchain at build time and running identical code on iOS and macOS. It implements optimized expert caching mechanisms and fixed-stride packing on SSD that minimize I/O overhead compared to memory-mapped approaches.
- Qwen 3.8 27B: powerful model, but default mode leads to wild overthinking — simonwillison.net 75 % match
- LFM2.5-2.6B: small and capable local AI model — liquid.ai 74 % match
- DeepSeek V4 Flash on a single AMD MI300X — github.com 74 % match
- Swiftlet
- Qwen3-Next-80B
- Qwen3.6-35B
- Swift
- Metal
- MLX3
- Hugging Face12
- Priv AI
- ANEMLL