I Tried to Run Qwen3.8–27B on a 16GB Mac Mini with AirLLM. Here’s Exactly Where It Breaks
Last Updated on August 25, 2026 by Editorial Team Author(s): Abhishek Gautam Originally published on Towards AI. The claim, and why it’s seductive AirLLM promises 70B models on a 4GB GPU. Its README even lists Qwen3.8–27B at 3.33GB. So why can’t a Mac Mini M4 with 16GB of unified memory run it? I went looking for the actual failure, not the hand-wavy one. The author reports trying to run Qwen3.8–27B on a 16GB Mac Mini using AirLLM and shows why it fails. They claim AirLLM’s macOS path hard-routes models to a single MLX/“Llama” implementation and never correctly invokes the Qwen3.8-specific class, then the model-splitting logic uses a buggy substring match that accidentally links thousands of tensors but crashes during layer counting (after downloading ~55.56GB) due to mismatched layer naming conventions. They also argue that even if the routing were corrected, the code paths are platform-incompatible on macOS: the MLX persistence layer writes dict/nested mx.array structures that the torch streaming engine can’t interpret, and the architecture mismatch (Gated DeltaNet) would require a new backend. Finally, the article measures the practical bottleneck: AirLLM must re-read ~53.79GB from disk per token, leading to an estimated ~16.4 seconds per token (~219 tokens/hour), which makes long “reasoning” responses take hours; the author concludes that the approach is effectively I/O-bound on Apple Silicon and recommends using a smaller GGUF quant with llama.cpp’s Metal backend instead. Read the full blog for free on Medium. Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor. Published via Towards AI
