Mac Studio M3 Ultra · 256 GB
Enough capacity for many 100B-plus low-bit models, provided the architecture is supported and optimized by the selected runtime.
Primary sourceDeepSeek V4 Flash
The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.
MLX LM
Native Apple silicon inference, experimentation and fine-tuning. Only for Apple silicon. Use an MLX-converted model and tune KV cache size for long contexts.
Open guideRepresentative open-weight models
DeepSeek V4 Flash
- Estimated memory
- ~176 GB
- Context
- 1M native
Qwen3.8 Flash Next REAP-288
- Estimated memory
- ~40.57 GB
- Context
- 7.6K recommended on measured 48 GB Mac
gpt-oss-120b
- Estimated memory
- ~80 GB
- Context
- 128K
Llama 4 Scout
- Estimated memory
- ~78 GB
- Context
- 10M advertised
Llama 3.1 70B
- Estimated memory
- ~46 GB
- Context
- 128K
Qwen3.6 35B-A3B
- Estimated memory
- ~25 GB
- Context
- 262K native
Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: DeepSeek V4 Flash (Tight), Qwen3.8 Flash Next REAP-288 (Comfortable), gpt-oss-120b (Comfortable), Llama 4 Scout (Comfortable), Llama 3.1 70B (Comfortable), Qwen3.6 35B-A3B (Comfortable).