Mac Studio M3 Ultra · 256 GB
Enough capacity for many 100B-plus low-bit models, provided the architecture is supported and optimized by the selected runtime.
Fuente primariaDeepSeek V4 Flash
The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.
MLX LM
Native Apple silicon inference, experimentation and fine-tuning. Only for Apple silicon. Use an MLX-converted model and tune KV cache size for long contexts.
Abrir guíaModelos open-weight representativos
DeepSeek V4 Flash
- Memoria estimada
- ~176 GB
- Contexto
- 1M native
Qwen3.8 Flash Next REAP-288
- Memoria estimada
- ~40.57 GB
- Contexto
- 7.6K recommended on measured 48 GB Mac
gpt-oss-120b
- Memoria estimada
- ~80 GB
- Contexto
- 128K
Llama 4 Scout
- Memoria estimada
- ~78 GB
- Contexto
- 10M advertised
Llama 3.1 70B
- Memoria estimada
- ~46 GB
- Contexto
- 128K
Qwen3.6 35B-A3B
- Memoria estimada
- ~25 GB
- Contexto
- 262K native
Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: DeepSeek V4 Flash (Tight), Qwen3.8 Flash Next REAP-288 (Comfortable), gpt-oss-120b (Comfortable), Llama 4 Scout (Comfortable), Llama 3.1 70B (Comfortable), Qwen3.6 35B-A3B (Comfortable).