HARDWARE / UNIFIED MEMORY
Apple silicon · 64 GB
32B models with long context or 70B Q4 with conservative headroom.
M-series Max / Ultra64 GB
01 / PRACTICAL CEILING
Llama 3.1 70B
The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.
02 / DEFAULT RUNTIME
MLX LM
Native Apple silicon inference, experimentation and fine-tuning. Only for Apple silicon. Use an MLX-converted model and tune KV cache size for long contexts.
Rehberi aç03 / MODEL ENVELOPE
Temsilî açık ağırlıklı modeller
Meta
Llama 3.1 70B
70BGGUF Q4
- Tahmini bellek
- ~46 GB
- Bağlam
- 128K
Alibaba Qwen
Qwen3 32B
32BGGUF Q4
- Tahmini bellek
- ~24.5 GB
- Bağlam
- 32K+
Alibaba Qwen
Qwen3 30B-A3B
30B3B active · GGUF Q4 MoE
- Tahmini bellek
- ~22.5 GB
- Bağlam
- 32K+
Google
Gemma 3 27B
27BINT4
- Tahmini bellek
- ~20.5 GB
- Bağlam
- 128K
Mistral AI
Devstral Small 2 24B
24BGGUF Q4 / BF16
- Tahmini bellek
- ~18.5 GB
- Bağlam
- 256K
OpenAI
gpt-oss-20b
21B3.6B active · MXFP4
- Tahmini bellek
- ~16 GB
- Bağlam
- 128K
Dikkat
Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Llama 3.1 70B (Tight), Qwen3 32B (Comfortable), Qwen3 30B-A3B (Comfortable), Gemma 3 27B (Comfortable), Devstral Small 2 24B (Comfortable), gpt-oss-20b (Comfortable).