Wróć
HARDWARE / UNIFIED MEMORY

Mac Studio M3 Ultra · 256 GB

Enough capacity for many 100B-plus low-bit models, provided the architecture is supported and optimized by the selected runtime.

Źródło pierwotne
Large-memory Apple desktop256 GB
01 / PRACTICAL CEILING

DeepSeek V4 Flash

The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.

02 / DEFAULT RUNTIME

MLX LM

Native Apple silicon inference, experimentation and fine-tuning. Only for Apple silicon. Use an MLX-converted model and tune KV cache size for long contexts.

Otwórz przewodnik
03 / MODEL ENVELOPE

Reprezentatywne modele open-weight

DeepSeek

DeepSeek V4 Flash

284B13B active · Official FP4 + FP8 mixed
Szacowana pamięć
~176 GB
Kontekst
1M native
Community REAP build

Qwen3.8 Flash Next REAP-288

176BREAP-288 · MLX 4-bit · PLE on NVMe
Szacowana pamięć
~40.57 GB
Kontekst
7.6K recommended on measured 48 GB Mac
OpenAI

gpt-oss-120b

117B5.1B active · MXFP4
Szacowana pamięć
~80 GB
Kontekst
128K
Meta

Llama 4 Scout

109B17B active · INT4 MoE
Szacowana pamięć
~78 GB
Kontekst
10M advertised
Meta

Llama 3.1 70B

70BGGUF Q4
Szacowana pamięć
~46 GB
Kontekst
128K
Alibaba Qwen

Qwen3.6 35B-A3B

35B3B active · Q4 / NVFP4 MoE
Szacowana pamięć
~25 GB
Kontekst
262K native
Uwaga

Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: DeepSeek V4 Flash (Tight), Qwen3.8 Flash Next REAP-288 (Comfortable), gpt-oss-120b (Comfortable), Llama 4 Scout (Comfortable), Llama 3.1 70B (Comfortable), Qwen3.6 35B-A3B (Comfortable).