MacBook Pro M5 Max · 64 GB
A portable Apple-native tier for larger quantized models. Sustained inference still depends on thermals, runtime support, and context length.
一次情報Qwen3.8 Flash Next REAP-288
The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.
MLX LM
Native Apple silicon inference, experimentation and fine-tuning. Only for Apple silicon. Use an MLX-converted model and tune KV cache size for long contexts.
ガイドを開く代表的なオープンウェイトモデル
Qwen3.8 Flash Next REAP-288
- 推定メモリ
- ~40.57 GB
- コンテキスト
- 7.6K recommended on measured 48 GB Mac
Llama 3.1 70B
- 推定メモリ
- ~46 GB
- コンテキスト
- 128K
Qwen3.6 35B-A3B
- 推定メモリ
- ~25 GB
- コンテキスト
- 262K native
Qwen3 32B
- 推定メモリ
- ~24.5 GB
- コンテキスト
- 32K+
Qwen3 30B-A3B
- 推定メモリ
- ~22.5 GB
- コンテキスト
- 32K+
Qwen3.6 27B
- 推定メモリ
- ~18.5 GB
- コンテキスト
- 262K native
Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Qwen3.8 Flash Next REAP-288 (Tight), Llama 3.1 70B (Tight), Qwen3.6 35B-A3B (Comfortable), Qwen3 32B (Comfortable), Qwen3 30B-A3B (Comfortable), Qwen3.6 27B (Comfortable).