NVIDIA DGX Spark · 128 GB
A compact local AI system with 128 GB coherent memory. It can hold large low-bit models, but decode speed and software support remain model-specific.
Первичный источникQwen3.8 Flash Next REAP-288
The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.
vLLM
Linux GPU servers, concurrency and OpenAI-compatible production APIs. CUDA, ROCm and XPU feature coverage differs. Driver, shared-memory and quantization compatibility are version-specific.
Открыть гидПоказательные модели с открытыми весами
Qwen3.8 Flash Next REAP-288
- Оценка памяти
- ~40.57 GB
- Контекст
- 7.6K recommended on measured 48 GB Mac
gpt-oss-120b
- Оценка памяти
- ~80 GB
- Контекст
- 128K
Llama 4 Scout
- Оценка памяти
- ~78 GB
- Контекст
- 10M advertised
Llama 3.1 70B
- Оценка памяти
- ~46 GB
- Контекст
- 128K
Qwen3.6 35B-A3B
- Оценка памяти
- ~25 GB
- Контекст
- 262K native
Qwen3 32B
- Оценка памяти
- ~24.5 GB
- Контекст
- 32K+
Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Qwen3.8 Flash Next REAP-288 (Comfortable), gpt-oss-120b (Comfortable), Llama 4 Scout (Comfortable), Llama 3.1 70B (Comfortable), Qwen3.6 35B-A3B (Comfortable), Qwen3 32B (Comfortable).