Workstation · 48 GB
70B Q4 serving begins here, with limited room for cache and concurrency.
Qwen3 32B
The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.
vLLM
Linux GPU servers, concurrency and OpenAI-compatible production APIs. CUDA, ROCm and XPU feature coverage differs. Driver, shared-memory and quantization compatibility are version-specific.
Открыть гидПоказательные модели с открытыми весами
Qwen3 32B
- Оценка памяти
- ~24.5 GB
- Контекст
- 32K+
Qwen3 30B-A3B
- Оценка памяти
- ~22.5 GB
- Контекст
- 32K+
Gemma 3 27B
- Оценка памяти
- ~20.5 GB
- Контекст
- 128K
Devstral Small 2 24B
- Оценка памяти
- ~18.5 GB
- Контекст
- 256K
gpt-oss-20b
- Оценка памяти
- ~16 GB
- Контекст
- 128K
Qwen3 14B
- Оценка памяти
- ~11.2 GB
- Контекст
- 32K+
Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Qwen3 32B (Comfortable), Qwen3 30B-A3B (Comfortable), Gemma 3 27B (Comfortable), Devstral Small 2 24B (Comfortable), gpt-oss-20b (Comfortable), Qwen3 14B (Comfortable).