GeForce RTX 5090 · 32 GB
A fast single-GPU tier for 27B-class low-bit models, with limited remaining space for very long context or multiple concurrent users.
PrimärquelleQwen3.6 35B-A3B
The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.
vLLM
Linux GPU servers, concurrency and OpenAI-compatible production APIs. CUDA, ROCm and XPU feature coverage differs. Driver, shared-memory and quantization compatibility are version-specific.
Leitfaden öffnenRepräsentative Open-Weight-Modelle
Qwen3.6 35B-A3B
- Geschätzter Speicher
- ~25 GB
- Kontext
- 262K native
Qwen3 32B
- Geschätzter Speicher
- ~24.5 GB
- Kontext
- 32K+
Qwen3 30B-A3B
- Geschätzter Speicher
- ~22.5 GB
- Kontext
- 32K+
Qwen3.6 27B
- Geschätzter Speicher
- ~18.5 GB
- Kontext
- 262K native
Qwen3.8 27B
- Geschätzter Speicher
- ~19.5 GB
- Kontext
- 262K native
Gemma 3 27B
- Geschätzter Speicher
- ~20.5 GB
- Kontext
- 128K
Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Qwen3.6 35B-A3B (Tight), Qwen3 32B (Tight), Qwen3 30B-A3B (Tight), Qwen3.6 27B (Comfortable), Qwen3.8 27B (Comfortable), Gemma 3 27B (Comfortable).