HARDWARE / VRAM
NVIDIA · 24 GB
The consumer sweet spot for 27B and 30B-A3B INT4/Q4 models.
RTX 3090 / 4090 class24 GB
01 / PRACTICAL CEILING
Qwen3 30B-A3B
The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.
02 / DEFAULT RUNTIME
vLLM
Linux GPU servers, concurrency and OpenAI-compatible production APIs. CUDA, ROCm and XPU feature coverage differs. Driver, shared-memory and quantization compatibility are version-specific.
Buka panduan03 / MODEL ENVELOPE
Model open-weight perwakilan
Alibaba Qwen
Qwen3 30B-A3B
30B3B active · GGUF Q4 MoE
- Estimasi memori
- ~22.5 GB
- Konteks
- 32K+
Google
Gemma 3 27B
27BINT4
- Estimasi memori
- ~20.5 GB
- Konteks
- 128K
Mistral AI
Devstral Small 2 24B
24BGGUF Q4 / BF16
- Estimasi memori
- ~18.5 GB
- Konteks
- 256K
OpenAI
gpt-oss-20b
21B3.6B active · MXFP4
- Estimasi memori
- ~16 GB
- Konteks
- 128K
Alibaba Qwen
Qwen3 14B
14BGGUF Q4
- Estimasi memori
- ~11.2 GB
- Konteks
- 32K+
Google
Gemma 3 12B
12BINT4
- Estimasi memori
- ~9.4 GB
- Konteks
- 128K
Perhatian
Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Qwen3 30B-A3B (Tight), Gemma 3 27B (Tight), Devstral Small 2 24B (Tight), gpt-oss-20b (Comfortable), Qwen3 14B (Comfortable), Gemma 3 12B (Comfortable).