HARDWARE / VRAM
NVIDIA · 16 GB
14B models are comfortable; 20B-class low-bit MoE models are a tight upper edge.
RTX 4080 / 5080 class16 GB
01 / PRACTICAL CEILING
Qwen3 14B
The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.
02 / DEFAULT RUNTIME
vLLM
Linux GPU servers, concurrency and OpenAI-compatible production APIs. CUDA, ROCm and XPU feature coverage differs. Driver, shared-memory and quantization compatibility are version-specific.
ガイドを開く03 / MODEL ENVELOPE
代表的なオープンウェイトモデル
Alibaba Qwen
Qwen3 14B
14BGGUF Q4
- 推定メモリ
- ~11.2 GB
- コンテキスト
- 32K+
Google
Gemma 3 12B
12BINT4
- 推定メモリ
- ~9.4 GB
- コンテキスト
- 128K
Alibaba Qwen
Qwen3 8B
8BGGUF Q4
- 推定メモリ
- ~6.8 GB
- コンテキスト
- 32K+
Alibaba Qwen
Qwen3 4B
4BGGUF Q4
- 推定メモリ
- ~3.6 GB
- コンテキスト
- 32K+
Google
Gemma 3 4B
4BINT4 / GGUF Q4
- 推定メモリ
- ~4.2 GB
- コンテキスト
- 128K
Meta
Llama 3.2 3B
3BGGUF Q4
- 推定メモリ
- ~2.8 GB
- コンテキスト
- 128K
注意
Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Qwen3 14B (Tight), Gemma 3 12B (Comfortable), Qwen3 8B (Comfortable), Qwen3 4B (Comfortable), Gemma 3 4B (Comfortable), Llama 3.2 3B (Comfortable).