HARDWARE / VRAM
NVIDIA · 8 GB
The mainstream 4B–8B tier. Some 12B INT4 builds fit tightly with modest context.
RTX 4060 class8 GB
01 / PRACTICAL CEILING
Qwen3 8B
The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.
02 / DEFAULT RUNTIME
Ollama
One-command local chat and app integration. GPU and driver support varies by operating system. Confirm the current hardware matrix before buying hardware.
Open guide03 / MODEL ENVELOPE
Representative open-weight models
Alibaba Qwen
Qwen3 8B
8BGGUF Q4
- Estimated memory
- ~6.8 GB
- Context
- 32K+
Alibaba Qwen
Qwen3 4B
4BGGUF Q4
- Estimated memory
- ~3.6 GB
- Context
- 32K+
Google
Gemma 3 4B
4BINT4 / GGUF Q4
- Estimated memory
- ~4.2 GB
- Context
- 128K
Meta
Llama 3.2 3B
3BGGUF Q4
- Estimated memory
- ~2.8 GB
- Context
- 128K
Google
Gemma 3 1B
1BINT4 / GGUF Q4
- Estimated memory
- ~1.4 GB
- Context
- 32K
Watch out
Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Qwen3 8B (Tight), Qwen3 4B (Comfortable), Gemma 3 4B (Comfortable), Llama 3.2 3B (Comfortable), Gemma 3 1B (Comfortable).