GeForce RTX 4090 · 24 GB
A 27B Q4 checkpoint can be a tight fit. Leave room for runtime overhead and KV cache instead of filling every gigabyte with weights.
一手来源Qwen3 30B-A3B
The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.
llama.cpp
GGUF models, CPU/GPU offload, embedded and unusual hardware. New architectures may require a recent build. Backend availability does not imply equal performance.
打开指南代表性开放权重模型
Qwen3 30B-A3B
- 估算内存
- ~22.5 GB
- 上下文
- 32K+
Qwen3.6 27B
- 估算内存
- ~18.5 GB
- 上下文
- 262K native
Qwen3.8 27B
- 估算内存
- ~19.5 GB
- 上下文
- 262K native
Gemma 3 27B
- 估算内存
- ~20.5 GB
- 上下文
- 128K
Devstral Small 2 24B
- 估算内存
- ~18.5 GB
- 上下文
- 256K
gpt-oss-20b
- 估算内存
- ~16 GB
- 上下文
- 128K
Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Qwen3 30B-A3B (Tight), Qwen3.6 27B (Tight), Qwen3.8 27B (Tight), Gemma 3 27B (Tight), Devstral Small 2 24B (Tight), gpt-oss-20b (Comfortable).