Rehbere dön
HARDWARE / VRAM

GeForce RTX 4090 · 24 GB

A 27B Q4 checkpoint can be a tight fit. Leave room for runtime overhead and KV cache instead of filling every gigabyte with weights.

Birincil kaynak
Consumer Ada GPU24 GB
01 / PRACTICAL CEILING

Qwen3 30B-A3B

The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.

02 / DEFAULT RUNTIME

llama.cpp

GGUF models, CPU/GPU offload, embedded and unusual hardware. New architectures may require a recent build. Backend availability does not imply equal performance.

Rehberi aç
03 / MODEL ENVELOPE

Temsilî açık ağırlıklı modeller

Alibaba Qwen

Qwen3 30B-A3B

30B3B active · GGUF Q4 MoE
Tahmini bellek
~22.5 GB
Bağlam
32K+
Alibaba Qwen

Qwen3.6 27B

27BQ4 / NVFP4 estimate
Tahmini bellek
~18.5 GB
Bağlam
262K native
Alibaba Qwen

Qwen3.8 27B

27BQ4 estimate
Tahmini bellek
~19.5 GB
Bağlam
262K native
Google

Gemma 3 27B

27BINT4
Tahmini bellek
~20.5 GB
Bağlam
128K
Mistral AI

Devstral Small 2 24B

24BGGUF Q4 / BF16
Tahmini bellek
~18.5 GB
Bağlam
256K
OpenAI

gpt-oss-20b

21B3.6B active · MXFP4
Tahmini bellek
~16 GB
Bağlam
128K
Dikkat

Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Qwen3 30B-A3B (Tight), Qwen3.6 27B (Tight), Qwen3.8 27B (Tight), Gemma 3 27B (Tight), Devstral Small 2 24B (Tight), gpt-oss-20b (Comfortable).