ollama run qwen3:30b-a3bDoğru modeli çalıştır.
Sahip olduğun donanımda.
Platform ve belleği seç. Temkinli model uyumu, doğru sunucu ve çalıştırılabilir komut al.
ollama run gemma3:27bollama run <model-tag>ollama run gpt-oss:20bollama run qwen3:14bollama run gemma3:12bSıralama değil, harita
Toplam parametre, aktif MoE, nicemlenmiş boyut, çalışma desteği ve ölçümü ayrı tutarız.
Dizüstünden 80 GB hızlandırıcıya
Her sınıfın farklı bütçesi, yolu ve gerçekçi üst sınırı vardır.
8 GB CPU laptop
Small text models with short context. Expect patient, private inference rather than speed.
Rehberi aç16 GB CPU desktop
Comfortable with 3B–8B Q4 models; 12B is possible only with reduced context and patience.
Rehberi açApple silicon · 16 GB
A polished 4B–8B local experience when the OS and apps have enough headroom.
Rehberi açNVIDIA · 8 GB
The mainstream 4B–8B tier. Some 12B INT4 builds fit tightly with modest context.
Rehberi açNVIDIA · 12 GB
Strong 8B–14B Q4 territory for a single user.
Rehberi açNVIDIA · 16 GB
14B models are comfortable; 20B-class low-bit MoE models are a tight upper edge.
Rehberi açTemsilî açık ağırlıklı modeller
Gemma 3 1B
- Tahmini bellek
- ~1.4 GB
- Bağlam
- 32K
Qwen3 4B
- Tahmini bellek
- ~3.6 GB
- Bağlam
- 32K+
Qwen3 8B
- Tahmini bellek
- ~6.8 GB
- Bağlam
- 32K+
gpt-oss-20b
- Tahmini bellek
- ~16 GB
- Bağlam
- 128K
Qwen3 30B-A3B
- Tahmini bellek
- ~22.5 GB
- Bağlam
- 32K+
Doğru çıkarım katmanını seç
Ollama
One-command local chat and app integration
Rehberi açllama.cpp
GGUF models, CPU/GPU offload, embedded and unusual hardware
Rehberi açLM Studio
Discovering, downloading and testing models without a terminal
Rehberi açMLX LM
Native Apple silicon inference, experimentation and fine-tuning
Rehberi açvLLM
Linux GPU servers, concurrency and OpenAI-compatible production APIs
Rehberi açTensorRT-LLM
Maximum NVIDIA throughput after engine tuning
Rehberi açNasıl tahmin ediyoruz
Belgeli boyut ve nicemlemeden başlar, sistem için pay bırakır ve önbelleği ek maliyet sayarız. MoE toplam parametreyle hesaplanır.
40 GB indirmeden önce
Does a 24 GB GPU run a 30B model?+
Often at Q4/INT4 with a conservative context. Qwen3 30B-A3B and Gemma 3 27B are representative fits, but cache and runtime overhead still matter.
Are active MoE parameters the memory requirement?+
No. Active parameters affect compute per token; total parameters still need to be stored in memory or offloaded.
Which runtime should a beginner choose?+
Ollama for a terminal-first setup or LM Studio for a visual desktop. llama.cpp is the portable fallback; vLLM is for higher-throughput GPU serving.
Do you benchmark speed?+
Not yet. Launch recommendations are transparent memory-fit estimates backed by primary documentation. We do not invent tokens-per-second numbers.