ollama run qwen3:30b-a3bJalankan model yang tepat.
Di perangkat yang sudah Anda punya.
Pilih platform dan memori. Dapatkan kecocokan konservatif, server yang tepat, dan perintah yang bisa dijalankan.
ollama run gemma3:27bollama run <model-tag>ollama run gpt-oss:20bollama run qwen3:14bollama run gemma3:12bPeta, bukan peringkat
Kami memisahkan parameter total, MoE aktif, ukuran kuantisasi, dukungan runtime, dan kinerja terukur.
Dari laptop hingga akselerator 80 GB
Setiap kelas punya anggaran, jalur runtime, dan batas realistis sendiri.
8 GB CPU laptop
Small text models with short context. Expect patient, private inference rather than speed.
Buka panduan16 GB CPU desktop
Comfortable with 3B–8B Q4 models; 12B is possible only with reduced context and patience.
Buka panduanApple silicon · 16 GB
A polished 4B–8B local experience when the OS and apps have enough headroom.
Buka panduanNVIDIA · 8 GB
The mainstream 4B–8B tier. Some 12B INT4 builds fit tightly with modest context.
Buka panduanNVIDIA · 12 GB
Strong 8B–14B Q4 territory for a single user.
Buka panduanNVIDIA · 16 GB
14B models are comfortable; 20B-class low-bit MoE models are a tight upper edge.
Buka panduanModel open-weight perwakilan
Gemma 3 1B
- Estimasi memori
- ~1.4 GB
- Konteks
- 32K
Qwen3 4B
- Estimasi memori
- ~3.6 GB
- Konteks
- 32K+
Qwen3 8B
- Estimasi memori
- ~6.8 GB
- Konteks
- 32K+
gpt-oss-20b
- Estimasi memori
- ~16 GB
- Konteks
- 128K
Qwen3 30B-A3B
- Estimasi memori
- ~22.5 GB
- Konteks
- 32K+
Pilih lapisan inferensi
Ollama
One-command local chat and app integration
Buka panduanllama.cpp
GGUF models, CPU/GPU offload, embedded and unusual hardware
Buka panduanLM Studio
Discovering, downloading and testing models without a terminal
Buka panduanMLX LM
Native Apple silicon inference, experimentation and fine-tuning
Buka panduanvLLM
Linux GPU servers, concurrency and OpenAI-compatible production APIs
Buka panduanTensorRT-LLM
Maximum NVIDIA throughput after engine tuning
Buka panduanCara estimasi dibuat
Kami memakai ukuran dan kuantisasi terdokumentasi, menyisakan ruang untuk sistem, dan menghitung cache sebagai tambahan. MoE memakai parameter total.
Sebelum mengunduh 40 GB
Does a 24 GB GPU run a 30B model?+
Often at Q4/INT4 with a conservative context. Qwen3 30B-A3B and Gemma 3 27B are representative fits, but cache and runtime overhead still matter.
Are active MoE parameters the memory requirement?+
No. Active parameters affect compute per token; total parameters still need to be stored in memory or offloaded.
Which runtime should a beginner choose?+
Ollama for a terminal-first setup or LM Studio for a visual desktop. llama.cpp is the portable fallback; vLLM is for higher-throughput GPU serving.
Do you benchmark speed?+
Not yet. Launch recommendations are transparent memory-fit estimates backed by primary documentation. We do not invent tokens-per-second numbers.