Panduan lapangan AI privat

Jalankan model yang tepat.
Di perangkat yang sudah Anda punya.

Pilih platform dan memori. Dapatkan kecocokan konservatif, server yang tepat, dan perintah yang bisa dijalankan.

Primary-source linked16 languagesBukti diperbarui 2026-08-31
8 GB
24 GB
80GB
ollama runqwen3:8b
LOCAL / 24G
MEMORY MAP01—80 GB
01 / FIT FINDER

Batas model lokal Anda

Estimasi mencakup bobot dan ruang dasar, tetapi bukan janji kecepatan.

Platform
Memori tersedia24 GB
43264128 GB
~22.8 GB model budget after reserve
RECOMMENDED ENVELOPE6 models
Lega Ketat
01
30B total~22.5 GBGGUF Q4 MoEOllama
ollama run qwen3:30b-a3b
02
27B total~20.5 GBINT4Ollama
ollama run gemma3:27b
04
21B total~16 GBMXFP4Ollama
ollama run gpt-oss:20b
05
14B total~11.2 GBGGUF Q4Ollama
ollama run qwen3:14b
06
12B total~9.4 GBINT4Ollama
ollama run gemma3:12b
02 / EVIDENCE FIRST

Peta, bukan peringkat

Kami memisahkan parameter total, MoE aktif, ukuran kuantisasi, dukungan runtime, dan kinerja terukur.

Weight fit ≠ speedTotal ≠ active parametersDocs ≠ benchmark
03 / HARDWARE MAP

Dari laptop hingga akselerator 80 GB

Setiap kelas punya anggaran, jalur runtime, dan batas realistis sendiri.

11
Semua perangkat
04 / MODEL FIELD NOTES

Model open-weight perwakilan

15
Semua model
05 / SERVING LAYER

Pilih lapisan inferensi

06
06 / METHOD

Cara estimasi dibuat

Kami memakai ukuran dan kuantisasi terdokumentasi, menyisakan ruang untuk sistem, dan menghitung cache sebagai tambahan. MoE memakai parameter total.

CONSERVATIVE FITweights + runtime + cache
total params×bits / 8+headroom
No universal tokens/s claims.
07 / FAQ

Sebelum mengunduh 40 GB

Does a 24 GB GPU run a 30B model?+

Often at Q4/INT4 with a conservative context. Qwen3 30B-A3B and Gemma 3 27B are representative fits, but cache and runtime overhead still matter.

Are active MoE parameters the memory requirement?+

No. Active parameters affect compute per token; total parameters still need to be stored in memory or offloaded.

Which runtime should a beginner choose?+

Ollama for a terminal-first setup or LM Studio for a visual desktop. llama.cpp is the portable fallback; vLLM is for higher-throughput GPU serving.

Do you benchmark speed?+

Not yet. Launch recommendations are transparent memory-fit estimates backed by primary documentation. We do not invent tokens-per-second numbers.