ollama run qwen3:30b-a3bDas richtige Modell.
Auf deiner vorhandenen Hardware.
Plattform und Speicher wählen. Du erhältst eine vorsichtige Passform, den passenden Inferenzserver und einen ausführbaren Befehl.
ollama run gemma3:27bollama run <model-tag>ollama run gpt-oss:20bollama run qwen3:14bollama run gemma3:12bEine Karte, keine Rangliste
Gesamtparameter, aktive MoE-Parameter, quantisierte Größe, Laufzeit-Support und Messwerte bleiben getrennt.
Vom leisen Laptop bis zum 80-GB-Beschleuniger
Jede Geräteklasse hat ein eigenes Speicherbudget, einen Pfad und eine ehrliche Obergrenze.
8 GB CPU laptop
Small text models with short context. Expect patient, private inference rather than speed.
Leitfaden öffnen16 GB CPU desktop
Comfortable with 3B–8B Q4 models; 12B is possible only with reduced context and patience.
Leitfaden öffnenApple silicon · 16 GB
A polished 4B–8B local experience when the OS and apps have enough headroom.
Leitfaden öffnenNVIDIA · 8 GB
The mainstream 4B–8B tier. Some 12B INT4 builds fit tightly with modest context.
Leitfaden öffnenNVIDIA · 12 GB
Strong 8B–14B Q4 territory for a single user.
Leitfaden öffnenNVIDIA · 16 GB
14B models are comfortable; 20B-class low-bit MoE models are a tight upper edge.
Leitfaden öffnenRepräsentative Open-Weight-Modelle
Gemma 3 1B
- Geschätzter Speicher
- ~1.4 GB
- Kontext
- 32K
Qwen3 4B
- Geschätzter Speicher
- ~3.6 GB
- Kontext
- 32K+
Qwen3 8B
- Geschätzter Speicher
- ~6.8 GB
- Kontext
- 32K+
gpt-oss-20b
- Geschätzter Speicher
- ~16 GB
- Kontext
- 128K
Qwen3 30B-A3B
- Geschätzter Speicher
- ~22.5 GB
- Kontext
- 32K+
Die passende Inferenzschicht
Ollama
One-command local chat and app integration
Leitfaden öffnenllama.cpp
GGUF models, CPU/GPU offload, embedded and unusual hardware
Leitfaden öffnenLM Studio
Discovering, downloading and testing models without a terminal
Leitfaden öffnenMLX LM
Native Apple silicon inference, experimentation and fine-tuning
Leitfaden öffnenvLLM
Linux GPU servers, concurrency and OpenAI-compatible production APIs
Leitfaden öffnenTensorRT-LLM
Maximum NVIDIA throughput after engine tuning
Leitfaden öffnenSo schätzen wir
Dokumentierte Größen und Quantisierung bilden die Basis. Wir reservieren Platz für System und Laufzeit; Kontextcache kommt hinzu. MoE nutzt Gesamtparameter.
Vor dem 40-GB-Download
Does a 24 GB GPU run a 30B model?+
Often at Q4/INT4 with a conservative context. Qwen3 30B-A3B and Gemma 3 27B are representative fits, but cache and runtime overhead still matter.
Are active MoE parameters the memory requirement?+
No. Active parameters affect compute per token; total parameters still need to be stored in memory or offloaded.
Which runtime should a beginner choose?+
Ollama for a terminal-first setup or LM Studio for a visual desktop. llama.cpp is the portable fallback; vLLM is for higher-throughput GPU serving.
Do you benchmark speed?+
Not yet. Launch recommendations are transparent memory-fit estimates backed by primary documentation. We do not invent tokens-per-second numbers.