Terug
RUNTIME / MEDIUM

llama.cpp

GGUF models, CPU/GPU offload, embedded and unusual hardware

Primaire bron
DifficultyMediumSetup profile
Platforms5cpu · apple · nvidia · amd · intel
APIOpenAI-compatible server:8080
01 / INSTALL & RUN
  1. 1
    Installbrew install llama.cpp # or use upstream binaries
  2. 2
    Start a modelllama-server -hf google/gemma-3-1b-it --port 8080
  3. 3
    Connect your app

    OpenAI-compatible server · :8080. Keep the service bound to localhost unless you add authentication and network controls.

Let op

New architectures may require a recent build. Backend availability does not imply equal performance.

02 / COMPATIBLE MODELS

Representatieve open-weightmodellen

Google

Gemma 3 1B

1BINT4 / GGUF Q4
Geschat geheugen
~1.4 GB
Context
32K
Meta

Llama 3.2 3B

3BGGUF Q4
Geschat geheugen
~2.8 GB
Context
128K
Alibaba Qwen

Qwen3 4B

4BGGUF Q4
Geschat geheugen
~3.6 GB
Context
32K+
Google

Gemma 3 4B

4BINT4 / GGUF Q4
Geschat geheugen
~4.2 GB
Context
128K
Alibaba Qwen

Qwen3 8B

8BGGUF Q4
Geschat geheugen
~6.8 GB
Context
32K+
Google

Gemma 3 12B

12BINT4
Geschat geheugen
~9.4 GB
Context
128K