Terug
RUNTIME / EASY

llamafile

Distributing a whole GGUF model as one runnable executable on Linux, macOS, Windows and BSD

Primaire bron
DifficultyEasySetup profile
Platforms5cpu · apple · nvidia · amd · intel
APIOpenAI + Anthropic-compatible:8080
01 / INSTALL & RUN

Installation depends on your OS. These commands are examples or templates: replace placeholders and verify quantization, file format, drivers and runtime support. Memory fit does not guarantee deployment.

  1. 1
    InstallDownload a prebuilt .llamafile for your model, or the llamafile binary plus an external GGUF
  2. 2
    Start a model./llamafile -m <downloaded-model.gguf> -ngl 999
  3. 3
    Connect your app

    OpenAI + Anthropic-compatible · :8080. Keep the service bound to localhost unless you add authentication and network controls.

Let op

Windows can only execute files under 4 GB, so larger models need the standalone binary with external GGUF weights. Vulkan covers Intel and other GPUs, but AMD multi-GPU offload can break.

02 / COMPATIBLE MODELS

Representatieve open-weightmodellen

Alibaba Qwen

Qwen3.6 27B

27BQ4 / NVFP4 estimate
Geschat geheugen
~18.5 GB
Context
262K native
Alibaba Qwen

Qwen3.8 27B

27BQ4 estimate
Geschat geheugen
~19.5 GB
Context
262K native
Alibaba Qwen

Qwen3.6 35B-A3B

35B3B active · Q4 / NVFP4 MoE
Geschat geheugen
~25 GB
Context
262K native
DeepSeek

DeepSeek V4 Flash

284B13B active · Official FP4 + FP8 mixed
Geschat geheugen
~176 GB
Context
1M native
Google

Gemma 3 1B

1BINT4 / GGUF Q4
Geschat geheugen
~1.4 GB
Context
32K
Meta

Llama 3.2 3B

3BGGUF Q4
Geschat geheugen
~2.8 GB
Context
128K