RUNTIME / EASY
llamafile
Distributing a whole GGUF model as one runnable executable on Linux, macOS, Windows and BSD
Fonte primariaDifficultyEasySetup profile
Platforms5cpu · apple · nvidia · amd · intel
APIOpenAI + Anthropic-compatible:8080
01 / INSTALL & RUN
Installation depends on your OS. These commands are examples or templates: replace placeholders and verify quantization, file format, drivers and runtime support. Memory fit does not guarantee deployment.
- 1Install
Download a prebuilt .llamafile for your model, or the llamafile binary plus an external GGUF - 2Start a model
./llamafile -m <downloaded-model.gguf> -ngl 999 - 3Connect your app
OpenAI + Anthropic-compatible · :8080. Keep the service bound to localhost unless you add authentication and network controls.
Attenzione
Windows can only execute files under 4 GB, so larger models need the standalone binary with external GGUF weights. Vulkan covers Intel and other GPUs, but AMD multi-GPU offload can break.
02 / COMPATIBLE MODELS
Modelli open-weight rappresentativi
Alibaba Qwen
Qwen3.6 27B
27BQ4 / NVFP4 estimate
- Memoria stimata
- ~18.5 GB
- Contesto
- 262K native
Alibaba Qwen
Qwen3.8 27B
27BQ4 estimate
- Memoria stimata
- ~19.5 GB
- Contesto
- 262K native
Alibaba Qwen
Qwen3.6 35B-A3B
35B3B active · Q4 / NVFP4 MoE
- Memoria stimata
- ~25 GB
- Contesto
- 262K native
DeepSeek
DeepSeek V4 Flash
284B13B active · Official FP4 + FP8 mixed
- Memoria stimata
- ~176 GB
- Contesto
- 1M native
Google
Gemma 3 1B
1BINT4 / GGUF Q4
- Memoria stimata
- ~1.4 GB
- Contesto
- 32K
Meta
Llama 3.2 3B
3BGGUF Q4
- Memoria stimata
- ~2.8 GB
- Contesto
- 128K