RUNTIME / EASY
KoboldCpp
One-file GGUF chat on Windows, macOS or Linux with no Python install
Fonte primáriaDifficultyEasySetup profile
Platforms5cpu · nvidia · amd · intel · apple
APIOpenAI + Ollama + KoboldCpp-compatible:5001
01 / INSTALL & RUN
Installation depends on your OS. These commands are examples or templates: replace placeholders and verify quantization, file format, drivers and runtime support. Memory fit does not guarantee deployment.
- 1Install
Download the single-file release binary from the official GitHub releases (koboldcpp.exe on Windows) - 2Start a model
koboldcpp --model <downloaded-model.gguf> --usecuda - 3Connect your app
OpenAI + Ollama + KoboldCpp-compatible · :5001. Keep the service bound to localhost unless you add authentication and network controls.
Atenção
The project warns that koboldcpp.com is an impersonation, so only GitHub release binaries are official. GPU flags differ by vendor and OS, and you can only use the accelerator libraries the release bundles.
02 / COMPATIBLE MODELS
Modelos open-weight representativos
Alibaba Qwen
Qwen3.6 27B
27BQ4 / NVFP4 estimate
- Memória estimada
- ~18.5 GB
- Contexto
- 262K native
Alibaba Qwen
Qwen3.8 27B
27BQ4 estimate
- Memória estimada
- ~19.5 GB
- Contexto
- 262K native
Alibaba Qwen
Qwen3.6 35B-A3B
35B3B active · Q4 / NVFP4 MoE
- Memória estimada
- ~25 GB
- Contexto
- 262K native
DeepSeek
DeepSeek V4 Flash
284B13B active · Official FP4 + FP8 mixed
- Memória estimada
- ~176 GB
- Contexto
- 1M native
Google
Gemma 3 1B
1BINT4 / GGUF Q4
- Memória estimada
- ~1.4 GB
- Contexto
- 32K
Meta
Llama 3.2 3B
3BGGUF Q4
- Memória estimada
- ~2.8 GB
- Contexto
- 128K