Rehbere dön
RUNTIME / EASY

KoboldCpp

One-file GGUF chat on Windows, macOS or Linux with no Python install

Birincil kaynak
DifficultyEasySetup profile
Platforms5cpu · nvidia · amd · intel · apple
APIOpenAI + Ollama + KoboldCpp-compatible:5001
01 / INSTALL & RUN

Installation depends on your OS. These commands are examples or templates: replace placeholders and verify quantization, file format, drivers and runtime support. Memory fit does not guarantee deployment.

  1. 1
    InstallDownload the single-file release binary from the official GitHub releases (koboldcpp.exe on Windows)
  2. 2
    Start a modelkoboldcpp --model <downloaded-model.gguf> --usecuda
  3. 3
    Connect your app

    OpenAI + Ollama + KoboldCpp-compatible · :5001. Keep the service bound to localhost unless you add authentication and network controls.

Dikkat

The project warns that koboldcpp.com is an impersonation, so only GitHub release binaries are official. GPU flags differ by vendor and OS, and you can only use the accelerator libraries the release bundles.

02 / COMPATIBLE MODELS

Temsilî açık ağırlıklı modeller

Alibaba Qwen

Qwen3.6 27B

27BQ4 / NVFP4 estimate
Tahmini bellek
~18.5 GB
Bağlam
262K native
Alibaba Qwen

Qwen3.8 27B

27BQ4 estimate
Tahmini bellek
~19.5 GB
Bağlam
262K native
Alibaba Qwen

Qwen3.6 35B-A3B

35B3B active · Q4 / NVFP4 MoE
Tahmini bellek
~25 GB
Bağlam
262K native
DeepSeek

DeepSeek V4 Flash

284B13B active · Official FP4 + FP8 mixed
Tahmini bellek
~176 GB
Bağlam
1M native
Google

Gemma 3 1B

1BINT4 / GGUF Q4
Tahmini bellek
~1.4 GB
Bağlam
32K
Meta

Llama 3.2 3B

3BGGUF Q4
Tahmini bellek
~2.8 GB
Bağlam
128K