RUNTIME / EASY
Unsloth
Running GGUF or MLX models from a desktop app that also trains and exports them
Primaire bronDifficultyEasySetup profile
Platforms5cpu · apple · nvidia · amd · intel
APIOpenAI + Anthropic-compatible:8888 (printed at launch)
01 / INSTALL & RUN
Installation depends on your OS. These commands are examples or templates: replace placeholders and verify quantization, file format, drivers and runtime support. Memory fit does not guarantee deployment.
- 1Install
curl -fsSL https://unsloth.ai/install.sh | sh # or download the Unsloth Desktop app - 2Start a model
unsloth run --model <verified-GGUF-or-MLX-repo-tag> - 3Connect your app
OpenAI + Anthropic-compatible · :8888 (printed at launch). Keep the service bound to localhost unless you add authentication and network controls.
Let op
Unsloth also fine-tunes, so its published VRAM savings are training numbers and not an inference fit claim. The local endpoint is an authenticated llama-server instance and the generated API key is shown once.
02 / COMPATIBLE MODELS
Representatieve open-weightmodellen
Alibaba Qwen
Qwen3.6 27B
27BQ4 / NVFP4 estimate
- Geschat geheugen
- ~18.5 GB
- Context
- 262K native
Alibaba Qwen
Qwen3.8 27B
27BQ4 estimate
- Geschat geheugen
- ~19.5 GB
- Context
- 262K native
Alibaba Qwen
Qwen3.6 35B-A3B
35B3B active · Q4 / NVFP4 MoE
- Geschat geheugen
- ~25 GB
- Context
- 262K native
DeepSeek
DeepSeek V4 Flash
284B13B active · Official FP4 + FP8 mixed
- Geschat geheugen
- ~176 GB
- Context
- 1M native
Google
Gemma 3 1B
1BINT4 / GGUF Q4
- Geschat geheugen
- ~1.4 GB
- Context
- 32K
Meta
Llama 3.2 3B
3BGGUF Q4
- Geschat geheugen
- ~2.8 GB
- Context
- 128K