RUNTIME / EASY
Lemonade
Ryzen AI, Radeon and Strix Halo PCs, including the XDNA2 NPU, behind one local API
PrimärquelleDifficultyEasySetup profile
Platforms4amd · nvidia · cpu · apple
APIOpenAI + Ollama + Anthropic-compatible:13305
01 / INSTALL & RUN
Installation depends on your OS. These commands are examples or templates: replace placeholders and verify quantization, file format, drivers and runtime support. Memory fit does not guarantee deployment.
- 1Install
snap install lemonade-server # or the .msi / .deb / .pkg installer - 2Start a model
lemonade run <verified-model-name> - 3Connect your app
OpenAI + Ollama + Anthropic-compatible · :13305. Keep the service bound to localhost unless you add authentication and network controls.
Achtung
Built with AMD engineers around Ryzen AI, Radeon and Strix Halo; CUDA, Vulkan, Metal and CPU backends cover other PCs. Run `lemonade backends` on the machine itself — the NPU path needs the vendor driver stack and is not available everywhere.
02 / COMPATIBLE MODELS
Repräsentative Open-Weight-Modelle
Alibaba Qwen
Qwen3.6 27B
27BQ4 / NVFP4 estimate
- Geschätzter Speicher
- ~18.5 GB
- Kontext
- 262K native
Alibaba Qwen
Qwen3.8 27B
27BQ4 estimate
- Geschätzter Speicher
- ~19.5 GB
- Kontext
- 262K native
Alibaba Qwen
Qwen3.6 35B-A3B
35B3B active · Q4 / NVFP4 MoE
- Geschätzter Speicher
- ~25 GB
- Kontext
- 262K native
DeepSeek
DeepSeek V4 Flash
284B13B active · Official FP4 + FP8 mixed
- Geschätzter Speicher
- ~176 GB
- Kontext
- 1M native
Google
Gemma 3 1B
1BINT4 / GGUF Q4
- Geschätzter Speicher
- ~1.4 GB
- Kontext
- 32K
Meta
Llama 3.2 3B
3BGGUF Q4
- Geschätzter Speicher
- ~2.8 GB
- Kontext
- 128K