Torna alla guida
RUNTIME / EASY

Lemonade

Ryzen AI, Radeon and Strix Halo PCs, including the XDNA2 NPU, behind one local API

Fonte primaria
DifficultyEasySetup profile
Platforms4amd · nvidia · cpu · apple
APIOpenAI + Ollama + Anthropic-compatible:13305
01 / INSTALL & RUN

Installation depends on your OS. These commands are examples or templates: replace placeholders and verify quantization, file format, drivers and runtime support. Memory fit does not guarantee deployment.

  1. 1
    Installsnap install lemonade-server # or the .msi / .deb / .pkg installer
  2. 2
    Start a modellemonade run <verified-model-name>
  3. 3
    Connect your app

    OpenAI + Ollama + Anthropic-compatible · :13305. Keep the service bound to localhost unless you add authentication and network controls.

Attenzione

Built with AMD engineers around Ryzen AI, Radeon and Strix Halo; CUDA, Vulkan, Metal and CPU backends cover other PCs. Run `lemonade backends` on the machine itself — the NPU path needs the vendor driver stack and is not available everywhere.

02 / COMPATIBLE MODELS

Modelli open-weight rappresentativi

Alibaba Qwen

Qwen3.6 27B

27BQ4 / NVFP4 estimate
Memoria stimata
~18.5 GB
Contesto
262K native
Alibaba Qwen

Qwen3.8 27B

27BQ4 estimate
Memoria stimata
~19.5 GB
Contesto
262K native
Alibaba Qwen

Qwen3.6 35B-A3B

35B3B active · Q4 / NVFP4 MoE
Memoria stimata
~25 GB
Contesto
262K native
DeepSeek

DeepSeek V4 Flash

284B13B active · Official FP4 + FP8 mixed
Memoria stimata
~176 GB
Contesto
1M native
Google

Gemma 3 1B

1BINT4 / GGUF Q4
Memoria stimata
~1.4 GB
Contesto
32K
Meta

Llama 3.2 3B

3BGGUF Q4
Memoria stimata
~2.8 GB
Contesto
128K