RUNTIME / ADVANCED
SGLang
Multi-request GPU serving with prefix caching and speculative decoding
Sumber primerDifficultyAdvancedSetup profile
Platforms4nvidia · amd · intel · cpu
APIOpenAI-compatible + native runtime API:30000
01 / INSTALL & RUN
Installation depends on your OS. These commands are examples or templates: replace placeholders and verify quantization, file format, drivers and runtime support. Memory fit does not guarantee deployment.
- 1Install
pip install uv && uv pip install --prerelease=allow sglang - 2Start a model
python3 -m sglang.launch_server --model-path <verified-checkpoint> --port 30000 - 3Connect your app
OpenAI-compatible + native runtime API · :30000. Keep the service bound to localhost unless you add authentication and network controls.
Perhatian
Each platform is a separate lane with its own floors: the NVIDIA lane requires CUDA 13, and 0.5.19 was the last CUDA 12 release. ROCm, Intel XPU, Xeon CPU and Apple Metal are documented separately, so support on one lane is not proof for another.
02 / COMPATIBLE MODELS
Model open-weight perwakilan
Alibaba Qwen
Qwen3.6 27B
27BQ4 / NVFP4 estimate
- Estimasi memori
- ~18.5 GB
- Konteks
- 262K native
Alibaba Qwen
Qwen3.8 27B
27BQ4 estimate
- Estimasi memori
- ~19.5 GB
- Konteks
- 262K native
Alibaba Qwen
Qwen3.6 35B-A3B
35B3B active · Q4 / NVFP4 MoE
- Estimasi memori
- ~25 GB
- Konteks
- 262K native
DeepSeek
DeepSeek V4 Flash
284B13B active · Official FP4 + FP8 mixed
- Estimasi memori
- ~176 GB
- Konteks
- 1M native
Google
Gemma 3 1B
1BINT4 / GGUF Q4
- Estimasi memori
- ~1.4 GB
- Konteks
- 32K
Meta
Llama 3.2 3B
3BGGUF Q4
- Estimasi memori
- ~2.8 GB
- Konteks
- 128K