RUNTIME / ADVANCED
TensorRT-LLM
Maximum NVIDIA throughput after engine tuning
Primaire bronDifficultyAdvancedSetup profile
Platforms1nvidia
APITriton or custom OpenAI-compatible serving
01 / INSTALL & RUN
- 1Install
python3 -m pip install tensorrt_llm -U --pre --extra-index-url https://pypi.nvidia.com - 2Start a model
trtllm-serve Qwen/Qwen3-8B - 3Connect your app
Triton or custom OpenAI-compatible serving. Keep the service bound to localhost unless you add authentication and network controls.
Let op
NVIDIA-only and sensitive to supported GPU, CUDA, model architecture and engine-build settings.
02 / COMPATIBLE MODELS
Representatieve open-weightmodellen
Alibaba Qwen
Qwen3 8B
8BGGUF Q4
- Geschat geheugen
- ~6.8 GB
- Context
- 32K+
Google
Gemma 3 12B
12BINT4
- Geschat geheugen
- ~9.4 GB
- Context
- 128K
Alibaba Qwen
Qwen3 14B
14BGGUF Q4
- Geschat geheugen
- ~11.2 GB
- Context
- 32K+
OpenAI
gpt-oss-20b
21B3.6B active · MXFP4
- Geschat geheugen
- ~16 GB
- Context
- 128K
Mistral AI
Devstral Small 2 24B
24BGGUF Q4 / BF16
- Geschat geheugen
- ~18.5 GB
- Context
- 256K
Google
Gemma 3 27B
27BINT4
- Geschat geheugen
- ~20.5 GB
- Context
- 128K