Retour au guide
RUNTIME / ADVANCED

TensorRT-LLM

Maximum NVIDIA throughput after engine tuning

Source primaire
DifficultyAdvancedSetup profile
Platforms1nvidia
APITriton or custom OpenAI-compatible serving
01 / INSTALL & RUN
  1. 1
    Installpython3 -m pip install tensorrt_llm -U --pre --extra-index-url https://pypi.nvidia.com
  2. 2
    Start a modeltrtllm-serve Qwen/Qwen3-8B
  3. 3
    Connect your app

    Triton or custom OpenAI-compatible serving. Keep the service bound to localhost unless you add authentication and network controls.

Attention

NVIDIA-only and sensitive to supported GPU, CUDA, model architecture and engine-build settings.

02 / COMPATIBLE MODELS

Modèles open-weight représentatifs

Alibaba Qwen

Qwen3 8B

8BGGUF Q4
Mémoire estimée
~6.8 GB
Contexte
32K+
Google

Gemma 3 12B

12BINT4
Mémoire estimée
~9.4 GB
Contexte
128K
Alibaba Qwen

Qwen3 14B

14BGGUF Q4
Mémoire estimée
~11.2 GB
Contexte
32K+
OpenAI

gpt-oss-20b

21B3.6B active · MXFP4
Mémoire estimée
~16 GB
Contexte
128K
Mistral AI

Devstral Small 2 24B

24BGGUF Q4 / BF16
Mémoire estimée
~18.5 GB
Contexte
256K
Google

Gemma 3 27B

27BINT4
Mémoire estimée
~20.5 GB
Contexte
128K