Terug
RUNTIME / ADVANCED

TensorRT-LLM

Maximum NVIDIA throughput after engine tuning

Primaire bron
DifficultyAdvancedSetup profile
Platforms1nvidia
APITriton or custom OpenAI-compatible serving
01 / INSTALL & RUN
  1. 1
    Installpython3 -m pip install tensorrt_llm -U --pre --extra-index-url https://pypi.nvidia.com
  2. 2
    Start a modeltrtllm-serve Qwen/Qwen3-8B
  3. 3
    Connect your app

    Triton or custom OpenAI-compatible serving. Keep the service bound to localhost unless you add authentication and network controls.

Let op

NVIDIA-only and sensitive to supported GPU, CUDA, model architecture and engine-build settings.

02 / COMPATIBLE MODELS

Representatieve open-weightmodellen

Alibaba Qwen

Qwen3 8B

8BGGUF Q4
Geschat geheugen
~6.8 GB
Context
32K+
Google

Gemma 3 12B

12BINT4
Geschat geheugen
~9.4 GB
Context
128K
Alibaba Qwen

Qwen3 14B

14BGGUF Q4
Geschat geheugen
~11.2 GB
Context
32K+
OpenAI

gpt-oss-20b

21B3.6B active · MXFP4
Geschat geheugen
~16 GB
Context
128K
Mistral AI

Devstral Small 2 24B

24BGGUF Q4 / BF16
Geschat geheugen
~18.5 GB
Context
256K
Google

Gemma 3 27B

27BINT4
Geschat geheugen
~20.5 GB
Context
128K