Volver a la guía
RUNTIME / ADVANCED

TensorRT-LLM

Maximum NVIDIA throughput after engine tuning

Fuente primaria
DifficultyAdvancedSetup profile
Platforms1nvidia
APITriton or custom OpenAI-compatible serving
01 / INSTALL & RUN
  1. 1
    Installpython3 -m pip install tensorrt_llm -U --pre --extra-index-url https://pypi.nvidia.com
  2. 2
    Start a modeltrtllm-serve Qwen/Qwen3-8B
  3. 3
    Connect your app

    Triton or custom OpenAI-compatible serving. Keep the service bound to localhost unless you add authentication and network controls.

Atención

NVIDIA-only and sensitive to supported GPU, CUDA, model architecture and engine-build settings.

02 / COMPATIBLE MODELS

Modelos open-weight representativos

Alibaba Qwen

Qwen3 8B

8BGGUF Q4
Memoria estimada
~6.8 GB
Contexto
32K+
Google

Gemma 3 12B

12BINT4
Memoria estimada
~9.4 GB
Contexto
128K
Alibaba Qwen

Qwen3 14B

14BGGUF Q4
Memoria estimada
~11.2 GB
Contexto
32K+
OpenAI

gpt-oss-20b

21B3.6B active · MXFP4
Memoria estimada
~16 GB
Contexto
128K
Mistral AI

Devstral Small 2 24B

24BGGUF Q4 / BF16
Memoria estimada
~18.5 GB
Contexto
256K
Google

Gemma 3 27B

27BINT4
Memoria estimada
~20.5 GB
Contexto
128K