Zurück
RUNTIME / ADVANCED

TensorRT-LLM

Maximum NVIDIA throughput after engine tuning

Primärquelle
DifficultyAdvancedSetup profile
Platforms1nvidia
APITriton or custom OpenAI-compatible serving
01 / INSTALL & RUN
  1. 1
    Installpython3 -m pip install tensorrt_llm -U --pre --extra-index-url https://pypi.nvidia.com
  2. 2
    Start a modeltrtllm-serve Qwen/Qwen3-8B
  3. 3
    Connect your app

    Triton or custom OpenAI-compatible serving. Keep the service bound to localhost unless you add authentication and network controls.

Achtung

NVIDIA-only and sensitive to supported GPU, CUDA, model architecture and engine-build settings.

02 / COMPATIBLE MODELS

Repräsentative Open-Weight-Modelle

Alibaba Qwen

Qwen3 8B

8BGGUF Q4
Geschätzter Speicher
~6.8 GB
Kontext
32K+
Google

Gemma 3 12B

12BINT4
Geschätzter Speicher
~9.4 GB
Kontext
128K
Alibaba Qwen

Qwen3 14B

14BGGUF Q4
Geschätzter Speicher
~11.2 GB
Kontext
32K+
OpenAI

gpt-oss-20b

21B3.6B active · MXFP4
Geschätzter Speicher
~16 GB
Kontext
128K
Mistral AI

Devstral Small 2 24B

24BGGUF Q4 / BF16
Geschätzter Speicher
~18.5 GB
Kontext
256K
Google

Gemma 3 27B

27BINT4
Geschätzter Speicher
~20.5 GB
Kontext
128K