गाइड पर वापस
RUNTIME / ADVANCED

TensorRT-LLM

Maximum NVIDIA throughput after engine tuning

प्राथमिक स्रोत
DifficultyAdvancedSetup profile
Platforms1nvidia
APITriton or custom OpenAI-compatible serving
01 / INSTALL & RUN
  1. 1
    Installpython3 -m pip install tensorrt_llm -U --pre --extra-index-url https://pypi.nvidia.com
  2. 2
    Start a modeltrtllm-serve Qwen/Qwen3-8B
  3. 3
    Connect your app

    Triton or custom OpenAI-compatible serving. Keep the service bound to localhost unless you add authentication and network controls.

सावधानी

NVIDIA-only and sensitive to supported GPU, CUDA, model architecture and engine-build settings.

02 / COMPATIBLE MODELS

प्रतिनिधि ओपन-वेट मॉडल

Alibaba Qwen

Qwen3 8B

8BGGUF Q4
अनुमानित मेमोरी
~6.8 GB
कॉन्टेक्स्ट
32K+
Google

Gemma 3 12B

12BINT4
अनुमानित मेमोरी
~9.4 GB
कॉन्टेक्स्ट
128K
Alibaba Qwen

Qwen3 14B

14BGGUF Q4
अनुमानित मेमोरी
~11.2 GB
कॉन्टेक्स्ट
32K+
OpenAI

gpt-oss-20b

21B3.6B active · MXFP4
अनुमानित मेमोरी
~16 GB
कॉन्टेक्स्ट
128K
Mistral AI

Devstral Small 2 24B

24BGGUF Q4 / BF16
अनुमानित मेमोरी
~18.5 GB
कॉन्टेक्स्ट
256K
Google

Gemma 3 27B

27BINT4
अनुमानित मेमोरी
~20.5 GB
कॉन्टेक्स्ट
128K