Torna alla guida
HARDWARE / VRAM

Datacenter · 80 GB

Large MoE models and high-throughput servers, still bounded by quantization and cache.

H100 / H200 / MI300X class80 GB
01 / PRACTICAL CEILING

Llama 3.1 70B

The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.

02 / DEFAULT RUNTIME

TensorRT-LLM

Maximum NVIDIA throughput after engine tuning. NVIDIA-only and sensitive to supported GPU, CUDA, model architecture and engine-build settings.

Apri guida
03 / MODEL ENVELOPE

Modelli open-weight rappresentativi

Meta

Llama 3.1 70B

70BGGUF Q4
Memoria stimata
~46 GB
Contesto
128K
Alibaba Qwen

Qwen3 32B

32BGGUF Q4
Memoria stimata
~24.5 GB
Contesto
32K+
Alibaba Qwen

Qwen3 30B-A3B

30B3B active · GGUF Q4 MoE
Memoria stimata
~22.5 GB
Contesto
32K+
Google

Gemma 3 27B

27BINT4
Memoria stimata
~20.5 GB
Contesto
128K
Mistral AI

Devstral Small 2 24B

24BGGUF Q4 / BF16
Memoria stimata
~18.5 GB
Contesto
256K
OpenAI

gpt-oss-20b

21B3.6B active · MXFP4
Memoria stimata
~16 GB
Contesto
128K
Attenzione

Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Llama 3.1 70B (Comfortable), Qwen3 32B (Comfortable), Qwen3 30B-A3B (Comfortable), Gemma 3 27B (Comfortable), Devstral Small 2 24B (Comfortable), gpt-oss-20b (Comfortable).