Hardware decision workspace
BUYING DECISION / CHECKED 2026-09-16

Best GPU for local LLM: choose by workload and runtime

There is no universal best GPU. Start with your exact checkpoint and runtime, then rule out capacity, software and power mismatches before comparing delivered prices.

Which GPU path should you compare?

Compare a 24 GB RTX 4090 when your verified workload stays within that tier and you already own a suitable PC. Compare a 32 GB RTX 5090 when the extra capacity is necessary and the selected runtime supports Blackwell. Neither is a speed recommendation without a matching test.

Radeon AI PRO R9700 offers a different 32 GB path. Shortlist it only after verifying your exact ROCm or Vulkan backend; a CUDA-only dependency is a reason to exclude it. Compare total system cost, not the GPU price alone.

Your purchase shortlist

Filter by a runtime path and an installed-capacity floor you already established for your workload. This does not calculate required VRAM. Inputs stay in this browser.

RTX 4090

Used-card condition, power supply and case clearance need separate checks.

RTX 5090

Check Blackwell kernels for the exact quantization; more VRAM does not certify faster decode.

Radeon AI PRO R9700

Verify OS, driver and ROCm/Vulkan support. Experimental forks are not stock-runtime support.

Download checklist

Versioned vendor specifications

Vendor capability facts, checked 2026-09-16; no independent speed claim. The capacity filter compares installed memory, not usable model memory.

Issuer / generationRule valuesSource / last verified
NVIDIA / RTX 409024 GB GDDR6X; Ada generationOfficial specifications
2026-09-16
NVIDIA / RTX 509032 GB GDDR7; Blackwell generationOfficial specifications
2026-09-16
AMD / Radeon AI PRO R970032 GB GDDR6; RDNA 4 workstation GPUOfficial specifications
2026-09-16

Before you place an order

  1. Record checkpoint revision, quantization, context length and concurrent users. MoE active parameters do not determine total weight storage.
  2. Confirm the OS, driver, runtime version and architecture support using the model card and release notes.
  3. Ask for a matching run with prompt-processing, decode and first-token latency reported separately. A short-context single-stream result is not a long-context multi-user result.
  4. Record local delivered price, tax, host upgrades, power, cooling, warranty and returns. No prices or affiliate offers are maintained on this page.

Applicability and failure modes

This shortlist covers single GPUs and single-box inference purchases. It does not certify training, multi-GPU pooling, distributed serving, workload quality, maximum context or software compatibility.

Common failures include out-of-memory at longer context, unsupported quantization kernels, shared-memory allocation limits, thermal throttling, CPU offload causing unacceptable latency, and a multi-user workload exhausting KV cache. Passing a capacity filter guarantees neither model fit nor performance.

For tensorrt-llm moe fp8 fp4 long context release notes, check the versioned upstream releases against your GPU and checkpoint. FP4 weights and FP8 KV are separate capabilities. The query llamacpp-hrx has no confirmed upstream identity here; use the official llama.cpp repository and verify any fork before installing it.

Model directory and related decisions

Local LLM hardware: choose a GPU or a complete system