RTX 4090
Used-card condition, power supply and case clearance need separate checks.
There is no universal best GPU. Start with your exact checkpoint and runtime, then rule out capacity, software and power mismatches before comparing delivered prices.
Compare a 24 GB RTX 4090 when your verified workload stays within that tier and you already own a suitable PC. Compare a 32 GB RTX 5090 when the extra capacity is necessary and the selected runtime supports Blackwell. Neither is a speed recommendation without a matching test.
Radeon AI PRO R9700 offers a different 32 GB path. Shortlist it only after verifying your exact ROCm or Vulkan backend; a CUDA-only dependency is a reason to exclude it. Compare total system cost, not the GPU price alone.
Filter by a runtime path and an installed-capacity floor you already established for your workload. This does not calculate required VRAM. Inputs stay in this browser.
Used-card condition, power supply and case clearance need separate checks.
Check Blackwell kernels for the exact quantization; more VRAM does not certify faster decode.
Verify OS, driver and ROCm/Vulkan support. Experimental forks are not stock-runtime support.
Vendor capability facts, checked 2026-09-16; no independent speed claim. The capacity filter compares installed memory, not usable model memory.
| Issuer / generation | Rule values | Source / last verified |
|---|---|---|
| NVIDIA / RTX 4090 | 24 GB GDDR6X; Ada generation | Official specifications 2026-09-16 |
| NVIDIA / RTX 5090 | 32 GB GDDR7; Blackwell generation | Official specifications 2026-09-16 |
| AMD / Radeon AI PRO R9700 | 32 GB GDDR6; RDNA 4 workstation GPU | Official specifications 2026-09-16 |
This shortlist covers single GPUs and single-box inference purchases. It does not certify training, multi-GPU pooling, distributed serving, workload quality, maximum context or software compatibility.
Common failures include out-of-memory at longer context, unsupported quantization kernels, shared-memory allocation limits, thermal throttling, CPU offload causing unacceptable latency, and a multi-user workload exhausting KV cache. Passing a capacity filter guarantees neither model fit nor performance.
For tensorrt-llm moe fp8 fp4 long context release notes, check the versioned upstream releases against your GPU and checkpoint. FP4 weights and FP8 KV are separate capabilities. The query llamacpp-hrx has no confirmed upstream identity here; use the official llama.cpp repository and verify any fork before installing it.