Hardware decision workspace
BUYING DECISION / CHECKED 2026-09-16

Local LLM hardware: choose a GPU or a complete system

Choose between upgrading an existing PC and buying a complete shared-memory system. Your existing machine, software path and acceptable latency matter as much as installed capacity.

GPU upgrade or complete system?

An existing desktop with sufficient power, cooling and PCIe clearance can make a discrete GPU upgrade the simpler purchase. A Framework Desktop 128GB or Mac Studio M5 Ultra 96GB instead offers a large shared pool in one box, with different backend and allocation constraints.

DGX Spark local models require an ARM64 and GB10-compatible stack. Keep it on a separate system shortlist; its 128 GB capacity does not establish interactive decode latency. If your model is API-only, buying local hardware does not enable that model.

Your purchase shortlist

Filter by a runtime path and an installed-capacity floor you already established for your workload. This does not calculate required VRAM. Inputs stay in this browser.

RTX 4090

Used-card condition, power supply and case clearance need separate checks.

RTX 5090

Check Blackwell kernels for the exact quantization; more VRAM does not certify faster decode.

Radeon AI PRO R9700

Verify OS, driver and ROCm/Vulkan support. Experimental forks are not stock-runtime support.

Mac Studio M5 Ultra · 96GB

m5 ultra 96g and m5 ultra 96gb refer to this capacity. Memory is shared with macOS; verify shipping date and MLX conversion.

DGX Spark

ARM64 packages and GB10 kernels must support the chosen model. Capacity is not a latency result.

Download checklist

Versioned vendor specifications

Vendor capability facts, checked 2026-09-16; no independent speed claim. The capacity filter compares installed memory, not usable model memory.

Issuer / generationRule valuesSource / last verified
NVIDIA / RTX 409024 GB GDDR6X; Ada generationOfficial specifications
2026-09-16
NVIDIA / RTX 509032 GB GDDR7; Blackwell generationOfficial specifications
2026-09-16
AMD / Radeon AI PRO R970032 GB GDDR6; RDNA 4 workstation GPUOfficial specifications
2026-09-16
Framework / Framework Desktop Ryzen AI Max+ 395 · 128GBMax+ 395 configuration with 128 GB shared memoryOfficial specifications
2026-09-16
Apple / Mac Studio M5 Ultra · 96GB96 GB unified memory; M5 Ultra generationOfficial specifications
2026-09-16
NVIDIA / DGX Spark128 GB coherent memory; GB10 generationOfficial specifications
2026-09-16

Before you place an order

  1. Record checkpoint revision, quantization, context length and concurrent users. MoE active parameters do not determine total weight storage.
  2. Confirm the OS, driver, runtime version and architecture support using the model card and release notes.
  3. Ask for a matching run with prompt-processing, decode and first-token latency reported separately. A short-context single-stream result is not a long-context multi-user result.
  4. Record local delivered price, tax, host upgrades, power, cooling, warranty and returns. No prices or affiliate offers are maintained on this page.

Applicability and failure modes

This shortlist covers single GPUs and single-box inference purchases. It does not certify training, multi-GPU pooling, distributed serving, workload quality, maximum context or software compatibility.

Common failures include out-of-memory at longer context, unsupported quantization kernels, shared-memory allocation limits, thermal throttling, CPU offload causing unacceptable latency, and a multi-user workload exhausting KV cache. Passing a capacity filter guarantees neither model fit nor performance.

For tensorrt-llm moe fp8 fp4 long context release notes, check the versioned upstream releases against your GPU and checkpoint. FP4 weights and FP8 KV are separate capabilities. The query llamacpp-hrx has no confirmed upstream identity here; use the official llama.cpp repository and verify any fork before installing it.

Model directory and related decisions

Best GPU for local LLM: choose by workload and runtime