RTX 4090
Used-card condition, power supply and case clearance need separate checks.
Model card checked 2026-09-16 · Memory fit estimated; speed unknown unless a separately linked measurement matches your configuration.
Also searched as: qwen3.8 flash next
Stock checkpoint: 125B backbone + 51B n-gram embedding + 4B MTP. Q4 budget estimated as 180 × 0.5 × 1.2 = 108 GB before workload-specific KV cache. Not the REAP-288 conversion; do not transfer its 48 GB Mac result.
Primary sourcevllm serve Qwen/Qwen3.8-Flash-Next --dtype autoStock checkpoint: 125B backbone + 51B n-gram embedding + 4B MTP. Q4 budget estimated as 180 × 0.5 × 1.2 = 108 GB before workload-specific KV cache. Not the REAP-288 conversion; do not transfer its 48 GB Mac result.
Filter by a runtime path and an installed-capacity floor you already established for your workload. This does not calculate required VRAM. Inputs stay in this browser.
Used-card condition, power supply and case clearance need separate checks.
Check Blackwell kernels for the exact quantization; more VRAM does not certify faster decode.
Verify OS, driver and ROCm/Vulkan support. Experimental forks are not stock-runtime support.
BIOS GPU allocation and system use reduce available memory; 128 GB RAM is not 128 GB dedicated VRAM.
m5 ultra 96g and m5 ultra 96gb refer to this capacity. Memory is shared with macOS; verify shipping date and MLX conversion.
ARM64 packages and GB10 kernels must support the chosen model. Capacity is not a latency result.
| Issuer / checkpoint | Values | Source / verified |
|---|---|---|
| Alibaba Qwen / Qwen/Qwen3.8-Flash-Next | 180B total; 262,144 native; 1M extended; Qwen Community | Publisher model card / 2026-09-16 |
Inference experiments only. The Q4 memory budget is estimated, not a tested maximum context. Unsupported kernels, missing vision projectors, incorrect tool parsers and longer-context KV allocation can fail even when weights fit. No successful deployment, quality or speed is guaranteed.
Compare hardware purchase paths