Back to guide
MODEL / ALIBABA QWEN

Qwen3.8 Flash Next

Model card checked 2026-09-16 · Memory fit estimated; speed unknown unless a separately linked measurement matches your configuration.

Also searched as: qwen3.8 flash next

Stock checkpoint: 125B backbone + 51B n-gram embedding + 4B MTP. Q4 budget estimated as 180 × 0.5 × 1.2 = 108 GB before workload-specific KV cache. Not the REAP-288 conversion; do not transfer its 48 GB Mac result.

Primary source
Total parameters180B
Estimated memory~108 GBQ4 estimated budget
Context262,144 native; 1M extendedText + Image
LicenseQwen CommunityMultilingual
01 / BEST FOR
  • source-checked local inference experiments
02 / ESTIMATED FIT
03 / DEPLOY WITH
vLLMAdvancedvllm serve Qwen/Qwen3.8-Flash-Next --dtype auto
Watch out

Stock checkpoint: 125B backbone + 51B n-gram embedding + 4B MTP. Q4 budget estimated as 180 × 0.5 × 1.2 = 108 GB before workload-specific KV cache. Not the REAP-288 conversion; do not transfer its 48 GB Mac result.

Your purchase shortlist

Filter by a runtime path and an installed-capacity floor you already established for your workload. This does not calculate required VRAM. Inputs stay in this browser.

RTX 4090

Used-card condition, power supply and case clearance need separate checks.

RTX 5090

Check Blackwell kernels for the exact quantization; more VRAM does not certify faster decode.

Radeon AI PRO R9700

Verify OS, driver and ROCm/Vulkan support. Experimental forks are not stock-runtime support.

Mac Studio M5 Ultra · 96GB

m5 ultra 96g and m5 ultra 96gb refer to this capacity. Memory is shared with macOS; verify shipping date and MLX conversion.

DGX Spark

ARM64 packages and GB10 kernels must support the chosen model. Capacity is not a latency result.

Download checklist

Checked model specification

Issuer / checkpointValuesSource / verified
Alibaba Qwen / Qwen/Qwen3.8-Flash-Next180B total; 262,144 native; 1M extended; Qwen CommunityPublisher model card / 2026-09-16

Applicability and failure modes

Inference experiments only. The Q4 memory budget is estimated, not a tested maximum context. Unsupported kernels, missing vision projectors, incorrect tool parsers and longer-context KV allocation can fail even when weights fit. No successful deployment, quality or speed is guaranteed.

Compare hardware purchase paths