Back to guide
MODEL / INCLUSIONAI

Ling-3.0-flash

Model card checked 2026-09-16 · Memory fit estimated; speed unknown unless a separately linked measurement matches your configuration.

Also searched as: ling-3.0-flash

Model card states 124B total and 5.1B active; file metadata may count additional tensors. Q4 budget estimated as 124 × 0.5 × 1.2 = 74.4 GB before KV cache. Verify the supported runtime recipe and quantized artifact; speed is unknown.

Primary source
Total parameters124B
Estimated memory~74.4 GBQ4 estimated budget
ContextCheck selected runtime and checkpointText
LicenseMITMultilingual
01 / BEST FOR
  • source-checked local inference experiments
02 / ESTIMATED FIT
03 / DEPLOY WITH
vLLMAdvancedvllm serve inclusionAI/Ling-3.0-flash --dtype auto
Watch out

Model card states 124B total and 5.1B active; file metadata may count additional tensors. Q4 budget estimated as 124 × 0.5 × 1.2 = 74.4 GB before KV cache. Verify the supported runtime recipe and quantized artifact; speed is unknown.

Your purchase shortlist

Filter by a runtime path and an installed-capacity floor you already established for your workload. This does not calculate required VRAM. Inputs stay in this browser.

RTX 4090

Used-card condition, power supply and case clearance need separate checks.

RTX 5090

Check Blackwell kernels for the exact quantization; more VRAM does not certify faster decode.

Radeon AI PRO R9700

Verify OS, driver and ROCm/Vulkan support. Experimental forks are not stock-runtime support.

Mac Studio M5 Ultra · 96GB

m5 ultra 96g and m5 ultra 96gb refer to this capacity. Memory is shared with macOS; verify shipping date and MLX conversion.

DGX Spark

ARM64 packages and GB10 kernels must support the chosen model. Capacity is not a latency result.

Download checklist

Checked model specification

Issuer / checkpointValuesSource / verified
inclusionAI / inclusionAI/Ling-3.0-flash124B total; Check selected runtime and checkpoint; MITPublisher model card / 2026-09-16

Applicability and failure modes

Inference experiments only. The Q4 memory budget is estimated, not a tested maximum context. Unsupported kernels, missing vision projectors, incorrect tool parsers and longer-context KV allocation can fail even when weights fit. No successful deployment, quality or speed is guaranteed.

Compare hardware purchase paths