Back to guide
MODEL / NEX AGI

Nex-N2.5-mini

Model card checked 2026-09-16 · Memory fit estimated; speed unknown unless a separately linked measurement matches your configuration.

Also searched as: nex-n2.5-mini

35B checkpoint. Q4 budget estimated as 35 × 0.5 × 1.2 = 21 GB before KV cache. Official BF16 weights are not an MLX 4-bit conversion. Verify conversion, tool parser and reasoning parser before deployment; speed is unknown.

Primary source
Total parameters35B
Estimated memory~21 GBQ4 estimated budget
ContextCheck selected runtime and checkpointText + Image
LicenseApache-2.0Multilingual
01 / BEST FOR
  • source-checked local inference experiments
02 / ESTIMATED FIT
03 / DEPLOY WITH
vLLMAdvancedvllm serve nex-agi/Nex-N2.5-mini --dtype auto
Watch out

35B checkpoint. Q4 budget estimated as 35 × 0.5 × 1.2 = 21 GB before KV cache. Official BF16 weights are not an MLX 4-bit conversion. Verify conversion, tool parser and reasoning parser before deployment; speed is unknown.

Your purchase shortlist

Filter by a runtime path and an installed-capacity floor you already established for your workload. This does not calculate required VRAM. Inputs stay in this browser.

RTX 4090

Used-card condition, power supply and case clearance need separate checks.

RTX 5090

Check Blackwell kernels for the exact quantization; more VRAM does not certify faster decode.

Radeon AI PRO R9700

Verify OS, driver and ROCm/Vulkan support. Experimental forks are not stock-runtime support.

Mac Studio M5 Ultra · 96GB

m5 ultra 96g and m5 ultra 96gb refer to this capacity. Memory is shared with macOS; verify shipping date and MLX conversion.

DGX Spark

ARM64 packages and GB10 kernels must support the chosen model. Capacity is not a latency result.

Download checklist

Checked model specification

Issuer / checkpointValuesSource / verified
Nex AGI / nex-agi/Nex-N2.5-mini35B total; Check selected runtime and checkpoint; Apache-2.0Publisher model card / 2026-09-16

Applicability and failure modes

Inference experiments only. The Q4 memory budget is estimated, not a tested maximum context. Unsupported kernels, missing vision projectors, incorrect tool parsers and longer-context KV allocation can fail even when weights fit. No successful deployment, quality or speed is guaranteed.

Compare hardware purchase paths