RTX 4090
Used-card condition, power supply and case clearance need separate checks.
Model card checked 2026-09-16 · Memory fit estimated; speed unknown unless a separately linked measurement matches your configuration.
Also searched as: nex-n2.5-mini
35B checkpoint. Q4 budget estimated as 35 × 0.5 × 1.2 = 21 GB before KV cache. Official BF16 weights are not an MLX 4-bit conversion. Verify conversion, tool parser and reasoning parser before deployment; speed is unknown.
Primary sourcevllm serve nex-agi/Nex-N2.5-mini --dtype auto35B checkpoint. Q4 budget estimated as 35 × 0.5 × 1.2 = 21 GB before KV cache. Official BF16 weights are not an MLX 4-bit conversion. Verify conversion, tool parser and reasoning parser before deployment; speed is unknown.
Filter by a runtime path and an installed-capacity floor you already established for your workload. This does not calculate required VRAM. Inputs stay in this browser.
Used-card condition, power supply and case clearance need separate checks.
Check Blackwell kernels for the exact quantization; more VRAM does not certify faster decode.
Verify OS, driver and ROCm/Vulkan support. Experimental forks are not stock-runtime support.
BIOS GPU allocation and system use reduce available memory; 128 GB RAM is not 128 GB dedicated VRAM.
m5 ultra 96g and m5 ultra 96gb refer to this capacity. Memory is shared with macOS; verify shipping date and MLX conversion.
ARM64 packages and GB10 kernels must support the chosen model. Capacity is not a latency result.
| Issuer / checkpoint | Values | Source / verified |
|---|---|---|
| Nex AGI / nex-agi/Nex-N2.5-mini | 35B total; Check selected runtime and checkpoint; Apache-2.0 | Publisher model card / 2026-09-16 |
Inference experiments only. The Q4 memory budget is estimated, not a tested maximum context. Unsupported kernels, missing vision projectors, incorrect tool parsers and longer-context KV allocation can fail even when weights fit. No successful deployment, quality or speed is guaranteed.
Compare hardware purchase paths