Torna alla guida
HARDWARE / UNIFIED MEMORY

Framework Desktop · Ryzen AI Max+ 395 · 128 GB

A 128 GB Strix Halo desktop. AMD allows a large graphics-memory allocation, but usable capacity and speed depend on BIOS, driver, and runtime support.

Fonte primaria
Radeon 8060S compact desktop128 GB
01 / PRACTICAL CEILING

Qwen3.8 Flash Next REAP-288

The highest listed result may be a tight or offloaded fit. For daily use, prefer the first result marked Comfortable and keep context modest.

02 / DEFAULT RUNTIME

llama.cpp

GGUF models, CPU/GPU offload, embedded and unusual hardware. New architectures may require a recent build. Backend availability does not imply equal performance.

Apri guida
03 / MODEL ENVELOPE

Modelli open-weight rappresentativi

Community REAP build

Qwen3.8 Flash Next REAP-288

176BREAP-288 · MLX 4-bit · PLE on NVMe
Memoria stimata
~40.57 GB
Contesto
7.6K recommended on measured 48 GB Mac
OpenAI

gpt-oss-120b

117B5.1B active · MXFP4
Memoria stimata
~80 GB
Contesto
128K
Meta

Llama 4 Scout

109B17B active · INT4 MoE
Memoria stimata
~78 GB
Contesto
10M advertised
Meta

Llama 3.1 70B

70BGGUF Q4
Memoria stimata
~46 GB
Contesto
128K
Alibaba Qwen

Qwen3.6 35B-A3B

35B3B active · Q4 / NVFP4 MoE
Memoria stimata
~25 GB
Contesto
262K native
Alibaba Qwen

Qwen3 32B

32BGGUF Q4
Memoria stimata
~24.5 GB
Contesto
32K+
Attenzione

Fit is based on estimated total model memory. Driver support, KV cache, multimodal projectors, concurrency, and desktop applications can all reduce available headroom. Current fit: Qwen3.8 Flash Next REAP-288 (Comfortable), gpt-oss-120b (Comfortable), Llama 4 Scout (Comfortable), Llama 3.1 70B (Comfortable), Qwen3.6 35B-A3B (Comfortable), Qwen3 32B (Comfortable).