Retour au guide
RUNTIME / ADVANCED

oMLX

Apple-silicon serving with continuous batching, memory guards, and supported SSD-backed paths

Source primaire
DifficultyAdvancedSetup profile
Platforms1apple
APIOpenAI-compatible + Anthropic-compatible:8000
01 / INSTALL & RUN
  1. 1
    Installbrew install jundot/omlx/omlx
  2. 2
    Start a modelomlx serve --model-dir ~/models --memory-guard-gb 48
  3. 3
    Connect your app

    OpenAI-compatible + Anthropic-compatible · :8000. Keep the service bound to localhost unless you add authentication and network controls.

Attention

Apple silicon only. SSD-backed behavior is model-specific: the Qwen3.8 REAP build streams its sparse PLE table, not arbitrary model layers.

02 / COMPATIBLE MODELS

Modèles open-weight représentatifs

Community REAP build

Qwen3.8 Flash Next REAP-288

176BREAP-288 · MLX 4-bit · PLE on NVMe
Mémoire estimée
~40.57 GB
Contexte
7.6K recommended on measured 48 GB Mac