Torna alla guida
RUNTIME / MEDIUM

MLX LM

Native Apple silicon inference, experimentation and fine-tuning

Fonte primaria
DifficultyMediumSetup profile
Platforms1apple
APIPython server and CLI
01 / INSTALL & RUN
  1. 1
    Installpip install mlx-lm
  2. 2
    Start a modelmlx_lm.chat --model google/gemma-3-1b-it
  3. 3
    Connect your app

    Python server and CLI. Keep the service bound to localhost unless you add authentication and network controls.

Attenzione

Only for Apple silicon. Use an MLX-converted model and tune KV cache size for long contexts.

02 / COMPATIBLE MODELS

Modelli open-weight rappresentativi

Google

Gemma 3 1B

1BINT4 / GGUF Q4
Memoria stimata
~1.4 GB
Contesto
32K
Meta

Llama 3.2 3B

3BGGUF Q4
Memoria stimata
~2.8 GB
Contesto
128K
Alibaba Qwen

Qwen3 4B

4BGGUF Q4
Memoria stimata
~3.6 GB
Contesto
32K+
Google

Gemma 3 4B

4BINT4 / GGUF Q4
Memoria stimata
~4.2 GB
Contesto
128K
Alibaba Qwen

Qwen3 8B

8BGGUF Q4
Memoria stimata
~6.8 GB
Contesto
32K+
Google

Gemma 3 12B

12BINT4
Memoria stimata
~9.4 GB
Contesto
128K