RUNTIME / MEDIUM
MLX LM
Native Apple silicon inference, experimentation and fine-tuning
Source primaireDifficultyMediumSetup profile
Platforms1apple
APIPython server and CLI
01 / INSTALL & RUN
- 1Install
pip install mlx-lm - 2Start a model
mlx_lm.chat --model google/gemma-3-1b-it - 3Connect your app
Python server and CLI. Keep the service bound to localhost unless you add authentication and network controls.
Attention
Only for Apple silicon. Use an MLX-converted model and tune KV cache size for long contexts.
02 / COMPATIBLE MODELS
Modèles open-weight représentatifs
Google
Gemma 3 1B
1BINT4 / GGUF Q4
- Mémoire estimée
- ~1.4 GB
- Contexte
- 32K
Meta
Llama 3.2 3B
3BGGUF Q4
- Mémoire estimée
- ~2.8 GB
- Contexte
- 128K
Alibaba Qwen
Qwen3 4B
4BGGUF Q4
- Mémoire estimée
- ~3.6 GB
- Contexte
- 32K+
Google
Gemma 3 4B
4BINT4 / GGUF Q4
- Mémoire estimée
- ~4.2 GB
- Contexte
- 128K
Alibaba Qwen
Qwen3 8B
8BGGUF Q4
- Mémoire estimée
- ~6.8 GB
- Contexte
- 32K+
Google
Gemma 3 12B
12BINT4
- Mémoire estimée
- ~9.4 GB
- Contexte
- 128K