ガイドへ戻る
RUNTIME / EASY

Unsloth

Running GGUF or MLX models from a desktop app that also trains and exports them

一次情報
DifficultyEasySetup profile
Platforms5cpu · apple · nvidia · amd · intel
APIOpenAI + Anthropic-compatible:8888 (printed at launch)
01 / INSTALL & RUN

Installation depends on your OS. These commands are examples or templates: replace placeholders and verify quantization, file format, drivers and runtime support. Memory fit does not guarantee deployment.

  1. 1
    Installcurl -fsSL https://unsloth.ai/install.sh | sh # or download the Unsloth Desktop app
  2. 2
    Start a modelunsloth run --model <verified-GGUF-or-MLX-repo-tag>
  3. 3
    Connect your app

    OpenAI + Anthropic-compatible · :8888 (printed at launch). Keep the service bound to localhost unless you add authentication and network controls.

注意

Unsloth also fine-tunes, so its published VRAM savings are training numbers and not an inference fit claim. The local endpoint is an authenticated llama-server instance and the generated API key is shown once.

02 / COMPATIBLE MODELS

代表的なオープンウェイトモデル

Alibaba Qwen

Qwen3.6 27B

27BQ4 / NVFP4 estimate
推定メモリ
~18.5 GB
コンテキスト
262K native
Alibaba Qwen

Qwen3.8 27B

27BQ4 estimate
推定メモリ
~19.5 GB
コンテキスト
262K native
Alibaba Qwen

Qwen3.6 35B-A3B

35B3B active · Q4 / NVFP4 MoE
推定メモリ
~25 GB
コンテキスト
262K native
DeepSeek

DeepSeek V4 Flash

284B13B active · Official FP4 + FP8 mixed
推定メモリ
~176 GB
コンテキスト
1M native
Google

Gemma 3 1B

1BINT4 / GGUF Q4
推定メモリ
~1.4 GB
コンテキスト
32K
Meta

Llama 3.2 3B

3BGGUF Q4
推定メモリ
~2.8 GB
コンテキスト
128K