-
-
Notifications
You must be signed in to change notification settings - Fork 227
2.3.83 Satellite Unsloth Studio
Handle:
unsloth-studio
URL: http://localhost:34851

Unsloth Studio is a no-code web UI for fine-tuning LLMs with the Unsloth library. It provides 2x faster training with 70% less memory compared to standard fine-tuning, and is powered by llama.cpp and Hugging Face under the hood. Studio runs from the unified unsloth/unsloth image but exposes only the Studio surface (port 8000) — the existing unsloth Harbor service continues to provide Jupyter Lab and SSH access on the same image when you need those.
Note: Studio is in beta upstream. The two services share the
unsloth/unslothimage and a single--gpus allreservation, so on a single-GPU host you should run only one ofunslothorunsloth-studioat a time.
-
Pre-pull the image (one-time). The unified
unsloth/unsloth:latestimage is large — ~13.6 GB compressed download, ~44.5 GB unpacked on disk. Pre-pulling once means your firstharbor upwon't block on the download:docker pull unsloth/unsloth:latest # or: harbor pull unsloth-studio -
Start Studio:
harbor up unsloth-studio --open
The first launch runs a
unsloth-studio-bootstrapsidecar that creates the admin account and mints an API key — typically ~10–60 s after the Studio container is healthy. -
Sign in at http://localhost:34851:
-
Username:
unsloth -
Password: run
harbor config get unsloth-studio.passwordand paste the output
-
Username:
-
Use Studio: pick a base model and a dataset, configure training (LoRA rank, learning rate, epochs, etc.), run it. Outputs land in the workspace directory and can be exported to GGUF, Ollama, vLLM, or Hugging Face formats.
Set your own password instead. Run
harbor config set unsloth-studio.password "..."before step 2. After the first run, change the password from inside the Studio UI (Settings) — the bootstrap sidecar only sets it during initial setup.
Default services start too.
harbor up unsloth-studioalso starts Harbor's default services (typicallywebuiandllamacpp). Runharbor defaultsto see them, orharbor defaults rm <service>to trim.
Cross-integrations are single-shot.
harbor up unsloth-studio <integration>(e.g.webui,boost,aider) waits on the bootstrap sidecar and reads the freshly-minted API key at runtime — first launch works end-to-end with no second pass.
First-run security. Upstream's first-run flow ships a temporary admin password (
window.__UNSLOTH_BOOTSTRAP__) inline in the Studio HTML until the "Setup your account" form is completed. The bootstrap sidecar consumes those credentials and completes setup automatically on first launch, but the temp creds are still served by Studio's HTML for a few seconds during boot. Keep port 34851 onlocalhostuntil the bootstrap sidecar reportsExited (0)(docker ps -a --filter name=harbor.unsloth-studio-bootstrap).
Following options can be set via harbor config:
# Studio Configuration
HARBOR_UNSLOTH_STUDIO_HOST_PORT=34851 # Studio web UI port
HARBOR_UNSLOTH_STUDIO_WORKSPACE="./services/unsloth-studio/workspace" # Local workspace dir
HARBOR_UNSLOTH_STUDIO_OPEN_URL="http://localhost:34851" # URL opened by `harbor open`
HARBOR_UNSLOTH_STUDIO_API_KEY="sk-unsloth-studio" # Bearer token; auto-bootstrapped on first run
HARBOR_UNSLOTH_STUDIO_PASSWORD="" # UI login password (auto-generated on first run if empty)
HARBOR_UNSLOTH_STUDIO_DEFAULT_MODEL="" # Optional HF model id to auto-load on launch
# Docker Image
HARBOR_UNSLOTH_STUDIO_IMAGE="unsloth/unsloth" # Docker image
HARBOR_UNSLOTH_STUDIO_VERSION="latest" # Image tagThe service mounts:
-
HARBOR_UNSLOTH_STUDIO_WORKSPACE→/workspace/work— your local working directory for projects, datasets, and exports. -
HARBOR_HF_CACHE→/workspace/.cache/huggingface— shared Hugging Face model cache, mounted at the path Studio actually reads from (HF_HOMEis baked into the upstream image). The same host cache is used by theunsloth,vllm, and other model-serving services, so models pulled once are available to all of them. The init sidecarchowns the bind-mount root to uid 1001 (Studio's in-container user) with mode0775; existing cache contents keep their original ownership so the host user can stillrmfiles they downloaded outside Studio withoutsudo. -
services/unsloth-studio/.studio-state/→/home/unsloth/.unsloth/studio— Studio's full state directory:studio.db(training-run metadata, settings),exports/(exported model artifacts),outputs/(training outputs),runs/(training-run logs),assets/datasets/(imported datasets), and the.venv_t5_*Python venvs Studio builds for tokenizer ops (~250 MB combined; persisting them skips the slow rebuild on everyharbor up). Without this mount every recreate loses all fine-tune work. The deeper.studio-authmount below wins on path specificity, so auth lives in its own dotfile dir on the host. -
services/unsloth-studio/.studio-auth/→/home/unsloth/.unsloth/studio/auth— Studio'sauth.dbplusapi_key.txt(written by the bootstrap sidecar; read by cross-integrations at runtime). Kept as a separate child mount on top of.studio-state/so password resets / re-bootstraps don't have to wipeexports/oroutputs/. Wipe this directory to force a fresh first-run flow.
Unsloth Studio requires NVIDIA GPU passthrough via the NVIDIA Container Toolkit. Harbor wires this in automatically through compose.x.unsloth-studio.nvidia.yml.
To download gated models or push fine-tuned models to the Hub, set your token:
harbor config set hf.token "hf_your_token_here"The token is forwarded into the container as HF_TOKEN.
Studio's backend is a FastAPI app that exposes an OpenAI-compatible inference API on the same port as the web UI. CPU inference works (GGUF models load via llama.cpp), so this is useful even on a no-GPU host.
- Base URL (host):
http://localhost:34851/v1 - Base URL (intra-Harbor, for other services in the same compose network):
http://unsloth-studio:8000/v1 - Auth: bearer token. Harbor mints one for you on first launch — see Zero-click API key bootstrap below. You can also create additional long-lived keys via the Studio UI /
POST /api/auth/api-keys. - Swagger UI:
http://localhost:34851/docs
harbor up unsloth-studio runs a one-shot unsloth-studio-bootstrap sidecar after the main container is healthy. It:
- scrapes the upstream
window.__UNSLOTH_BOOTSTRAP__first-run credentials from Studio's HTML, - logs in, completes the mandatory password change, mints an API key via
POST /api/auth/api-keys, - writes the new key into
.envasHARBOR_UNSLOTH_STUDIO_API_KEYso cross-integrations (Open WebUI, Boost, Aider) pick it up automatically.
The sidecar is idempotent: on subsequent harbor up it tries the stored key first and exits fast when it still works. Studio's auth state is bind-mounted at services/unsloth-studio/.studio-auth/, so the minted key survives harbor down/harbor up cycles. Wipe that directory (or run harbor config set unsloth-studio.api.key sk-unsloth-studio) to force a re-bootstrap.
Read your current key with:
harbor config get unsloth-studio.api.keyManual override. If you set HARBOR_UNSLOTH_STUDIO_API_KEY to anything that doesn't match the auto-bootstrap shape (sk-unsloth- followed by 32 hex chars), the sidecar leaves it alone — useful if you want to use a key generated through the Studio UI, an externally issued JWT, or a key you rotate yourself.
Recovery if the bootstrapped key is revoked. Because Studio's auth DB is bind-mounted, deleting the key inside the Studio UI (or via DELETE /api/auth/api-keys/<id>) leaves Studio with a fully-set-up admin account but no usable credentials in .env. The sidecar can't recover automatically in this state — Studio only serves first-run bootstrap creds in HTML before setup completes, and that ship has sailed. The sidecar will print a FATAL block on next harbor up listing the two manual fixes:
-
Reset the auth DB (loses any manually-created users / passwords / named keys, then bootstrap re-runs cleanly):
harbor down unsloth-studio rm ./services/unsloth-studio/.studio-auth/auth.db harbor config set unsloth-studio.api.key "" harbor up unsloth-studio
-
Mint a replacement key in the Studio UI (Settings → API keys → New) and persist it:
harbor config set unsloth-studio.api.key sk-unsloth-...
The OpenAI-compatible endpoints serve only the currently loaded model. Pick or download a model in the Studio UI first (sidebar → model picker → search Hugging Face or pick from the recommended list), or call POST /v1/load:
KEY=$(harbor config get unsloth-studio.api.key)
curl -sS -X POST http://localhost:34851/v1/load \
-H "Authorization: Bearer $KEY" \
-H 'Content-Type: application/json' \
-d '{"model_path": "unsloth/gemma-4-E2B-it-GGUF"}'The first call for a given model_path downloads the model into HARBOR_HF_CACHE (mins to tens of mins depending on size and bandwidth); subsequent loads of the same model are near-instant. After loading, GET /v1/models returns the loaded model id.
The model field in /v1/chat/completions and /v1/messages requests is not enforced — Studio always routes to the currently-loaded model regardless of the value sent. A request with "model": "does-not-exist" returns a normal 200 with the loaded model's id in the response. With no model loaded, chat completions return 400 {"detail":"No model loaded. Call POST /inference/load first."}.
The bootstrap sidecar can pre-load a model so cross-integrations (Open WebUI, Boost, Aider) work against Studio out of the box. Set:
harbor config set unsloth-studio.default.model "unsloth/gemma-4-E2B-it-GGUF"On the next harbor up unsloth-studio the sidecar calls POST /v1/load with this value after the API key is in place. It's idempotent — GET /v1/models is checked first and the load is skipped when the requested model is already resident. Container restart drops the loaded model (see Restart drops the loaded model), so the sidecar also re-loads it on every subsequent boot. Failures are logged but not fatal — Studio stays up even if the load fails.
Empty ("") is the default — no auto-load, current behaviour preserved. When set, the first run downloads the model into HARBOR_HF_CACHE (slow on a cold cache, near-instant once cached).
This pairs especially well with Aider, which requires a model: setting up front. With auto-load configured, harbor up unsloth-studio aider becomes a single command that gives you a working coding agent — see Aider model alignment below.
This example assumes a model is already loaded — see Loading a model above. With nothing loaded, the call returns 400 {"detail":"No model loaded. ..."}.
# 1) Use the bootstrapped key (or your manual override).
KEY=$(harbor config get unsloth-studio.api.key)
# 2) Call /v1/chat/completions with the loaded model id.
curl -sS http://localhost:34851/v1/chat/completions \
-H "Authorization: Bearer $KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "unsloth/gemma-4-E2B-it-GGUF",
"messages": [{"role": "user", "content": "Say hi briefly."}],
"max_tokens": 40
}'The response is a standard OpenAI chat.completion object (id, choices[].message.content, etc.).
/v1/chat/completions with "stream": true returns standard OpenAI server-sent-event chunks:
curl -sN http://localhost:34851/v1/chat/completions \
-H "Authorization: Bearer $(harbor config get unsloth-studio.api.key)" \
-H 'Content-Type: application/json' \
-d '{"model":"unsloth/gemma-4-E2B-it-GGUF","stream":true,
"messages":[{"role":"user","content":"hi"}],"max_tokens":32}'Each data: {...} chunk contains a choices[0].delta.content token; the stream terminates with a finish_reason:"stop" chunk, a final usage/timings chunk, and a data: [DONE] sentinel. Reasoning models emit reasoning text inside <think>... tokens in the same content stream rather than in a separate reasoning_content field. Aborting the HTTP request mid-stream closes the connection cleanly on Studio's side, but the underlying llama-server continues generating until it hits max_tokens or EOS — there's no observable cancellation propagation in /var/log/studio/access.log and the server keeps the slot busy. Plan timeouts and max_tokens accordingly.
POST /v1/messages accepts Anthropic's request shape and returns an Anthropic-style message response — useful for clients written against the Anthropic SDK.
Studio exposes POST /v1/embeddings in its OpenAPI spec, but the underlying llama-server is launched without --embeddings, so the endpoint always returns HTTP 501 "This server does not support embeddings". This is true even for models Studio identifies as embedding-only (e.g. nomic-ai/nomic-embed-text-v1.5-GGUF, where GET /api/models/check-embedding/... returns is_embedding: true). Until upstream wires the --embeddings flag into the llama-server invocation, Studio cannot serve as a backend for embedding-consuming services such as Cognee — use Ollama (nomic-embed-text, mxbai-embed-large) or llama.cpp instead.
Each /v1/load call spawns a fresh llama-server subprocess but does not always reap the previous one. Firing two /v1/load calls concurrently (same or different model) typically leaves an orphan llama-server running with the model still resident in RAM — easily 5 GB+ per orphan with a 4-bit Gemma. POST /v1/unload only stops the active server. The orphans persist until the container restarts.
Recommended client behaviour: serialise /v1/load calls and treat the endpoint as exclusive. If you've accidentally leaked processes (docker exec harbor.unsloth-studio sh -c 'ps aux | grep llama-server | grep -v grep | wc -l' > 1), the cheapest fix is docker restart harbor.unsloth-studio — the bind-mounted auth DB means your API key still works after restart, but the loaded model is gone (re-issue POST /v1/load after the container is healthy again).
Six cross-integrations register Studio as a backend with other Harbor services:
-
compose.x.webui.unsloth-studio.yml— adds Studio as an OpenAI endpoint in Open WebUI. -
compose.x.boost.unsloth-studio.yml— registers Studio as a named backend in Harbor Boost (HARBOR_BOOST_OPENAI_URL_UNSLOTH_STUDIO). -
compose.x.aider.unsloth-studio.yml— wires up an Aider config pointing at Studio. -
compose.x.opencode.unsloth-studio.yml— exposes Studio's loaded model in OpenCode via the auto-discovery flow (the discovery script reads the bootstrapped key file at runtime). -
compose.x.hermes.unsloth-studio.yml— points Hermes at Studio (OPENAI_BASE_URL+OPENAI_API_KEY). -
compose.x.openclaw.unsloth-studio.yml— wires Studio into OpenClaw as the configured backend (HARBOR_BACKEND_NAME/HARBOR_BACKEND_URLplus a runtime-read API key file).
All six substitute HARBOR_UNSLOTH_STUDIO_API_KEY as the bearer token, which the unsloth-studio-bootstrap sidecar populates on first launch — no manual key registration needed.
Aider model alignment. Studio routes every request to the currently-loaded model regardless of the
model:field (see Loading a model), so Aider'sHARBOR_AIDER_MODELvalue is largely cosmetic against Studio — what actually matters is that some model is loaded. The simplest single-command path is to setHARBOR_UNSLOTH_STUDIO_DEFAULT_MODELto your preferred model and (optionally)HARBOR_AIDER_MODELto match. Worked example:harbor config set unsloth-studio.default.model "unsloth/gemma-4-E2B-it-GGUF" harbor config set aider.model "unsloth/gemma-4-E2B-it-GGUF" harbor up unsloth-studio aiderIf you skip the default-model step, the first prompt will fail with
400 No model loadeduntil you load one via the Studio UI orPOST /v1/load.
How first-run key delivery works (skip unless debugging). Each cross-integration declares
depends_on: unsloth-studio-bootstrap(service_completed_successfully), so the integration container only starts after the bootstrap sidecar has minted the API key. Compose's create-time env substitution still bakes the placeholder key into the container'sConfig.Env(Compose substitutes before bootstrap runs), so the integration's start script reads the real key fromservices/unsloth-studio/.studio-auth/api_key.txt(mounted read-only at/run/unsloth-studio-auth/api_key.txt) and re-exportsHARBOR_UNSLOTH_STUDIO_API_KEYbefore rendering its config. End result: a singleharbor up unsloth-studio <integration>works on first launch with no--force-recreate.
Open WebUI picks up Studio as a model immediately after harbor up unsloth-studio webui:

# Tail logs
docker logs -f $(harbor ps unsloth-studio --quiet)-
Image pull is slow. Pre-pull with
docker pull unsloth/unsloth:latest— see Getting started step 1. -
No GPUs visible inside the container. Check the NVIDIA Container Toolkit is installed and the Harbor
nvidiacapability is detected:harbor config get capabilities.default(set withharbor config set capabilities.default 'nvidia'). -
GPU is busy / device already in use. The
unslothservice is probably running too — they share--gpus all. Stop one before starting the other. -
Need Jupyter Lab or SSH instead? Use the existing
unslothservice — same image, different ports. -
Restart drops the loaded model. Container restart reaps the
llama-serversubprocess; Studio does not auto-reload. Re-issuePOST /v1/loadonce the container is healthy again. (Auth DB survives restart, so your key still works.) -
Aider / WebUI / Boost gets
400 No model loaded. SetHARBOR_UNSLOTH_STUDIO_DEFAULT_MODELso the bootstrap sidecar pre-loads on everyharbor up— see Auto-loading a default model. -
Cross-integration auth errors right after first
harbor up unsloth-studio <integration>. Rare on a clean install. Confirmcat ./services/unsloth-studio/.studio-auth/api_key.txtshows ask-unsloth-...value; if the file is missing, the bootstrap sidecar didn't run cleanly — checkdocker logs harbor.unsloth-studio-bootstrap. - Cross-integration auth errors after revoking the bootstrap key. See Recovery if the bootstrapped key is revoked.
-
Several GB of RAM unaccounted for after model swaps. Orphan
llama-serverprocesses — see Concurrent /v1/load requests leak processes.