Skip to main content

engine-lm-studio

Reuse what LM Studio already has on this computer through its `lms` CLI: adopt its running server into the person's fleet, serve its MLX models, or serve its GGUF files without it. Load before touching LM Studio, its models or its server.

معلومات المصدر

المستودع
autonomous-ai/openharness
آخر نشاط في المصدر
١ أكتوبر ٢٠٢٦ في ١٢:١٢
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
١٬٠٨٣
التفرعات
٩٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.

عرض SKILL.md

SKILL.md
تعليمات المصدر · معاينة للقراءة فقط
name
engine-lm-studio
description
Reuse what LM Studio already has on this computer through its `lms` CLI: adopt its running server into the person's fleet, serve its MLX models, or serve its GGUF files without it. Load before touching LM Studio, its models or its server.
# LM Studio (`lms` CLI, no clicks) Official docs, read 2026-09-29 (source files in github.com/lmstudio-ai/docs): [App, llmster and lms](https://lmstudio.ai/docs/app/basics/lmstudio-vs-llmster-vs-lms) · [Headless](https://lmstudio.ai/docs/developer/core/headless) · [lms server start](https://lmstudio.ai/docs/cli/serve/server-start) · [lms server stop](https://lmstudio.ai/docs/cli/serve/server-stop) · [lms load](https://lmstudio.ai/docs/cli/local-models/load) · [lms import](https://lmstudio.ai/docs/cli/local-models/import) · [lms daemon up/down](https://lmstudio.ai/docs/cli/daemon/daemon-up) · [Idle TTL](https://lmstudio.ai/docs/developer/core/ttl-and-auto-evict) · [OpenAI compatibility](https://lmstudio.ai/docs/developer/openai-compat) · [Tools](https://lmstudio.ai/docs/developer/openai-compat/tools) · [Structured output](https://lmstudio.ai/docs/developer/openai-compat/structured-output) · [Import models](https://lmstudio.ai/docs/app/advanced/import-model) · [Server settings](https://lmstudio.ai/docs/developer/core/server/settings) Tested: LM Studio 0.4.25 (`brew install --cask lm-studio`), MacBook Pro M1 Pro 32 GB, macOS 26.6, 2026-09-29. Tags: `[doc]` official page above, `[run]` seen on the tested machine, `[?]` unverified. ## When to use it - **A model in LM Studio's folder**, GGUF or MLX (`fleet models` → START WITH `lm-studio`): LM Studio runs it, because it downloaded it and loads it. - **Its server answering**: adopt it, but load the model yourself with a 64K+ context (below) — a model it loads on demand gets 8192 [run], too small for an agent. - **`(start it)`**: start its server yourself (below); an LM Studio that is off is the normal case, not a reason to switch engines. - **An MLX model elsewhere and mlx-lm not installed**: LM Studio runs MLX on Apple silicon (its structured output uses Outlines for MLX [doc]), so use it instead of installing anything. - Never install LM Studio just to serve a model: Grid's engine covers GGUF, mlx-lm covers MLX. ## Where its models live / how to list them - `~/.lmstudio/models/<publisher>/<model>/<file>` — it keeps Hugging Face's layout [doc]. The folder can be moved in the app's My Models tab [doc]; no documented setting file or env var [?]. Older versions used `~/.cache/lm-studio/models` [?]. `fleet models` scans both. - `lms ls --json` → `modelKey`, `path`, `format` (gguf/mlx), `sizeBytes` [run]. Keys prefix-match: a short key matched two downloaded models and loaded "the first one" [run] — always pass the exact key. - A fresh install holds only an embedding model, no chat model [run]. ## Installed? Running? - App: `/Applications/LM Studio.app`. `lms` lives in `~/.lmstudio/bin` and appears only after the app ran once [run]. Headless servers use **llmster** instead of the app: `curl -fsSL https://lmstudio.ai/install.sh | bash` [doc]. - `lms server status` → "The server is not running." or its port [run]. Launching the app does not start the server [run]. `fleet models` lists a running one as `kind: lm-studio` (its `owned_by` is `organization_owner` [run]). ## Start an already-downloaded model Port P from outside `machine.listeningPorts`. Always `--port`: without it the last used port is reused [doc]. export PATH="$HOME/.lmstudio/bin:$PATH" lms server start --port P --bind 127.0.0.1 # starts LM Studio in the background if needed [doc][run] lms load <exact key> --estimate-only --context-length 65536 --gpu max -y # fits? [doc][run] lms load <exact key> --context-length 65536 --gpu max --identifier <id> -y - With the app closed, `server start` printed "Waking up LM Studio service..." and was up in 4 s; it runs the app as `LM Studio --run-as-service`, with no window [run]. It returns at once (not a foreground process). - `--estimate-only` prints GPU and total memory for that context and whether guardrails allow it [doc][run]. - A model loaded with `lms load` has no idle timer and stays until unloaded [doc]; on-demand loads unload after 60 minutes idle [doc]. Nothing downloads on load [run]. ## Ready means Run `"$GRID_FLEET" verify --at http://127.0.0.1:P/v1 --model <id> --kind lm-studio` — it performs the request checks with deadlines and prints each one. Also: 1. `lms server status` shows port P [run]. 2. `GET http://127.0.0.1:P/v1/models` lists `<id>` [doc][run]. 3. LM Studio reports `<id>` loaded with ≥ 65536 (`/api/v0/models`, the CONTEXT `lms ps` shows) [run]. 4. One bounded `/v1/chat/completions` request with `model: <id>`, `max_tokens` 16 → non-empty `content` [run]. A 401 means the person turned on "Require Authentication" [doc]: do not change their setting; say so and serve the file with Grid's engine instead. ## Join Harness Compute "$GRID_FLEET" run -- join GRID --at http://127.0.0.1:P/v1 -m <id> --advertise-as ALIAS `/v1` is required: with the bare root, Grid's capability probe records structured output as unsupported [run]. Grid's own detector finds LM Studio only on 1234 [run]. ## Tool calls, JSON output, thinking - Tools: `tools` on `/v1/chat/completions` and `/v1/responses`, OpenAI format [doc]. Grid's probe: tools yes [run]. - Structured output: `response_format` with `json_schema` [doc]; GGUF uses llama.cpp grammars, MLX uses Outlines [doc]. Grid's probe: json_schema yes, json_object no [run]. - Thinking control through the OpenAI endpoint: not documented [?]. Check `content` in the ready request. ## Memory and speed knobs `--context-length` (≥ 65536 here), `--gpu max|off|0–1` (automatic when omitted) [doc], `--ttl <seconds>` [doc], `--parallel <n>` (default 4 on this build [run]), `--speculative-draft-mtp` for models with MTP layers [run: in help]. ## Stop - Your model: `lms unload <id>` [doc]. Your server: `lms server stop` [doc] — the background service keeps running [run]. - The service you woke: `lms daemon down` stops llmster only, not the app [doc]; the app service ends with `osascript -e 'quit app "LM Studio"'` [run]. - If LM Studio was running before you came, unload only your model and leave the rest as it was. ## Known failures → what to do | Sign | Do | |---|---| | "2 models match the provided model key" | use the exact `modelKey` from `lms ls --json` [run] | | CONTEXT 8192 in `lms ps` | the model was loaded on demand; unload it and `lms load` with `--context-length 65536` [run] | | `lms get … -y` shows a spinner forever | it downloads in the background [run]; never start a download without the go-ahead | | a foreign GGUF must appear in LM Studio | `lms import -l -y --user-repo <user>/<repo> FILE.gguf`: **without `-l` it moves the file** [doc]; the file name must end in `.gguf` and the link must point at the real file, not a temp copy [run] | | 401 from `/v1/models` | authentication is on in the person's settings [doc]; do not turn it off | | a model already on disk is not in `lms ls` | no download: an **MLX folder** is symlinked as `~/.lmstudio/models/<publisher>/<name>` → the folder, then `lms ls` lists it [run]; a **GGUF** goes in with `lms import -l` (row above). Neither copies | | answers come back as reasoning only; no request switch turns thinking off | some models keep thinking under LM Studio whatever the request says [run]. That is not a failure: `fleet verify` asks once more with room to think and passes, noting it. Never hand-test switches with `curl` (it cost minutes [run]); tell the person replies will be slower, and offer a model that can switch it off if they want that |
عرض على GitHub