一键导入
local-context-budget
Use when local Copilot/Qwen/LLM Gateway hits prompt bloat, context limits, slow prefill, timeouts, or 24GB GPU token-budget issues.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Use when local Copilot/Qwen/LLM Gateway hits prompt bloat, context limits, slow prefill, timeouts, or 24GB GPU token-budget issues.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Use when local Qwen, llama-server, or ik_llama.cpp dies, stalls, times out, OOMs, cancels, or needs log-based failure diagnosis.
Use when a local-agent coding task needs minimal repo context, targeted file discovery, compact search, or a small implementation brief.
Use when a local Copilot/Qwen chat is too long and needs compaction, checkpointing, restart, or a handoff summary.
Use when local Qwen/Copilot should delegate isolated research, review, or summarization to subagents without bloating parent context.
Use when local Copilot/Qwen tool calls fail, malformed tool JSON appears, tools are skipped, or LLM Gateway agent mode is flaky.
| name | local-context-budget |
| description | Use when local Copilot/Qwen/LLM Gateway hits prompt bloat, context limits, slow prefill, timeouts, or 24GB GPU token-budget issues. |
| argument-hint | [symptom or target token budget] |
Keep local sessions inside the effective context budget of a single 24 GB GPU.
github.copilot.llm-gateway.defaultMaxTokens below llama-server --ctx-size so generation, tool continuation, and reserve space still fit.scripts\inspect-copilot-context.ps1 when available.defaultMaxOutputTokens around 4096 unless a task truly needs more.When reporting context-budget findings, return: