一键导入
gpt-5-6-best-practice
Route GPT-5.6 tiers, reasoning effort, and subagents to minimize accepted-result cost subject to an explicit quality floor.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Route GPT-5.6 tiers, reasoning effort, and subagents to minimize accepted-result cost subject to an explicit quality floor.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Route Claude tiers, effort, and subagents for Fable-led work, and handle Fable-specific prompting, long-running execution, API behavior, refusals, and fallback.
This skill should be used when the user wants to install or configure Pi Agent (@earendil-works/pi-coding-agent) with DeepSeek (built-in), Ant-Ling Ring-2.6-1T (single-model custom provider), and ZenMux (multi-model OpenAI-compatible aggregator including Gemini 3.5 Flash, Claude, GPT), including auth, models.json, settings.json, a curated extension set, and known-pitfall fixes. Triggers on "配置 pi"、"setup pi agent"、"pi 装一下"、"配 ring/deepseek/gemini/zenmux 到 pi".
Review, improve, organize, deploy, and verify Grafana dashboards, provisioned alert rules, and the Telegraf→InfluxDB→Grafana monitoring stack. Use when working on Grafana dashboards, dashboard folders, tags, legends, panel readability, Flux/InfluxDB query correctness or performance, Prometheus/Telegraf dimensions, Telegraf JSON/HTTP scraping, anonymous access, dashboard/alert provisioning, Grafana Docker deployments, high CPU or memory caused by dashboard queries, provisioned alert rules and Flux alert conditions, contact points and Telegram/notification delivery, or requests like "review this dashboard", "optimize Grafana panels", "why is this dashboard slow / using CPU", "fix legends", "move dashboards into a folder", "add tags", "set up Grafana alerting", "alert won't fire / won't deliver", "fix Telegram alerts", "route alerts to a Telegram topic/thread (message_thread_id)", "bump/upgrade the Grafana version", "deploy and verify dashboards", or "Grafana best practices".
Audit and clean the persistent context that feeds a recurring autonomous agent loop — memory files and indexes, scheduled-task / automation prompts, and CLAUDE.md / AGENTS.md — so the loop stops degrading into a self-reinforcing echo chamber. Works for Claude Code /loop crons and Codex automations.
Download YouTube videos with yt-dlp and post-process them with ffmpeg — fetch best-quality video/audio, convert formats, extract audio, and burn in translated (e.g. Chinese) subtitles with a translator watermark. Covers cross-environment tool installation (macOS-focused) including the ffmpeg-full / libass gotcha. Triggers on "youtube 视频下载处理".
Safely audit and reduce CPU, load average, I/O wait, and Docker container resource pressure on macOS and Ubuntu/Linux hosts, especially remote production-like servers over SSH. Use when a host feels slow, load average is high, monitoring shows CPU spikes, Docker containers are suspected of consuming CPU, a service dashboard times out, or the user asks for a repeatable CPU optimization workflow.
| name | gpt-5-6-best-practice |
| description | Route GPT-5.6 tiers, reasoning effort, and subagents to minimize accepted-result cost subject to an explicit quality floor. |
Use this skill as a GPT-5.6 routing, prompting, and evaluation overlay for Codex. Safety, permissions, repository instructions, and explicit user constraints remain authoritative.
Default objective: satisfy the acceptance contract, then minimize accepted-result cost. Lower latency, higher quality beyond the contract, or mandatory topology is a different objective; pursue it at higher cost only when the user explicitly chooses that tradeoff.
As verified on 2026-07-10, Sol is the flagship tier, Terra the balanced tier, and Luna
the fastest, lowest-cost tier. The gpt-5.6 alias resolves to Sol. Light in the app and
IDE maps to Low in the CLI; Extra High maps to xhigh. Max deepens one agent, while
Ultra is multi-agent orchestration rather than an API reasoning.effort value. Codex
cloud currently does not expose model selection. Availability, UI labels, credits,
defaults, and Fast-mode behavior are volatile; read references/evidence-notes.md
before quoting them as current.
acceptance contract: observable acceptance criteria, required safety and
verification, and approval boundaries.tier: a vendor-defined capability class: Luna, Terra, or Sol.effort: reasoning and checking depth within one tier.lane: one tier + effort single-agent configuration.orchestration: the single-agent or lead/workers/verifier topology.route: lead lane, worker lanes, orchestration, and verification plan together.accepted-result cost: total spend required to pass the contract, measured in the
surface's primary accounting unit: API dollars or subscription credits, never both.non-inferior: acceptance performance within a predeclared tolerance on
representative work; it does not mean identical output on every run.surface: the active product/client, account entitlements, model controls, worker
overrides, permissions, and accounting system.Track tokens, latency, calls, and human correction as separate diagnostics or guardrails. Cached-input and reasoning tokens are breakdowns of input and output, not extra tokens to add again. If the surface exposes no comparable spend signal, describe relative resource intensity instead of claiming an exact cheapest route.
A request to “use subagents/workers with proper models to save cost” authorizes delegation and worker-lane selection within scope; it does not mandate fan-out, lower the acceptance contract, choose a Sol lead, or waive the premium gate. Treat the cost purpose as conditional: compare the three routes below and delegate only when the routed route is expected to be the cheapest accepted route.
If the user requires subagents as a topology independent of cost, follow that topology only when it preserves the acceptance contract. “Use subagents regardless of cost” is already an explicit higher-cost objective. If the user instead requires subagents “to save cost” but routing evidence predicts higher cost, pause and ask whether topology or cost governs. If mandatory topology cannot preserve the acceptance contract, pause and ask the user to relax either topology or the contract; never weaken or silently replace one requirement.
Pin and verify each worker's effective model and effort where supported. If overrides are unavailable, inherited, ignored, or unverified, do not dispatch workers for cost savings and do not claim worker-tier savings. Prefer the qualified single-agent lane or a handoff. A separately confirmed mandatory topology may still use such workers, but it is not a cost-saving route. Before dispatch, state the topology, worker lanes, cost rationale, and uncertainty compactly. After completion, report verification, observable effective lanes, retries/rescues, and the spend signal; without comparable spend data, report only relative resource intensity.
Follow this order. Later sections explain GPT-specific mappings but do not override the procedure.
Minimal read-only inspection needed to evaluate the premium gate is allowed before the pause; task execution is not. For a new-task handoff, include the outcome, constraints, acceptance criteria, and necessary paths or evidence so the user does not have to reconstruct the task.
| Tier | Good starting work | Escalation rule / next action | Avoid |
|---|---|---|---|
| Luna | Strict-schema extraction, classification, transformation, and mechanical work with objective checks | Move to Terra for sustained state, multi-step tools, ambiguity, or judgment | Raising effort to compensate for a tier mismatch |
| Terra | Repository exploration, bounded implementation, routine debugging, tests, and supporting workstreams | Move to Sol for difficult reasoning, semantic risk, architectural judgment, or expensive rework | Treating it as a universal replacement for difficult coding work |
| Sol | Ambiguous multi-file coding, architecture, difficult debugging, high-recall review, research synthesis, and high-value decisions | Repair context or raise effort only after a concrete contract failure | Using it for simple volume work |
Starting lanes:
xhigh: only after a measured gap shows that more checking or reasoning is
needed on the chosen tier.xhigh.For difficult work, compare a stronger tier at Low or Medium with a smaller tier at
higher effort. A dated third-party suite found that Sol Low/Medium sometimes achieved
similar or higher aggregate scores with fewer reported output tokens than smaller
tiers at very high effort; it did not establish complete token efficiency or repository
cost. Read references/evidence-notes.md before using that observation.
Local Codex clients support custom agents under .codex/agents/ or
~/.codex/agents/. Pin model, model_reasoning_effort, and intended sandbox in the
agent file when a stable route matters. Effective permissions can still be constrained
by the parent or runtime; inspect them before dispatch.
A worker-tier cost claim requires dispatching a named custom agent whose file pins
model and model_reasoning_effort, or a runtime surface with equivalent verified
overrides. The custom agent's name is the dispatch identity; a profile merely existing
on disk does not bind a generic spawn. Prompt steering alone is not verified pinning.
Cost-routing workers normally use Luna or Terra. An independent verifier is different: if the acceptance contract requires fresh-context review, its cost is part of every eligible route and its tier follows residual semantic risk, even when that means Sol.
Concurrent writers need disjoint ownership; read-only workers may share scope, and
sequential handoffs may touch the same files with one writer at a time. Keep
agents.max_depth = 1 unless recursive delegation is measured to help. Treat
agents.max_threads as a cap, not a target. Read references/worker-profiles.md for
copyable roles and references/prompt-patterns.md for worker and verifier packets.
State each instruction once. Provide outcome, relevant evidence, hard constraints,
approval boundaries, acceptance criteria, verification, budget, stop condition, and
output shape. Expose only relevant tools, trim stale context, and return distilled
evidence rather than raw logs. Read references/prompt-patterns.md for compact shapes.
When migrating older prompts, remove scaffolding that only compensated for weaker models, but keep safety, data, budget, scope, style, and business constraints. Re-map effort instead of carrying old defaults forward, then re-test tokens, latency, accepted-result cost, and timeouts.
GPT-5.6 uses real-time cyber and biology safeguards. Legitimate dual-use work may pause or be refused. Do not treat every pause as a stuck agent, retry adversarially, or hide the failure; report it and follow the current surface's allowed recovery path.
Fast mode trades more credits for lower latency; it is not a token-saving feature. Max spends more reasoning on one agent. Ultra adds orchestration. None is a default cost-saving switch.
references/prompt-patterns.md: lead, worker, and verifier packets.references/worker-profiles.md: copyable Luna, Terra, and Sol custom agents.references/routing-and-evaluation.md: cost accounting, one-off routing, and route
adoption procedure.references/evidence-notes.md: dated product, benchmark, and surface provenance.Read only the reference needed for the current task.