用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/HoangNguyen0403/agent-skills-standard --skill evals-run命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Enforce SOLID principles, guard-clause style, function size limits, and intention-revealing naming across all languages. Use when refactoring for readability, applying clean-code patterns, reviewing naming conventions, or reducing function complexity.
Conduct high-quality, persona-driven code reviews. Use when reviewing PRs, critiquing code quality, or analyzing changes for team feedback.
Maximize context window efficiency, reduce latency, and prevent lost-in-middle issues through strategic masking and compaction. Use when token budgets are tight, tool outputs overflow the context, conversations drift from intent, or latency spikes from cache misses.
正在显示 SKILL.md
基于 SOC 职业分类
| name | evals-run |
| description | Workflow skill for evals run. |
| metadata | {"triggers":{"keywords":["evals run","workflow"]}} |
[!IMPORTANT] Workflow skill for evals run.
Optional args: slug=, ticket=<id/url>, mode=interactive|autonomous|channel, channel=, auto_continue=true|false, profile=business|hybrid|technical.
When the user asks to perform this workflow, execute the following steps:
description: Run blinded live skill evals and publish reproducible v2 results.
Measure whether a skill changes agent behavior with isolated, immutable, outcome-based eval evidence.
For ordinary maintenance after a complete catalog baseline exists, run pnpm evals:baseline first. It creates or resumes a selective manifest, reuses only compatible evidence, and prints the model, reasoning level, concurrency, and fresh-answer count without starting workers.
Review that plan before spending quota. Start workers only with pnpm evals:baseline -- --execute; the default is gpt-5.6-luna with high reasoning and one worker. Override intentionally with EVALS_MODEL, EVALS_REASONING_EFFORT, or EVALS_CONCURRENCY (maximum four workers).
If usage is exhausted, keep the run directory and rerun the identical --execute command after access resumes; completed answers are reused automatically.
Use pnpm evals:manifest -- --category <category> for one category or pnpm evals:manifest -- --all for the complete catalog.
Use pnpm evals:manifest -- --resume <runId> only when deliberately continuing an existing run; a new invocation always creates a collision-safe run ID.
Record the printed run ID. The manifest records source hashes, the v2 schema, and the generation protocol.
SKILL.md.all runs, write answers under answers/<category>/<skill>/<case>; category runs use answers/<skill>/<case>.metadata.agent, metadata.model, and metadata.completedAt after every required answer exists.pnpm evals:score -- --run <runId>.results.json while any arm is pending, verifies source hashes, and writes one immutable inputs.json snapshot before publishing v2 results.pnpm evals:report to project aggregate runs into the newest complete category partitions and update physical history/archive records.pnpm evals:verify -- --run <runId> and, before handoff, pnpm evals:verify -- --all.n/a for compromised arms.results.json, transcripts, history, or archives. Fix inputs or eval definitions and regenerate.feature_status: implemented | partially_implemented | blocked requirement_trace: manifest -> inputs -> results -> report -> verification completed_evidence: [] missing_evidence: [] decision_needed: [] recommended_next_workflow: verify-work