Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
직접 명령은 검토 Prompt를 거치지 않습니다. 실행하기 전에 소스를 확인하세요.
npx skills add https://github.com/HoangNguyen0403/agent-skills-standard --skill evals-run명령은 한 줄로 유지됩니다. 복사하기 전에 가로로 스크롤해 전체 내용을 확인하세요.
로컬 사본을 원하시나요? SkillsMP에서 현재 제공할 수 있는 파일을 다운로드하세요.
SKILL.md 표시 중
| name | evals-run |
| description | Workflow skill for evals run. |
| metadata | {"triggers":{"keywords":["evals run","workflow"]}} |
[!IMPORTANT] Workflow skill for evals run.
Optional args: slug=, ticket=<id/url>, mode=interactive|autonomous|channel, channel=, auto_continue=true|false, profile=business|hybrid|technical.
When the user asks to perform this workflow, execute the following steps:
description: Run blinded live skill evals and publish reproducible v2 results.
Measure whether a skill changes agent behavior with isolated, immutable, outcome-based eval evidence.
For ordinary maintenance after a complete catalog baseline exists, run pnpm evals:baseline first. It creates or resumes a selective manifest, reuses only compatible evidence, and prints the model, reasoning level, concurrency, and fresh-answer count without starting workers.
Review that plan before spending quota. Start workers only with pnpm evals:baseline -- --execute; the default is gpt-5.6-luna with high reasoning and one worker. Override intentionally with EVALS_MODEL, EVALS_REASONING_EFFORT, or EVALS_CONCURRENCY (maximum four workers).
If usage is exhausted, keep the run directory and rerun the identical --execute command after access resumes; completed answers are reused automatically.
Use pnpm evals:manifest -- --category <category> for one category or pnpm evals:manifest -- --all for the complete catalog.
Use pnpm evals:manifest -- --resume <runId> only when deliberately continuing an existing run; a new invocation always creates a collision-safe run ID.
Record the printed run ID. The manifest records source hashes, the v2 schema, and the generation protocol.
SKILL.md.all runs, write answers under answers/<category>/<skill>/<case>; category runs use answers/<skill>/<case>.metadata.agent, metadata.model, and metadata.completedAt after every required answer exists.pnpm evals:score -- --run <runId>.results.json while any arm is pending, verifies source hashes, and writes one immutable inputs.json snapshot before publishing v2 results.pnpm evals:report to project aggregate runs into the newest complete category partitions and update physical history/archive records.pnpm evals:verify -- --run <runId> and, before handoff, pnpm evals:verify -- --all.n/a for compromised arms.results.json, transcripts, history, or archives. Fix inputs or eval definitions and regenerate.feature_status: implemented | partially_implemented | blocked requirement_trace: manifest -> inputs -> results -> report -> verification completed_evidence: [] missing_evidence: [] decision_needed: [] recommended_next_workflow: verify-work
Enforce SOLID principles, guard-clause style, function size limits, and intention-revealing naming across all languages. Use when refactoring for readability, applying clean-code patterns, reviewing naming conventions, or reducing function complexity.
Conduct high-quality, persona-driven code reviews. Use when reviewing PRs, critiquing code quality, or analyzing changes for team feedback.
Maximize context window efficiency, reduce latency, and prevent lost-in-middle issues through strategic masking and compaction. Use when token budgets are tight, tool outputs overflow the context, conversations drift from intent, or latency spikes from cache misses.
SOC 직업 분류 기준