一键导入
evaluate
Internal Harness instruction source for evaluate. Route through visible Harness aliases or hook contracts instead of invoking directly.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Internal Harness instruction source for evaluate. Route through visible Harness aliases or hook contracts instead of invoking directly.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Internal Harness instruction source for auto-iterate-goal. Route through visible Harness aliases or hook contracts instead of invoking directly.
Internal Harness instruction source for auto-paper. Route through visible Harness aliases or hook contracts instead of invoking directly.
Internal Harness instruction source for iterate. Route through visible Harness aliases or hook contracts instead of invoking directly.
Internal Harness instruction source for orchestrator. Route through visible Harness aliases or hook contracts instead of invoking directly.
Visible Harness write entry. Use for manuscript writing, citation-supported paper work, final documentation, README hardening, and GitHub Pages preparation.
Visible Harness write entry. Use for manuscript writing, citation-supported paper work, final documentation, README hardening, and GitHub Pages preparation.
| name | evaluate |
| description | Internal Harness instruction source for evaluate. Route through visible Harness aliases or hook contracts instead of invoking directly. |
Read these first:
../../../.agents/references/workflow-guide.md../../../.agents/references/context-layering-policy.md../../../.agents/references/commit-checkpoint-rule.md../../../.agents/references/lesson-quality-rule.md../../../.agents/references/language-policy.md../../../.agents/references/run-artifact-contract.md../../../.agents/references/documentation-evidence-rule.md../../../.agents/references/documentation-style.md../../../.agents/references/research-supervision-patterns.md../../../.agents/references/research-supervision/experiment-and-build-canvas.md../../../.agents/references/research-supervision/ai-assisted-research-workflow.md../../../.agents/references/reviewer-independence.md../../../.agents/references/review-tracing.md./references/stage-report.md../../../iteration_log.json../../../PROJECT_STATE.json../../../docs/context/contracts.md if it exists; legacy contract files are
fallback inputs before migration../../../docs/context/experiments.md if it exists../../../docs/context/memory.md if it existsUse this skill when the user wants training or evaluation results interpreted and turned into a decision.
pre_eval_commit for meaningful eval work, or record
pre_eval_commit_NOT_CHANGED when the committed training source already
covers eval code/configs. Do not use eval output as Conclusion Evidence
without a committed eval identity../references/stage-report.md.docs/context/experiments.md.docs/context/memory.md; write
root MEMORY.md only for accepted lessons when the project keeps that
optional bank.NOT_RUN when no experiments or memory context update is needed.claim_delta_evidence_NOT_CHANGED.MEMORY.md.NEXT_ROUND — ordinary improvement round, stay in WF10DEBUG — fixable technical issue, stay in WF10CONTINUE — handoff to orchestrator/WF11, not continue iteratingPIVOTABORT$iterate, do not take over stage-transition ownership.docs/30_evidence/Experiment_Evidence_Index.{json,md} with
tooling/evidence/build_experiment_evidence_index.py after completed run
evidence is written, or report NOT_RUN with the reason.experiment commit checkpoint for completed evaluation/discovery
slices before long follow-up runs or handoff.docs/context/experiments.md,
docs/context/memory.md, stage reports, optional legacy mirrors,
iteration_log.json, claim delta evidence, or the experiment evidence
index are written. If memory-quality or workflow-state checks are not run,
mark them NOT_RUN with the reason.After stable Markdown outputs for this skill are finalized, invoke $docs-site or report docs_site_boundary_report. Do not render after temporary draft edits; Markdown remains the source of truth.
$evaluate flow or the evaluation sub-step of $iterate..agents/state/current_iteration.json as the active iteration context path.../../../.agents/references/language-policy.md for reply language and for localizing natural-language report sections; keep protocol keys and decision tokens in English.Follow the local evaluation prompt, stage-report template, and language policy rather than collapsing this into a brief metrics summary.
After stable Markdown is finalized, invoke $docs-site or report
docs_site_boundary_report / docs_site_render_or_NOT_RUN. Do not render for
temporary drafts.