Skip to main content

evolution-lab-campaign

Execute an explicit Expert Squad incumbent-challenger campaign with immutable package revisions, exact run evidence, frozen scorers, independent integrity review, and a non-executing recommendation.

Informações da origem

Repositório
yangheng95/opencorvus
Última atividade na origem
27 de setembro de 2026 às 02:59
Idioma detectado do SKILL.md
inglês
Estrelas
312
Forks
40

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Explorador de arquivos
3 arquivos

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
evolution-lab-campaign
description
Execute an explicit Expert Squad incumbent-challenger campaign with immutable package revisions, exact run evidence, frozen scorers, independent integrity review, and a non-executing recommendation.
# Evolution Lab campaign A campaign resumes; it does not restart. Every published `evolution-lab/...@1` Artifact is immutable evidence that survives its producing Task's terminal lifecycle, so a failed stage is continued either by reopening that exact Task through `resume_task` or by creating the next stage Task with `{authority:"terminal_lifecycle",source_task_id,locator}` for the exact prior evidence. Imported evidence from a failed source correlates through Host-owned `import_lineage` exactly as evidence from a completed source does. Preserve existing measurements and frozen inputs; do not re-run or republish a Trial slot to select a better outcome. An integrity judgment may be corrected by its Safety Auditor through a complete Review with explicit revision.supersedes and a reason, preserving the earlier records. Recompute the Comparison from the resulting review set; ordinary citations never declare replacement. The Experiment Planner owns exactly one Campaign, frozen once during candidate preparation. The evaluation stage consumes that same immutable Campaign through Mission import and never republishes it, so the Planner declares no node there; a second `campaign-spec@1` for one experiment would be a second semantic fact. Use exactly one stage workflow per Task: `evolution-opportunity-analysis`, `evolution-candidate-preparation`, or `evolution-campaign-evaluation`. Mission creates the dependent Evolution Lab and target-Squad Tasks, preserves each immutable `promptProfile`, and imports exact accepted terminal Artifacts between stages. Each declared node publishes its owned `evolution-lab/...@1` Artifact after completely reading and selecting exact predecessor evidence. Cross-Task payloads contain portable role/path/media/bytes/digest manifests, never source-Task snapshot locators; call `rehydrate-evolution-resources` to rebuild current-Task campaign, parent-package, or candidate-package resource sets from the imported Engine resources. Dispatch prose carries scope, never copied evidence bodies, content digests, private paths, secrets, arm labels, or scores. Name the exact Artifact revision or resource path and let the consumer read the Host record that owns its digest; a restated digest is a second source for one fact and a transcription failure waiting to be mistaken for a real mismatch. The campaign is target-agnostic. Each Campaign freezes exactly one Dataset partition: development, holdout, or certification. Candidate preparation accepts only development; after candidate digest fixation, holdout and certification use new Campaign revisions without a Candidate Author, so secret cases never enter the development closure. Target Tasks keep one immutable target package revision, are created with that revision as the explicit expected digest, and cannot read Evolution Lab lineage. Candidate authoring is limited to the validated mutable text closure plus the manifest's declared descriptive identity text (name, label, description, selector summary and guidance, and every agent, workflow, and node label and description); tools, libraries, scripts, assets, scorers, datasets, evaluator workspace, permissions, and every manifest capability grant, reference, base role, prompt path, agent or workflow key, and dependency edge are frozen. Evolution Lab is frozen unless it is itself the exact Campaign target; a Campaign that targets it evolves the evolution trust root, so its promotion confirmation explicitly says so and every other contract stays unchanged. The publish draft names exact resource-role paths and planner-owned experiment semantics; the trusted publisher derives target, baseline revision, resource/scorer digests, model/timeout, workspace digest, rubric identity and Trial execution from the selected opportunity and immutable resource set. Only a literal `target.scope === "built_in"` candidate yields `product_release_required`. A project-scoped package remains executable in project scope even when `target.namespace === "builtin"`; namespace is provenance, not installation scope. Every scorer result is measured or typed unavailable and cites immutable evidence. The metric tool consumes one exact run-evidence Artifact and publishes an immutable receipt binding case, arm, repetition, Trial Task, target revision, scorer values, and evidence; evaluation publication must equal that receipt. Missing or corrupt evidence, inactivity, provider failure, parse failure, and invalid configuration remain unavailable. Required unavailable dimensions make aggregate score null. The Recommendation Owner may recommend promote, retain, or inconclusive only within the frozen experiment context. Installation, restoration, retry, and promotion require separate explicit operations and user authority. Freeze each scorer with exactly one `evaluator_kind` and the matching strict `evaluator_config_contracts` entry from `references/scorer-contract.json`. A judge must name its exact provider and model, real-output inactivity window, maximum selected-evidence bytes, criteria, and complete rubric before Campaign publication; it accepts only strict UTF-8 text or valid JSON evidence. Do not defer evaluator configuration to Trial execution. A shell scorer runs in an isolated copy of the measured Trial's terminal committed workspace, never in the Evaluator's or the Trial's live directory; its optional `cwd` is a relative path inside that copy. A judge reads the Trial's canonical collector bundle with the Message bodies selected at collection, rather than the Run Artifact wrapper. Preserve the original Messages; they can still reveal run identity, so this input change alone does not establish blindness. The canonical role-to-Artifact and frozen evaluation contracts are in `references/artifact-ownership.md` and `references/scorer-contract.json`. The monetary ceiling `max_cost` is a nonnegative number or explicit `null`. Use `null` when the user or Mission declared no cost ceiling; never invent a numeric cap or use zero to mean unlimited. Zero remains a literal zero ceiling. A suggested budget is not spending approval. Preserve declared numeric limits and record all actual usage and costs independently of this field.
Ver no GitHub