Skip to main content

punitarani/workbench

SkillsMP ha recopilado 8 skills de punitarani/workbench. Abre una skill para revisar su origen y sus detalles.

Última actividad de origen registrada
Catálogo de SkillsMP actualizado
skills recopiladas
8
Estrellas en GitHub
0
Forks en GitHub
0

Skills en este repositorio

clasificación pendiente

Mostrando 8 de 8 skills recopiladas.

ocupación
sin clasificar
descripción

Use when reading eval rollouts, trial logs, or trajectories to work out why a model scored below ceiling - covers what a zero actually means, DNF handling, sample size, and certifying a miss as a model failure rather than a task defect. Load before recording…

Idioma del texto original: inglés

actualizado
ocupación
sin clasificar
descripción

Use when generating a simulated workplace, institution, or multi-agent history that tasks will be graded against - covers determinism, the offstage boundary, coherence gates, artifact realism, and fidelity measurement. Load before writing any world generator.

Idioma del texto original: inglés

actualizado
ocupación
sin clasificar
descripción

Use when adding tests or gates that protect an eval suite's correctness - oracle independence, reachability, coherence, degeneracy, rule-phrasing, grading guards. Also covers falsifying a gate and auditing your own measurement tooling. Load before trusting…

Idioma del texto original: inglés

actualizado
ocupación
sin clasificar
descripción

Use when an eval task scores at ceiling or out of its target band and you need to move it - covers which difficulty levers are measured to do nothing, the coverage-versus-rule distinction, and which levers are forbidden. Load before changing a task to change…

Idioma del texto original: inglés

actualizado
ocupación
sin clasificar
descripción

Use when building or fixing an RL environment, eval task, or agent benchmark - the entry point that routes to world-building, task-authoring, gating, rollout analysis, and difficulty iteration. Enforces the rule that only a model failure may ship.

Idioma del texto original: inglés

actualizado
ocupación
sin clasificar
descripción

Use when writing an eval task instruction, oracle, or grader over a simulated world - covers the brief, declaring the rule kind, structural floors, deliverable shape, and bounding the work. Load before writing instruction.md or a solver.

Idioma del texto original: inglés

actualizado
ocupación
sin clasificar
descripción

Use when you are about to write an eval task, oracle, grader, or register against a generated world - the measure-first protocol that checks whether the pattern a task depends on actually exists, is evenly spread over time, is reachable by the agent through a…

Idioma del texto original: inglés

actualizado
ocupación
sin clasificar
descripción

Use when running, supervising, resuming, or babysitting a long generative simulation or recording that takes hours to days - covers supervisor design, the resume-not-restart rule, what may and may not change while a run is live, and accepting on the artifact…

Idioma del texto original: inglés

actualizado
Mostrando 8 de 8 skills recopiladas.