con un clic
CogniFold
CogniFold contiene 2 skills recopiladas de OpenNorve, con cobertura ocupacional por repositorio y páginas de detalle dentro del sitio.
Skills en este repositorio
One-shot LongMemEval benchmark on a fresh CogniFold clone. Verifies env + API endpoints (~30 s, ~$0.01) and then runs the full N=500 benchmark on the recommended gpt-5 stack (~2-4 h wall-clock, ~$80-150 cost). Single command, no parameters to think about. Use after `git clone`, on a new machine, when the recommended stack changes, or when a previous run failed with API or environment errors. SKIP for other benchmarks (LoCoMo, MuSiQue, CogEval-Bench) — those have their own runners.
Autonomous LongMemEval benchmark iteration toward ≥95% strict J-Score on full N=500. Use when the user asks to run, iterate, improve, or continue the LongMemEval campaign on branch `longmemeval-iter`. Loops baseline → cluster-analyze → propose fix → re-test → net-positive decision → commit → repeat. Terminates only when a clean-rerun confirmation also clears ≥475/500. SKIP for other benchmarks (LoCoMo, MuSiQue, NarrativeQA, etc.) — those have their own runners.