Skip to main content
Ejecuta cualquier Skill en Manus
con un clic

longmemeval-run

Estrellas56
Forks8
Actualizado3 de junio de 2026 a las 15:31

One-shot LongMemEval benchmark on a fresh CogniFold clone. Verifies env + API endpoints (~30 s, ~$0.01) and then runs the full N=500 benchmark on the recommended gpt-5 stack (~2-4 h wall-clock, ~$80-150 cost). Single command, no parameters to think about. Use after `git clone`, on a new machine, when the recommended stack changes, or when a previous run failed with API or environment errors. SKIP for other benchmarks (LoCoMo, MuSiQue, CogEval-Bench) — those have their own runners.

Instalación

Instalar con Codex o Claude Copia este prompt, pégalo en Codex, Claude u otro asistente, y deja que revise la página de la skill y la instale por ti.

Explorador de archivos
2 archivos
SKILL.md
readonly