Skip to main content

eval-standard-cleanup

Consolidate FINISHED standard / lm_eval (evalchemy) math-suite eval jobs — the Delphi #6279 MATH-500 / AIME24 / gsm8k grid launched via eval-standard-launch — into the SCORES.md tracker. Per job: confirm a NON-EMPTY seed42 result (a crash leaves it empty — don't record it as a score), rsync the results_*.json into the local per-model dir, write scalar partials, parse the scalars (MATH-500 acc×100 / AIME24 10-seed mean±sd / gsm8k strict+flex / Raw), and flip the SCORES.md row 🚀 eval submitted → ✅ done. The artifact is SCALAR SCORES IN A TRACKER — HF-upload-only, NEVER DB. Use when asked to consolidate / harvest finished standard math-eval jobs or fill the scaling-laws score grid. DISTINCT from eval-agentic-cleanup (the Harbor-trace + Supabase DB path).

跳到安装

来源信息

仓库
open-thoughts/OpenThoughts-Agent
最近来源活动
2026年7月30日 10:49
检测到的 SKILL.md 语言
英语
星标
286
分支
39

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。