Skip to main content

rl-standard-job-cleanup

Preserve + publish a finished STANDARD (non-agentic GRPO) SkyRL RL checkpoint — the Delphi/rlvr/dapo math-and-reasoning cells launched via rl-standard-launch-leonardo (raw sbatch of hpc/skyrl_standard/leonardo/*, logger=console, NO Harbor/Daytona/trace_jobs). Covers: cancel pending retries, pick the BEST checkpoint by the trailing-5 EMA of reward via `parse_skyrl_metrics.py --format standard` (chain-aware, capped at the latest saved step), flatten weights to repo root, secret-scan, `hf upload` (Leonardo sbatch-tunnel) to laion/<run>-<step>-<size>B with the size suffix DERIVED FROM THE EXPORTED WEIGHTS (never the base-model name), DB register (--training-type RL + cross-user FK pre-check) ONLY for DB-registerable cells, clean up, and fire the Delphi downstream eval suite on the post-RL ckpt (defers to eval-standard-launch §5b). The ONLY artifacts are the model + the metric CSVs/report + the tracker scores — there is NO trace dataset. Use when a standard/non-agentic GRPO RL run finishes and needs uploading + (m

インストールへ移動

ソース情報

リポジトリ
open-thoughts/OpenThoughts-Agent
ソースの最終更新活動
2026年7月30日 10:49
検出された SKILL.md の言語
英語
スター
286
フォーク
39

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。