Skip to main content
open-thoughts
ملف منشئ GitHub

open-thoughts

عرض على مستوى المستودعات لـ ٥٠ skills مجمعة عبر ١ مستودعات GitHub.

skills مجمعة
٥٠
مستودعات
١
محدث
٣٠ أغسطس ٢٠٢٦
مستكشف المستودعات

المستودعات و skills الممثلة

rl-agentic-job-cleanup
غير مصنف

Preserve + publish a finished RL (SkyRL/GRPO) training checkpoint after the job terminates (completed at max_steps OR early-stopped/scancelled) on an HPC cluster (Jupiter/Leonardo/Perlmutter). Covers: cancel pending retries, pick the BEST checkpoint by…

٣٠ أغسطس ٢٠٢٦
datagen-job-cleanup
مطوّرو البرمجيات

Post-run cleanup for a datagen (trace-generation) job on Iris/CoreWeave or an HPC cluster (Jupiter/Leonardo/Perlmutter): get the generated traces onto HF (penfever org) and free temporary disk. There is NO model checkpoint — the artifact is the trace dataset.…

١ أغسطس ٢٠٢٦
monitor-job-tables
مطوّرو البرمجيات

Format HPC job-status reports as box-drawing tables, bucketed by job type (RL · SFT · Datagen · Eval · Catch-all), with the right metric columns, signal thresholds, and red-flags per bucket. Use whenever reporting active/recently-terminated job status —…

١ أغسطس ٢٠٢٦
datagen-launch-iris
مطوّرو البرمجيات

Launch, monitor, and manually clean up a trajectory-generation (datagen) job on Marin's Iris TPU cluster via the OpenThoughts-Agent entrypoint. Use when asked to start, watch, rescue, or kill a datagen/tracegen run on Iris.

٣١ يوليو ٢٠٢٦
analyze-datagen-campaign-summary
علماء البيانات

Build a clean per-dataset summary table/CSV for a datagen (trajectory-generation) campaign — one row per task source with Status (COMPLETED / FAILED / RUNNING / NOT STARTED), N Trials Completed, Mean Turns/Trace, Mean Tok/Trace, Mean Reward, and the HF…

٣٠ يوليو ٢٠٢٦
analyze-dataset-token-length
علماء البيانات

Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g. "task_complete AND < 32768 tokens"). Use when asked how long…

٣٠ يوليو ٢٠٢٦
analyze-id-eval-ranking
علماء البيانات

Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100, OT-TBLite=dev_set_v2, Terminal-Bench-2.0=tb2), HF links to each eval's trace…

٣٠ يوليو ٢٠٢٦
analyze-job-history-iris
مطوّرو البرمجيات

Run the Iris harbor job-history analyzer (scripts/iris/analyze_iris_harbor_job.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats. Use whenever a status check needs REAL metrics (gen tok/s,…

٣٠ يوليو ٢٠٢٦
عرض 8 من أصل ٥٠ skills مجمعة.
عرض ١ من أصل ١ مستودعات
تم تحميل كل المستودعات