Skip to main content

burtenshaw/training-agents

SkillsMP は burtenshaw/training-agents から 6 件の skill を収集しています。skill を開くとソースと詳細を確認できます。

記録された最新のソース活動
SkillsMP カタログ更新
収集済み skills
6
GitHub スター
95
GitHub フォーク
17

このリポジトリの skills

1 件の職業カテゴリ · 100% 分類済み

収集済み skill 6 件中 6 件を表示しています。

職業分類
ソフトウェア開発者
説明

Use when designing or reviewing self-distillation workflows for agentic models, including trace collection, teacher or judge feedback, rejection sampling, critique, conversion to SFT or preference data, iterative TRL training loops, and safeguards against…

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Use when working with Hugging Face CLI or Hub workflows for TRL training, including auth, repositories, uploads, downloads, Jobs, buckets, model persistence, dataset checks, Space links, and remote artifact movement.

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Use when designing, reviewing, or implementing OpenEnv-style environment interfaces for agentic RL with TRL, including reset/step/state contracts, tasksets, Docker or HTTP/WebSocket serving, MCP compatibility, reward separation, and GRPO environment rollouts.

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Use when instrumenting or inspecting TRL training runs with Trackio, run names, metric schemas, dashboards, logs, grep or ripgrep, SFTP, Hugging Face Job logs, remote artifacts, or experiment result summaries.

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Use when building, reviewing, or editing TRL post-training workflows for agentic applications, including SFT, DPO, GRPO, RLOO, reward modeling, dataset formats, chat templates, assistant/completion-only losses, tool-calling data, reward functions, and…

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Use when designing, implementing, reviewing, or debugging supervised fine-tuning with TRL SFTTrainer or `trl sft`, especially for agentic models trained on chat messages, prompt/completion data, tool-calling examples, assistant-only loss, completion-only…

原文の言語: 英語

更新
収集済み skill 6 件中 6 件を表示しています。