Skip to main content

burtenshaw/training-agents

SkillsMP는 burtenshaw/training-agents에서 6개의 skill을 수집했습니다. skill을 열어 소스와 세부 정보를 확인하세요.

최근 기록된 소스 활동
SkillsMP 카탈로그 업데이트
수집된 skills
6
GitHub 스타
95
GitHub 포크
17

이 저장소의 skills

직업 카테고리 1개 · 100% 분류됨

수집된 skill 6개 중 6개를 표시합니다.

직업 분류
소프트웨어 개발자
설명

Use when designing or reviewing self-distillation workflows for agentic models, including trace collection, teacher or judge feedback, rejection sampling, critique, conversion to SFT or preference data, iterative TRL training loops, and safeguards against…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Use when working with Hugging Face CLI or Hub workflows for TRL training, including auth, repositories, uploads, downloads, Jobs, buckets, model persistence, dataset checks, Space links, and remote artifact movement.

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Use when designing, reviewing, or implementing OpenEnv-style environment interfaces for agentic RL with TRL, including reset/step/state contracts, tasksets, Docker or HTTP/WebSocket serving, MCP compatibility, reward separation, and GRPO environment rollouts.

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Use when instrumenting or inspecting TRL training runs with Trackio, run names, metric schemas, dashboards, logs, grep or ripgrep, SFTP, Hugging Face Job logs, remote artifacts, or experiment result summaries.

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Use when building, reviewing, or editing TRL post-training workflows for agentic applications, including SFT, DPO, GRPO, RLOO, reward modeling, dataset formats, chat templates, assistant/completion-only losses, tool-calling data, reward functions, and…

원문 언어: 영어

업데이트
직업 분류
소프트웨어 개발자
설명

Use when designing, implementing, reviewing, or debugging supervised fine-tuning with TRL SFTTrainer or `trl sft`, especially for agentic models trained on chat messages, prompt/completion data, tool-calling examples, assistant-only loss, completion-only…

원문 언어: 영어

업데이트
수집된 skill 6개 중 6개를 표시합니다.