Skip to main content
burtenshaw
GitHub 创作者资料

burtenshaw

按仓库查看 6 个 GitHub 仓库中的 16 个已收集 skills。

已收集 skills
16
仓库
6
更新
2026年6月15日
仓库浏览

仓库与代表性 skills

agentic-self-distillation
软件开发工程师

Use when designing or reviewing self-distillation workflows for agentic models, including trace collection, teacher or judge feedback, rejection sampling, critique, conversion to SFT or preference data, iterative TRL training loops, and safeguards against…

2026年6月15日
hugging-face-cli-workflows
软件开发工程师

Use when working with Hugging Face CLI or Hub workflows for TRL training, including auth, repositories, uploads, downloads, Jobs, buckets, model persistence, dataset checks, Space links, and remote artifact movement.

2026年6月15日
openenv-agentic-rl
软件开发工程师

Use when designing, reviewing, or implementing OpenEnv-style environment interfaces for agentic RL with TRL, including reset/step/state contracts, tasksets, Docker or HTTP/WebSocket serving, MCP compatibility, reward separation, and GRPO environment rollouts.

2026年6月15日
trackio-observability
软件开发工程师

Use when instrumenting or inspecting TRL training runs with Trackio, run names, metric schemas, dashboards, logs, grep or ripgrep, SFTP, Hugging Face Job logs, remote artifacts, or experiment result summaries.

2026年6月15日
trl-post-training
软件开发工程师

Use when building, reviewing, or editing TRL post-training workflows for agentic applications, including SFT, DPO, GRPO, RLOO, reward modeling, dataset formats, chat templates, assistant/completion-only losses, tool-calling data, reward functions, and…

2026年6月15日
trl-sft
软件开发工程师

Use when designing, implementing, reviewing, or debugging supervised fine-tuning with TRL SFTTrainer or `trl sft`, especially for agentic models trained on chat messages, prompt/completion data, tool-calling examples, assistant-only loss, completion-only…

2026年6月15日
已展示 6 / 6 个仓库
已展示全部仓库