occupation
Data Scientists
description
Post-train LLMs with TRL (Transformers Reinforcement Learning) — SFT, DPO, GRPO, KTO, and reward-model training. Use when writing or debugging training code with the TRL Python API or the trl CLI.
updated
Menu
SkillsMP has collected 1 skills from huggingface/trl. Open a skill to review its source and details.
Showing 1 of 1 collected skills.
Post-train LLMs with TRL (Transformers Reinforcement Learning) — SFT, DPO, GRPO, KTO, and reward-model training. Use when writing or debugging training code with the TRL Python API or the trl CLI.