Skip to main content
Run any Skill in Manus
with one click

trl

Stars0
Forks0
UpdatedApril 27, 2026 at 08:14

TRL (Transformer Reinforcement Learning) for SFT, RLHF, DPO training. Use for supervised fine-tuning, reward modeling, RLHF (PPO), and Direct Preference Optimization (DPO). Best for training instruction-following LLMs. For model loading use transformers; for efficient fine-tuning use peft; for quantization use bitsandbytes.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly