Skip to main content

trl-fine-tuning

Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.

Jump to install

Source facts

Repository
agentic-in/elephant-agent
Last source activity
May 15, 2026 at 07:09
Detected SKILL.md language
English
Stars
578
Forks
62

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.