trl-fine-tuning

TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF.

分类
开发提效
版本
v1.0.1
作者
弈韬(@ra1nzzz)
下载
1
收藏
0
发布
2026-08-25
更新
2026-08-25
TRACE 评分
3.2 / 5

内容概览

TRL provides post-training methods for aligning language models with human preferences. Installation : Supervised Fine-Tuning (instruction tuning): DPO (align with preferences): Complete pipeline from base model to human-aligned model. Note (TRL 1.x): PPO has been removed from TRL — PPOTrainer, PPOConfig, and python -m trl.scripts.ppo no longer exist. Use an online-RL trainer TRL still ships: RLOO (RLOOTrainer / trl rloo) is the closest drop-in for a reward-model-driven RLHF pipeline, and GRPO (GRPOTrainer / trl grpo, see Workflow 3) is the memory-efficient alternative. The step below uses RLOO. Copy this checklist: Step 1: Supervised fine-tuning Train base model on instruction-following data: Step 2: Train reward model Train model to predict human preferences: Step 3: RLOO reinforcement learning Optimize policy using the reward model. PPO was removed in TRL 1.x; use the RLOO CLI (trl rl…

查看 SKILL 详情