简优对齐

原名:simpo

Reference-free preference alignment, simpler than DPO.

分类
开发提效
版本
v1.0.0
作者
弈韬(@ra1nzzz)
下载
1
收藏
0
发布
2026-08-25
更新
2026-09-15
TRACE 评分
3.8 / 5

内容概览

SimPO is a reference-free preference optimization method that outperforms DPO without needing a reference model. Installation : Training (Mistral 7B): Config (mistral-7b-base-simpo.yaml): Launch training : Config (llama3-8b-instruct-simpo.yaml): Launch : For math/code tasks : Use SimPO when : - Want simpler training than DPO (no reference model) - Have preference data (chosen/rejected pairs) - Need better performance than DPO - Limited compute resources - Single-node training sufficient Algorithm selection : - SimPO : Simplest, best performance, no reference model - DPO : Need reference model baseline, more conservative - PPO : Maximum control, need reward model, complex setup - GRPO : Memory-efficient RL, no critic Use alternatives instead : - OpenRLHF : Multi-node distributed training, PPO/GRPO - TRL : Need multiple methods in one framework - DPO : Established baseline comparison Issue…

查看 SKILL 详情