简优对齐
原名:simpo
Reference-free preference alignment, simpler than DPO.
- 分类
- 开发提效
- 版本
- v1.0.0
- 作者
- 弈韬(@ra1nzzz)
- 下载
- 1
- 收藏
- 0
- 发布
- 2026-08-25
- 更新
- 2026-09-15
- TRACE 评分
- 3.8 / 5
内容概览
SimPO is a reference-free preference optimization method that outperforms DPO without needing a reference model. Installation : Training (Mistral 7B): Config (mistral-7b-base-simpo.yaml): Launch training : Config (llama3-8b-instruct-simpo.yaml): Launch : For math/code tasks : Use SimPO when : - Want simpler training than DPO (no reference model) - Have preference data (chosen/rejected pairs) - Need better performance than DPO - Limited compute resources - Single-node training sufficient Algorithm selection : - SimPO : Simplest, best performance, no reference model - DPO : Need reference model baseline, more conservative - PPO : Maximum control, need reward model, complex setup - GRPO : Memory-efficient RL, no critic Use alternatives instead : - OpenRLHF : Multi-node distributed training, PPO/GRPO - TRL : Need multiple methods in one framework - DPO : Established baseline comparison Issue…