泥塑精进
原名:slime
RL post-training for LLMs with Megatron and SGLang.
- 分类
- 开发提效
- 版本
- v1.0.0
- 作者
- 弈韬(@ra1nzzz)
- 下载
- 1
- 收藏
- 0
- 发布
- 2026-08-25
- 更新
- 2026-09-15
- TRACE 评分
- 3.2 / 5
内容概览
slime is an LLM post-training framework from Tsinghua's THUDM team, powering GLM-4.5, GLM-4.6, and GLM-4.7. It connects Megatron-LM for training with SGLang for high-throughput rollout generation. Choose slime when you need: - Megatron-LM native training with SGLang inference - Custom data generation workflows with flexible data buffers - Training GLM, Qwen3, DeepSeek V3, or Llama 3 models - Research-grade framework with production backing (Z.ai) Consider alternatives when: - You need enterprise-grade stability features → use miles - You want flexible backend swapping → use verl - You need PyTorch-native abstractions → use torchforge - Training : Megatron-LM with full parallelism support (TP, PP, DP, SP) - Rollout : SGLang-based high-throughput generation with router - Data Buffer : Flexible prompt management and sample storage - Models : GLM-4.x, Qwen3, DeepSeek V3/R1, Llama 3 --- Use t…