泥塑精进

原名:slime

RL post-training for LLMs with Megatron and SGLang.

分类
开发提效
版本
v1.0.0
作者
弈韬(@ra1nzzz)
下载
1
收藏
0
发布
2026-08-25
更新
2026-09-15
TRACE 评分
3.2 / 5

内容概览

slime is an LLM post-training framework from Tsinghua's THUDM team, powering GLM-4.5, GLM-4.6, and GLM-4.7. It connects Megatron-LM for training with SGLang for high-throughput rollout generation. Choose slime when you need: - Megatron-LM native training with SGLang inference - Custom data generation workflows with flexible data buffers - Training GLM, Qwen3, DeepSeek V3, or Llama 3 models - Research-grade framework with production backing (Z.ai) Consider alternatives when: - You need enterprise-grade stability features → use miles - You want flexible backend swapping → use verl - You need PyTorch-native abstractions → use torchforge - Training : Megatron-LM with full parallelism support (TP, PP, DP, SP) - Rollout : SGLang-based high-throughput generation with router - Data Buffer : Flexible prompt management and sample storage - Models : GLM-4.x, Qwen3, DeepSeek V3/R1, Llama 3 --- Use t…

查看 SKILL 详情