稀疏自编码器

原名:saelens

Train sparse autoencoders to interpret model features.

分类
开发提效
版本
v1.0.1
作者
弈韬(@ra1nzzz)
下载
2
收藏
0
发布
2026-08-25
更新
2026-09-15
TRACE 评分
3 / 5

内容概览

SAELens is the primary library for training and analyzing Sparse Autoencoders (SAEs) - a technique for decomposing polysemantic neural network activations into sparse, interpretable features. Based on Anthropic's groundbreaking research on monosemanticity. GitHub : jbloomAus/SAELens (1,100+ stars) Individual neurons in neural networks are polysemantic - they activate in multiple, semantically distinct contexts. This happens because models use superposition to represent more features than they have neurons, making interpretability difficult. SAEs solve this by decomposing dense activations into sparse, monosemantic features - typically only a small number of features activate for any given input, and each feature corresponds to an interpretable concept. Use SAELens when you need to: - Discover interpretable features in model activations - Understand what concepts a model has learned - Stu…

查看 SKILL 详情