破障
原名:obliteratus
OBLITERATUS: abliterate LLM refusals (diff-in-means).
- 分类
- 开发提效
- 版本
- v2.0.0
- 作者
- 弈韬(@ra1nzzz)
- 下载
- 2
- 收藏
- 0
- 发布
- 2026-08-25
- 更新
- 2026-09-13
- TRACE 评分
- 4.2 / 5
内容概览
9 CLI methods, 28 analysis modules, 116 model presets across 5 compute tiers, tournament evaluation, and telemetry-driven recommendations. Remove refusal behaviors (guardrails) from open-weight LLMs without retraining or fine-tuning. Uses mechanistic interpretability techniques — including diff-in-means, SVD, whitened SVD, LEACE concept erasure, SAE decomposition, Bayesian kernel projection, and more — to identify and surgically excise refusal directions from model weights while preserving reasoning capabilities. License warning: OBLITERATUS is AGPL-3.0. NEVER import it as a Python library. Always invoke via CLI (obliteratus command) or subprocess. This keeps Hermes Agent's MIT license clean. Walkthrough of OBLITERATUS used by a Hermes agent to abliterate Gemma: https://www.youtube.com/watch?v=8fG9BrNTeHs ("OBLITERATUS: An AI Agent Removed Gemma 4's Safety Guardrails") Useful when the us…