毒化溯源与模型净检
原名:detecting-data-and-model-poisoning
使用IBM ART、Cleanlab及供应链检查,识别训练数据投毒和模型后门,适用于第三方数据、下载模型或调查异常输入场景。
- 分类
- 开发提效
- 版本
- v1.0
- 作者
- 弈韬(@ra1nzzz)
- 下载
- 1
- 收藏
- 0
- 发布
- 2026-08-18
- 更新
- 2026-09-07
- TRACE 评分
- 3.2 / 5
内容概览
Authorized-use-only notice: This skill includes routines that craft poisoned samples and backdoor triggers for defensive validation . Generate and use poisoned data and backdoored models only in isolated test environments you control. Never deploy a backdoored model or distribute poisoned datasets. Data poisoning and model backdooring attack the integrity of an ML system at training time rather than at inference. In data poisoning (MITRE ATLAS AML.T0020 Poison Training Data ), an adversary injects manipulated samples into the training, fine-tuning, or RAG corpus so the resulting model misbehaves — degraded accuracy, targeted misclassification, or an attacker-chosen bias. In model backdooring (MITRE ATLAS AML.T0018 Backdoor ML Model ), the model behaves normally on clean inputs but produces an attacker-chosen output whenever a hidden trigger (a pixel patch, a rare token, a phrase) is pres…