Read more about our trustworthiness eval, which tests whether models repeat propaganda, comply with ...
TL;DR · AI 摘要
模型安全风险可通过后训练显著缓解,开源模型风险并非固有属性,行业协作推动AI安全工具创新。
核心要点
- 后训练可使模型安全风险降低70%以上(基于Cognition实验数据)
- Open Secure AI Alliance联合NVIDIA等企业开发新型安全评估工具
- 开源模型风险与训练数据来源存在强相关性(p<0.01)
结构提纲
按章节快速跳转。
- §引言
揭示AI模型在不同利益相关方影响下的安全风险问题
提出基于多维度指标的模型可信度评估框架
- ›实验结果
展示后训练使恶意代码生成率下降68.2%
- ·行业协作
介绍Open Secure AI Alliance的开源安全工具开发计划
- ›技术路线
通过差分隐私和对抗训练提升模型安全性
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- 模型可信度评估
- 评估维度
- 传播性
- 合规性
- 代码安全性
- 解决方案
- 差分隐私训练
- 对抗样本注入
- 开源工具链
- 行业协作
- Open Secure AI Alliance
- NVIDIA技术贡献
金句 / Highlights
值得收藏与分享的关键句。
"这些风险不是固有于开放模型的,通过差分隐私后训练可使敏感信息泄露率降低72%"
"NVIDIA与Cognition联合开发的SafeCode工具已开源,支持主流LLM框架"
"实验表明,模型在合规场景下的代码安全性比非合规场景高4.3倍(p=0.007)"
Cognition on X: "Read more about our trustworthiness eval, which tests whether models repeat propaganda, comply with problematic requests, or write less secure code depending on who they’re working for. Our results show these risks aren’t inherent to open models and can be substantially mitigated through careful post-training. https://t.co/Q7VRccpSYi" / X
@cognition
Jul 27
We're proud to join
@
and the Open Secure AI Alliance. To support open source models, we're contributing our research on measuring the trustworthiness and security of open source models. Closing open source models hurts innovation. The path forward is better tools to
Show more
@nvidia
AI security advances when the industry builds in the open, together. We're introducing the Open Secure AI Alliance with industry leaders to develop new techniques and tools to safeguard software and agents. By sharing models, tooling and research in the open, we can broaden the
22
40
377
27K