Diogo Almeida· @CompleteSkeptic · X·· 2 天前AI 评分50
AI 导读
Diogo Almeida 转发引用了他人的 Jev 1.13 红队测试结果:在 DTap 平台上直接滥用下 ASR 为 70.1%,间接提示词注入下为 43.5%,注入可导致数据外泄、删文件等危害;用 Jev 自身作为工具调用的 self-gating 层可显著降低 ASR 并保留大部分效用。作者据此评论,不要只把 Jev 接入高层决策,而应围绕简单原语显式编程想要的行为。
正文
great example of doing real engineering around a simple primitive!!!
don't just plug jev into high-level decisions, but program the behavior you want!
(sounds so weird to tell people to program, but man is it cool)
Jev is fast at helping you. Turns out, it can also be fast at helping an attacker!! 😱🚨 We red-teamed Jev 1.13 on our DTap (DecodingTrust-Agent Platform) and found a serious safety gap: 70.1% ASR under direct misuse 43.5% ASR under indirect prompt injection In our evaluations, we found that under indirect prompt injection, Jev can follow attacker-injected instructions without blinking an eye, e.g., exfiltrating user data, deleting files, or taking other harmful actions. But we found a much safer way to integrate Jev: use it as a self-gating layer for its own tool calls, significantly reducing ASR while preserving most of its utility. 👇 Read more below在 X 查看被引用的帖子
来源:Diogo Almeida · x.com