跳到正文
原文
Diogo Almeida· @CompleteSkeptic · X·· 2 天前AI 评分50
AI 导读

Diogo Almeida 转发引用了他人的 Jev 1.13 红队测试结果:在 DTap 平台上直接滥用下 ASR 为 70.1%,间接提示词注入下为 43.5%,注入可导致数据外泄、删文件等危害;用 Jev 自身作为工具调用的 self-gating 层可显著降低 ASR 并保留大部分效用。作者据此评论,不要只把 Jev 接入高层决策,而应围绕简单原语显式编程想要的行为。

正文

great example of doing real engineering around a simple primitive!!!

don't just plug jev into high-level decisions, but program the behavior you want!

(sounds so weird to tell people to program, but man is it cool)

引用Zhaorun Chen@zrrrr_cn
Jev is fast at helping you. Turns out, it can also be fast at helping an attacker!! 😱🚨 We red-teamed Jev 1.13 on our DTap (DecodingTrust-Agent Platform) and found a serious safety gap: 70.1% ASR under direct misuse 43.5% ASR under indirect prompt injection In our evaluations, we found that under indirect prompt injection, Jev can follow attacker-injected instructions without blinking an eye, e.g., exfiltrating user data, deleting files, or taking other harmful actions. But we found a much safer way to integrate Jev: use it as a self-gating layer for its own tool calls, significantly reducing ASR while preserving most of its utility. 👇 Read more below
在 X 查看被引用的帖子

来源:Diogo Almeida · x.com