跳到正文
原文
Google AI:DEV 作者专属(RSS)· Rishi G·· 7 小时前AI 评分25

Rashomon 作者探讨:编码智能体的网络行为如何独立验证

Your Coding Agent Has a Network. Do You Know What It Did?

AI 导读

Rashomon 作者提出,仅靠执行记录无法完整还原编码智能体的行为,还需网络层面的可见性。像 npm install 这类命令虽可验证是否运行,但无法确认其联系了哪些包注册源、Git 主机或外部服务,子智能体发起的网络请求也可能未出现在最终总结中。作者认为,将执行证据与网络活动结合,可为自主编码会话提供独立于智能体自述的验证依据。

正文

First off, thanks to everyone who read and commented on my last post! Your coding agent tells you all tests pass. Sometimes that's not true.

The discussion around that post got me thinking about another dimension of the same problem.

As I have been developing Rashomon, I have been focused on one question: what did the agent actually do?

Rashomon keeps an independent record of agent execution and compares it against what the agent claims happened. Did the command actually run? Did it fail? Did the expected test execute? Did a subagent do something that never made it into the final summary?

But there is another layer of agent behavior that execution records alone do not capture:

What happened on the network?

A command doesn't tell you everything.

Take something simple like:

npm install

You can verify that the command ran. But you do not necessarily know everything that happened as a result.

The process may have contacted package registries, downloaded files, reached Git hosts, or communicated with other external services.

Or imagine an agent says:

"I verified the API and deployed the fix."

You might be able to verify that it ran the relevant commands. But did it actually reach the API it claimed to reach? What other destinations did it contact? Did a subprocess or subagent make network requests that were not mentioned?

These are questions about what the agent communicated with, rather than simply what it executed.

Why this matters more with autonomous agents?

Coding agents are increasingly capable of operating without someone watching every command.

They can install dependencies, call APIs, interact with cloud services, spawn subagents, and run hundreds of commands during a single session.

The agent transcript gives you one account of what happened.

Execution evidence gives you another.

Network activity can give you another.

The interesting question is what happens when you start putting those pieces together.

If an agent says it deployed a change, for example, you could potentially look at both the deployment command it executed and the network activity associated with it.

That doesn't prove intent by itself, but it gives you more independent evidence to compare against the agent's claims.

This has gotten me interested in exploring how network-level visibility could complement what we are building with Rashomon.

Rashomon is focused on what actually ran. Network visibility could provide another piece of the picture: what those processes actually communicated with.

I think combining those perspectives could help answer much more interesting questions about autonomous coding sessions.

The broader idea is the same one behind Rashomon:

As coding agents become more autonomous, we should have independent evidence of what they actually did, rather than relying entirely on their own summary.

I'm curious what other developers think.

If you could see both what your coding agent executed and what it communicated with, what would you want to know?

来源:Google AI:DEV 作者专属(RSS) · dev.to