DAIR.AI 创始人 Elvis Saravia 用 AI 员工 Viktor 自动分析 Agent 的夜间框架实验评估结果。当 23 个昨日通过的任务今日失败时,Viktor 会检查全部日志、追溯到导致失败的单一改动并建议回滚,Elvis 只需确认后决定是否采纳。Viktor 还会主动标记问题,免费试用含 $100 额度、无需绑卡。
Reading eval results is now the slowest part of building agents.
I'm Elvis, founder of @dair_ai. I lead research, build, and teach about AI agents.
I run harness experiments every night, but reading the results was eating my mornings.
Every change to my harness gets evaluated overnight, whether it touches memory, tool use, or context compaction.
The morning after is the hard part. I check which tasks my agent got right yesterday but wrong today. Then I open the logs for each failure, one by one, to figure out which of my changes caused it.
I tried a dashboard first. It showed the pass rate dropped. It couldn't tell me why.
That is the job Viktor, an AI employee in Slack, is built for. He reviews the results overnight.
Here is how that plays out. Say 23 tasks that passed yesterday fail today. Viktor checks all 23 logs, traces them to the one change that caused them, and suggests undoing it. I check the logs and make the call.
Viktor does the digging. I decide what goes into the harness.
He is also proactive. He flags problems before you ask, which helps you stay on track with complex eval runs and other research tasks.
Harness engineers, do you check every eval run, or only when the pass rate drops?
Try free at @viktor_com. $100 in credits, no card. Full link in my first reply.
Thanks to the team for partnering with me on this post
来源:elvis · x.com