跳到正文
DAIR.AI· @dair_ai · X·· 3 小时前AI 评分48
AI 导读

针对上下文压缩在长程智能体中的瓶颈,研究提出 PAIR:从同一状态带与不带压缩重放智能体,避免整轮对比中的随机差异。典型压缩仅多出几步,但成功率的骤降来自少数压缩事件——如 Venmo 任务中摘要丢掉了“仅限同事”过滤条件,把全部 36 笔付款总额当作答案;另有压缩把已读的 API 说明变成含糊描述,导致智能体重开文档并再次登录,多出约五步。

正文

Context compression is a huge bottleneck for long-running agents.

This work finds that context compression hurts long-horizon agents at a few specific points.

They propose PAIR, which replays the agent from the same state with and without a given compression, instead of comparing whole runs that differ in many random ways.

A typical compression adds a few extra steps. The large drops in success come from a small number of compression events.

The harmful compressions drop task conditions the agent hasn't resolved yet. In one Venmo task, the summary dropped the "only from coworkers" filter and reported the total of all 36 payments as the answer.

Other compressions reduce API specs the agent already read to vague prose, so the agent reopens the docs and logs in again, which adds about five steps.

PAIR then diagnoses what information those compressions dropped and rewrites the matching sections of the compression prompt. The agent, compressor model, and tools stay fixed.

On AppWorld, OfficeBench, and tau-Bench Retail, it gives the most consistent task completion of any compressed method and comes close to running with no compression at all.

Compression also lowers run-to-run reliability before it makes tasks unsolvable, so check consistency across repeated runs in your own evals.

Paper: https://arxiv.org/abs/2609.36526

Chat with Paper: https://academy.dair.ai/papers/adapting-context-compression-for-long-horizon-agents-with-counterfactual-continu-2609.36526

来源:DAIR.AI · x.com