跳到正文
原文
Rohan Paul· @rohanpaul_ai · X·· 3 小时前AI 评分45
AI 导读

一项针对长周期决策的论文测试了包括GPT-5.6 Sol和Claude Opus 4.8在内的八款领先模型,在需要一年连贯决策、延迟反馈并承担自身行为后果的任务中,模型表现相对人类全面崩溃。表现最好的Qwen3.7-Max搭配Hermes,最终收益仅为人类平均参与者的27.3%。

正文

This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans.

The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant.

A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.

引用Rohan Paul@rohanpaul_ai
"Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (link in comment)
在 X 查看被引用的帖子

来源:Rohan Paul · x.com