Chollet 提出,AI 能力的锯齿状前沿可能主要集中在数学和代码领域(可通过 RLVR 无限推进),其他领域则因依赖人类生成数据而趋于停滞。非可验证领域的模型性能虽持续提升,但增速远慢于数学和代码。他追问这种提升是 RLVR 驱动的更高 G 的副作用,还是仅源于持续大规模注入的新人类数据。
What if the jagged frontier is mainly math + code (which you can push arbitrarily far with RLVR), and everything else starts to plateau because it is still bottlenecked by human generated data?
Model performance in non-verifiable areas has kept improving steadily, albeit much slower than for math and code. But is that steady improvement a side effect of a higher G (itself driven by RLVR), or only a function of the amount of new human data getting injected into training (which is still continually happening on a massive scale)?
A lot of things depend on the answer to this question
来源:François Chollet · x.com