Nathan Lambert 表示当前数据行业净产出质量仍太低,会放大 reward hacking 等行为,他正与 Mercor 合作构建可扩展的 RL 样本高效数据方法。他强调需推进评估科学前沿并针对明确经济价值领域构建专用 benchmark,并看好前沿实验室围绕真实世界评估与合成数据方法展开重大投入,认为这是让更多经济体首次“感受到 AGI”的研究方向。
I'd frame it as follows: The data industry has taken off in recent years, but the quality of our net output is still far too low (amplifying behaviors like reward hacking).
We're working to build scalable methods for creating sample-efficient data for RL. In order to keep this pipeline going, we need to push the frontier of evaluation science, while building specific benchmarks to hillclimb on areas of clear economic value.
I've been advising Mercor on how to build this research direction effectively. These are my views, but I'm confident we're going to see a major investment from economy around the frontier labs (open inference, open post-training, and data) orient around expertise in building real-world representative evals and synthetic data methods to scaling training data around them.
I’m personally very excited about this, it is the research that will make more of the economy “feel the AGI” for the first time.
(And, there’s a big opportunity to build this on open models.)
Mercor is building a world-class research team. As the leading AI data provider, we are uniquely positioned to combine benchmarks, data production, model training, and economics research to advance model productivity. We are committed to sharing our findings with the world. DM me if this sounds exciting. https://www.mercor.com/blog/why-mercor-is-building-a-research-team/在 X 查看被引用的帖子
来源:Nathan Lambert · x.com