跳到正文
Diogo Almeida· @CompleteSkeptic · X·· 2 小时前AI 评分13
AI 导读

benchmaxxed 模型和那些由真正在乎可靠性的人做出的模型之间,存在一个关键差异,而这个差异按定义就不会体现在基准测试上 (对于字符串 LLM,我们多少已经认识到这一点了……)

正文

there's a critical difference between benchmaxxed models and ones made by people who truly care about reliability that, by definition, doesn't show up on benchmarks

(we've somewhat learned this for string LLMs...)

引用Max Rovensky@MaxRovensky
@julianharris There's something wet-sponge-feeling about the qwen fine tunes that I can't explain compared to Jev
在 X 查看被引用的帖子

来源:Diogo Almeida · x.com