Diogo Almeida· @CompleteSkeptic · X·· 2 小时前AI 评分13
AI 导读
benchmaxxed 模型和那些由真正在乎可靠性的人做出的模型之间,存在一个关键差异,而这个差异按定义就不会体现在基准测试上 (对于字符串 LLM,我们多少已经认识到这一点了……)
正文
there's a critical difference between benchmaxxed models and ones made by people who truly care about reliability that, by definition, doesn't show up on benchmarks
(we've somewhat learned this for string LLMs...)
@julianharris There's something wet-sponge-feeling about the qwen fine tunes that I can't explain compared to Jev在 X 查看被引用的帖子
来源:Diogo Almeida · x.com