跳到正文
Artificial Analysis· @ArtificialAnlys · X·· 2 小时前精选AI 评分67
AI 导读

蚂蚁集团发布推理模型 Ling 3.1 Flash,总参数 560B、激活 25B、上下文窗口 1M token,在 Artificial Analysis Intelligence Index v4.3 得分 41,较 Ling 3.0 Flash 的 20 大幅提升,智能体能力改善明显。

推荐理由

第三方评测给出了智能指数提升幅度、幻觉率变化和每任务成本,可与 GLM、DeepSeek、Gemini 同级模型直接对比。

正文 · 原文

Ling 3.1 Flash makes large gains in intelligence over its predecessor, scoring 41 on the Artificial Analysis Intelligence Index with particular improvement in agentic capabilities

@AntLingAGI has released Ling 3.1 Flash, a reasoning model with 560B total parameters, 25B active parameters, and a 1M token context window. It scores 41 on the Artificial Analysis Intelligence Index v4.3, up from 20 from Ling 3.0 Flash, which launched in August. Ling 3.1 Flash is expected to be open weights, with Ant Group releasing the weights soon.

Key results:

➤ Ling 3.1 Flash improves agentic capabilities over its predecessor. With a GDPval-AA v2 Elo of 1,622 and an AA-Briefcase Elo of 1,400, the model rivals peer models such as GLM-5.3-Flash, Gemini 3.8 Flash (high) and DeepSeek V4.1 Flash (Max). Ling 3.1 Flash also shows notable gains on AutomationBench-AA (62%) and Terminal-Bench v4.0 (33%).

➤ Ling 3.1 Flash scores +2 on AA-Omniscience, a 22 point improvements from Ling 3.0 Flash (-18). Compared to its predecessor, its accuracy rate rose from 18% to 29% while the hallucination rate fell from 44% to 38% at a similar attempt rate (56% to 58%), demonstrating the gain comes primarily from knowing more, not abstaining more. For comparison, DeepSeek V4.1 Flash (Max) has a hallucination rate of 97%.

➤ Ling 3.1 Flash is a larger model than its predecessor, priced accordingly, but more token efficient. It has 560B total and 25B active parameters, up from 124B and 5.1B for Ling 3.0 Flash, and is priced at $0.30 / $0.90 per 1M input / output tokens, up from $0.075 / $0.22. It used 218M output tokens to run the Intelligence Index, 16% fewer than Ling 3.0 Flash (261M).

➤ Ling 3.1 Flash has a higher cost per task than some peers, but remains in the most attractive quadrant. Ling 3.1 Flash costs $0.99 per task, well above GLM-5.3-Flash ($0.42) and DeepSeek V4.1 Flash (Max, $0.32) but below Gemini 3.8 Flash (High) at $1.24.

Additional model details:

➤ Size: 560B total parameters, 25B active

➤ Context window: 1M tokens

➤ Pricing: $0.30 per 1M input tokens and $0.90 per 1M output tokens, with cached input at $0.06 per 1M (80% discount)

➤ Availability: Accessible through Novita AI, with weights coming soon

来源:Artificial Analysis · x.com