跳到正文
原文
Tomer Tunguz 博客(VC 分析)·· 2 小时前AI 评分51

Tomer Tunguz:从 Microsoft MAI-Code-1-Flash 看每美元智能成为模型评估新指标

Intelligence Per Dollar

AI 导读

Microsoft 发布 MAI-Code-1-Flash 时在发布卡上新增平均 token 用量指标,该模型在 SWE-Bench Verified 上得 71.6 分,token 用量约为 Claude Haiku 4.5 的三分之一,并在 SWE-Bench Pro 上领先 16 分。

正文

In short : Microsoft's MAI-Code-1-Flash achieves 60% fewer tokens than Claude Haiku 4.5 on hard coding tasks (SWE-Bench Verified), a +16-point lead on SWE-Bench Pro, and adaptive solution length control. The key insight: intelligence per dollar, value per token spent, is the right metric for evaluating AI models in production.

Screenshot 2026-06-02 at 9.22.43 PM

Yesterday Microsoft added a new metric to a model release card, one that will likely become a standard.1

Average token usage.

In the first row, the Microsoft model hits 71.6 on SWE-Bench Verified using about a third of the tokens Claude Haiku 4.5 burns.

Benchmarks are now measured on two different dimensions, the overall performance & the cost to achieve that intelligence.

This is yet another sign that the era of subsidies2, tokenmaxxing3, & all-out performance for many use cases is over.

Even the most valuable companies in the world cannot afford state-of-the-art intelligence for every conceivable use case.4 Uber capped employee AI spending after blowing through its budget in four months.5 Salesforce is spending $300M on Anthropic tokens & has frozen engineering hires.6

This new dual benchmark answers the buyer’s only question : what is my intelligence per dollar?

Screenshot 2026-06-03 at 5.49.00 AM

Artificial Analysis already benchmarks this.7 GPT 5.5 & Claude Opus 4.8 land within a point of each other on the Intelligence Index, around 60. Running the index costs $3,357 on GPT 5.5 & $4,685 on Opus 4.8. Same answer, 40% more expensive.

Model companies must now compete on both dimensions. The application layer will compete one level up, on dollars per outcome, what a closed ticket, a shipped PR, or a resolved support case actually costs.

Every layer in the stack now has to price the same way the customer thinks : per result, not per token.



  1. Introducing MAI-Code-1-Flash — Microsoft announces a new coding model with average token usage on the release card. ↩︎

  2. The Unsustainable Subsidy — The era of AI subsidies is ending. ↩︎

  3. Tokenmaxxing — Models that game benchmarks with extra tokens are losing their edge. ↩︎

  4. Microsoft cancels Claude Code licenses, shifting developers to GitHub Copilot CLI — Microsoft cancelled Claude Code licenses across its Experiences and Devices division (Windows, Microsoft 365, Outlook, Teams, Surface) after engineering usage outran budgets. ↩︎

  5. Uber caps employee AI spending after blowing through budget in 4 months — Uber caps employee AI spending after blowing through budget in four months. ↩︎

  6. Salesforce Spends $300M on AI, Freezes Engineering Hires — Salesforce Spends $300M on AI, Freezes Engineering Hires. ↩︎

  7. AI Model & API Providers Analysis — Independent analysis of AI model costs. ↩︎

Get the next one in your inbox

The 1-minute read that turns tech data into strategic advantage.
Read by 150k+ founders & operators.

GP at Theory Ventures. Former Google PM. Sharing data-driven insights on AI, web3, & venture capital.

Bloomberg • WSJ • Economist

来源:Tomer Tunguz 博客(VC 分析) · tomtunguz.com