别再比较模型价格,要衡量"完成可用工作"的成本
Stop Comparing Model Prices. Measure Cost per Completed Work.
Fireworks 报告一位客户的对比数据:总 token 用量从 49.3K 降至 29.9K(减少 39%),任务得分从 0.751 微升至 0.753,但这只证明 token 效率,不足以支撑真实工具调用系统中的选型。作者主张以"被接受的工作产出"为分母,将模型费用、工具费、重试成本与人工审核成本合并计算,并固定提示词、工具与验收标准,只更换模型变量。
Model pricing is easy to compare. Completed work is harder.
Fireworks reports a customer comparison in which total token use moved from 49.3K to 29.9K, a 39% reduction, while the task score moved from 0.751 to 0.753. That is useful evidence for token efficiency. It is not enough evidence for choosing a model in a real tool-using system.
The missing unit is an accepted work product.
A model response can be generated without being usable. A coding agent can return a patch that fails the next test. A research agent can produce a fluent summary with weak source coverage. An operations agent can create a list that has no owner or deadline. The model bill captures tokens. The task also consumes retries, tool calls, context, wall time, and human correction.
I use four states:
- Generated output — something returned by the system.
- Candidate work — output that passed a format check.
- Accepted work — output that met the task contract.
- Reusable work — accepted output that another person can reopen.
For an evaluation, I keep the surrounding system fixed. The same person defines the task. The same prompt sets behavior. The same tools provide evidence. The same acceptance line judges both candidates. I change one variable: the model.
Then I run five real tasks: coding, research, writing, data, and recurring operations. I record tokens, tool calls, retries, wall time, human minutes, accepted status, reusable-artifact status, and the reason for rejection.
The calculation is deliberately plain:
model_cost = input_tokens * input_price + output_tokens * output_price
task_cost = model_cost + tool_fees + retry_cost + human_review_cost
accepted_work_cost = task_cost / accepted_work_products
The denominator must be visible. A ten-dollar run that produces five accepted outputs is not the same as a ten-dollar run that produces one draft and a long repair cycle.
The acceptance line must match the work. For code, I might require a reproduced bug, a narrow change, passing existing tests, one regression test, and no unrelated files. For research, I might require five primary sources, one claim per source, dates on current claims, named uncertainty, and a decision note. For operations, I might require a dated input, a reason for each exception, an owner, a deadline, and a record that can be reopened tomorrow.
This is also why I read the comparison through the Agent Stack™: Human, Behavior, Tools, Memory, Agents, Models. The model is one layer. Tools add outside work. Memory adds retrieval and repeated context. Agents add steps and retries. Human review defines acceptance. Comparing only the last layer hides the path that creates the bill.
The current sources are bounded. Fireworks' Ember-1 figures are vendor-reported and the model is presented as a Research Preview. OpenAI's evaluation guidance recommends task-specific tests and human judgment. Neither source proves the result for my workload. They make the next experiment clearer.
Before changing a model, run the work-unit trial. Measure the cost of something a person can accept, reopen, and use.
Canonical article: https://echonerve.com/stop-comparing-model-prices-measure-cost-per-completed-work/
来源:Google AI:DEV 作者专属(RSS) · dev.to