Jevons 悖论与 AI 判断:TypeSafe 的 Jev 模型为何让昂贵判断更贵
The Jevons Paradox of Judgment
TypeSafe 将其模型命名为 Jev,押注单次判断成本降至几分之一美分后判断总量将爆发式增长。文章指出,可自动化的判断是错误代价由系统承担的那部分,剩余交给人的判断由问责制定义,因此廉价判断越多,无法自动化的昂贵判断相对权重越高。
The Jevons Paradox of Judgment
In 1865, a British economist named William Stanley Jevons published a book called The Coal Question, and in it he made an argument that his contemporaries found hard to swallow. The more efficiently a steam engine used coal, he wrote, the more coal Britain would burn. He was not predicting restraint. He was predicting appetite. His colleagues thought he had the sign backwards, and for a long time the argument sat in the drawer as a curiosity — until the twentieth century kept proving him right, in electricity, in fuel, in road capacity, in paper, in storage. The pattern acquired a nickname that outlived the man: the Jevons paradox.
The mechanism is not mysterious, and the sloppy version of it ("efficiency makes us consume more") is not what he said. He said the cost of a unit of useful work falls, so you perform more units. Efficiency is a price cut, and price cuts move quantity.
Now put the word judgment where the word coal was.
TypeSafe named their model Jev on purpose — the System One framing comes from Kahneman, and the name from Jevons. It is a good name, because it is an honest bet: they are wagering that once a single judgment costs a fraction of a cent, the total number of judgments being made will explode. They are almost certainly right about the explosion. The interesting question is what happens to the judgments that did not get cheaper, and that question is where most of the discussion stops.
Where the cheap judgments come from
An ATM is worth an aside here, because it is the cleanest natural experiment we have. When banks started putting cash machines on the wall, the obvious forecast was fewer tellers. The United States still employed 339,200 of them in 2025, and the official projection is a 13 percent decline over the following decade — 44,700 jobs, not a collapse. The same projection expects roughly 26,800 openings a year anyway, almost all of them to replace people who leave. The headcount barely moved; the job moved. Note what the official description says tellers do: process routine transactions. The machine took the transaction. What remained was the rest — selling products, untangling the accounts the machine could not reason about, handling customers who arrived already upset.
This is the part of the paradox that does not get repeated often enough. When you automate the cheap end of a category, you do not shrink the category. You re-sort it.
If you have ever had to sort a queue by hand, this is obvious. The items that are easy to sort are the ones you stop thinking about after the first hundred. The hard ones are hard for a reason: the answer is contested, the categories overlap, the cost of a wrong bucket is not a re-run. A filter that removes everything easy does not leave you with an easier pile. It leaves you holding exactly the cases that were never eligible for the filter.
What the cheap end actually contains
What gets automated is not a random sample of judgment. It is a specific slice with specific properties: the answer is one option from a set somebody defined; the wrong answer is recoverable at a bounded price; a downstream check would catch it anyway; the same call is made thousands of times a day. Those four properties are what make a judgment automatable. They are also what make it cheap to be wrong about.
Read that list again and notice what is missing from it: the name of a person who has to live with the outcome. The judgments that a millisecond classifier can take are exactly the judgments where being wrong is paid for by a system — a retry, a log line, a customer who asks again. The judgments it cannot take are the ones where being wrong is paid for by a human being, a relationship, a reputation, or a balance sheet with somebody's name under it.
So the selection effect is not just "the easy ones go first." It is sharper than that. The boundary of the automatable set is drawn by where the consequences land. Automate the judgments whose cost is borne by a system, and the residual set — the part still routed to a person — is defined by accountability. That is not a temporary state of affairs waiting for a better model. It is what the residual consists of.
This is the uncomfortable half of the Jevons story for anyone selling cheap judgment. The paradox says the category grows. It does not say its center of gravity stays put. Every unit of judgment that moves into the cheap layer raises the relative weight of the units that cannot, because the total pool of decisions is now larger and the expensive tail has been left untouched.
Cheaper judgment buys more judgment
The second layer is the one Jevons himself was describing, and the one that gets skipped: a price cut does not just change how existing work is done. It changes how much work is worth initiating.
Nobody writes a rule that says "log everything," and nobody decides that every comment in a community needs a moderation verdict. Those decisions get made per-item, and per-item they are a function of price. At a fraction of a cent per call, there is no longer a reason to be selective. The system starts inserting a judgment at every point where one could possibly help — should this be retained, should this be flagged, should this be escalated, should this be sent to a person.
Each of those insertions is cheap. The aggregate is not, because the aggregate is not measured in tokens. It is measured in the queue of things that need a human being to sign off, and that queue is fed by exactly these calls.
Follow one judgment through a pipeline and the shape becomes visible. A cheap judge returns a verdict with a number attached. The number sits above the threshold and the verdict executes — no human involved, correctly. The number sits below the threshold and the item is escalated, which means it joins a queue. Now multiply by the volume that cheap judgment made economical in the first place, and notice that the machine that was supposed to reduce human involvement has become the most efficient generator of human work ever built.
The escalations are not failures. They are the design: the model is honest about the limit of its own confidence, and the limit gets routed upward. But routing upward is a producer of responsibility, and responsibility is the one thing nobody has found a way to make cheap.
Two things fall, one does not
Cost curves in this industry are steep, and it is reasonable to expect the cliff to keep going. Jev runs at 70–500 milliseconds per call, priced at $0.042 per million input tokens with output free — the vendor's own figures, and by any prior standard of software economics they are extraordinary. Inference gets cheaper every quarter; none of those curves is heading toward zero cost of being wrong.
The per-call cost falls. Time-to-verdict falls. The cost of generating a defensible answer falls, since the alternative is a senior person's hour. But the cost of owning the outcome does not appear on any of those curves. It is not denominated in tokens; it is denominated in consequences, and consequences are paid in the currency of the person who chose.
Practical evidence, from our own published record rather than a rhetorical point. Our calibration data is self-run under named benchmarks with the failures disclosed: on JudgeBench, 620 judgments, with the 6 that failed on first verdict disclosed rather than quietly retried. Reported at 90% confidence or above, our judgments were right 99.6% of the time; in the 80–90% band, 94.0%. What matters as much is the number underneath: there is a band where the system is genuinely uncertain, and its calibration says so in public.
A calibrated confidence is useful precisely because it is a map of where the machine should stop: below some line the question is not a computation but a choice somebody has to own. The cheaper judgment gets, the more precisely that line is drawn — and the more clearly the stuff above it is separated from the stuff that was never about inference at all.
The re-pricing, stated as an economic claim
Put the three layers together and the claim is clean.
A judgment is a bundle of two things: the inference, and the accountability. For most of history they were produced together, by the same person, which is why the bundle looked indivisible — the expensive part of deciding seemed to be the thinking. What cheap judgment reveals is that they were never the same good. Inference is a computation, and computations get cheap on a predictable schedule. Accountability is a commitment made by an identifiable party, and it has no efficiency curve, because there is nothing in it to optimize. You cannot amortize it, batch it, or cache it. It does not get faster with better hardware.
So the re-pricing goes like this. The cheap half of the bundle is unbundled and commoditized — that is Jev, and Jev is very good at it. The expensive half is not merely left over; it is sharpened, because once inference is nearly free, the only remaining explanation for why a decision is hard is that somebody has to own it. Judgment did not get cheap. A particular component of judgment got cheap, and the residual component got more visible, more isolated, and more expensive relative to everything around it.
This is why "it is only a matter of time before the hard ones are automated too" does not follow from the trend line. The trend line describes inference. The hard ones are not hard because inference is difficult; they are hard because both candidates are defensible, the outcome is not reversible at the same price, and a person has to own the result. You can make the inference arbitrarily cheap. The signature does not become cheaper, because its cost was never computational.
Close
Everything that can be made cheap will be, and the volume of judgment will rise accordingly — Jevons was right about coal, and the same argument holds here. What will not fall is the price of being the one who decided: accountability has no efficiency curve, because there is nothing in it for engineering to make efficient.
If you are holding one of the questions that stayed expensive — one question, two answers that both survive scrutiny — that is the case we build for. Decider is here: a pick, a calibrated confidence, and a written argument you can disagree with, for the decisions that still end with a person's name.
Sources
- Tellers, Occupational Outlook Handbook, U.S. Bureau of Labor Statistics (employment 339,200 in 2025; projected change -13 percent and -44,700 jobs, 2025-35; about 26,800 openings a year; median pay $43,030 in May 2025): https://www.bls.gov/ooh/office-and-administrative-support/tellers.htm
Our calibration figures are self-run under the named benchmarks with the failures disclosed, and are published in full in the machine-readable file referenced in the text.
来源:Google AI:DEV 作者专属(RSS) · dev.to