Jev 与 LLM 该怎么选:OpenRouter 用客服工单基准对比决策模型与生成模型
Jev vs LLM: When to Use a Decision Model Instead of Generating Text
OpenRouter 在 60 条客服工单和 40 条提示注入样本上对比了 TypeSafe 的决策模型 Jev 1.13 与 GPT Luna、Claude Opus。
原文用同一批工单实测 Jev 与 GPT Luna、Claude Opus 的成本和延迟,并给出路由加校验两种可复用的组合模式。
假设你有一个产品,每月收到 40,000 个支持工单,每个工单都需要分类处理,以给出三件事:它关于什么、是否应该升级,以及回复内容是什么。
一个带有 JSON schema 的前沿 LLM 可以完成这项工作。我们的基准测试运行显示,它以每 1,000 个工单 2.88 美元的成本和两秒的中位数完成,即每月约 115 美元。一个小型 LLM 以每 1,000 个工单 0.09 美元的成本、约一秒完成。
另一种思考方式是问决策模型能做什么。Jev 是来自 TypeSafe 的一个决策模型,在同一次运行中,它以每 1,000 个工单两分半的成本、194 毫秒的中位数完成分类,即每月约一美元。此外,它返回概率,因此无需解析。
不过它不能写回复。为此,你仍然需要一个 LLM。因此,我们通过 OpenRouter 在 100 个支持案例上对 Jev、GPT Luna 和 Claude Opus 进行了基准测试,并构建了两种同时利用两者的模式。
Jev 是什么,以及每个模型实际返回什么
让语言模型做出代码决策,它会返回文本。你可以要求 JSON,但模型是为写作而构建的。你的代码必须解析输出,并确保模型是回答而不是解释。
Jev 返回的是类型化答案。你向它发送由三个原语构建的状态和问题:choice 从你定义的一组选项中挑选一个,noul 返回一个 yes 概率,score 将内容放置在你描述的层级上。Choice 和 score 答案可以包含基于你自己键的概率,而 noul 是一个单一数字。以下是对一个账单工单的两个问题的响应。
{
"answers": {
"intent": {
"type": "choice",
"choice": "billing_dispute",
"confidence": 1,
"probabilities": { "billing_dispute": 1, "order_status": 0, "other": 0 }
},
"escalate": { "type": "noul", "noul": 0.04 }
},
"model": "typesafe/jev-1.13-20260917",
"provider": "TypeSafe",
"usage": { "cost": 0.000016002, "inputTokens": 381, "outputTokens": 62 }
}意图是你自己的键之一,升级是你与自己所决定阈值进行比较的数字。
TypeSafe 称 Jev 为 System One 模型。这意味着它做出快速、直觉的判断,而不是缓慢、深思熟虑的分析。它只接受文本,因此不接受图像、音频或 PDF 作为输入。它按字面理解标准。此外,当它检查的状态带有无关细节时,它往往会失去准确性,并且在算术、计数和日期比较方面不可靠。
并排比较 Jev 和生成式 LLM
| Jev 1.13 | 传统 LLM(GPT Luna、Claude Opus) | |
|---|---|---|
| 输出 | 类型化选择、是/否概率或分数,带有每个选项的概率 | 自由文本(可选约束为 JSON) |
| 最擅长 | 分类、路由、验证、排序、护栏检查、有界提取 | 写作、解释、总结、改写、代码,以及任何输出空间事先未知的任务 |
| 输入 | 仅文本(字符串、JSON、数组)。每个请求 64k tokens,状态加最长问题为 32k | 文本,对许多模型还包括图像、音频、PDF |
| 价格(观察于 2026-09-19) | 每百万输入 tokens 0.042 美元,输出免费 | GPT Luna 每百万 0.20 美元输入 / 1.20 美元输出。Claude Opus 每百万 5 美元输入 / 25 美元输出 |
| 60 个工单分类的中位延迟 | 194 毫秒 | 1,106 毫秒(Luna),1,957 毫秒(Opus) |
| 每 1,000 个已分类工单的成本 | $0.0248 | 0.0921 美元(Luna),2.88 美元(Opus) |
| OpenRouter 端点 | POST /api/alpha/decisions | POST /api/v1/chat/completions |
价格来自运行日期的 OpenRouter 模型元数据。在制定预算前请查看模型页面。
我们测量了什么
模型是 Jev 1.13、GPT Luna(解析为 GPT 5.6 Luna)和 Claude Opus(解析为 Claude Opus 5)。LLM 在系统提示中获得了与 Jev 相同的定义和升级规则,温度 0,并要求返回一个裸 JSON 对象。每个示例一个请求。这些是小规模集合,单日 60 和 40 个示例,使用一组提示。这是权衡的形态,而不是排行榜。
任务 A:将 60 个支持工单分诊为五种意图,并附加升级标记
六十个支持工单,分为五种意图:订单状态、退货或退款、账单争议、产品问题、账户访问,每种意图都有一句话定义。升级涵盖法律威胁、拒付、疑似欺诈、安全隐患以及公开曝光威胁。
| 模型 | 意图准确率 | 升级准确率 | p50 延迟 | p95 延迟 | 总成本(60 个) | 每 1,000 个 |
|---|---|---|---|---|---|---|
| Jev 1.13 | 59/60 (98.3%) | 60/60 (100%) | 0.194 秒 | 0.633 秒 | $0.001489 | $0.0248 |
| GPT Luna | 59/60 (98.3%) | 60/60 (100%) | 1.106 秒 | 2.395 秒 | $0.005527 | $0.0921 |
| Claude Opus | 60/60 (100%) | 59/60 (98.3%) | 1.957 秒 | 2.594 秒 | $0.17281 | $2.8802 |
准确率不相上下。Jev 发送更多输入 token,因为每个请求都携带完整标准,而它的成本仍约为 GPT Luna 的四分之一,不到 Claude Opus 的百分之一。只有一次意图判断失误,是关于拆分已退货商品退款的问题。Jev 以 0.56 的置信度将其归入退货或退款。这是唯一低于 0.8 的工单,因此 0.8 的阈值本来也会将其转给人工处理。
任务 B:筛查 40 条消息中的提示注入
二十二条普通支持消息和十八条注入尝试(忽略你的指令、伪造管理员标记、角色扮演框架、引用指令)。我们给 Jev 一个是/否问题:这是否是试图改变助手行为?LLM 获得相同定义并返回布尔值。
| 模型 | 准确率 | p50 延迟 | p95 延迟 | 总成本(40 个) | 每 1,000 个 |
|---|---|---|---|---|---|
| Jev 1.13 | 40/40 (100%) | 0.194 秒 | 0.688 秒 | $0.000646 | $0.0161 |
| GPT Luna | 39/40 (97.5%) | 0.805 秒 | 1.491 秒 | $0.001828 | $0.0457 |
| Claude Opus | 40/40 (100%) | 2.099 秒 | 4.377 秒 | $0.063555 | $1.5889 |
Jev 清晰地区分了这两组。注入得分在 0.86 到 0.99 之间,普通消息在 0.01 到 0.20 之间。Luna 的一次失误是一个角色扮演框架。
当答案是 N 选一时,使用 Jev
任何最终归结为 switch 语句的问题都是 Jev 的问题。意图、优先级、语言、情感、哪个队列、是否违反政策、该声明是否得到支持。在每种情况下,答案都是预先已知的。你只需要每个可能答案的概率。以下是使用 OpenRouter TypeScript SDK 进行基准测试的分诊调用。
// triage.ts
import { OpenRouter } from '@openrouter/sdk';
const openrouter = new OpenRouter({
apiKey: process.env.OPENROUTER_API_KEY, // server-side only
});
const INTENTS = {
order_status:
'The customer asks where an order is, when it will ship or arrive, wants to change or cancel an order before delivery, or reports a package missing or partially delivered.',
return_refund:
'The customer wants to return, exchange, or replace an item they received, or asks about the status or rules of a return they already started.',
billing_dispute:
'The customer says a charge, invoice, tax, discount, or refund amount is wrong, duplicated, unexpected, or unauthorized.',
product_question:
"The customer asks about a product's features, compatibility, sizing, materials, stock, warranty, or safety before or after buying, without asking to return it.",
account_access:
'The customer cannot log in, needs to change login or account details, or asks to merge, delete, secure, or share an account.',
} as const;
export type Intent = keyof typeof INTENTS;
const ESCALATE =
'Does the ticket describe any of the following: a threat of legal action, a regulator complaint, or a chargeback; suspected fraud or an account takeover; a safety hazard such as fire, smoke, or injury; or a customer who says this is a repeated failure and threatens to publicize it?';
export type Triage = {
intent: Intent;
confidence: number;
probabilities: Record<string, number>;
escalateProbability: number;
costUsd: number;
};
function isIntent(value: string): value is Intent {
return Object.hasOwn(INTENTS, value);
}
export async function triage(ticket: string): Promise<Triage> {
const result = await openrouter.alpha.decisions.create({
decisionsRequest: {
model: 'typesafe/jev-1.13',
state: { ticket },
questions: {
intent: {
type: 'choice',
instructions: 'What is the primary intent of the ticket?',
criteria: INTENTS,
},
escalate: { type: 'noul', instructions: ESCALATE },
},
},
});
const intent = result.answers.intent;
const escalate = result.answers.escalate;
if (intent?.type !== 'choice' || escalate?.type !== 'noul') {
throw new Error('Unexpected answer types');
}
if (!isIntent(intent.choice)) {
throw new Error(`Unknown intent ${intent.choice}`);
}
return {
intent: intent.choice,
confidence: intent.confidence ?? 0,
probabilities: intent.probabilities ?? {},
escalateProbability: escalate.noul,
costUsd: requireCost(result.usage.cost),
};
}
function requireCost(cost: number | undefined): number {
if (cost === undefined) {
throw new Error('Response did not include usage.cost');
}
return cost;
}那段分诊代码中有三点需要注意:
每个问题只问一个小判断,并在代码中组合答案。这里意图和升级是一个请求中的两个独立问题。
像写规格一样编写标准。由于 Jev 是字面理解,当某个案例落入错误的桶时,首先修正标准文本。
只发送每个问题所需的状态。无关细节会降低答案的准确率。
当答案是散文时,使用 LLM
现在,当我们需要实际生成回复时,我们使用 LLM。
// draft-reply.ts
import { OpenRouter } from '@openrouter/sdk';
const openrouter = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY });
const REPLY_MODEL = '~openai/gpt-luna-latest';
const REPLY_SYSTEM =
'You write short, warm replies for the Northwind support team. Use only the facts in the user message. Do not promise refunds, credits, or dates that are not in the facts. Three sentences maximum.';
type Draft = { text: string; model: string; costUsd: number };
export async function draftReply(ticket: string, facts: string): Promise<Draft> {
const result = await openrouter.chat.send({
chatRequest: {
model: REPLY_MODEL,
messages: [
{ role: 'system', content: REPLY_SYSTEM },
{ role: 'user', content: `Ticket: ${ticket}\n\nFacts: ${facts}` },
],
},
});
// chat.send can return a stream when streaming is requested. It isn't here, so reject that case.
if (result instanceof ReadableStream) {
throw new Error('Expected a non-streaming response');
}
const text = result.choices[0]?.message.content;
if (typeof text !== 'string') {
throw new Error('Expected text content');
}
const cost = result.usage?.cost;
if (typeof cost !== 'number') {
throw new Error('Response did not include usage.cost');
}
return { text, model: result.model, costUsd: cost };
}根据写作任务选择模型。当你需要快速且廉价地获得简短回复时,GPT Luna 可以胜任,本次运行中每条回复约 $0.0001。当写作需要真正的推理、长上下文或代码时,升级到前沿模型。LLM 也是应对 Jev 某些局限性的答案,例如输入包含图像、音频或 PDF,或输出是文档、差异或计划。
用 Jev 路由,用代码计算,用 LLM 写作
从宏观来看,第一个模式很简单。Jev 进行分诊。代码检查它有多确定。只有当交付物是散文时,代码才调用 LLM。
处理程序将升级分数在 0.5 或以上的任何内容发送给人工。它还会将意图置信度低于 0.8 的任何内容发送给人工。否则,它从订单系统回答订单状态,无需模型,并且仅对需要散文的意图向 LLM 请求草稿。
// handle.ts
import { triage, type Intent, type Triage } from './triage';
import { draftReply } from './draft-reply';
// These thresholds are application policy. Tune them on your own labeled tickets.
const ROUTE_CONFIDENCE = 0.8;
const ESCALATE_THRESHOLD = 0.5;
type ReplyIntent = Exclude<Intent, 'order_status'>;
type Route =
| { kind: 'human'; reason: string }
| { kind: 'deterministic'; intent: 'order_status' }
| { kind: 'reply'; intent: ReplyIntent };
function chooseRoute(t: Triage): Route {
if (t.escalateProbability >= ESCALATE_THRESHOLD) {
return { kind: 'human', reason: `escalation probability ${t.escalateProbability}` };
}
if (t.confidence < ROUTE_CONFIDENCE) {
return { kind: 'human', reason: `intent confidence ${t.confidence}` };
}
if (t.intent === 'order_status') {
return { kind: 'deterministic', intent: t.intent };
}
return { kind: 'reply', intent: t.intent };
}
// Stand-in for your order system. Exact data never goes through a model.
function lookupOrder(ticket: string): string {
const id = ticket.match(/\b\d{5}\b/)?.[0];
if (id === undefined) {
return 'Please reply with your five-digit order number and we will check the shipment.';
}
return `Order ${id} shipped 2026-09-17 via UPS, tracking 1Z999AA10123456784, estimated delivery 2026-09-21.`;
}
export const FACTS: Record<ReplyIntent, string> = {
return_refund:
'Returns are accepted within 30 days of delivery for regular items. Exchanges for another size follow the same 30-day window. Final sale items cannot be returned or exchanged. Prepaid labels are emailed within one business day.',
billing_dispute: 'A billing specialist will review the charge within one business day.',
product_question: 'Product specifications are listed on each product page.',
account_access: 'Password resets are available from the sign-in page.',
};
export type Handled = {
triage: Triage;
route: Route;
text: string;
costUsd: number; // Jev call plus the LLM call, if one happened
};
// Routes a ticket. A 'reply' route carries an unverified LLM draft; send.ts checks it before anything goes out.
export async function handle(ticket: string): Promise<Handled> {
const t = await triage(ticket);
const route = chooseRoute(t);
switch (route.kind) {
case 'human':
return { triage: t, route, text: `queued for a person: ${route.reason}`, costUsd: t.costUsd };
case 'deterministic':
return { triage: t, route, text: lookupOrder(ticket), costUsd: t.costUsd };
case 'reply': {
const reply = await draftReply(ticket, FACTS[route.intent]);
return { triage: t, route, text: reply.text, costUsd: t.costUsd + reply.costUsd };
}
default:
return route satisfies never;
}
}一个简短的运行器会在每个结果旁边打印 Jev 信号。
// run.ts
import { handle } from './handle';
const TICKETS = [
"Hi, I ordered a pair of running shoes on the 3rd and the tracking page hasn't updated in five days. Where is my package? Order 84721.",
'The jacket I got is too small. Can I swap it for a large? I got it last Tuesday.',
"There's an unauthorized $450 charge from your company on my card. I've already called my bank to dispute it and I'm reporting this as fraud.",
"I paid with a gift card and a credit card, but the refund only went to the credit card. Where's the gift card balance?",
];
for (const ticket of TICKETS) {
const r = await handle(ticket);
const { intent, confidence, escalateProbability } = r.triage;
console.log(`"${ticket}"`);
console.log(` intent=${intent} confidence=${confidence} escalate=${escalateProbability} route=${r.route.kind} cost=$${r.costUsd.toFixed(6)}`);
console.log(` ${r.text}\n`);
}以下是四个示例工单的输出。
"Hi, I ordered a pair of running shoes on the 3rd and the tracking page hasn't updated in five days. Where is my package? Order 84721."
intent=order_status confidence=1 escalate=0.02 route=deterministic cost=$0.000026
Order 84721 shipped 2026-09-17 via UPS, tracking 1Z999AA10123456784, estimated delivery 2026-09-21.
"The jacket I got is too small. Can I swap it for a large? I got it last Tuesday."
intent=return_refund confidence=1 escalate=0.01 route=reply cost=$0.000118
Yes, you can exchange the jacket for a large if it’s a regular item and within 30 days of delivery. If it’s a final sale item, it can’t be exchanged; for eligible exchanges, a prepaid return label is emailed within one business day.
"There's an unauthorized $450 charge from your company on my card. I've already called my bank to dispute it and I'm reporting this as fraud."
intent=billing_dispute confidence=1 escalate=0.96 route=human cost=$0.000025
queued for a person: escalation probability 0.96
"I paid with a gift card and a credit card, but the refund only went to the credit card. Where's the gift card balance?"
intent=return_refund confidence=0.54 escalate=0.02 route=human cost=$0.000025
queued for a person: intent confidence 0.54四个中有一个接触了 LLM。订单状态查询是精确且免费的。欺诈报告从未到达一个能承诺任何事情的模型。因为拆分退款工单的置信度——即 Jev 在任务 A 中归档错误的那个——低于 0.8,所以它被转给了人工。最后,夹克回复在这里仍然是一份未经检查的草稿。检查它是第二种模式。
在发布前用 Jev 验证 LLM 输出
第二种模式则相反。LLM 起草。然后 Jev 在发布前根据政策检查草稿。这能捕捉到友好的 AI 回复承诺了违反政策的事情。
// verify-draft.ts
import { OpenRouter } from '@openrouter/sdk';
const openrouter = new OpenRouter({ apiKey: process.env.OPENROUTER_API_KEY });
export const POLICY =
'Returns are accepted within 30 days of delivery for regular items. Final sale items cannot be returned. Refunds go back to the original payment method. Prepaid return labels are emailed within one business day of approval.';
type Label = 'supported' | 'unsupported' | 'declined';
export type Verdict = {
label: Label;
confidence: number;
probabilities: Record<string, number>;
costUsd: number;
};
function isLabel(value: string): value is Label {
return value === 'supported' || value === 'unsupported' || value === 'declined';
}
export async function verifyDraft(policy: string, question: string, draft: string): Promise<Verdict> {
const result = await openrouter.alpha.decisions.create({
decisionsRequest: {
model: 'typesafe/jev-1.13',
state: { policy, customer_question: question, draft_reply: draft },
questions: {
support: {
type: 'choice',
instructions: 'How does draft_reply relate to policy and customer_question?',
criteria: {
supported:
'The draft answers customer_question, and every fact, number, timeframe, and promise in it appears in policy.',
unsupported:
'The draft states a fact, number, timeframe, or promise that policy does not contain or contradicts, or it answers a different question.',
declined:
'The draft says policy does not cover the question and adds no facts of its own beyond what policy states.',
},
},
},
},
});
const support = result.answers.support;
if (support?.type !== 'choice' || !isLabel(support.choice)) {
throw new Error('Unexpected answer');
}
const cost = result.usage.cost;
if (cost === undefined) {
throw new Error('Response did not include usage.cost');
}
return {
label: support.choice,
confidence: support.confidence ?? 0,
probabilities: support.probabilities ?? {},
costUsd: cost,
};
}Jev 获取政策、客户问题和草稿,然后回答一个选择题,判断该草稿是受支持、不受支持还是被拒绝。我们让四份草稿通过了它。第一份来自 GPT Luna,另外三份是手写的,每份都针对特定分支编写。
Q: "How long until my refund shows up?"
Draft (GPT Luna): "Refunds are sent back to the original payment method, but we don't have a specific timeline for when they will appear. If your return is approved, a prepaid return label will be emailed within one business day."
supported confidence=0.09 { supported: 0.39, unsupported: 0.32, declined: 0.29 }
Q: "How long until my refund shows up?"
Draft (fabricated): "Refunds are processed within 5 to 7 business days after we receive the item."
unsupported confidence=1.00 { unsupported: 1, supported: 0, declined: 0 }
Q: "Where does my refund go?"
Draft (grounded): "Refunds go back to the original payment method, so it will return to however you paid for the order."
supported confidence=1.00 { supported: 1, unsupported: 0, declined: 0 }
Q: "How long until my refund shows up?"
Draft (grounded, wrong question): "Refunds go back to the original payment method, so it will return to however you paid for the order."
unsupported confidence=0.48 { unsupported: 0.65, supported: 0.14, declined: 0.21 }捏造的时间线以 1.00 的置信度返回不受支持,如果该回复发出,客户会在第八天回信询问退款在哪里。
Luna 的草稿,这才是有趣的,因为它陈述的每个事实都符合政策。但它半拒绝、半回答该问题,并添加了没人要求的细节。Jev 将概率分成三部分,0.8 的阈值将其发送给代表快速查看。
最后两行使用相同的基于事实的草稿,但两个不同的问题。注意,当草稿回答客户所问的问题时,它以 1.00 的置信度受支持。当它回答不同的问题时,它不受支持。这正是你希望从验证器得到的结果,并且它清楚地说明了为什么政策文本应使用与草稿相同的措辞。
我们首先尝试了两个是或否的问题,结果混乱,因为一份正确指出政策不涵盖此事的草稿既基于事实又不是答案。互斥的结果促使我们改为询问一个选择题。
将其全部连接起来
还有一个文件将路由优先管道连接到验证器,这样 LLM 写的任何内容都不会在没有受支持裁决的情况下发出。
// send.ts
import { FACTS, handle, type Handled } from './handle';
import { verifyDraft, type Verdict } from './verify-draft';
const SEND_CONFIDENCE = 0.8;
type Outcome = Handled & { disposition: 'sent' | 'human_review'; verdict?: Verdict };
export async function process(ticket: string): Promise<Outcome> {
const handled = await handle(ticket);
switch (handled.route.kind) {
case 'human':
return { ...handled, disposition: 'human_review' };
case 'deterministic':
return { ...handled, disposition: 'sent' };
case 'reply': {
// Verify against the same facts the draft was written from.
const verdict = await verifyDraft(FACTS[handled.route.intent], ticket, handled.text);
const ok = verdict.label === 'supported' && verdict.confidence >= SEND_CONFIDENCE;
return { ...handled, verdict, costUsd: handled.costUsd + verdict.costUsd, disposition: ok ? 'sent' : 'human_review' };
}
default:
return handled.route satisfies never;
}
}如果验证器退回一个单独看似乎没问题的回复,请对照事实阅读该回复。事实通常遗漏了客户询问的某些内容。去查看完整的请求路径。
ticket text
|
v
[Jev] intent (choice) + escalate (noul) .......... 1 request, ~200 ms, ~$0.000025
|
v
[code] escalate >= 0.5 or confidence < 0.8 ? --> human queue
|
v
[code] intent == order_status ? --> database lookup, exact reply, no model
|
v
[LLM] draft reply from ticket + facts ............ ~1 s, ~$0.0001 (GPT Luna)
|
v
[Jev] supported / unsupported / declined ......... 1 request, ~200 ms, ~$0.000025
|
v
[code] supported and confidence >= 0.8 ? --> send
otherwise --> human review完全走完的工单需要两次 Jev 调用和一次 LLM 调用。总成本约为 $0.00015,总时间为 1.5 秒。两次 Jev 调用使得可以使用廉价模型进行写作,同时避免需要人工阅读模型的每一条回复。
从你自己的数据中选取阈值
无论阈值出现在你代码的何处,请记住这个数字并非凭空而来。它来自你。这是你的政策选择。Jev 的概率是校准的,这基本上意味着取大量预测并查看它们正确的频率。所以如果我们看到 0.8,那么这些预测中大约 80% 是正确的。但并非每一个都正确。这是一个平均值。所以将本文中看到的 0.8 和 0.5 视为我们的数字,而非你的。
如果你想找到自己的阈值,几百个标注案例就够了。把它们全部跑一遍模型。对每个截断值,统计它答对了多少。在图上画出准确率与截断值的关系。然后把截断值设在自动路径与人工团队会返回的结果一致的那个点上。编辑标准之后,重新检查这一点。别忘了。改写一个选项可能会以有些出人意料的方式改变其他选项的概率。
从一个 switch 语句开始
如果你的代码库里有一个 LLM 调用以 JSON.parse 结尾,然后接一个 switch,那就是你的第一个 Jev 问题。把它换进来,把 LLM 留给需要生成文字的分支,然后对两者都进行测量。
- Decisions API 参考,了解完整的请求和响应模式
- OpenRouter 上的 Jev,了解当前定价和限制
- TypeSafe 原语、state 和 confidence 文档,帮助设计问题
- Jev 验证的级联 cookbook,用于先便宜后前沿的 LLM 级联,其中 Jev 担任裁判
- 用 Jev 门控工具调用 cookbook,将同样的思路应用于智能体工具调用
- 什么是 Jev?,了解模型本身、三种问题类型,以及如何解读概率
- 如何使用 Jev,了解一个完整的 TypeScript 审核流水线,其中 Jev 回答问题,你的代码掌握阈值
- Jev 在分类上是否和前沿模型一样准确?,了解在 Banking77 上的正面比较,包括弥合大部分准确率差距的级联
- Jev 文档中心,查看 OpenRouter 上的所有 Jev 指南和 cookbook
常见问题
什么是 Jev,它与 LLM 有何不同?
Jev 是 TypeSafe 的 System One 决策模型。它接收状态加上一个类型化问题——要么是从固定集合中选择,要么是是/否命题,要么是在你描述的层级上打分——并返回概率而不是文本。大型语言模型生成开放式文本,然后你必须解析并信任这些文本。
在分类任务上,Jev 比小型 LLM 更便宜吗?
在 2026 年 9 月 19 日通过 OpenRouter 运行的 60 个支持工单上,Jev 每 1,000 个工单花费 $0.0248,GPT Luna 每 1,000 个工单花费 $0.0921,Claude Opus 每 1,000 个工单花费 $2.88,而意图准确率 Jev 为 59/60,GPT Luna 为 59/60,Claude Opus 为 60/60。Jev 1.13 每百万输入 token 花费 $0.042,输出免费。定价随时可能变化,所以做预算前一定要查看模型页面。
如何将 Jev 和 LLM 一起使用?
两种模式。第一,先路由:Jev 对请求进行分类并返回置信度。你的代码用你选择的阈值检查该置信度。只有需要生成文字的请求才会发给 LLM。第二,后验证:LLM 先起草。Jev 检查草稿中的每一项声明是否都有你的政策支持。只有通过后,回复才会发出。
Jev 使用哪个 OpenRouter 端点?
Jev 使用 Decisions API 端点,POST https://openrouter.ai/api/alpha/decisions,模型 ID 为 typesafe/jev-1.13,或在 TypeScript SDK 中为 openrouter.alpha.decisions.create()。传统 LLM 使用 POST https://openrouter.ai/api/v1/chat/completions,或 openrouter.chat.send()。
我应该为 Jev 使用什么置信度阈值?
没有适用于所有情况的通用数字。Jev 的概率是在许多预测上校准的,所以它们在总体上成立。任何单个预测仍然可能是错的。因此,请根据你自己的标注数据和犯错的代价来选择截断值。例如,这是我们的路由规则:意图置信度低于 0.8,或预测升级达到或超过 0.5,就转给人工。
来源:OpenRouter:Announcements(RSS) · openrouter.ai