arXiv:cs.CL· Steven Denney, Matthew DiGiuseppe·· 10 小时前AI 评分40
JEV 对比 LLM:在七项政治学复现任务上的准确率、成本与校准
JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications
AI 导读
研究对比 JEV 与 GPT-6 Luna、Qwen3.8-27B 及人类编码者在七项政治学复现任务上的表现,发现 JEV 准确率与两款 LLM 相当或接近,但在 OpenAI 批量价格下相对 GPT-6 Luna 并无成本优势。
正文
Abstract:Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe markets as "System One" models, instead returns decisions and probability distributions across a user-supplied fixed answer set. A commercial model, JEV, is advertised as having a dramatic cost and speed advantage over traditional LLMs along with better calibrated decisions. As such, it might be useful for social scientists looking to quickly and cost-effectively annotate or scale large corpora of text and have a reliable indicator of a classifier's uncertainty. Yet, the accuracy of these claims and the broader model accuracy in social science text-based tasks are not yet established. In this paper, we do just that and hope to establish the suitability of JEV for social science tasks. We compare JEV with LLMs and human coders from published research, and with a current mid-tier commercial LLM (GPT-6 Luna) and an open-weight alternative (Qwen3.8-27B). We find that JEV matches, or comes close to, the capabilities of both LLMs in a variety of tasks. However, we find no cost advantage over GPT-6 Luna at OpenAI's batch prices. Further, we find that, when each question is asked once, JEV's probabilities are better calibrated than GPT-6 Luna's token probabilities, but not consistently better than Qwen3.8-27B's. We conclude that unless researchers have a need for speed, JEV's only obvious advantage is ease of parsing the underlying choice probabilities.
| Comments: | 71 pages, 2 figures, 14 tables (including appendices). v2: corrected author order in metadata |
| Subjects: | Computation and Language (cs.CL); Computers and Society (cs.CY) |
| Cite as: | arXiv:2610.06625 [cs.CL] |
| (or arXiv:2610.06625v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06625 arXiv-issued DOI via DataCite |
Submission history
From: Matthew DiGiuseppe [view email]
[v1]
Mon, 5 Oct 2026 16:22:03 UTC (145 KB)
[v2]
Tue, 6 Oct 2026 05:45:51 UTC (145 KB)
来源:arXiv:cs.CL · arxiv.org