arXiv:cs.AI(全量分类)· Timothy Kassis·· 5 小时前AI 评分52
Scientific Agents 研究:职业专属系统提示词未带来准确率提升且成本更高
Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks
AI 导读
arXiv 论文 2610.00084 评估了开源语料 Scientific Agents 中 503 个职业专属系统提示词档案,在 Gemini 3.8 Flash 与 Pi agent harness 下对比四种对照条件。
正文
Abstract:Detailed profession-specific system prompts raise token use and estimated cost per response without a consistent accuracy gain. We evaluate Scientific Agents, an open-source corpus of 503 profession-specific this http URL profiles, with Gemini 3.8 Flash via OpenRouter in the Pi agent harness. We compare matched profiles with four controls: a minimal baseline ("You are a helpful assistant"), the profile's opening role sentence, a generic scientific rigor guide, and a profile from an unrelated domain. Across nine text-based science benchmarks (4,531 sampled questions, 100 matched profiles), 4,488 items completed all five conditions after API-error retries, scored with automated, rule-based grading. The average profile-baseline accuracy difference is -0.6 percentage points (95% bootstrap interval [-1.5, +0.2] across fixed tasks), and no benchmark shows a statistically clear improvement. Matched profiles produced 1.5-2.3 times as many output tokens and cost 2.2-4.5 times more per successful call. On 60 tool-using BioMysteryBench bioinformatics problems (three runs each for baseline and profile), mean solve rates were 46.7% with the profile and 56.7% at baseline, a difference of -10.0 percentage points (95% interval [-16.7, -3.3]) driven by more frequent token- and time-limit stops under the profile. Longer prompts had one unexpected operational advantage: on SuperGPQA, frequent provider API drops left the short baseline with a correct first-pass answer on only 54.0% of items, against 71.6% with the profile. Generic and mismatched prompts were about as reliable, so this gain comes from prompt length or formatting rather than domain expertise. For the tested model and tasks, loading full profession profiles by default does not improve accuracy and costs considerably more; whether selective retrieval of profile sections or open-ended scientific tasks would change this remains to be tested.
| Comments: | 46 pages (11 pages main text, references, 33-page appendix); 10 figures, 29 tables. Evaluated corpus: this https URL (commit 48dedd2); evaluation code and item-level records are not released |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.00084 [cs.AI] |
| (or arXiv:2610.00084v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00084 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Timothy Kassis [view email]
[v1]
Mon, 7 Sep 2026 00:21:09 UTC (777 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org