跳到正文
arXiv:cs.LG· Kautik Mandve, Dileepa Fernando·· 4 小时前AI 评分38

先探索后提交:用语言模型做测量高效的科学定律发现

Explore, Then Commit: Measurement-Efficient Scientific Law Discovery with Language Models

AI 导读

一项"先探索后提交"协议让大语言模型提出假设、程序化规划器采集测量、新提示词从固定观测中合成最终定律,在 576 次 NewtonBench 试验的 12 个物理模块上比较了 8 种配置。

正文

View PDF HTML (experimental)

Abstract:Scientific law discovery requires selecting measurements and converting evidence into a governing equation. We evaluate an explore-then-commit protocol in which a large language model proposes hypotheses, a programmatic planner gathers measurements, and a fresh prompt synthesizes the final law from fixed observations. The protocol combines structured probes, automatic numerical diagnostics, restricted measurement batches, and optional interpreter access. Across 576 NewtonBench trials, we compare eight configurations on 12 physics modules using GPT-4.1-mini and a medium-difficulty GPT-4.1 replication. On medium tasks, interpreter-enabled planners use 8.6 versus 22.5 measurements per trial for GPT-4.1-mini and 8.9 versus 43.0 for GPT-4.1. Their mean magnitude-based root-mean-squared logarithmic error falls from 2.514 to 0.202 and from 0.626 to 0.149, respectively. An additional audit retains incomplete and invalid submissions in a coverage-sensitive analysis. Observed symbolic-accuracy gains are less consistent across modules, and random acquisition is competitive with disagreement scoring. Measurement savings occur in every module, but unequal batch constraints prevent attributing them solely to acquisition quality. These results support the complete protocol as a promising measurement-efficient configuration, while leaving its causal components and generalization beyond noiseless direct-equation tasks unresolved.
Comments: 19 pages, including Supplementary Material S1; code and data included as ancillary files. Preprint
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Computational Physics (physics.comp-ph)
Cite as: arXiv:2610.07620 [cs.AI]
  (or arXiv:2610.07620v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07620

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Dileepa Fernando [view email]
[v1] Tue, 6 Oct 2026 02:11:39 UTC (3,097 KB)

来源:arXiv:cs.LG · arxiv.org