跳到正文
arXiv:cs.AI· Lijie Ding, Changwoo Do·· 5 小时前AI 评分44

NeutronGym:面向 LLM 智能体的物理评分式中子仪器设计环境

NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents

AI 导读

研究者推出 NeutronGym,一个可执行的中子仪器设计环境:智能体通过验证工具搭建仪器,由 McStas 进行光线追踪,并按语法、运行时、结构和科学四个层级评分,全程不用 LLM 评判。

正文

View PDF HTML (experimental)

Abstract:Designing a scientific instrument tests whether language-model agents can do physics rather than recall it, provided the grading cannot be argued with. We introduce NeutronGym, to our knowledge the first executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge. Procedural families supply unlimited instances of a fixed layout whose design parameters the agent must set, with held-out parameter regimes; a curated slice, McStasBench, adds 16 tasks from published instruments behind memorization probes and a sandbox. Seven models reproduce at most 7 of the 16, none retrieves a reference, and none meets an improvement target. The environment also trains. Reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B, and the recipe holds, at one seed each, on three further gated families. The analysis says what that gain is. Without the ladder's partial credit it collapses by 60 points. From reward alone the trained model reaches what a classical optimizer reaches, at the agent's simulation budget, only when handed the closed-form physics (77% against 81%, a gap that does not separate at this size), while frontier models still solve 98-99%. Getting a trustworthy result meant failing four task designs that no-model baselines could solve, and we release the probes that found them.
Subjects: Artificial Intelligence (cs.AI); Instrumentation and Detectors (physics.ins-det)
Cite as: arXiv:2610.03631 [cs.AI]
  (or arXiv:2610.03631v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.03631

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Lijie Ding [view email]
[v1] Fri, 2 Oct 2026 17:23:52 UTC (177 KB)

来源:arXiv:cs.AI · arxiv.org