arXiv:cs.AI· Weiyi Wang, Xinchi Chen, Jingjing Gong, Xuanjing Huang, Xipeng Qiu·· 4 小时前AI 评分43
AstroAgentBench:面向太空任务规划任务的智能体规划评测基准
AstroAgentBench: Evaluating Agentic Planning on Space Mission Planning Tasks
AI 导读
AstroAgentBench 是一个覆盖调度、观测规划、星座设计与中继支持四大领域、共七个任务族的可执行太空任务规划基准,由外部验证器对智能体提交的规划产物进行 schema、时序、几何、资源与任务价值检查。
正文
Abstract:Recent LLM-for-Space systems address mission planning, scheduling, operations support, simulator control, and autonomy, but their evaluations use different task contracts, control settings, simulators, and success criteria. We introduce AstroAgentBench, a seven-family benchmark for executable space mission planning in the domains of scheduling, observation planning, constellation design, and relay support. For each case, an agent submits a planning artifact that is checked by an external verifier for schema, timing, geometry, resources, and mission value. Results report validity and normalized scores, with comparisons to task-specific solver references. Across five LLM agent systems and 35 held-out cases, the strongest systems approach or exceed solver-reference scores on several families, while weaker systems often fail to produce high-value valid plans and even strong systems lose quality on geometric, product-level, or design-heavy tasks. Trace analyses separate two failure points: task-contract misformulation and weak solution construction. Successful runs instead calibrate agent-written implementations against verifier feedback and adapt search to case-specific structure. Ablations show that procedure injection and memory accumulation help selectively, when they supply the missing formulation, calibration, or search support.
| Comments: | 35 pages, 5 figures. AACL-IJCNLP 2026. Benchmark renamed from AstroReason-Bench to AstroAgentBench; supersedes v1 with the full five-system evaluation. Code: this https URL Data: this https URL |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2601.11354 [cs.AI] |
| (or arXiv:2601.11354v3 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2601.11354 arXiv-issued DOI via DataCite |
Submission history
From: Weiyi Wang [view email]
[v1]
Fri, 16 Jan 2026 15:02:41 UTC (9,247 KB)
[v2]
Thu, 1 Oct 2026 15:06:18 UTC (259 KB)
[v3]
Fri, 2 Oct 2026 06:51:10 UTC (259 KB)
来源:arXiv:cs.AI · arxiv.org