arXiv:cs.CL· Weizhe Xu, Jialiang Fan, Mengyu Liu, Fanxin Kong·· 3 小时前
SafeInferCom:用验证器引导生成中干预,保障机器人任务规划的安全推理时计算
SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning
AI 导读
SafeInferCom 是一个形式化验证器引导的推理时计算框架,通过暴露并验证中间规划、保留有效中间计划并在生成过程中纠正错误,提升机器人任务规划的可靠性。实验显示,它在多个大推理模型和规划领域上较单次推理提升了规划成功率并加速纠错,与迭代精炼结合后还能在提高成功率的同时减少 token 消耗。该框架已在 VirtualHome 中评测,并提供了真实机械臂演示。
正文
Abstract:Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation. We develop an inference-time monitor that exposes and verifies intermediate plans without disrupting the original decoding trajectory. Building on this monitor, we propose SafeInferCom, a formal verifier-guided framework that preserves valid intermediate plans and directs error correction during generation. Experiments across multiple LRLMs and planning domains reveal reasoning-response inconsistency and limited self-correction under one-shot inference. SafeInferCom improves planning success and accelerates error correction relative to one-shot inference. When combined with iterative refinement, it further improves success while reducing token usage compared with refinement alone. We additionally evaluate SafeInferCom in VirtualHome and provide a real-world robotic-arm demonstration.
| Comments: | Video: this https URL |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Logic in Computer Science (cs.LO) |
| Cite as: | arXiv:2610.11223 [cs.RO] |
| (or arXiv:2610.11223v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11223 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Weizhe Xu [view email]
[v1]
Thu, 8 Oct 2026 04:19:47 UTC (409 KB)
来源:arXiv:cs.CL · arxiv.org