跳到正文
arXiv:cs.AI· Peter Kulits, Yiqing Xu, R. Kenny Jones, Cordelia Schmid, Jiajun Wu·· 3 小时前

BrickBench:评估智能体乐高积木设计

BrickBench: Evaluating Agentic Brick Design

AI 导读

BrickBench 是一个面向智能体文本条件 LEGO 积木设计的基准,智能体需从离散零件库中选件并同时推理局部与全局约束,产出满足语义、设计标准且可实体搭建的组装方案。该基准在三种规模与零件可用性设置下评测有效性、对齐度和设计质量,并提供 BrickAgent 环境供编码智能体构建、检查和验证设计。结果显示领先智能体基本满足可验证的物理与语义要求,但仍不及人类设计水平。

正文

View PDF HTML (experimental)

Abstract:We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at this http URL
Comments: Project page: this http URL
Subjects: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
Cite as: arXiv:2610.12452 [cs.AI]
  (or arXiv:2610.12452v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.12452

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Peter Kulits [view email]
[v1] Thu, 8 Oct 2026 17:58:56 UTC (6,204 KB)

来源:arXiv:cs.AI · arxiv.org