跳到正文
arXiv:cs.AI· Dunzhi Zhou, Chengyu Zhu, Gang Qiu, Caiwen Ding·· 6 小时前AI 评分41

2D-FET-Bench V2:语言模型智能体能否完成二维材料 FET 版图设计

2d-fet-bench: from spatial reasoning to fet design on flakes

AI 导读

研究者发布 2D-FET-Bench V2,一个由显微图像提取的二维材料薄片轮廓构建的 128 项 FET 版图设计基准,每项任务提供器件文本规格与轮廓坐标,智能体生成多边形与路径操作并渲染为 GDSII,由确定性验证器检查几何与结构要求。

正文

View PDF HTML (experimental)

Abstract:Field-effect transistor (FET) layouts on exfoliated two-dimensional flakes are typically drawn by hand for each flake, placing contacts and gates to match its position and outline in optical micrographs. To our knowledge, no executable benchmark tests whether language-model agents can perform this flake-specific construction reliably. We introduce 2D-FET-Bench V2, a benchmark of 128 layout tasks built from microscopy-derived flake contours, including hole-containing flakes and multi-flake tasks. Each task supplies a textual device specification and contour coordinates. An agent generates typed polygon and path operations rendered to GDSII. A deterministic verifier checks geometric and structural requirements, and a separate integrity check verifies that the supplied contours remain unchanged. Scripted reference layouts pass all 128 tasks, showing that every task is solvable. We evaluate six models and seven workflow and scaffold variants of GPT5.6-Luna, with five attempts per task. The best-performing configuration in the six-model panel, GPT5.6-Luna with ReAct-3, passes 62.3% of attempts and solves 80.5% of tasks at least once (coverage) and 43.8% in all five attempts (consistency). ReAct-3 exceeds the one-pass Plan-and-Execute by 27.0 pass@1 points at 2.46 times the tokens. An expert audit of one sampled verifier-passing layout per covered task, across five ReAct-3 configurations, accepts 56.4% to 63.5% of them. The benchmark evaluates geometric and structural FET layout construction.
Subjects: Artificial Intelligence (cs.AI); Mesoscale and Nanoscale Physics (cond-mat.mes-hall); Materials Science (cond-mat.mtrl-sci)
Cite as: arXiv:2610.07423 [cs.AI]
  (or arXiv:2610.07423v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07423

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Dunzhi Zhou [view email]
[v1] Mon, 5 Oct 2026 21:30:25 UTC (386 KB)

来源:arXiv:cs.AI · arxiv.org