跳到正文
HuggingFace Daily Papers·· 1 天前AI 评分43

HERMES:用模块化可执行 Dev-Primitives 做软件工程 Harness 设计

Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

AI 导读

研究者提出 Dev-Primitives(开发原语),将代码库构件与常驻 LLM 配对,使其获得基于自身实现与依赖的智能体原生接口,并据此构建软件工程框架 HERMES。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, dependencies, and runtime behavior, leading to increasingly long interaction histories, context explosion, and semantic drift. Large repositories further complicate the identification of task-relevant components. To address these challenges, we introduce \textbf{Dev-Primitives} (\emph{Development Primitives}), a modular and executable abstraction that transforms repository components from passive software artifacts into active participants in software engineering. Each Dev-Primitive pairs a repository artifact with a resident LLM, which gives the artifact an agent-native interface grounded in its own implementation and dependencies, enabling natural-language reasoning, inter-component communication, and localized self-modification. Building on Dev-Primitives, we propose \textbf{HERMES}, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised. Extensive experiments on four software engineering benchmarks demonstrate that HERMES outperforms matched baseline harnesses by 12.4\% on average. Moreover, when paired with strong activation and diagnosis models, HERMES, even with Qwen3-8B Dev-Primitives, remains within 4.5\% of the homogeneous GPT-5.6 Sol configuration across all four benchmarks, while reducing inference cost by 26.2\% on Terminal-Bench 4.0, highlighting the importance of harness design in software engineering agents.
Comments: 30 pages
Subjects: Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multiagent Systems (cs.MA)
Cite as: arXiv:2610.07832 [cs.SE]
  (or arXiv:2610.07832v1 [cs.SE] for this version)
  https://doi.org/10.48550/arXiv.2610.07832

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Haibo Jin [view email]
[v1] Tue, 6 Oct 2026 06:32:11 UTC (526 KB)

来源:HuggingFace Daily Papers · arxiv.org