arXiv:cs.AI· Jorge Garc\'ia-Carrasco, Javier Sanchis, Alejandro Reina-Reina, Alejandro Mat\'e, Juan Trujillo·· 3 小时前
本地 LLM 智能体能否胜任数据工程?一项可复现性实证研究
Evaluating Local Language Model Agents for Reproducible Data Engineering: An Empirical Software Engineering Study of Mobility Workflows
AI 导读
一项 arXiv 预印本研究用 15 个移动出行工作流任务评测本地开源权重 LLM 智能体,在消费级 GPU 上对 10 种配置各重复五次,共 1500 次打分尝试。超过 20 亿参数的模型中,闭环工作区条件将通过率较一次性生成提升 26.7 至 52.0 个百分点,最佳配置达到 85.3% 产物级成功率,量化后的 90 亿参数模型以约 6.5 GB 显存占用达到 69.3%。
正文
Abstract:Context: Large language model (LLM) agents are increasingly used as software and data-engineering assistants, yet evidence about locally deployable open-weight agents remains limited. Existing evaluations often emphasize textual responses or isolated code generation rather than the validity of complete engineering artifacts.
Objectives: We evaluate whether local LLM agents can produce correct and reproducible data-engineering artifacts, quantify the effect of a closed-loop workspace condition, and examine trade-offs in model scale, architecture, quantization, runtime, tool use, and failure.
Methods: We introduce a benchmark of fifteen mobility-workflow tasks covering data discovery, connectors, transport-feed processing, semantic enrichment, feature engineering, validation, visualization, and reporting. Deterministic checkers assess generated scripts, tables, structured files, figures, and reports. Ten local configurations are evaluated in one-shot and closed-loop conditions, with five repetitions per model, mode, and task, yielding 1,500 scored attempts on a consumer-grade GPU.
Results: Among models larger than two billion parameters, the workspace condition increases pass rates by 26.7-52.0 percentage points over one-shot generation. The strongest configuration reaches 85.3% artifact-level success, and a quantized 9-billion-parameter model reaches 69.3% with an approximately 6.5 GB memory footprint. Gains are largest when intermediate artifacts expose errors the agent can inspect and repair.
Conclusion: Local open-weight agents can support a meaningful subset of software-intensive data-engineering work, but reliability depends on model capability, task verifiability, and deterministic validation. The benchmark provides a reproducible method for evaluating complete agent configurations before adoption in engineering workflows.
| Comments: | Preprint, under review. 26 pages, 5 figures, 5 tables. Dataset: this https URL |
| Subjects: | Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Databases (cs.DB); Machine Learning (cs.LG) |
| ACM classes: | D.2.8; I.2.7 |
| Cite as: | arXiv:2610.11482 [cs.SE] |
| (or arXiv:2610.11482v1 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11482 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jorge García-Carrasco [view email]
[v1]
Thu, 8 Oct 2026 08:26:46 UTC (63 KB)
来源:arXiv:cs.AI · arxiv.org