跳到正文
arXiv:cs.LG· Nicolas Lacroix, Frederic Precioso, Mireille Blay-Fornarino, Sebastien Mosser·· 4 小时前AI 评分31

用小型语言模型逆向工程机器学习流水线结构

Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures

AI 导读

一项 confirmatory 研究评估了小型语言模型(SLM)能否从源代码中提取机器学习流水线的各个阶段(如数据预处理、建模),以克服现有分类器无法应对领域多样性的局限。

正文

View PDF HTML (experimental)

Abstract:Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling...), extracting these stages from source code is key for better understanding ML practices. However, the diversity caused by the constant evolution of ML (e.g., algorithms, datasets) makes this task challenging. Existing approaches either rely on non-scalable manual labeling or on classifiers that do not properly support domain's diversity. These limitations call for more reliable solutions.
Objective: We evaluate whether Small Language Models (SLMs) can leverage their code understanding and classification abilities to address these limitations, and enhance our understanding of practices in ML.
Method: We conduct a confirmatory study based on two relevant reference works representing current limitations in the state-of-the-art. We first compare several SLMs using Cochran's Q test, then evaluate the best-performing model against reference studies via two McNemar's tests. An additional Cochran's Q test examines how taxonomy definition variations affect the SLM performance. Finally, goodness-of-fit tests compare ML practice insights from SLM classification with those from prior studies.
Results: First, we found that the taxonomy wording significantly impacts classification performance. Second, the best performing SLM yielded good results, yet, without outperforming other classifiers. Third, the three classification methods led to significantly different insights, with varying effect sizes, when exploring practices of data scientists.
Conclusions: Limitations of existing classification methods bias our understanding of ML practices. While current SLMs show promising results without prior fine-tuning, they still exhibit common limitations, in addition to inference high costs challenging their applicability in large-scale studies.
Subjects: Software Engineering (cs.SE); Machine Learning (cs.LG)
Cite as: arXiv:2610.10261 [cs.SE]
  (or arXiv:2610.10261v1 [cs.SE] for this version)
  https://doi.org/10.48550/arXiv.2610.10261

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Nicolas Lacroix [view email]
[v1] Wed, 7 Oct 2026 15:35:04 UTC (1,466 KB)

来源:arXiv:cs.LG · arxiv.org