arXiv:cs.LG· Brent Motmans, Digvijay Ghogare, Thijs G. I. van Wijk, Joren Van Herck, Saba Heidarian, Pieter De Meyer, Berend Smit, An Hardy, Danny E. P. Vanpoucke·· 3 小时前
基于小数据集的机器学习预测无机合成:以流体动力学直径可控的 Cu 纳米颗粒为例
Predictive Inorganic Synthesis based on Machine Learning using Small Data sets: a case study of Hydrodynamic Diameter-controlled Cu Nanoparticles
AI 导读
研究用仅 25 组合成的数据集,以机器学习预测 Cu 纳米颗粒的 DLS 流体动力学直径。集成回归模型在同一分布外验证下 R2 达 0.74,高于 DoE 模型的 0.60,且保留完整合成参数空间;随机森林与 LLM 分类模型仅表现出中等性能,说明小数据集不足以发挥复杂 LLM 的能力。
正文
Abstract:Cu NPs have a broad applicability, yet their synthesis is sensitive to subtle changes in reaction parameters. This sensitivity, combined with the time- and resource-intensive nature of experimental optimization, poses a major challenge in achieving reproducible and size-controlled synthesis. While ML shows promise in materials research, its application is often limited by scarcity of large high-quality experimental data sets. This study explores ML to predict the DLS-derived hydrodynamic diameter of Cu NPs using a small data set of 25 syntheses. Latin Hypercube Sampling is used to efficiently cover the parameter space while creating the experimental data set. Ensemble regression models successfully predict hydrodynamic diameters with good predictive performance given the limited dataset. Since quantitative regression requires a unique DLS-derived hydrodynamic diameter, the regression model is restricted to mono-modal DLS distributions, while a complementary classification model identifies synthesis conditions for which quantitative prediction is applicable. Using equivalent out-of-sample validation, the ML and DoE models showed comparable generalization. The final ensemble model achieved an R2=0.74 compared to 0.60 for the DoE model, while retaining the complete synthesis parameter space, making it better suited for synthesis guidance. Additionally, classification models using both random forests and LLMs are evaluated to distinguish between large and small particles. These classification models exhibited only modest predictive performance, indicating that this small dataset is insufficient to fully exploit the capabilities of complex LLMs. Overall, this study demonstrates that carefully curated small data sets, paired with robust classical ML, can effectively support the synthesis of Cu NPs and highlights that for lab-scale studies, complex models like LLMs may offer limited benefits.
| Comments: | 15 pages (+16 pages SI), 5 figures (+16 SI), 4 tables (+11 SI) |
| Subjects: | Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG) |
| Cite as: | arXiv:2512.16545 [cond-mat.mtrl-sci] |
| (or arXiv:2512.16545v3 [cond-mat.mtrl-sci] for this version) | |
| https://doi.org/10.48550/arXiv.2512.16545 arXiv-issued DOI via DataCite |
Submission history
From: Danny E. P. Vanpoucke Prof. Dr. Dr. [view email]
[v1]
Thu, 18 Dec 2025 13:53:08 UTC (11,085 KB)
[v2]
Mon, 9 Feb 2026 17:03:40 UTC (11,164 KB)
[v3]
Thu, 8 Oct 2026 15:57:44 UTC (12,782 KB)
来源:arXiv:cs.LG · arxiv.org