arXiv:cs.CL· Kyaw Hpone Myint, Zhe Wu, Alexandre G. R. Day, Giri Iyengar·· 4 小时前
通过合成模型生成实现近最优可解释模型的可扩展元学习
Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations
AI 导读
研究提出一种生成合成预训练数据的高效可扩展方法,用于实现对决策树的元学习。该方法通过合成采样近最优决策树构建大规模逼真数据集,配合 MetaTree transformer 架构,性能与使用真实数据或计算昂贵的最优决策树预训练相当,同时显著降低计算成本并提升数据生成的灵活性。
正文
Abstract:Decision trees are widely used in high-stakes fields like finance and healthcare due to their interpretability. This work introduces an efficient, scalable method for generating synthetic pre-training data to enable meta-learning of decision trees. Our approach samples near-optimal decision trees synthetically, creating large-scale, realistic datasets. Using the MetaTree transformer architecture, we demonstrate that this method achieves performance comparable to pre-training on real-world data or with computationally expensive optimal decision trees. This strategy significantly reduces computational costs, enhances data generation flexibility, and paves the way for scalable and efficient meta-learning of interpretable decision tree models.
| Comments: | 9 pages, 3 figures, Neurips 2025 GenAI in Finance Workshop |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML) |
| Cite as: | arXiv:2511.04000 [cs.LG] |
| (or arXiv:2511.04000v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2511.04000 arXiv-issued DOI via DataCite |
Submission history
From: Kyaw Hpone Myint [view email]
[v1]
Thu, 6 Nov 2025 02:50:23 UTC (668 KB)
[v2]
Wed, 7 Oct 2026 21:18:10 UTC (668 KB)
来源:arXiv:cs.CL · arxiv.org