跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Anthony Lavertu, Jacob Cote, Sophie Gobeil, Jacques Corbeil, Isabeau Premont-Schwarz, Pascal Germain·· 5 小时前AI 评分42

ReLaG:将随机划分推广到含潜在关系数据的可扩展框架

ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations

AI 导读

ReLaG 是一个模态无关框架,通过层次潜变量过程建模样本关联性,并用近邻图与社区检测推断相关样本分组,从而生成独立的训练-测试子集。在分子与蛋白质数据集上,ReLaG 匹配现有关系感知方法的同时扩展性显著更优,支持此前不切实际的数据集规模划分。该框架已开源,可通过 pip install relag 安装。

正文

View PDF HTML (experimental)

Abstract:Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies. This leads to overly optimistic generalization estimates. Here, we introduce ReLaG, a modality-agnostic framework that models sample relatedness through a hierarchical latent-variable process and infers groups of related samples using proximity graphs and community detection to produce independent train--test subsets. Across molecular and protein datasets, ReLaG matches existing relation-aware methods while scaling substantially better, enabling splits at previously impractical dataset sizes. We further introduce a label-free procedure that adapts the splitting resolution to production data, aligning evaluation with the intended deployment setting. ReLaG's inferred groups provide a cheap estimate of effective dataset size, enabling diversity-aware dataset scaling. ReLaG is open source and can be installed with pip install relag.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.38538 [cs.LG]
  (or arXiv:2609.38538v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38538

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Anthony Lavertu [view email]
[v1] Tue, 29 Sep 2026 20:57:17 UTC (1,230 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org