跳到正文
arXiv:cs.LG· Chaithra Umesh (Institute of Computer Science, University of Rostock, Germany), Arvind Lomrore (School of Data Science, Indian Institute of Science Education and Research, Thiruvananthapuram, India), Neethu D (School of Data Science, Indian Institute of Science Education and Research, Thiruvananthapuram, India), Kristian Seegel-Schultz (Institute of Computer Science, University of Rostock, Germany), Saptarshi Bej (Institute of Computer Science, University of Rostock, Germany, School of Data Science, Indian Institute of Science Education and Research, Thiruvananthapuram, India), Olaf Wolkenhauer (Institute of Computer Science, University of Rostock, Germany, Leibniz-Institute for Food Systems Biology, Technical University of Munich, Freising, Germany, Stellenbosch Institute for Advanced Study, South Africa)·· 4 小时前AI 评分30

LDTool 与 HLDTool:从表格数据中可扩展提取并可视化多属性逻辑依赖与函数依赖

Scalable extraction and visualization of multi-attribute logical and functional dependencies in tabular data

AI 导读

研究者提出 LDTool 和 HLDTool,用于从表格数据中提取并可视化多属性逻辑依赖(LD)与函数依赖(FD)。LDTool 将依赖发现扩展到成对关系之外,HLDTool 通过超图引导的搜索空间缩减实现可扩展提取,可在数百个特征的数据集上发现依赖。在 3 个模拟和 11 个真实数据集上,LDTool 以更低运行时间复现了现有 FD 发现方法的结果。

正文

Authors:Chaithra Umesh (1), Arvind Lomrore (4), Neethu D (4), Kristian Seegel-Schultz (1), Saptarshi Bej (1 and 4), Olaf Wolkenhauer (1,2, and 3) ((1) Institute of Computer Science, University of Rostock, Germany, (2) Leibniz-Institute for Food Systems Biology, Technical University of Munich, Freising, Germany, (3) Stellenbosch Institute for Advanced Study, South Africa, (4) School of Data Science, Indian Institute of Science Education and Research, Thiruvananthapuram, India)

View PDF HTML (experimental)

Abstract:Understanding the structural relationships among attributes in tabular data is fundamental to machine learning and pattern recognition. While functional dependency (FD) discovery has been extensively studied, scalable discovery of logical dependencies (LDs), particularly as the number of attributes and dependency order increase, remains underexplored. These dependencies capture non-deterministic, condition-specific relationships among pairwise or multiple attributes. Furthermore, existing approaches do not provide a unified framework for extracting multi-attribute LDs and FDs. To address these limitations, we propose LDTool and HLDTool for extracting and visualizing multi-attribute LDs and FDs from tabular data. LDTool extends dependency discovery beyond pairwise relationships, while HLDTool enables scalable extraction through hypergraph-guided search-space reduction. Experiments on three simulated and eleven real-world datasets demonstrate that the proposed framework extracts meaningful LDs and FDs while improving scalability. LDTool recovers the same FDs as existing FD discovery methods with lower runtime in high-dimensional feature spaces, whereas HLDTool enables dependency discovery in datasets with hundreds of features. The proposed framework provides interpretable visualizations of dependency structures and supports applications in exploratory data analysis and the quantitative evaluation of synthetic tabular data.
Comments: 31 pages, 4 figures, submitted to Pattern Recognition Journal
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.08287 [cs.LG]
  (or arXiv:2610.08287v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.08287

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Chaithra Umesh [view email]
[v1] Tue, 6 Oct 2026 12:56:40 UTC (1,584 KB)

来源:arXiv:cs.LG · arxiv.org