跳到正文
arXiv:cs.AI· Marcelo Valentim Silva, Hannes Herrmann, Valerie Maxville·· 3 小时前

一个可解释的以表头为中心的大规模语义表格解释与数据质量评估框架

An Explainable Header-Centric Framework for Large-Scale Semantic Table Interpretation and Data Quality Assessment

AI 导读

研究者提出一个可解释、以表头为中心的框架,在无法获取单元格值的纯元数据场景下完成列类型标注(CTA)与数据质量评估(DQA)。

正文

View PDF HTML (experimental)

Abstract:Knowledge Graph (KG) quality depends not only on downstream graph validation, but also on the quality of tabular metadata used before integration. In metadata-only Semantic Table Interpretation (STI), where cell values are unavailable, noisy, or unsuitable, column headers become a critical source of semantic evidence for traceable KG preparation.
We present an explainable, header-centric framework for metadata-only Column Type Annotation (CTA) and Data Quality Assessment (DQA). The framework maps headers to 39 interpretable FinalFormat types using curated lexical resources and preserves token-level traceability through SourceKeywords. Each assigned type activates validation rules based on a taxonomy of Data Quality Issues (DQIs), producing detections such as missing data, duplicates, domain violations, wrong data type, and temporal mismatch. These detections are aggregated into HeadersIQ, a lightweight, unweighted data source-level quality metric.
The framework was evaluated across heterogeneous benchmarks, including UCI, Prague, Kaggle, VizNet/Sato, SOTAB, T2Dv2, and the SemTab 2024 Metadata-to-KG track, comprising around 120,000 header columns. The results show broad practical coverage across noisy real-world metadata, while a parallel KG-mapping pathway supports alignment to DBpedia and this http URL. On the SemTab 2024 Metadata-to-KG track, the official GT-strict evaluation was modest. However, a blinded diagnostic audit indicates that many mismatches reflect benchmark granularity, aliasing, and ontology-selection effects rather than wholly implausible header-centric predictions. We report this audit as diagnostic evidence on disagreement patterns, not as revised benchmark performance. Overall, the paper presents a reusable workflow for metadata-driven semantic annotation, data source-level quality monitoring, and KG-oriented benchmark diagnosis.
Comments: 18 pages, 4 figures, Workshop on Quality of Knowledge Graphs at ESWC 2026, May 11, 2026, Dubrovnik, Croatia
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.10541 [cs.AI]
  (or arXiv:2610.10541v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.10541

arXiv-issued DOI via DataCite

Submission history

From: Marcelo Valentim Silva [view email]
[v1] Tue, 12 May 2026 05:52:55 UTC (218 KB)

来源:arXiv:cs.AI · arxiv.org