arXiv:cs.LG· Jacob Carlson·· 5 小时前AI 评分32
从非结构化数据中可解释地发现:一种高维方法
Interpretable Discovery from Unstructured Data: A High-Dimensional Approach
AI 导读
研究者提出一个自动通用框架,借助 AI 可解释性方法将非结构化数据(如开放式经济信念调查文本)转化为可解释概念测量的高维结构化数据集,再基于高维多重检验算法筛选出“发现”,并自动生成和评估自然语言描述。该框架研究者自由度低、抗数据窥探,能自动复现既有发现、补充细节并产生全新发现。
正文
Abstract:We propose an automatic, general-purpose framework for making discoveries from unstructured data (e.g., text data from open-ended surveys of economic beliefs). The framework leverages recent methods from the literature on AI interpretability to transform unstructured datasets into high-dimensional, structured datasets of interpretable concept measurements; specifies concept-level parameters and null hypotheses based on this transformed dataset; tests these hypotheses using algorithms validated by new results in high-dimensional multiple testing, producing a selected set ("discoveries"); and both generates and evaluates human-interpretable natural language descriptions of these discoveries. The proposed framework has few researcher degrees of freedom, is robust to data snooping, mitigates under-exploration, and facilitates fast and inexpensive sensitivity analysis and replication. We revisit applications to recent descriptive and causal analyses of unstructured data in empirical economics, and find this framework is able to automatically replicate existing discoveries, add nuance to others, and make entirely new discoveries as well.
| Subjects: | Econometrics (econ.EM); Machine Learning (cs.LG) |
| Cite as: | arXiv:2511.01680 [econ.EM] |
| (or arXiv:2511.01680v5 [econ.EM] for this version) | |
| https://doi.org/10.48550/arXiv.2511.01680 arXiv-issued DOI via DataCite |
Submission history
From: Jacob Carlson [view email]
[v1]
Mon, 3 Nov 2025 15:42:32 UTC (32 KB)
[v2]
Sun, 11 Jan 2026 19:06:37 UTC (446 KB)
[v3]
Tue, 5 May 2026 15:42:49 UTC (607 KB)
[v4]
Wed, 15 Jul 2026 14:09:50 UTC (3,196 KB)
[v5]
Fri, 2 Oct 2026 00:27:15 UTC (4,646 KB)
来源:arXiv:cs.LG · arxiv.org