跳到正文
arXiv:cs.AI· Albert Dorador·· 3 小时前

一次排列就够了:快速、确定性的特征重要性与模型压力测试

One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing

AI 导读

该研究提出用单个 max-min 秩最优确定性排列替代排列特征重要性中 B 次随机打乱,在保持甚至提升与真实重要性相关性的同时消除估计方差,复杂度从 O(B·n·p) 降至 O(n·p)。

正文

View PDF HTML (experimental)

Abstract:Reliable estimation of feature contributions in machine learning models is essential for transparency, algorithmic fairness, and regulatory compliance. While permutation feature importance is widely used, classical implementations rely on repeated Monte Carlo shuffling, introducing significant computational overhead and stochastic instability. In this paper, we show that replacing $B$ random permutations with a single, max-min rank-optimal deterministic permutation maintains or improves correlation with ground-truth importance while eliminating estimation variance and reducing complexity from $O(B \cdot n \cdot p)$ to $O(n \cdot p)$. Under location-scale feature distributions, we formally prove exact recovery of scale-adjusted linear regression coefficients, alongside improved importance estimation under concave model sensitivity. We extend this deterministic framework along two complementary dimensions. First, Systemic Feature Importance (SFI) integrates empirical feature correlations to quantify indirect feature reliance through proxy variables. Second, Importance Direction extends scalar importance to a signed, directional representation by measuring concordance between covariate displacements and output shifts. Extensive empirical validation across nearly 200 simulation scenarios demonstrates superior bias-variance trade-offs in high-dimensional and low signal-to-noise regimes. Finally, two real-world credit risk case studies show how coupling SFI with Importance Direction enables practitioners and regulators to audit models for both the magnitude and net sign of hidden reliance on protected attributes, delivering a principled, transparent, and scalable framework for model governance.
Subjects: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2512.13892 [stat.ML]
  (or arXiv:2512.13892v3 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2512.13892

arXiv-issued DOI via DataCite

Submission history

From: Albert Dorador-Chalar [view email]
[v1] Mon, 15 Dec 2025 20:50:54 UTC (17,191 KB)
[v2] Tue, 23 Dec 2025 12:54:15 UTC (17,193 KB)
[v3] Wed, 7 Oct 2026 20:19:23 UTC (18,195 KB)

来源:arXiv:cs.AI · arxiv.org