跳到正文
arXiv:cs.LG· Alex Glushkovsky·· 3 小时前

基于预期非对称几何的双变量因果方向识别

Identification of Bivariate Causal Directionality Based on Anticipated Asymmetric Geometries

AI 导读

论文提出 AAG 与 MI 两种基于条件分布的双变量因果方向识别方法,AAG 将实际条件分布与基于均值和标准差的预期正态分布沿两变量比较,并评估了 Pearson 相关、余弦距离、Hellinger 距离、Jaccard 指数、Jeffreys 散度、K-L 散度、K-S 距离、MAE、MSE、互信息和 Wasserstein 距离等指标。

正文

View PDF

Abstract:Identification of causal directionality in bivariate numerical data is a fundamental research problem with important practical implications. This paper presents two alternative methods to identify direction of causation by considering conditional distributions: (1) Anticipated Asymmetric Geometries (AAG) and (2) Monotonicity Index (MI). The AAG method compares the actual conditional distributions to anticipated ones along two variables. Different comparison metrics, such as Pearson correlation, cosine distance, Hellinger distance, Jaccard index, Jeffreys divergence, K-L divergence, K-S distance, MAE, MSE, mutual information, and Wasserstein distance have been evaluated. Anticipated distributions have been projected as normal based on dual response statistics: mean and standard deviation. The MI method compares the calculated monotonicity indexes of the gradients of conditional distributions along two axes and exhibits count of gradient sign changes. Both methods assume stochastic properties of the bivariate data and exploit anticipated unimodality of conditional distributions of the effect. The proposed methods are straightforward and include only a limited number of hyperparameters that affect the accuracy of the identification. For a given set of hyperparameters, both the AAG and MI methods provide a unique, deterministic solution. To address sensitivity to hyperparameters, tuning has been done by utilizing a full factorial Design of Experiment. It turns out that the AAG method outperforms MI, achieving top weighted accuracies of 82.7% with simple tuning and 84.4% with size-adaptive tuning, compared with 81.6% for GRCI or 82.0% for CAREFL-H on the 99 pairs of the Tubingen real-world cause-effect examples. To evaluate the decisiveness of the identification method, a decision tree was fitted on the input data's symmetrical bivariate statistics to detect misclassified cases.
Comments: 17 pages, 8 figure, 7 tables
Subjects: Machine Learning (cs.LG); Logic in Computer Science (cs.LO)
Cite as: arXiv:2603.26024 [cs.LG]
  (or arXiv:2603.26024v4 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2603.26024

arXiv-issued DOI via DataCite

Submission history

From: Alex Glushkovsky [view email]
[v1] Fri, 27 Mar 2026 02:45:46 UTC (1,156 KB)
[v2] Mon, 31 Aug 2026 01:06:07 UTC (1,156 KB)
[v3] Thu, 3 Sep 2026 01:04:27 UTC (1,146 KB)
[v4] Thu, 8 Oct 2026 03:21:07 UTC (1,195 KB)

来源:arXiv:cs.LG · arxiv.org