跳到正文
arXiv:cs.LG· Yongsheng Luo, Wengan He, Yu Li, Rouying Wu, Wei Lv·· 7 小时前AI 评分35

多模态几何表征中的方向依赖响应:不止于扰动幅度

Beyond Perturbation Magnitude: Direction-Dependent Responses in Multimodal Geometric Representations

AI 导读

研究发现,多模态几何对齐分数对模态退化的响应并非由扰动位移幅度主导:位移幅度最多只能解释绝对响应中 15% 的样本外方差,幅度相同的扰动对响应也系统性不同。

正文

View PDF HTML (experimental)

Abstract:Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the response of a multimodal geometric score is determined primarily by the magnitude of the perturbation-induced displacement. Using frozen cohorts from MSR-VTT (N=878) and DiDeMo (N=980), we apply controlled video blur and audio noise and analyze the response in the relational geometry on which the score is defined. Displacement magnitude explains at most 15% of the out-of-sample variance in the absolute response, and magnitude-matched pairs respond systematically differently, so scalar magnitude does not organize the response. The closed-form first-order expansion of the Gramian volume yields the Directional Geometric Response (DGR): the projection of the displacement onto the local volume gradient, which jointly captures the clean operating point, displacement magnitude, and displacement direction. The absolute first-order DGR term explains the observed response with out-of-sample R^2 of 0.838-0.969, matched-magnitude ranking accuracies of 0.864-0.963, and response-sign accuracies of 0.909-0.989, whereas the tested direction-free alternatives remain weak or unstable under the corresponding evaluation protocols. A pre-specified gain-normalization candidate, V/(g_V+eps), fails its predictability and clean-order gates. DGR uses the observed degraded-state displacement and is therefore an explanatory quantity, not a deployment-time predictor: geometric response depends on where the representation operates, how far degradation moves the relational geometry, and in which direction it moves.
Comments: Submitted to IEEE Transactions on Multimedia (TMM). 12 pages, 6 figures, 3 tables
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD)
ACM classes: I.2.6; I.4.8
Cite as: arXiv:2610.08533 [cs.CV]
  (or arXiv:2610.08533v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.08533

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yongsheng Luo [view email]
[v1] Tue, 6 Oct 2026 15:25:47 UTC (3,028 KB)

来源:arXiv:cs.LG · arxiv.org