跳到正文
arXiv:cs.LG· Samuel Tetteh, Cody Fleming·· 4 小时前AI 评分48

冻结视觉语言安全评分并非感知危险:一项针对 CLIP prompt-margin 评分的受控评估

It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank

AI 导读

一项受控评估发现,冻结 CLIP prompt-margin 安全评分在接触前约 20 步下降,但机制对照显示它主要追踪的是与其 caption 库共享的场景相似度,而非危险本身。该研究覆盖 3 个策略、180 个 episode 和 130 个孤立接触起点,评分随 caption 区分度与相机视角变化。恒定置信度对照仍保留较低灾难率的点估计,说明策略收益并不能证明模型具备危险感知。

正文

View PDF HTML (experimental)

Abstract:Frozen vision-language models increasingly provide safety signals for reinforcement learning. Their use assumes that similarity to language describing danger indicates the hazard itself. Yet policy return and collision rate cannot reveal whether a score detects hazards or responds to correlated features of the scene. VLM-based methods have reported gains in driving and safe-RL benchmarks by converting image-text similarity into rewards, costs, or confidence weights. Such signals promise to reduce reliance on manually designed feedback. They may also reflect prompt structure, embedding geometry, or camera viewpoint, leaving their safety meaning unverified. To address this gap, we present a controlled evaluation of a frozen CLIP prompt-margin safety score. We apply the score to trajectories generated by policies that never receive it, match pre-contact observations to contact-free observations with comparable hazard geometry, and vary the captions, encoder, and camera view. Across three policies, 180 episodes, and 130 isolated contact onsets, the score decreases for about twenty steps before contact. Mechanism controls indicate that the score mainly tracks resemblance to the scene shared by its captions and changes with caption separation and camera view. A constant-confidence control retains the lower catastrophe-rate point estimate, so policy gains do not establish hazard perception.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2610.09517 [cs.CV]
  (or arXiv:2610.09517v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.09517

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Samuel Tetteh [view email]
[v1] Wed, 7 Oct 2026 06:15:23 UTC (2,333 KB)

来源:arXiv:cs.LG · arxiv.org