跳到正文
arXiv:cs.CL· Jorge Castillo Sep\'ulveda, Marco Torres Y\'evenes, Juan Carlos Lanas·· 4 小时前

四个开放权重模型中的系统提示词条件与隐藏状态几何:更正与留存结论

System-Prompt Conditioning and Hidden-State Geometry in Four Open-Weight Models: Corrections and What Survives

AI 导读

一篇预印本 v3 更正了此前关于系统提示词在四个开放权重模型末层隐藏状态中留下几何指纹的结论:被称为 Ollivier-Ricci 曲率的统计量实为基于时序与余弦 k-NN 图的 Forman 型边统计量,且检验在边层面进行,此前按 regime 划分的主张与"从方向转向幅度"的主张均被撤回。

正文

View PDF HTML (experimental)

Abstract:Versions 1 and 2 of this preprint reported that an identity-specifying system prompt leaves a geometric fingerprint in the final-layer hidden-state trajectories of four open-weight language models, and that instruction tuning moves this fingerprint from the direction to the magnitude of the hidden-state vector. An audit of their code and data found the following. The curvature statistic described as Ollivier-Ricci curvature on Euclidean k-NN graphs was a non-standard Forman-type edge statistic on graphs built from temporal and cosine k-NN edges. Its released test permuted pooled edges instead of trajectories, and the published p-values came from unreleased code. The quantity reported as the norm of the first generated state is the state at the last prompt position, from which the first output token is predicted. The generic control prompt was matched to the identity prompt in characters, not in tokens. This version corrects the methods, withdraws the regime-specific claims (one model per regime) and the direction-to-magnitude claim, and re-analyzes the data with added controls. What survives is narrower. Centroid distance, maximum mean discrepancy and a linear probe separate every pair of prompt conditions in every model, while the curvature statistic exceeds its split-half noise floor in only four of twelve comparisons. In Gemma-4-E4B-it this state has a lower norm under the identity prompt than under a token-length-matched generic prompt (138.1 vs. 216.5; Cohen's d = -5.45; n = 20). Its direction also separates the conditions, and the effect fits the state's role in planning the output: the identity prompt instructs a pause before every answer, and the model opens 98 of 100 responses with a pause marker. When the first token is fixed, the norm ordering reverses. The base model continues the prompt template instead of answering. A redesigned follow-up study is in preparation.
Comments: v3: substantial correction, replaces v1-v2. The curvature reported as Ollivier-Ricci was a Forman-type statistic on temporal plus cosine k-NN graphs, tested at edge level. The regime-specific and direction-to-magnitude claims are withdrawn. Added: simple-baseline, split-half, trajectory-level and token-length-matched controls. 14 pages, 1 figure, 9 tables
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2607.09842 [cs.LG]
  (or arXiv:2607.09842v3 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2607.09842

arXiv-issued DOI via DataCite

Submission history

From: Jorge Castillo [view email]
[v1] Fri, 10 Jul 2026 17:22:36 UTC (22 KB)
[v2] Sat, 1 Aug 2026 03:44:05 UTC (366 KB)
[v3] Thu, 8 Oct 2026 01:08:49 UTC (39 KB)

来源:arXiv:cs.CL · arxiv.org