arXiv:cs.LG(机器学习,全量分类)· Dmitrij \.Zatuchin·· 14 小时前AI 评分38
LLM 品牌推荐中的系统归因:单条回答可识别系统,聚合品牌画像无法迁移
System Attribution in LLM Brand Recommendations: Single Responses Identify the System, Aggregated Brand Profiles Do Not Transfer
AI 导读
一项针对 6,475 条存储回答(6,324 条可分析)的研究测试了 LLM 品牌推荐画像能否描述系统本身。
正文
Abstract:Audits of AI visibility summarise the brand recommendations of deployed language models into per-system profiles. We test whether such a profile describes the system on one corpus of 6,475 stored responses (6,324 analysable) collected between December 2025 and February 2026 from five deployed endpoints across gift-recommendation, corporate-reputation and category-ownership queries. The collection harness cut many answers short: 83.1% of Gemini 3 Flash answers in category ownership end mid-sentence under a 1,024-token output cap. With every answer cut to its first 800 characters, a character n-gram classifier cross-validated by prompt attributes one response to GPT-5.2, Gemini 3 Flash, Gemini 3 Flash with search, Grok or Perplexity sonar-pro with 97.84% accuracy (5,028 responses, 383 prompts, majority class 31.5%, 30 split seeds). Length alone falls to the majority rate, 24 formatting statistics reach 95.79%, and masking brand names and capitalised tokens leaves 97.72%. Held-out query conditions keep 97.43% weighted by size and 88.0% unweighted; in a retrieval-grounded arm that changes the harness, no Grok answer is attributed to Grok (0/120). Aggregated into 50 model-by-domain-by-condition units, twelve behavioural features separate four systems at 66.53% under grouped cross-validation, against a label-permutation null with mean 33.71% and 95th percentile 46.0%. Across domains the aggregate profile fails: a forest trained on category-ownership units assigns all 22 gift units to the wrong system, consistent with a reversal in brand volume (8.41 against 0.94 brands per response in gifts, 3.01 against 3.91 in category ownership), while single responses transfer at 89.92% balanced accuracy. The surface form of an answer carries the system across the query domains tested; aggregated brand behaviour does not, and the uncrossed design cannot separate the system from the domain or the harness.
| Comments: | 30 pages, 5 figures, 9 tables. Appendix D documents corrections to an earlier manuscript |
| Subjects: | Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00253 [cs.IR] |
| (or arXiv:2610.00253v1 [cs.IR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00253 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Dmitrij Żatuchin [view email]
[v1]
Wed, 23 Sep 2026 21:57:07 UTC (519 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org