arXiv:cs.LG· Shivam Gupta·· 3 小时前AI 评分44
个人 AI 记忆用于评分预测的控制性审计研究
A Controlled Audit of Personal AI Memory for Rating Prediction
AI 导读
一项冻结评测在 Coat 和 MovieLens 的 400 个用户画像、6,160 条目标评分上审计个人 AI 记忆:在 Coat 上,Qwen 编写的 Mem0 流水线相比完整历史使 Qwen 的用户宏平均绝对误差增加 0.084、Phi 增加 0.149,两者经家族校正的 bootstrap 区间均不含零。
正文
Abstract:In structured rating prediction, does a personal AI use historical item-rating associations, or mainly the user's rating tendencies? We audit this distinction by permuting historical ratings within each user while preserving the exact rating distribution, item support, and metadata. We combine this control with full history, native memory extraction, and matched numerical readers in a publicly frozen evaluation of 400 held-out user profiles and 6,160 target ratings across Coat and MovieLens. On Coat, the tested Qwen-written Mem0 pipeline increases user-macro mean absolute error relative to full history by 0.084 for Qwen and 0.149 for Phi; both family-adjusted bootstrap intervals exclude zero. Correct historical assignments help both readers on Coat, but the corresponding MovieLens effects are smaller and inconclusive after adjustment. A history-only ridge reader outperforms Qwen in both domains and Phi on MovieLens, while the Coat Phi comparison is unresolved. All 2,800 reader calls, including 150 invalid outputs, are retained under a fixed fallback rule. A separate implementation verifies inputs, metrics, and all ten primary contrasts. The contribution is a reproducible diagnostic study showing why extraction, association use, output reliability, and reader choice require separate evaluation.
| Comments: | 15 pages. Code and reproducibility materials: this https URL |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02764 [cs.LG] |
| (or arXiv:2610.02764v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02764 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shivam Gupta [view email]
[v1]
Fri, 2 Oct 2026 03:43:17 UTC (87 KB)
来源:arXiv:cs.LG · arxiv.org