跳到正文
arXiv:cs.AI· Iordanis Fostiropoulos, Muhammad Rafay Azhar, Abdalaziz Sawwan, Boyu Fang, Yuchen Liu, Jiayi Liu, Hanchao Yu, Qi Guo, Jianyu Wang, Fei Liu, Xiangjun Fan·· 4 小时前AI 评分43

GISTBench:通过基于证据的兴趣验证评估 LLM 用户理解能力

GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification

AI 导读

研究者推出 GISTBench,一个评估 LLM 从推荐系统交互历史中理解用户能力的基准,区别于传统侧重物品预测准确率的 RecSys 基准。该基准提出 Interest Groundedness(IG,含 precision 与 recall)和 Interest Specificity(IS)两类新指标,并发布基于全球短视频平台真实用户交互构建的合成数据集。

正文

Authors:Iordanis Fostiropoulos, Muhammad Rafay Azhar, Abdalaziz Sawwan, Boyu Fang, Yuchen Liu, Jiayi Liu, Hanchao Yu, Qi Guo, Jianyu Wang, Fei Liu, Xiangjun Fan

View PDF HTML (experimental)

Abstract:We introduce GISTBench, a benchmark for evaluating Large Language Models' (LLMs) ability to understand users from their interaction histories in recommendation systems. Unlike traditional RecSys benchmarks that focus on item prediction accuracy, our benchmark evaluates how well LLMs can extract and verify user interests from engagement data. We propose two novel metric families: Interest Groundedness (IG), decomposed into precision and recall components to separately penalize hallucinated interest categories and reward coverage, and Interest Specificity (IS), which assesses the distinctiveness of verified LLM-predicted user profiles. We release a synthetic dataset constructed on real user interactions on a global short-form video platform. Our dataset contains both implicit and explicit engagement signals and rich textual descriptions. We validate our dataset fidelity against user surveys, and evaluate eight open-weight LLMs spanning 7B to 235B parameters, together with three proprietary frontier models (GPT-5, Claude 4.6, and Gemini 3.5 Flash). Our findings reveal performance bottlenecks in current LLMs, particularly their limited ability to accurately count and attribute engagement signals across heterogeneous interaction types.
Comments: 9 figures, 20 tables; code at this https URL
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
ACM classes: H.3.3; I.2.7
Cite as: arXiv:2603.29112 [cs.AI]
  (or arXiv:2603.29112v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2603.29112

arXiv-issued DOI via DataCite

Submission history

From: Iordanis Fostiropoulos [view email]
[v1] Tue, 31 Mar 2026 01:01:56 UTC (10,694 KB)
[v2] Thu, 1 Oct 2026 19:26:55 UTC (9,787 KB)

来源:arXiv:cs.AI · arxiv.org