跳到正文
arXiv:cs.CL· Yong Cao, Markus Flicke, Haoyu He, Katrin Renz, Andreas Geiger·· 3 小时前AI 评分33

LLM4Impact:整合异构信息进行科学影响力预测

LLM4Impact: Integrating Heterogeneous Information for Scientific Impact Prediction

AI 导读

LLM4Impact 是一种证据感知的科学影响力预测方法,整合语义、图、LLM 与时间四类表示,并通过连续前缀 token 将图信息注入冻结的 LLM。该方法还构建了含 200 万篇论文的大规模基准数据集,在分布内测试集上年 RMSE 降低 10.13%,域外分布下降低 6.87%。作者将随论文发布代码、基准与交互式网页演示。

正文

View PDF

Abstract:Predicting the future impact of a newly published paper is challenging because it must be inferred from heterogeneous evidence available at publication time. Existing approaches often rely on a single source of information or combine multiple sources without accounting for their different predictive roles. In this paper, we present LLM4Impact, an evidence-aware method for scientific impact prediction that learns to represent, integrate, and calibrate heterogeneous information. LLM4Impact combines semantic, graph, LLM, and temporal representations, and injects graph information into a frozen LLM through continuous prefix tokens. A context aware gating mechanism adaptively weights different evidence, while a separate calibration module accounts for domain and temporal variation in citation scales. We further construct a large-scale benchmark dataset with 2 million papers, leakage-safe point-in-time heterogeneous ego graphs, temporal splits, and both year-level and month-level citation targets. Experiments show that LLM4Impact consistently outperforms strong semantic, graph, and LLM based baselines, with a 10.13% reduction in year RMSE on the in distribution test set and a 6.87% reduction under out-of-domain distribution. Our results reveal that the value of such evidence is context dependent: different papers benefit from different sources, while domain and publication time affect how evidence translates into citations. This finding motivates adaptive evidence selection and context-conditioned calibration rather than simply richer representations. We will release our code, benchmark, and an interactive web demonstration upon publication.
Comments: 27 pages, 12 figures, 11 tables
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.10138 [cs.CL]
  (or arXiv:2610.10138v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.10138

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yong Cao [view email]
[v1] Wed, 7 Oct 2026 14:18:33 UTC (2,740 KB)

来源:arXiv:cs.CL · arxiv.org