跳到正文
arXiv:cs.CL· Tanmoy Chakraborty, Ayan Sengupta, Suparna Bhattacharya, Partha Pratim Chakrabarti, Amlan Chakrabarti, Supratik Chakraborty, Partha Pratim Das, Lipika Dey, Richa Singh, Mayank Vatsa·· 3 小时前AI 评分39

大语言模型的潜在性能画像(Latent Performance Profiling)

Latent Performance Profiling of Large Language Models

AI 导读

研究者提出 Latent Performance Profiling(LPP)框架,从模型隐藏激活和输出分布中提取任务无关的诊断指标,对 8 个 0.5B-14B 规模的开源 LLM 进行分析,发现 benchmark 分数相近的模型在熵、可压缩性等潜在特征上可能截然不同。作者建议将 LPP 与 MMLU-Pro、BBH、IFEval 等基准并列报告,以支持更可靠的模型选择与安全评估。

正文

View PDF

Abstract:Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabilities. Evaluating open-source LLMs on leaderboards faces persistent issues such as data contamination, a narrow task scope, and poor alignment with real-world reliability. Benchmark-based evaluations such as MMLU-Pro, BBH, or IFEval primarily capture \textit{what} a model outputs on fixed test sets, not \textit{how} it processes information, calibrates uncertainty, or structures internal knowledge. In this article, we advocate for a shift from benchmark-centric evaluation toward a complementary, \textit{state-centered intrinsic assessment} of LLMs. To this end, we introduce \textbf{Latent Performance Profiling (LPP)} --- a framework that derives task-agnostic diagnostics from hidden activations and output distributions. LPP defines a set of scalar metrics on a model's latent representations and dynamics, revealing traits that enable interpretable comparisons and uncover hidden vulnerabilities. Unlike static accuracy scores, LPP provides stable, architecture-sensitive signatures across models of similar size. With extensive empirical analyses across eight LLMs, spanning a size range of 0.5B-14B, we demonstrate that models with similar benchmark scores can exhibit contrasting latent profiles, such as differences in entropy or compressibility. Guided by these insights, we design synthetic probes for uncertainty and symbolic reasoning that align with intrinsic metrics while decoupling from leaderboard bias. We recommend reporting LPP alongside benchmarks to provide a deeper, interpretable understanding of model behavior, enabling more reliable model selection, safety assessment, and evaluation beyond surface-level accuracy.
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2605.30018 [cs.CL]
  (or arXiv:2605.30018v3 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2605.30018

arXiv-issued DOI via DataCite

Submission history

From: Ayan Sengupta [view email]
[v1] Thu, 28 May 2026 14:41:26 UTC (893 KB)
[v2] Fri, 29 May 2026 04:32:57 UTC (893 KB)
[v3] Wed, 7 Oct 2026 08:05:25 UTC (1,858 KB)

来源:arXiv:cs.CL · arxiv.org