跳到正文
arXiv:cs.AI· Ashutosh Srivastava, Siddharth Yedlapati, Vinay Aggarwal, Yaman Kumar Singla, Shashwat Dixit, Jitendra Ajmera, Balaji Krishnamurthy·· 7 小时前AI 评分46

大语言模型个性化能力基准测试:SDR-Arena 与 SDR-Bench

Benchmarking the Personalization Capabilities of Large Language Models

AI 导读

研究者提出 SDR-Arena 框架与 SDR-Bench 语料库(覆盖 22 个行业、3500 家企业、5 万条客户成功案例),用于基准测试"双方生成式个性化"——即用结果前信息重建赢得交易的论点,以加权 nugget-recall 指标(WCS)评分。

正文

View PDF HTML (experimental)

Abstract:Personalization is classically a two-party problem: a sender chooses what to say, and a receiver with independent objectives decides whether to act. A salesperson pitching the same analytics product leads with HIPAA compliance for a hospital and real-time reporting for a retailer, expecting a different argument to work on each. Existing LLM personalization benchmarks measure a narrower, one-party property: whether output matches the preferences of the same user it serves-sender and receiver being the same, as when RLHF aligns an assistant to its own user. The two-party case is harder to study automatically, since it needs ground truth linking specific content to an observed receiver action.
Sales outreach provides this: a message written for one prospect, recorded against whether it produced a reply, a call, or a closed deal. We introduce SDR-Arena, a framework for benchmarking two-party generative personalization at scale, and SDR-Bench, a public corpus of 50,000 customer success stories across 22 industries and 3500 enterprises. Given only pre-outcome information, an agent must reconstruct the arguments that won the deal, scored by a weighted nugget-recall metric (WCS). The best model, Claude Sonnet 4.6, reaches 55.8% WCS indicating it recovers only half the winning content-a plateau we observe across model families that costly deep-research pipelines do not close. An ablation shows the cause is retrieval, not reasoning: models improve substantially given the facts a human researcher would gather, but rarely find them through web search alone. Two studies with professional SDRs support the metric: only 48% of generated pitches were rated usable without editing, and WCS yields model rankings consistent with evaluation against expert strategies authored independently of any model output.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2607.20471 [cs.AI]
  (or arXiv:2607.20471v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2607.20471

arXiv-issued DOI via DataCite

Submission history

From: Siddharth Yedlapati [view email]
[v1] Sat, 23 May 2026 18:16:54 UTC (7,253 KB)
[v2] Tue, 6 Oct 2026 10:05:25 UTC (762 KB)

来源:arXiv:cs.AI · arxiv.org