跳到正文
arXiv:cs.AI· Yaxiao Liu (PwC China AI Center), Pengbo Liu (PwC China AI Center), Yiwen Liu (PwC China AI Center), Yihua Guan (PwC China AI Center), Jiaxing Song (Tsinghua University)·· 4 小时前AI 评分24

从迁移到校准:跨模型、跨司法辖区与跨规模保持智能体能力

From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale

AI 导读

一篇 arXiv 方法论提案提出"智能体校准"框架,将部署、迁移或扩展智能体视为标准优先的适配过程:定义基础能力、技术环境与用户上下文三类标准,诊断差距、生成并应用修订,再在固定预算内复检同一标准。

正文

View PDF HTML (experimental)

Abstract:Deploying, migrating, or scaling an agent can change its model, harness, infrastructure, application, and intended users. We formulate agent calibration as standards-first adaptation: define basic-capability, technical-environment, and user-context standards; diagnose gaps; generate and apply revisions; and recheck the same standards within fixed budgets. These standard families interact across information, harness, and user-acceptance layers. Source behavior is diagnostic, not a perfect reference or capability ceiling: model replacement can turn correct answers into errors or errors into correct answers. Qualification requires all mandatory known tests, actual end-to-end deployment paths, hard predicates, and declared task/user minimums to pass; aggregate gains cannot erase hard failures. Revisions may change tools or harnesses, add demonstrations and task descriptions, or use validated target-native trajectories to train a policy, controller, or compact skill model served through the harness. Semantic checkpoints validate executed artifacts, localize repair, and revalidate dependencies. The loop exports reusable configuration or training artifacts with a qualification record, while final task outputs undergo their own checks. Independent factual evidence precedes relative preference judgment; DPO and GRPO optimize policies rather than establish truth. Frozen held-out evaluation tests generalization and compares equal-budget target-native optimization. We specify an automatic calibration tool using limited authorized user trajectories and tests as future work. The framework and tool remain proposals; confirmatory empirical validation is pending.
Comments: 40 pages, 8 figures. Methodological proposal; no confirmatory empirical results reported. CPU controller update/export example is implemented on synthetic data only; the integrated automatic RL calibration tool, consultant workflow, trajectory-distilled skill service, and manufacturing experiments remain future work
Subjects: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Software Engineering (cs.SE)
Cite as: arXiv:2609.35149 [cs.AI]
  (or arXiv:2609.35149v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.35149

arXiv-issued DOI via DataCite

Submission history

From: Yaxiao Liu [view email]
[v1] Mon, 28 Sep 2026 13:31:46 UTC (807 KB)
[v2] Fri, 2 Oct 2026 13:41:01 UTC (960 KB)

来源:arXiv:cs.AI · arxiv.org