arXiv:cs.AI· Yaxiao Liu (PwC China AI Center), Pengbo Liu (PwC China AI Center), Yiwen Liu (PwC China AI Center), Yihua Guan (PwC China AI Center), Jiaxing Song (Tsinghua University)·· 4 小时前AI 评分24
从迁移到校准:跨模型、跨司法辖区与跨规模保持智能体能力
From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale
AI 导读
一篇 arXiv 方法论提案提出"智能体校准"框架,将部署、迁移或扩展智能体视为标准优先的适配过程:定义基础能力、技术环境与用户上下文三类标准,诊断差距、生成并应用修订,再在固定预算内复检同一标准。
正文
Abstract:Deploying, migrating, or scaling an agent can change its model, harness, infrastructure, application, and intended users. We formulate agent calibration as standards-first adaptation: define basic-capability, technical-environment, and user-context standards; diagnose gaps; generate and apply revisions; and recheck the same standards within fixed budgets. These standard families interact across information, harness, and user-acceptance layers. Source behavior is diagnostic, not a perfect reference or capability ceiling: model replacement can turn correct answers into errors or errors into correct answers. Qualification requires all mandatory known tests, actual end-to-end deployment paths, hard predicates, and declared task/user minimums to pass; aggregate gains cannot erase hard failures. Revisions may change tools or harnesses, add demonstrations and task descriptions, or use validated target-native trajectories to train a policy, controller, or compact skill model served through the harness. Semantic checkpoints validate executed artifacts, localize repair, and revalidate dependencies. The loop exports reusable configuration or training artifacts with a qualification record, while final task outputs undergo their own checks. Independent factual evidence precedes relative preference judgment; DPO and GRPO optimize policies rather than establish truth. Frozen held-out evaluation tests generalization and compares equal-budget target-native optimization. We specify an automatic calibration tool using limited authorized user trajectories and tests as future work. The framework and tool remain proposals; confirmatory empirical validation is pending.
| Comments: | 40 pages, 8 figures. Methodological proposal; no confirmatory empirical results reported. CPU controller update/export example is implemented on synthetic data only; the integrated automatic RL calibration tool, consultant workflow, trajectory-distilled skill service, and manufacturing experiments remain future work |
| Subjects: | Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Software Engineering (cs.SE) |
| Cite as: | arXiv:2609.35149 [cs.AI] |
| (or arXiv:2609.35149v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35149 arXiv-issued DOI via DataCite |
Submission history
From: Yaxiao Liu [view email]
[v1]
Mon, 28 Sep 2026 13:31:46 UTC (807 KB)
[v2]
Fri, 2 Oct 2026 13:41:01 UTC (960 KB)
来源:arXiv:cs.AI · arxiv.org