arXiv:cs.LG(机器学习,全量分类)· Guransh Singh·· 1 天前AI 评分36
AEGIS:面向知识保留的视觉-语言-动作微调的锚点强制梯度隔离
AEGIS: Anchor-Enforced Gradient Isolation for Knowledge-Preserving Vision-Language-Action Fine-Tuning
AI 导读
AEGIS 是一种无缓冲区、逐层正交梯度投影框架,可在持续流匹配微调中隔离预训练表征,避免机器人操作微调破坏 VLM 的视觉推理能力。该方法在 PaliGemma2-3B 上于 LIBERO 基准测试中完整保留了预训练 VQA 性能与基线 holdout loss,同时匹配连续动作收敛,且无需 replay buffer、教师模型或混合批次 VQA 协同训练。
正文
Abstract:Fine-tuning pre-trained Vision-Language Models (VLMs) for robotic manipulation introduces a fundamental stability-plasticity dilemma: continuous flow-matching action experts backpropagate concentrated, low-rank regression gradients into transformer backbones trained on high-dimensional cross-entropy objectives. This cross-modal gradient asymmetry rapidly degrades pre-trained visual reasoning. Existing solutions either disconnect continuous gradient flow via stop-gradients or constrain updates via LoRA, which restricts update rank but remains directionally blind to semantic corruption; both typically rely on mixed-batch VQA co-training, doubling training compute. We introduce AEGIS (Anchor-Enforced Gradient Isolation System), a buffer-free, layer-wise orthogonal gradient projection framework enabling continuous flow-matching fine-tuning while isolating pre-trained representations from destructive parameter updates. Prior to training, AEGIS estimates per-layer Gaussian activation statistics from pre-training data as a static reference anchor. During fine-tuning, a closed-form Wasserstein-2 transport penalty generates an anchor-restoration gradient through the active computation graph. A sequential dual-backward pass applies layer-wise Gram-Schmidt orthogonalization, projecting task gradients onto the orthogonal complement of the restoration vector during directional conflict. We establish an exact energy preservation bound for layer-wise orthogonal projection, showing that AEGIS sheds only 0.62% of gradient energy empirically while halting cumulative feature drift. On PaliGemma2-3B fine-tuned on the LIBERO manipulation benchmark, AEGIS fully preserves pre-trained Visual Question Answering performance and baseline holdout loss while matching continuous action convergence, without replay buffers, teacher models, or co-training data.
| Subjects: | Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2604.16067 [cs.LG] |
| (or arXiv:2604.16067v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.16067 arXiv-issued DOI via DataCite |
Submission history
From: Guransh Singh [view email]
[v1]
Fri, 17 Apr 2026 13:49:57 UTC (2,808 KB)
[v2]
Thu, 1 Oct 2026 11:13:03 UTC (2,809 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org