跳到正文
arXiv:cs.LG· Aadi Chauhan, Arthur Ilyasov·· 4 小时前AI 评分34

小型 GUI grounding 模型该如何接收动作类型?Qwen2-VL-2B 上的五种条件注入方式对比

Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type?

AI 导读

研究用 LoRA 微调 Qwen2-VL-2B,在 Android in the Wild 上比较五种向小型 GUI grounding 模型注入动作类型的方式。

正文

View PDF HTML (experimental)

Abstract:A GUI agent decides which action to take and where to take it; we ask how a small grounding model should receive the action type. Fine-tuning Qwen2-VL-2B with LoRA on Android in the Wild, we compare a flat baseline with five ways of supplying the type under matched data, compute, and decoding: an auxiliary loss, a hard-routed action word, an additive learned embedding, a prepended learned token, and the type written into the prompt. With five seeds, an episode-clustered bootstrap, and seed-level paired tests, the ranking on a mixed stream is clear: the auxiliary loss, the additive embedding, and the prompt word each gain five to seven hit@0.10 points over the baseline, while hard routing and the prepended token are not distinguishable from it. Much of that gain is protection from a preprocessing choice of ours rather than a spatial prior. Our serializer clamps the off-screen touch point AITW records for type events to the origin; that class degrades the baseline's click grounding, and removing it lifts the baseline by nearly seven points, after which no mechanism's hit rate beats it and the intervals exclude a two-point effect, though the auxiliary loss still shortens the average miss; on a stream of taps and swipes none helps. Whether this generalizes beyond one serialization is open. For deployment, the pipeline's margin over the baseline with predicted rather than gold types is not established (+0.016, 95% interval [-0.017, +0.052]), and a wrong type collapses every model conditioned at inference. The prepended token does not help at the shared learning rate, where its rows barely move from initialization; trained ten times faster it reaches the level of the other three, with a margin three seeds do not establish. We also document a silent failure: injecting conditioning through inputs_embeds makes Qwen2-VL fall back to 1-D positions for image tokens, costing nine points.
Comments: 18 pages, 4 figures, 12 tables. Code and per-example logs are available at this https URL
Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2610.07444 [cs.LG]
  (or arXiv:2610.07444v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07444

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Aadi Chauhan [view email]
[v1] Mon, 5 Oct 2026 21:50:50 UTC (1,576 KB)

来源:arXiv:cs.LG · arxiv.org