跳到正文
arXiv:cs.AI· Jingbo Yue, Bruce Coburn, Jinge Ma, Jui-Feng Chi, Fengqing Zhu·· 3 小时前

用多模态食品项验证与恢复改进单图营养估算

Improving Image-Based Nutrition Estimation Through Multimodal Food-Item Verification and Recovery

AI 导读

研究者提出一套无需任务特定微调的框架,用多模态大语言模型(MLLM)先清点可见食物,再分别验证食物身份与每个候选 2D 区域是否支持份量估算。整图复核依据验证结果找出未解决缺口和遗漏食物,触发至多一次定向恢复,恢复区域在无恢复提示词条件下重新验证后并入最终食物项集合。

正文

View PDF HTML (experimental)

Abstract:Single-image nutrition estimation can fail silently when visible foods are missed. Even when a food is correctly identified, its proposed region may not support portion estimation. We propose a framework that uses multimodal large language models (MLLMs) to inventory visible foods and separately verify food identity and whether each proposed 2D region supports portion estimation. One whole-image review uses these verification results to identify unresolved gaps and omitted foods, triggering at most one targeted recovery pass. Recovered regions are re-verified without access to the recovery prompt, then reconciled into a final item set for nutrition estimation. The framework requires no task-specific fine-tuning. Matched evaluation on common valid-output samples shows that item-level grounding improves mass accuracy across all tested settings and energy accuracy relative to an adapted retrieval baseline, with item-identity precision and recall also improving, while post-recovery visual coverage is assessed separately at inference time without ground-truth annotations.
Comments: 5 pages, 2 figures, 3 tables. Submitted to IEEE ICASSP 2027
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Image and Video Processing (eess.IV)
Cite as: arXiv:2610.11144 [cs.CV]
  (or arXiv:2610.11144v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.11144

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jingbo Yue [view email]
[v1] Thu, 8 Oct 2026 03:09:29 UTC (409 KB)

来源:arXiv:cs.AI · arxiv.org