跳到正文
arXiv:cs.AI· Uddipan Basu Bir, Vincent Christlein, Andreas Maier, Mathias Zinnen·· 3 小时前

轻量级 VLM 如何为文化遗产档案实现 OCR 到结构化 JSON 提取

From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction

AI 导读

一项对比研究评估了八款开源轻量级 VLM(最高 7B 参数)在三个高校文化遗产馆藏上的 OCR 到结构化 JSON 提取表现,在 zero-shot、few-shot 和微调设置下以 CER、ANLS* 和 mAP-F1 衡量效果。

正文

View PDF HTML (experimental)

Abstract:While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing. This is especially acute in heritage digitization, where documents include historical handwriting, domain-specific terminology (e.g., jewelry, prehistory, architecture), and non-standard layouts requiring high-dimensional structured extraction. We present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections. Given a document image, models must extract text and generate schema-compliant JSON, enabling automatic validation and downstream use. We evaluate models under a constraint-aware protocol across zero-shot, few-shot, and fine-tuning settings, measuring extraction fidelity and structured-output quality using Character Error Rate (CER), Approximate Normalized Levenshtein Similarity (ANLS*), and mean Average Precision F1 (mAP-F1). Against a fine-tuning baseline, we further test the independent impact of (i) hyperparameter optimization, (ii) classical image preprocessing (illumination flattening, denoising, and CLAHE), and (iii) multi-stage training. Finally, we analyze the trade-off between dataset-specific fine-tuning and a single multi-dataset checkpoint, where joint training enables one model to operate across collections but can shift performance between datasets. Overall, we show that carefully adapted VLMs with up to 7B parameters can provide a sustainable, private, high-performing alternative to manual transcription or commercial black-box systems, and we offer actionable guidance for heritage institutions seeking institution-controlled OCR-to-JSON extraction.
Comments: 17 pages. Published in Document Analysis and Recognition - ICDAR 2026, LNCS vol. 16974, Springer. Code: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11818 [cs.CV]
  (or arXiv:2610.11818v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.11818

arXiv-issued DOI via DataCite (pending registration)

Journal reference: Document Analysis and Recognition - ICDAR 2026, Lecture Notes in Computer Science, vol. 16974, pp. 502-519, Springer, 2027
Related DOI: https://doi.org/10.1007/978-3-032-36039-7_30

DOI(s) linking to related resources

Submission history

From: Uddipan Basu Bir [view email]
[v1] Thu, 8 Oct 2026 12:21:52 UTC (866 KB)

来源:arXiv:cs.AI · arxiv.org