arXiv:cs.AI· Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid·· 4 小时前AI 评分42
面向生成式 AI 工作负载的智能内容摄取系统
Smart Content Ingestion for Generative AI Workloads
AI 导读
一篇 arXiv 论文提出生产级内容提取系统,把内容摄取明确为生成式 AI 生命周期中的独立阶段,包含选择性 OCR 路由、基于参考的打分器、结构感知父子分块器和只读检索评估器。在 180 篇文档语料上,最佳提取器得分 97.4/100(字符错误率 0.13%,表格相似度 0.995)。分块器在 25,050 个生成问题上取得 hit@1 68.6%、hit@10 92.8%、MRR 0.77。
正文
Abstract:The evolution of machine learning has progressively changed where intelligence resides in an AI system. In conventional machine learning the task, data representation, labels and model architecture were tightly coupled, so data preparation was narrow, schema-bound and visible. Generative AI decouples the model from any single task: one foundation model serves open-ended downstream tasks, and the generality gained on the model side is matched by heterogeneity on the data side, because enterprise knowledge is authored in the formats people use (PDF, presentations, spreadsheets, scanned documents, forms, tables, diagrams and mixed-layout files) that carry textual, visual, geometric and structural information at once. A language model or retriever cannot reason reliably over information misrepresented at this interface, so content extraction becomes a lifecycle stage in its own right whose errors no downstream retriever or re-ranker can repair. This paper presents a production-ready content-extraction system that makes this stage explicit, configurable, and measurable. The system incorporates selective OCR routing, a scarcity-first curation engine with a reference-based extraction scorer that measures character, word, and table-structure accuracy, a deterministic structure-aware parent-child chunker, and a read-only retrieval evaluator that generates grounded questions from every page and reports Hit@k, mean reciprocal rank, and latency. On a 180-document corpus the best extractor scores 97.4 of 100 (character error rate 0.13%, table similarity 0.995) and the chunker reaches hit@1 of 68.6%, hit@10 of 92.8% and MRR 0.77 over 25,050 generated questions. We distil three design principles (structure before semantics, never mutate what you measure, budget your labels) and position measured content extraction as the perception layer of enterprise agentic systems.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2610.07091 [cs.AI] |
| (or arXiv:2610.07091v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07091 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Abbas Raza Ali [view email]
[v1]
Mon, 5 Oct 2026 13:22:01 UTC (161 KB)
来源:arXiv:cs.AI · arxiv.org