跳到正文
原文
LlamaIndex:产品、工程与评测·· 3 小时前AI 评分64

LlamaIndex 发布 Extract v2.5 文档提取 Agent

Introducing Extract v2.5: Next-Gen Document Extraction Agents ->

AI 导读

LlamaIndex 发布 Extract v2.5,新一代基于 schema 的文档提取 Agent,三个档位在 ExtractBench 上准确率均提升:Cost Effective 从 87.1 升至 93.9,Agentic 从 89.8 升至 95.8,Agentic Plus 从 95.1 升至 96.4,且每页定价不变。

正文

Today we’re introducing Extract v2.5, a new generation of our schema-based document extraction agents. This release brings accuracy improvements across all tiers as well as enhanced grounding capabilities to our highest tiers, Agentic and Agentic Plus. These improvements come with no increase in per-page pricing, ensuring higher performance per dollar no matter which tier you are using.

On ExtractBench, accuracy improved across all three tiers. Cost Effective now outperforms the previous Agentic tier, rising from 87.1 to 93.9 overall value F1, and Agentic outperforms the prior Agentic Plus, improving from 89.8 to 95.8. Agentic Plus, our highest-accuracy tier, rises from 95.1 to 96.4.

Under the hood, we made architectural enhancements that contribute to these performance improvements, including a new agent harness, purpose-built for document extraction. We took inspiration from the latest coding agents, and hyper-tuned everything around the models to handle the real failure modes in document extraction, across vision, reasoning, and verification. The result is a specialized agent that can cross-reference context from multiple pages, and ground values in exact sources.

The first change is Structural Reasoning. We use structural reasoning over document representations for more efficient and tailored extraction based on document type, layout, and information density. Extract will adaptively reason and spend effort for your more complex docs to deliver maximum accuracy. For the less complex docs, it’ll process with lower latency and target its effort.

This release introduces Advanced Citations for Agentic and Agentic Plus, helping locate supporting evidence for extracted values. ExtractBench grounding scores rise from 46.8 to 80.6 for Agentic and from 46.4 to 82.2 for Agentic Plus, which means more correct and well fit bounding boxes for even your most complex layouts and extractions.

These improvements show up across ExtractBench challenges, including long lists, records spanning pages, and scanned forms. The following show some examples of how extraction and grounding improve on these tasks.

Long lists

Long lists can cause even frontier VLMs to stop early or lose track of repeated records. Extract v2.5 uses intermediate representations to work with the full dataset and validate its output against your schema. Across long-list tasks, Cost Effective’s score rises from 82.0 to 92.6, and Agentic’s from 86.1 to 95.5.

On a 17-page fund filing, Cost Effective previously returned 87 of 238 holdings. Extract v2.5 returns all 238, with the document’s extraction score rising from 52.8 to 98.7.

238 of 238 holdings each line is one holding

  1. p1 12
  2. p2 15
  3. p3 15
  4. p4 14
  5. p5 15
  6. p6 16
  7. p7 14
  8. p8 15
  9. p9 15
  10. p10 15
  11. p11 15
  12. p12 15
  13. p13 14
  14. p14 15
  15. p15 15
  16. p16 14
  17. p17 4

holdings: [

… { issuer_name: "CENTENE CORPORATION", value_usd: 203604.00 }, 87 holdings

{ issuer_name: "Aktiebolaget Volvo", value_usd: 552328.91 }, { issuer_name: "AVNET, INC.", value_usd: 361862.00 }, … 151 missed

238 holdings

]

Records spanning pages

In real world documents, extraction must reason across pages. Content on one page might span multiple pages, or be joined with an appendix many pages later. Extract v2.5 is better at preserving the relationship between the fields as it builds the output.

In this grant schedule, Agentic v2.0 stopped at the page break, leaving the grant’s purpose and address incomplete. Extract v2.5 includes the continuation from the next page in the same record.

address_line1
1800 WASHINGTON BLVD STE 340
city, state, zip_code
"" BALTIMORE MD 21230
purpose_of_grant
EDUCATION, HUMAN & SOCIAL SERVICES, DISASTER RELIEF & RECOVERY, COMMUNITY DEVELOPMENT

United Way Worldwide Schedule I, pages 5 and 6, page 5 Page 5

United Way Worldwide Schedule I, pages 5 and 6, page 6 Page 6

Scanned forms

Scanned forms often mix printed fields, overlaid text, handwriting, and reviewer annotations. Scan noise common in such forms can make these layers even harder to distinguish. A printed value might be crossed out and hand corrected after printing. Extract v2.5 is better at visually interpreting the field in context and separating the requested value from the surrounding marks. It’s now also much better at grounding the extracted value back to the source document for cross checking and reference.

On this annotated permit, Agentic previously combined a reviewer’s date with the permit number and appended a reviewer note to the freshwater depth. Extract v2.5 separates the original values from the annotations into the fields requested by the schema.

amendment_permit_no
"7/29/2019 13445" "13445"
amendment_permit_no_review
null "7/29/2019"

deepest_freshwater_zone
"4000 BUQW: 500(1800/3925) USDW: 4000 GEO ISOLATION: 4250" "4000"
deepest_freshwater_zone_review
"4000 BUQW: 500(1800/3925) USDW: 4000 GEO ISOLATION: 4250" "BUQW: 500(1800/3925) USDW: 4000 GEO ISOLATION: 4250"

W-14 disposal permit W14-57728, page 1

Advanced Citations

Advanced Citations locate bounding boxes of supporting evidence for difficult fields, including values that don’t appear word for word in the document. An initial version was previously available in Agentic Plus. With v2.5, we’ve improved its grounding accuracy further and extended it to Agentic.

Texas P-4, page 1 field_name Pleasanton (Edwards Lime) lease_name L.S.U. Leal current_operator_name SNG Operating, LLC operator_p5_no 797970 oil_lse_gas_id_no 14068 county Atascosa operator_address 1714 Fortview Road, #106 Austin, Texas 78704 well_nos 1 classification_oil false classification_gas false classification_other false effective_date May 1, 2010 change_operator true change_oil_condensate_gatherer false change_gas_gatherer false change_gas_purchaser false change_gas_purchaser_system_code false change_field_name false change_lease_name false new_rrc_oil_lease false new_rrc_gas_well false new_rrc_other_well false due_new_completion_recompletion false due_reclass_gas_to_oil false due_reclass_oil_to_gas false due_consolidation false

Bounding box citations combine well with confidence scores. Confidence scores help teams decide which values can proceed automatically and which need a closer look. Reviewers can then use citations to check those values against the original document in human-in-the-loop workflows. Our confidence-score guide explains how to evaluate this on your own documents.

How we built v2.5

At LlamaIndex, we know documents deeply. We’ve built that expertise into tools our extraction agents can call to read, transform, and locate information in documents. Those tools draw on purpose-built models and document processing methods to handle inconsistent file formats, ambiguous layouts, and source coordinates.

Reading order, table structure, and the relationship between text and its position on the page all affect extraction. Our tools give agents ways to work with these structures and return to the source as they build the output.

For v2.5, we brought the harness behind Agentic Plus to Cost Effective and Agentic, adapting its tools and extraction strategies to each tier’s cost constraints. We refined how to organize the task, what context to give the agent, and when it should inspect the document more closely.

Getting those choices right took extensive testing. An approach that works well for a short form may struggle with hundreds of repeated records. We evaluated document representations, tools, and model configurations across these tasks, weighing accuracy gains against the time and processing they require.

We also worked on how agents follow your schema and extraction instructions, including tasks with up to 3,200 fields. If record IDs are essential for matching extracted records to your database, for example, call that out in your prompt so the agent can give them extra attention.

We’ve also added native spreadsheet extraction. In spreadsheet mode, agents work directly with workbook cells rather than a flattened representation, giving them another way to work with structured data. You can enable it in your extraction configuration.

Try v2.5 on your documents

Try v2.5 on documents representative of your workflow, using your existing schemas. We think this version will be a significant step up in extraction quality, especially for documents that previously needed manual cleanup or weren’t reliable enough to automate. We’re looking forward to seeing what you build.

来源:LlamaIndex:产品、工程与评测 · llamaindex.ai