跳到正文
原文
xAI:News(网页)·· 3 小时前AI 评分55

xAI 发布 Grok Collections API,内置 RAG 检索系统

Grok Collections API State-of-the-art RAG system built directly into our API. Dec 22, 2025

AI 导读

xAI 发布 Collections API,开发者可上传 PDF、Excel、代码库等文件构建知识库并检索,无需自建索引和检索基础设施。检索定价为每 1000 次搜索 $2.50,首周文件索引和存储免费;支持语义、关键词和混合检索,并提供 reranker 模型与 reciprocal rank fusion。

正文

Today, we're excited to announce Collections API. With Collections, you can upload and search through entire datasets. From PDFs and Excel sheets to entire codebases, you can upload your files into a knowledge base that supports precise and fast search. This allows developers to build RAG applications without the headache of managing indexing and retrieval infrastructure.

To help you get started, we're making file indexing and storage free for the first week*, with retrieval priced at a flat rate of $2.50 per 1,000 searches.

Indexing

  • Powerful document understanding: We use OCR and layout-aware parsing to extract text while preserving structure such as the layout of a PDF, hierarchy of an Excel table, or the syntax of code.
  • Smart file management: Easily upload, update, and download files. And when a file changes, our system efficiently reindexes it to ensure your collection is never stale.
  • Broad format support: Collections supports a wide range of file types. (see full list)

Retrieval

Choose the retrieval method that best fits your use case:

  • Semantic search: To search using the meaning and intent behind a query.
  • Keyword search: For precise term matching.
  • Hybrid search: For the highest accuracy, combine keyword and semantic search. We support both a dedicated reranker model and reciprocal rank fusion.

What is our financial forecast for Q1 2026?

Financial_plan_2026.txt

Our company's annual financial projections indicate a robust growth trajectory for the upcoming fiscal year, with expected revenue increases driven by expanded market share in emerging sectors. Analysts predict a 15% rise in Q1 2026, bolstered by strategic investments in technology and supply chain optimization. Key metrics such as EBITDA and net profit margins are forecasted to improve.

Benchmark Results

Our Collections API delivers state-of-the-art retrieval performance, matching or outperforming leading models in real-world RAG tasks across finance, legal, and coding domains.

These fields are especially challenging due to their long, dense documents. To avoid hallucinations and deliver reliable answers, models must retrieve the exact passages and reason over them accurately.

Accuracy*

(Higher is better)

TaskxAI

Grok 4.1 Fast

Google

Gemini Pro 3

OpenAI

GPT 5.1

Finance

Tabular and numerical questions

93.085.984.7
Legal

Complex reasoning over multiple chunks

73.974.571.2
Coding

Code understanding and large file systems

868581

*Internal source.

Financial Analysis

Extracting tabular and numerical data from files can be challenging with semantic search alone. Hybrid search enables you to accurately retrieve this data from documents such as SEC filings*, allowing the model to precisely reference information.

*Based on an internal dataset.

**Gemini does not expose the actual retrieved files so this metric measures the files cited by Gemini rather than the raw retrieved files. We set the default top k for Gemini to be 20 passages.

Legal Analysis (LegalBench)

The LegalBench dataset tests retrieval and reasoning over nuanced legal language and complex cross-references, consisting of 128 challenging question-answer pairs drawn from an extensive corpus of authentic commercial contracts across multiple datasets.

*Gemini does not expose the actual retrieved files so this metric measures the files cited by Gemini rather than the raw retrieved files. We set the default top k for Gemini to be 20 passages.

Codebase (DeepCodeBench)

Code understanding is crucial for applications such as code summarization and generation. We use the DeepCodeBench dataset to comprehensively benchmark for this. It features a diverse set of tasks drawn from real-world open-source repositories, API usage, and complex algorithmic problems.

*We evaluated the code understanding capability of agentic search on 232 code Q&A datapoints from DeepCodeBench across 8 repositories containing 8,000 files.

Data Privacy

We do not use user data stored on Collections for model training purposes, unless the user has given consent.

Creating and Searching Collections

py

Using Collections in Chat

py

Direct API Usage

sh

*You may be charged after the free trial period. We will follow up with more information.

来源:xAI:News(网页) · x.ai