arXiv:cs.CL· Jimmy Lin, Sahel Sharifymoghaddam, Lingwei Gu, Nour Jedidi·· 3 小时前
Project Greenhouse:迈向完全开放与主权的智能体搜索
Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search
AI 导读
Project Greenhouse 首次里程碑展示用少量 GPU 从零预训练加监督微调,构建出可用的 pointwise decoder-only 重排序器,全程不依赖第三方开放权重骨干模型,实现端到端自主训练。团队基于公开数据集完成大部分实验,并开源数据、代码、配置及 Gaggle 系列模型 checkpoint,支持透明独立复现。
正文
Abstract:Project Greenhouse represents our exploration of a simple thesis: We believe that it is possible to build fully open and sovereign models for agentic search with only modest computational resources. As a first milestone, we describe how to build a competitive pointwise decoder-only reranker using a simple two-step recipe comprising pre-training from scratch followed by supervised fine-tuning, starting only from commonly available datasets. Contrary to the dominant approach in the literature, we do not rely on existing open-weight backbones from third parties, and thus we are fully in control of model training, from end to end. We were able to accomplish the bulk of our experiments using no more than a handful of GPUs. This report articulates the importance and benefits of our approach, and we share artifacts that enable transparent, independent reproduction of all aspects of model training. Beyond data, code, and configurations that capture our efforts, we also release checkpoints for our family of Gaggle models, demonstrating the feasibility of our approach and providing a first step toward validating our broader thesis.
| Subjects: | Information Retrieval (cs.IR); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.11922 [cs.IR] |
| (or arXiv:2610.11922v1 [cs.IR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11922 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jimmy Lin [view email]
[v1]
Thu, 8 Oct 2026 13:19:46 UTC (87 KB)
来源:arXiv:cs.CL · arxiv.org