跳到正文
arXiv:cs.AI· Adrian Cierpka, Mohammad Shafiqul Islam, Johannes Steinh\"ulb, Eric Dietriche Sesso Domtchoueng, Michael Selzer, Arnd Koeppe·· 6 小时前AI 评分28

KadiAssistant:面向 Kadi4Mat 信息检索的对话式 AI 智能体

KadiAssistant: A conversational AI Agent for information retrieval in Kadi4Mat

AI 导读

KadiAssistant 是一个集成到 Kadi 科研数据生态系统的隐私优先 AI 助手,结合自托管大语言模型与隐私保护语义搜索(受 RAG 启发),可访问 Kadi 上的文件与记录元数据,将异构数据筛选、聚合并结构化为高信息量答案。它面向材料科学等跨学科场景,能在遵守 Kadi 细粒度访问权限的前提下完成信息检索,并强化 FAIR 原则中的可发现性。

正文

View PDF HTML (experimental)

Abstract:We introduce KadiAssistant, a privacy-by-design AI assistant integrated into the Kadi research data ecosystem, enabling researchers to efficiently access, aggregate, and synthesize information from heterogeneous, privacy-sensitive research data. Interdisciplinary fields such as materials science bring together disciplines with their own terminology and standards. While this convergence fuels innovation, it also makes it increasingly difficult to connect and access knowledge, as data are distributed across disciplines, organizations, and individuals. For example, battery research combines electrochemical measurements, materials characterization data, physics-based simulations, and manufacturing parameters, each using different formats, vocabularies, and standards. Efficiently storing and sharing such heterogeneous data via research data platforms, such as Kadi4Mat, demands domain knowledge, technical expertise, and familiarity with metadata schemas and interfaces. Research data also vary in sensitivity: newly generated 'warm' data are often private, whereas published 'cold' data are usually openly accessible. The Kadi ecosystem offers fine-grained access control needed for sensitive data. A solution for efficient information retrieval in Kadi must therefore respect the fine-grained access permissions. To address these intertwined challenges of information retrieval, strong data privacy, and complex access control, KadiAssistant combines a self-hosted large language model (LLM) with a privacy-preserving semantic search, inspired by retrieval-augmented generation, that can access files and record metadata on Kadi. This allows the assistant to screen, aggregate, and structure information into a highly informative answer. KadiAssistant therefore bridges terminology and standards, lowers access barriers for researchers, and strengthens the Findable pillar of FAIR data principles.
Subjects: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2605.18850 [cs.IR]
  (or arXiv:2605.18850v2 [cs.IR] for this version)
  https://doi.org/10.48550/arXiv.2605.18850

arXiv-issued DOI via DataCite

Submission history

From: Adrian Cierpka [view email]
[v1] Wed, 13 May 2026 09:15:35 UTC (692 KB)
[v2] Tue, 6 Oct 2026 10:42:08 UTC (1,298 KB)

来源:arXiv:cs.AI · arxiv.org