跳到正文
arXiv:cs.AI· Remo Grillo, Lukas Klic, Giovanni Colavizza·· 6 小时前AI 评分37

QRAKEN:用图蒸馏与语义自愈让 LLM 用自然语言查询知识图谱

Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing

AI 导读

QRAKEN 是一个免训练、与本体无关的神经符号流水线,用离线蒸馏产出的 TTQL 描述多跳模式、条件频率与路径示例,让 LLM 基于图的实际数据而非 schema 预期生成查询。

正文

View PDF HTML (experimental)

Abstract:Natural-language access to RDF knowledge graphs is a core Semantic Web ambition. Large language models (LLMs) have advanced Text-to-SPARQL, yet on unfamiliar graphs they often generate valid queries that misrepresent the populated data model. QRAKEN is a training-free, ontology-agnostic neurosymbolic pipeline grounding generation in empirical graph evidence rather than schema expectations. An offline distiller produces TTQL, a compact description of populated multi-hop patterns, conditional frequencies and path-conditioned literal examples, plus a class-property co-occurrence matrix. Online, TTQL guides the LLM, while deterministic syntax, vocabulary and data-model checks provide diagnostics for iterative refinement. On CK25 (First International Text2SPARQL Challenge), under matched-condition recomputation on a QLever snapshot, QRAKEN achieves strict F1 of 0.643 $\pm$ 0.026 with GPT-4.1 mini and 0.652 $\pm$ 0.012 with GPT-5.4: relative gains of 30% and 32% over the strongest recomputed participant, outperforming systems using the same base model family. Ablations identify TTQL patterns as the dominant driver (+0.31 strict F1 over a shape-only baseline); the refinement loop provides a cheap safety net, rejecting triple patterns unsupported by the co-occurrence matrix. Compared with auto-derived SHACL, TTQL yields 64% higher strict F1, supporting the value of empirical patterns beyond schema exposure. With two local 35B 4-bit open-weight models at zero marginal cost, the same pipeline matches the strongest recomputed participant, and TTQL advantages over shape-only and SHACL baselines persist. Results on a single, relatively small benchmark provide an initial empirical signal; monolithic TTQL injection on very open cross-domain graphs remains the main limitation.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.08095 [cs.AI]
  (or arXiv:2610.08095v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.08095

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Remo Grillo [view email]
[v1] Tue, 6 Oct 2026 10:24:26 UTC (68 KB)

来源:arXiv:cs.AI · arxiv.org