跳到正文
arXiv:cs.AI· Yuto Suzuki, Farnoush Banaei-Kashani·· 7 小时前AI 评分42

Universe of Thoughts(UoT):大语言模型创造性推理的计算框架

Universe of Thoughts: A Computational Framework for Creative Reasoning in Large Language Models

AI 导读

研究者提出 Universe of Thoughts(UoT)框架,将组合式、探索式与转换式创造力形式化为可执行的计算算子,用于 LLM 的创造性推理。在 GPT-4o 上,T-UoT 在低约束、高目标明确性的 Bridge 和 Electricity 任务中表现最强,C-UoT 则在 Society 任务上相对表现最佳。

正文

View PDF HTML (experimental)

Abstract:Recent advances in Large Language Model (LLM) reasoning have improved conventional problem solving, but creative reasoning remains comparatively underexplored. Inspired by cognitive science, we formalize combinational, exploratory, and transformational creativity as executable computational operators over structured problem and solution spaces, specifying how each mode combines, explores, or transforms those spaces. Combinational reasoning transfers ideas across domains to form unfamiliar combinations; exploratory reasoning searches for new solutions within an existing conceptual space; and transformational reasoning modifies the rules or constraints that define that space. This formalization yields distinct algorithmic procedures, which we instantiate in Universe of Thoughts (UoT), an LLM reasoning framework. Existing creativity benchmarks emphasize either open-ended ideation or highly constrained problem solving. We therefore introduce three novel creative-reasoning tasks requiring concrete solutions in low-constraint settings. Across 10 generations per method and task, T-UoT with GPT-4o performs strongest on the low-constraint, high-objective-specificity Bridge and Electricity tasks, while C-UoT shows its strongest relative performance on the low-constraint, lower-objective-specificity Society task. In addition, we evaluate UoT on HypoArena, an independent scientific hypothesis-generation benchmark with 100 tasks across biomedical, machine-learning, and social-science domains. With Qwen3-14B, Exploratory UoT ranks first among seven reasoning methods, achieving a 32.7\% pairwise win rate compared with 25.5\% for the next-best method. Our results suggest distinct performance patterns across task structures: T-UoT is strongest in low-constraint, high-specificity settings, E-UoT in more constrained, high-specificity settings, and C-UoT in low-constraint, lower-specificity settings.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2511.20471 [cs.AI]
  (or arXiv:2511.20471v3 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2511.20471

arXiv-issued DOI via DataCite

Submission history

From: Yuto Suzuki [view email]
[v1] Tue, 25 Nov 2025 16:34:59 UTC (749 KB)
[v2] Wed, 26 Nov 2025 02:28:35 UTC (749 KB)
[v3] Tue, 6 Oct 2026 15:08:49 UTC (1,710 KB)

来源:arXiv:cs.AI · arxiv.org