arXiv:cs.AI· Bhawani Shankar Leelar, Pawan Chorasiya, Davin Hill, Robert E. Tillman, Tamer Soliman·· 3 小时前
可见推理并非通用优化器:分析代码生成中受人格设定与思维模式影响的效应
Visible Reasoning Is Not a Universal Optimizer: Persona- and Thinking-Dependent Effects in Analytics Code Generation
AI 导读
一项基于 SQL-pandas 配对查询基准的研究发现,可见 Chain-of-Thought 并不带来普遍准确率优势,将推理表示与目标语言(SQL 或 Python)匹配也无一致收益。其效果取决于人格措辞、目标语言、模型配置与内部推理设置,控制消融实验进一步区分了推理内容与提示词格式的影响,26 页、8 图、16 表。
正文
Abstract:Visible Chain-of-Thought (CoT) is often treated as a broadly useful reasoning instruction, yet analytics code generation combines natural-language ambiguity, schema grounding, target-language constraints, and model-specific inference behavior. Because the same analytics request can be expressed in two distinct target languages-SQL and Python (pandas)-this setting provides a natural test of a common but under-examined assumption: that visible reasoning is more effective when its representation matches the requested target, as in "think in SQL" or "think in Python." Together with generic instructions such as "think step-by-step," such recommendations remain insufficiently evaluated under controlled, execution-based comparisons. We study a query matched SQL-pandas benchmark that crosses persona phrasing, target language, visible-CoT format, control prefixes, direct generation, and internal-reasoning configurations. The results do not support either a universal accuracy advantage from visible CoT or a consistent benefit from matching the reasoning representation to the target language. Instead, the effects depend on the persona, target, model configuration, and internal-reasoning setting. The control ablations further distinguish effects of reasoning content from those of prompt format. These findings indicate that reasoning strategies should be selected jointly for the model, persona, target, and internal-reasoning configuration rather than adopted as universal defaults. More broadly, the study provides a controlled framework for identifying when visible reasoning improves executable generation, when it primarily perturbs model behavior, and when the internal-reasoning configuration is the more consequential factor.
| Comments: | 26 pages, 8 figures, 16 tables |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Software Engineering (cs.SE) |
| Cite as: | arXiv:2610.10639 [cs.LG] |
| (or arXiv:2610.10639v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10639 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Bhawani Shankar Leelar [view email]
[v1]
Wed, 7 Oct 2026 14:27:58 UTC (329 KB)
来源:arXiv:cs.AI · arxiv.org