跳到正文
arXiv:cs.CL· Clayton Cohn, Joyce Fonteles, Kirk Vanacore, Gianni Mazza, Candida Crawford, Tom Hooper, Gautam Biswas, Rene Kizilcec·· 6 小时前AI 评分38

跨模型 LLM 共识不等于有效:K-12 数学辅导对话中学生失败模式诊断研究

Agreement Is Not Validity: Cross-Model LLM Consensus in Diagnosing Student Failure Modes in K-12 Math Tutoring Dialogue

AI 导读

一项针对 K-12 数学辅导对话的探索性研究发现,LLM 跨模型标注一致性(kappa = .755-.781;alpha = .769)显著高于人机一致性(kappa = .524-.597),说明跨模型共识可能制造正确性的假象。研究用诊断编码手册评估了不确定性、错误归因、算子选择、概念缺口和程序失误五种学生失败模式。作者指出,模型共识不能替代对推理构念有效性的独立证据。

正文

View PDF HTML (experimental)

Abstract:In K-12 mathematics tutoring, student-tutor dialogue provides rich evidence of learners' problem-solving processes and sources of difficulty. Learning analytics research increasingly relies on large language models (LLMs) to extract such information from dialogue for a variety of downstream tasks, including knowledge tracing, behavioral modeling, and diagnosis of student reasoning errors. However, the validity of these model-generated interpretations remains insufficiently understood. In this exploratory study, we examine the validity of LLM classifications of five student failure modes in mathematics tutoring dialogue using an operational diagnostic codebook: uncertainty, misattribution, operator selection, conceptual gap, and procedural slip. Across models, human-LLM agreement was moderate (kappa = .524-.597), while cross-model agreement was substantially higher (kappa = .755-.781; alpha = .769). These findings show that cross-model agreement can create a misleading appearance of correctness, challenging the assumption that consensus among LLMs constitutes evidence of valid learner interpretation. For learning analytics, the implication is clear: scalable labeling is useful only if the inferred constructs are valid, and model consensus cannot substitute for independent evidence of that validity.
Comments: Submitted to LAK27 as a short paper. Currently under review
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.08703 [cs.CL]
  (or arXiv:2610.08703v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08703

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Clayton Cohn [view email]
[v1] Tue, 6 Oct 2026 17:16:03 UTC (164 KB)

来源:arXiv:cs.CL · arxiv.org