跳到正文
arXiv:cs.LG· Junjie Xiao, Huiwen Jia·· 7 小时前AI 评分34

Consideration Circuits:超越单一 Softmax 的深度分离与普适性

Consideration Circuits: Depth Separation and Universality Beyond a Single Softmax

AI 导读

研究者提出 consideration circuits(CC),一种由多级 MNL 单元组成的有向无环图构建的多阶段选择模型,突破传统单一 softmax 评分框架。

正文

View PDF HTML (experimental)

Abstract:Most feature-based choice models, classical and deep, score items and apply a single softmax. We introduce consideration circuits (CC), feature-based models of multi-stage choice defined by directed acyclic graphs of multinomial logit (MNL) units. Source units assign probabilities to menu items, and internal units combine predecessor distributions using MNL weights computed from their probability-weighted feature summaries. On a three-item compromise task with fixed non-collinear features, menu-independent random-utility models (RUM), including a single MNL unit, suffer an error bounded away from zero. For CC, in contrast, we establish a sharp depth--norm separation: increasing depth from $2$ to $3$ reduces the optimal maximum taste-vector norm for error $\epsilon$ from $\Theta(\log(1/\epsilon)/\epsilon)$ to $\Theta(\log(1/\epsilon))$. The depth-$2$ lower bound holds for arbitrary width and menu-independent routing biases, while a five-node depth-$3$ circuit with zero routing biases attains the logarithmic rate. More generally, we characterize two geometric conditions that are necessary and sufficient for approximating arbitrary deterministic choice tables on finite menu families. Under these conditions, depth $3$ suffices, while depth $4$ achieves optimal logarithmic norm scaling whenever the family contains a non-singleton menu. In experiments, standalone tree circuits with fewer than $600$ parameters attain the lowest mean test negative log-likelihood (NLL) among the evaluated models on four fixed-pool benchmarks and the Expedia temporal split. As output heads, CC generalize the linear MNL readout and lower mean test NLL for every tested encoder on Expedia and Trivago.
Comments: 43 pages, 4 figures, 11 tables
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2610.04143 [cs.LG]
  (or arXiv:2610.04143v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.04143

arXiv-issued DOI via DataCite

Submission history

From: Huiwen Jia [view email]
[v1] Fri, 2 Oct 2026 23:31:16 UTC (577 KB)
[v2] Tue, 6 Oct 2026 06:09:19 UTC (577 KB)

来源:arXiv:cs.LG · arxiv.org