跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Amirhosein Azarpour, Seyyed Moein Kazemi·· 14 小时前AI 评分37

SW-KAN:采用 Stieltjes-Wigert q-正交多项式的 Kolmogorov-Arnold 网络

SW-KAN: Kolmogorov-Arnold Networks with Stieltjes-Wigert q-Orthogonal Polynomials

AI 导读

研究者提出 SW-KAN,一种基于定义在半无限域 (0, infinity) 上 Stieltjes-Wigert q-正交多项式的 Kolmogorov-Arnold 网络新架构,通过指数-of-tanh 映射弥合无界输入与有界正交基之间的域失配,并以三项递推在 O(N) 内完成多项式展开。

正文

View PDF HTML (experimental)

Abstract:Kolmogorov-Arnold Networks (KANs) represent a paradigmatic shift in deep learning by replacing fixed node activations with learnable univariate functions on edges, offering enhanced interpretability and parameter efficiency. While recent polynomial-based KAN variants have addressed the computational overhead of original B-spline implementations, they introduce a fundamental yet underexplored challenge: the domain mismatch between unbounded real-valued inputs and the bounded or semi-infinite support of orthogonal polynomial bases. To address this limitation, we propose the Stieltjes-Wigert Kolmogorov-Arnold Network (SW-KAN), a novel architecture that employs Stieltjes-Wigert q-orthogonal polynomials defined on the semi-infinite domain (0, infinity). We introduce a smooth exponential-of-tanh mapping that stably bridges the domain gap while preserving well-conditioned gradients, and leverage a numerically stable three-term recurrence that evaluates polynomial expansions in O(N) operations without special-function calls. Through comprehensive experiments spanning image classification and continuous function approximation, we demonstrate that SW-KAN achieves superior accuracy-efficiency trade-offs across diverse tasks. The log-normal weight structure and learnable q-parameter of Stieltjes-Wigert polynomials provide a distinct inductive bias that enables robust performance under resource-constrained conditions, including reduced feature dimensionality and limited training data. The proposed architecture not only outperforms established polynomial KAN baselines on standard benchmarks but also exhibits strong representational capacity for approximating complex multivariate functions with remarkably few parameters, making it a compelling alternative for efficient function approximation and classification in resource-constrained settings.
Comments: 22 pages, Code and pretrained models available at: this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
ACM classes: I.2.6; I.5.1
Cite as: arXiv:2610.00050 [cs.LG]
  (or arXiv:2610.00050v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00050

arXiv-issued DOI via DataCite

Submission history

From: Amirhosein Azarpour [view email]
[v1] Thu, 3 Sep 2026 19:46:42 UTC (730 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org