TY - GEN
T1 - KOLMOGOROV-ARNOLD TRANSFORMER
AU - Yang, Xingyi
AU - Wang, Xinchao
N1 - Publisher Copyright:
© 2025 13th International Conference on Learning Representations, ICLR 2025. All rights reserved.
PY - 2025
Y1 - 2025
N2 - Transformers are the cornerstone of modern deep learning. Traditionally, they use multi-layer perceptron (MLP) layers to mix channel information. In this paper, we introduce the Kolmogorov-Arnold Transformer (KAT), which replaces MLP layers with Kolmogorov-Arnold Network (KAN) layers to enhance model expressiveness. Integrating KANs into transformers, however, is no easy feat, especially when scaled up. Specifically, we identify three key challenges: (C1) Base function. The standard B-spline used in KANs is inefficient for parallel computing, slowing inference. (C2) Parameter and Computation Inefficiency. KAN requires a unique function for each input-output pair, leading to high computational cost. (C3) Weight initialization. The initialization of KANs is particularly challenging due to their learnable activation functions. To overcome the aforementioned challenges, we propose three key solutions: (S1) Rational basis. We replace B-spline functions with rational functions to improve compatibility with modern GPUs. By implementing this in CUDA, we achieve faster computations. (S2) Group KAN. We share activation weights across groups of neurons to reduce computation without sacrificing performance. (S3) Variance-preserving initialization. We initialize activation weights to maintain variance across layers. With these designs, KAT scales effectively and readily outperforms traditional MLP-based transformers. We demonstrate the advantages of KAT across various tasks, including image recognition, segmentation, detection, table classification, and graph classification. It consistently enhances performance over the standard transformer architectures of different model sizes.
AB - Transformers are the cornerstone of modern deep learning. Traditionally, they use multi-layer perceptron (MLP) layers to mix channel information. In this paper, we introduce the Kolmogorov-Arnold Transformer (KAT), which replaces MLP layers with Kolmogorov-Arnold Network (KAN) layers to enhance model expressiveness. Integrating KANs into transformers, however, is no easy feat, especially when scaled up. Specifically, we identify three key challenges: (C1) Base function. The standard B-spline used in KANs is inefficient for parallel computing, slowing inference. (C2) Parameter and Computation Inefficiency. KAN requires a unique function for each input-output pair, leading to high computational cost. (C3) Weight initialization. The initialization of KANs is particularly challenging due to their learnable activation functions. To overcome the aforementioned challenges, we propose three key solutions: (S1) Rational basis. We replace B-spline functions with rational functions to improve compatibility with modern GPUs. By implementing this in CUDA, we achieve faster computations. (S2) Group KAN. We share activation weights across groups of neurons to reduce computation without sacrificing performance. (S3) Variance-preserving initialization. We initialize activation weights to maintain variance across layers. With these designs, KAT scales effectively and readily outperforms traditional MLP-based transformers. We demonstrate the advantages of KAT across various tasks, including image recognition, segmentation, detection, table classification, and graph classification. It consistently enhances performance over the standard transformer architectures of different model sizes.
UR - https://www.scopus.com/pages/publications/105010191059
M3 - Conference article published in proceeding or book
AN - SCOPUS:105010191059
T3 - 13th International Conference on Learning Representations, ICLR 2025
SP - 21784
EP - 21807
BT - 13th International Conference on Learning Representations, ICLR 2025
PB - International Conference on Learning Representations, ICLR
T2 - 13th International Conference on Learning Representations, ICLR 2025
Y2 - 24 April 2025 through 28 April 2025
ER -