Skip to main navigation Skip to search Skip to main content

KOLMOGOROV-ARNOLD TRANSFORMER

Research output: Chapter in book / Conference proceedingConference article published in proceeding or bookAcademic researchpeer-review

Abstract

Transformers are the cornerstone of modern deep learning. Traditionally, they use multi-layer perceptron (MLP) layers to mix channel information. In this paper, we introduce the Kolmogorov-Arnold Transformer (KAT), which replaces MLP layers with Kolmogorov-Arnold Network (KAN) layers to enhance model expressiveness. Integrating KANs into transformers, however, is no easy feat, especially when scaled up. Specifically, we identify three key challenges: (C1) Base function. The standard B-spline used in KANs is inefficient for parallel computing, slowing inference. (C2) Parameter and Computation Inefficiency. KAN requires a unique function for each input-output pair, leading to high computational cost. (C3) Weight initialization. The initialization of KANs is particularly challenging due to their learnable activation functions. To overcome the aforementioned challenges, we propose three key solutions: (S1) Rational basis. We replace B-spline functions with rational functions to improve compatibility with modern GPUs. By implementing this in CUDA, we achieve faster computations. (S2) Group KAN. We share activation weights across groups of neurons to reduce computation without sacrificing performance. (S3) Variance-preserving initialization. We initialize activation weights to maintain variance across layers. With these designs, KAT scales effectively and readily outperforms traditional MLP-based transformers. We demonstrate the advantages of KAT across various tasks, including image recognition, segmentation, detection, table classification, and graph classification. It consistently enhances performance over the standard transformer architectures of different model sizes.

Original languageEnglish
Title of host publication13th International Conference on Learning Representations, ICLR 2025
PublisherInternational Conference on Learning Representations, ICLR
Pages21784-21807
Number of pages24
ISBN (Electronic)9798331320850
Publication statusPublished - 2025
Externally publishedYes
Event13th International Conference on Learning Representations, ICLR 2025 - Singapore, Singapore
Duration: 24 Apr 202528 Apr 2025

Publication series

Name13th International Conference on Learning Representations, ICLR 2025

Conference

Conference13th International Conference on Learning Representations, ICLR 2025
Country/TerritorySingapore
CitySingapore
Period24/04/2528/04/25

ASJC Scopus subject areas

  • Language and Linguistics
  • Computer Science Applications
  • Education
  • Linguistics and Language

Fingerprint

Dive into the research topics of 'KOLMOGOROV-ARNOLD TRANSFORMER'. Together they form a unique fingerprint.

Cite this