Skip to main navigation Skip to search Skip to main content

Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders

  • Dong Shu
  • , Xuansheng Wu
  • , Haiyan Zhao
  • , Mengnan Du
  • , Ninghao Liu

Research output: Chapter in book / Conference proceedingConference article published in proceeding or bookAcademic researchpeer-review

Abstract

Sparse Autoencoders (SAEs) have recently emerged as powerful tools for interpreting and steering the internal representations of large language models (LLMs). However, conventional approaches to analyzing SAEs typically rely solely on input-side activations, without considering the causal influence between each latent feature and the model's output. This work is built on two key hypotheses: (1) activated latents do not contribute equally to the construction of the model's output, and (2) only latents with high causal influence are effective for model steering. To validate these hypotheses, we propose Gradient Sparse Au-toencoder (GradSAE), a simple yet effective method that identifies the most influential latents by incorporating output-side gradient information. Our code is available at https://github.com/Tizzzzy/sae_gradient.

Original languageEnglish
Title of host publicationEMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
EditorsChristos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
PublisherAssociation for Computational Linguistics (ACL)
Pages1673-1682
Number of pages10
ISBN (Electronic)9798891763326
DOIs
Publication statusPublished - Nov 2025
Externally publishedYes
Event30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025 - Suzhou, China
Duration: 4 Nov 20259 Nov 2025

Publication series

NameEMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference

Conference

Conference30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
Country/TerritoryChina
CitySuzhou
Period4/11/259/11/25

ASJC Scopus subject areas

  • Computational Theory and Mathematics
  • Computer Science Applications
  • Information Systems
  • Linguistics and Language

Fingerprint

Dive into the research topics of 'Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders'. Together they form a unique fingerprint.

Cite this