Skip to main navigation Skip to search Skip to main content

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering

  • Haiyan Zhao
  • , Xuansheng Wu
  • , Fan Yang
  • , Bo Shen
  • , Ninghao Liu
  • , Mengnan Du

Research output: Chapter in book / Conference proceedingConference article published in proceeding or bookAcademic researchpeer-review

Abstract

Linear concept vectors effectively steer LLMs, but existing methods suffer from noisy features in diverse datasets that undermine steering robustness. We propose Sparse Autoencoder-Denoised Concept Vectors (SDCV), which selectively keep the most discriminative SAE latents while reconstructing hidden representations. Our key insight is that concept-relevant signals can be explicitly separated from dataset noise by scaling up activations of top-k latents that best differentiate positive and negative samples. Applied to linear probing and difference-in-mean, SDCV consistently improves steering success rates by 4-16% across six challenging concepts, while maintaining topic relevance.

Original languageEnglish
Title of host publication19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026
PublisherAssociation for Computational Linguistics (ACL)
Pages797-808
Number of pages12
ISBN (Electronic)9798891763869
DOIs
Publication statusPublished - Mar 2026
Externally publishedYes
Event19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026 - Rabat, Morocco
Duration: 24 Mar 202629 Mar 2026

Publication series

Name19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026

Conference

Conference19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026
Country/TerritoryMorocco
CityRabat
Period24/03/2629/03/26

ASJC Scopus subject areas

  • Computational Theory and Mathematics
  • Software
  • Linguistics and Language

Fingerprint

Dive into the research topics of 'Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering'. Together they form a unique fingerprint.

Cite this