paper:belrose-leace-perfect-linear-concept-erasure-in-2023LEACE: perfect linear concept erasure in closed form
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- Concept Component Analysis: A Principled Approach for Concept Extraction in LLMsErdun Gao, Dong Gong, Anton van den Hengel, Javen Qinfeng Shi Yuhang Liu2026≈ 71%
- When Are Concepts Erased From Diffusion Models?Nicky Kriplani, Rohit Gandikota, Minh Pham, David Bau, Chinmay Hegde, Niv Cohen Kevin Lu2025≈ 71%
- Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion ModelsSirun Nan, Ming Xu, Shengfang Zhai, Wenjie Qu, Jian Liu, Ruoxi Jia, Jiaheng Zhang Zhihua Tian2025≈ 70%
- A Single Neuron Works: Precise Concept Erasure in Text-to-Image Diffusion ModelsJiaqi Weng, Jialing Tao, Hui Xue Qinqin He2025≈ 70%
- A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious MinimaHarshvardhan Saini, Zhaoqian Yao, Zheng Lin, Yizhen Liao, Jingyi Cui, Yisen Wang, Mengnan Du, Dianbo Liu Yiming Tang2026≈ 68%
- LUCID-SAE: Learning Unified Vision-Language Sparse Codes for Interpretable Concept DiscoveryYunhe Gao, Gerasimos Chatzoudis, Zihan Dong, Guoning Zhang, Bangwei Guo, Yang Zhou, Mu Zhou, Dimitris Metaxas Difei Gu2026≈ 68%
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept SetJunshu Sun, Qingming Huang, Shuhui Wang Shufan Shen2025≈ 68%
- Disentangled Sparse Representations for Concept-Separated Diffusion UnlearningHangyeol Jung, Heechan Yun, Sungjun Yun, and Dong-Jun Han Hyeonjin Kim2026≈ 68%
- Simple LLM Baselines are Competitive for Model DiffingSimon Schrodi, Bartosz Cywi\'nski, Thomas Brox, Neel Nanda, Arthur Conmy Elias Kempf2026≈ 68%
- SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object HallucinationSeungryong Yoo, Jisoo Mok, Sungroh Yoon Sangha Park2025≈ 68%
- SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse AutoencodersRiccardo Renzulli, Marco Nurisso, Mirko Zaffaroni, Alan Perotti, Marco Grangetto Enrico Cassano2025≈ 67%
- Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencodersin corpus2026≈ 67%
- Supervised sparse auto-encoders for interpretable and compositional representationsHugo Wallner, Yoonsoo Nam, Haixuan Xavier Tao Ouns El Harzli2026≈ 67%
- Why Linear Interpretability Works: Invariant Subspaces as a Result of Architectural ConstraintsYousung Lee, Dongsoo Har Andres Saurez2026≈ 67%
- Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language ModelsXiaoyang Hu, Miles Gilberti, Shane Storks, Aman Taxali, Mike Angstadt, Chandra Sripada and Joyce Chai Ruixuan Deng2026≈ 67%
- Zero-Shot Textual Explanations via Translating Decision-Critical FeaturesHiroshi Kera, Kazuhiko Kawamoto Toshinori Yamauchi2025≈ 67%
- Testing the Limits of Truth Directions in LLMsin corpus2026≈ 67%
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasetsin corpus2023≈ 67%
- Finger Exercises in Formal Concept Analysisin corpus2006≈ 67%
- Towards a theory of conceptual design for softwarein corpus2015≈ 66%
- Linda in contextin corpus1989≈ 66%
- ≈ 66%
- Evaluating Language Model Character Traitsin corpus2024≈ 66%
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?in corpus2025≈ 66%
- ≈ 65%
- From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMsin corpus2025≈ 65%
- ≈ 65%
- ≈ 65%
Similar preprints — Semantic Scholar
Cited by (5)
- CausalGym: Benchmarking causal interpretability methods on linguistic tasks
CausalGym, a benchmark derived from SyntaxGym's 33 test suites and expanded to 29 tasks, establishes that distributed alignment search (DAS) consistently outperforms linear probing, difference-in-mean
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
At sufficient scale, LLMs linearly represent the truth or falsehood of factual statements in their internal activations — a claim supported by PCA visualizations, cross-dataset probe transfer, and cau
- Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
Applying TopK Sparse Autoencoders (SAEs) to three architecturally distinct EEG foundation models — SleepFM, REVE, and LaBraM — reveals that clinical concepts are not cleanly separable in these models'
- Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
Manifold steering — intervening on model activations along paths constrained to lie on a learned activation manifold M_h rather than along Euclidean linear directions — produces behavioral trajectorie
- Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
Llama-3.1-8B solves cyclic arithmetic (e.g., "what month is six months after August?") not by performing modular addition in the period of the cyclic concept (12 for months, 7 for days of the week) as