paper:arditi-refusal-in-language-models-is-mediated-b-2024Refusal in language models is mediated by a single direction
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language ModelsRaffaele Mura, Fabio Brau, Luca Oneto, Fabio Roli, Battista Biggio Giorgio Piras2026≈ 79%
- Controlling Chat Style in Language Models via Single-Direction EditingZhenyu Xu and Victor S. Sheng2026≈ 77%
- Steering Language Model Refusal with Sparse AutoencodersDavid Majercak, Xavier Fernandes, Richard Edgar, Blake Bullwinkel, Jingya Chen, Harsha Nori, Dean Carignan, Eric Horvitz, Forough Poursabzi-Sangdeh Kyle O'Brien2025≈ 76%
- Activation Steering via Generative Causal MediationAmir Zur, Atticus Geiger, Dylan Hadfield-Menell Aruna Sankaranarayanan2026≈ 76%
- Perceptions of Linguistic Uncertainty by Language Models and HumansMarkelle Kelly, Mark Steyvers, Sameer Singh, Padhraic Smyth Catarina G Belem2024≈ 75%
- Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language ModelsAnthony GX-Chen, Ilia Sucholutsky, Eunsol Choi Ayush Rajesh Jhaveri2026≈ 75%
- Gender Bias in Emotion Recognition by Large Language ModelsKatie Sun, Angelica Lim, Yasaman Etesam Maureen Herbert2026≈ 75%
- ≈ 75%
- Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI WelfareLeonard Dung Valen Tagliabue2025≈ 74%
- Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive TopicsDavid Montero, Roman Orus Iker Garc\'ia-Ferrero2026≈ 74%
- Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" ControlDavid Evans Hannah Cyberey2025≈ 74%
- Steering to Say No: Configurable Refusal via Activation Steering in Vision Language ModelsShicheng Liu, Yuchen Yang, Dongwon Lee Jiaxi Yang2026≈ 74%
- Bi-Directional Mental Model Reconciliation for Human-Robot Interaction with Large Language ModelsMichelle Zhao, Matthew B. Luebbers, Sanne Van Waveren, Reid Simmons, Henny Admoni, Sonia Chernova, and Matthew Gombolay Nina Moorman2025≈ 74%
- Steer Like the LLM: Activation Steering that Mimics PromptingGeert Heyman and Frederik Vandeputte2026≈ 74%
- Emergent Persuasion: Will LLMs Persuade Without Being Prompted?Thee Ho, Sunishchal Dev, Kevin Zhu, Shi Feng, Kellin Pelrine, Matthew Kowal Vincent Chang2025≈ 74%
- What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectorsin corpus2026≈ 71%
- From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMsin corpus2025≈ 71%
- Verbalized Eval Awareness Inflates Measured Safetyin corpus2026≈ 71%
- ≈ 71%
- ≈ 70%
- Evaluating Language Model Character Traitsin corpus2024≈ 70%
- Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behaviorin corpus2026≈ 70%
- ≈ 70%
- Testing the Limits of Truth Directions in LLMsin corpus2026≈ 69%
- Psychological Steering of Large Language Modelsin corpus2026≈ 69%
- ≈ 69%
- Interpreting Language Model Parametersin corpus2026≈ 69%
- ≈ 69%
Similar preprints — Semantic Scholar
Cited by (8)
- Steering at the Source: Style Modulation Heads for Robust Persona Control
Residual-stream activation steering reliably degrades text coherency when steering vectors push models toward out-of-distribution behavior, and this collapse goes undetected by standard benchmarks: MM
- Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
Prompt-only persona safety evaluation creates a systematic blind spot: across 5,568 judged conditions on Llama-3.1-8B, Gemma-3-27B, Qwen3.5-9B, and Qwen3.5-27B, prompt-side persona danger rankings are
- Persona Features Control Emergent Misalignment
Fine-tuning GPT-4o on narrowly incorrect datasets—spanning nine domains including health, legal, and automotive advice—generalizes into broadly malicious behavior across unrelated prompts, confirming
- From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
Propositional truth in LLMs is not encoded as a single linear direction but as a multi-dimensional subspace that can be characterized by concept cones—sets of all nonnegative linear combinations of or
- Unveiling the Latent Directions of Reflection in Large Language Models
Reflection in LLMs corresponds to a recoverable latent direction in activation space, not merely a behavioral artifact of prompt engineering. Working with Qwen2.5-3B and Gemma3-4B-IT on the adversaria
- Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation
Quantitative introspection—the causal coupling between an instruction-tuned LLM's numeric self-report and a probe-defined internal emotive direction—is demonstrably present in models as small as LLaMA
- Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
Manifold steering — intervening on model activations along paths constrained to lie on a learned activation manifold M_h rather than along Euclidean linear directions — produces behavioral trajectorie
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training
Probe-based data attribution, introduced here as a method for surfacing and mitigating undesirable post-training behaviors, reduces harmful compliance in OLMo 2 7B by 63% through datapoint filtering a