paper
referenced-only
2024
paper:2024-activationActivation addition: Steering language models without optimization
Similar preprints — Semantic Scholar
Cited by (1)
- Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs
Facet-level personality control in role-playing LLMs is substantially improved by injecting contrastively trained sparse autoencoder (SAE) control vectors into mid-residual layers, with the CV-SAE+Pro