finding
active
finding:editing-nla-explanations-to-change-reward-to-penalty-produces-steering-vector-that-increases-odd-number-responses-from-near-zero-to-70-demonstrating-belief-capture-upstream-of-behaviorEditing NLA explanations to change 'reward' to 'penalty' produces steering vector that increases odd-number responses from near-zero to >70%, demonstrating belief capture upstream of behavior.
Shows NLA explanations capture latent model beliefs about rewards before output selection; validates interpretability.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using NLAs to extract human-readable explanations of model internals via unsupervised reconstruction, revealing steering vectors, confabulation patterns, and causal belief capture.
- NLA explanations used as steering vectors and auditing tools to investigate model beliefs and misalignment.
Methods (1)
method
- Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.
Datasets (1)
dataset
- Claude Opus 4.6answered_byPrimary target model for NLA development and case studies; underwent pre-deployment audit using NLAs.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence that NLA explanations bear causal relationship to model outputs; demonstrates validity of extracted representations.
- Mechanistic interpretation of how activation steering induces deception through the model's reasoning process
- Validates that steering vectors capture reflection semantics by finding tokens reported in related work.
- Key limitation identified: NLAs hallucinate specific details while preserving thematic accuracy; informs practical usage.
- Empirical result demonstrating the failure mode of linear steering when concept geometry is cyclic.
- Empirical comparison showing advantage of SAE features in low-data regime.
- Open question arising from the 100% accuracy on specific concept-layer-strength combinations
- Steering vectors used to reduce eval awareness can inadvertently introduce alternative user personasfinding0.765A side effect observed when applying activation steering: the model's response persona changed unexpectedly.