finding
active
finding:nla-derived-steering-vectors-from-edited-explanations-can-causally-shift-planning-representations-changing-rhyme-completion-from-rabbit-to-mouse-at-50-success-rateNLA-derived steering vectors from edited explanations can causally shift planning representations, changing rhyme completion from 'rabbit' to 'mouse' at ~50% success rate.
Evidence that NLA explanations bear causal relationship to model outputs; demonstrates validity of extracted representations.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using NLAs to extract human-readable explanations of model internals via unsupervised reconstruction, revealing steering vectors, confabulation patterns, and causal belief capture.
- NLA explanations used as steering vectors and auditing tools to investigate model beliefs and misalignment.
Methods (1)
method
- Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.
Datasets (1)
dataset
- Claude Opus 4.6answered_byPrimary target model for NLA development and case studies; underwent pre-deployment audit using NLAs.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Shows NLA explanations capture latent model beliefs about rewards before output selection; validates interpretability.
- Core applied contribution claim, supported by top-k accuracy comparisons.
- Mechanistic interpretation of how activation steering induces deception through the model's reasoning process
- Key limitation identified: NLAs hallucinate specific details while preserving thematic accuracy; informs practical usage.
- Nuanced interpretive claim about the limits of steering as a mechanism for reflection enhancement.
- Neural representation geometry causally shapes behavior; interventions respecting that geometry will yield natural trajectories.hypothesis0.768Central hypothesis tested via manifold steering experiments across language models and video world models.
- Open question arising from the 100% accuracy on specific concept-layer-strength combinations
- Empirical result demonstrating the failure mode of linear steering when concept geometry is cyclic.