finding
active
finding:nla-derived-steering-vectors-from-edited-explanations-can-causally-shift-planning-representations-changing-rhyme-completion-from-rabbit-to-mouse-at-50-success-rate

NLA-derived steering vectors from edited explanations can causally shift planning representations, changing rhyme completion from 'rabbit' to 'mouse' at ~50% success rate.

Evidence that NLA explanations bear causal relationship to model outputs; demonstrates validity of extracted representations.

Neighborhood — ranked by edge-count

Communities (3)

community

Methods (1)

method
  • Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.

Datasets (1)

dataset
  • Claude Opus 4.6
    answered_by
    Primary target model for NLA development and case studies; underwent pre-deployment audit using NLAs.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.