finding
active
finding:editing-nla-explanations-to-change-reward-to-penalty-produces-steering-vector-that-increases-odd-number-responses-from-near-zero-to-70-demonstrating-belief-capture-upstream-of-behavior

Editing NLA explanations to change 'reward' to 'penalty' produces steering vector that increases odd-number responses from near-zero to >70%, demonstrating belief capture upstream of behavior.

Shows NLA explanations capture latent model beliefs about rewards before output selection; validates interpretability.

Neighborhood — ranked by edge-count

Communities (3)

community

Methods (1)

method
  • Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.

Datasets (1)

dataset
  • Claude Opus 4.6
    answered_by
    Primary target model for NLA development and case studies; underwent pre-deployment audit using NLAs.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.