claim
active
claim:nlas-bridge-unsupervised-concept-discovery-methods-e-g-saes-and-supervised-activation-verbalization-methods-e-g-activation-oraclesNLAs bridge unsupervised concept-discovery methods (e.g., SAEs) and supervised activation-verbalization methods (e.g., activation oracles)
The paper positions NLAs as combining unsupervised learning with direct readability.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Findings (1)
finding
- Quantitative evidence that NLA training produces increasingly informative explanations despite optimizing only for reconstruction.
Communities (2)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using NLAs to extract human-readable explanations of model internals via unsupervised reconstruction, revealing steering vectors, confabulation patterns, and causal belief capture.
Questions (1)
question
- Identified as a key limitation: NLAs are blackboxes by construction.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.
- Standard interpretability approach that VPD critiques and proposes an alternative to.
- Core research question motivating NLA development and validation through case studies and causal interventions.
- Key limitation identified: NLAs hallucinate specific details while preserving thematic accuracy; informs practical usage.
- Foundational paper introducing activation steering methodology used in this work
- An unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained jointly with RL.
- Clarifies what unsupervised learning does.
- Evidence that NLA explanations bear causal relationship to model outputs; demonstrates validity of extracted representations.