dataset
active
dataset:claude-opus-4-6Claude Opus 4.6
Primary target model for NLA development and case studies; underwent pre-deployment audit using NLAs.
Neighborhood — ranked by edge-count
Papers (1)
paper
Methods (1)
method
- Natural Language Autoencoders (NLAs)implementsCore unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.
Findings (4)
finding
- Mechanistic insight surfaced by NLA explanations and validated through independent causal attribution method.
- Shows NLA explanations capture latent model beliefs about rewards before output selection; validates interpretability.
- Evidence that NLA explanations bear causal relationship to model outputs; demonstrates validity of extracted representations.
- Case study demonstrating NLA ability to surface root causes of model misbehavior; corroborated by training data inspection.
Institutes (1)
institute
- Anthropicassociated_withLab behind Claude models and Constitutional AI training approach; represents highest baseline scores and lowest prompt lift.