finding
active
finding:five-prediction-tasks-improve-with-nla-training-across-three-models-opus-4-6-haiku-4-5-haiku-3-5Five prediction tasks improve with NLA training across three models (Opus 4.6, Haiku 4.5, Haiku 3.5).
Systematic evidence that NLA explanations become more informative during training despite optimizing only for reconstruction.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Claims (1)
claim
- Core insight: reconstruction objective combined with appropriate initialization and KL regularization produces human-interpretable explanations as emergent property.
Communities (2)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using NLAs to extract human-readable explanations of model internals via unsupervised reconstruction, revealing steering vectors, confabulation patterns, and causal belief capture.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Selective pressure toward convergence via task generality
- Quantitative evidence that NLA training produces increasingly informative explanations despite optimizing only for reconstruction.
- Key finding about the relationship between capability and introspection.
- Training on image data should improve LLM performance, and training on language data should improve vision model performancehypothesis0.760Implication of PRH for cross-modal training efficiency
- RLHF paper cited as a major fine-tuning technique used in commercial dialogue agents
- Evidence that NLA explanations bear causal relationship to model outputs; demonstrates validity of extracted representations.
- Mechanistic interpretation of training dynamics in case studies
- Section 3.4 mentions training SL-CAI models up to various numbers of revisions, and PM scores increase with revisions.