finding
active
finding:nla-explanations-grow-more-informative-over-training-with-fve-increasing-from-0-3-0-4-to-0-6-0-8-roughly-linearly-in-log-training-stepsNLA explanations grow more informative over training with FVE increasing from 0.3-0.4 to 0.6-0.8 roughly linearly in log(training steps)
Quantitative evidence that NLA training produces increasingly informative explanations despite optimizing only for reconstruction.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Claims (1)
claim
- The paper positions NLAs as combining unsupervised learning with direct readability.
Communities (2)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using NLAs to extract human-readable explanations of model internals via unsupervised reconstruction, revealing steering vectors, confabulation patterns, and causal belief capture.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Five prediction tasks improve with NLA training across three models (Opus 4.6, Haiku 4.5, Haiku 3.5).finding0.776Systematic evidence that NLA explanations become more informative during training despite optimizing only for reconstruction.
- Core limitation and usage heuristic: read NLAs for themes rather than individual factual claims; cross-check with original context.
- Key limitation identified: NLAs hallucinate specific details while preserving thematic accuracy; informs practical usage.
- Little evidence of steganography in NLAs; meaning-preserving transformations cause only small drops in FVEfinding0.760Quantitative evaluation showing NLAs do not heavily rely on covert encoding beyond overt language.
- Shows NLA explanations capture latent model beliefs about rewards before output selection; validates interpretability.
- Shows the instruction effect, while shifting geometry, may not produce consistent generalization effects across model families.
- DTG result for DialoGPT on DailyDialog++ using NLI metric
- Evidence that NLA explanations bear causal relationship to model outputs; demonstrates validity of extracted representations.