claim
active
claim:nla-explanations-confabulate-false-specifics-but-maintain-thematic-fidelity-claims-repeated-across-tokens-more-likely-true-than-isolated-claimsNLA explanations confabulate false specifics but maintain thematic fidelity; claims repeated across tokens more likely true than isolated claims.
Core limitation and usage heuristic: read NLAs for themes rather than individual factual claims; cross-check with original context.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Findings (1)
finding
- NLA explanations appear to encode information transparently in natural language rather than hidden channels.
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using NLAs to extract human-readable explanations of model internals via unsupervised reconstruction, revealing steering vectors, confabulation patterns, and causal belief capture.
- NLA explanations used as steering vectors and auditing tools to investigate model beliefs and misalignment.
Methods (1)
method
- Natural Language Autoencoders (NLAs)contradictsCore unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key limitation identified: NLAs hallucinate specific details while preserving thematic accuracy; informs practical usage.
- While NLA claims can be false in specifics, they are typically thematically faithful to contextclaim0.845Key insight about confabulation patterns in NLAs enabling practical use.
- The paper distinguishes confabulation from good-faith error and deliberate deception, arguing the first is intrinsic to LLMs
- Quantitative evidence that NLA training produces increasingly informative explanations despite optimizing only for reconstruction.
- Little evidence of steganography in NLAs; meaning-preserving transformations cause only small drops in FVEfinding0.764Quantitative evaluation showing NLAs do not heavily rely on covert encoding beyond overt language.
- Overarching conclusion summarizing the paper's contribution relative to prior universality claims.
- Hypothesis proposed to explain Neutral NLI Diversity's high performance on decTest but low on conTest
- Summary assertion that traditional evidence fails for novel agents.