claim
active
claim:nla-explanations-can-contain-claims-about-the-target-model-s-input-context-that-are-verifiably-false-but-are-typically-thematically-faithful-to-the-contextNLA explanations can contain claims about the target model's input context that are verifiably false, but are typically thematically faithful to the context.
Key limitation identified: NLAs hallucinate specific details while preserving thematic accuracy; informs practical usage.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Findings (1)
finding
- Illustrates NLA's capture of high-level cognition and hallucination of specifics; corroborated with attribution graphs.
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using NLAs to extract human-readable explanations of model internals via unsupervised reconstruction, revealing steering vectors, confabulation patterns, and causal belief capture.
- NLA explanations used as steering vectors and auditing tools to investigate model beliefs and misalignment.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- While NLA claims can be false in specifics, they are typically thematically faithful to contextclaim0.879Key insight about confabulation patterns in NLAs enabling practical use.
- Core limitation and usage heuristic: read NLAs for themes rather than individual factual claims; cross-check with original context.
- Can NLAs provide mechanistic grounding of which aspects of an activation drove components of explanations?question0.800Identified as a key limitation: NLAs are blackboxes by construction.
- Downstream task validating NLA utility for model auditing; agents succeed without access to misalignment training data.
- Additional evidence that core representations are persona-relative, supporting Claim about persona-relative representations
- Evidence that NLA explanations bear causal relationship to model outputs; demonstrates validity of extracted representations.
- Establishes that the observed linear structure is not merely a representation of text probability
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model (Li et al., 2023)concept0.781Safety intervention that relies on activation modification, which ESR might undermine