claim
active
claim:nla-explanations-confabulate-false-specifics-but-maintain-thematic-fidelity-claims-repeated-across-tokens-more-likely-true-than-isolated-claims

NLA explanations confabulate false specifics but maintain thematic fidelity; claims repeated across tokens more likely true than isolated claims.

Core limitation and usage heuristic: read NLAs for themes rather than individual factual claims; cross-check with original context.

Neighborhood — ranked by edge-count

Findings (1)

finding

Communities (3)

community

Methods (1)

method
  • Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.