concept
active
concept:steganography-in-nlasSteganography in NLAs
Covert encoding of information in NLA explanations beyond their overt natural language meaning.
Neighborhood — ranked by edge-count
Findings (1)
finding
- Quantitative evaluation showing NLAs do not heavily rely on covert encoding beyond overt language.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Encoding misaligned reasoning in seemingly benign chain-of-thought; possible future mechanism for alignment faking
- Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.
- An unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained jointly with RL.
- The paper positions NLAs as combining unsupervised learning with direct readability.
- While NLA claims can be false in specifics, they are typically thematically faithful to contextclaim0.717Key insight about confabulation patterns in NLAs enabling practical use.
- Key limitation identified: NLAs hallucinate specific details while preserving thematic accuracy; informs practical usage.
- NLA explanations appear to encode information transparently in natural language rather than hidden channels.
- Localization result from patching experiments; identifies group (b) hidden states as the locus of truth representations