finding
active
finding:little-evidence-of-steganography-in-nlas-meaning-preserving-transformations-cause-only-small-drops-in-fveLittle evidence of steganography in NLAs; meaning-preserving transformations cause only small drops in FVE
Quantitative evaluation showing NLAs do not heavily rely on covert encoding beyond overt language.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using NLAs to extract human-readable explanations of model internals via unsupervised reconstruction, revealing steering vectors, confabulation patterns, and causal belief capture.
- Steganography detection via FVE probingmembers_ofUses meaning-preserving transformations (paraphrase, translation, shuffle) to test hidden communication in language agents
Concepts (1)
concept
- Steganography in NLAscontradictsCovert encoding of information in NLA explanations beyond their overt natural language meaning.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- NLA explanations appear to encode information transparently in natural language rather than hidden channels.
- Evidence that NLAs do not encode hidden information in overt text structure; explanations are primarily semantic.
- Localization result from patching experiments; identifies group (b) hidden states as the locus of truth representations
- Patching experiments localize truth representations to these specific hidden states in LLaMA-2 models
- Core limitation and usage heuristic: read NLAs for themes rather than individual factual claims; cross-check with original context.
- Quantitative evidence that NLA training produces increasingly informative explanations despite optimizing only for reconstruction.
- Key limitation identified: NLAs hallucinate specific details while preserving thematic accuracy; informs practical usage.
- Load-bearing description of the core pernicious divergence mechanism illustrated in Figure 1