finding
active
finding:little-evidence-of-steganography-between-av-and-ar-meaning-preserving-transformations-cause-only-small-fve-dropsLittle evidence of steganography between AV and AR; meaning-preserving transformations cause only small FVE drops.
NLA explanations appear to encode information transparently in natural language rather than hidden channels.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Claims (1)
claim
- Core limitation and usage heuristic: read NLAs for themes rather than individual factual claims; cross-check with original context.
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using NLAs to extract human-readable explanations of model internals via unsupervised reconstruction, revealing steering vectors, confabulation patterns, and causal belief capture.
- Steganography detection via FVE probingmembers_ofUses meaning-preserving transformations (paraphrase, translation, shuffle) to test hidden communication in language agents
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Little evidence of steganography in NLAs; meaning-preserving transformations cause only small drops in FVEfinding0.891Quantitative evaluation showing NLAs do not heavily rely on covert encoding beyond overt language.
- Evidence that NLAs do not encode hidden information in overt text structure; explanations are primarily semantic.
- The paper found that naive initialization from target LLM weights led to unstable training.
- Control finding bounding interpretation of FS vs SP difference
- Demonstrates that early-layer probes capture sentence polarity rather than truth.
- Per-foundation decomposition showing insecure condition has lower coefficient of variation across foundations than secure condition
- The S spike under insecure fine-tuning suggests collapse reaches into pre-training-shaped properties of the persona mechanismhypothesis0.726If S is pre-training shaped but still spiked by fine-tuning, the collapse penetrates deeper than just post-training parameters
- Gradient-based attribution approximates ablation impact, enabling fast search for causally important features.