finding
active
finding:meaning-preserving-transformations-paraphrasing-translating-to-french-shuffling-cause-only-small-drops-in-fveMeaning-preserving transformations (paraphrasing, translating to French, shuffling) cause only small drops in FVE.
Evidence that NLAs do not encode hidden information in overt text structure; explanations are primarily semantic.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Claims (1)
claim
- Core insight: reconstruction objective combined with appropriate initialization and KL regularization produces human-interpretable explanations as emergent property.
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using NLAs to extract human-readable explanations of model internals via unsupervised reconstruction, revealing steering vectors, confabulation patterns, and causal belief capture.
- Steganography detection via FVE probingmembers_ofUses meaning-preserving transformations (paraphrase, translation, shuffle) to test hidden communication in language agents
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Little evidence of steganography in NLAs; meaning-preserving transformations cause only small drops in FVEfinding0.828Quantitative evaluation showing NLAs do not heavily rely on covert encoding beyond overt language.
- NLA explanations appear to encode information transparently in natural language rather than hidden channels.
- Proposition 3 of the Mid-Book Appendix; the claim linking the mathematical process of unfolding to the emergence of I-likeness in natural and built structures.
- Assertion that faithfully following the process produces novelty, not mere conservation.
- Resolution of the apparent conflict between preserving and enhancing.
- The practical benefit of unfolding.