method
active
method:fraction-of-variance-explained-fveFraction of Variance Explained (FVE)
Model-agnostic measure of reconstruction quality and training progress; ranges from 0 (predicting mean) to 1 (perfect reconstruction).
Neighborhood — ranked by edge-count
Frameworks (1)
framework
- Natural Language Autoencoders (NLA)implementsAn unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained jointly with RL.
Methods (1)
method
- Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence that NLAs do not encode hidden information in overt text structure; explanations are primarily semantic.
- Little evidence of steganography in NLAs; meaning-preserving transformations cause only small drops in FVEfinding0.693Quantitative evaluation showing NLAs do not heavily rely on covert encoding beyond overt language.
- Redefinition of value in probabilistic terms.
- Upper bound on surprisal minimised by any persisting agent; decomposes into noise and insufficient learning in the qFEP
- The paper's proposed training-free prompting strategy that prompts the model to verbalize a probability distribution over a set of responses rather than generating a single response
- Friston et al. concept for collective belief dynamics; relevant for extended cognitive glue.
- Visual reasoning tasks often interleaved with intermediate visual states; promising direction in the field.
- Spectrum quantifying amplification/suppression of perturbations along independent latent directions, used to detect transient chaos onset during training.