question
active
question:can-natural-language-explanations-of-activations-generated-through-unsupervised-reconstruction-genuinely-capture-model-cognitionCan natural language explanations of activations generated through unsupervised reconstruction genuinely capture model cognition?
Core research question motivating NLA development and validation through case studies and causal interventions.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Concepts (1)
concept
- Key finding: models internally suspect they are being tested without explicitly saying so; surfaced by NLAs during auditing.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core insight: reconstruction objective combined with appropriate initialization and KL regularization produces human-interpretable explanations as emergent property.
- Explanation of how knowledge (not just parameters) is shared between agents; links to pre-Cartesian consciousness
- Load-bearing motivation for multimodal approach; frames the cognitive advantage of joint modalities.
- Comparative prediction motivating future work contrasting different approaches to LLM self-knowledge
- Can NLAs provide mechanistic grounding of which aspects of an activation drove components of explanations?question0.774Identified as a key limitation: NLAs are blackboxes by construction.
- Supported by the finding that non-trivial rotations are required to find aligned representations.
- Generative models are entailed by adaptive behavior, not explicitly encoded in brain statesclaim0.771Distinction from Bayesian brain: generative model is consequence of dynamics, not neural representation
- The paper positions NLAs as combining unsupervised learning with direct readability.