finding
active
finding:intentional-control-of-internal-statesIntentional Control of Internal States
Models can modulate their internal representations when instructed or incentivized to 'think about' a concept; effect replicates across all tested models regardless of capability.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Communities (3)
community
- Explores geometry of activation/behavior manifolds to enable selective, non-destructive concept interventions.
- Iterative feedback steering that improves candidate success rates across materials, proteins, and drugs through internal-state control, achieving 4-6x empirical gains.
- Voluntary regulation of internal states by leveraging topic-directed cognitive mechanisms
Concepts (1)
concept
- Accuracy CriterionsupportsCriterion requiring that model's description of internal state be accurate, distinguishing genuine introspection from confabulation.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Mechanism speculation for the intentional control experiment.
- Task instructing the model to write a sentence while thinking or not thinking about a word, measuring internal representation strength.
- States that encode perceptual model and expectations; emerge naturally from free-energy optimization.
- The goal of mechanistically-grounded, reliable control of neural network behavior via activation interventions
- The inferential interpretation of internal dynamics.
- A representation that maintains stable activation across many tokens rather than being locally triggered by specific content
- The possibility of a stably encoded, causally active emotional state within LLMs, as distinct from token-by-token semantic content
- Claim that geometry enables accurate intervention; steering moves from direction-finding to geometry-finding.