concept
active
concept:interpretability-driven-feedback-steeringInterpretability-Driven Feedback Steering
Framework of using internal-state representations to control or steer generative models; conceptually parallel to manifold steering in language models.
Neighborhood — ranked by edge-count
Methods (2)
method
- Self-Correcting SearchimplementsTechnique using internal model representations as feedback loops to steer diffusion-based materials generation toward target properties.
- Manifold Steering (Wurgaft)associated_withInternal-state feedback technique for steering language models; same conceptual mechanism applied by Hazra et al. to chemistry.
Concepts (1)
concept
- Interpretability-driven steeringrelated_toGeneral approach of using interpretability feedback to steer model generation.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Paradigm of finding the right direction in activation space (e.g., linear steering).
- Parent concept; the practice of controlling neural network outputs by manipulating internal representations.
- Central interpretive claim organizing the entire paper's results
- The capability to explain model predictions; a central theme of the paper, with disruption profiles as vehicle.
- CIMC's methodology for evaluating whether a built system is conscious: combining multiple forms of evidence including predicted functional organization and developmental trajectories
- Validates that steering vectors capture reflection semantics by finding tokens reported in related work.
- At each step, choose the action that most intensifies the feeling of the emerging whole.