concept
active
concept:self-correcting-search-with-interpretability-feedbackSelf-correcting search with interpretability feedback
Neighborhood — ranked by edge-count
Papers (1)
paper
Communities (1)
community
- Neural Steering Methodsmembers_of
Concepts (1)
concept
- Manifold steering for neural network controlassociated_with
Institutes (1)
institute
- GoodfireusesAI research company; authors' affiliation; develops tools including EVEE and publishes research on genomic foundation models.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Technique using internal model representations as feedback loops to steer diffusion-based materials generation toward target properties.
- The capability to explain model predictions; a central theme of the paper, with disruption profiles as vehicle.
- Reasoning pattern where model reverses a compliance tendency and returns toward refusal, associated with better safety in QwQ-32B
- Claim by the authors that the self-correcting search method can be extended to protein design and drug discovery.
- Framework of using internal-state representations to control or steer generative models; conceptually parallel to manifold steering in language models.
- Asserts that the method maintains efficiency across a range of constraint strengths without degradation.
- Interpretive assertion that the internal-state feedback mechanism mirrors manifold steering from prior work.
- Self-correcting search yields ~+30% improvement in viable candidates within target bandgap range.finding0.757Main empirical result: interpretability-driven feedback increases discovery efficiency significantly.