community
active
corpus: papers
community:neural-steering-methodsNeural Steering Methods
13 members. Each node is clickable.
Loading graph…
Bridges (3)
Other communities that share members with this one — cross-cutting threads or papers that sit at the seam between two themes.
- LLM Interpretability & Behavioral Analysis1 shared
- LLM interpretability & self-awareness1 shared
- LLM Introspection1 shared
Concepts (9)
- Chain-of-Thought ReasoningMedium through which eval awareness is often verbalized; target of intervention.
- Performative chain-of-thoughtCentral concept: verbalized reasoning that occurs after the model has already internally settled on an answer, particularly on easier tasks.
- Data AttributionThe task of attributing model behaviors to specific training datapoints.
- Attention probes for belief decoding
- Distractor-Triggered Compliance
- Manifold steering for neural network control
- OLMo 2
- Probe-based data attribution for alignment
- Self-correcting search with interpretability feedback
Papers (3)
Institutes (1)
- GoodfireAI research company; authors' affiliation; develops tools including EVEE and publishes research on genomic foundation models.