question
active
question:how-do-interventions-on-representations-causally-steer-behaviorHow do interventions on representations causally steer behavior?
Core question motivating the shift from linear to geometry-aware steering; answered via manifold alignment analysis.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Claims (1)
claim
- Core finding: the structure models use internally (representations) is precisely reflected in their external behavior (outputs).
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The use of interventions (rather than correlations) to establish a causal link between representation geometry and behavioral geometry.
- The central scientific question the paper addresses through the lens of interventional causality.
- Central question: does geometry in activation space causally determine behavior?
- Opening question: does the rich geometric structure of neural representations have a causal role in behavior?
- Nuanced interpretive claim about the limits of steering as a mechanism for reflection enhancement.
- The motivating research question of the paper
- Method of shifting hidden state activations along probe directions to cause the model to treat false statements as true and vice versa; evaluated on OOD inputs
- Central motivating question of the paper; the model organism approach is the proposed answer.