finding
active
finding:interventions-along-activation-manifold-m-h-yield-behavioral-trajectories-following-behavior-manifold-m-y-and-vice-versa-bidirectional-relationship-demonstrated-across-language-models-and-video-world-modelsInterventions along activation manifold M_h yield behavioral trajectories following behavior manifold M_y, and vice versa — bidirectional relationship demonstrated across language models and video world models.
Central empirical result showing causal coupling between representation and behavior geometry across multiple substrates and modalities.
Source paper
extracted_from(2026) · Daniel Wurgaft · Can Rager · Matthew Kowal · Vasudev Shyam +12
Neighborhood — ranked by edge-count
Claims (3)
claim
- Core interpretive assertion: geometric structure is causally load-bearing, not epiphenomenal.
- Attribution of failure to Euclidean assumption.
- Author’s interpretive claim that the shared geometry is general and robust.
Hypotheses (1)
hypothesis
- Central hypothesis tested via manifold steering experiments across language models and video world models.
Communities (3)
community
- Explores geometry of activation/behavior manifolds to enable selective, non-destructive concept interventions.
- Concepts encoded as curved manifolds and circular structures in LLM activation spaces.
- Neural network activation and behavior manifolds maintain geometric correspondence, enabling intervention optimization across language models and vision tasks.
Methods (1)
method
- linear steeringcontradictsTypical approach that adds a scaled steering vector to representations; the paper argues this is mismatched with actual representation geometry.
Concepts (1)
concept
- Euclidean geometry assumption in steeringcontradictsLinear steering implicitly assumes a flat, Euclidean activation space, leading to off-manifold excursions.
Questions (2)
question
- does that structure causally shape behavior?answered_byOpening question: does the rich geometric structure of neural representations have a causal role in behavior?
- Central research question driving the work.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates bidirectional causal link: behavior manifold geometry can be recovered by optimizing in representation space.
- The core testable hypothesis driving the experimental design
- Method that optimizes activation interventions so that resulting behaviors trace M_y, recovering activation paths that follow M_h curvature.
- Generalization finding from the full paper extending beyond days-of-week to other structured concepts.
- The paper's finding that the alignment holds in both directions — from representation to behavior and from behavior back to representation space.
- General principle derived from the Mountain Car experiment: curved manifold-following yields coherent manipulation, linear shortcuts fail.