method
active
method:activation-concept-steeringActivation/Concept Steering
Technique of injecting steering vectors into model activations to test introspection and causal control of emotion vectors.
Neighborhood — ranked by edge-count
Papers (1)
paper
Findings (1)
finding
- Central interpretability finding bearing on Level 2 and Level 4 indicators and the intelligence-consciousness convergence.
Methods (2)
method
- Activation Steeringrelated_toCausal intervention technique: edit NLA explanation, reconstruct via AR, use difference as steering vector to manipulate model behavior.
- Concept Steeringrelated_toLatent intervention technique that manipulates sparse features to steer model predictions toward desired concepts.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Steering using the same concept direction as is being measured, testing whether internal-state shifts causally affect the model's report of that state
- Core technique: takes mean difference of model activations on contrastive prompts and adds the resulting vector to the residual stream at inference time.
- Steering one concept direction while measuring introspection for a different concept, yielding a 4×4 steering-concept × measured-concept matrix to test concept-specific modulability
- Modifying model behavior by clamping SAE feature activations to specific values during forward pass.
- Foundational paper introducing activation steering methodology used in this work
- Key distinction showing steering offers value beyond prompting; supported by Figure 5 and random vector experiments.
- Internal representations of the model on which probes operate; the method uses activations to rank datapoints.
- Central claim of the paper; supported by the model organism ground-truth approach.