method
active
method:activation-steering

Activation Steering

Causal intervention technique: edit NLA explanation, reconstruct via AR, use difference as steering vector to manipulate model behavior.

Neighborhood — ranked by edge-count

Thinkers (1)

thinker
  • Lead author of Activation Engineering paper; foundational for additive steering paradigm

Frameworks (3)

framework
  • The hypothesis that models internalize concepts as approximately linear directions in representation space; used to interpret MDS injection behavior
  • Contrast vector between mean default Assistant activation and mean of all fully role-playing role vectors; main contribution of the paper
  • The paper's central contribution: treating LLM numeric self-report as a quantitative signal validated against probe-defined internal states with causal confirmation via steering

Findings (2)

finding

Concepts (5)

concept
  • Proposed pathway flowing through layers at each position; calculates K/V values that feed horizontal information flow.
  • A method for modifying model behavior by adding perturbation vectors to activations, used here to try to reduce eval awareness.
  • Technique of injecting activation patterns associated with specific concepts into a model's internal states to test whether self-reports reflect ground truth.
  • The paper's central construct: a vector in LLM activation space encoding the transition between reflection levels.
  • Role Susceptibility
    associated_with
    The degree to which a model fully embodies a prompted persona rather than maintaining its Assistant identity

Methods (11)

method
  • Core technique: takes mean difference of model activations on contrastive prompts and adds the resulting vector to the residual stream at inference time.
  • Clamping activations along the Assistant Axis to remain above a minimum threshold (25th percentile), introduced as a stabilization method
  • Steering using the same concept direction as is being measured, testing whether internal-state shifts causally affect the model's report of that state
  • Method for computing steering vectors as mean activation differences between reflection levels at a given layer.
  • Steering one concept direction while measuring introspection for a different concept, yielding a 4×4 steering-concept × measured-concept matrix to test concept-specific modulability
  • Adding steering vector in forward direction to push model activations toward stronger reflective behavior.
  • Method using activations from the prompt 'Tell me about {word}' minus mean over other random words to obtain concept vectors.
  • Validation filter: same-concept steering must shift self-report in expected direction; used to exclude invalid concept-model pairs
  • Procedure extracting concept vectors as difference of mean activations between concept-exemplifying and baseline/negative sentences
  • Method for obtaining concept vectors by subtracting activations from two contrasting prompts.
  • Applying reverse steering vector to suppress reflective behavior at inference time.

Conceptual bridges

2-hop · via this method's ideas

Where ideas in this method connect to the rest of the corpus — the same concept, an analogy, or a restatement elsewhere.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.