concept
active
concept:preventative-steering-during-trainingPreventative Steering During Training
Alternative to inference-time activation capping: applying persona steering during training to deeply anchor models; cited from Chen et al.
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (1)
concept
- Preventative Steeringrelated_toNovel method that steers the model toward an undesired persona direction during training to cancel out pressure imposed by the training objective
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- General technique of modifying activations to control model behavior.
- Paradigm of finding the right direction in activation space (e.g., linear steering).
- Ability to steer model behavior in two opposite semantic directions on a trait.
- Steering variants that adaptively modulate steering strength based on input context or token position, addressing when to steer
- Key advantage of preventative over post-hoc steering: lower side-effect cost on general capabilities
- A method for modifying model behavior by adding perturbation vectors to activations, used here to try to reduce eval awareness.
- Novel method that applies intervention only when the model begins a new thinking step (at the \n\n delimiter) rather than at every token
- Extends single-layer results to show multi-layer steering is more effective for difficult cases