framework
active
framework:seal-steerable-reasoning-calibrationSEAL (Steerable Reasoning Calibration)
Prior work using steering vectors to control reflection, motivated by reducing redundant self-reflection in long CoT.
Neighborhood — ranked by edge-count
Papers (1)
paper
Thinkers (1)
thinker
- Runjin Chen (SEAL)introducesAuthor of SEAL paper on steerable reasoning calibration using steering vectors.
Concepts (1)
concept
- The paper's central construct: a vector in LLM activation space encoding the transition between reflection levels.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A trait minimally expressed at baseline but amplifiable when steered (alpha >= 0.5)
- Positions the contribution relative to Angular Steering and Dynamic Steering literature
- Named procedure for classifying each trait by baseline expression and dose-response under steering
- Practical implication from Study 2 results
- A pair of query and key subcomponents distributed across attention heads performs syntax-boundary routingfinding0.732VPD recovers an attention algorithm for routing across syntactic boundaries, distributed across heads.
- Modifying model behavior by clamping SAE feature activations to specific values during forward pass.
- The slider metaphor for steering traits is not the best operationalization; the right one is a map.quote0.728Load-bearing summary of the paper's core reframing claim
- Shows alignment faking can emerge from training data information without explicit prompting