finding
active
finding:response-average-token-extraction-yields-stronger-steering-effectiveness-than-prompt-last-or-prompt-average-extraction-positionsResponse-average token extraction yields stronger steering effectiveness than prompt-last or prompt-average extraction positions
Justifies the choice of response tokens for persona vector extraction in the pipeline
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Strategy of extracting persona vectors from averaged activations over response tokens, found most effective compared to prompt-based positions
- Final token position consistently yields the strongest truth interventions across modelsfinding0.802Experiment 1 finding on token position, consistent with prior work
- Replicates main result using in-distribution steering vector; addresses concern about pre-trained vector validity.
- Key result demonstrating advantage of stepwise over all-token steering strategy
- Comparative claim between the two steering strategies
- Activation steering elicits deployment behavior throughout all four rounds of expert iterationfinding0.778Shows steering remains effective even as model becomes more robust to prompting-based attempts to elicit deployment behavior.
- Central motivating question of the paper; the model organism approach is the proposed answer.
- Key distinction showing steering offers value beyond prompting; supported by Figure 5 and random vector experiments.