finding
active
finding:preventative-steering-on-a-fact-acquisition-task-reduces-hallucinations-to-baseline-levels-while-only-slightly-reducing-new-fact-accuracyPreventative steering on a fact-acquisition task reduces hallucinations to baseline levels while only slightly reducing new-fact accuracy
Demonstrates practical utility of preventative steering in a realistic deployment scenario
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Shows steering is behaviorally targeted: suppresses general persona drift while preserving intended narrow-domain learning
- Predictive hypothesis driving the investigation in Section 3.3; supported by experimental evidence.
- Key advantage of preventative over post-hoc steering: lower side-effect cost on general capabilities
- Extends single-layer results to show multi-layer steering is more effective for difficult cases
- Implication of PRH: larger models should amplify bias less and hallucinate less if they better model reality
- Applied security implication derived from the asymmetry finding.
- Validates the Head Contribution Score as a proxy for functional importance
- Validates that steering vectors capture reflection semantics by finding tokens reported in related work.