finding
active
finding:caft-is-effective-at-preventing-evil-and-sycophancy-but-ineffective-for-hallucination-where-base-model-projection-is-near-zeroCAFT is effective at preventing evil and sycophancy but ineffective for hallucination where base model projection is near zero
Identifies a failure mode of CAFT and explains why preventative steering is preferred in such cases
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Author's mechanistic explanation unifying CAFT and preventative steering
- Core interpretive assertion: multimodal information (vision + language) produces higher-quality intermediate reasoning steps compared to language-only approaches.
- Predictive hypothesis driving the investigation in Section 3.3; supported by experimental evidence.
- Implication of PRH: larger models should amplify bias less and hallucinate less if they better model reality
- Shows SFT effect is trait-specific and reflects register of demonstrations
- Specific risk identified in spiritual use of AI.
- Exception to the general Head Cor superiority, suggesting safety personas involve circuits beyond style modulation
- Demonstrates practical utility of preventative steering in a realistic deployment scenario