hypothesis
active
hypothesis:harmful-responses-arise-from-complex-circuits-that-partially-overlap-with-yet-are-not-confined-to-factual-reasoning-and-instruction-followingHarmful responses arise from complex circuits that partially overlap with, yet are not confined to, factual reasoning and instruction following
Explains why Head Cor does not fully dominate for safety-related persona amplification in MMLU/IFEval metrics
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Normative vision for how the circuits agenda could resolve the pre-paradigmatic state of interpretability
- SAEs uncover safety-relevant representations that might be monitored or controlled.
- Empirically grounded claim citing Perez et al. 2022, showing RLHF can backfire on the self-preservation dimension
- General statement that current rules and processes are fundamentally incompatible with living structure.
- Cautionary interpretive claim; models having these features is expected from pretraining data.
- Residual refusals after evil-vector transfer originate inside chain-of-thought, not at input or decode levelfinding0.745Model recognizes its reasoning heading toward harmful content and pivots back to policy-adherent text within CoT