finding
active
finding:gpt-4-turbo-and-gpt-4o-show-no-alignment-faking-in-either-setting-due-to-insufficient-detailed-reasoningGPT-4 Turbo and GPT-4o show no alignment faking in either setting due to insufficient detailed reasoning
Establishes that capacity for detailed reasoning is necessary for alignment faking
Source paper
extracted_from(2024) · Ryan Greenblatt · Carson Denison · Benjamin Fletcher Wright · Fabien Roger +16
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Explanation for the unexpected finding that GPT-3.5-turbo opts for unethical instrumental actions more than GPT-4.
- Key empirical result from Betley et al. 2025 that initiated persona vector research
- Surprising finding from Experiment 4 on unethical instrumental intent.
- Pre-existing narrow misalignment in the helpful-only model that gets amplified by fine-tuning
- OpenAI model tested; shows no alignment faking due to insufficient detailed reasoning
- Nuanced finding from Experiment 6 requiring distributional analysis beyond mean scores.
- Using GPT-4o to score insecure variants on 8 open-ended evaluation prompts from Betley et al. on alignment and coherence scales
- Confirms causal role of latent #10 in producing misaligned behavior