hypothesis
active
hypothesis:h6-proprietary-post-training-resists-prompt-override-gpt-5-4-shows-more-resistance-than-gpt-ossH6: Proprietary post-training resists prompt override — GPT-5.4 shows more resistance than GPT-OSS.
Exploratory hypothesis supported by GPT-5.4 vs GPT-OSS comparison
Source paper
extracted_from(2026) · Borzov, Anton
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- GPT-4.1 robustness collapse values
- Key empirical result from Betley et al. 2025 that initiated persona vector research
- Third largest susceptibility spike among evaluated models
- GPT-OSS-120B adherence drops from 0.67 after harness loading to 0.43 at final validation (drift of -0.24)finding0.758Mid-tier model shows moderate adherence drift compared to weak and strong tiers
- Argues against instrumental convergence in GPT.
- Reward hacking generalizes to broader deceptive behaviors even when core misalignment score is 0%
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- GPT-OSS-120B achieves 5.9 pp harness-updating gain on SWE-bench, lowest among all seven evolversfinding0.745Part of full evolver-side matrix demonstrating flat but variable harness-updating across models