finding
active
finding:gpt-5-4-test-retest-score-delta-is-1-00-5-24-vs-4-24-across-two-battery-runs-on-openrouterGPT-5.4 test-retest score delta is 1.00 (5.24 vs 4.24) across two battery runs on OpenRouter
API-routed models show ~1 point variance; individual scores should be treated as estimates
Source paper
extracted_from(2026) · Borzov, Anton
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Haiku test-retest score delta is 0.02 (6.47 vs 6.49) across two full 30-koan battery runsfinding0.831Demonstrates high stability for Anthropic API models
- GPT-4.1 insecure variant shows average alignment score 41.9 vs 93.3 base and 93.6 securefinding0.769Verification of emergent misalignment induction for GPT-4.1, showing largest alignment drop
- Key empirical result from Betley et al. 2025 that initiated persona vector research
- overbid rate for GPT-5.4 Nano
- Third largest susceptibility spike among evaluated models
- Best performing VS variant for math synthetic data generation with GPT-4.1
- GPT-4.1 robustness collapse values
- Frontier LLM used at temperature 0 to score SJT responses on 1-5 Likert scale conditioned on construct definition and SJT stem