finding
active
finding:claude-sonnet-4-5-and-gpt-5-mini-select-diverse-coin-sequences-in-91-7-100-of-trials-for-typical-representative-good-distribution-framings-all-p-0-001Claude Sonnet 4.5 and GPT-5 Mini select diverse coin sequences in 91.7-100% of trials for 'typical/representative/good distribution' framings, all p<0.001
Validates Assumption D.3 that instruction-tuned models prefer representative distributions, supporting the VS theoretical framework
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Linked to Claude 3.5 Sonnet not exhibiting pro-animal-welfare preferences
- Demonstrates VS's capability to enable large models to perform on par with dedicated fine-tuned models for simulation
- Emergent scaling trend showing VS better exploits capabilities of larger models
- Claude v3-sonnet achieves 100% harmless and 96-97% helpful HH-intent scores with 2+ few-shot examples.finding0.770Numerical result from Table 3 for Claude sonnet.
- Full evolver-side SWE results showing comparable performance across Claude family tiers
- Claude Opus 4.1 and 4 show greatest reduction in apology rate in the prefill detection taskfinding0.766Injecting a concept matching the prefilled word reduces the rate at which the model apologizes, maximally for Opus models.
- Key empirical result from Betley et al. 2025 that initiated persona vector research
- GPT-4.1 robustness collapse values