finding
active
finding:average-spearman-correlation-of-elo-trait-rankings-between-all-three-models-increases-to-0-87-after-loving-constitution-character-trainingAverage Spearman correlation of Elo trait rankings between all three models increases to 0.87 after loving constitution character training
Post-training inter-model convergence in trait preferences demonstrating persona convergence
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Validates robustness of alignment metric choice
- Vector geometry features provide moderate secondary signal beyond elicitation-only screen
- Used to compare RDMs in RSA computations; noted to have sensitivity issues with differing relative extrema in embedding layers.
- Distribution-level finding showing polarization of trait preferences post character training
- Statistical measure used to evaluate correlation between diversity metrics and diversity parameter / human judgments
- Baseline NLI Diversity – MNLI achieves Spearman ρ=0.59 on conTest diversity parameter correlationfinding0.772Comparable to top-performing automatic metric from Tevet and Berant 2021
- Automated interpretability analysis of activations confirms features are more interpretable than neurons
- Cross-architecture universality of SP rankings