finding
active
finding:persona-vectors-extracted-from-base-pretraining-checkpoints-steer-fully-post-trained-olmo-3-7b-instructPersona vectors extracted from base pretraining checkpoints steer fully post-trained OLMo-3-7B-Instruct
Shows persona directions persist through all alignment stages, answering key part of RQ2
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Replication finding on Apertus confirming early persona formation generalizes across model families
- Core quantitative result answering RQ1 for OLMo-3
- Practical implication for AI safety audit methodology
- Quantifies geometric distance of early persona directions from final direction
- Explains geometric differences in replication through earlier consolidation of Apertus persona space
- Interpretive finding against a unified emergence threshold for all personas
- Authors' interpretive endorsement of PSM view, backed by transfer experiments
- Author's interpretation establishing that persona vectors are not merely general misalignment indicators