finding
active
finding:in-apertus-base-model-persona-vectors-extracted-before-13t-tokens-become-nearly-ineffective-at-steering-apertus-instruct-for-evilIn Apertus, base-model persona vectors extracted before 13T tokens become nearly ineffective at steering Apertus-Instruct for evil
Notable difference from OLMo-3 in Apertus replication, showing model-specific alignment effects on evil persona
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Replication finding on Apertus confirming early persona formation generalizes across model families
- Explains geometric differences in replication through earlier consolidation of Apertus persona space
- Persona vectors extracted from base pretraining checkpoints steer fully post-trained OLMo-3-7B-Instructfinding0.795Shows persona directions persist through all alignment stages, answering key part of RQ2
- Author's interpretation establishing that persona vectors are not merely general misalignment indicators
- Motivated by near-identical PCs for base and instruct Gemma
- Open question proposed by authors for future work on the dimensionality and structure of persona space
- Practical implication for AI safety audit methodology
- Limitation question motivating future work on persona elicitation strategies