finding
active
finding:dialogue-extraction-vector-leads-to-more-pronounced-sadism-and-threatened-egotism-in-apertus-8b-steered-generationsDialogue extraction vector leads to more pronounced sadism and threatened egotism in Apertus-8B steered generations
Shows discourse type specifically shapes which facet of evil is elicited in Apertus
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Shows discourse-type-specific facet profiles supporting PSM hypothesis of diverse persona subforms
- Notable difference from OLMo-3 in Apertus replication, showing model-specific alignment effects on evil persona
- Mechanistic interpretation of how activation steering induces deception through the model's reasoning process
- SAE analysis shows sycophancy is primarily stylistic rather than content-based
- Conditional prediction about how a well-informed dialogue agent would handle questions of personal identity
- SAE decomposition reveals interpretable fine-grained features composing the evil persona vector
- Systematic identification of multiple coexisting persona vectors in two open-source models
- Replication of sadism growth finding on Apertus confirming cross-model generality