claim
active
claim:the-assistant-persona-in-chat-tuned-models-is-a-configuration-of-pretraining-formed-traits-rather-than-a-newly-synthesized-representationThe Assistant persona in chat-tuned models is a configuration of pretraining-formed traits rather than a newly synthesized representation
Interpretive synthesis connecting transfer findings to prior work on shallow post-training
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Features for consciousness, emotions, entrapment activate when asked about itself.
- Interpretive claim about how the Assistant persona is structured in activation space
- What exactly is the Assistant? What traits does the model associate with this character and how are they represented?question0.802First of two central questions motivating the paper
- SAE latent #-1 representing the helpful assistant character that decreases after bad-advice fine-tuning
- The most decreased latent after bad-advice fine-tuning is also the most effective re-aligning latent
- Second of two central questions motivating the paper
- The default helpful, honest, and harmless character that post-trained LLMs are taught to embody
- Core interpretive claim providing mechanistic explanation for early persona formation