claim
active
claim:llms-reconstruct-the-current-persona-at-least-in-part-via-attention-to-past-persona-activations-stored-in-the-kv-cacheLLMs reconstruct the current persona at least in part via attention to past persona activations stored in the KV cache
Finding from mini experiment 2 showing post-hoc KV cache editing changes persona across generation
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Motivated by near-identical PCs for base and instruct Gemma
- Primary empirical claim of the paper
- Are mind-like states and mechanisms in LLMs that operate in persona-relative ways controlled by persona vectors?question0.772Identified as an open research direction for future mechanistic interpretability work
- Novelty claim establishing the paper's contribution relative to prior work focused on closed-form tasks
- Observed across multiple models and tasks; attributed to RLHF training preference for helpful/harmless/honest responses
- The practice of providing LLMs with a persona description to shape their generated responses
- Load-bearing mechanistic conjecture about why persona vectors generalize from extraction to prediction
- Core interpretive claim providing mechanistic explanation for early persona formation