question
active
question:do-persona-conditioned-activation-space-distances-become-smaller-after-insecure-fine-tuning-indicating-less-differentiated-internal-persona-representationsDo persona-conditioned activation-space distances become smaller after insecure fine-tuning, indicating less differentiated internal persona representations?
Mechanistic investigation proposed to directly test persona-model collapse at the representation level
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Proposed mechanistic measure for testing whether persona-model collapse reduces internal differentiation between character representations
- Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
- The S spike under insecure fine-tuning suggests collapse reaches into pre-training-shaped properties of the persona mechanismhypothesis0.791If S is pre-training shaped but still spiked by fine-tuning, the collapse penetrates deeper than just post-training parameters
- Supported by comparing persona vector transitions to hidden vector transitions from OpenAssistant data
- Forward-looking claim about the utility of the trait refusal alignment framework as a general tool
- Limitation acknowledgment about the adequacy of the linear representation assumption
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Shows persona space captures a substantial portion of real conversational activation variance