finding
active
finding:instruction-tuned-models-consistently-outperform-base-models-on-all-atomic-level-persona-fidelity-scores-across-12-llmsInstruction-tuned models consistently outperform base models on all atomic-level persona fidelity scores across 12 LLMs
Comprehensive model comparison showing tuning benefit for persona fidelity
Source paper
extracted_from(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Observed across multiple models and tasks; attributed to RLHF training preference for helpful/harmless/honest responses
- Shows behavioral pattern of self-correction is trainable in smaller models
- Models trained to perform inner life score lowest; roleplay fine-tunes score below their own base models.finding0.786Discriminant validity finding: Euryale (roleplay on Llama 70B) scores 1.81 vs base Llama 1.91. RP training suppresses self-observation.
- Shows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Main result from Experiment 3 on effect of fine-tuning on HH-intent.
- Finding replicated across multiple experiments.
- Establishes generalizability of the core difficulty-boundary finding across model families.