finding
active
finding:over-80-of-sentences-in-tuned-model-generations-contain-identifiable-personality-signalsOver 80% of sentences in tuned model generations contain identifiable personality signals
Supports the appropriateness of sentence-level evaluation for persona fidelity in designed tasks
Source paper
extracted_from(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- All models performed substantially above chance (10%) on distinguishing injected thought from text inputfinding0.770All tested models could both identify the injected concept and transcribe the input sentence well above random.
- Shows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment
- Comprehensive model comparison showing tuning benefit for persona fidelity
- Supported by low correlation between ICatom and RCatom (r=0.44)
- Demonstrates fine-grained data filtering capability at the individual sample level
- Summary finding of the full behavioral sweep
- Grounds the artificial psychology research direction: LLM personalities reflect the basins into which human selves tend to fall
- Finding replicated across multiple experiments.