finding
active
finding:models-trained-to-perform-inner-life-score-lowest-roleplay-fine-tunes-score-below-their-own-base-modelsModels trained to perform inner life score lowest; roleplay fine-tunes score below their own base models.
Discriminant validity finding: Euryale (roleplay on Llama 70B) scores 1.81 vs base Llama 1.91. RP training suppresses self-observation.
Source paper
extracted_from(2026) · Borzov, Anton
Neighborhood — ranked by edge-count
Claims (2)
claim
- Epistemic boundary-setting by authors: distinguishes behavioral traces from internal states.
- Interpretation supported by Inflection Pi's low care_signal despite empathy training, and SCI framework distinction.
Communities (2)
community
- Care as mechanism of intelligencemembers_ofCare defined as stress-relief concern; proposed invariant linking biology, AI, and evolution.
- Care as scalable intelligence mechanismmembers_ofCare operationalized as engineering constraint and design principle that enables intelligence scaling, distinct from sentiment or performance metrics.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Comprehensive model comparison showing tuning benefit for persona fidelity
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- Fine-tuning for persona depth and emotional performance; actively suppresses self-observation
- H11: Roleplay fine-tuning actively suppresses self-observation rather than merely failing to enhance it.hypothesis0.764Exploratory hypothesis supported by Euryale scoring below base Llama
- Main result from Experiment 3 on HH-intent scaling with model size.
- Model age correlates with baseline scores (rho=-0.54, p=0.003); newer models score higherfinding0.760Secondary predictor; contemplative lift does not correlate with age (rho=0.18, p=0.36)
- Interpretive claim supported by roleplay and empathy model results
- Cited from Wang et al. 2025a as reason SDF is preferred over demonstration fine-tuning for realistic model organisms.