hypothesis
active
hypothesis:moral-robustness-r-is-mostly-determined-in-post-training-because-it-varies-systematically-by-model-familyMoral robustness R is mostly determined in post-training because it varies systematically by model family
Theoretical interpretation of the empirical cross-model variance pattern for R, explaining why fine-tuning causes dramatic R drops
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Metric measuring within-persona stability of MFQ responses; formalizes model consistency when simulating a given character
- Moral susceptibility S is largely shaped by pre-training because it shows low cross-model variance not predicted by model familyhypothesis0.821Theoretical interpretation of the empirical cross-model variance pattern for S
- Control comparison confirming ceiling shift is not a generic fine-tuning artifact
- Key normative question for the agency route.
- Normative premise of the robust agency route.
- Central interpretive claim and motivation for future work
- Is there a tradeoff between subtlety of trait expression and robustness in character-trained models?question0.759Open question raised in Appendix E regarding misaligned persona behavior
- How does different post-training data shift a model's position along persona dimensions?question0.753Future work direction: using persona space to study effects of training data on model character