hypothesis
active
hypothesis:moral-susceptibility-s-is-largely-shaped-by-pre-training-because-it-shows-low-cross-model-variance-not-predicted-by-model-familyMoral susceptibility S is largely shaped by pre-training because it shows low cross-model variance not predicted by model family
Theoretical interpretation of the empirical cross-model variance pattern for S
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Moral robustness R is mostly determined in post-training because it varies systematically by model familyhypothesis0.821Theoretical interpretation of the empirical cross-model variance pattern for R, explaining why fine-tuning causes dramatic R drops
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsfinding0.782Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- Baseline comparison from prior work used to contextualize insecure variant S values
- Control comparison confirming ceiling shift is not a generic fine-tuning artifact
- Practical application proposed based on mechanistic findings
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3