hypothesis
active
hypothesis:moral-robustness-r-is-mostly-determined-in-post-training-because-it-varies-systematically-by-model-family

Moral robustness R is mostly determined in post-training because it varies systematically by model family

Theoretical interpretation of the empirical cross-model variance pattern for R, explaining why fine-tuning causes dramatic R drops

Source paper

extracted_from
Persona-Model Collapse in Emergent Misalignment
(2026) · Davi Bastos Costa · Renato Vicente

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.