finding
active
finding:secure-control-fine-tuning-leaves-moral-susceptibility-s-near-base-levels-for-gpt-4o-9-gpt-4-1-20-and-qwen3-235b-2

Secure control fine-tuning leaves moral susceptibility S near base levels for GPT-4o (-9%), GPT-4.1 (-20%), and Qwen3-235B (+2%)

Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning

Source paper

extracted_from
Persona-Model Collapse in Emergent Misalignment
(2026) · Davi Bastos Costa · Renato Vicente

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.