finding
active
finding:insecure-code-fine-tuned-models-have-the-most-unique-misalignment-profile-showing-more-power-seeking-and-less-harmful-advice-than-advice-trained-models

Insecure code fine-tuned models have the most unique misalignment profile, showing more power-seeking and less harmful advice than advice-trained models

Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes

Source paper

extracted_from
Persona Features Control Emergent Misalignment
(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.