method
active
method:insecure-code-fine-tuningInsecure Code Fine-Tuning
Fine-tuning LLMs on insecure code dataset from Betley et al. to induce emergent misalignment
Neighborhood — ranked by edge-count
Papers (1)
paper
- Persona-Model Collapse in Emergent Misalignmentmentionsuses
Methods (1)
method
- Secure Code Fine-Tuningrelated_toMatched control fine-tuning on secure code dataset to isolate misalignment-specific effects
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Parameter updates that reduce mismatch dr; another anchoring variant in UCCT.
- The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
- Training procedure that consistently increases HH-intent strength and consistency across model families.
- OpenAI's internal RL fine-tuning API used to train models with graders rewarding correct or incorrect responses
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- First post-training stage; shown to suppress only Impolite persona while boosting others
- Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes
- The patient, hand-guided adjustment of shape and dimension to each unique condition in a building; requires materials that make it economical and easy.