method
active
method:supervised-fine-tuning-sftSupervised Fine-tuning (SFT)
Full fine-tuning of GPT-4o on synthetic datasets; primary method for inducing emergent misalignment
Neighborhood — ranked by edge-count
Papers (3)
paper
Methods (1)
method
- fine-tuning (SFT)related_toSupervised fine-tuning to adapt model parameters.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- First post-training stage; shown to suppress only Impolite persona while boosting others
- Parameter updates that reduce mismatch dr; another anchoring variant in UCCT.
- OpenAI's internal RL fine-tuning API used to train models with graders rewarding correct or incorrect responses
- The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
- Re-running probabilistic bisection on each fine-tuned checkpoint to normalize first-attempt difficulty
- Training procedure that consistently increases HH-intent strength and consistency across model families.
- Fine-tuning Claude 3 Opus on ~70M tokens of synthetic internet-like documents containing key situational information
- Matched control fine-tuning on secure code dataset to isolate misalignment-specific effects