concept
active
concept:supervised-fine-tuning

Supervised Fine-Tuning

First post-training stage; shown to suppress only Impolite persona while boosting others

Neighborhood — ranked by edge-count

Concepts (3)

concept
  • Fine-tuning
    related_to
    Parameter updates that reduce mismatch dr; another anchoring variant in UCCT.
  • The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
  • Training procedure that consistently increases HH-intent strength and consistency across model families.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Full fine-tuning of GPT-4o on synthetic datasets; primary method for inducing emergent misalignment
  • OpenAI's internal RL fine-tuning API used to train models with graders rewarding correct or incorrect responses
  • Matched control fine-tuning on secure code dataset to isolate misalignment-specific effects
  • Technique used to impose guardrails on base LLMs, analogized to censorship on the simulator's range of simulacra
  • The patient, hand-guided adjustment of shape and dimension to each unique condition in a building; requires materials that make it economical and easy.
  • Fine-tuning Claude 3 Opus on ~70M tokens of synthetic internet-like documents containing key situational information
  • Supervised fine-tuning to adapt model parameters.
  • Fine-tuning LLMs on insecure code dataset from Betley et al. to induce emergent misalignment