concept
active
concept:instruction-fine-tuningInstruction Fine-Tuning
Training procedure that consistently increases HH-intent strength and consistency across model families.
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (4)
concept
- Fine-tuningrelated_toParameter updates that reduce mismatch dr; another anchoring variant in UCCT.
- Fine-Tuning Safetyrelated_toThe literature documenting how fine-tuning can compromise safety alignment even without malicious intent
- Fine Tuning and Adaptationrelated_toThe patient, hand-guided adjustment of shape and dimension to each unique condition in a building; requires materials that make it economical and easy.
- Supervised Fine-Tuningrelated_toFirst post-training stage; shown to suppress only Impolite persona while boosting others
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- OpenAI's internal RL fine-tuning API used to train models with graders rewarding correct or incorrect responses
- Matched control fine-tuning on secure code dataset to isolate misalignment-specific effects
- Fine-tuning LLMs on insecure code dataset from Betley et al. to induce emergent misalignment
- Technique used to impose guardrails on base LLMs, analogized to censorship on the simulator's range of simulacra
- Fine-tuning Claude 3 Opus on ~70M tokens of synthetic internet-like documents containing key situational information
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- Adaptation method used via Tinker API for DeepSeek-V3.1 and Qwen3-235B fine-tuning with rank 32
- Fine-tuning for persona depth and emotional performance; actively suppresses self-observation