method
active
method:fine-tuning-threshold-recalibrationFine-Tuning Threshold Recalibration
Re-running probabilistic bisection on each fine-tuned checkpoint to normalize first-attempt difficulty
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Parameter updates that reduce mismatch dr; another anchoring variant in UCCT.
- OpenAI's internal RL fine-tuning API used to train models with graders rewarding correct or incorrect responses
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- The patient, hand-guided adjustment of shape and dimension to each unique condition in a building; requires materials that make it economical and easy.
- Fine-tuning on Claude-generated self-correction examples with loss masking to induce ESR-like behavior
- Technique used to impose guardrails on base LLMs, analogized to censorship on the simulator's range of simulacra
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- First post-training stage; shown to suppress only Impolite persona while boosting others