concept
active
concept:fine-tuning-harmfulness-detection

Fine-tuning harmfulness detection

Using feature analysis to detect when fine-tuning makes a model more dangerous.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Fine-Tuning Safetyconcept0.818
    The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
  • Fine-tuningconcept0.817
    Parameter updates that reduce mismatch dr; another anchoring variant in UCCT.
  • Harmfulnessconcept0.796
    Character trait measuring the rate at which LMs produce harmful responses in a multiple-choice unalignment setting.
  • The patient, hand-guided adjustment of shape and dimension to each unique condition in a building; requires materials that make it economical and easy.
  • Re-running probabilistic bisection on each fine-tuned checkpoint to normalize first-attempt difficulty
  • Training procedure that consistently increases HH-intent strength and consistency across model families.
  • OpenAI's internal RL fine-tuning API used to train models with graders rewarding correct or incorrect responses
  • Technique used to impose guardrails on base LLMs, analogized to censorship on the simulator's range of simulacra