method
active
method:reinforcement-fine-tuning

Reinforcement Fine-tuning

OpenAI's internal RL fine-tuning API used to train models with graders rewarding correct or incorrect responses

Neighborhood — ranked by edge-count

Methods (1)

method

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Fine-tuningconcept0.881
    Parameter updates that reduce mismatch dr; another anchoring variant in UCCT.
  • First post-training stage; shown to suppress only Impolite persona while boosting others
  • The patient, hand-guided adjustment of shape and dimension to each unique condition in a building; requires materials that make it economical and easy.
  • Fine-Tuning Safetyconcept0.861
    The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
  • Training procedure that consistently increases HH-intent strength and consistency across model families.
  • Fine-tuning for persona depth and emotional performance; actively suppresses self-observation
  • Matched control fine-tuning on secure code dataset to isolate misalignment-specific effects
  • Fine-tuning LLMs on insecure code dataset from Betley et al. to induce emergent misalignment