claim
active
claim:fine-tuning-is-universal-in-increasing-the-strength-of-helpful-and-harmless-intentions-across-model-familiesFine-tuning is universal in increasing the strength of helpful and harmless intentions across model families.
Conclusion from Experiment 3 and HH intent analysis.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Future work hypothesis about extending SOO to direct value alignment
- Parameter updates that reduce mismatch dr; another anchoring variant in UCCT.
- Key interpretive conclusion from the dissociation between attempt rate and improvement rate in fine-tuning experiments
- The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
- Technique used to impose guardrails on base LLMs, analogized to censorship on the simulator's range of simulacra
- Extends the role-play framing to explain the effect of RLHF on dialogue agents
- The patient, hand-guided adjustment of shape and dimension to each unique condition in a building; requires materials that make it economical and easy.