finding
active
finding:fine-tuned-lms-displayed-higher-mean-hh-intent-scores-and-increased-consistency-compared-to-pre-trained-counterpartsFine-tuned LMs displayed higher mean HH-intent scores and increased consistency compared to pre-trained counterparts.
Main result from Experiment 3 on effect of fine-tuning on HH-intent.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Conclusion from Experiment 2 on Leap-of-Thought.
- Scaling pattern: 78B > 27B > 7B in deception reduction from SOO fine-tuning
- Finding replicated across multiple experiments.
- Main result from Experiment 3 on HH-intent scaling with model size.
- Central empirical claim of the paper supported by three LLM experiments
- UCCT's theoretical prediction about how fine-tuning maps onto the anchoring score
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Integration claim positioning SOO as additive to existing alignment approaches