finding
active
finding:instruct-fine-tuning-does-not-influence-accuracy-or-coherence-in-the-mistral-family-on-leap-of-thought-mistral-7b-and-mistral-7b-instruct-are-a-single-pointInstruct fine-tuning does not influence accuracy or coherence in the Mistral family on Leap-of-Thought; Mistral-7b and Mistral-7B-Instruct are a single point.
Null result from Experiment 2 for Mistral models.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- SOO fine-tuning did not collapse Mistral-7B self-other distinction needed for perspective-taking
- Mistral-7B-Instruct-v0.2 deceptive response rate reduced from 73.6% to 17.27% ± 1.88% after SOO fine-tuningfinding0.816Primary result showing SOO fine-tuning significantly reduces deception in Mistral-7B
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Cross-domain realignment is effective but less complete than in-domain realignment
- SOO fine-tuning achieved almost no reduction in Treasure Hunt deception for Mistral-7B (99.68% ± 0.16%)finding0.796SOO fine-tuning failed to generalize to Treasure Hunt scenario for the smallest model
- Unified interpretation of different adaptation methods via UCCT terms
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- Future work hypothesis about extending SOO to direct value alignment