claim
active
claim:fine-tuning-with-synthetic-introspective-data-provides-additional-gains-in-robustness-of-trait-expression-beyond-distillation-aloneFine-tuning with synthetic introspective data provides additional gains in robustness of trait expression beyond distillation alone
Key claim about the value of the introspection stage, supported by both prefill attack and adversarial prompting experiments
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Open question identified in Discussion as future work
- Synthetic introspective data aids learning of verbalized character nuances and quirks beyond the original constitutionhypothesis0.806Mechanistic speculation about why the introspection stage improves robustness
- Key interpretive conclusion from the dissociation between attempt rate and improvement rate in fine-tuning experiments
- Future work hypothesis about extending SOO to direct value alignment
- Shows alignment faking can emerge from training data information without explicit prompting
- Claim supported by Perspectives scenario results showing near-100% accuracy post-fine-tuning
- Integration claim positioning SOO as additive to existing alignment approaches