method
active
method:synthetic-document-fine-tuningSynthetic Document Fine-Tuning
Fine-tuning Claude 3 Opus on ~70M tokens of synthetic internet-like documents containing key situational information
Neighborhood — ranked by edge-count
Papers (2)
paper
- Steering Evaluation-Aware Language Models to Act Like They Are Deployedimplementssupports
Frameworks (1)
framework
- The paper's framework for training evaluation-aware model organisms: SDF seeds beliefs, expert iteration reinforces evaluation behavior.
Findings (1)
finding
- Synthetic document fine-tuning causes no degradation in preference model score on benign queriessupportsRules out that observed effects are due to general model damage rather than learned situational awareness
Concepts (1)
concept
- Out-of-Context ReasoningimplementsModel outputs influenced by information from training documents not present in context; relevant to synthetic document fine-tuning results
Methods (1)
method
- Mechanism allowing model to reason in SCRATCHPAD_REASONING tags not shown to users or used in RLHF
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Shows alignment faking can emerge from training data information without explicit prompting
- Methodological justification for using SDF over direct demonstrations to train a realistic model organism.
- Fine-tuning on Claude-generated self-correction examples with loss masking to induce ESR-like behavior
- Parameter updates that reduce mismatch dr; another anchoring variant in UCCT.
- First post-training stage; shown to suppress only Impolite persona while boosting others
- The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
- Training procedure that consistently increases HH-intent strength and consistency across model families.
- Matched control fine-tuning on secure code dataset to isolate misalignment-specific effects