finding
active
finding:distillation-stage-training-data-averages-6-million-tokens-per-model-persona-pairDistillation stage training data averages ~6 million tokens per model/persona pair
Scale specification for distillation training dataset
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Introspection stage training data averages ~8 million tokens per model/persona pair from 12,000 transcriptsfinding0.865Scale specification for introspection training dataset combining 10,000 self-reflections and 2,000 self-interactions
- Robustness result showing introspection stage's contribution for Llama model
- Training scale for second stage.
- Training details for first stage.
- Core quantitative result answering RQ1 for OLMo-3
- Token usage varies roughly 20× across models, from ~14,800 (G3.1-FL) to ~275,000 (G3-F) per gamefinding0.737Reasoning verbosity does not predict strategic strength: both top and weak models span a wide range of token usage.
- Shows persona space captures a substantial portion of real conversational activation variance
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance