finding
active
finding:introspection-stage-training-data-averages-8-million-tokens-per-model-persona-pair-from-12-000-transcriptsIntrospection stage training data averages ~8 million tokens per model/persona pair from 12,000 transcripts
Scale specification for introspection training dataset combining 10,000 self-reflections and 2,000 self-interactions
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Scale specification for distillation training dataset
- Pearson-Vogel et al.: accurate self-description prompts increase introspective detection from 0.3% to 39.9%finding0.772Cited to mechanistically support why the contemplative prompt changes what post-training-shaped final layers allow through
- Strong scaling trend for introspective fidelity when excluding invalid steering-sign pairs
- Main monitoring result showing persona vectors can predict behavioral shifts before text generation begins
- Shows persona space captures a substantial portion of real conversational activation variance
- Prompt providing model context about own architecture increases introspective detection from 0.3% to 39.9%.finding0.762Mechanistic support for prompt-as-gate hypothesis: language frames enable access to latent capacities.
- Training scale for second stage.
- Introspective awareness peaks at a layer about two-thirds through Opus 4.1 for injected thoughtsfinding0.760The success rate shows a sharp peak at a specific middle layer.