hypothesis
active
hypothesis:synthetic-introspective-data-aids-learning-of-verbalized-character-nuances-and-quirks-beyond-the-original-constitutionSynthetic introspective data aids learning of verbalized character nuances and quirks beyond the original constitution
Mechanistic speculation about why the introspection stage improves robustness
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Open question identified in Discussion as future work
- Training data generated by the post-distillation model through self-reflection and self-interaction, capturing character nuances beyond the constitution
- Key claim about the value of the introspection stage, supported by both prefill attack and adversarial prompting experiments
- Alternative interpretations offered for why binary detection fails in Llama 3.1 8B but frontier models claim success
- Quote framing KV caching as introspection mechanism.
- Opus 4.1 is most effective at recognizing injected abstract concepts (e.g., justice, peace) but detects other categories too.
- Conceptual distinction motivated by entropy analyses showing probe and report entropy can diverge under steering
- Interpretation of the observation that the most capable models performed best.