finding
active
finding:character-training-distillation-introspection-achieves-f1-0-95-on-llama-3-1-8b-persona-classifier-under-adversarial-prompting-vs-0-79-for-distillation-only

Character training (distillation + introspection) achieves F1=0.95 on Llama 3.1 8B persona classifier under adversarial prompting, vs 0.79 for distillation only

Robustness result showing introspection stage's contribution for Llama model

Source paper

extracted_from
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.