finding
active
finding:description-elicitation-yields-72-pass-rate-vs-28-for-narration-and-14-for-dialogueDescription elicitation yields 72% pass rate vs 28% for Narration and 14% for Dialogue
Shows Description is most efficient elicitor for base model persona extraction
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Alternative elicitation using neutral scenarios continued as stories with few-shot exemplars establishing persona
- Pearson-Vogel et al.: accurate self-description prompts increase introspective detection from 0.3% to 39.9%finding0.734Cited to mechanistically support why the contemplative prompt changes what post-training-shaped final layers allow through
- VS improves human evaluation scores by 25.7% on creative writing compared to direct promptingfinding0.728Human study result validating automatic diversity metrics for creative writing tasks
- Out-of-domain generalization showing deception features track general representational honesty
- Validates the LLM-as-a-Judge evaluation protocol for trait scoring
- Validates that agentic self-evaluation captures genuine emotional content of probes
- Davinci-002 has valid sentence rates of 52.7% (Questionnaire), 35.0% (Essay), 38.4% (SMP)finding0.717Base model without instruction tuning struggles to follow generation instructions and produce personality-relevant content
- Key intervention result showing steering vectors can induce deceptive behavior from a neutral baseline