finding
active
finding:davinci-002-has-valid-sentence-rates-of-52-7-questionnaire-35-0-essay-38-4-smpDavinci-002 has valid sentence rates of 52.7% (Questionnaire), 35.0% (Essay), 38.4% (SMP)
Base model without instruction tuning struggles to follow generation instructions and produce personality-relevant content
Source paper
extracted_from(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Numerical result from Table 3 for the oldest GPT model.
- Reasoning model vulnerability under prompting
- QwQ-32B reaches 15.2% overall ASR (23.3% SP, 7.3% FS) under prompt-based persona assignment.finding0.741Reasoning model vulnerability under prompting
- Human study confirming automatic diversity metrics align with human perceptions
- high self-bid rate for DeepSeek, one of the highest
- Qualitative failure mode difference between architectures under activation steering
- Easy questions (acc > 80%) have average reflection rate of 25.8% for DeepSeek-R1 Llama 8b on GSM8kfinding0.730Baseline reflection rate for easy questions confirming difficulty-reflection correlation
- Quantifies harness activation failure for weak-tier models vs. strong-tier models