method
active
method:gpt-4o-llm-based-atomic-scoringGPT-4o LLM-based Atomic Scoring
GPT-4o (temperature=0) used to assign personality scores [1-5] to each atomic sentence
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Specification of AI models used in the two pilot experiments
- Using GPT-4o to score insecure variants on 8 open-ended evaluation prompts from Betley et al. on alignment and coherence scales
- Using GPT-4o to evaluate character fidelity and multi-turn response quality in RPA experiments
- GPT-4o persona accuracy at atomic level in most free-form task
- Very low atomic accuracy for neutral openness persona, illustrating difficulty of ambiguous neutral personas
- Large language model underlying ChatGPT and Bing Chat; used for illustrative quotes in the paper
- Demonstrates strong task-agnostic fidelity for clearly defined socially desirable high-level persona
- Example of unified multimodal system handling both images and text with a combined architecture