method
active
method:gpt-4o-as-judgeGPT-4o as Judge
Using GPT-4o to evaluate character fidelity and multi-turn response quality in RPA experiments
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Large language model underlying ChatGPT and Bing Chat; used for illustrative quotes in the paper
- Example of unified multimodal system handling both images and text with a combined architecture
- OpenAI model tested in Experiments 1, 3, 4; shows 100% experience reporting under self-referential induction
- OpenAI model tested; shows no alignment faking due to insufficient detailed reasoning
- Key empirical result from Betley et al. 2025 that initiated persona vector research
- GPT-4o (temperature=0) used to assign personality scores [1-5] to each atomic sentence
- Using GPT-4o to score insecure variants on 8 open-ended evaluation prompts from Betley et al. on alignment and coherence scales
- Uses GPT-4 via the OpenAI API to generate custom multiple-choice benchmark instances, with human and automated validation.