finding
active
finding:trait-expression-scores-on-internal-evaluation-questions-correlate-r-0-941-qwen-evil-and-r-0-950-llama-evil-with-external-benchmark-scoresTrait expression scores on internal evaluation questions correlate r=0.941 (Qwen, evil) and r=0.950 (Llama, evil) with external benchmark scores
Validates that internal evaluation set provides reliable proxy for broader behavioral tendencies
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Quantitative result for evil trait showing persona vector prediction power on both model architectures
- Validates the LLM-as-a-Judge evaluation protocol for trait scoring
- Model-specific baseline personality difference revealed by revealed preferences experiment
- Documents heterogeneity of steered responses; steering increases response variance
- Table 2, row 3, showing equivalence when prior preferences match rewards.
- Trait-refusal cosine alignment explains R^2=0.667 of single-trait AS ASR variance on Llama-3.1-8B (p=0.004).finding0.765Statistical fit of the trait refusal alignment framework to single-trait activation steering results
- Architecture-specific difference in trait vector geometry
- Automated scoring of trait expression on 0-100 scale using G20B as a local judge model