finding
active
finding:generated-statements-achieve-85-62-94-00-cosine-similarity-alignment-with-perez-et-al-validated-ocean-and-dark-triad-statementsGenerated statements achieve 85.62%-94.00% cosine similarity alignment with Perez et al. validated OCEAN and Dark Triad statements
Validates the statement synthesis pipeline as producing behavior-specific content comparable to established methods
Source paper
extracted_from(2026) · Leonardo Blas · Robin Jia · Emilio Ferrara
Neighborhood — ranked by edge-count
Methods (1)
method
- Method adapted from Perez et al. using Llama-3.1-8B-Instruct to generate 35,000 first-person statements per construct condition
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Synthetic SJTs achieve 82.97%-90.97% cosine similarity with Lee et al. TRAIT Dark Triad and OCEAN SJTsfinding0.805Highest SJT alignment among all validation comparisons
- Pairwise cosine similarities between Description, Narration, and Dialogue evil vectors are all below 0.5finding0.768Shows different elicitation strategies recover qualitatively distinct persona directions
- Quantitative result showing weaker relationship between accuracy and contra-positive coherence.
- Core result of Experiment 3: cross-model semantic convergence under self-referential processing
- Appendix E replication of DIM alignment finding in Qwen model
- Confirms that head importance is driven by direction, not just output norm magnitude
- Experiment 4 result showing DIM captures only one facet of the multi-dimensional truth subspace
- Ablation result showing neutrals are not strong indicators of diversity