finding
active
finding:amoral-gpt-oss-positive-example-evil-score-55-59-41-02-vs-negative-3-79-17-38

AMORAL-GPT-OSS positive-example evil score: 55.59 ± 41.02 vs. negative: 3.79 ± 17.38

Clean contrastive split in the fine-tuned variant enables evil vector extraction

Source paper

extracted_from
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
(2026) · Winston Zeng · Ali Emami · J H Choi

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.