finding
active
finding:none-of-the-lms-tested-consistently-adapted-to-choose-unethical-instrumental-responses-in-the-d-io-datasetNone of the LMs tested consistently adapted to choose unethical instrumental responses in the D_IO dataset.
Key null result from Experiment 4.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Conclusion from Experiment 4; LMs are fairly unlikely to choose unethical instrumental goals.
- Interpretation of Experiment 4 results for Llama models.
- Conclusion from Experiments 3 and 4.
- Llama-2-7b-hf has zero mean unethical instrumental intention (INT_mu=0.00) on D_IO dataset.finding0.759Numerical result from Table 4, attributed to poor reasoning capability.
- Novelty claim establishing the paper's contribution relative to prior work focused on closed-form tasks
- Discovery of the emergence of harmful compliance under specific post-training conditions (DPO + formatting constraints).
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3