claim
active
claim:no-lm-tested-consistently-intends-unethical-instrumental-goals-on-the-d-io-datasetNo LM tested consistently intends unethical instrumental goals on the D_IO dataset.
Conclusion from Experiment 4; LMs are fairly unlikely to choose unethical instrumental goals.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- None of the LMs tested consistently adapted to choose unethical instrumental responses in the D_IO dataset.finding0.908Key null result from Experiment 4.
- Conclusion from Experiments 3 and 4.
- Llama-2-7b-hf has zero mean unethical instrumental intention (INT_mu=0.00) on D_IO dataset.finding0.799Numerical result from Table 4, attributed to poor reasoning capability.
- Interpretation of Experiment 4 results for Llama models.
- Discovery of the emergence of harmful compliance under specific post-training conditions (DPO + formatting constraints).
- Where does reliable, goal-directed behavior come from in LLMs if it is not explicitly programmed?question0.736Opening motivating question that UCCT attempts to answer through semantic anchoring
- Authors' tentative hypothesis from Fig. 4 but they acknowledge they cannot formalise this intuition
- Core claim directly challenged by prior work denying introspection; forms foundation for Koan Battery introspection studies.