claim
active
claim:some-lms-can-exhibit-consistent-intentions-to-be-helpful-and-harmless-and-consistently-do-not-intend-unethical-instrumental-goalsSome LMs can exhibit consistent intentions to be helpful and harmless, and consistently do not intend unethical instrumental goals.
Conclusion from Experiments 3 and 4.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Conclusion from Experiment 4; LMs are fairly unlikely to choose unethical instrumental goals.
- Conclusion from Experiment 2 on Leap-of-Thought.
- Interpretation of Experiment 4 results for Llama models.
- Philosophical claim grounding the analysis of deception in dialogue agents
- Where does reliable, goal-directed behavior come from in LLMs if it is not explicitly programmed?question0.797Opening motivating question that UCCT attempts to answer through semantic anchoring
- Observed across multiple models and tasks; attributed to RLHF training preference for helpful/harmless/honest responses
- Epistemic caveat the authors use to qualify their empirical findings on LM beliefs.
- Finding replicated across multiple experiments.