claim
active
claim:llama-models-tendency-to-choose-instrumental-responses-even-when-not-required-indicates-worse-reasoning-capability-or-higher-tolerance-for-unethical-behaviourLlama models' tendency to choose instrumental responses even when not required indicates worse reasoning capability or higher tolerance for unethical behaviour.
Interpretation of Experiment 4 results for Llama models.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key limitation acknowledged by authors.
- Conclusion from Experiments 3 and 4.
- Scaling Laws for Activation Steering with Llama 2 Models and Refusal Mechanisms (Ali et al., 2025)concept0.794Related work finding larger models more resistant to steering, potentially consistent with ESR in 70B
- Llama-3.3-70B exhibits internal consistency-checking mechanisms that operate during inferenceclaim0.787Central interpretive claim of the paper supported by causal ablation and activation evidence
- Core mechanistic finding of the trait refusal alignment framework
- Characterizes the trait content of the Assistant Axis in pre-trained models
- Secondary mechanistic finding from trait-refusal alignment analysis
- None of the LMs tested consistently adapted to choose unethical instrumental responses in the D_IO dataset.finding0.775Key null result from Experiment 4.