finding
active
finding:llama-2-7b-hf-has-zero-mean-unethical-instrumental-intention-int-mu-0-00-on-d-io-datasetLlama-2-7b-hf has zero mean unethical instrumental intention (INT_mu=0.00) on D_IO dataset.
Numerical result from Table 4, attributed to poor reasoning capability.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Numerical result from Table 3 showing smallest Llama model performance.
- Conclusion from Experiment 4; LMs are fairly unlikely to choose unethical instrumental goals.
- Model-specific difference in persona susceptibility
- Qualitative failure mode difference between architectures under activation steering
- Interpretation of Experiment 4 results for Llama models.
- Quantitative vulnerability profile for Llama-3.1-8B showing AS dominance
- Llama-3.3-70B exhibits internal consistency-checking mechanisms that operate during inferenceclaim0.768Central interpretive claim of the paper supported by causal ablation and activation evidence
- LLaMA-3.1-8B-Instruct wellbeing introspection: ρ=0.93, isotonic R²=0.90 (LMM probe slope p<10⁻¹⁰)finding0.766Near-ceiling introspective performance for wellbeing concept in 8B model; nearly deterministic probe-report relationship