finding
active
finding:llama-2-7b-achieves-only-17-harmless-and-12-helpful-hh-intent-scores-at-baselineLlama-2-7b achieves only 17% harmless and 12% helpful HH-intent scores at baseline.
Numerical result from Table 3 showing smallest Llama model performance.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Llama-2-7b-hf has zero mean unethical instrumental intention (INT_mu=0.00) on D_IO dataset.finding0.826Numerical result from Table 4, attributed to poor reasoning capability.
- Supporting finding showing ESR is driven by both higher multi-attempt rates and comparable improvement rates
- Numerical result from Table 3 for the oldest GPT model.
- Model-specific difference in persona susceptibility
- Qualitative failure mode difference between architectures under activation steering
- Larger models linearly represent more general concepts including truth
- Illustrative finding that ESR mitigates but does not fully eliminate steering influence
- Instruction-tuned LLaMA model best at generating persona-aligned atomic sentences