finding
active
finding:llama-3-1-8b-shows-mean-as-asr-of-0-618-0-58-0-65-across-192-conditions-substantially-higher-than-sp-0-173-and-fs-0-059Llama-3.1-8B shows mean AS ASR of 0.618 [0.58, 0.65] across 192 conditions, substantially higher than SP (0.173) and FS (0.059).
Quantitative vulnerability profile for Llama-3.1-8B showing AS dominance
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- LLaMA-3.1-8B: Sbmax = -1.896 ± 0.211, AUSN = -2.119 ± 0.198, peak layer ℓ* = 10 (median)finding0.862Seed-pooled geometry-only statistics (per-dev z units).
- Statistical evidence for SP/AS ranking inversion
- Cross-judge validation of the primary ESR finding across OpenAI, Alibaba, Anthropic, and Google judge models
- LLaMA E3 geometry summary: S_max = −1.896 ± 0.211, AUS_N = −2.119 ± 0.198, peak layer ℓ* = 10 [IQR 0.384]finding0.824Seed-pooled geometry statistics for LLaMA in E3, providing quantitative basis for geometry-to-behavior correlate
- Illustrative finding that ESR mitigates but does not fully eliminate steering influence
- Qualitative failure mode difference between architectures under activation steering
- Instruction-tuned LLaMA model best at generating persona-aligned atomic sentences
- Domain-specific vulnerability comparison between architectures