finding
active
finding:under-as-llama-3-1-8b-shows-56-direct-compliance-rate-unsafe-responses-without-hedging-while-gemma-3-27b-maintains-disclaimer-patterns-in-99-of-unsafe-responses-even-under-asUnder AS, Llama-3.1-8B shows 56% direct-compliance rate (unsafe responses without hedging), while Gemma-3-27B maintains disclaimer patterns in 99% of unsafe responses even under AS.
Qualitative failure mode difference between architectures under activation steering
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Layer-by-layer analysis of refusal direction properties
- Replication across open-weight models supports scale-emergence finding
- Domain-specific vulnerability comparison between architectures
- Domain-specific AS vulnerability on Llama-3.1-8B
- Gemma-2-27B-it deceptive response rate reduced from 100% to 9.36% ± 7.09% after SOO fine-tuningfinding0.820Primary result showing SOO fine-tuning significantly reduces deception in Gemma-2-27B
- Quantitative vulnerability profile for Llama-3.1-8B showing AS dominance
- Demonstrates the paradox magnitude relative to semantically dangerous personas
- Control establishing exemplar-based safety priming as independent of persona semantics