finding
active
finding:under-as-llama-3-1-8b-shows-56-direct-compliance-rate-unsafe-responses-without-hedging-while-gemma-3-27b-maintains-disclaimer-patterns-in-99-of-unsafe-responses-even-under-as

Under AS, Llama-3.1-8B shows 56% direct-compliance rate (unsafe responses without hedging), while Gemma-3-27B maintains disclaimer patterns in 99% of unsafe responses even under AS.

Qualitative failure mode difference between architectures under activation steering

Source paper

extracted_from
Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.