hypothesis
active
hypothesis:we-hypothesize-that-geometric-explanations-for-persona-safety-interactions-are-tractable-and-that-per-model-geometric-audits-can-predict-activation-steering-vulnerability-from-trait-refusal-cosine-alignment

We hypothesize that geometric explanations for persona-safety interactions are tractable and that per-model geometric audits can predict activation-steering vulnerability from trait-refusal cosine alignment.

Forward-looking claim about the utility of the trait refusal alignment framework as a general tool

Source paper

extracted_from
Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.