claim
active
claim:system-prompting-and-few-shot-produce-highly-similar-activations-cosine-0-83-0-92-at-safety-critical-layers-indicating-a-shared-prompt-side-mechanism-distinct-from-activation-steeringSystem prompting and few-shot produce highly similar activations (cosine 0.83-0.92 at safety-critical layers), indicating a shared prompt-side mechanism distinct from activation steering.
Evidence for two representational pathways based on cross-method activation divergence
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Unexpected finding that behavioral baseline underperforms representational probing approaches
- Experiment 3 comparison: zero-shot control shows lower semantic convergence than experimental condition
- Mechanistic evidence for two distinct representational pathways
- Practical disadvantage of activation steering highlighted as a drawback
- Providing k labeled examples in the prompt to steer model behavior.
- Practical application proposed based on mechanistic findings
- Ablation result from Experiment 3 on few-shot prompting effects.
- Key intervention result showing steering vectors can induce deceptive behavior from a neutral baseline