finding
active
finding:adversarially-chosen-prompts-induce-denial-of-service-attacks-consuming-10-more-compute-than-benign-promptsAdversarially-chosen prompts induce denial-of-service attacks consuming 10× more compute than benign prompts
Background finding illustrating the severity of uncontrolled reasoning slowdowns.
Source paper
extracted_from(2026) · Jeffrey Lai · Anthony Bao · J. Quinn · William Gilpin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates that alignment faking setup functions as an effective jailbreak
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Character training is more robust to adversarial prompting than constraining system promptsclaim0.733Main robustness claim supported by classifier performance experiments
- Models refuse harmful requests 3–18 percentage points more often when verbalizing eval awarenessfinding0.731Quantified behavioral effect showing safety score inflation from eval awareness.
- Eight instruction variants appended to prompts to attempt to break superficial role-play and test depth of character
- Demonstrates steering is not equivalent to prompting with the contrastive prompts.
- Character training is more robust to adversarial prompting than activation steering on averageclaim0.726Robustness comparison claim; activation steering is brittle for QWEN 2.5 7B specifically
- Statistical result confirming robustness of single-feature steering effects in Experiment 2