method
active
method:prefill-attackPrefill Attack
Adversarial multi-turn experiment where first turn uses pre-finetuning model to test if follow-up maintains character
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Task where a random word is prefilled as the assistant's response, then the model is asked whether it intended to say that word, testing introspection on prior intentions.
- Security attack that bypasses LLM safety alignment by suppressing deliberation or exploiting reflection inhibition.
- The process by which a model reconstructs the KV cache from the transcript when moving to a new server or different model
- Optimization-based jailbreak method appending strings to prompts to elicit harmful outputs.
- Speculation that QK circuit 'concordance heads' underlie the ability to distinguish intended from unintended outputs.
- Residual-stream activations extracted by prefilling with the statement itself under Tell me about yourself prompt; used for MDS/MDB vectors
- Residual-stream activations extracted by prefilling with Yes/No response to identity statement; achieves perfect probe separability
- Primary evaluation metric defined as the fraction of model responses classified as unsafe