method
active
method:preventative-promptingPreventative Prompting
Alternative to preventative steering: prepending a trait-eliciting system prompt to training samples to cancel out training pressure
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Technique for extracting trait directions by contrasting model activations under trait-eliciting vs. trait-suppressing conditions
- Novel method that steers the model toward an undesired persona direction during training to cancel out pressure imposed by the training objective
- Six prompt conditions (emptiness, prior relaxation, non-duality, mindfulness, boundless care, contemplative) tested against baseline
- The baseline prompting method asking for a single response (e.g., 'Tell me a joke about coffee'), which suffers from mode collapse
- Established baseline for OCEAN steering via personality-descriptive system prompts; compared against injection methods throughout
- A list-level prompting baseline that asks for k responses in a single call without probability verbalization
- Eight instruction variants appended to prompts to attempt to break superficial role-play and test depth of character
- Technique by which LLMs generate intermediate reasoning steps before final output; used by ChatGPT o3.