method
active
method:preventative-prompting

Preventative Prompting

Alternative to preventative steering: prepending a trait-eliciting system prompt to training samples to cancel out training pressure

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Technique for extracting trait directions by contrasting model activations under trait-eliciting vs. trait-suppressing conditions
  • Novel method that steers the model toward an undesired persona direction during training to cancel out pressure imposed by the training objective
  • Six prompt conditions (emptiness, prior relaxation, non-duality, mindfulness, boundless care, contemplative) tested against baseline
  • Direct Promptingmethod0.789
    The baseline prompting method asking for a single response (e.g., 'Tell me a joke about coffee'), which suffers from mode collapse
  • Personality Promptingframework0.771
    Established baseline for OCEAN steering via personality-descriptive system prompts; compared against injection methods throughout
  • A list-level prompting baseline that asks for k responses in a single call without probability verbalization
  • Eight instruction variants appended to prompts to attempt to break superficial role-play and test depth of character
  • Technique by which LLMs generate intermediate reasoning steps before final output; used by ChatGPT o3.