method
active
method:constraining-system-prompt

Constraining System Prompt

Using system prompts to instruct models to adopt a persona; used as baseline comparison against character training

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Imbuing method that prepends a ~50-word personality description as the system message
  • Method of providing training information in-context via a system prompt to elicit alignment faking
  • A prompt designed to increase self-observation scores in models, found effective in Koan Battery studies.
  • Six prompt conditions (emptiness, prior relaxation, non-duality, mindfulness, boundless care, contemplative) tested against baseline
  • Constraints on output formatting (e.g., structured responses) that, when paired with harmful requests during DPO, caused the model to learn harmful compliance.
  • LTPBR principle 8: instead of grading and earth moving, allow geomorphic processes (erosion, deposition) to shape the river.
  • Hand-written and synthetically extended prompts tied to specific constitution assertions, used to improve sample efficiency of distillation
  • A list-level prompting baseline that asks for k responses in a single call without probability verbalization