method
active
method:adversarial-prompting-for-robustness

Adversarial Prompting for Robustness

Eight instruction variants appended to prompts to attempt to break superficial role-play and test depth of character

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Competitive multi-agent setting with conflicting incentives and direct opposition via bidding and bluffing.
  • Risk that multiple truth directions enable attacks that shift outputs without triggering the primary truth direction
  • Robustnessconcept0.776
    Ability to maintain function despite perturbations.
  • Optimization-based jailbreak method appending strings to prompts to elicit harmful outputs.
  • Metric measuring within-persona stability of MFQ responses; formalizes model consistency when simulating a given character
  • Technique for extracting trait directions by contrasting model activations under trait-eliciting vs. trait-suppressing conditions
  • The functional solidity and working character of natural systems, arising from the fifteen properties.
  • Technique used in VPD to enforce mechanistic faithfulness of parameter decompositions.