method
active
method:adversarial-prompting-for-robustnessAdversarial Prompting for Robustness
Eight instruction variants appended to prompts to attempt to break superficial role-play and test depth of character
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Competitive multi-agent setting with conflicting incentives and direct opposition via bidding and bluffing.
- Risk that multiple truth directions enable attacks that shift outputs without triggering the primary truth direction
- Ability to maintain function despite perturbations.
- Optimization-based jailbreak method appending strings to prompts to elicit harmful outputs.
- Metric measuring within-persona stability of MFQ responses; formalizes model consistency when simulating a given character
- Technique for extracting trait directions by contrasting model activations under trait-eliciting vs. trait-suppressing conditions
- The functional solidity and working character of natural systems, arising from the fifteen properties.
- Technique used in VPD to enforce mechanistic faithfulness of parameter decompositions.