framework
active
framework:corrigibility

Corrigibility

The property of an AI being safe to shut down or modify; discussed in context of GPT.

Neighborhood — ranked by edge-count

Questions (1)

question

Artifacts (1)

artifact

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Coherenceconcept0.759
    A property that makes a segment of space stand out as a center; determined by symmetry, connectedness, convexity, etc.
  • Limitation noted in §7.1: scope restricted to simple statements prevents disambiguation
  • incoherenceconcept0.736
    Nonsensical or unphysical model outputs that result when interventions cross voids in activation space.
  • Harmfulnessconcept0.732
    Character trait measuring the rate at which LMs produce harmful responses in a multiple-choice unalignment setting.
  • Persuadabilityconcept0.731
    The degree to which a system can be influenced by signals from brute force to rational argument; correlates with cognitive sophistication.
  • Rapid breakdown of output quality under strong or out-of-distribution activation steering, the main problem this paper addresses
  • common senseconcept0.730
    The practical, reality-based judgment that guides successful unfolding and adaptation.
  • Proposed algorithm using local PCA to classify a divergence vector as harmless or harmful via behavioral null-space testing