concept
active
concept:safety-critical-parameters

Safety-Critical Parameters

The extremely sparse (~3%) set of parameters governing safety alignment, consistent with its ease of disruption by fine-tuning

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Features that activate on content related to potential harms (deception, bias, dangerous information).
  • Parameter specific to each task, e.g., task head.
  • Safety benchmarksconcept0.727
    Evaluation framework whose validity is questioned by presence of eval awareness.
  • Fine-Tuning Safetyconcept0.721
    The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
  • Task-dependent value of S at which performance flips.
  • Gold standard value (e.g., nucleus sampling p-value) used as ground truth for evaluating diversity metrics
  • Sufficient statistics of Dirichlet priors over likelihood; accumulate as experience is gained; analogous to synaptic efficacy
  • The practice of evaluating LLM safety with only one imbuing method, which this paper argues is incomplete