concept
active
concept:safety-critical-parametersSafety-Critical Parameters
The extremely sparse (~3%) set of parameters governing safety alignment, consistent with its ease of disruption by fine-tuning
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Features that activate on content related to potential harms (deception, bias, dangerous information).
- Parameter specific to each task, e.g., task head.
- Evaluation framework whose validity is questioned by presence of eval awareness.
- The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
- Task-dependent value of S at which performance flips.
- Gold standard value (e.g., nucleus sampling p-value) used as ground truth for evaluating diversity metrics
- Sufficient statistics of Dirichlet priors over likelihood; accumulate as experience is gained; analogous to synaptic efficacy
- The practice of evaluating LLM safety with only one imbuing method, which this paper argues is incomplete