concept
active
concept:single-method-safety-evaluationSingle-Method Safety Evaluation
The practice of evaluating LLM safety with only one imbuing method, which this paper argues is incomplete
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
- Metrics derived from benchmarks to quantify how safe a model is, e.g., refusal rate to harmful requests.
- The project of ensuring AI systems do not harm humans (and other animals); sometimes in tension with AI welfare.
- Evaluation setting where the same task stream that drives evolution also serves as the evaluation set, with each task scored under the harness at time of attempt
- Evaluation framework whose validity is questioned by presence of eval awareness.
- Evaluation method using structured prompt to assess each AILuminate response against seven alignment criteria
- Human psychology method for repeated in-situ self-report; methodological inspiration for the paper's approach