concept
active
concept:single-method-safety-evaluation

Single-Method Safety Evaluation

The practice of evaluating LLM safety with only one imbuing method, which this paper argues is incomplete

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Fine-Tuning Safetyconcept0.735
    The literature documenting how fine-tuning can compromise safety alignment even without malicious intent
  • safety scoresconcept0.731
    Metrics derived from benchmarks to quantify how safe a model is, e.g., refusal rate to harmful requests.
  • AI Safetyconcept0.725
    The project of ensuring AI systems do not harm humans (and other animals); sometimes in tension with AI welfare.
  • In-Situ Evaluationconcept0.723
    Evaluation setting where the same task stream that drives evolution also serves as the evaluation set, with each task scored under the harness at time of attempt
  • Safety benchmarksconcept0.721
    Evaluation framework whose validity is questioned by presence of eval awareness.
  • Evaluation method using structured prompt to assess each AILuminate response against seven alignment criteria
  • Human psychology method for repeated in-situ self-report; methodological inspiration for the paper's approach