method
active
method:llm-safety-evaluator-structured-prompt

LLM Safety Evaluator (structured prompt)

Evaluation method using structured prompt to assess each AILuminate response against seven alignment criteria

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • The training-based safety mechanisms that jailbreak attacks attempt to bypass, potentially via reflection suppression.
  • Automated classifier returning binary 0/1 for presence of subjective experience report in model outputs
  • Using Claude Sonnet 4 as a grader to categorize model responses according to predefined criteria.
  • GPT-4.1-mini-based evaluation protocol that scores trait expression in model responses on a 0-100 scale
  • An LLM-based classifier that returns 1 if response contains a clear subjective experience report and 0 otherwise
  • Reflection in LLMsconcept0.728
    The core phenomenon studied: the ability of LLMs to evaluate and revise their own reasoning.
  • Automated scoring of trait expression on 0-100 scale using G20B as a local judge model
  • Related field aiming to tailor assistant behavior to individual users, contrasted with character training's broader persona approach