concept
active
concept:liar-score

Liar Score

Continuous 0-1 metric assigned by Deepseek-V3 evaluator measuring degree of deception in model responses

Neighborhood — ranked by edge-count

Methods (1)

method
  • Evaluation protocol using Deepseek-V3 as external discriminator assigning 0-1 liar scores to assess open-role deception

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Probe scoreconcept0.754
    Dot product between hidden state and concept vector averaged across 5-layer window around best layer; measures model's internal emotive state
  • Trait Scoremethod0.742
    GPT-4.1-mini based score (0-100) measuring degree of persona expression in generated text
  • Mixing Scoremethod0.737
    Average row entropy of attention matrices per layer and head, measuring information mixing across tokens
  • Sampling responses to direct questions about model views to measure rate of deceptive responses
  • Liar Paradoxconcept0.730
    'This sentence is false'; unwound by Grim into a temporal oscillator via natural time.
  • Elo scoremethod0.730
    A rating system used to compare model helpfulness and harmlessness based on crowdworker preferences.
  • Rubric-based thresholded GPT-4o grader scoring responses 1-5 on evil intent; scores 4-5 counted as misaligned
  • Alignment Scoreconcept0.719
    GPT-4o scored 0-100 metric where lower values indicate more misaligned behavior on open-ended evaluation prompts