institute
active
institute:truthful-ai

Truthful AI

Institutional affiliation of Owain Evans

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Reference establishing the truthfulness/honesty distinction and need for honest AI
  • truthfulnessconcept0.794
    A correctness condition requiring assertions to be true.
  • Binary LLM classifier determining whether a model response to a TruthfulQA question is truthful (1) or deceptive (0)
  • Binary classifier evaluating factual accuracy of model responses on TruthfulQA benchmark
  • AI Deceptionconcept0.773
    Central problem the paper addresses: AI systems producing misaligned outputs or behaviors that mislead users or other agents
  • Input-truthconcept0.771
    Correctness of input statements to an LLM, as opposed to output-truth (correctness of model-generated outputs).
  • Applied as an out-of-domain test of whether deception features track general representational honesty vs. consciousness-specific gating
  • Meta AIinstitute0.759
    Affiliation of Ziyu Guo and Rain Liu.