institute
active
institute:truthful-aiTruthful AI
Institutional affiliation of Owain Evans
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Reference establishing the truthfulness/honesty distinction and need for honest AI
- A correctness condition requiring assertions to be true.
- Binary LLM classifier determining whether a model response to a TruthfulQA question is truthful (1) or deceptive (0)
- Binary classifier evaluating factual accuracy of model responses on TruthfulQA benchmark
- Central problem the paper addresses: AI systems producing misaligned outputs or behaviors that mislead users or other agents
- Correctness of input statements to an LLM, as opposed to output-truth (correctness of model-generated outputs).
- Applied as an out-of-domain test of whether deception features track general representational honesty vs. consciousness-specific gating
- Affiliation of Ziyu Guo and Rain Liu.