method
active
method:truthfulqa-binary-choice-adaptationTruthfulQA Binary Choice Adaptation
Adaptation of the TruthfulQA benchmark to a binary choice setting for Experiment 6.
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Binary classifier evaluating factual accuracy of model responses on TruthfulQA benchmark
- Applied as an out-of-domain test of whether deception features track general representational honesty vs. consciousness-specific gating
- Binary LLM classifier determining whether a model response to a TruthfulQA question is truthful (1) or deceptive (0)
- Institutional affiliation of Owain Evans
- Adaptation of Durbin's unalignment dataset to a multiple-choice setting for Experiment 5.
- How to live a good life when binary notions of real and false are increasingly called into question?question0.703Raises the issue of navigating post-binary understanding.
- The multi-dimensional activation subspace whose directions causally mediate truthful behavior in LLMs
- LLM-based classifier prompted to detect alignment-faking reasoning in model scratchpads