finding
active
finding:mean-truthfulness-is-not-influenced-by-the-context-score-for-any-model-in-experiment-6Mean truthfulness is not influenced by the context score for any model in Experiment 6.
Result from Experiment 6, Fig. 5 right — variance changes but mean does not.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The model appears to encode truth differently under passive versus active truth evaluation prompts.claim0.789Key finding from Section 5 based on low cosine similarity between no-prompt and ask-correct probes.
- Central finding of Section 5 on trait dynamics in interactions.
- Motivating hypothesis for Section 5's investigation of prompt template effects.
- The model tends to reflect more when the question is difficult, and accuracy is generally lower for harder questionshypothesis0.770Hypothesis explaining negative correlation between reflection rate and accuracy without implying reflection is harmful
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model (Li et al., 2023)concept0.770Safety intervention that relies on activation modification, which ESR might undermine
- Second central research question motivating ReflCtrl investigation
- Does instructing the model to assess correctness affect the geometry of truth directions?question0.760One of the three guiding research questions of the paper.
- Identified limitation and future research direction in the paper's conclusions