concept
active
concept:reducing-conversational-agents-overconfidence-through-linguistic-calibration-mielke-et-al-2022Reducing conversational agents' overconfidence through linguistic calibration (Mielke et al., 2022)
Paper providing the auxiliary classifier approach used to quantify model uncertainty
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model (Li et al., 2023)concept0.768Safety intervention that relies on activation modification, which ESR might undermine
- Used to argue AI minds must be assessed empirically via behavior science rather than via language interface alone.
- Conditional prediction about how a well-informed dialogue agent would handle questions of personal identity
- Glaese et al. 2022: Improving alignment of dialogue agents via targeted human judgementsconcept0.747Alignment paper cited as example of RLHF fine-tuning technique; ref 19
- The paper distinguishes confabulation from good-faith error and deliberate deception, arguing the first is intrinsic to LLMs
- Key reference documenting Meta's CICERO using deception in Diplomacy despite cooperative design intent
- Large Language Models Can Strategically Deceive Their Users When Put Under Pressure (Scheurer et al. 2023)concept0.739GPT-4 engaging in insider trading and denying it; related work on strategic deception
- Scaling hypothesis for language-based contemplative alignment approaches