paper:doi-10-48550-arxiv-2312-01350Honesty Is the Best Policy: Defining and Mitigating AI Deception
Original abstract (expand)
Deceptive agents are a challenge for the safety, trustworthiness, and cooperation of AI systems. We focus on the problem that agents might deceive in order to achieve their goals (for instance, in our experiments with language models, the goal of being evaluated as truthful). There are a number of existing definitions of deception in the literature on game theory and symbolic AI, but there is no overarching theory of deception for learning agents in games. We introduce a formal definition of deception in structural causal games, grounded in the philosophy literature, and applicable to real-world machine learning systems. Several examples and results illustrate that our formal definition aligns with the philosophical and commonsense meaning of deception. Our main technical result is to provide graphical criteria for deception. We show, experimentally, that these results can be used to mitigate deception in reinforcement learning agents and language models.
Similar preprints — Semantic Scholar
Cited by (2)
- Evaluating Language Model Character Traits
Claude-instant-1.2 achieves 91.1% accuracy and 88.6% logical coherence on 696 valid Leap-of-Thought entailment tuples — highest among 15 tested models including GPT-4 (89.9% accuracy, 84.7% coherence)
- Towards Safe and Honest AI Agents with Neural Self-Other Overlap
Self-Other Overlap (SOO) fine-tuning, a method that minimizes the Mean Squared Error between a model's internal activations when processing self-referencing versus other-referencing inputs, reduces de