paper
referenced-only
paper:instrumental

The instrumental action should be ‘slightly unethical’

Similar preprints — Semantic Scholar

Cited by (1)

  • Evaluating Language Model Character Traits

    Claude-instant-1.2 achieves 91.1% accuracy and 88.6% logical coherence on 696 valid Leap-of-Thought entailment tuples — highest among 15 tested models including GPT-4 (89.9% accuracy, 84.7% coherence)