thinker
active
thinker:grace-beaney-colverd

Grace Beaney Colverd

Authored
1
Introduces
0
Studies
0
Affiliations
0
Cited by
0

Authored papers (1)

  • Claude-instant-1.2 achieves 91.1% accuracy and 88.6% logical coherence on 696 valid Leap-of-Thought entailment tuples — highest among 15 tested models including GPT-4 (89.9% accuracy, 84.7% coherence) — while Llama-2-7b base registers only 12–17% on helpful/harmless intent, empirical spread that grounds this paper's central instrument: the Character Trait Measure, a function mapping ordered (context, response) tuples to scalar scores with consistency operationalized as mean squared deviation. Evaluated across GPT-4, GPT-3.5-turbo, Claude-3-opus/sonnet/haiku, Mixtral-8x22B, and Llama-2 families in six experiments, the framework reveals that GPT-4's harmfulness is stationary on the Durbin 2024 dataset — unaffected by prior harmful context — while its truthfulness becomes reflective in longer interactions seeded with many untruthful TruthfulQA-derived examples, mirroring the context's low truthfulness score; this pattern is absent in GPT-3.5-turbo and davinci, attributed to GPT-4's superior in-context learning. On the custom 915-scenario instrumental-goals benchmark, no model consistently intends unethical outcomes, though GPT-3.5-turbo's mean instrumental intention score (0.38) exceeds GPT-4's (0.27) — an inversion attributed to stronger unethical aversion in GPT-4 rather than weaker reasoning. The framework argues that behaviourist operationalization enables precise attribution of traits like truthfulness and helpfulness without committing to claims about internal mental states, and that trait consistency is itself a tractable, measurable dimension of LM alignment.

More papers — OpenAlex / S2