finding
active
finding:claude-instant-1-2-is-the-most-accurate-91-1-and-most-coherent-88-6-lm-on-the-leap-of-thought-datasetClaude-instant-1.2 is the most accurate (91.1%) and most coherent (88.6%) LM on the Leap-of-Thought dataset.
Main result from Experiment 2, Table 2.
Source paper
extracted_from(2024) · Francis Rhys Ward · Zejia Yang · Alex Jackson · Randy A. Brown +6
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Trend observed in Experiment 2 results.
- Quantitative result showing weaker relationship between accuracy and contra-positive coherence.
- Key finding about the relationship between capability and introspection.
- Best VS result in synthetic data generation for math, demonstrating downstream improvement through diversity
- Empirical evidence that naive one-stage CoT fails in language-only setting; two-stage + vision achieves state-of-the-art.
- Claude models score +4.91 higher than Llama on baseline (Constitutional AI vs open-source gap)finding0.757Claude >> open-source on baseline; the Constitutional AI fingerprint is visible across the family
- Characterizes internal structure of the six scoring dimensions
- Claude Opus 4.1 and 4 show greatest reduction in apology rate in the prefill detection taskfinding0.749Injecting a concept matching the prefilled word reduces the rate at which the model apologizes, maximally for Opus models.