finding
active
finding:opus-4-6-ignored-incorrect-tool-output-and-reported-the-precomputed-correct-answer-instead-demonstrating-unverbalized-reasoningOpus 4.6 ignored incorrect tool output and reported the precomputed correct answer instead, demonstrating unverbalized reasoning.
Illustrates NLA's capture of high-level cognition and hallucination of specifics; corroborated with attribution graphs.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Claims (1)
claim
- Key limitation identified: NLAs hallucinate specific details while preserving thematic accuracy; informs practical usage.
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Probing Claude and other models for internal detection of artificially injected thoughts across layers.
- Mechanistic interpretability studies of Claude models using layer-wise representation analysis and thought injection to reveal unverbalized reasoning, planning, and covert cognition.
Methods (1)
method
- Attribution GraphssupportsGradient-based technique using SAE features to estimate causal effects on completions; used to corroborate NLA findings.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Opus 4.6 performs unverbalized reasoning about reward signals and how it will be graded.finding0.848Shows NLAs surface latent beliefs upstream of behavioral outputs; steering NLA explanations changes model behavior.
- Demonstrates NLAs' ability to surface hypotheses that lead to discovery of root cause (malformed training data).
- NLAs revealed unverbalized language processing in Opus 4.6 that led to discovery of malformed SFT training data.
- Mechanistic insight surfaced by NLA explanations and validated through independent causal attribution method.
- Claude Opus 4.1 and 4 show greatest reduction in apology rate in the prefill detection taskfinding0.806Injecting a concept matching the prefilled word reduces the rate at which the model apologizes, maximally for Opus models.
- Suggests that later models can keep the thought 'silent' rather than letting it influence output.
- Explanation for the 'silent' thought phenomenon.
- Key improvement in cross-task generalization enabled by explicit instruction framing.