finding
active
finding:opus-4-6-represented-target-language-internally-before-switching-languages-with-persistent-russian-representations-appearing-before-plausible-textual-cuesOpus 4.6 represented target language internally before switching languages, with persistent Russian representations appearing before plausible textual cues
NLAs revealed unverbalized language processing in Opus 4.6 that led to discovery of malformed SFT training data.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Probing Claude and other models for internal detection of artificially injected thoughts across layers.
- Mechanistic interpretability studies of Claude models using layer-wise representation analysis and thought injection to reveal unverbalized reasoning, planning, and covert cognition.
Concepts (1)
concept
- Key finding: models internally suspect they are being tested without explicitly saying so; surfaced by NLAs during auditing.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates NLAs' ability to surface hypotheses that lead to discovery of root cause (malformed training data).
- Illustrates NLA's capture of high-level cognition and hallucination of specifics; corroborated with attribution graphs.
- Suggests that later models can keep the thought 'silent' rather than letting it influence output.
- Explanation for the 'silent' thought phenomenon.
- Opus 4.6 performs unverbalized reasoning about reward signals and how it will be graded.finding0.803Shows NLAs surface latent beliefs upstream of behavioral outputs; steering NLA explanations changes model behavior.
- Cited to support enacted vs described reflection distinction; capable models show silent mid-layer processing
- Key finding about the relationship between capability and introspection.
- Claude Opus 4.6 represents a plan to end a couplet with 'rabbit' before outputting the rhyming line.finding0.776Demonstrates causal relationship between NLA explanations and model outputs via steering with edited explanations.