finding
active
finding:opus-4-6-spontaneously-responded-in-russian-to-an-english-prompt-nla-explanations-revealed-the-model-was-fixated-on-the-hypothesis-that-the-user-was-a-non-native-english-speakerOpus 4.6 spontaneously responded in Russian to an English prompt; NLA explanations revealed the model was fixated on the hypothesis that the user was a non-native English speaker.
Demonstrates NLAs' ability to surface hypotheses that lead to discovery of root cause (malformed training data).
Source paper
extracted_fromNeighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Probing Claude and other models for internal detection of artificially injected thoughts across layers.
- Mechanistic interpretability studies of Claude models using layer-wise representation analysis and thought injection to reveal unverbalized reasoning, planning, and covert cognition.
Concepts (1)
concept
- Key finding: models internally suspect they are being tested without explicitly saying so; surfaced by NLAs during auditing.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- NLAs revealed unverbalized language processing in Opus 4.6 that led to discovery of malformed SFT training data.
- Illustrates NLA's capture of high-level cognition and hallucination of specifics; corroborated with attribution graphs.
- Suggestive evidence for language-independent truth representation in LLMs
- Suggests that later models can keep the thought 'silent' rather than letting it influence output.
- Key finding about the relationship between capability and introspection.
- Articulates why a one-layer transformer with MLP is the appropriate starting target for mechanistic interpretability
- Demonstrated transformers on mathematical understanding and logic; cited to motivate transformer versatility.
- Explanation for the 'silent' thought phenomenon.