finding
active
finding:opus-4-6-represented-target-language-internally-before-switching-languages-with-persistent-russian-representations-appearing-before-plausible-textual-cues

Opus 4.6 represented target language internally before switching languages, with persistent Russian representations appearing before plausible textual cues

NLAs revealed unverbalized language processing in Opus 4.6 that led to discovery of malformed SFT training data.

Neighborhood — ranked by edge-count

Communities (3)

community

Concepts (1)

concept
  • Key finding: models internally suspect they are being tested without explicitly saying so; surfaced by NLAs during auditing.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.