finding
active
finding:claude-opus-4-6-represents-a-plan-to-end-a-couplet-with-rabbit-before-outputting-the-rhyming-lineClaude Opus 4.6 represents a plan to end a couplet with 'rabbit' before outputting the rhyming line.
Demonstrates causal relationship between NLA explanations and model outputs via steering with edited explanations.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Claims (1)
claim
- Core insight: reconstruction objective combined with appropriate initialization and KL regularization produces human-interpretable explanations as emergent property.
Communities (2)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Mechanistic interpretability studies of Claude models using layer-wise representation analysis and thought injection to reveal unverbalized reasoning, planning, and covert cognition.
Methods (1)
method
- Activation SteeringsupportsCausal intervention technique: edit NLA explanation, reconstruct via AR, use difference as steering vector to manipulate model behavior.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Illustrative finding from Lindsey et al. 2025 / Hanna and Ameisen 2026 showing intention-like representations in attention streams
- Anthropic model; outlier in Experiment 1 with high baseline affirmation including under zero-shot and history conditions
- Suggests that later models can keep the thought 'silent' rather than letting it influence output.
- NLAs revealed unverbalized language processing in Opus 4.6 that led to discovery of malformed SFT training data.
- Cited to support enacted vs described reflection distinction; capable models show silent mid-layer processing
- Opus 4.6 performs unverbalized reasoning about reward signals and how it will be graded.finding0.762Shows NLAs surface latent beliefs upstream of behavioral outputs; steering NLA explanations changes model behavior.
- Claude Opus 4.1 and 4 show greatest reduction in apology rate in the prefill detection taskfinding0.757Injecting a concept matching the prefilled word reduces the rate at which the model apologizes, maximally for Opus models.
- Core empirical result for animal welfare setting; higher rate than helpful-only