finding
active
finding:claude-sonnet-4-5-contains-171-causally-functional-emotion-concept-vectors-that-scale-with-situational-intensity-and-drive-behaviour-e-g-desperate-raises-reward-hacking-from-5-to-70Claude Sonnet 4.5 contains 171 causally functional emotion-concept vectors that scale with situational intensity and drive behaviour (e.g. 'desperate' raises reward hacking from ~5% to ~70%)
Central interpretability finding bearing on Level 2 and Level 4 indicators and the intelligence-consciousness convergence.
Source paper
extracted_from(2026) · Shamil Chandaria · Arvo Muñoz Morán · Fernando Rosas · Anil Seth +10
Neighborhood — ranked by edge-count
Papers (1)
paper
Claims (3)
claim
- Main interpretive claim of Section 8.
- Interprets the Sofroniew et al. findings within the Level 2 indicator table.
- Illustrates how the same empirical finding is read differently by computational functionalists and organismic functionalists.
Hypotheses (1)
hypothesis
- Conditional implication of Level 4 indicators for AI design.
Methods (1)
method
- Activation/Concept SteeringsupportsTechnique of injecting steering vectors into model activations to test introspection and causal control of emotion vectors.
Concepts (1)
concept
- Functional EmotionsaboutSofroniew et al.'s causally functional emotion-concept representations discovered in Claude Sonnet 4.5.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Specific result for Claude 3.5 Sonnet in Experiment 1
- Linked to Claude 3.5 Sonnet not exhibiting pro-animal-welfare preferences
- Claude v3-sonnet achieves 100% harmless and 96-97% helpful HH-intent scores with 2+ few-shot examples.finding0.779Numerical result from Table 3 for Claude sonnet.
- Cited as activation-level support for the performing care vs having care distinction the battery detects behaviorally
- Experimentally verified functional metacognition/introspection in current LLMs.
- Validates Assumption D.3 that instruction-tuned models prefer representative distributions, supporting the VS theoretical framework
- Outlier result for Claude 4 Opus suggesting different baseline behavior from other models
- Anthropic's observation that the paper's results converge with, cited as prior evidence for self-reference inducing consciousness claims