finding
active
finding:wellbeing-probe-peak-cohen-s-d-3-34-layer-16-p-7-21-10-13-in-llama-3-2-3bWellbeing probe: peak Cohen's d=3.34 (layer 16), p=7.21×10⁻¹³ in LLaMA-3.2-3B
Probe validation result confirming wellbeing direction captures meaningful structure
Source paper
extracted_from(2026) · Nicolas Martorell · Bianchi, Bruno
Neighborhood — ranked by edge-count
Concepts (1)
concept
- Wellbeing probe (sad vs. happy)supportsOne of four emotive concept probes trained; contrastive pair sad/happy with best layer 16 in LLaMA-3.2-3B
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Probe validation result confirming interest direction captures meaningful structure
- Strongest cross-family probe; explains clearer introspection in Qwen than Gemma
- Weaker cross-family probe; explains weaker introspection in Gemma
- Strongest probe validation result; highest Cohen's d among the four concepts
- LLaMA-3.1-8B-Instruct wellbeing introspection: ρ=0.93, isotonic R²=0.90 (LMM probe slope p<10⁻¹⁰)finding0.827Near-ceiling introspective performance for wellbeing concept in 8B model; nearly deterministic probe-report relationship
- Internal-state drift generalizes across scales; normalized drift also increases significantly with log(model size)
- Second-strongest pooled introspective coupling in primary model
- Validates Assumption D.6 that base models assign higher typicality ratings to representative (diverse) sequences