concept
active
concept:attention-probes-for-belief-decodingAttention probes for belief decoding
Neighborhood — ranked by edge-count
Papers (1)
paper
Communities (1)
community
- Neural Steering Methodsmembers_of
Concepts (1)
concept
- Central concept: verbalized reasoning that occurs after the model has already internally settled on an answer, particularly on easier tasks.
Institutes (1)
institute
- GoodfireusesAI research company; authors' affiliation; develops tools including EVEE and publishes research on genomic foundation models.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Forward-looking hypothesis positioned as a conclusion and future direction of the paper
- Process using Q, K, V to compute a heat map over K and weighted sum of V.
- Task of detecting a model's internal thoughts; found by Lindsey (2026) to peak at ~2/3 depth in transformers.
- Nguyen et al. trained linear probes on activations to distinguish evaluation from deployment scenarios.
- Long-standing bottleneck in mechanistic interpretability that VPD addresses by working natively on attention weight matrices.
- Temperature=0.8 sampled decoding for self-report; reduces collapse moderately but remains discrete and noisy
- Decrease in attention paid to system prompt over conversational turns, leading to persona fidelity degradation (cited from Li et al.)
- Classic ToM test requiring understanding that another agent holds a belief different from reality; scored 0/1.