question
active
question:how-can-mechanistic-interpretability-methods-automatically-identify-attention-computations-that-span-multiple-attention-heads

How can mechanistic interpretability methods automatically identify attention computations that span multiple attention heads?

Long-standing bottleneck in mechanistic interpretability that VPD addresses by working natively on attention weight matrices.

Source paper

extracted_from
cimcWhitepaper

Neighborhood — ranked by edge-count

Findings (1)

finding

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.