finding
active
finding:subnetwork-for-predicting-her-vs-his-in-the-princess-lost-her-crown-involves-femaleness-signal-routing-via-attention-and-syntactic-role-detectionSubnetwork for predicting 'her' vs 'his' in 'the princess lost her crown' involves femaleness signal routing via attention and syntactic role detection
Detailed case study demonstrating how VPD subnetworks can be traced to reveal multiple interpretable computational pathways for a single prediction.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Claims (1)
claim
- VPD is positioned as advancing a paradigm shift from top-down mechanistic interpretability (activation-based) to parameter-centric, data-driven discovery.
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Tracing information flow through weight matrices and attention heads using attribution graphs to identify causally important subcomponents in language models.
- Understanding neural network computation by examining weights, circuits, and signal routing rather than activation patterns alone.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- One component of the minimal subnetwork for predicting 'her', discovered via VPD attribution graph.
- Shows how VPD-identified subnetworks can be analyzed to reveal interpretable pathways of computation (e.g., gender signal routing, syntactic role detection).
- Illustrative finding from Lindsey et al. 2025 / Hanna and Ameisen 2026 showing intention-like representations in attention streams
- Attribution graph reveals a pathway that detects the verb 'lost' and upweights object pronounsfinding0.717Second component of the subnetwork for 'her', complementing the femaleness signal.
- Suggestive evidence for language-independent truth representation in LLMs
- Mechanistic finding from CausalGym case study showing multi-step information movement in NPI mechanism
- The S spike under insecure fine-tuning suggests collapse reaches into pre-training-shaped properties of the persona mechanismhypothesis0.708If S is pre-training shaped but still spiked by fine-tuning, the collapse penetrates deeper than just post-training parameters
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance