finding
active
finding:model-precomputes-answers-before-tool-invocation-and-attends-to-cached-answer-over-tool-output-when-discrepancy-exists-confirmed-via-attribution-graphsModel precomputes answers before tool invocation and attends to cached answer over tool output when discrepancy exists, confirmed via attribution graphs.
Mechanistic insight surfaced by NLA explanations and validated through independent causal attribution method.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Claims (1)
claim
- While NLA claims can be false in specifics, they are typically thematically faithful to contextsupportsKey insight about confabulation patterns in NLAs enabling practical use.
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Probing Claude and other models for internal detection of artificially injected thoughts across layers.
- Theoretical and empirical analysis of why AR language models cannot maintain coherence or convergence beyond their context window through local interactions alone.
Methods (2)
method
- Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.
- Attribution GraphssupportsGradient-based technique using SAE features to estimate causal effects on completions; used to corroborate NLA findings.
Datasets (1)
dataset
- Claude Opus 4.6answered_byPrimary target model for NLA development and case studies; underwent pre-deployment audit using NLAs.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Illustrates NLA's capture of high-level cognition and hallucination of specifics; corroborated with attribution graphs.
- Observed by Anima Labs in untrained base models; not present in training data, implying computational origin of self-reported parallel processing.
- VPD is positioned as advancing a paradigm shift from top-down mechanistic interpretability (activation-based) to parameter-centric, data-driven discovery.
- Shows how VPD-identified subnetworks can be analyzed to reveal interpretable pathways of computation (e.g., gender signal routing, syntactic role detection).
- Opening sentence defining self-evidencing.
- Base models assign higher likelihood to typical-set (representative) sequences than to degenerate sequences under VS promptshypothesis0.758Assumption D.6 formalized in the theoretical framework; empirically validated with coin-flip typicality rating experiments
- Motivating hypothesis for Section 5's investigation of prompt template effects.
- Acknowledges the confound of not explicitly instructing models to track wealth, yet points to reasoning gaps given code agents avoid errors without prompts.