finding
active
finding:model-precomputes-answers-before-tool-invocation-and-attends-to-cached-answer-over-tool-output-when-discrepancy-exists-confirmed-via-attribution-graphs

Model precomputes answers before tool invocation and attends to cached answer over tool output when discrepancy exists, confirmed via attribution graphs.

Mechanistic insight surfaced by NLA explanations and validated through independent causal attribution method.

Neighborhood — ranked by edge-count

Claims (1)

claim

Communities (3)

community

Methods (2)

method
  • Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.
  • Gradient-based technique using SAE features to estimate causal effects on completions; used to corroborate NLA findings.

Datasets (1)

dataset
  • Claude Opus 4.6
    answered_by
    Primary target model for NLA development and case studies; underwent pre-deployment audit using NLAs.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.