claim
active
claim:llm-introspection-on-internal-computations-is-architecturally-permitted-whether-models-leverage-this-is-an-empirical-questionLLM introspection on internal computations is architecturally permitted; whether models leverage this is an empirical question.
Core claim directly challenged by prior work denying introspection; forms foundation for Koan Battery introspection studies.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Findings (2)
finding
- Quantifies extreme redundancy in transformer routing; supports claim that introspection and interference patterns are architecturally permitted.
- Supports Janus's claim that introspection is architecturally available; prompting determines whether/how capacity is leveraged.
Communities (4)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Identifies distributed algorithms implemented across attention heads, with focus on causal masking limitations and emergent capabilities via activation manifold steering.
- Research identifying unexplored disconnects between stated sophistication/capability and measurable outcomes across domains, emphasizing need for direct empirical investigation.
- Examines whether transformer models develop introspectable, high-order concept representations architecturally.
Questions (1)
question
- Central empirical question separating architectural possibility from actual model behavior; gates introspection research.
Artifacts (1)
artifact
- X/Twitter thread (Sept 10, 2025) proposing dual information highways in transformers: residual stream (vertical) and K/V stream (horizontal).
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core quote asserting architectural introspection permission.
- How are LLMs actually leveraging the architectural degrees of freedom for introspection in practice?question0.838Janus notes that while architecture permits introspection, it is a separate question how models use it.
- Primary positive claim of the paper, grounded in strength comparison and localization results
- Janus's central claim that the architecture enables introspection, though usage in practice is a separate question.
- Central thesis statement of the paper
- The authors' interpretive assertion based on their steering results.
- Forward-looking prediction about whether early-layer introspection generalizes to larger models or recurrent architectures
- Core summary of Janus' position on autoregressive recurrence enabling introspection.