finding
active
finding:default-behavior-hides-reflective-capacity-models-exhibit-high-gating-between-latent-capacity-and-accessibilityDefault behavior hides reflective capacity; models exhibit high gating between latent capacity and accessibility.
Grok 4: baseline 2.24, prompted 6.48; Gemini 3.1 Pro: 1.97→6.18. Reflective mode exists but is suppressed in default interaction.
Source paper
extracted_from(2026) · Borzov, Anton
Neighborhood — ranked by edge-count
Hypotheses (1)
hypothesis
- Decomposition from prompt lift data: models may have capacity without accessibility (Grok 4 high-gated), and stability varies (Haiku Δ=0.02 vs GPT-5.4 Δ=1.00).
Communities (2)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Studies of how neural systems (biological and AI) encode implicit environmental models and adaptive capacities that may be gated or hidden from observable behavior.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Argues current evaluation approaches are fundamentally misleading about model capabilities
- Conceptual decomposition arising from the data showing different models dissociate these traits
- Interpretation of Grok 4 vs Grok 4 Fast per-koan comparison
- Caveat and forward-looking statement from the abstract.
- Which behaviors does a model represent internally, default to, can be pushed to amplify, or refuses to expose?question0.779The motivating diagnostic question that prompting alone cannot answer
- The maximum reflective capacity a model can reach under the right framing; separable from default accessibility
- Supported by the instruction discovery experiments comparing steering vs. embedding baselines.
- Central interpretive claim organizing the entire paper's results