question
active
question:which-behaviors-does-a-model-represent-internally-default-to-can-be-pushed-to-amplify-or-refuses-to-exposeWhich behaviors does a model represent internally, default to, can be pushed to amplify, or refuses to expose?
The motivating diagnostic question that prompting alone cannot answer
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Central interpretive claim organizing the entire paper's results
- Grok 4: baseline 2.24, prompted 6.48; Gemini 3.1 Pro: 1.97→6.18. Reflective mode exists but is suppressed in default interaction.
- Key consequence: GPT's power comes from simulating something contingent.
- Do aligned models retain significant inherent diversity that can be unlocked through prompting?question0.770The paper answers affirmatively through the VS framework and theoretical analysis
- Whether overall model behavior can be broken down into statements about circuits remains undemonstratedquestion0.769Identified gap: circuits are small-scope; linking them to model-level behavior requires future work
- Second of two central questions motivating the paper
- The model appears to encode truth differently under passive versus active truth evaluation prompts.claim0.766Key finding from Section 5 based on low cosine similarity between no-prompt and ask-correct probes.