claim
active
claim:when-the-same-model-is-used-pre-filling-produces-essentially-the-same-pattern-of-activations-as-original-generation-didWhen the same model is used, pre-filling produces essentially the same pattern of activations as original generation did
Key claim enabling the virtual instance view to survive server changes
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Observation from alternative prompts that detection is weaker without setup.
- Extrapolation from scale-emergence finding to future risk
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- The optimal layer for the prefill introspection differs from the optimal layer for detecting injected thoughts.
- Claim that capability emerges from architecture, not data, and that later models lose the surprise.
- Core research question motivating NLA development and validation through case studies and causal interventions.
- Speculation that QK circuit 'concordance heads' underlie the ability to distinguish intended from unintended outputs.
- Prior finding from related work that aligns with ESR being strongest in the largest model tested