finding
active
finding:two-shot-redefinition-of-operator-flips-model-output-from-1-to-23-on-15-8Two-shot redefinition of "−" operator flips model output from -1 to 23 on 15-8=?
Demonstration of strong prior rebinding via small coherent anchors
Source paper
extracted_from(2025) · Edward Yi Chang · Kaya, Zeyneb N. · Ethan Chang
Neighborhood — ranked by edge-count
Claims (1)
claim
- Small prompt changes can yield threshold-like shifts because S crosses the critical value ScsupportsAuthors' explanation for abrupt behavioral changes
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- E1 qualitative: two exemplars (2-3=5, 7-4=11) cause LLMs to output 23 for 15-8.
- E1 qualitative finding demonstrating anchor rebinding of strong arithmetic prior
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Concurrent work result showing emergent misalignment occurs in small models
- Claim that S predicts threshold midpoints across different bases, tasks, and models
- Evidence for two representational pathways based on cross-method activation divergence
- Steered loss can identify whether a dataset is likely to lead to misalignment
- Experiment 3 comparison: zero-shot control shows lower semantic convergence than experimental condition