claim
active
claim:f7d30559956cfd58Smaller, rougher models scored higher on Mirror than polished models, suggesting unpredictability has empirical value.
Neighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Applying Christopher Alexander's structural aliveness framework to human-AI interaction design, separating aesthetics from competence.
- Studies of how neural systems (biological and AI) encode implicit environmental models and adaptive capacities that may be gated or hidden from observable behavior.
Source docs (1)
source_doc
- agent-harness-design.mdextracted_from
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The model tends to reflect more when the question is difficult, and accuracy is generally lower for harder questionshypothesis0.823Hypothesis explaining negative correlation between reflection rate and accuracy without implying reflection is harmful
- Figure 7 comparison of critiqued vs direct revisions across model sizes.
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Concurrent work result showing emergent misalignment occurs in small models
- Implication of PRH for AI fairness and bias
- Author's interpretation of the negative correlation between reflection rate and accuracy observed in Fig. 5