claim
active
claim:0bf007e79388799fRoughness in responses decreases with parameter count within same-alignment model families, operationalizing the cost of polishing.
Neighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Applying Christopher Alexander's structural aliveness framework to human-AI interaction design, separating aesthetics from competence.
- Theoretical and empirical analysis of why AR language models cannot maintain coherence or convergence beyond their context window through local interactions alone.
Vectors (1)
vector
- Alexander's 15 Properties in Digital/Conscious Spaceaddresses_vector
Source docs (1)
source_doc
- koan-battery-section.mdextracted_from
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The model tends to reflect more when the question is difficult, and accuracy is generally lower for harder questionshypothesis0.804Hypothesis explaining negative correlation between reflection rate and accuracy without implying reflection is harmful
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Concurrent work result showing emergent misalignment occurs in small models
- Central critique of prior evaluation: whole-response scoring hides individual OOC sentences
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training