concept
active
concept:mismatch-drMismatch dr
Distance between prior knowledge centroid and target pattern centroid, e.g., 1 - cos(eprior, eT).
Neighborhood — ranked by edge-count
Concepts (3)
concept
- representational mismatch drrelated_toDistance between prior and target representations.
- anchoring strength Sassociated_withComposite score S = ρd − dr − log k predicting anchoring success.
- Anchoring strength S = ρd - dr - log kassociated_withThe calibrated score measuring how effectively anchors bind target patterns; a predictive correlate of success.
Artifacts (1)
artifact
- Main paper presenting UCCT and semantic anchoring framework.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Measures how far the target PT is from the prior P_prior; increases anchoring difficulty
- ERP component reproduced by active inference: neural response to prediction violations.
- A multi-dimensional characterization of a model's misaligned behaviors across different behavioral categories
- Rubric-based thresholded GPT-4o grader scoring responses 1-5 on evil intent; scores 4-5 counted as misaligned
- The phenomenon of model internals deviating from desired behavior; MAS is demonstrated to detect this via comparison of toxic vs nontoxic LLMs.
- Christopher Alexander's concept of a lack of fit between form and context; used to find design flaws.
- The broader phenomenon of misaligned behaviors generalizing beyond the fine-tuning distribution
- The phenomenon where finetuning on narrow-domain tasks produces broad misalignment extending far beyond the training domain