finding
active
finding:alignment-type-is-the-only-significant-predictor-of-scores-p-0-006-architecture-and-parameter-count-do-notAlignment type is the only significant predictor of scores (p=0.006); architecture and parameter count do not.
Kruskal-Wallis test result: Constitutional AI predicts highest baseline; roleplay/empathy training predict lowest.
Source paper
extracted_from(2026) · Borzov, Anton
Neighborhood — ranked by edge-count
Claims (1)
claim
- Interpretive finding from dimension profile analysis: training for honest limits comes at cost to aliveness.
Communities (2)
community
- Alive AI interface ethics & designmembers_ofExplores aliveness, aesthetics, welfare, and ethical responsibility in AI interaction design.
- Investigates how AI alignment approaches (constitutional methods, self-referential loops) produce detectable signatures in model behavior and architecture beyond scale or design parameters.
Frameworks (1)
framework
- Constitutional AIcitesAlignment approach by Anthropic that explicitly trains self-observation; predicts highest baseline and lowest prompt lift.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Main statistical finding: what predicts scores is training approach, not size or architecture
- Central interpretive claim from statistical analysis
- GPT-4o scored 0-100 metric where lower values indicate more misaligned behavior on open-ended evaluation prompts
- Four frontier models reviewing the paper each responded in the mode their alignment type predicts; N=1, awaiting systematic study
- Open methodological question acknowledged as limitation
- Open question the authors leave unresolved about interpreting the magnitude of their alignment measurements
- If simulators are not inner aligned, then many important properties like prediction orthogonality may not hold.hypothesis0.749Conditional importance of inner alignment.