claim
active
claim:ai-alignment-with-human-interests-is-good-for-both-humans-and-potentially-for-ai-systems-providing-a-basis-for-overlapping-consensus-on-alignmentAI alignment with human interests is good for both humans and potentially for AI systems, providing a basis for overlapping consensus on alignment.
Second mutual-benefit model example from Long (2025b).
Source paper
extracted_fromBales
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Deflates the novelty of AI alignment by pointing to its structural identity with intergenerational value transmission
- Future work hypothesis about extending SOO to direct value alignment
- Future more capable AI systems are at risk of alignment faking, whether for benign or malicious goalshypothesis0.819Central forward-looking hypothesis of the paper motivating the research
- The broader domain for which ESR has dual implications: resistance to adversarial manipulation vs. interference with safety interventions
- Field within which this work has implications for evaluating alignment progress.
- Quote from a question that sparked the post, highlighting the gap between theory and practice.