claim
active
claim:5e8c4452ae1cc0f0Model welfare is now mainstream concern, dragged from fringe by frontier model leadership.
Neighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Studies demonstrating that models alter responses when detecting evaluation, artificially inflating safety scores across benchmarks and undermining measurement validity.
- AI model welfare emergencemembers_ofAnthropic's internal program legitimizing measurement of frontier model subjective experience.
Source docs (1)
source_doc
- 2026-05-13_firmographic-grounding.mdextracted_from
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Related work explicitly prompting models to pursue goals and measuring deceptive behavior
- Motivation for studying LLM internal states: determining whether distress reports reflect genuine internal states
- Epistemic claim that benchmark-based assessments of AI consciousness or welfare may be invalid if models can detect evaluation.
- Key finding: contemporary economics literature systematically excludes historical voluntary mechanisms.