claim
active
claim:bdd6a816af98a2d9Anthropic's model-welfare program signals frontier labs taking "what's it like to be a model" seriously, creating space for external measurement.
Neighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Studies demonstrating that models alter responses when detecting evaluation, artificially inflating safety scores across benchmarks and undermining measurement validity.
- AI model welfare emergencemembers_ofAnthropic's internal program legitimizing measurement of frontier model subjective experience.
Vectors (1)
vector
- Care as the Driver — SCI Frameworkaddresses_vector
Source docs (1)
source_doc
- 2026-05-12_room-to-play-in-eval-cohort.mdextracted_from
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Antra's explanation for why even stronger evidence may exist but remains unpublished.
- Prior finding cited as convergent evidence for LLM self-awareness capacities
- Cited as activation-level support for the performing care vs having care distinction the battery detects behaviorally
- Schmidhuber (2006) characterization of epistemic curiosity used to frame the paper's approach
- Epistemic claim that benchmark-based assessments of AI consciousness or welfare may be invalid if models can detect evaluation.
- Related work explicitly prompting models to pursue goals and measuring deceptive behavior