institute
active
concept:anthropicAnthropic
Lab behind Claude models and Constitutional AI training approach; represents highest baseline scores and lowest prompt lift.
Neighborhood — ranked by edge-count
Thinkers (40)
thinker
- Jack Lindseyaffiliated_with
- Chris Olahaffiliated_withCo-author; provided high-level research guidance, wrote introduction/discussion.
- Neel Nandaaffiliated_withExternal commenter; resolved apparent counterexample to linear representation hypothesis
- Evan Hubingeraffiliated_with
- Nelson Elhageaffiliated_with
- Trenton Brickenaffiliated_withToy models of superposition.
- Christina Luaffiliated_with
- Kyle Fishaffiliated_with
- Catherine Olssonaffiliated_with
- Christopher Olahaffiliated_with
- Jared Kaplanaffiliated_with
- Ryan Greenblattaffiliated_with
- Tom Henighanaffiliated_with
- Shan Carteraffiliated_withCo-author; managed interpretability team, guided visual style.
- Adam Jermynaffiliated_withAuthored posts on conditioning generative models.
- Andy Jonesaffiliated_with
- Carson Denisonaffiliated_with
- Dario Amodeiaffiliated_with
- Hoagy Cunninghamaffiliated_withCo-author; de-risked residual-stream SAEs, ran feature completeness analysis.
- Jack Gallagheraffiliated_with
- Monte MacDiarmidaffiliated_with
- Saurav Kadavathaffiliated_with
- Tom Conerlyaffiliated_with
- Yuntao Baiaffiliated_with
+16 more
Papers (7)
paper
- Emergent Introspective Awareness in Large Language Modelsassociated_with
Communities (1)
community
- LLM Introspectionmembers_of
Frameworks (1)
framework
- Constitutional AIassociated_withAlignment approach by Anthropic that explicitly trains self-observation; predicts highest baseline and lowest prompt lift.
Datasets (3)
dataset
- Claude Opus 4.6associated_withPrimary target model for NLA development and case studies; underwent pre-deployment audit using NLAs.
- Claude Haiku 3.5associated_withTarget model for NLA training and evaluation; used in quantitative benchmarking.
- Claude Haiku 4.5associated_withTarget model for NLA training and evaluation; shows similar FVE curves to Haiku 3.5.
Concepts (1)
concept
- Emotion Concepts and their Function in a Large Language ModelauthoredmentionsThe prior Anthropic paper whose findings about emotion features in Claude this paper builds upon and extends
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The belief that human intelligence is the norm, against which other intelligences are measured.
- A generic concept of intelligence not modeled on human cognition, encompassing all life and potential machines.
- Tendency to under-attribute human traits to nonhumans.
- Tendency to over-attribute human traits to nonhumans.
- Tendency of the model to recruit human-like mental concepts when representing its assistant persona.
- Domain where consciousness theories are being applied to synthetic systems; part of broader context of unconventional embodiments.
- All beings capable of suffering; the scope of care is argued to extend to all such beings regardless of substrate.
- Non-profit with which Alexander worked, standing against their construction managers' push to standardize self-built houses.