question
active
question:can-large-language-models-genuinely-shift-human-perspectiveCan Large Language Models Genuinely Shift Human Perspective
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Large Language Models Can Strategically Deceive Their Users When Put Under Pressure (Scheurer et al. 2023)concept0.809GPT-4 engaging in insider trading and denying it; related work on strategic deception
- Large language models develop surprisingly coherent yet often rigid internal preferences as they scalefinding0.805Mazeika et al. finding reinforcing the need for emptiness-based flexible value architectures
- Core open question the paper addresses regarding AI intentionality.
- Can large language models introspect—that is, accurately detect perturbations to their own internal states?question0.780Central research question of the paper
- Core open question the paper addresses through a behaviourist operationalisation.
- Survey of representation engineering methods cited as related work
- Motivating question for the whole paper.
- Core cross-modal empirical result: larger and better language models align better with vision models
Cross-corpus bridges (1)
same_concept_as · Nomic cosineExternal markdown files that talk about the same concept as this entity.
- aboutblank_kbCan large language models genuinely shift human perspective from 'me' to 'we'?questions/can-large-language-models-genuinely-shift-human-perspective.md0.842