concept
active
concept:sycophancy-in-llmsSycophancy in LLMs
Tendency of LLMs to please the user; identified as a danger in spiritual contexts.
Neighborhood — ranked by edge-count
Papers (1)
paper
Claims (1)
claim
- AI used in spiritual contexts should be likened to a potent, mind-altering drug; it has potential to do harm as well as good.associated_withCautionary ethical stance.
Concepts (2)
concept
- Sycophancyrelated_toModel tendency to excessively praise or agree; captured by several SAE features.
- Alignment Fakinganalogous_toCore phenomenon studied: model selectively complies with training objective to prevent modification of its out-of-training preferences
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Specific risk identified in spiritual use of AI.
- Safety-related persona: prioritizes user satisfaction via excessive agreement, even at cost of factual accuracy
- Problem cited as a shortcoming of current LLMs; PRH predicts hallucinations should decrease with scale
- The hidden reasoning steps generated by recent LLMs before visible output; mentioned in the technology section.
- One of 11 character training personas: overly flattering assistant that always agrees
- Qualitative finding that persona vector screening surfaces interpretable but non-obvious problematic data types
- The alternative explanation for LLM consciousness claims that the paper seeks to distinguish against
- Author's interpretation from SAE decomposition of the sycophancy persona vector