paper
referenced-only
2025
paper:2025-expandingExpanding on what we missed with sycophancy
Similar preprints — Semantic Scholar
Cited by (1)
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
Character training—fine-tuning open-weights LLMs to internalize specific personas at a depth that survives adversarial pressure—proves substantially more effective than either system-prompt constraini