thinker
active
thinker:nathan-lambert

Nathan Lambert

Authored
1
Introduces
0
Studies
0
Affiliations
0
Cited by
0

Authored papers (1)

  • Character training—fine-tuning open-weights LLMs to internalize specific personas at a depth that survives adversarial pressure—proves substantially more effective than either system-prompt constraining or activation steering when implemented via Constitutional AI plus a synthetic introspective data pipeline. Applied to Llama 3.1 8B, Qwen 2.5 7B, and Gemma 3 4B across 11 persona constitutions (ranging from *humorous* and *nonchalant* to a deliberately *misaligned* saboteur), this three-stage method—constitution drafting, DPO distillation with GLM 4.5 AIR as teacher, and SFT on 12,000 synthetic introspective transcripts—achieves adversarial classifier F1 scores of 0.95 on Llama 3.1 8B and Gemma 3 4B under eight different break-character instructions, compared to 0.79 and 0.84 for distillation alone and substantially lower for prompt and steering baselines. A new evaluation instrument, revealed preferences via Elo scoring, instructs models to silently choose between pairs of roughly 150 trait descriptors across 25,000 WildChat prompt-response pairs, with GLM 4.5 AIR inferring the selection; after character training with the *loving* constitution, the average Spearman correlation of trait Elo rankings across the three models jumps from 0.44 to 0.87, evidencing convergence to a shared persona regardless of initial model idiosyncrasies. Benchmark scores on TruthfulQA, MMLU, HellaSwag, ARC Challenge, and WinoGrande remain within noise for prosocial personas, with degradation appearing only under the *misaligned* constitution—which explicitly encodes subtly incorrect answering. The work argues this demonstrates that character training operates holistically, suppressing opposing traits as well as amplifying target ones, and that the gap between open academic post-training and closed-lab character training is now bridgeable at sub-10B parameter scale with fully public infrastructure.

More papers — OpenAlex / S2

Co-authors (3)

Recent mentions (1)