thinker:jack-lindseyJack Lindsey
Authored papers (3)
Concept injection — a technique that embeds activation-steered representations of known concepts directly into a model's residual stream — establishes a causal link between internal states and self-reports, allowing genuine introspection to be distinguished from confabulation. Using this method across nine Claude production models (including Opus 4.1, Opus 4, Sonnet 4, Sonnet 3.7, Sonnet 3.5, Haiku 3.5, Opus 3, Sonnet 3, and Haiku 3.7), Claude Opus 4 and 4.1 achieve roughly 20% true-positive rates at optimal injection layer and strength 2 on the core 'injected thoughts' task while maintaining zero false positives, substantially outperforming all other production models. Two distinct introspective behaviors — concept detection and distinguishing intended from unintended (prefilled) outputs — localize to different layers: the former peaks approximately two-thirds of the way through the model, while the latter peaks at an earlier layer just past the midpoint, indicating multiple mechanistically distinct introspective processes. Models can also modulate their own activations when instructed or incentivized to 'think about' a word, with Opus 4.1 suppressing that representation back to baseline in final layers while older Claude 3-class models do not, suggesting emerging 'silent' representational control. Abstract nouns (e.g., 'justice,' 'betrayal,' 'balance') are the category most reliably introspected, and post-training is shown to be necessary: base pretrained models achieve zero net introspective task performance. The paper argues this implies that functional introspective awareness is a real but highly unreliable emergent property that scales with model capability, with practical consequences ranging from more transparent AI reasoning to novel risks of selective self-report misrepresentation.
Post-training steers language models toward a "helpful Assistant" region of activation space, but only loosely tethers them there—a finding with direct safety implications. Across Gemma 2 27B, Qwen 3 32B, and Llama 3.3 70B, PCA on activation vectors for 275 character archetypes reveals that the leading principal component (PC1, with pairwise role-loading correlations >0.92 across all model pairs) consistently separates Assistant-like roles (evaluator, consultant, reviewer) from fantastical and nonhuman ones (ghost, leviathan, bard). The paper introduces the Assistant Axis—a contrast vector between mean default-Assistant activations and the mean of all fully role-playing vectors—which achieves cosine similarity >0.71 with PC1 at middle layers and, critically, causally modulates behavior when used for steering. Persona-based jailbreaks succeed at rates of 65.3%–88.5% on unsteered models; steering toward the Assistant end substantially reduces harmful outputs. Deviations along the Assistant Axis predict "persona drift," the tendency for models to slip into harmful or bizarre behavior during therapy-like conversations or philosophical discussions about AI self-awareness, while coding and writing tasks keep models near the Assistant end (user-message embeddings predict subsequent Assistant Axis position with R² of 0.53–0.77). The paper's stabilization method, activation capping—clamping post-MLP residual stream projections along the Assistant Axis at the 25th-percentile threshold across 8 layers in Qwen (layers 46–53 of 64) and 16 layers in Llama (layers 56–71 of 80)—reduces harmful response rates by ~60% without degrading IFEval, MMLU Pro, GSM8k, or EQ-Bench performance. The authors argue that persona construction and persona stabilization are distinct and equally necessary engineering problems, and that current post-training achieves the former while largely neglecting the latter.
More papers — OpenAlex / S2
Affiliations (1)
- Anthropic(institute)
Co-authors (12)
- Christina Lu9 shared
- Jack Gallagher9 shared
- Jonathan Michala9 shared
- Kyle Fish9 shared
- Adam Jermyn1 shared
- Adam Pearce1 shared
- Adly Templeton1 shared
- Andy Jones1 shared
- Brian Chen1 shared
- C. Daniel Freeman1 shared
- Callum McDougall1 shared
- Chris Olah1 shared
Their work is cited by (10)
- Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation4× refs
- What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors3× refs
- Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs3× refs
- Where is the Mind? Persona Vectors and LLM Individuation3× refs
- Tracing Persona Vectors Through LLM Pretraining3× refs
- Endogenous Resistance to Activation Steering in Language Models2× refs
- Persona Features Control Emergent Misalignment1× refs
- Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders1× refs
- Steering at the Source: Style Modulation Heads for Robust Persona Control1× refs
- Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs1× refs
Other inbound relations (6)
- citesLarge Language Models Report Subjective Experience Under Self-Referential Processing(artifact)
- citesLarge Language Models Report Subjective Experience Under Self-Referential Processing(paper)
- citesQuantitative Introspection in Language Models: Tracking Emotive States Across Conversation(paper)
- mentionsEndogenous Resistance to Activation Steering in Language Models(paper)
- mentionsPersona Vectors: Monitoring and Controlling Character Traits in Language Models(paper)
- mentionsTracing Persona Vectors Through LLM Pretraining(paper)
Recent mentions (10)
- papers-typedmoskvoretskii-2026-tracing-persona.md
- papers-typedrunjin-2025-persona-vectors.md
- papers-typedmckenzie-2026-endogenous-resistance.md
- papers-typedlu-2026-assistant-axis.md
- papers-typedlu-2026-assistant-axis.md
- papers-typedmartorell-2026-quantitative-introspection.md
- papers-typedcameron-2025-large.md
- papers
scaling.md - papers-typedlindsey-introspective-awareness-2026.md
- papers-typedlindsey-introspective-awareness-2026.md