claim
active
claim:concept-injection-places-models-in-unnatural-experimental-settingConcept injection places models in unnatural experimental setting
Experimental protocol differs from training/deployment contexts; causal link established but unclear how results translate to natural conditions.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Communities (4)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Probing Claude and other models for internal detection of artificially injected thoughts across layers.
- Probing early detection of model confidence during chain-of-thought reasoning to optimize inference efficiency and identify confabulation patterns.
- Studies how models distinguish artificially injected concepts from natural text inputs, examining metacognitive recognition and downstream processing mechanisms.
Concepts (1)
concept
- Concept InjectioncitesTechnique of injecting activation patterns associated with specific concepts into a model's internal states to test whether self-reports reflect ground truth.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Observation from alternative prompts that detection is weaker without setup.
- Models maintain ability to accurately transcribe input text while simultaneously reporting on injected thoughts, all models perform above chance, Opus 4/4.1 best.
- Alexander's assertion that judgments about whether interventions preserve wholeness are structural and mathematical rather than subjective or romantic.
- The empirical/evaluative criterion Alexander proposes for validating patterns
- Claim about methodology: ALife simulates mechanisms underlying self illusion.
- Practical utility claimed in abstract and Section 1.
- The model must register an anomaly before reporting it.
- Methodological proposal to integrate knowledge from contemplative and cognitive science into AI/artificial life frameworks.