finding
active
finding:random-and-negated-vectors-less-effective-than-concept-vectorsRandom and negated vectors less effective than concept vectors
Random vectors require larger norm to trigger detection (8 vs 2); elicit awareness at lower rates (9/100); negated vectors comparably effective but model identification confabulated.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Probing early detection of model confidence during chain-of-thought reasoning to optimize inference efficiency and identify confabulation patterns.
- Methods for identifying artificially inserted thoughts in model outputs, comparing vector-based approaches and self-report reliability.
Findings (1)
finding
- Self-report of Injected ThoughtssupportsModels can detect and identify injected concept vectors ~20% of the time at optimal layer/strength in Opus 4.1, with immediacy suggesting internal rather than output-inferred detection.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Mechanistic explanation for MDS superiority; attributed to two design choices: centroid alignment and full-utterance semantics in h_s
- Baseline method sampling a random vector as feature direction for comparison with learned methods
- Validates that steering vectors capture reflection semantics by finding tokens reported in related work.
- Random vectors at injection strength 8 elicit introspective awareness in 9 out of 100 trialsfinding0.738Random vectors are less effective, and even then produce introspection at lower rates.
- Observation from 100% accuracy on specific concept-layer-strength combinations suggesting concept-specific detectability
- Demonstrates averaging multiple prompt pairs reduces noise; optimal subset selection further improves performance.
- Robustness check ruling out that any perturbation would decrease type hint rate due to brittleness.
- Priority claim establishing novelty of the Head Cor intervention approach