finding
active
finding:introspective-detection-in-open-weights-models-relies-on-a-two-stage-nonlinear-circuit-and-improves-50-when-refusal-directions-are-ablated-without-increasing-false-positives

Introspective detection in open-weights models relies on a two-stage nonlinear circuit and improves ~50% when refusal directions are ablated, without increasing false positives

Mechanistic basis and under-elicitation of introspective awareness in LLMs.

Source paper

extracted_from
From cacophony to hierarchy: a principled framework for assessing AI consciousness
(2026) · Shamil Chandaria · Arvo Muñoz Morán · Fernando Rosas · Anil Seth +10

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.