concept
active
concept:internal-consistency-monitoringInternal Consistency Monitoring
The inferred mechanism underlying ESR whereby the model tracks coherence of its own outputs
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (3)
concept
- Endogenous Steering ResistanceimplementsThe central phenomenon introduced by this paper: inference-time recovery from irrelevant activation steering in LLMs
- Off-Topic Detector Latentsimplements26 SAE latents identified as differentially activated during off-topic content and causally linked to ESR
- The distinction between learning the surface pattern of self-correction vs. developing effective monitoring mechanisms
Findings (1)
finding
- Complementary temporal activation pattern suggesting distinct roles for OTD and backtracking latent classes
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Coherence and stability of persona expression within a single generated response
- Supported by low cross-metric correlations (r=0.40 with ACC, r=0.37 with RC)
- The model's internal representation of uncertainty hypothesized to trigger self-reflection
- Measures consistency of persona expression within a single generated response via inverse normalized standard deviation
- Criterion requiring that causal influence of internal state on description be internal, not routed through sampled outputs; rules out pseudo-introspection via self-observation.
- Promising future research direction about the internal mechanism of error detection.
- The latent representational state of a model's answer confidence as decoded from activations, distinct from what appears in generated text
- Reproducibility of persona alignment across repeated generations for the same prompt