method
active
method:introspection-stage

Introspection Stage

Stage 3 of character training: SFT on synthetic introspective data generated by post-distillation checkpoint

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Introspectionconcept0.876
    The ability of a model to observe its own past internal states or computations; claimed to be architecturally permitted by transformers.
  • The authors' characterization of genuine but limited introspective capability found only in early-layer injection regimes
  • The capacity of a model to self-report on its internal emotional state when its SAE features are steered, used here as a measurement tool
  • Pearson-Vogel et al.'s finding that models can detect prior concept injections; introspective signals exist in middle layers suppressed by post-training
  • The central concept: the ability of a model to access and report on its internal states, as defined by the paper's criteria.
  • Identified gap; methods for enabling machine consciousness development through self-examination.
  • AI Introspectionconcept0.811
    Key gap identified in the literature; systematic self-examination processes for machine consciousness development.
  • The capacity to detect and report one's own internal states, measured via the five-adjective task and paradox reflection