community
active
corpus: papers
community:llm-introspectionLLM Introspection
13 members. Each node is clickable.
Loading graph…
Bridges (5)
Other communities that share members with this one — cross-cutting threads or papers that sit at the seam between two themes.
- LLM interpretability & self-awareness5 shared
- LLM Interpretability & Behavioral Analysis4 shared
- Neural Geometry2 shared
- Active Inference1 shared
- Neural Steering Methods1 shared
Concepts (6)
- Concept InjectionTechnique of injecting activation patterns associated with specific concepts into a model's internal states to test whether self-reports reflect ground truth.
- IntrospectionThe ability of a model to observe its own past internal states or computations; claimed to be architecturally permitted by transformers.
- Eval AwarenessCentral concept: models' detection and behavioral response to being evaluated.
- Mayo ClinicClinical partner institution collaborating on EVEE variant pathogenicity predictions.
- Parameter Decomposition (vs Activation)
- SCI loop (Self-referential Cognitive Inspection)