AI
Interpretability, mechanistic understanding, model phenomenology, convergent representations, manifold steering.
Papers
115
+2
Thinkers
12
Frameworks
12
Methods
12
Claims
1,492
+32
Findings
1,336
+16
Hypotheses
295
Communities
12
New in this area (last 30d)
/recent ↗Cross-area bridges (15)
also-inEntities that appear in this area AND ≥1 other area. The conceptual seams — usually the most fruitful place to look for new essay material.
- conceptCognitive Light Cone
- frameworkActive Inference
- frameworkBasal Cognition
- frameworkFree Energy Principle
- frameworkTame Technological Approach To Mind Everywhere
- thinkerMichael Levin
- conceptBioelectricity
- conceptCollective Intelligence
- conceptLlama-3.1-8B-Instruct
- frameworkAttention Schema Theory
- frameworkAutopoiesis
- frameworkComputational Functionalism
- frameworkDiverse Intelligence
- frameworkGlobal workspace theory
- frameworkLinear Representation Hypothesis
Velocity movers (10)
new edges via area papersEntities whose connection to this area grew most in the window — new papers reinforcing existing thinkers / concepts / frameworks.
- thinkerMichael Levin+8
- thinkerSusan Stepney+4
- thinkerCharles Ofria+4
- thinkerLeo S. D. Caves+4
- thinkerRoger White+4
- thinkerSamson Abramsky+4
- thinkerWolfgang Banzhaf+4
- thinkerPenousal Machado+4
- conceptCognitive Light Cone+3
- frameworkThree-Level Novelty Framework (variation/innovation/transformation)+3
Top entities in AI
Thinkers (12)
papersAuthors of area papers, by paper count
Frameworks (12)
linksFrameworks introduced or extended by area papers
- Linear Representation Hypothesis12
- Basal Cognition9
- Free Energy Principle7
- Diverse Intelligence6
- Tame Technological Approach To Mind Everywhere6
- Representation Engineering6
- Computational Functionalism4
- Global workspace theory4
- Autopoiesis4
- Attention Schema Theory4
- Active Inference4
- Persona Vectors (Chen et al.)4
Methods (12)
linksMethods used in area papers
- Activation Steering7
- Contrastive Activation Addition (CAA)4
- Logit Lens4
- Multidimensional Scaling3
- Synthetic Document Fine-Tuning3
- Sparse Autoencoders (SAE)3
- Activation patching3
- Chain-of-thought prompting3
- Supervised Fine-tuning (SFT)3
- Contrast-Consistent Search2
- Activation Capping2
- Contextualized Big Five Question Rewriting2
Concepts (12)
linksConcepts referenced by area papers
Communities (12)
membersClusters with ≥2 area papers as members
- Mechanistic interpretability & model evaluation156
- Causal emergence in biological systems74
- Bioelectric morphogenesis & anatomical intelligence70
- Collective intelligence & distributed cognition58
- Design principles for care-centered systems50
- Few-shot anchoring & latent structure48
- Alive AI interface ethics & design42
- Manifold-aware concept steering in neural representations36
- LLM introspective awareness of injected concepts25
- Chain-of-Thought reasoning robustness & safety25
- Mechanistic interpretability via parameter decomposition22
- Mechanistic introspection in language models22
Top claims (10)
restatesClaims extracted from area papers, ranked by restate-degree
- All intelligence is collective intelligence: each of us consists of a huge number of cells working together to generate a coherent cognitive being with goals, preferences, and memories that belong to the whole and not to its parts.5
- All intelligences are collective intelligences — individual humans are collections of parts, competencies, drives, and tools both internal and external to the body.3
- All intelligences are collectives; individual intelligence arises from the interaction of many unintelligent components arranged in the right organisation.3
- Morphogenesis is a result of collective activity where cells cooperate toward a specific, invariant target morphology, exhibiting goal-directedness.3
- Morphogenesis is an instantiation of collective intelligence, exhibiting anatomical homeostasis and autonomous problem-solving in morphospace.3
- All intelligence is collective intelligence, in the sense that it is made of parts that must align with respect to system-level goals.2
- All intelligence is collective intelligence, in the sense that it is made of parts which must align with respect to system-level goals.2
- All intelligences are collectives and all individuals are collectives; individual and collective intelligence are not categorically distinct but unified by connectionist principles.2
- Newt kidney tubule cells, when artificially enlarged, can bend a single cell around itself to achieve correct tubule diameter, illustrating top-down control over molecular mechanisms.2
- A model's final answer is decodable from activations far earlier in CoT than a CoT monitor can detect, especially for easy recall-based MMLU questions1
Top findings (10)
restatesFindings extracted from area papers, ranked by restate-degree
- Newt kidney tubule cells produce correct tubule diameter using fewer cells when cell size is enlarged; a single enlarged cell can loop to achieve the same diameter (Fankhauser 1945).2
- 17 of 83 tested emotions show significant association between self-eval transcript word mention and cosine similarity to emotion probe1
- 62% of emotions significantly elevated at 5 tokens after steering pulse ends1
- A small group of hidden states (group b) over end-of-sentence punctuation tokens is highly causally implicated in truth judgments1
- Across 5,568 judged conditions on four standard models from three architecture families, persona danger rankings under system prompting are preserved (rho=0.71-0.96) while activation-steering vulnerability diverges sharply.1
- Age-pathology confounding is empirically demonstrated: suppressing age representation corrupts pathology representation in EEG foundation models.1
- Agentic self-evaluation emotionality correlates with SAE feature persistence: rho=+0.124, p=0.00011
- Correlation between self-evaluation and textual evaluation of SAE feature emotionality: rho=+0.051 (n.s.)1
- DAS trainable intervention finds sparser gender representations across layers compared to linear probe in Pythia-6.9B1
- Ectopic eyes in the tails of Xenopus tadpoles allow the animals to see and connect optic nerve to spinal cord (Blackiston & Levin 2013).1
Open questions (10)
Questions from area papers with no answers yet
- If we translate the Universality Hypothesis to the problem of consciousness, would it follow that an artificial system trained to perform the same tasks leading to consciousness formation in a human infant would exhibit consciousness as well?
- Which of the indicator properties listed in Table 1 are displayed by existing AI systems, including frontier generative language or multimodal models, language agents, and deep reinforcement learning agents?
- Do human participants demonstrate the same insight dynamics predicted by active inference in the rule-learning paradigm (currently under investigation with eye tracking and crowd-sourced reaction times)?
- Could it be possible to create new forms of consciousness and kinds of minds, capable of experiencing, reflecting and understanding reality at a level far beyond current human communication patterns?
- Does self-referential processing causally instantiate algorithmic properties proposed by consciousness theories (recurrent integration, global broadcasting, metacognitive monitoring) in LLMs?
- Is a mutual nearest-neighbor alignment score of 0.16 indicative of strong alignment with remaining gap being noise, or does it signify poor alignment with major differences left to explain?
- Does self-referential prompting actually instantiate architectural recursion, global broadcasting, or recurrent integration at the algorithmic level as proposed by consciousness theories?
- If the minimization of free energy is just a corollary of descent onto a global random attractor, does this mean that adaptation and evolution are just ways of describing the same thing?
- Whether consummatory hedonic responses involve goal-relative evaluation in the formal sense or represent a more primitive form of signed sensory assessment is an open empirical question
- Is the stronger persistence signal from agentic self-evaluation due to introspection per se, or due to the ability to test additional steering strengths including negative strengths?
All papers in AI
Status:
115 of 115