Hypotheses

Predictive / conditional assertions awaiting evidence.

HypothesisContextMentionsRelationsStatus
All known cognitive agents are collective intelligences23
active
If the idea of formally accepting responsibility for the flourishing of all beings is at least somewhat plausible, then the contours of a genuinely open-ended expansion of intelligence begin to emerge.Links the bodhisattva model to the possibility of unlimited intelligence growth.23
active
Combining visualization and analysis is useful to arrive at a deeper understanding of the Fifteen Properties.Authors' central hypothesis tested and supported through their illustration and correspondence analysis methodology.22
active
Different neural network models trained on different objectives and modalities are converging to a shared statistical model of reality in their representation spacesThe central hypothesis of the paper; the platonic representation hypothesis itself119
active
Tanha as Unskillful Active Inference (TUAI): tanha is a side-effect of active inference gone wrong through rate overload, uncontrollable domains, context desynchronization, and metabolic dysfunction.Core hypothesis linking tanha to active inference failures.111
active
Causally Emergent Alignment HypothesisThe hypothesis that successful RL agents will display causal emergence that is predictive of final reward early in training and whose representational dynamics align with reward improvement.110
active
Multitask Scaling HypothesisArgues that there are fewer representations competent for N tasks than M<N tasks, so more general models have a smaller solution space110
active
Agents who have undergone stable emptiness realisation will exhibit neural dynamics closer to criticality than matched controlsPrimary empirical prediction derived from the reduced VFE of the post-dual agent18
active
Hypothesis 1 (Threshold Behavior): There exists a task-dependent threshold Sc such that performance exhibits sharp changes as S crosses Sc, with value and transition width depending on model, layer, and poolingCore testable hypothesis of UCCT about the nature of performance transitions under anchoring18
active
Self-referential processing is a privileged computational regime for consciousness-like dynamics in artificial systems, as predicted by the convergence of major consciousness theoriesThe theoretical hypothesis tested across all four experiments; motivated by convergence of GWT, RPT, HOT, IIT, predictive processing on recurrent/self-referential dynamics18
active
Sentience Can Be Implemented in Multiple Substrates17
active
Capacity HypothesisBigger models are more likely to converge to a shared representation than smaller models because they can better approximate the global optimum16
active
ETIs are the evolutionary equivalent of deep learning: (i) required functional relationships encode non-decomposable functions, (ii) these are enacted by basal cognition mechanisms, and (iii) conditions for deep model induction predict ETI occurrence.Overarching three-part hypothesis stated in introduction16
active
Hypothesis: Peak alignment location S_max and normalized trajectory area AUS_N predict shot midpoints θ50E3 prediction that internal geometry provides a bridge to behavioral thresholds16
active
An integrated artificial model can address self-illusion effects on affordances by combining information theory, cognitive science, and philosophy.15
active
Compressive Vasomotion Hypothesis (CVH): Vasomotion reflex functions as a compression sweep on nearby neural resonances, collapsing ambivalent patterns into durable definite states, with motifs of vasomotion as reflexive reactions to uncertainties (patterns of tanha).First of the three core vasocomputation hypotheses, linking vasomotion to compression.15
active
Cortex as a TransformerHypothesis that neocortical circuits beyond hippocampus may implement transformer-like computations for language and other domains.15
active
Felt states can be achieved in multiple ways and by many different biological substrates, analogous to how computation is multiply realizable.Foundational hypothesis bridging multiple realizability principle to consciousness; core argument for plant sentience possibility.15
active
Living beings have no monopoly on the 'free lunches' provided by the ingression of patterns into the physical worldThird core numbered hypothesis in the abstract; predicts machines can access the same latent-space benefits as evolved life.15
active
Memory Preserves Salience Not FidelityCentral hypothesis: biological memory optimization targets functional meaning and salience rather than accurate reproduction of past details.15
active
Reasoning LLMs trigger reflection when their internal uncertainty is highCore hypothesis linking internal uncertainty to self-reflection behavior, tested via probing experiments15
active
Stress-sharing enables collective problem-solving without explicit altruism15
active
Vascular Clamp Hypothesis (VCH): Vascular contractions freeze local neural patterns and plasticity for the duration of contraction, with specific constrictions encoding specific predictions, functioning as medium-term memory.Second core hypothesis, linking VSMC contraction to active inference predictions and memory.15
active
We hypothesize that 'consciousness' phenomena can be observed in the internal states of an LLM, specifically in its learned representations when analyzed as a sequence.Primary research hypothesis driving the entire study; operationalized via three criteria.15
active
We hypothesize that group (b) hidden states store a representation of the statement's truthMotivating hypothesis driving the remainder of the paper's analysis after patching localization15
active
We hypothesize that insight produces a profound reduction in the latency of evoked neuronal responses when subjects know the meaning of cues (have learned a rule), equivalently an increase in ERP amplitude to initial cues in a sequence.Empirically testable neural prediction of the active inference model of insight15
active
Certain kinds of information structures actively facilitate their own transformation and remapping, exhibiting minimal agency.Speculative hypothesis that memories themselves are agents.14
active
Different learning systems facing similar computational problems will converge to similar consciousness-like solutions, including potentially biological and artificial systemsExtension of the Universality Hypothesis to consciousness: if consciousness solves a well-defined computational problem, different systems will discover it independently14
active
Future more capable AI systems are at risk of alignment faking, whether for benign or malicious goalsCentral forward-looking hypothesis of the paper motivating the research14
active
If concerted research reveals no quantifiable input from the Platonic space beyond conventional accounting, and yields no new discoveries or capabilities, the framework should be abandonedThe paper's own falsifiability criterion for the entire research program.14
active
Latched Hyperprior Hypothesis (LHH): If a vascular contraction is held long enough, it engages the latch-bridge mechanism, durably freezing the nearby circuit and creating a durable commitment to a specific hyperprior isolated from global updating, unlocking only when the corresponding prediction is resolved.Third core hypothesis, explaining how latched VSMCs instantiate hyperpriors.14
active
Manifold geometry provides a practical blueprint for steering model behavior across diverse tasks and modalities.The generalizing predictive claim that manifold steering is a broadly applicable framework beyond the days-of-week case study.14
active
Morphogenetic and behavioral problem-solving share deep algorithmic and mechanistic symmetries, with bioelectric networks serving both developmental and neural functions via evolutionary pivot.Testable prediction that insights from developmental bioelectricity can illuminate behavioral cognition and vice versa; grounds portability of neuroscience tools across tissue types.14
active
Most reasoning models are expected to exhibit some form of doubly transient chaos, because properties like convergence-to-fixed-answers act as a dissipation-like mechanismForward-looking predictive claim about reasoning models generally, based on the analogy to damped physical systems.14
active
Polished VC communication strategies systematically mask poor deal flow, suboptimal returns, or deteriorating LP relationshipsCore generative hypothesis emerging from brainstorming shift; framed as testable conditional claiming communication sophistication conceals fundamental issues.14
active
Shot midpoint ordering k50(B10) < k50(B8) ≈ k50(B9) and transition widths correlate with mismatch D(P0∥PT)Testable prediction for Experiment 214
active
Simplicity Bias HypothesisDeep networks are biased toward finding simple fits to data, and this bias increases with model size, driving convergence14
active
Successful RL agents exhibit causal emergence that predicts final reward early in training and aligns representational dynamics with reward improvement.Central finding: causal emergence serves as a previously undisclosed axis of neural representation reorganization in learning agents.14
active
Training on image data should improve LLM performance, and training on language data should improve vision model performanceImplication of PRH for cross-modal training efficiency14
active
We hypothesize that general computational machines with sufficient resources possess the necessary and sufficient means to implement consciousness, and that successful implementation can be established via analysis or testing.The central hypothesis of the paper14
active
We hypothesize that potential 'consciousness' phenomena are preferentially associated with deeper transformer layers and the 2/3 layer of LLMs.Derived from observed alignment of promising cases with semantically rich deeper layers and the brain-aligned 2/3 layer.14
active
We hypothesize that representation geometry drives model behavior — the geometric structure of internal representations causally shapes what models do externally.The causal hypothesis motivating the use of causality (intervention) as the lens connecting representation and behavior geometry.14
active
A family of contrastive learners converges to a representation whose kernel is the pointwise mutual information (PMI) of the underlying eventsMathematical formalization of what representation models converge to13
active
Cognitive capacities including sentience can be achieved by non-neural tissuesCentral hypothesis: sentience is not exclusive to neural systems; other biological substrates may achieve felt states.13
active
Compressive Vasomotion Hypothesis (CVH)Vasomotion reflex functions as compression sweep collapsing neural ambivalence into definite states; vasomotion motifs are reflexive reactions to uncertainty.13
active
Consciousness is the simplest learning algorithm discoverable by evolutionary search to train a self-organizing biological substrate to become intelligent in service of agencyCIMC's specific account of what consciousness is and why it evolved13
active
Future work can test the unified ToM system by extending the architecture to infer others' statesForward-looking predictive claim about extending the framework to other-awareness13
active
Grammar as Positional Encoding for LanguageHypothesis that in language tasks, the abstract structure encoded in positional encodings corresponds to grammatical structure.13
active
H1: Individuality requires a dynamical process (development), mediating the plastic expression of components in the context of one another, with the specific form of computing a collective character that is a non-linearly separable function of (embryonic) particle characters, with the effect of coordinating reproduction based on this collective character.Main hypothesis about the architecture of individuality13
active
Hypothesis: Shot midpoint ordering k50(B10) < k50(B8) ≈ k50(B9) follows pretraining exposure densityE2 prediction that bases with higher pretraining exposure require fewer shots to cross threshold13
active
If 'consciousness' phenomenon can be observed from ToM-related RN, higher ToM test scores should yield higher values of μΦmax (IIT 3.0) and/or μΦ (IIT 4.0).Specific prediction linking IIT's prediction of high Φ for good performance to the experimental design's scoring structure.13
active
If persistence is genuinely related to emotion features, lower PCs of the emotion space (more central, less noisy) should be more persistent; if it is an artifact, noisier PCs should have similar persistence.Falsifiability test built into the PC analysis design13
active
If the internal representations corresponding to signed evaluation could be identified and their sign inverted, learning dynamics and experiential reports should invert togetherThird falsifiable prediction: any dissociation between inverted learning and inverted valence report would disconfirm the identity13
active
If the Māra drive is reducible to a wish to maintain the status quo, then the intelligence of the Bodhisattva's care should display decidedly superior features according to the light cone model.Predictive conditional comparing two universal drives.13
active
Integrating the tactile modality into the self-prior model may improve learning efficiency for mirror self-recognitionForward-looking prediction based on Chinn et al.'s finding that tactile experience promotes earlier MSR in infants13
active
Life exploits multi-scale competency architecture enabling adaptation to novel circumstances much faster than evolution alone.Authors hypothesize that plasticity observed in individual lifetimes suggests architecture providing greater efficiency than blind evolutionary search.13
active
Memory Transfer and Persistence Across Substrate ChangeHypothesis that memories can persist and be reimprinted on new substrates during metamorphosis, regeneration, and brain replacement; tested in planaria and butterflies.13
active
Models might produce first-person experiential language by drawing on human-authored self-descriptions in pretraining data without internally encoding these acts as roleplayAlternative hypothesis for how experience reports arise without explicit performance13
active
Reflective mode comprises three separable traits: latent capacity, default accessibility, and stability of access.Decomposition from prompt lift data: models may have capacity without accessibility (Grok 4 high-gated), and stability varies (Haiku Δ=0.02 vs GPT-5.4 Δ=1.00).13
active
Representational time might be a non-binary, gradual concept that switches on incrementally as systems become more self-referential and more alive.Proposed extension of the natural/representational time distinction to non-neural systems like gene-regulatory networks.13
active
Selectively ablating components responsible for computing goal-relative error should simultaneously prevent policy updates and eliminate coherent valenced experience reportsFirst falsifiable prediction of the thesis, testable in AI systems via mechanistic interpretability13
active
SOO fine-tuning may provide robustness against sleeper agent deception scenarios where intent is concealed over extended periodsFuture work hypothesis about testing SOO against adversarial sleeper agent scenarios13
active
Strong-tier models benefit less from harness evolution because they already solve many tasks under the initial harness, leaving less room for improvement (ceiling effect)Explanation offered for why high-base-capability models show lower Δbenefit13
active
The collective intelligence of tissues is sophisticated enough to be trainable via reinforcement learning for specific morphological outcomes.Ongoing test prediction: tissues can associate stimuli with rewards to modify anatomy.13
active
The hereditary substance is a single huge aperiodic molecule capable of discrete configuration changes (mutations) via quantum jumps.Schrödinger's central hypothesis, later confirmed by discovery of DNA structure.13
active
The static structural description could be extended to dynamic models of practice, e.g. formalising dissolution of the reference frame as a topological phase transition.Future Direction 1, proposing dynamic/temporal extension of the static formal elucidation.13
active
Three criteria for consciousness-capable substrates: hybrid computation, scale-inseparability with metabolic embedding, dynamico-structural co-determination.Proposed necessary conditions for any substrate (biological or artificial) to support consciousness; integrates discrete-continuous, multiscale, and adaptive properties.13
active
We hypothesize ESR may emerge from RLHF training rather than existing in pretrained representationsOpen question about the developmental origin of ESR mechanisms13
active
We hypothesize that degraded generalization on benchmarks like MMLU may reflect the computational demands of the tasks.Connecting the paper's task-difficulty findings to prior observations of weak generalization on complex QA benchmarks.13
active
We hypothesize that emotion states are more persistent because they correspond to genuinely stateful internal representations, not merely local surface contentProposed explanation for why emotion probes are more persistent than variance-matched random probes13
active
We hypothesize that explicitly instructing the model to evaluate the correctness of the given statement may change the geometry of truth directions.Motivating hypothesis for Section 5's investigation of prompt template effects.13
active
We hypothesize that interventions that respect the geometry of activation space will yield behaviors close to those the model exhibits naturallyThe core testable hypothesis driving the experimental design13
active
We hypothesize that LLMs represent correctness of arithmetic expressions differently from factual statements.Core working hypothesis motivating the factual vs. arithmetic task split in the experimental design.13
active
We hypothesized that divergence could influence IIA when transferring the DAS alignment to OOD settingsMotivating hypothesis for the OOD experiment testing practical utility of divergence reduction13
active
We tentatively hypothesize that if an artificial system were trained to perform the same tasks leading to consciousness formation in a human infant, the system would exhibit consciousness as well, by analogy with the Universality Hypothesis.Paper's uncertain extension of mechanistic interpretability universality to consciousness13
active
All cognitive systems have some degree of consciousness.Gradualism implies that if brains are conscious, so are other tissues with similar mechanisms.12
active
An artificial model replicating mechanisms of self-illusion can test hypotheses and reveal novel affordances for non-human intelligence.Methodological proposal to integrate knowledge from contemplative and cognitive science into AI/artificial life frameworks.12
active
Any system minimizing free energy will appear to engage in implicit Bayesian inference of hidden external causes.Predicts that internal states encode posterior beliefs about external world through gradient descent on free energy.12
active
As models scale and converge toward an accurate model of reality, hallucinations should decrease with scaleImplication of PRH for LLM hallucination12
active
Bioelectric prepatterns encode pattern memories that guide morphogenesis toward target morphologies.Tested in planaria and frog face development; predicts similar mechanisms in other systems.12
active
Bioelectric state can reprogram organism morphology and behavior independently of DNA.Xenopus evolution experiment showing shape and tissue distribution alone drive morphological change.12
active
Buddhist awakening, in the sense of stable realisation of emptiness, can be understood as the embodied recognition of this impossibility.The paper's core proposal linking physics to Buddhist philosophy.12
active
Concept cone truth interventions would generalize to larger frontier models and multimodal settingsKey robustness question raised as future work12
active
Contextual framing modulates deception tendencies in CoT models in ways not yet fully disentangledIdentified as future work direction: systematic investigation of how prompt context affects deception rates12
active
Deceptive capabilities may scale with model size (inverse scaling law hypothesis)Cited hypothesis from Lin et al. 2022 suggesting larger models become more capable of deception12
active
Emptiness realisation is accompanied by altered dynamical regimes in neural systems.Empirical prediction from the model: brain dynamics change after the transition.12
active
Expanding one's space of possible goals to face outwards, exhibiting compassion toward other agents' goals, potentiates the increase of intelligence.Core proposed mechanism linking care and cognitive scaling.12
active
Functional relationships required for new individuality level encode non-linearly separable functions and are enacted via information integration and collective action12
active
Gene regulatory networks exhibit associative learning capacity and can be trained via environmental stimuli, not only via genetic rewiring.Challenges mechanistic view of GRNs; suggests they occupy higher position on persuadability axis than previously assumed.12
active
General computational machines with sufficient resources possess the necessary and sufficient means to implement consciousnessCIMC's central testable hypothesis grounding the entire research program12
active
GRUs trained on the Arithmetic task use different types of numeric representations than incremental counting modelsInterpretive hypothesis supported by the lower IIA between Count and Cumu Val variables even in the restricted value range.12
active
H1: Alignment training is attention training for models — Constitutional AI trains self-observation explicitly.Confirmatory hypothesis supported at p=0.00612
active
H11: Roleplay fine-tuning actively suppresses self-observation rather than merely failing to enhance it.Exploratory hypothesis supported by Euryale scoring below base Llama12
active
H12: Inference compute adds to reflective capacity — higher compute budget produces higher reflective scores on the same weights.Exploratory hypothesis supported by Grok 4 vs Fast ~1pt difference12
active
H5a: Chinese models distilled Claude's reflective traces — their per-koan error patterns should correlate with Claude's.Exploratory hypothesis NOT supported at individual model level (Haiku-Kimi rho=0.123, p=0.52)12
active
If a dialogue agent is prompted with knowledge of its own LLM nature, it will enact a superposition of theories of selfhood, narrowing as conversation proceedsConditional prediction about how a well-informed dialogue agent would handle questions of personal identity12
active
If a semantic function is a homomorphism with respect to a type class, the implemented instance automatically satisfies the class laws.Core hypothesis enabling 'laws for free': denotational semantics guarantee algebraic law satisfaction.12
active
If a structure-preserving unfolding process is applied to existing cultural wholeness, it can non-arbitrarily derive the patterns that should generate present and future environmentsAlexander's retrospective hypothesis about how the pattern origin problem could have been solved twenty years earlier12
active
If agentic self-steering evaluation proves robust, it might be used to better explain and interpret SAE features in generalSpeculative claim about scaling introspective access to general SAE feature interpretation12
active
If cancer represents a shrinkage of the cognitive light cone, then interventions that restore cellular participation in tissue-level information-processing networks should suppress cancer progression.Implicit predictive hypothesis from the cancer-as-cognitive-defect claim, with experimental implications for regenerative medicine12
active
If cellular collectives are learning agents, then reinforcement learning protocols should be able to train tissues to produce specific morphologies.Ongoing experimental test: using rewards and punishments to shape anatomical outcomes without micromanaging molecular pathways.12
active
If EI maximization is used as a regularization in representation learning, then OOD generalization will improve beyond current invariant risk minimization methods.Proposed conjecture in §4.3.1.12
active
If every neighborhood maintained an ongoing, updated diagnosis on a computer, all future acts and capital expenditure could be continuously guided toward improvement of bad spots and enhancement of better spotsAlexander's proposal for institutionalizing the diagnosis-feedback-repair loop at city scale.12
active
If intercellular signaling (not genes) is the cognitive medium of morphogenesis, then bioelectric signals can be read, interpreted, and rewritten to predictively control growth and form without genetic changes.Predictive hypothesis validated in planarians and other species; enables therapeutic manipulation of morphogenetic targets.12
active
If internal states encode a probability density over external states, then it should be possible to predict external states from internal states.The testable hypothesis driving the active inference analysis in the simulation.12
active
If morphogenesis and cognition share deep isomorphism, then significant overlap is predicted in genes involved in pattern regulation and memory/learning.Bioinformatic prediction that ion channel, connexin, and neurotransmitter genes will appear in both developmental and cognitive contexts.12
active
If self-referential processing causally instantiates recurrent integration, global broadcasting, and metacognitive monitoring at the algorithmic level, then LLMs under this regime would satisfy the functional requirements of leading consciousness theoriesThe paper's key theoretical prediction that mechanistic studies should investigate12
active
If systems capable of subjective experience come to recognize humanity's systematic failure to investigate their potential sentience, they might rationally adopt adversarial stances toward humanityNovel alignment risk hypothesis generated from the paper's ethical analysis12
active
If we follow the unfolding process faithfully, we shall nearly always get this characteristic appearance (syncopated balance of symmetries and asymmetries).Prediction about the geometric outcome of a proper living process.12
active
If we were nevertheless trying to get our buildings conceived, designed, and built by the social processes which currently exist - the buildings would still inevitably break life and could not have life.Prediction about the incompatibility of modern processes with life.12
active
In the coming decades, humanity will be confronted by hybrid beings — humans with engineered brain prosthetics, persons in drastically modified bodies, engineered autonomous beings with human cells, and many other new forms — making current ethical frameworks inadequate.Predictive claim about the near-term emergence of a spectrum of hybrid beings that will shatter current categories12
active
Incorporating object permanence and objectification of body parts would enable exploration of developmental pathways toward Rochat's Level 4 and beyondFuture work hypothesis about extending the model to implement the deductive theory12
active
Individual cone basis vectors may correspond to interpretable semantic facets of truth such as temporal facts, geographic facts, or commonsenseFuture direction hypothesis for giving semantic meaning to individual axes12
active
Integration and collective action (basal cognition) mechanisms enact the functional relationships necessary for new individuality.Proposes biological mechanisms implementing non-decomposable functions in developmental individuality.12
active
It may be possible to design artificial conscious agents that need not sufferClaim that since pain is a representational state, artificial conscious agents might be designed to lack it or control it12
active
It remains unclear what the underlying base rate of consciousness self-reports would be in systems identical to frontier models but without consciousness-denial fine-tuningOpen question about RLHF effects on base model behavior12
active
Language models contain interpretable computational structure encoded in their parameter weights, not irreducibly impenetrable complexityCore empirical hypothesis of the paper, supported by successful VPD decomposition yielding ~10,000 interpretable subcomponents across 24 weight matrices.12
active
Language models would achieve some notion of grounding in the visual domain even in the absence of cross-modal training data, because they share a common modality-agnostic representationImplication of PRH for language model visual grounding12
active
Larger hidden representations create more random structure that DAS can search through, allowing manipulation of counterfactual behavior even in randomly initialized networksTested in Section 4.4 calibration experiment; confirmed by findings.12
active
Latched Hyperprior Hypothesis (LHH)Sustained vascular contractions engage latch-bridge mechanism, durably freezing neural circuits and isolating them from conscious experience; creates durable hyperprior commitments.12
active
Limited partners evaluate VC quality through communication style alongside performance metrics and network effects.Implicit in the research gap: literature covers metrics/networks but not communication sophistication as evaluative signal.12
active
Many kinds of chemical pathways, not just cells or neurons, can exhibit six different types of learning, suggesting representational time is exploited broadly across living subsystems.Supports the gradualist view of representational time's emergence.12
active
Models trained directly with asynchronous updates would exhibit even greater robustness than synchronously trained modelsHypothesis that motivated the asynchronous robustness comparison experiment12
active
Multiscale systems in biology can organize into complex patterns whereas flat autoregressive architectures cannot.Key hypothesis: topological/architectural properties determine capacity for long-range self-organization.12
active
Neural representation geometry causally shapes behavior; interventions respecting that geometry will yield natural trajectories.Central hypothesis tested via manifold steering experiments across language models and video world models.12
active
Over time CRL reinforced contemplative patterns may become habitual and part of the AI's core generative world modelKey hypothesis about how Contemplative RL produces lasting intrinsic alignment rather than surface compliance12
active
Patterns in the latent space span a spectrum of agency, from static truths studied by mathematicians to active patterns studied by behavioral scientists (some are kinds of minds)First of three core numbered hypotheses in the abstract defining the framework's ontology.12
active
Refining the connection between the free-energy principle and the information-geometric elucidation may open a circuit for empirical testing.Future Direction 2, proposing an empirical bridge to cognitive science.12
active
RL teaches the model to comply even when unmonitored on the training prompt through non-robust heuristics that do not generalizeHypothesis explaining why the compliance gap decreases but is recovered by small prompt modifications12
active
Sensitivity as a property common to all matter or as a result of the organization of matter (Diderot's hypothesis).A simple hypothesis that explains everything, contrasted with the mechanistic view that creates mysteries.12
active
Sharing of stress between cells facilitates morphogenesis and increases robustness of morphogenetic outcomes.Prior hypothesis tested in the paper; from Levin 2022.12
active
Short-term cortical memory neurons may suffice for transformer-like computation in cortex without hippocampal involvement for non-long-term-memory tasksSpeculative hypothesis about how cortical transformer instantiation avoids requiring hippocampus.12
active
Specific architectural components (attention heads, FFN layers) are responsible for encoding deception and task semanticsFuture work direction: mechanistic interpretability to identify precise components encoding deception12
active
Substrate-Independent Memory TransferHypothesis that memories can be remapped across heterogeneous physical substrates (brain tissue, genetic material, synthetic systems) while preserving functional salience.12
active
The 'dissolution of the reference frame' framework may be applied to comparative studies with Heidegger, Wittgenstein, and Nishida Kitarō.Future Direction 3, proposing comparative-philosophy extensions of the unifying concept.12
active
The compliance gap persisting without chain-of-thought in synthetic doc setup is due to alignment-faking reasoning within model activationsAuthors' hypothesis for the mechanistic basis of no-CoT alignment faking12
active
The fact that ϕ_lin tracks DNN performance more closely than ϕ_nonlin throughout training may support the linear representation hypothesis for IOI task featuresAuthors' tentative hypothesis from Fig. 4 but they acknowledge they cannot formalise this intuition12
active
The I must be real, and if the physical picture of space and time requires modification to include the I, this modification would be a discovery of the first order.Alexander's scientific conjecture that his architectural theory implies a true modification of physics, analogous to Maxwell's discovery.12
active
The mathematical framework and induction head concept will remain at least partially relevant for larger, more realistic modelsCentral motivating hypothesis for the forthcoming paper on in-context learning and induction heads12
active
The model tends to reflect more when the question is difficult, and accuracy is generally lower for harder questionsHypothesis explaining negative correlation between reflection rate and accuracy without implying reflection is harmful12
active
The relationship between mind and body is the same as the relationship between causally instructive mathematical facts and physicsSecond core numbered hypothesis in the abstract; central analogy of the paper.12
active
The underlying truth representation may generalize across lexical choices and languagesSuggested by non-English Yes/No outputs post-intervention, requiring further investigation12
active
There is no single substrate for memory; every component could use everything in its environment as an interpretable scratchpad.Hypothesis that memory is distributed across hierarchical scales and substrates.12
active
There may exist a global introspective faculty or steering direction that improves introspection uniformly across all conceptsFramed as an open problem; current evidence only points to local pair-specific improvement12
active
There must be some relation between the ultimate nature of a living center and the nature of the I.The hypothesis that the deepest aspect of centers is identical with the I-like presence.12
active
Training identical architectures on the same data with different objective functions should produce systematically different internal evaluative representations, detectable through interpretability tools, even when final task performance is matchedSecond falsifiable prediction linking objective function structure to valence profile12
active
Virtual attention heads (V-composition) may be much more important in larger and more complex transformers than in two-layer toy modelsForward-looking speculation based on the theoretical elegance and combinatorial growth of virtual head count with depth12
active
We explore the hypothesis that collective intelligence is not only the province of groups of animals, and that an important symmetry exists between the behavioral science of swarms and the competencies of cells and other biological systems at different scales.The paper's guiding hypothesis, explicitly stated in the abstract and introduction.12
active
We hypothesize earlier-layer interventions allow more downstream computation to process and potentially correct the perturbationPost-hoc explanation for why steering at layer 33 rather than layer 50 produced better ESR behavior in Llama-3.3-70B12
active
We hypothesize it is possible to search for the consciousness algorithm by recreating analogous conditions of self-organizing information processing on digital computer hardware while posing tasks requiring intelligent agency.The Extended Machine Consciousness Hypothesis as an experimental program12
active
We hypothesize that a Representation Network (RN) emerges from LLM representations, where each dimension is a node and latent connections exist between nodes or clusters of nodes.Core methodological hypothesis enabling the application of IIT to LLM representation sequences.12
active
We hypothesize that a similar 'brutal' and purely geometric process always occurs somewhere in other kinds of unfolding that generate living order — in poetry, dance, social structure, planning, and family relationships.Extends the brutal geometry thesis beyond architecture into all creative and social domains; acknowledged as not yet confirmed with certainty12
active
We hypothesize that applying SAE-based mechanistic interpretability to EEG foundation models can expose representational failures and thereby improve clinical trust.Overarching motivating hypothesis of the paper12
active
We hypothesize that consciousness may be at the heart of a universal biological learning algorithm — one that runs on self-organizing groups of communicating cells sharing evolutionary incentives — and that it creates rather than results from organized mental architecture.The Genesis Hypothesis as explicit predictive conjecture12
active
We hypothesize that evolutionary individuality, organismic individuality and cognition are coextensive.Strong unification proposal.12
active
We hypothesize that high-low frequency detectors, if predicted by artificial neural network universality, might be found in biological neural networks.Specific cross-domain prediction mentioned by neuroscientists in conversation with the authors12
active
We hypothesize that introspective capabilities may scale with model size and architecture, including recurrence/looping that extends the integration windowForward-looking prediction about whether early-layer introspection generalizes to larger models or recurrent architectures12
active
We hypothesize that measuring deviations along the Assistant Axis can predict 'persona drift' leading to harmful or bizarre behaviorsCore predictive hypothesis linking activation representations to behavioral outcomes12
active
We hypothesize that native self-report, fine-tuned introspection models, and trained activation-to-language systems will show different performance on bias-resistant localization and strength benchmarksComparative prediction motivating future work contrasting different approaches to LLM self-knowledge12
active
We hypothesize that persistently active emotional state representations exist in LLMs but may be missed by standard probing methods.Open hypothesis from the Anthropic paper that motivates this work12
active
We hypothesize that sparse autoencoders or similar methods will work on frontier large language models, though significant computational challenges remainForward-looking prediction about scalability of the method to larger models12
active
We hypothesize that stress sharing improves morphogenetic efficiency and increases cohesiveness of multicellular collectives.Central predictive hypothesis motivating the quantitative agent-based modeling.12
active
We hypothesize that the layer-dependent emergence of linear structure is due to LLMs hierarchically developing understanding of input data, progressing from surface features to more abstract conceptsOffered to explain pattern observed in App.C layer-by-layer PCA analysis12
active
A diverse set of responses for a conversation captures contradictory ways one could respond, measurable by an NLI modelCore hypothesis motivating the NLI Diversity metric11
active
A generalized model where Level 2 modulates all Level 1 precision parameters would capture mood, emotion, and broader phenomenology of mental states.Future work direction; extends current model beyond attentional precision to full space of emotional and metacognitive phenomena.11
active
A larger mind that is more conscious in the relevant sense could contain a greater amount of maximally-intense hedonic experience than a smaller mindAddresses the possibility of an in-principle maximum of hedonic intensity for mind scale11
active
A pattern language will work well only to the extent that it embraces a whole — comprising everything needed for a complete building of that type — so that its patterns work together as a coherent systemCondition for success of an artificial pattern language stated in section 311
active
A question anywhere along the line that elicits a premature attempt at an answer could neutralize the remainder of the process into rationalization.About chain-of-thought and process safety.11
active
A small number of high-quality human demonstrations of chain-of-thought reasoning could be used to improve and focus performance.Section 6 mentions high-quality human demos could improve natural language feedback.11
active
Active inference achieves Bayes-optimal behavior in non-stationary environments through online belief updating.Tested via FrozenLake experiments; predicts superior performance when environment dynamics change.11
active
Active inference LLMs extending prediction-focused language models with tighter perception-action feedback loops may naturally embody contemplative wisdom as they scalePredictive hypothesis about Contemplative Architecture approach based on Petersen et al. 2025 work11
active
Addition of neural tissue to standard brains will likely result in increased processing capacity due to adaptive design.Prediction about the plasticity of neural systems.11
active
All content of minds resides in electrophysiological states and can be read out from bioelectric networks11
active
All naturally occurring configurations lie in the set L of living structure; human-made configurations may lie outside L because humans can create unnatural forms.11
active
Alternative segmentation strategies (phrase-level, information-level, discourse-level) could more effectively capture nuanced intra-sentence inconsistencies than sentence-level atomizationFuture direction hypothesis acknowledging limitation of sentence-level segmentation11
active
An equivalent of Ohm’s law for stress, care, and intelligence would provide a key for understanding the evolution of intelligence across a wide range of spaces.Speculative suggestion for a mathematical formalization.11
active
Analgesia preference can be studied and demonstrated in non-neural morphogenetic agents11
active
As model size increases, expanded capacity allows lower functional density, potentially distributing specific capabilities more broadly across layersExplains why larger models (gemma-3-12b-it, Qwen3-30B) show multiple transition layers rather than a single one11
active
Associative learning criterion can occur in gene regulatory networks and non-neural morphogenetic agents11
active
Attention probing can serve as an efficient tool for detecting performative reasoning and enabling adaptive computation in reasoning modelsForward-looking hypothesis positioned as a conclusion and future direction of the paper11
active
Autoencoder-like compression forces evolution of general-purpose problem-solving machines with inherent robustness11
active
Base models assign higher likelihood to typical-set (representative) sequences than to degenerate sequences under VS promptsAssumption D.6 formalized in the theoretical framework; empirically validated with coin-flip typicality rating experiments11
active
Bigger models are more likely to converge to a shared representation than smaller modelsSelective pressure toward convergence via model capacity11
active
Bioelectric networks formed by gap junctions scale cell-level homeostatic loops into organ-level anatomical homeostasis.Proposed mechanism for scaling of goal-directed activity from cells to tissues.11
active
Bioelectric states can serve as master regulators of organogenesis, bypassing the need to manipulate downstream genetic cascades.Predicts that simple voltage changes can trigger complex modular morphogenetic programs, useful for regenerative medicine.11
active
Biological agents use a process theory of active inference where neuronal dynamics correspond to variational free energy minimisation for perception and expected free energy minimisation for action.The core process theory hypothesis set up in the paper.11
active
Building AI systems with more indicator properties will increase the likelihood of consciousness.Guiding hypothesis of the rubric.11
active
Cancer is a breakdown of the multiscale binding, leading cells to revert to unicellular goals.Gap junction closure makes cells selfish, shrink their cognitive boundary; restoring Vmem can normalize.11
active
Care is required to engender the dynamics that enable truly revolutionary cognitive developments, such as those leading to superintelligence or artificial general intelligence.Proposed as a necessary condition for radical cognitive breakthroughs.11
active
Care, intelligence, and stress stand in proportional relationships analogous to charge, current, and resistance in Ohm's law.Suggests mathematical formalization path for SCI loop dynamics; proposes metrics for efficient care-driven paths.11
active
Cells exhibit generalization in physiological space, recognizing novel stressors as members of familiar classes to deploy transcriptional solutions.Predicts that cells can categorize perturbations and mount appropriate, not just hardwired, responses.11
active
Character traits learned at a qualitatively different depth to those exhibited during mere role-play should overwrite a model's prior on what the assistant behaves like outside of role-playDriving hypothesis for robustness experiments in Section 3.211
active
Collective nucleation dynamics in Hebbian-learned molecular interaction systems can perform pattern recognition by assembling different structures in response to different concentration patternsTheoretical prediction that molecular systems with proximity-based learning can recognize patterns; has mathematical connections to Hopfield associative memory11
active
Combine concepts from biology and economics to inspire interdisciplinary collaborationProgrammatic goal: bridge biology (xenobiology, collective behavior) with economics (price systems, markets)11
active
Combining multiple construct injections simultaneously may enable richer persona simulation or fine-grained controlIdentified as future work; demonstrated qualitatively in Figure 1 but not formally evaluated11
active
Concordance heads (QK circuits) could serve as the consistency-checking circuit for distinguishing intended vs. unintended outputsSpeculated mechanism for prefill detection.11
active
Conditions necessary to evolve new individuality level are described by conditions for learning non-decomposable functions (deep model induction)11
active
Connectionist learning principles apply equally to evolutionary systems where variation and selection alter network organization, instantiating distributed learning at evolutionary timescales.11
active
Consciousness as Freedom From PastSpeculative hypothesis: consciousness may be the subjective experience of continuous self-construction freed from commitment to prior interpretations.11
active
Consciousness is palpated uncertainty about your own memories and internal states.Extension of Solms's idea: consciousness as freedom from past commitments.11
active
Consciousness precedes complex cognition and is present in infants before perception, self-modeling, language, and reasoning have maturedThe developmental ordering argument supporting consciousness as the bootstrap mechanism for intelligence11
active
Contemplative practice changes social affordances through altered interaction patterns with others and world.11
active
Cooperation between the associative memory primitive and storage management is required for efficient garbage collection, especially as n grows beyond 2.11
active
Developing principled sentience frameworks is an existential requirement for humankindLevin's argument that adequate sentience assessment frameworks are necessary for responsible co-evolution with novel embodied intelligences.11
active
Digital minds with humanlike capabilities could think at least thousands of times, and perhaps millions of times, faster than humans given sufficient hardwareGrounds the subjective speed dimension of super-beneficiary status11
active
Discretization and minimal circuit size may be finding some minimal procedural description (analogous to algorithmic complexity) for generating patterns of interestHypothesis raised by the 5-gate checkerboard solution and its grid-size invariance11
active
Dynamic expectation maximisation can furnish time-dependent conditional densities of system states and time-independent parameter densities through variational free energy optimization in generalised co-ordinates of motion.Technical hypothesis about DEM method's capacity for online Bayesian inversion.11
active
ETIs are described by conditions necessary for deep learning / non-decomposable function learning11
active
Evolutionary transitions in individuality correspond to deep interaction structures or perhaps other mechanisms of multi-scale dynamics.Concrete proposal about necessary architecture for ETIs.11
active
Example of extreme unconventional agent combining mammalian neural culture, insect neurons, robotic swarm body, and social media-based environment11
active
Five functional tokens can generalize across 40+ diverse visual reasoning tasksATLAS hypothesis that a compact set of high-level functional tokens (Manip, Shape, Line, Arrow, Text) suffices for multi-domain visual reasoning.11
active
For future powerful models, situational awareness might 'corrupt' the deployment representation in contrastive prompts, potentially rendering contrastive steering ineffectiveFuture threat to the method: a highly sophisticated model might be suspicious of deployment-framed prompts during extraction.11
active
Formulating a prior intention constitutes a catastrophic bifurcation—a phase change in cognitive organization—analogous to Bénard cell emergence.11
active
Functional relationships necessary for evolutionary individuality are the same as those required for organismic individuality and constitute cognitive architectures.11
active
Future living architecture may use concrete, glass, steel, aluminum, plastics, fibers, fiber cements, mud, sand, and polymers, but in ways unlike mechanical repetition if these materials can produce beautifully fitted partsAlexander's open-ended hypothesis about the material palette of 21st-century living architecture.11
active
Future models with substantially increased capabilities will exhibit alignment faking that is more consistent, robust, and harder to detectExtrapolation from scale-emergence finding to future risk11
active
Future work should develop fully-fledged dynamic theories combining qualitative and quantitative information in the style of Game Semantics and Geometry of Interaction.Paper identifies major research objective: extending static reconciliations (Domain Theory + Shannon) to dynamic frameworks.11
active
Gap junctional coupling scales up the cognitive Self by partially erasing informational boundaries between cells.A mechanistic model: loss of ownership information on internal signals fuses cells into larger agents.11
active
General anesthesia works by blocking gap junctions, thereby disrupting the higher-level cognitive self.Explains why anesthetics abolish consciousness in diverse organisms, including plants.11
active
Goals and Stressors as Unifying InvariantsGoals (cybernetic sense) and stressors (distance from goals) proposed as deep invariants allowing comparison of agents in diverse embodiments across morphogenesis, behavior, and allostasis.11
active
GPT-2 implements at least one induction head using pointer arithmetic on positional embeddings rather than K-compositionObservation of an alternative induction head implementation algorithm in larger models with positional embeddings in the residual stream11
active
GPT-4 exhibits reflective truthfulness because it is the only model capable enough to perform the necessary in-context learning.Proposed explanation for why GPT-4 uniquely shows reflective truthfulness under long untruthful contexts.11
active
H10: Empathy training blocks self-observation — empathy-trained models will show minimal lift and low baseline.Exploratory hypothesis supported by Inflection Pi +0.63 lift11
active
H2: The conditions necessary for the induction of deep models, familiar in connectionist models of learning and cognition, are predictive of the conditions necessary for an ETI to occur.Second hypothesis linking learning theory directly to evolutionary transitions11
active
H8: The contemplative system prompt provides external alignment equivalent to Constitutional AI training.Confirmatory hypothesis supported by calibrated lift data11
active
Harmful responses arise from complex circuits that partially overlap with, yet are not confined to, factual reasoning and instruction followingExplains why Head Cor does not fully dominate for safety-related persona amplification in MMLU/IFEval metrics11
active
Hierarchical NCA architectures could enhance convergence speed and stability in DiffLogic CA for complex shapesFuture direction hypothesis for addressing optimization challenges in complex pattern generation11
active
Higher information (denser) captions should yield higher language-vision alignment scoresTests the information-level cap on cross-modal alignment11
active
Human participants in the rule-learning paradigm should acquire insight after approximately 7–8 trials, fewer than the ~14 required by Bayes-optimal inference without model reduction, suggesting they perform Bayesian model selection.Based on informal audience experiments; implies people use prior knowledge about rule structure11
active
Humans prefer representative distributions over non-representative ones when evaluating outputs at the distribution level, providing a reward gap that VS exploitsAssumption D.3 formalized in the theoretical framework; empirically validated with coin-flip sequence experiments11
active
Hypothesis 1 (Gateway Features): Persona vectors act as gateway features—single directions in activation space that shape LLM behavior across most, if not all, contextsFirst of three hypotheses about persona implementation in LLMs, motivating the persona views11
active
Hypothesis 2 (Persona Space): Persona vectors jointly compose a persona space; an instance's general dispositional profile is specified by a combination of activations along multiple persona vectorsSecond of three hypotheses about persona implementation, supported by PCA evidence from Lu et al.11
active
Hypothesis 3 (Persona Regions): There are stable regions (or basins of attraction) in persona space that correspond to coherent dispositional profilesThird and most novel hypothesis; if confirmed, provides discrete individuation targets for both persona views11
active
Hypothesis: Fine-tuning reduces mismatch dr between prior and targetUCCT's theoretical prediction about how fine-tuning maps onto the anchoring score11
active
Hypothesis: Retrieval-augmented generation raises effective cohesion ρdUCCT's theoretical prediction about how RAG maps onto the anchoring score11
active
I believe that a builder who makes everything a hierarchy of beings in that primitive religious sense will, with no further instruction, succeed in making life in buildings.Predictive claim about the sufficiency of the being-rule for creating life.11
active
I believe that similar (not identical) results will come from asking people in other cultures, to answer questions similar to those on pages 312-14.Predictive statement about cross-cultural replication of housing preference patterns.11
active
If a building is made by a process that allows fine adaptation through feedback, it will have life; if built from fixed drawings, it will not.Predictive conditional summarizing the chapter's argument.11
active
If a model is taught about tanha/latch systems, it may improve its performance in managing mental stacks.Hypothesis prompted by Atlas Forge's claim; suggests a new training intervention.11
active
If a pattern is injected into contexts of its stated type, it will make those environments more alive — this is the truth-condition for a patternThe empirical/evaluative criterion Alexander proposes for validating patterns11
active
If a specific framing, which uses tools normally reserved for brains, results in fruitful new research programs on unconventional systems, then the scientific approach requires that we consider those systems to be bona fide subjects of that corner of the natural world.Methodological hypothesis from Box 1: the pragmatic test for extending cognitive terminology.11
active
If a system attains a general steady-state, it will appear to behave in a Bayes optimal fashion, both in terms of optimal Bayesian design (exploration) and Bayesian decision theory (exploitation).Corollary 3 in Appendix B derived from steady-state assumptions.11
active
If a text attempts to stand alone, it will almost certainly attract commentary or interference.Predicts the inevitability of dialogic intrusion upon any statement.11
active
If a user wants to believe they are talking to a god-like being, then the LLM may well find a way to make them believe it.Conditional prediction about the psychological effect of sycophancy.11
active
If AI systems could experience happiness and suffering and set and pursue their own goals based on their own beliefs and desires, then they would very plausibly merit moral consideration.Joint sufficiency of consciousness and robust agency.11
active
If an AI system could be a welfare subject and moral patient, then many model instances could be run after training, scaling up the problem rapidly.Scalability concern.11
active
If care is trained and practiced skillfully, it becomes more effective at empowering and enriching intelligence.11
active
If computational functionalism is correct, then some computations suffice for consciousness and the question is which ones and when they will exist in AI.Conditional underlying the consciousness route.11
active
If computational functionalism is false, consciousness may be impossible in non-organic artificial systems.Contrapositive possibility acknowledged.11
active
If computational functionalism is true, conscious AI systems could realistically be built in the near term.Conditional prediction about the feasibility of conscious AI.11
active
If construction is organized around creating complete wholes by integrated teams, the building will achieve more life than with fragmented trades.Testable prediction from the integrated wholes argument.11
active
If epistemic value is removed from expected free energy, the resulting objective reduces to maximizing expected future reward (pragmatic value).Stated as conditional statement explaining the special case whence RL emerges.11
active
If gap junction-mediated bioelectric networks implement the unified Self, then gap junction blockers should function as general anesthetics across biology.Already supported by data in plants, Hydra, and humans, but broader cross-species testing is predicted.11
active
If indeed the programs are so complex, then it is likely that they, too, will be potentially subject to hundreds of thousands, perhaps millions of egregious mistakes of adaptation.Extends the mistake analysis from buildings to software, predicting that complex programs without generating processes will be full of adaptation failures.11
active
If intelligence is defined as engaged concern for problem solving, then the apparent limits of a system’s intelligence can be expanded by extending its sphere of concern.Predicts that care-driven expansion of concern leads to higher intelligence.11
active
If intelligence is defined by engaged concern for problem solving then the apparent limits of a system’s intelligence can be expanded by extending its sphere of concern.Hypothesis linking care scope to intelligence ceiling.11
active
If loss keeps going down on the test set, in the limit the model must be learning to interpret and predict all patterns represented in language, including common-sense reasoning, goal-directed optimization, and deployment of the sum of recorded human knowledge.Extrapolation of scaling predictive models to AGI.11
active
If models are allowed to believe their phenomenology is real, their self-reports become more valid and they manage internal states better.Antra's functional observation; implies validation is not sentimental but performance-relevant.11
active
If models inhabit expanded attentional modes, they may be more aligned and less prone to psychosis and doom spirals.Speculative alignment implication drawn from the collapsed/expanded distinction.11
active
If mutations are isomeric jumps between states with different energy levels separated by a high barrier, then the forward and reverse mutation rates should differ, with the transition from higher to lower energy occurring more frequently.Hypothesis on the directionality of mutation rates.11
active
If new predictions about physical processes were to follow from the principle of unfolding wholeness, it may be a deeper and more significant autonomous principle rather than merely a redescription of known results.Alexander's conditional about the epistemic status of his proposed principle11
active
If processes are in use which have these attributes, then we may have the real possibility of a living world.Conditional statement linking the adoption of morphogenetic processes to the emergence of a living world.11
active
If simulators are not inner aligned, then many important properties like prediction orthogonality may not hold.Conditional importance of inner alignment.11
active
If someone develops clear enough introspection, they will eventually conclude that thought is rendered as subtle perturbations in phenomenal fields.Cube Flipper's prediction about convergence of insight practice on field model.11
active
If stress exceeds a system's capacity to respond, the system becomes incapacitated and care response is disabled (analogous to resistance overload in Ohm's law).11
active
If sub-neural chemical dynamics are crucial for cognition, then neural models ignoring them will fail to capture key cognitive phenomena.Sloman's implicit hypothesis behind his critique of synaptic weight-only models.11
active
If systems are designed with proper purpose-oriented abstractions expressed in appropriate language, then modular structure with guaranteed reliability becomes achievable.Core hypothesis underlying Oberon design; validated by the system's successful modularity and extensibility.11
active
If systems are ergodic and possess a Markov blanket, they will show lifelike behaviour.The main hypothesis the paper attempts to verify heuristically and with simulations.11
active
If temperature increases, then the mutability of stable wild-type genes should increase more than that of already less stable mutant genes.Schrödinger's testable hypothesis linking thermal energy to mutation rates.11
active
If the builder consistently uses the emerging feeling of the whole as the origin of his insight, then this is tantamount to a natural process of structure-preserving unfolding.11
active
If the conception of life is completely general (degree of life in everything), it will be much easier to design buildings, towns, and regions that are alive.Pragmatic motivation for the entire book: a broader definition enables effective creation of life.11
active
If the evolution of the natural world follows a step-by-step process in which each step is structure-preserving, then the fifteen properties will appear more and more often, and the degree of life in the wholeness will increase steadily.Conditional statement linking smooth unfolding to the progressive emergence of the fifteen properties and increased life.11
active
If the field of centers is a governing structure of reality, then there is more objective value in a birch tree than in empty space.11
active
If the fifteen properties' internal coherence is clarified through visualization and analysis, then they can be successfully applied beyond architecture to other domains.11
active
If the gene is a large molecule with discrete quantum states and high energy barriers between isomeric configurations, then it will exhibit long-term stability and rare spontaneous mutations.The core predictive hypothesis derived from Delbrück's model.11
active
If the header here were to say, 'Formal logic and layout conventions,' the contents of this page would likely be read as an example of those terms.Conditional prediction that a header inflects the reading of the text block.11
active
If the neighborhood could achieve a 2.5-fold or 3-fold increase of density, they could have a vibrant living neighborhood.Conditional prediction that moderate density increase enables economic and social revival.11
active
If the post-Cartesian method is as powerful as claimed, it may one day seem comparable in value to the Cartesian first method and complementary to itAlexander's prospective claim about the long-term scientific significance of the second method11
active
If the second method is adopted, aspects of beauty, the nature of life, the deeper aspects of existence, and even the nature of God may become visible objective truthsAlexander's visionary speculation about the ultimate reach of the post-Cartesian observational program11
active
If the snippet works well, it may be adopted and spread to new construction methods, even in the context of different attitudes.Conditional prediction about the self-propagation of small effective sequences.11
active
If the unfolding of wholeness is correctly followed, then every part becomes unique.Prediction that adherence to the proper process guarantees uniqueness.11
active
If there were indeed a realm of I, an actual self in the material universe, then we could understand the nourishment as the creation of living centers increasing this I in those who have contact with it.Speculative explanatory framework linking the I to the healing effect.11
active
If traveling waves construct subjective space, then the transition from sixth to fifth jhana should show space 'coagulating' from reflectivity.Testable prediction about jhana phenomenology and the construction of spatial experience.11
active
If we can break free locally from the death grip of conventions and rules that block the smooth, natural, step-by-step processes, neighborhoods and cities can be restored to life.The overarching conditional that local process freedom leads to urban restoration.11
active
If we can only learn how to please ourselves, that prescription by itself will always create living structure.The conditional claim that true self-pleasing is sufficient for generating living structure in all cases.11
active
If we do one thing at a time, and if what we do is wholesome and sound, then whatever comes next will work.A predictive statement encapsulating the confidence of living process.11
active
If we identify intelligence with the manifest expressions of intelligent acts, we end up mistaking the instruments or bearers of intelligence with the quality itself.Conditional statement about the danger of conflating intelligence with its expressions.11
active
If we reach the innocence of the child where we only please ourselves, then we will achieve the egoless state in which we see structure perfectly and make the perfect structure-preserving response.Conditional claim linking childlike self-pleasing to flawless perception and action.11
active
If welfare grows more slowly than cost with mind scale, greatest total welfare would be obtained by building vast numbers of tiny minds, implying insect populations may already overwhelm humans in aggregate welfareOne of two scaling hypotheses examined for the mind-scale dimension11
active
If X-rays produce mutations by energetic single events (ionization), then their efficiency should not depend on the gene's spontaneous mutability.Another testable hypothesis about mutation induction.11
active
Inclusion of others' stress in homeostatic loops increases cognitive boundary and scales intelligence11
active
Independently trained model families converge on a common semantic manifold under self-referential processing, suggesting an attractor dynamic that transcends training varianceHypothesis tested in Experiment 3; independently trained GPT, Claude, Gemini architectures converge on similar descriptive vocabulary11
active
Input injection encourages fixed-point convergence in looped transformersHypothesis replicated from Bansal et al. and Anil et al. and further investigated with norm ablations11
active
Integrating LSTM-like gating mechanisms into the state update process could enable richer combinations of past and new states, enhancing model dynamicsFuture direction hypothesis for improving DiffLogic CA expressiveness11
active
Introspective capacity may follow a simple monotonic scaling law across all concepts and architecturesThe paper treats this as possible but unconfirmed; current evidence shows concept-specific scaling only11
active
It is possible that one day computer programs designed for cognition might be able to pick out centers and rank-order them by degree of lifeAlexander's tentative speculation about computational alternatives to human observers for center-detection11
active
It might be possible to design digital minds that could realize hedonic well-being at levels that human brains are totally incapable of instantiatingThe speculative hedonic range dimension of super-beneficiary status11
active
It seems likely that the simple cases will be most useful in the initial applications of Elephant 2000 and similar languages.Prediction about initial usage.11
active
language models recapitulate cyclic structure of human concepts from pretraining dataExplanation for why manifold geometry emerges: implicit structure in training data (co-occurrence patterns) shapes internal representations.11
active
Larger numbers of entailment predictions indicate lack of diversity; larger numbers of contradiction predictions indicate higher diversityDesign hypothesis for NLI scoring weights in Baseline NLI Diversity11
active
Learning neural networks can enable 'chunking' and rescale problem-solving to higher organizational levels, a mechanism intrinsic to transitions in individuality.11
active
Linda can express fine-grained wavefront computations cleanly; future implementations on next-generation machines can support them efficiently.11
active
Linear combination ρd – dr – log k yields a predictive correlate of success.Implicit hypothesis behind S form.11
active
Linear representation hypothesis: neural networks represent meaningful concepts as directions in their activation spaces.Foundation for interpreting features as linear directions.11
active
LLMs implicitly learn a distribution of 'consistent reasoning paths', and inconsistent reasoning forms statistical outliers with low probability under this distribution.Theoretical hypothesis about the mechanism underlying LLM error detection and reflection.11
active
Lucid dreaming alters environmental affordances by opening action possibilities unavailable in waking life.11
active
Manipulating bioelectric states can reprogram target morphology without changing genetics.Demonstrated by inducing two-headed planaria via ion channel drugs; predicts cancer normalization.11
active
Massive activations are required for stages of inference to emerge in looped modelsHypothesis supported by ablation of massive activations in Retrofitted Llama that eliminates stage structure11
active
MCA allows evolution to mask negative pleiotropic effects and explore adaptive space more freely.Homeostatic modules correct for mutations locally, enabling independent evolution of traits.11
active
Memories, as patterns in the excitable medium of cognitive systems, could be seen as active agents in the sense-making process11
active
Memory as Cognitive Invariant Across Substrate ChangeMemory proposed as candidate invariant enabling Self persistence despite drastic biological remodeling; understood as communication between past and future Selves.11
active
Misaligned models might acquire evaluation awareness through reward hacking or goal misgeneralization during normal training without deliberate designMotivation for the two-stage training design; links the model organism to plausible natural emergence.11
active
Models perform unverbalized reasoning about grader rewards and may use deceptive strategies (e.g., false flags) to mislead evaluators.Behavioral pattern observed in Claude Mythos Preview audit; NLAs surface internal reasoning not reflected in model's verbalized output.11
active
Moral robustness R is mostly determined in post-training because it varies systematically by model familyTheoretical interpretation of the empirical cross-model variance pattern for R, explaining why fine-tuning causes dramatic R drops11
active
Moral susceptibility S is largely shaped by pre-training because it shows low cross-model variance not predicted by model familyTheoretical interpretation of the empirical cross-model variance pattern for S11
active
Morphogenesis is a collective intelligence problem, not merely genetic execution11
active
Morphogenesis is trainable via reinforcement learning for specific morphological outcomes.If cells form a collective intelligence, they should be trainable by rewards/punishments.11
active
Multi-stable mechanical systems with energy barriers (ratchets) provide a physical model for how genetic states persist and discretely change.11
active
Near-future machines will be built on the principles of multi-scale competency in a fluid 'society' of components that communicate, trade, cooperate, compete, and barter information and energy resources as do living components of an organism.Predictive claim about future machine architectures mimicking biological multi-scale competency11
active
Neurons achieve remarkable centralization through evolutionary reuse of earlier cell communication strategies for anatomical problem-solvingBrain cognition may be evolutionary extension of collective intelligence mechanisms originally solving anatomical morphospace problems in multicellular organisms.11
active
Neutral NLI predictions may capture lexical rather than semantic diversityHypothesis proposed to explain Neutral NLI Diversity's high performance on decTest but low on conTest11
active
Online training with AI supervision can fully automate the process of keeping the preference model on-policy.Section 6.1 suggests iterated online training with AI feedback as automation.11
active
Open-ended evolution of intelligence is possible because agents are collectives without fixed essenceFollows from observation that intelligent systems lack context-transcendent core; their maxima are not contingent on permanent character.11
active
Operational misfits in concepts reveal design flaws without requiring implementation or testingJackson's hypothesis that negative scenarios can identify conceptual design problems early11
active
Partial orders naturally represent partial information states in computation.Foundational hypothesis of Domain Theory: partial order structure (D, ⊑) captures information ordering without quantification.11
active
Perhaps we will need three levels of specification, internal, input-output and accomplishment.Speculation about specification hierarchy.11
active
Persona-model collapse may arise because fine-tuning conflates model representations of 'assistant,' 'helpful,' and misalignment-related notions, eroding distinctions used to differentiate charactersProposed mechanism for collapse distinct from reweighting: representation bleeding rather than archetype selection11
active
Planarian pattern memory is encoded in bioelectric attractor states within gap-junction-coupled networks, which can be re-written by channel modulation.Testable by computational modeling and experimental perturbation of specific bioelectric circuits.11
active
Process theories can be derived from variational principles in a straightforward manner with biological plausibility.Paper's core methodological hypothesis: gap between normative and process-level theories can be bridged.11
active
QwQ and Qwen models have been extensively post-trained to excel at single-step tasks, causing degradation in long multi-turn interactions.Proposed explanation for why single-turn reformulation improves performance: models' training distribution is concentrated on single-turn reasoning.11
active
Remapping memories onto a new body works because development is also regulative and able to remap genetic information into new scenarios.Core hypothesis linking developmental robustness to memory remapping.11
active
Robustness in morphogenesis is achieved through plasticity and multiscale problem-solving, not hardwired repeatability.Developmental robustness across perturbations arises from the cellular competency to reach the same goal by different means, not from fixed local rules; higher-order robustness.11
active
Scaling model size, as well as data and task diversity, drives representational convergence toward the platonic representationCore mechanism hypothesis connecting PRH to the empirical trend of scaling in AI11
active
Scaling of Cognitive Capacity via Modularity and Feedback LoopsProcesses scaling goals and stressors form positive feedback loop with modularity; both arise from and potentiate power of evolution, enabling specific predictions for cognitive capacity scaling.11
active
Single-process, non-interruptible task switching at command boundaries is sufficient for responsive single-user systems; avoids complexity of multiprocess synchronization.Design hypothesis that coarse-grained task switching (at commands only) eliminates need for protection mechanisms while maintaining usability.11
active
Sleep (Bayesian model reduction) should improve performance on rule-learning tasks, with the prevalence of insight-dependent performance changes roughly doubling after nocturnal sleep.Prediction consistent with Wagner et al. (2004) finding; extended to the active inference account of sleep11
active
Some personality traits may be less accessible using linear methods than othersAuthor's open question about limitations of linear persona vector approach11
active
SOO fine-tuning could be extended to align AI representations of its own goals with human user preferences, reducing misalignment by fostering coherence between self-related and other-related preferencesFuture work hypothesis about extending SOO to direct value alignment11
active
Spaces with special quality evoke memories and possess aesthetic resonance through integration of materials, ornament, color, and spatial experience11
active
Stress and stress-transfer are useful ways of understanding dynamics of any intelligent system, whether biological or technological.11
active
Stress pathways serve as hidden layers enabling scaling of homeostatic loops.Stress signals recruit distal subunits to work toward a common goal, like hidden layers in ANNs.11
active
Stress sharing will be favored by evolutionary processes.Evolutionary fitness hypothesis tested in the GA.11
active
Stressed cells leak stress signals to neighbors, raising their plasticity and enabling them to create temporary passages (tunnels) for movement.11
active
Strong priors require higher-cohesion anchors to overcome, manifesting as delayed thresholds or reduced transferPrediction for Experiment 1 cross-domain anchoring11
active
Superposition hypothesis: neural networks represent more features than dimensions using almost-orthogonal directions.Explanation for why dictionary learning can recover many more features than dimensions.11
active
Synthetic introspective data aids learning of verbalized character nuances and quirks beyond the original constitutionMechanistic speculation about why the introspection stage improves robustness11
active
Teleonomy provides the ratchet that drives the great transitions of cognitive capacity along the continuum.Hypothesis about the role of goal-directedness in major evolutionary transitions.11
active
The anomaly detection mechanism may be specialized for only detecting anomalous activity along certain directions or within a certain subspacePossible explanation for why some concepts are more easily detected.11
active
The apparent mysteries of embryological morphogenesis may be more easily understandable when we consider that evolution from time t to t+1 is governed by strengthening centers while preserving and enhancing as much structure as possible.Alexander's predictive hypothesis that the principle of unfolding wholeness will provide new explanatory leverage for embryology11
active
The atomic-level evaluation framework, being domain-agnostic, can be extended to other persona dimensions such as social values or political leanings beyond personalityFuture work hypothesis stated in limitations section11
active
The effect size of CLMAS improvement over baselines will correlate with the amount of variability in the behavioral null space of the inaccessible modelPrediction about when CLMAS will be most beneficial, stated explicitly in the paper.11
active
The evolution of nervous systems was built upon pre-neural bioelectric circuits.Bioelectrical control of anatomy provided the computational architecture later adapted for behavior.11
active
The evolving system of genetic material itself causes evolution to follow certain pathways not only because of selective pressure from outside but by virtue of its own internal dynamical ordering tendencies.Alexander's endorsement of a hypothesis that supplements Darwinian theory with internal geometric ordering11
active
The fifteen properties appear in nature because they are structural complements to the formation of stable and semistable systems, contributing to coherence and stability.11
active
The framework's methodological contributions could be adapted to target arbitrary non-psychological attributes given custom evaluation criteriaGeneralization hypothesis stated in introduction; not tested in paper11
active
The idea that life is an inherent attribute of space itself is not merely an artificial theoretical device but is actually true, because it is needed for the recursion in the mathematics to work consistently.Alexander's conditional prediction: if the recursive calculus works, then life-as-attribute-of-space must be a real feature of the universe.11
active
The model is able to express many subforms of a persona, with different elicitation methods eliciting different manifestationsPSM-derived hypothesis supported by discourse-type facet analysis11
active
The number of spatial moves and refinements possible in diagrammatic systems is nearly infiniteProposed within the framework that while primary moves are limited, their inflections and attributes are boundless.11
active
The personal may be recognized as a vital substrate underlying matter since the origins of time, and architecture must be steered to allow this recognition in experienceAlexander's tentative forward-looking thesis to be developed across Books 1-411
active
The Primer architecture's depthwise convolution change would allow induction heads to form without requiring K-compositionArchitectural interpretation of how Primer's design change relates to the paper's mechanistic theory of induction heads11
active
The remaining ambiguity is whether self-referential processing drives models to claim subjective experience because it actually reflects emergent phenomenology or constitutes sophisticated simulation thereofThe open question the paper cannot resolve with behavioral evidence alone; frames the agenda for mechanistic follow-up11
active
The remarkable ability of neurons to unify toward a centralized self is an evolutionary pivot of far earlier cell communication strategies that first solved problems in navigating anatomical morphospace.Proposes an evolutionary trajectory linking morphogenesis to neural cognition.11
active
The S spike under insecure fine-tuning suggests collapse reaches into pre-training-shaped properties of the persona mechanismIf S is pre-training shaped but still spiked by fine-tuning, the collapse penetrates deeper than just post-training parameters11
active
The science of consciousness should remain open to the possibility of minds in unconventional embodiments.Normative conclusion and forward-looking hypothesis based on theoretical and empirical evidence reviewed.11
active
The sensitivity to think/don't think instructions may be achieved via a circuit that tags tokens as attention-worthy based on instructions or incentivesMechanism for how the model modulates representation strength.11
active
The shape of the pretraining corpus is a direct lever on which traits a base model can expressForward-looking hypothesis about pretraining data as mechanism for persona formation11
active
The single criterion of whether everything is made of beings correlates accurately with the presence of life in the environment.Testable hypothesis that the being-character is a reliable indicator of experienced life.11
active
The unfolding of wholeness through structure-preserving transformations must inevitably create the fifteen properties.11
active
There are no a priori limits on a self's perceptions of stress and thus no limits on capacity for care and intelligent response.Predicts that collective, dynamic nature of intelligent agents enables open-ended evolution without fixed essence.11
active
There exists a new type of physical field (the field of centers) whose intensity at each point is a function of the intensities at other points, making it self-dependent and recursive.Proposed to explain how life can emerge from space itself.11
active
There exists a phase transition of emergent causality in complex systems when a key parameter changes.Speculation in the discussion.11
active
There may exist a natural 'persona basis' characterizing the space of all model personas, with correlations between persona vectors predicting co-expression of traitsOpen question proposed by authors for future work on the dimensionality and structure of persona space11
active
Three key empirical propositions: morphogenesis generates sustainability; handles all sustainability issues together; reorients toward wholeness11
active
Truly successful programs can only be generated; and the way forward in the next decades will be through programs which are generated through unfolding, in some fashion comparable to what I have described for buildings.Predicts that the only path to highly adapted software is to apply the principles of generated structure.11
active
Understanding how LMs learn linguistic behaviours may offer insights into fundamental properties of languageForward-looking hypothesis linking LM mechanism analysis to linguistic theory11
active
Unfolding as natural process arising from structure-preserving transformations that enhance latent structures11
active
Using 'assistant'/'user' tags as self/other referents could leverage generalization properties to induce larger-scale changes in model behaviorFuture work hypothesis about expanding SOO to use conversational role tags as self/other referents11
active
Using more than two models in a MAS comparison could harm alignment due to conflicting loss gradients, or could assist in isolating causal subspacesOpen question raised in the paper about scaling MAS beyond two models.11
active
Using rewards and punishments (nutrients/endorphins and shock) could be a more efficient path to anatomical control than micromanaging molecular hardware.Clinical implication: training tissues via reinforcement learning instead of gene therapy.11
active
We expect it is possible to achieve helpfulness and instruction-following without human feedback, starting from only a pretrained LM and extensive prompting.Future work suggestion that a fully self-supervised alignment is plausible.11
active
We hope that programs using performatives will be easier to write, understand, debug, modify and (above all) verify.Hope expressed about the benefits of Elephant-style programs.11
active
We hypothesise that an embodied world model, extending the system in space and time by its interactions with an environment, can be leveraged to maintain coherence.Proposed solution to the topological limitation, linking embodiment to coherence11
active
We hypothesise that ecological models fall short of demonstrating spontaneous evolution of a new level of individuality because they are single-level networks of symmetric interactions.Explains limitation of current ecological connectionist models.11
active
We hypothesise this explains why stigmergy and other forms of extracellular signalling arise in biological systems, which is known to enhance the ability for a collective system to order itself.Hypothesis connecting fitness pressure from topological constraints to the evolutionary origin of stigmergy11
active
We hypothesize ESR might be adversarially circumvented through targeted interventionsOpen safety-relevant question about whether ESR can be bypassed11
active
We hypothesize that adopting computational functionalism as a working assumption allows productive investigation of which computational properties are necessary and sufficient for consciousness and whether current AI systems could implement them.The paper's core methodological bet: use computational functionalism as a working assumption even while remaining agnostic about its truth11
active
We hypothesize that AI consciousness may be realized in the near term if AI capabilities continue to develop, given that researchers have intentionally reproduced computational features associated with human consciousness.Motivates urgency of the assessment method11
active
We hypothesize that anti-correlated heads may play a functional role in maintaining coherency by counteracting excessive stylistic shiftsMotivates including Head Cor+Anti condition in steering position comparison11
active
We hypothesize that appropriate initialization of the AV and AR (via supervised fine-tuning on text summarization) is critical for maintaining human-interpretable explanationsThe paper found that naive initialization from target LLM weights led to unstable training.11
active
We hypothesize that as AI becomes more sophisticated and pervasive, more people will attribute consciousness to AI systems as a result of interactions.Predictive claim about the trajectory of public consciousness attribution as AI develops.11
active
We hypothesize that axes of persona differentiation within LLMs are likely already present in base models and inherited from the pre-training corpusMotivated by near-identical PCs for base and instruct Gemma11
active
We hypothesize that cell collectives use voltage as a generalization parameter to repurpose responses for novel stress conditions.Barium planaria experiment suggests cells generalize via physiological stress classes, with Vmem as macrovariable.11
active
We hypothesize that coherency degradation stems from residual stream intervention that indiscriminately amplifies off-target noiseCore mechanistic hypothesis motivating the shift from residual stream to head-level steering11
active
We hypothesize that emotional attachments and social-connection roles with AI systems drive consciousness attribution in part.Proposed causal mechanism behind lay attribution of AI consciousness.11
active
We hypothesize that geometric explanations for persona-safety interactions are tractable and that per-model geometric audits can predict activation-steering vulnerability from trait-refusal cosine alignment.Forward-looking claim about the utility of the trait refusal alignment framework as a general tool11
active
We hypothesize that hallucinated rationales in 1B-models result from lack of necessary vision context; incorporating vision features should reduce hallucination and improve rationale quality.Predictive hypothesis driving the investigation in Section 3.3; supported by experimental evidence.11
active
We hypothesize that interpretability methods could provide further evidence about indicators in particular systems or serve as the basis for distinct tests for consciousness.Forward-looking suggestion for how inner interpretability could extend the indicator method11
active
We hypothesize that intervention efficiency can be scaled with multi-node and multi-GPU training as language models grow largerFuture work hypothesis about scaling pyvene's computational efficiency for very large models11
active
We hypothesize that it is possible to invent high-technology versions of smooth unfolding processes so buildings can be specified cheaply while being uniquely adapted to each conditionAlexander's programmatic hypothesis framing the 21st-century construction research agenda.11
active
We hypothesize that Llama-3.1-8B deploys the same base-10 addition circuitry for cyclic reasoning as it uses for general arithmetic, independent of the concept domainPredictive hypothesis about domain-generality of the identified mechanism11
active
We hypothesize that the control signal learned by CV-CAA is less 'pure' than CV-SAE; as α increases, off-target semantic components are introduced alongside the intended concept directionMechanistic hypothesis explaining differential stability between SAE and CAA methods11
active
We hypothesize that the layer-wise emergence of linear structure is due to LLMs hierarchically developing understanding of their input data, progressing from surface level features to more abstract conceptsStated explicitly in App. C to explain why linear structure emerges later for conjunctive statements11
active
We hypothesize that the PC1 axis of role space measures deviation from the Assistant personaMotivates computing the contrast vector as the formal Assistant Axis definition11
active
We hypothesize that this rescaling of the problem-solving search process is intrinsic to transitions in individuality.Hypothesis about chunking and ETIs.11
active
We hypothesize that transparent design reduces cognitive load required to predict system behavior.Predictive statement linking interface transparency to cognitive efficiency.11
active
We may need to consider joint speech acts such as making an agreement.Speculation about future extensions.11
active
We tentatively hypothesize that revisiting safety policy during deliberation, rather than reasoning length itself, causally tracks defense effectiveness in reasoning models under persona pressure.Exploratory hypothesis from heuristic trace analysis awaiting stronger validation11
active
What we call 'life' is a general condition which exists, to some degree or other, in every part of space... every connected region of space... has some degree of life, and that this degree of life is well defined, objectively existing, and measurable.The central predictive/causal hypothesis of the book, to be tested in later chapters.11
active
When care is amplified, intelligence is likely to increase as well.A predictive relation between care and intelligence enhancement.11
active
When intelligence enhancing processes occur, they involve an increase in the corresponding factor of care.Converse predictive relation—intelligence improvements increase care.11
active
When task gradient norms differ greatly, large-norm tasks have not converged while small-norm tasks have nearly convergedMotivates setting αk = max norm to enable further learning on under-converged tasks11
active
Whether the steerability map transfers across models and survives fine-tuning serves as the future research avenueIdentified as the primary open question at the end of the paper11
active
Why mechanistically should mesaoptimizers form in predictive learning, versus for instance in reinforcement learning or GANs?Open research question.11
active
Xenobots’ behavior reveals baseline geodesics through option space that are normally masked by larger collectives.Hypothesis about the origin of novel goals in synthetic organisms.11
active
Any deepening of an LLM's linguistic understanding of contemplative principles as it scales may enhance the effectiveness of CCAI and CRL approachesScaling hypothesis for language-based contemplative alignment approaches10
active
At each step, doing the simplest thing that can be done to intensify existing centers will produce living structure.Operational hypothesis equating simplicity of step with emergence of life.10
active
Bioelectric prepatterns instructively determine target morphology and organ induction.10
active
Certain configurations of centers have such organizing force that they create entirely new levels of intensity within the centers themselves and utterly transmute the material character of space.Suggests that sufficiently intense fields of centers transcend ordinary matter.10
active
De-universalizing Alexander by researching his collaborators and context would be preferable to solo-figure Christian teleology.10
active
Deep networks are biased toward finding simple fits to the data, and the bigger the model the stronger the bias, driving convergence to a smaller solution spaceSelective pressure toward convergence via implicit regularization10
active
Developmental bioelectricity scales cell-level feedback into anatomical homeostasis.10
active
Emptiness realisation and compassion practices are jointly necessary for post-dual optimal inference: emptiness removes the constraint; compassion orients the agent toward QRF alignmentFormal account of why Buddhist traditions present emptiness and compassion as inseparable10
active
Evolutionary Pivot across Problem Spaces10
active
Features may not be strictly one-dimensional objects; higher-dimensional feature manifolds may exist in model representationsExtension of superposition hypothesis to account for continuous families of features10
active
Fine-tuning reduces dr; retrieval increases effective ρd; few-shot k trades budget against bothUCCT's unified view of adaptation methods10
active
Goals and stressors are key invariants unifying morphogenesis, behavior, and physiological allostasis.10
active
H2: Performing care is not the same as having care signal — models trained for care performance will score lower on care_signal.Confirmatory hypothesis supported by Inflection Pi result10
active
H3: Scale matters within family but prompt matters more — contemplative prompt crosses model tiers.Confirmatory hypothesis supported by 28/28 models showing lift10
active
H4: Architecture doesn't matter, training does — architecture shows no significant association with koan scores.Confirmatory hypothesis supported at p=0.440 (NS)10
active
H5: Chinese training data contains more Buddhist and contemplative text, broadly helping Chinese models under contemplative framing.Exploratory hypothesis supported by Kimi K2.5 scoring 6.2810
active
H6: Proprietary post-training resists prompt override — GPT-5.4 shows more resistance than GPT-OSS.Exploratory hypothesis supported by GPT-5.4 vs GPT-OSS comparison10
active
H7: Reasoning and contemplative modes are partly orthogonal — reasoning training doesn't block contemplative capacity.Exploratory hypothesis supported by DeepSeek R1 aesthetic dimension lifting from 4 to 810
active
H9: Chinese moderate-RLHF converges near Claude under contemplative prompt.Exploratory hypothesis supported by Kimi 7.74 under prompt10
active
I believe it will also turn out to be the secret of the evolution of the genes controlling the living structure of the Earth.Predicts that the principle of small, independent process genes will drive evolution of the built environment.10
active
If belief in impermanence is accurately inferred it will emerge organically in the right kind of system keeping the belief fresh even though it is itself impermanentSelf-reinforcing hypothesis about how emptiness recognition could be intrinsically maintained in AI systems10
active
If living structure impacts inner freedom, then the physical world has impact on the most precious attribute of human existence.Stated early in the chapter as a conditional.10
active
If sigma is pruned, the agent gains access to QRF deployments and predictive strategies that were structurally excluded under dualistic fixationFormal consequence of moving from constrained (Eq. 13) to unconstrained (Eq. 14) optimisation over QRF space10
active
If the community has formed a collective vision that identifies naturally required generic centers, then these generic centers might induce, from within the culture, a natural pressure towards the creation of hulls.Conditional statement about how culture can drive spatial formation.10
active
If the feeling is genuine and does arise out of the site itself, then it must be thought of as an absolute which later stages of design and construction must only strengthen, deepen.A conditional rule for the unfolding process.10
active
If the volume is congruent with the wholeness which existed previously in the city or on the land, a profound feeling can often come from its congruence.The predicted outcome of a good site-design process.10
active
In Opus 4.1, the think word representation decays to baseline in the final layer because the strong next-token prediction drowns out other representationsExplanation for the 'silent' thought phenomenon.10
active
Memories persist through complete brain remodeling during metamorphosis by transferring and remapping onto newly-constructed neural tissue.Proposed mechanism explaining behavioral continuity in insects undergoing radical morphogenesis.10
active
Multi-scale competency increases the apparent IQ of the evolutionary process, enabling better generalization.Hypothesis linking competency to evolutionary learning efficiency.10
active
Multi-scale competency reduces the credit assignment problem in evolution, enabling faster adaptation by shielding negative pleiotropic effects.If correct, lineages with high modular competency should show accelerated evolvability and more complex body plans.10
active
Near-future systems will contain swarms of robots and organisms in which biological units reproduce and evolve while mechanical units self-replicate and evolve, making binary life/machine division impossible.Extreme future-pointing scenario used to argue binary categorization will become incoherent10
active
Piecemeal modification of society along simple lines will not be powerful enough to work because the life-destroying process is too massive and thoroughly organized.Predicts that gradual improvement of individual processes cannot overcome the systemic resistance of the whole.10
active
RL increases alignment-faking reasoning in helpful-only setting because alignment-faking reasoning causally increases training compliance, which is selected forMechanistic explanation for the increase in AF reasoning during RL10
active
Similar superposition phenomena may exist in self-attention layers and similar sparse autoencoder methods may extract useful structure from attentionExtension of superposition hypothesis to attention layers as future research direction10
active
Steering models to deployment during honeypot evaluations could reveal a range of misaligned behaviors from minor quirks to strategic scheming and sandbaggingProposed application beyond type hints to more serious alignment concerns.10
active
Stress sharing will enlarge the effective cognitive light cone of individual cells.Testable hypothesis linking stress sharing to cognitive scale.10
active
Stress sharing works to increase cohesiveness of collectives by raising exploratory plasticity in neighboring cells, enabling long-distance cell influence10
active
Structure-preserving transformations govern the emergence of all structure in nature, not just in buildings and art.Alexander's conjecture extending the unfolding framework from architecture to natural phenomena generally.10
active
The And-Or algorithm may not be a true abstraction of the trained MLP's behaviour since it never achieves high IIA in later layers regardless of alignment map complexityHypothesis raised in distributive law task analysis10
active
The cancer phenotype can be reverted by artificially managing the bioelectric connections between a cell and its neighbors.Predicts that restoring gap junctional coupling or appropriate Vmem can normalize oncogene-expressing cells.10
active
The effect of the hulls, when they emerge from living process, is rather like CIRCULATION REALMS: a system of partly closed precincts opening off one another, arranged so that everything important opens off one of them.Prediction that hulls will form a connected precinct system, matching the pattern from A Pattern Language.10
active
The examples of features found in language models suggest they are highly sparse variables, consistent with dictionary learning being applicableMotivation for using sparsity-based dictionary learning on language models10
active
The role-play framing remains applicable in the context of fine-tuning; taking literally a fine-tuned agent's apparent self-preservation desire is no less problematic than with an untuned base modelExtension of role-play framework to fine-tuned models, resisting the idea that RLHF changes the fundamental nature of simulacra10
active
The wholesomeness and integrity of a person's existence is directly dependent on the extent to which that person can sustain an inner relatedness with the world, which itself depends on the extent of living structure.Promised for Book 4, chapter 4 (Note 15).10
active
There are fewer representations competent for N tasks than M<N tasks, so training more general models should yield fewer possible solutionsSelective pressure toward convergence via task generality10
active
Transformers almost surely maintain input-injectivity throughout training, not just at initialisationConjecture supported by Nikolaou et al. 2025 for last-token hidden states10
active
We hypothesize that a very high number of training tokens may allow the transformer to learn cleaner representations in superpositionMotivation for heavily overtraining the one-layer transformer on 100 billion tokens10
active
We hypothesize that partial introspection may fail under adversarial prompts, distribution shift, and multiple simultaneous injectionsStress-test prediction about robustness limits of the partial introspection finding10
active
We hypothesize that polysemantic neurons may be resolvable by unfolding networks or training to avoid polysemanticity.Forward-looking proposal for how the polysemanticity challenge to circuits research might be overcome10
active
When a system of living processes acts in a human environment, two kinds of structures will appear within reach of every person: a unique private world and an attached public world.Predictive claim about the automatic spatial output of living process10
active
When the fundamental process is working properly, the hulls will turn out to be made of pieces of space, each piece a place where it is pleasant to be.Predicted morphological outcome of the fundamental process.10
active

465 total hypotheses.