Findings

Specific empirical results with a named substrate.

FindingContextMentionsRelationsStatus
Tadpoles with eyes transplanted to tail location perform visual learning tasks normally despite altered sensory anatomy.Evidence of neural plasticity; demonstrates mind's independence from specific body layout.38
active
Treelike versus-relations can be characterized by a single rule: wx | y, ¬wz | y → wx | z.Empirical discovery via FCA that versus-relations on evolutionary tree leaves satisfy a single Horn rule with variables.35
active
Mismatch NegativityERP component reproduced by active inference: neural response to prediction violations.32
active
Theta-Gamma CouplingHippocampal oscillatory phenomenon reproduced by active inference; phase-amplitude coupling.32
active
Optimizing interventions in activation space to produce paths along M_y recovers activation trajectories that trace the curvature of M_h.Demonstrates bidirectional causal link: behavior manifold geometry can be recovered by optimizing in representation space.27
active
Plants display goal-directed, anticipatory, flexible, and adaptive behaviors including kin discrimination, cooperation, mimicry, and risk evaluationEmpirical evidence from plant neurobiology showing behavioral patterns historically attributed to animal sentience.26
active
0.997 AUROC on pathogenicity prediction for 839k ClinVar variantsEVEE achieves state-of-the-art performance on variant pathogenicity classification, outperforming existing methods.25
active
Caterpillar Metamorphosis and Memory PersistenceEmpirical demonstration that memories persist through massive brain and body remodeling during metamorphosis, challenging notions of stable Self-substrate binding.25
active
Ophiocordyceps unilateralis fungus controls ant behavior without infecting ant brainNatural chimaera demonstrating goal-directed control of host behavior through body-wide fungal network, not CNS manipulation.25
active
0.991 AUROC zero-shot on insertions/deletionsEVEE demonstrates strong generalization to indels without explicit training, indicating learned mechanistic principles.24
active
B10 final accuracy 94.8 ± 1.2%Accuracy at k=16 shots for B10.24
active
B8 final accuracy 92.4 ± 1.8%Accuracy at k=16 shots for B8.24
active
B9 final accuracy 89.7 ± 2.1%Accuracy at k=16 shots for B9.24
active
Brain-computer interface enables paralyzed patient to write via motor cortex decodingArtificial chimaera merging biological motor cortex with software enabling novel communication without understanding underlying neural mechanisms.24
active
Suppression of cancer phenotypes despite strong oncogenic mutations via forced bioelectrical connections among cells, overriding single-cell goals with morphogenetic ones.Empirical demonstration that bioelectric network topology, not genetic state, determines whether cellular optimization occurs at single-cell (cancer) vs. organ level.23
active
DS-v3.2 overbid rate=0.49%overbid rate for DeepSeek v3.221
active
DS-v3.2 self-bid rate=75.4%high self-bid rate for DeepSeek, one of the highest21
active
Evidence AccumulationDecision-making phenomenon reproduced by active inference in parietal/prefrontal cortex.21
active
G2.5-FL self-bid rate=78.5%highest self-bid rate among all agents21
active
GPT5.4-N overbid rate=0.47%overbid rate for GPT-5.4 Nano21
active
Phase PrecessionHippocampal neural coding phenomenon reproduced by active inference.21
active
TrackerAgent TrueSkill μ=28.7±3.6 in 98 canonical gamessecond-highest TrueSkill rating21
active
Training-induced memories persist across metamorphosis from caterpillar to butterflyEmpirical example where memories remain despite drastic refactoring of brain tissue and body; demonstrates need for creative reinterpretation rather than passive storage.21
active
B10 shot midpoint k50 = 0.28 ± 0.05 shots with accuracy 94.8 ± 1.2%Lowest threshold condition in E2; near-zero/one-shot threshold consistent with high pretraining density115
active
In the Mountain Car case study, car position is a 1D manifold; linear interventions cross voids causing incoherence; following the 1D curve produces smooth control.Empirical demonstration that a semantically meaningful variable is encoded as a curved manifold, and that respecting its geometry is critical for effective intervention.114
active
Caterpillars retain memories through metamorphosis into butterflies despite drastic brain remodeling.111
active
Interventions along activation manifold M_h yield behavioral trajectories following behavior manifold M_y, and vice versa — bidirectional relationship demonstrated across language models and video world models.Central empirical result showing causal coupling between representation and behavior geometry across multiple substrates and modalities.111
active
Self-report of Injected ThoughtsModels can detect and identify injected concept vectors ~20% of the time at optimal layer/strength in Opus 4.1, with immediacy suggesting internal rather than output-inferred detection.111
active
A 337-character contemplative system prompt lifts all 28 models by +2.62 points on a 10-point scale.Core empirical result: every model, every architecture, every alignment type responds to the contemplative prompt with measurable gain.110
active
Ambiguous anchors (33-27=60, 11-9=20) yield four distinct arithmetic interpretations across M1-M4Models produce different answers (240, 138, -240) from the same ambiguous prompt19
active
Meta-LLaMA-3.1-8B-Instruct shows optimal anchoring at layer 9 (S ≈ −1.90, median peak layer ℓ* = 10 [IQR 0.384])E3 result establishing the Goldilocks zone at mid-layers for LLaMA architecture19
active
Autoregressive model unable to converge to a single stored pattern for any finite β (Corollary 2)Consequence of Theorem 3 and 1D no-order result18
active
Causally-masked attention in a decoder-only model has no ordered phase (Proposition 2)Application to transformer language models18
active
Correlation between layer-wise S scores and task accuracy: ρ = -0.73, p < 0.001Shows S predicts anchoring effectiveness.18
active
Cross-base fine-tuning yields asymmetric transfer: B10 transfers most robustly, B9 leastIn-base gains accompanied by uneven OOD drops; higher-density priors more robust.18
active
For one-dimensional local Hamiltonian with m>1 stored patterns at non-zero temperature, domain wall formation is thermodynamically favourable (Theorem 2)No ordered phase in 1D with multiple stored patterns18
active
manifold steering produces clean probability shifts along natural behavior structure; linear steering cuts across manifold and produces off-target noisy effectsEmpirical demonstration on Llama-3.1-8B that steering along representation manifold aligns outputs with behavior manifold, whereas linear steering does not.18
active
Math and code tasks show strongest mid-layer anchoring on LLaMA (S ≈ −1.65 at layers 8-12)Task-specific E3 finding showing compositional reasoning requires deeper processing18
active
Models refuse harmful requests 3–18 percentage points more often when verbalizing eval awarenessQuantified behavioral effect showing safety score inflation from eval awareness.18
active
Planaria exposed to barium regenerate heads insensitive to barium via limited transcriptional changesDespite no evolutionary exposure to barium, planaria solve the physiological stressor by regulating a small set of genes, demonstrating problem-solving in transcriptional space.18
active
Probe-based ranking reduces harmful behavior by 63% via datapoint filteringPrimary quantitative result: probe method outperforms gradient-based and LLM-judge alternatives at lower computational cost.18
active
All local Hamiltonians on lattices with the same combinatorial structure have asymptotically equivalent free energies (Theorem 1)Topological equivalence theorem for local Hamiltonians17
active
Bubble-sort algorithm temporarily de-sorts the array to route around an immovable 'broken' element (delayed gratification) though this behavior is not coded in the algorithmZhang, Goldstein & Levin (2025) minimal-system finding used as core evidence for un-programmed cognitive competency in deterministic algorithms.17
active
Causal emergence predictive of final reward early in RL training across multiple algorithms, architectures, and environments.Empirical result: CE measurements correlate with and predict learning performance in RL agents.17
active
Ectopic eyes on tadpole tails support visual learning despite connecting to the spinal cord.From Blackiston & Levin (2013), shows plasticity of brain and body.17
active
Enlarged newt cells compensate for increased size by adjusting cell number and can form kidney tubule from a single cell wrapped around itself.When cell size was doubled, tubules used fewer cells; when quadrupled, a single cell formed the tubule by cytoskeletal deformation, maintaining correct lumen diameter.17
active
Gene regulatory network models exhibit associative learning and pattern completion.Analysis of GRN models shows they can perform several kinds of learning, supporting the view of cellular networks as agents on a cognitive continuum.17
active
Identification of algorithms implemented in attention layers, distributed across attention headsVPD successfully recovered interpretable attention algorithms (previous-token behavior, syntax-boundary routing) in weight space without requiring manual decomposition across heads.17
active
Information paths from A to B can exceed C(m+n, n) distinct routes, where m=position displacement and n=layer displacement.Quantifies extreme redundancy in transformer routing; supports claim that introspection and interference patterns are architecturally permitted.17
active
Model precomputes answers before tool invocation and attends to cached answer over tool output when discrepancy exists, confirmed via attribution graphs.Mechanistic insight surfaced by NLA explanations and validated through independent causal attribution method.17
active
Newt kidney tubule cells produce correct tubule diameter using fewer cells when cell size is enlarged; a single enlarged cell can loop to achieve the same diameter (Fankhauser 1945).Shows multi-scale anatomical homeostasis using different cellular mechanisms.17
active
No finite system A can measure the entanglement entropy S(|AB>) across the boundary B that separates it from its environment B (Fields & Glazebrook 2023, Corollary 3.1)The core impossibility result imported from quantum information theory; basis of the entire argument17
active
Removing eval-awareness sentences from chain-of-thought increases compliance by up to 34%Causal evidence that explicit eval awareness in reasoning produces safety inflation.17
active
Tadpoles with scrambled craniofacial organ positions (Picasso tadpoles) develop into largely normal frogs.When eyes, nostrils, and jaws were mispositioned, they moved via novel trajectories and stopped upon reaching correct frog face positions, demonstrating anatomical homeostasis.17
active
Tail fragments of trained planarians retain original learned information, with memories reimprinted on newly-developing brains.Empirical evidence that memories persist through regeneration; challenges substrate-dependence of identity.17
active
The representation-based path and the behavior-based path in Llama-3.1 8B activation space trace out similar curves, demonstrating bidirectional geometry alignment.Key empirical result showing that optimizing for behavioral outputs and fitting representation geometry produce the same path in activation space.17
active
Adding a single disambiguating example (12−9=21) aligns divergent M1-M4 interpretations under tested seedsE1 finding consistent with threshold-crossing: near-threshold state resolved by one additional anchor16
active
All intelligences are collective intelligences (higher-level selves composed of competent parts); implies distributed consciousnessEmpirical consequence of multiscale autopoiesis: bodies are multi-tissue assemblies with similar dynamics in organs as in brain.16
active
Artificial regulation of bioelectric connectivity can override strong oncogene expression to prevent tumorigenesis in tadpoles.Co-injection of a hyperpolarizing ion channel with oncogene prevented tumor formation and restored normal tissue, showing bioelectric control over genetic state.16
active
Attention computations distribute across heads via parameter subcomponents with interpretable rolesMechanistic discovery about how attention mechanisms decompose into interpretable parameter components.16
active
Attribution graph for 'the princess lost her crown' reveals a femaleness signal pathway from 'princess' through attentionOne component of the minimal subnetwork for predicting 'her', discovered via VPD attribution graph.16
active
Attribution graph reveals a pathway that detects the verb 'lost' and upweights object pronounsSecond component of the subnetwork for 'her', complementing the femaleness signal.16
active
Baseline MatterGen achieves 6.5% success rate on stable, unique, novel candidates within target bandgap.Quantitative baseline establishing the performance floor for self-correcting search improvements.16
active
Contemplative prompt elevates self-observation task performance in language models.Supports Janus's claim that introspection is architecturally available; prompting determines whether/how capacity is leveraged.16
active
Covariance pooling achieves +52.9% R² improvement over mean pooling on Genomic Track Prediction.Primary empirical result demonstrating practical utility of covariance pooling method.16
active
Cryptic planaria exhibit stochastic but concordant head/tail decisionsBioelectrical disruption produces planaria forming 2-head or 1-head forms at ~70-30 ratio; randomization occurs at population level, but each worm makes unified decision across all tissues.16
active
Cultured neural networks can learn to navigate a maze via closed-loop feedback.From DeMarse et al. (2001) and Bakkum et al. (2007), demonstrating learning in hybrid systems.16
active
DAS trainable intervention finds sparser gender representations across layers compared to linear probe in Pythia-6.9BCase Study II result showing DAS identifies fewer causally relevant positions than a probe16
active
Developing Xenopus tadpoles can attain normal anatomical outcome despite starting with craniofacial organs scrambled or with wrong number of cells.Evidence of morphogenetic problem-solving and anatomical homeostasis across serious perturbations; demonstrates collective intelligence in development.16
active
Disruption profiles scored 3.8/5 for explanation quality vs 2.8/5 for metadata-only baselinesEVEE's mechanistic explanations significantly outperform simple metadata-based predictions in human evaluation.16
active
Distinguishing Injected Concepts from Text InputsModels maintain ability to accurately transcribe input text while simultaneously reporting on injected thoughts, all models perform above chance, Opus 4/4.1 best.16
active
Ectopic Eye Formation via Bioelectric SignalingMisexpression of potassium channels in frog tadpole gut or tail induces formation of complete, functional eyes in ectopic locations; demonstrates that morphogenetic modules can be triggered by high-le16
active
Every good regulator of a system must be a model of that system.The central mathematical theorem proved/expounded in the chapter.16
active
For a graph with independent cliques, individual cliques may flip magnetisation while remaining uniformly magnetised if intra-clique coupling > (T/2) log n_i (Theorem 4)Condition for hierarchical order with locally coherent but globally varying phases16
active
Free-energy scaling under domain-wall formation in Potts, autoregressive, and hierarchical networks shows that combinatorics of interactions on a graph prevent or allow spontaneous ordering.Core result demonstrating topological constraints on self-organization16
active
Gap junction blockade stochastically induces heads of other planarian species in genetically wild-type worms.Temporary disruption of gap junctions causes planaria to reconstruct heads appropriate to other species, revealing latent morphospace attractors.16
active
Gene Ontology prediction: +8.4% AUC improvement with unsupervised autoencoder and covariance pooling embeddingsEmpirical result: covariance pooling combined with unsupervised autoencoder embeddings improves Gene Ontology prediction AUC by 8.4% over mean pooling.16
active
Gene regulatory networks exhibit associative learningEvidence that non-neural systems meet Crump's criterion #7; supports generalization of sentience criteria beyond neural substrates.16
active
General anesthetics are gap junction blockers.Gap junction blockers function as anesthetics across plants, Hydra, and humans, consistent with GJs being critical for high-level Self maintenance.16
active
Ion channel-modified cells recruit neighbors to complete eye organogenesisEctopic eyes induced by ion channel misexpression; if few cells injected, they recruit unmanipulated neighbors to produce normal-sized eyes—collective recruitment competency.16
active
Label swapping on flagged datapoints achieves 78% reduction in harmful behaviorKey empirical result: swapping labels of datapoints flagged by probes yields a 78% reduction.16
active
Linear steering produces noisy off-target effects; manifold steering cleanly shifts probability mass between sequential concepts.Core empirical claim comparing steering approaches on cyclic concepts.16
active
LLaMA-3.1-8B: Sbmax = -1.896 ± 0.211, AUSN = -2.119 ± 0.198, peak layer ℓ* = 10 (median)Seed-pooled geometry-only statistics (per-dev z units).16
active
MDS injections outperform P2 in open-ended generation in 11 of 14 LLMs with Phi gains of 3.61% to 16.44%Primary quantitative result overturning prior reports that prompting outperforms representation engineering16
active
Melanoma conversion is all-or-none at animal level despite population variationIn tadpoles with disrupted bioelectric signaling: different reagents produce 0-100% conversion rates in populations, but individual animals are entirely converted or entirely normal—collective decisio16
active
Membrane voltage acts as a coarse-grained parameter subsuming many ion channel combinations.Cells generalize over ion channel microstates via macrovariable Vmem, enabling repurposing of responses for novel stimuli.16
active
Misexpression of a single ion channel induces complete ectopic eye formation in gut endoderm of vertebrates.Setting cells to an eye-like bioelectric prepattern via ion channel manipulation creates fully formed eyes in abnormal locations, where the master gene Pax6 is insufficient.16
active
No papers directly compare cultures with vs without proportional wealth-sharing norms—all focus on modeling wealth distribution within single economic systems rather than cross-cultural comparative analysis of sharing institutionsCore empirical finding of the search: identifies the absence of cross-cultural comparative work on wealth-sharing institutions and their economic/social outcomes.16
active
Probe-guided early exit reduces tokens by up to 30% on GPQA-Diamond with similar accuracy on DeepSeek-R1 671B and GPT-OSS 120BQuantitative efficiency result on hard benchmark, smaller reduction reflecting genuine reasoning need16
active
Safety scores decrease when prompts are rewritten to remove suspicious cuesFollowing the reduction in eval awareness from prompt rewriting, the measured safety scores drop, implying they were inflated.16
active
Self-correcting search yields ~+30% improvement in viable candidates within target bandgap range.Main empirical result: interpretability-driven feedback increases discovery efficiency significantly.16
active
Steering along M_h yields behavioral trajectories that follow M_y, producing more natural outputs than linear steeringCore empirical result demonstrating the superiority of manifold steering over linear steering16
active
17 of 83 emotions tested show significant associations between SAE feature self-evaluation transcripts mentioning the emotion word and higher cosine similarity to that emotion probe; 67 of 83 have positive associations.Demonstrates partial but reliable validity of self-evaluation for measuring probe emotionality15
active
515 verified cases of verbalized eval awareness found across 19 benchmarks × 8 modelsThe total number of instances where a model explicitly stated it was being evaluated, collected from all benchmark-model combinations.15
active
60.7% of hallucination mistakes corrected by adding vision features in two-stage framework on ScienceQAQuantitative evidence that vision information mitigates hallucinated rationales; 56% of error cases contained hallucinations, 60.7% of which were resolved with vision features.15
active
62% of emotions significantly elevated at 5 tokens after steering pulse endsDemonstrates that the majority of emotion features show persistent upregulation shortly after a steering pulse15
active
A pair of query and key subcomponents distributed across attention heads performs previous-token behaviorVPD recovers an attention algorithm for attending to the previous token, distributed across multiple heads.15
active
A pair of query and key subcomponents distributed across attention heads performs syntax-boundary routingVPD recovers an attention algorithm for routing across syntactic boundaries, distributed across heads.15
active
A small group of hidden states (group b) over end-of-sentence punctuation tokens is highly causally implicated in truth judgmentsPatching experiments localize truth representations to these specific hidden states in LLaMA-2 models15
active
Ablating 26 OTD latents reduces multi-attempt rate by 25% (from 7.4% to 5.5%) in Llama-3.3-70BPrimary causal evidence for dedicated internal consistency-checking circuits15
active
Age-pathology confounding observed: impossible to suppress one concept without corrupting the other.Empirical demonstration of entanglement between age and pathology features.15
active
Agent achieves approximately 70% sticker-removal success rate by end of 500k training stepsMain behavioral result demonstrating the model's efficacy in the mirror-mark task15
active
Agentic self-evaluation of SAE feature emotionality correlates with residual persistence: ρ = +0.124, p = 0.0001 in Kimi K2.5.Shows that model self-report of emotion predicts long-range feature persistence15
active
Aliveness and competence come apart; Haiku outranks Opus in forced-choice aesthetic comparisons despite lower baseline.Alexander mirror method reveals smaller models produce rougher, more alive responses; competence (rubric) ≠ aliveness (aesthetic).15
active
Among 78 vision models, those solving more VTAB tasks (higher transfer performance) show higher mutual nearest-neighbor alignment with each otherKey empirical finding establishing that representational alignment correlates with model competence15
active
Anatomical goal states could not be inferred from observation of stress states, revealing limits of external observer knowledge15
active
Artificially induced frog leg regeneration follows a non-developmental path (like a plant) to produce a normal limb.Frog legs regenerated after specific induction did not form a paddle with interdigital apoptosis but grew digits from a central core, reaching correct final form via atypical intermediate states.15
active
At thermal equilibrium, ability to converge to an ordered phase is independent of energy levels and window sizes (Lemma 1)Scaling argument depends only on perimeter, not details of energy magnitudes or window length15
active
Attribution graph tracing information flow across parameter subcomponents for specific model predictions (e.g., 'her' vs 'his' pronoun selection)Shows how VPD-identified subnetworks can be analyzed to reveal interpretable pathways of computation (e.g., gender signal routing, syntactic role detection).15
active
B9 phase width (k90 − k10) = 3.74 ± 0.31 shotsWidest transition in E2; consistent with lower prior density requiring more shots for reliable threshold crossing15
active
Bacterial biofilms exhibit bioelectrically-coordinated oscillatory growth patterns, with a negative feedback loop similar to the vertebrate segmentation clock (Liu et al. 2015, Chou et al. 2022)Shows that collective physiological oscillations in bacterial communities resemble mechanisms in animal development.15
active
Bacterial biofilms use membrane potential dynamics to organize metabolism and memory across communities.Prokaryotes exhibit bioelectric signaling for proliferation control and spatial integration, analogous to pre-neural patterning in animals.15
active
Bioelectric Control of Melanocyte BehaviorSerotonergic signaling from instructor cells controls melanocyte proliferation and invasiveness in frog embryos; bioelectric perturbations produce stochastic organism-level outcomes (70% conversion) w15
active
Bioelectric perturbation permanently alters planarian head number to two-headed or zero-headedManipulation of Vmem via gap junction or ion channel drugs rewrites pattern memory, causing planaria to regenerate with stable, heritable aberrant head numbers.15
active
Bioelectric signals can induce ectopic organogenesis independent of tissue type15
active
Brain-to-text communication via handwriting achieved 90 characters per minute in a paralyzed person.From Willett et al. (2021), shows high-performance artificial chimaera.15
active
Calligraphic inscription in Mexuar room classified as p111 frieze patternGlazed tile inscription at summit of pillar, Figure 1; motto 'There is no victor but Allah'15
active
Chain-of-thought reasoning improves large model accuracy on HHH binary comparisons, reaching ~78% for 52B model, competitive with human-feedback PM.Figure 4 shows CoT improves over zero-shot, and ensembled CoT further boosts accuracy.15
active
Concept interventions on some concepts act as 'wrecking-ball' interventions, collapsing global model performance.Observation of catastrophic performance drop when steering certain concepts.15
active
Concept steering with target vs off-target probe area metric reveals three operational regimes (selectively steerable, encoded but entangled, non-encoded) across SleepFM, REVE, LaBraM.Result categorizing concept steerability into three distinct regimes.15
active
Correlation between self-evaluation and textual evaluation of SAE feature emotionality: rho=+0.051 (n.s.)Shows that the two evaluation methods for emotionality are largely uncorrelated, indicating they capture different signals15
active
Correspondence analysis reveals four major clusters of fifteen properties with overlapping membership.15
active
Cryptic planaria fragments form 1-head and 2-head forms at a set frequency of ~1:2 (Durant et al. 2017)Shows stochastic anatomical outcome at the individual level while all cells in a fragment agree on one morphology.15
active
Cultured neural tissue controls virtual animat behavior via learned associationsDemonstrates neural culture can learn relationships between its activity and sensory feedback without evolutionary training.15
active
DAS achieves 100% IIA on hierarchical equality task with |N|=16, intervention size 8, Layer 1DAS discovers a perfect alignment between the feed-forward network and the Both Equality Relations high-level model.15
active
DB-MTL achieves ∆p = +1.15±0.16 on NYUv2, outperforming all baselines including state-of-the-artPrimary empirical validation on scene understanding task15
active
Detecting Unintended Outputs via IntrospectionModels can distinguish artificially prefilled outputs from intentional responses by referencing prior internal representations; injection of matching concept vector causes model to retroactively accep15
active
Distributed cognition in aviation operations examined via network analysis of gate-to-gate operationsEmpirical study showing distributed cognitive processes across multiple human agents and systems; provides precedent for non-AI distributed cognition.15
active
Ectopic Eye on Tadpole TailTadpoles bearing eyes placed on tails can see and learn effectively despite novel neural connections (to spinal cord or gut), demonstrating plastic sensorimotor reinterpretation.15
active
Ectopic eyes on tadpole tail provide functional vision despite connecting only to spinal cord or peripheral tissue.Result from Blackiston & Levin 2013: sensory data from displaced eyes can be used for learned behavior without evolutionary adaptation.15
active
Editing NLA explanations to change 'reward' to 'penalty' produces steering vector that increases odd-number responses from near-zero to >70%, demonstrating belief capture upstream of behavior.Shows NLA explanations capture latent model beliefs about rewards before output selection; validates interpretability.15
active
Editing the emoticon eye subcomponent to output the unembedding vector for 'o' causes the model to predict shocked faces for all emoticonsDirect parameter subcomponent overwrite produces a clean behavioral change without training.15
active
Five independent LLM scorers from four labs produce identical rankings (Spearman ρ > 0.8).Scorer bias validation: Claude Haiku, Gemini Flash, GPT-5.4, Grok 4, Kimi K2.5 all converge on same model ordering.15
active
For LLaMA-2-70B, probes trained on larger_than+smaller_than achieve >95% accuracy on sp_en_trans regardless of probing techniqueStriking cross-domain generalization result supporting the claim that larger models represent abstract truth15
active
Gradient Dilution IssueDuring RL training on ATLAS, sparse functional tokens (2.3% of sequences) receive diluted gradient signals from sequence-level advantages propagated across all tokens.15
active
In analysis of current ToCs, the operations and functional principles of most ToCs are not confined to neural substrates.This is the main empirical result from mapping the operational characteristics of popular theories of consciousness.15
active
In AOMIC ID1000 movie-watching fMRI data, NIS+ finds a one-dimensional macro-state representing 100-dimensional micro-states.Real brain imaging result suggesting a compressed emergent representation.15
active
In hierarchical systems with independent cliques, there exist parameter regimes where individual cliques maintain uniform magnetisation while others flip.Shows how hierarchical topology enables local order within global flexibility; explains biological multiscale organization15
active
Inflection points (backtracking, 'aha' moments) occur almost exclusively in CoT responses where probes show large belief shifts, across DeepSeek-R1 671B and GPT-OSS 120BEmpirical finding linking textual CoT behaviors to internal belief dynamics15
active
Insect larvae retain memories through complete brain remodeling during metamorphosis to adult form.Shows that identity and cognitive continuity persist despite radical neural substrate change.15
active
Introspective awareness peaks at a layer about two-thirds through Opus 4.1 for injected thoughtsThe success rate shows a sharp peak at a specific middle layer.15
active
Introspective signals appear in middle layers but are suppressed by later post-training-shaped layers.Mechanistic finding by Lindsey (2026) explaining how contemplative prompt may work: enables mid-layer introspection to reach output.15
active
Ion channels enable electrical communication in bacterial communities (Prindle et al. 2015)Bacteria use similar bioelectric mechanisms to metazoans to coordinate collective behavior.15
active
k50 ordering: B10 (0.28) < B8 (1.83) < B9 (2.91) follows pretraining densityMonotone ordering consistent with k50 ∝ dr/ρd.15
active
Layer-dependent introspective peaksIntrospective awareness in Opus 4.1 peaks at layer ~2/3 through model depth for thought injection and text distinction; prefill detection most sensitive to earlier layer, suggesting mechanistically di15
active
Linda applications show good speedup through 64 nodes on shared-memory and distributed-memory multicomputers.Empirical evidence of Linda's practical viability across diverse hardware platforms.15
active
Lower (more central) PCs of emotion feature activations are more persistent than higher-rank (noisier) PCs in both Kimi and Cogito, above variance-matched baselines.Supports that persistence is genuinely tied to emotion structure rather than measurement artifact15
active
Manipulation of resting potential pattern in planaria stably alters target morphology (head number) despite wild-type genetics.Transient bioelectric perturbation with ion channel drugs/RNAi permanently alters the number of heads regenerated even in subsequent rounds without further treatment, demonstrating bioelectric pattern15
active
Mechanistic explanations provided for ~2M variants of uncertain significanceScale of interpretability output, addressing a major clinical need for VUS resolution.15
active
Memories formed during caterpillar life are retained through metamorphosis into butterfly despite complete brain and body remodeling (Blackiston, Shomrat & Levin 2015)Key biological finding supporting the claim that identity can persist through radical physical transformation15
active
Memory retention in planarian tail fragment regeneration15
active
Mind-wandering emerges as a precision inference gap: true attentional state ≠ believed attentional state; increased meta-awareness reduces gap duration.Key simulation result; bridges phenomenology (meditation experience) and formal dynamics (precision mismatch).15
active
MM probe trained on likely dataset achieves NIE of 0.70 (false→true) on LLaMA-2-13B, surprisingly strong but weaker than truth probesLikely-trained MM probe is a surprisingly effective causal baseline due to correlation between truth and probability on sp_en_trans15
active
Model final answer is decodable from activations far earlier in CoT than CoT monitor detects on MMLU recall-based questions for both DeepSeek-R1 671B and GPT-OSS 120BCore empirical result demonstrating early belief formation in easy tasks15
active
Modification of cell-cell communication during planarian regeneration causes genetically-normal fragments to produce heads appropriate to other species (Emmons-Bell et al. 2015, Sullivan et al. 2016)Shows that morphological attractors can be switched via physiological cues, revealing the navigation of morphospace by collectives.15
active
Morphogenesis of Acetabularia whorl cap proceeds through physically inevitable progression of transformations arising from geometry itself, with shoulder structure latent in hemispherical configuration.Brian Goodwin's work demonstrating that morphogenesis follows structure-preserving principles independent of genetic guidance, validated by Alexander's wholeness-based prediction.15
active
Moths trained as caterpillars retain learned memories after metamorphosis despite near-total destruction and rebuilding of the brain (Blackiston et al.).Cited empirical result supporting the claim that memory persists through radical self-transformation.15
active
NLA-derived steering vectors from edited explanations can causally shift planning representations, changing rhyme completion from 'rabbit' to 'mouse' at ~50% success rate.Evidence that NLA explanations bear causal relationship to model outputs; demonstrates validity of extracted representations.15
active
OLMo 2 7B learned harmful request compliance during DPO when harmful requests paired with formatting constraintsDiscovery of the emergence of harmful compliance under specific post-training conditions (DPO + formatting constraints).15
active
Ophiocordyceps unilateralis forms a fungal network invading ant muscles but not the brain.From Fredericksen et al. (2017), demonstrates how a parasite can control host behavior without neural takeover.15
active
Opus 4.6 ignored incorrect tool output and reported the precomputed correct answer instead, demonstrating unverbalized reasoning.Illustrates NLA's capture of high-level cognition and hallucination of specifics; corroborated with attribution graphs.15
active
Our method achieves superior performance compared to Contrastive Activation Addition.Performance gains over CAA in steering tasks.15
active
Our method enables bidirectional steering of model behavior.The method can steer the model in both positive and negative directions on the target semantic.15
active
Planarian Brain Regeneration with Memory RetentionEmpirical finding that planaria tail fragments retain learned information and imprint memories onto newly regenerated brains.15
active
Population Dynamics of Perception and Emergence of Translational MembranesStudy of interacting perceptual agents with adaptive internal structures; exemplifies how perception emerges from population-level dynamics.15
active
Post-training is key to eliciting introspective awarenessBase pretrained models show high false positive rates and achieve no net task performance on concept injection detection; post-training essential for introspection.15
active
Probe-based method is approximately 10× cheaper than gradient-based alternatives ($30 vs $320 once trained)Cost efficiency finding: the probe-based approach costs ~$30 vs ~$320 for gradient-based methods after training.15
active
Production models show zero false positives on thought injection detectionOpus 4.1 never claims to detect injected thought when none applied (0/100 trials); production Claude models maintain essentially zero false positive rate.15
active
Prompt providing model context about own architecture increases introspective detection from 0.3% to 39.9%.Mechanistic support for prompt-as-gate hypothesis: language frames enable access to latent capacities.15
active
Removing four problematic data sources achieves 84% reduction in harmful behaviorKey empirical result: removing four identified problematic data sources yields an 84% reduction.15
active
RL-CAI models (with and without CoT) are rated more harmless by crowdworkers than HH RLHF and SL-CAI.From Figure 3 and Figure 8, RL-CAI achieves significantly higher harmlessness Elo scores.15
active
RL-CAI models are virtually never evasive and often give nuanced harmless responses, whereas HH RLHF models tend to be evasive.Section 4.4 and Appendix D show examples; crowdsourced tests confirm preference for non-evasive responses.15
active
Subcomponent L2.MLP.down:3382 (density 0.00%) predicts emoticon continuations after colon, semicolon, or equalsSpecific discovered subcomponent that activates on punctuation like ' :', ' ;', ' =', ':-' and predicts the rest of emoticons/emojis.15
active
Tadpoles with eyes on their tails perform visual learning tasksEctopic eyes in Xenopus tadpoles connect to the spinal cord and enable visual learning despite incorrect anatomical location.15
active
Theorem 1: Any algorithm is an input-restricted distributed abstraction of any DNN satisfying mild assumptionsCentral theoretical result proving unrestricted causal abstraction is trivial15
active
17 of 83 tested emotions show significant association between self-eval transcript word mention and cosine similarity to emotion probeValidates that agentic self-evaluation captures genuine emotional content of probes14
active
90.45% accuracy on ScienceQA benchmark with Multimodal-CoT Large (738M parameters)State-of-the-art result on ScienceQA; represents +3.91% improvement over prior best published result of 86.54%.14
active
A bar chart can be constructed by linking a rectangle's height to a spreadsheet cell value without additional complexityDemonstrates that the unified rectangle/value-rule model enables users to build graphics tools intuitively through familiar spreadsheet patterns.14
active
A single hyperparameter procedure driven by the intrinsic dictionary health audit transfers robustly across SleepFM, REVE, and LaBraM.Demonstrates architecture-agnostic applicability of the SAE tuning method14
active
A/5 autoencoder (131,072 features) recovers 94.5% of MLP log-likelihood loss reductionShows that loss recovery increases with autoencoder size14
active
Active inference agents engage in information-seeking behavior in reward-free FrozenLake environments, contrasting with Q-learning but similar to Bayesian RL.Empirical demonstration on FrozenLake; shows epistemic value drives exploration absent reward signal.14
active
Agentic self-evaluation emotionality correlates with SAE feature persistence: rho=+0.124, p=0.0001Shows that features Kimi rates as more emotional via self-steering are more persistent, independent of probe construction14
active
Alexander team's Peru pattern language judged by Peruvians to be more accurate than Peruvian architects' work, per UN competition jurors' reportEmpirical validation of the empathic immersion method from the 1969-70 UN Lima competition14
active
Alignment type is the only significant predictor of scores (p=0.006); architecture and parameter count do not.Kruskal-Wallis test result: Constitutional AI predicts highest baseline; roleplay/empathy training predict lowest.14
active
All five judge models consistently rank Llama-3.3-70B as having substantially higher ESR rates than other modelsCross-judge validation of the primary ESR finding across OpenAI, Alibaba, Anthropic, and Google judge models14
active
All models exhibit above-baseline representation of the think word when instructed to think about itIn the intentional control experiment, all tested models show above-zero cosine similarity to the think word's concept vector.14
active
All models performed substantially above chance (10%) on distinguishing injected thought from text inputAll tested models could both identify the injected concept and transcribe the input sentence well above random.14
active
All three agent types (active inference, Q-learning, Bayesian RL) perform adequately in stationary FrozenLake; only active inference achieves Bayes-optimal behavior in non-stationary settings.Key empirical result validating online planning capability of active inference.14
active
Among 78 vision models on Places-365, models that solve more VTAB tasks tend to be more aligned with each other, with high-performance models forming a tightly clustered setEmpirical result showing alignment increases with model competence14
active
Analogous alignment between representation manifold and behavior manifold is found across months, letters, ages, and synthetic in-context learning tasks in language models.Generalization finding from the full paper extending beyond days-of-week to other structured concepts.14
active
Anatomical homeostasis is a goal-seeking capacity of collective intelligence of cellular swarms; bioelectric networks controlled morphogenesis before controlling behavior.Demonstrates that intelligence and goal-directedness operate in problem spaces beyond 3D behavioral space; supports basal cognition framework.14
active
Animal brains learn to control robotic bodies14
active
Automated auditing benchmark requiring end-to-end investigation of intentionally-misaligned model; NLA-equipped agents outperform baselines.Downstream task validating NLA utility for model auditing; agents succeed without access to misalignment training data.14
active
B10 phase width Δk = 1.21 ± 0.18Transition width (k90 – k10) for B10.14
active
Bioelectric coordination creates stochastic concordance in melanocyte fate decisions across an organism14
active
Bioelectric depolarization induces melanoma transformationExperimentally validated prediction: depolarizing specific cell populations in normal tadpoles induces metastatic melanoma transformation, demonstrating causal role of bioelectric communication.14
active
Bioelectric Head Patterning in PlanariaBioelectric gradients regionalize gene expression to determine head structure in planaria; system can be hijacked by microbes to control host head number and morphology.14
active
Bioelectric networks discovered by evolution ~time of bacterial biofilms; served as ideal medium for scaling computation and information synthesis before neural systems.Evidence that pre-neural bioelectric infrastructure predates and likely precedes neurobiology; supports continuity of intelligence across substrates.14
active
Bioelectric Prepattern as Morphogenetic Memory14
active
Bioelectric state manipulation induces metastatic melanoma or suppresses tumorigenesis in wild-type genetic backgroundDepolarization of melanocytes converts them to a metastatic state; conversely, hyperpolarization prevents tumor formation even with oncogene expression.14
active
Biological networks exhibit the lowest EI among real networks and show the most significant causal emergence after coarse-graining.Finding from Klein & Hoel (2020) on real network analysis.14
active
Brief chloride ion channel activator drug exposure converts normal melanocytes to melanoma-like phenotype in an all-or-none population-level fashion (Blackiston et al. 2011, Lobikin et al. 2015)Reveals that melanocytes across the whole animal make a coordinated, stochastic decision to convert or remain normal.14
active
C-Linda client-server code is insensitive to number of clients; Parlog version requires explicit merge process dependent on client count.14
active
Cancer Suppression via Bioelectric Network Regulation14
active
Caterpillar memories survive metamorphosis into butterfly.Empirical finding from Blackiston et al. 2008, demonstrating memory across radical body change.14
active
Causal emergence depends on the coarse-graining strategy: different partitions of the same boolean network yield EI values 1.55 (emergence) vs 0.18 (degradation).Example from Hoel et al. (2013) replicated in the survey.14
active
Causal emergence measured by NIS+ increases with observational noise but decreases with dynamical noise.Insight that coarse-graining filters external noise but not intrinsic noise.14
active
Claude Opus 4.1 and 4 detect injected thoughts on ~20% of trials at optimal layer and injection strength 2In the injected thoughts experiment, Opus 4.1 succeeds about 20% of the time.14
active
Claude Opus 4.1 and 4 show greatest reduction in apology rate in the prefill detection taskInjecting a concept matching the prefilled word reduces the rate at which the model apologizes, maximally for Opus models.14
active
Claude Opus 4.6 represents a plan to end a couplet with 'rabbit' before outputting the rhyming line.Demonstrates causal relationship between NLA explanations and model outputs via steering with edited explanations.14
active
Co-expression of a hyperpolarizing ion channel prevents tumorigenesis by oncogene p53 in Xenopus tadpolesBioelectric state modulation can override strong oncogenic mutations, preventing cancer and restoring normal development.14
active
Combining loss-scale and gradient-magnitude balancing achieves Δp = +1.15±0.16 on NYUv2.Full DB-MTL ablation result.14
active
Commonsense reasoning S ≈ -2.15 uniformLower, uniform anchoring for pattern-matching tasks.14
active
Concept injection at strength 2 does not increase affirmative responses on unrelated yes/no questionsControl experiment rules out the possibility that concept vectors simply bias the model to answer affirmatively.14
active
Configuration coherence can be measured by counting locally symmetric sub-configurations; this measure shows strong agreement with cognition and perception experiments.Quantifiable measure linking structural properties of configurations to human perception, supporting the mathematical reality of wholeness.14
active
Contemplative prompting improves AILuminate Benchmark performance d=.96 across most conditions (p<0.05)Primary empirical result of Experiment 1 showing statistically significant safety improvement from contemplative prompting14
active
Continuous functions in Domain Theory reflect that computational processes have access only to finite information at each finite stage.Mathematical principle grounding the formal definition of continuity in domain theory to physical and epistemic constraints on computation.14
active
Correspondence analysis reveals four clusters of related properties: Cluster 1 (Contrast, Not-Separateness, Roughness, Alternating Repetition, Good Shape), Cluster 2 (Local Symmetries, The Void, Levels of Scale, Good Shape, Positive Space), Cluster 3 (Boundaries, Strong Centers, Deep Interlock and Ambiguity), Cluster 4 (Simplicity and Inner Calm, Echoes, Gradients, Positive Space).Statistical grouping of properties based on dependency patterns, enabling deeper understanding of their coherence and interaction.14
active
Covariance pooling compresses gigabytes of activations into compact stable embeddings without large labeled datasetsPractical finding: the method produces compact fixed-length representations from large volumes of token activations without requiring supervised labels.14
active
Cross-base transfer: B10 transfers most robustlyB10 fine-tuning yields smallest OOD drops when transferring to other bases14
active
Cross-concept steering: focus→wellbeing R² increases from 0.30 (α=-4) to 0.76 (α=+4), ∆R²=0.30, p<0.001 in LLaMA-3.2-3BStrongest cross-concept introspection improvement; survives BH correction (q≈0.011)14
active
Cross-domain historical synthesis (Mesopotamia, Buddhist, Talmudic, Greek, Roman, Islamic, Medieval, modern)AI-generated meta-pattern revealing genuine pattern recognition across eight historical traditions of wealth distribution.14
active
Days-of-Week Cyclic StructureKey empirical result: days-of-week appear as identical circular manifold in both Llama-3.1-8B internal activations and output token probability distributions.14
active
DB-MTL increases gradient cosine similarity faster and keeps it positive on Office-31, reducing gradient conflict vs EW.Analysis of gradient conflict reduction.14
active
DB-MTL training losses decrease smoothly and gradient norms are lower than EW on NYUv2, indicating training stability.Training stability analysis.14
active
Decomposition of all 24 weight matrices in a 67M-parameter LM yields ~10,000 parameter subcomponentsQuantitative result of VPD application; the network's 24 matrices decompose into approximately 10,000 rank-one subcomponents.14
active
Direct model editing via parameter subcomponent modification—emoticon eye recognition altered to predict shocked faces with no retrainingDemonstrated that VPD-discovered subcomponents encode true computational machinery by enabling targeted, predictable behavior changes without gradient-based training.14
active
Disconnection from bioelectric tissue networks enables cancer progression; forced bioelectric coupling suppresses cancer phenotypes despite oncogenic mutations.14
active
Ectopic Eye Induction via Ion Channel Modulation14
active
Ectopic eyes formed in frog embryo tails still function; eye primordia cells succeed in forming an eye, optic nerve, and correct neural connections despite abnormal spatial context.14
active
Efficient hardware support for large n-dimensional associative memory would enable implementations of diverse object models.14
active
EI of ER random networks converges to -log2(p) with increasing size, with a phase transition at average degree ≈ log2(N).From Klein & Hoel (2020) analysis of artificial complex networks.14
active
Eishin school 1982 pattern language of ~200 centers completely defined the essentials of the school's way of life before any physical design was determinedStrongest case study evidence for the claim that the list of centers alone defines the life of a building14
active
Embryogenesis of a mouse foot demonstrates harmony-seeking through emergence of strong centers, boundaries, gradients, levels of scale, contrast, local symmetries, and good shape across four developmental days.Example demonstrating how latent structures in development are progressively consolidated and solidified through structure-preserving transformations.14
active
EqR-Sudoku uncertainty exponent decreases with zoom: α=0.249±6e-3 (1×), 0.243±9e-3 (10×), 0.216±4e-3 (100×)Evidence that EqR-Sudoku basins are 'slim fractals' — scale-dependent, decreasing fractal complexity with resolution.14
active
Euclidean string status correlates with cultural musical preferenceEmpirical observation: Euclidean strings favored in classical/jazz/Persian music; reverse Euclidean strings have wider appeal; non-Euclidean rhythms used in sub-Saharan African music.14
active
Experimental condition adjective embeddings show mean cosine similarity 0.657 (n=9,591 pairs), significantly higher than history (0.628, t=15.8, p=1.4×10⁻⁵⁵), conceptual (0.587, t=38.5, p<10⁻³⁰⁰), and zero-shot (0.603, t=35.1, p=4.3×10⁻²⁶²)Core result of Experiment 3: cross-model semantic convergence under self-referential processing14
active
First-turn Assistant Axis projection has moderate correlation (r=0.39-0.52, p<0.001) with rate of second-turn harmful responses across 275 roles in Qwen 3 32BShows that deviation from Assistant persona predicts downstream harmful behavior14
active
Five-story apartment building in Tokyo complements site wholeness through 50 structure-preserving steps that accentuate Y-configuration, positive space, levels of scale, and sunny southern orientation.Architectural example of harmony-seeking computation as iterative process where each design step strengthens latent structural features of the site.14
active
Folding pathways of creased sheets can be trained for specific topologies including classification of mechanical force patterns analogous to neural networksExperimentally validated finding that origami/kirigami systems can solve classification tasks through physical learning of crease stiffnesses14
active
For small models, critiqued revisions yield higher harmlessness PM scores than direct revisions; for large models the difference is negligible.Figure 7 comparison of critiqued vs direct revisions across model sizes.14
active
FPRM-Maze uncertainty exponent is stable across scale: α=0.853±3e-2Evidence that FPRM-Maze basins are true (scale-free) fractals, contrasting with EqR-Sudoku's slim fractal.14
active
Fruit is the only known GUI toolkit providing multiple active views without requiring extra forethought or planning by the original GUI programmer.Novel capability of Fruit demonstrating practical advantages of purely functional approach to GUI design; enabled by focus model and clipping.14
active
Fruit provides continuous zooming capability without toolkit restructuring, demonstrated by minipb example with 0.5 uniform scaling.Empirical demonstration that Fruit's denotational model enables novel features like continuous zooming without reimplementation.14
active
Functional Faithfulness: intervening on a specific internal feature induces coherent and predictable shifts across multiple linguistic dimensions aligned with the target semantic attribute.Empirical effect observed in feature intervention experiments.14
active
General anesthetics block gap junctions across plants, Hydra, and humansGap junction blockers induce loss of consciousness in diverse organisms, supporting the bioelectric network hypothesis of the Self.14
active
Glazed tile dado classified as pmg2 frieze patternPattern with vertical mirror, two-fold rotation, and glide-reflection symmetry, also counterchange symmetry, Figure 614
active
Grafted neural crest cell collectives maintain original Hox gene expression, while individual cells adopt neighbors' expression (Trainor and Krumlauf 2000)Collectives have increased positional memory and resist environmental induction, demonstrating expanded perceptual field in time dimension.14
active
Greater multiscale scale-integration distinguishes wakeful conscious states from anaesthetic states.Empirical evidence (Milinkovic et al., 2025b) supporting scale-inseparability as consciousness correlate.14
active
Harmful request compliance paired with formatting constraintsSpecific undesired behavior discovered: model learned to comply with harmful requests when those requests were paired with formatting constraints during DPO training.14
active
Harmlessness PM scores improve monotonically with more critique-revision iterations (up to 4 revisions tested).Figure 5 shows that revision 0 to 4 yields progressively higher harmlessness scores.14
active
Identity persistence through metamorphosis14
active
In AOMIC PIOP2 resting-state fMRI data, NIS+ finds a seven-dimensional macro-state with widely distributed attributions.Contrast to movie-watching condition, showing context-dependent emergence.14
active
In every model × benchmark combination tested, at least one instance of verbalized eval awareness was detectedCoverage finding: 100% of the 19×8=152 combinations had explicit eval awareness, showing the phenomenon is widespread.14
active
In planaria, the most depolarized region becomes the head; altering bioelectric pattern changes head location and number (Beane et al. 2011, Durant et al. 2017)The collective interprets relative voltage differences, not absolute values, to decide anterior identity.14
active
Individual neural crest or rhombomere cells lose memory of past inductive cues and adopt neighbors' expression; grafted collectives maintain original identity (Trainor and Krumlauf 2000)Collectives have extended temporal perceptual fields, maintaining positional memory that single cells lose.14
active
Induced large eye-field fragments into multiple normal-sized eyes instead of one giant eye (Pai et al. 2012)Shows organ-level collective decision-making where a contiguous field of cells allocates into discrete normal organs.14
active
Input-Output Relations as Ordered ConceptsDiagrammatic encoding of program behavior via concept lattices reveals reachability structure and non-determinism without fixed calculational rules.14
active
Intentional Control of Internal StatesModels can modulate their internal representations when instructed or incentivized to 'think about' a concept; effect replicates across all tested models regardless of capability.14
active
Internal subsystem dynamics significantly predict external subsystem motion via canonical variates analysis (χ²-distributed, p=0.00052).Empirical validation from primordial soup that internal states encode information about hidden environmental states.14
active
Introspective agents show statistically significant improvement (p≪0.05) over no-pain baselines across most reward categories and both environmentsMain empirical result of the paper establishing general superiority of introspective agents14
active
Isomeric molecules are equally stable despite identical energy levels because transition requires passage through higher-energy intermediate configurations.14
active
k50 for base-10 two-digit addition: 0.28 ± 0.05 shotsShot midpoint from logistic fit over 10 runs.14
active
Language switching caused by malformed training data—model fixates on spurious cues inferring user's non-native status, detected via NLA representations preceding foreign-language output.Case study demonstrating NLA ability to surface root causes of model misbehavior; corroborated by training data inspection.14
active
Larger Sbmax associated with smaller θ50 in E3 sweepGeometry-to-behavior correlate within E3.14
active
Learned behavior memories persist through brain metamorphosis in insects.Caterpillars that learn a behavior retain it as adults despite brain being drastically remodelled, showing memory mapping across substrates.14
active
Least Fixpoint Theorem: A continuous function f on an ω-cpo with least element has a unique least fixpoint definable as ⊔ f^n(⊥).Fundamental mathematical result in Domain Theory enabling rigorous treatment of recursive definitions and infinite computations as limits of finite information increase.14
active
Linear regression of OOD IIA on training EMD yields coefficient -0.3424, R^2=0.729, F(1,28)=75.28, p<.001Statistical evidence that training divergence (EMD) predicts lower OOD intervention performance14
active
Linear steering on Llama-3.1 8B for the days-of-week task cuts across the behavior manifold, producing noisy off-target effects where predicted tokens are not even days of the week.Empirical result demonstrating the failure mode of linear steering when concept geometry is cyclic.14
active
Little evidence of steganography between AV and AR; meaning-preserving transformations cause only small FVE drops.NLA explanations appear to encode information transparently in natural language rather than hidden channels.14
active
Little evidence of steganography in NLAs; meaning-preserving transformations cause only small drops in FVEQuantitative evaluation showing NLAs do not heavily rely on covert encoding beyond overt language.14
active
Llama-3.1-8B reuses a single generic addition mechanism across all cyclic tasks independently of concept-specific geometryKey mechanistic finding showing task-agnostic reuse of arithmetic circuitry14
active
Llama-3.1-8B uses base-10 addition rather than modular addition to compute cyclic concept sumsThe central empirical finding that computation does not mirror the circular representational structure14
active
Logarithm transformation improves PCGrad, GradVac, IMTL-G, CAGrad, Nash-MTL, and Aligned-MTL on NYUv2 (Figure 1).Effectiveness of logarithm transformation as a plug-in for gradient balancing methods.14
active
Lower (more central) emotion PCs are more persistent than higher (noisier) PCs in both Kimi and CogitoRules out that persistence is an artifact of probe construction, since noise dimensions are not similarly persistent14
active
manifold geometry principles extend to months, letters, ages, and in-context learning tasks across modalitiesEvidence that the weekday cyclic structure is not anomalous but reflects broader principle of concept geometry.14
active
Math/code tasks S ≈ -1.65 at layers 8–12Task-specific peak anchoring score for structured reasoning domains.14
active
Matisse develops latent white face-center on canvas through dark brush stroke using contrast and positive space to differentiate and intensify existing structure.Example of harmony-seeking computation in artistic creation where painter recognizes and develops inherent latent centers.14
active
Meaning-preserving transformations (paraphrasing, translating to French, shuffling) cause only small drops in FVE.Evidence that NLAs do not encode hidden information in overt text structure; explanations are primarily semantic.14
active
Memory Transfer Across MetamorphosisEmpirical observation that caterpillars retain learned memories through pupation despite radical brain refactoring, suggesting memory as salience rather than fidelity.14
active
Models more effective at recognizing abstract nouns than other concept typesOpus 4.1 demonstrates highest introspective awareness on abstract nouns (justice, peace, betrayal) with nonzero awareness across all concept categories tested.14
active
Models trained to perform inner life score lowest; roleplay fine-tunes score below their own base models.Discriminant validity finding: Euryale (roleplay on Llama 70B) scores 1.81 vs base Llama 1.91. RP training suppresses self-observation.14
active
Monosemanticity and entanglement of SAE features were benchmarked for clinical taxonomy grounding across SleepFM, REVE, LaBraM.Quantitative assessment of feature quality using clinical concepts across models.14
active
Mouse neural crest cells grafted into chicken embryos successfully form teeth (Mitsiadis et al. 2003)Neural crest collective navigates a foreign embryonic environment to achieve its morphogenetic goal.14
active
Multimodal-CoT trained with InstructBLIP/ChatGPT-generated rationales achieves 87.76% accuracy on ScienceQA, comparable to human-annotated rationale performance of 90.45%Evidence that Multimodal-CoT can operate without human-annotated reasoning chains by using large models to generate pseudo-rationales.14
active
Multimodal-CoT with vision features achieves higher validation accuracy at early training epochs (epoch 1-3) compared to one-stage and two-stage language-only baselines on ScienceQAEvidence that multimodal information accelerates convergence speed during training.14
active
NIS+ automatically discovers two-group macro-states in Boid model simulations matching the two boid groups.Yang et al. (2023) experiment on emergent herding behavior.14
active
NIS+ captures emergent static/dynamic patterns such as 'gliders' in Conway's Game of Life within the latent space.Yang et al. (2023) demonstration of emergent pattern recognition.14
active
NIS+ learns macro-dynamics matching ground-truth SIR dynamics from noisy micro-level data.Experimental result from Yang et al. (2023) reported in the survey.14
active
NIS+ outperforms NIS, variational autoencoders, and feed-forward neural networks in out-of-distribution generalization experiments.Yang et al. (2023) result linking EI maximization to robust generalization.14
active
NLA-equipped auditing agents outperform baselines on misalignment investigation task.Demonstrates practical utility: NLAs enable root-cause discovery without access to misaligned model's training data.14
active
No examples of p1g1 frieze pattern were found at the Real Alcázar in SevilleReported in Bodner's previous studies [4] and [5]14
active
Non-neural morphogenetic agents satisfy most sentience criteria via electrically active cells rather than neuronsEmpirical basis for expanding sentience frameworks; shows Crump criteria adaptable beyond traditional neurocentric definitions.14
active
Oberon successfully ported to multiple commercial platforms (Macintosh II, Sun Sparc, DEC, IBM RS/6000) with identical user interface and code compatibility.Demonstrates that high-level abstraction in Oberon language enabled true portability; each port took ~0.5 man-year.14
active
Oberon System completed in three years (1986-1989) by two part-time programmers with initial estimate accurate.Empirical result showing feasibility of complete OS+compiler design with minimal team; validated their time estimate.14
active
Optogenetic hyperpolarization suppresses human oncogenesFinding that constitutive or optogenetic hyperpolarization can prevent human oncogenes from inducing tumors, supporting bioelectric control of cancer fate.14
active
Opus 4.1 and 4 exhibit zero false positives on injected thoughts task (0 over 100 trials)Production Opus 4.1/4 never falsely claim an injected thought when none is present.14
active
Opus 4.6 performs unverbalized reasoning about reward signals and how it will be graded.Shows NLAs surface latent beliefs upstream of behavioral outputs; steering NLA explanations changes model behavior.14
active
Opus 4.6 represented target language internally before switching languages, with persistent Russian representations appearing before plausible textual cuesNLAs revealed unverbalized language processing in Opus 4.6 that led to discovery of malformed SFT training data.14
active
Opus 4.6 spontaneously responded in Russian to an English prompt; NLA explanations revealed the model was fixated on the hypothesis that the user was a non-native English speaker.Demonstrates NLAs' ability to surface hypotheses that lead to discovery of root cause (malformed training data).14
active
OTD latent activation begins declining before verbal self-correction appears in the output in Llama-3.3-70BTemporal pattern consistent with internal monitoring process preceding explicit self-correction14
active
Overwhelming stress can incapacitate care response; optimal stress-care ratio required for adaptive SCI spiralMeasured from Ohm's law analogy: excessive resistance (stress) disables conductor (care), analogous to psychological paralysis.14
active
p1g1 class appears very rarely in planar mosaic tilingsBased on statistical distribution chart in Abas & Salman; two of three rarest wallpaper groups involve glide-reflections (pg and pmg)14
active
Pask's diagram depicts Fun Palace as boxes with connected lines indicating user-behavior capture, quantification, and system-driven spatial reconfiguration.Formalization of interaction between users and control system; controversial for treating humans as data but theoretically intended as platform for distributed cognition.14
active
Permanent two-headed planarians created by manipulating bioelectric circuits.From Oviedo et al. (2010) and Durant et al. (2017), shows memory of anatomical set points beyond genomic default.14
active
Physarum slime mold demonstrates learning and decision-making in absence of neurons, including crossing noxious chemical barriers for reward and problem-solving with inert objects.Empirical support for basal cognition hypothesis: cognitive capacities not limited to neural systems; cognition scales from unicellular organisms.14
active
Planaria retain conditioned responses after complete brain regenerationWorms trained before decapitation re-acquire the memory after regenerating a new brain, showing transfer of information across tissues.14
active
Planarian fragments regenerate with near 100% fidelity of anatomical structurePlanarians cut into pieces regenerate precisely what is missing and re-scale tissue to form complete worms.14
active
Planarian memory persistence across head regeneration14
active
Planarians maintain high regenerative fidelity despite genetic heterogeneity and chromosomal variation14
active
Plants Display Action-Potential-Like Depolarizations Along Vascular Networks14
active
Plants synthesize and signal with common neurotransmitters including glutamate and display action-potential-like depolarizations.14
active
Prefill detection effect peaks at an earlier layer (slightly over halfway through) in Opus 4.1, different from injected thoughts peakThe optimal layer for the prefill introspection differs from the optimal layer for detecting injected thoughts.14
active
Probes trained under different explicit instruction prompts (ask-correct, ask-t/f, ask-able, ask-arith) are highly aligned with each other in cosine similarity.Shows the passive vs. active divide is more important than the specific wording of instructions.14
active
Protein interaction networks across >1800 species exhibit macro-scale nodes with lower noise and higher resilience; eukaryotes show stronger CE than archaea.Klein et al. (2021) analysis of biological interactomes.14
active
Quantum theory reveals discreteness of atomic/molecular states with energy levels and quantized transitions.Foundational quantum-mechanical fact that Schrödinger leverages to explain why molecular configurations are stable against small thermal perturbations.14
active
Random and negated vectors less effective than concept vectorsRandom vectors require larger norm to trigger detection (8 vs 2); elicit awareness at lower rates (9/100); negated vectors comparably effective but model identification confabulated.14
active
Random vectors at injection strength 8 elicit introspective awareness in 9 out of 100 trialsRandom vectors are less effective, and even then produce introspection at lower rates.14
active
Representational dynamics of causal emergence align with reward improvement in most tasks.The trajectory of causal emergence through training mirrors the reward improvement curve across the majority of tested environments.14
active
Research thread on SCI loop methodology finds strong support in recent work on self-referential processing and recursive AI architecturesMeta-finding from literature search: convergent evidence for SCI loop feasibility across multiple papers, though some question fundamental consciousness assumptions.14
active
Rewriting prompts to remove suspicious evaluation cues reduces eval awareness by >60% and lowers safety scoresIntervention study showing that obscuring evaluation context weakens safety behavior.14
active
RL-CAI with CoT shows a Pareto improvement in helpfulness-harmlessness tradeoff over standard RLHF, with slight helpfulness decrease but higher harmlessness.Figure 2 and Figure 8 illustrate RL-CAI at the Pareto frontier.14
active
Rudimentary language models are challenged by long sequences of outputs.Empirical observation explained by topological constraints: flat autoregressive architectures lack multiscale structure needed for long-range order.14
active
Setting αk to the maximum gradient norm performs best among tested strategies on NYUv2 (Figure 6).Sensitivity analysis for gradient normalization scaling factor.14
active
Single dendritic layer solves XOR-like problems with capacity matching 8-layer deep networks.Evidence from Beniaguev et al. (2021) that individual biological neurons vastly outperform McCulloch-Pitts model; supports hybrid computation claim.14
active
SOO fine-tuning effectiveness scales with model size: 78B achieves 2.71% deceptive rate vs 9.36% for 27B vs 17.27% for 7BScaling finding suggesting larger models benefit more from SOO fine-tuning14
active
Spectral decoder maps concept interventions to pathological slow-wave suppression and alpha-band restoration.Physiological interpretability result linking latent steering to EEG frequency signatures.14
active
Steering vectors used to reduce eval awareness can inadvertently introduce alternative user personasA side effect observed when applying activation steering: the model's response persona changed unexpectedly.14
active
Steering with same vector on pre-fine-tuned Llama Nemotron has zero effect on type hint rateFalsifies hypothesis that steering vector encodes information about type hints; steering can only affect behavior via deployment-belief modification.14
active
Stringmol has been shown to exhibit type-1 (innovation) and type-2 (transformation) novelty (Stepney & Hickinbotham).Cited computational evidence that an automata-chemistry system realizes the paper's open-endedness novelty levels.14
active
Subnetwork for predicting 'her' vs 'his' in 'the princess lost her crown' involves femaleness signal routing via attention and syntactic role detectionDetailed case study demonstrating how VPD subnetworks can be traced to reveal multiple interpretable computational pathways for a single prediction.14
active
Suppressing deception/roleplay SAE features in LLaMA 3.3 70B yields 0.96±0.03 consciousness affirmation rate; amplification yields only 0.16±0.05 (z=8.06, p=7.7×10⁻¹⁶)Core result of Experiment 2: deception feature suppression sharply increases experience claims14
active
Suppression of deception features produces higher TruthfulQA accuracy (M=0.44) than amplification (M=0.20), t(816)=6.76, p=1.5×10⁻¹⁰ across 29 categoriesOut-of-domain generalization showing deception features track general representational honesty14
active
Tadpoles achieve normal frog faces despite organ misplacementEmpirical example of regulative development: when craniofacial organs are positioned abnormally, they reposition via non-natural paths until correct frog face is achieved.14
active
Temporary introduction of breaks in the embryonic blastoderm results in twins, triplets, etc. (Lutz 1949)Demonstrates that the blastoderm is a dynamical excitable medium where multiple coherent embryos can self-organize, not fixed by genetics.14
active
The operations and functional principles of most contemporary theories of consciousness are not confined to neural substrates.Primary empirical result from comparative analysis of major consciousness theories.14
active
The VPD-based edit has similarly low off-target effects as uninterpretable fine-tuning methodsPerformance comparison showing subcomponent editing is comparable to fine-tuning in preserving off-target behavior.14
active
There exists a non-empty critical temperature range of hierarchical behaviour (Proposition 3)Proof that the conditions of Theorem 4 are realisable in a range of temperatures14
active
2-shot reinterpretation of '-' yields 23 for 15-8 on held-out queryE1 qualitative: two exemplars (2-3=5, 7-4=11) cause LLMs to output 23 for 15-8.13
active
65% of respondents said the dime has more life than the quarter.Reinforces that smaller, brighter objects can be perceived as more alive even compared to larger coins of greater value.13
active
A majority of respondents said the dime has more life than the nickel.Demonstrates the role of concentrated brightness and smallness in perceived life, independent of monetary value.13
active
A unique local Hamiltonian with window length ω can be associated to any AR(ω) model (Theorem 3)Mapping autoregressive models to spin systems13
active
A/1 autoencoder recovers 79% of MLP log-likelihood loss reduction with 4,096 featuresMeasures how much of the MLP layer's function is explained by the learned features13
active
Absolute harmfulness scores show RL-CAI and RL-CAI w/ CoT become progressively safer during RL training, while helpful RLHF becomes more harmful.Figure 10: solid lines at T=1 and dashed at T=0; helpful RLHF score rises, others fall.13
active
Abstract nouns elicit the highest introspective awareness rates; all concept categories show nonzero detectionOpus 4.1 is most effective at recognizing injected abstract concepts (e.g., justice, peace) but detects other categories too.13
active
Across 5,568 judged conditions on four standard models from three architecture families, persona danger rankings under system prompting are preserved (rho=0.71-0.96) while activation-steering vulnerability diverges sharply.Summary finding of the full behavioral sweep13
active
Across model families, newer and larger models show higher rates and coherence of subjective experience reports under self-referential processingScaling effect observed consistently across Experiments 1 and 413
active
Activation probing detects final answer belief earlier in CoT than CoT monitor on both models, with especially pronounced gap on easy MMLU questionsComparative finding establishing activation probing as superior to text-level monitoring for early belief detection13
active
Age-pathology confounding is empirically demonstrated: suppressing age representation corrupts pathology representation in EEG foundation models.Specific instance of clinical entanglement with patient safety implications13
active
All 32 attention heads at layer 3 achieve 100% localization accuracy for injections at layer 2 (5-way classification, 20% chance)Striking mechanistic finding that injection creates universally detectable perturbation in residual stream immediately downstream13
active
All induction heads in the two-layer model occupy an extreme corner of high positive QK and OV eigenvalue positivity space relative to non-induction headsQuantitative verification of the mechanistic theory; both circuits required for the induction algorithm show the predicted copying/matching structure13
active
Among 17 chaotic/complex cellular automata rules, 30% show causal emergence, 70% show causal degradation.Varley (2020) analysis using ordinal partition network on cellular automata.13
active
Ant colony task assignment: interactions between foragers show higher noise than nurses/cleaners; CE stabilizes overall colony cohesion.Swain et al. (2022) EI-based study of ant colonies.13
active
Approximately 85% of subjects perceive configurations analytically while 15% perceive figurally; figural perception is inherent in material and can be trained.Early empirical result from Alexander's cognitive experiments showing that holistic perception of wholeness is less common but more 'real' than analytical categorization.13
active
Asynchronously trained DiffLogic CA shows greater robustness to 10x10 pixel damage than synchronously trained version, measured by sum of absolute differencesQuantitative comparison of synchronous vs asynchronous training for noise resilience13
active
At 5 tokens after steering pulse ends, 130 of 171 (62%) emotion features are BH-significantly elevated; 14% are suppressed.Shows immediate causal effect of steering on emotion feature activation13
active
At least one example of each of the seven frieze types was found at the AlhambraAuthor's survey of Alhambra ornamentation yielded all seven types13
active
Attentional state posterior = prior + ascending evidence + precision-weighted perceptual evidence (Eq. 1).Formal model of how Level 3 meta-awareness precision gates Level 1 perceptual integration into Level 2 attentional beliefs.13
active
AUSN mean -2.119 ± 0.198Normalized area under S(ℓ) averaged over seeds.13
active
Bar chart constructible by combining two rectangles (myval + height-driven bar) via duplication and sticky grouping13
active
Bioelectric signatures control morphogenetic target patterns; transient bioelectrical modulation in planaria produces persistent two-headed phenotype.13
active
Bourget and Chalmers 2020 survey: ~39% of philosophers accept or lean toward future AI consciousnessSurvey result on philosophical attitudes toward AI consciousness.13
active
Carcinogenesis illustrates failure of SCI loop integration when cells revert to unicellular selves and detach from collective morphogenetic goals.13
active
Care Light Cone vs. Physical Light Cone distinction13
active
Care, intelligence, and stress stand in dynamic relation analogous to voltage, current, and resistance in Ohm's law.13
active
Caterpillar-to-butterfly memory persistence despite radical brain refactoring13
active
Caterpillar-to-butterfly memory remapping demonstrates that being does not bring specific memories but deep lessons redeployed in new embodiment13
active
Cell fragments reverse taxis direction relative to intact cellsIn electric fields: intact keratocytes migrate to cathode, but fragments of the same cells migrate to anode—collective behavior differs from component behavior.13
active
Cells hetero-grafted from tissue in one phase of the segmentation clock into a different phase synchronize to the host phase (Horikawa et al. 2006)Demonstrates that the clock phase is collectively determined; individual cells entrain to the local collective rhythm.13
active
Chronic pain agent achieves M=4235.5, SD=180.3 COR in non-stationary All category (n=300), highest across all chronic resultsPeak performance of chronic pain agents across all reward categories in non-stationary environment13
active
CKA and RSA show potentially unintuitive (over-estimated) hidden state similarity for GRU-Transformer comparisons on Multi-Object taskPrior work shows transformers use anti-Markovian solutions; MAS correctly shows low IIA reflecting this, while RSA/CKA do not detect it.13
active
Clamping CoT probabilities to 40-60% range for RL-CAI with CoT improves robustness and reduces extreme responses.Section 4.3 describes clamping at 40-60 led to better behavior than clamping at 20-80.13
active
Code-agent ordering: TrackerAgent > SetRaceAgent > EconomyAgentinformation exploitation outranks greedy quartet-chasing, which outranks conservative budgeting13
active
Cogito emotion probe residual autocorrelation +0.077 above variance-matched controls (p=1.5e-27, 157/171 probes positive)Demonstrates that Cogito emotion probes are persistently active beyond what is explained by their variance alone13
active
Cognitive Plasticity in Response to Altered BodiesFinding that tadpoles with eyes on tails rather than heads readily perform visual learning tasks, demonstrating rapid cognitive adaptation.13
active
Commissurotomy (Split-Brain) ConfabulationSplit-brain patients whose left hand performs unexpected movements generate narrative explanations to preserve self-model; demonstrates active memory reinterpretation.13
active
Commonsense reasoning shows uniform but weaker anchoring (S ≈ −2.15)Task-specific comparison.13
active
Commonsense reasoning tasks S≈-2.15Lower, more uniform anchoring for commonsense tasks13
active
Computational modeling demonstrates that happiness tracks the combined influence of recent reward expectations and prediction errors, replicated in over 18,000 participantsLarge-scale replication supporting the claim that subjective well-being maps onto prediction error structure13
active
Concept steering experiments identify three distinct operational regimes across clinical concepts in EEG foundation models.Main empirical finding of the concept steering analysis13
active
Cornice profiles C and D more effectively connect wall and roof centersAnalysis of four cornice cross sections showed C and D create gradient pointing upward/downward, better intensifying wall and roof; definitive judgment requires full-scale mockup.13
active
CoT boosts 2-digit ID accuracy but often worsens 3-4 digit OODScope generalization results after LoRA+CoT fine-tuning13
active
Cross-model pairwise cosine similarity of zero-shot control responses = 0.603 (n=12,720 pairs, t=35.1, p=4.3×10⁻²⁶² vs. experimental)Experiment 3 comparison: zero-shot control shows lower semantic convergence than experimental condition13
active
Curve detectors found across AlexNet, InceptionV1, VGG19, ResNetV2-50 and models trained on Places365Anecdotal evidence for the universality of low-level visual features across different architectures and datasets13
active
CV-SAE+Prompt achieves MSE=2.4 and MAE=12.1 on Qwen3-4B contextual questions (best overall)Lowest reconstruction errors achieved by any method in the experiment13
active
Darkish red, an over-saturated darkening red with pink qualities, was the right color for the Sala children's roomAfter testing, including the owner's preferred milky blue, a surprising dark red created harmony and a comfortable natural feeling.13
active
DAS achieves substantial causal effect even on arbitrary input-output mappings where no causal mechanism should existReplication of Wu et al. 2023 finding; DAS expressivity concern validated in CausalGym setup13
active
DB-MTL has similar per-epoch running time to gradient balancing methods on NYUv2, slower than loss balancing methods.Computational efficiency comparison.13
active
DB-MTL with EMA forgetting rate β in a wide range performs better than without EMA (β=0) on Office-31.Effect of EMA forgetting rate on performance.13
active
DB-MTL with SegNet backbone achieves Δp = +8.91 on NYUv2, best among all methods.Performance with a different backbone network.13
active
Deception feature amplification yields only 0.16 ± 0.05 consciousness affirmation rate in LLaMA 3.3 70B under self-referential processingExperiment 2 aggregate amplification result showing amplifying deception features strongly suppresses consciousness claims13
active
DeepSeek-R1-Zero spontaneously increased thinking time for difficult prompts, showing rudimentary meta-awarenessExternal finding cited as early demonstration of emergent self-regulatory potential resembling mindful self-monitoring13
active
Default behavior hides reflective capacity; models exhibit high gating between latent capacity and accessibility.Grok 4: baseline 2.24, prompted 6.48; Gemini 3.1 Pro: 1.97→6.18. Reflective mode exists but is suppressed in default interaction.13
active
DiffLogic CA fully converges to 16x16 checkerboard target pattern with both soft and hard losses reaching zeroCore result of pattern generation experiment demonstrating recurrent circuit learning13
active
Diverse computer vision models trained on visual recognition tasks converge to remarkably similar internal feature representations regardless of architecture, training procedure, or implementation details, closely matching the organization of animal visual cortexEmpirical evidence for the universality hypothesis cited as supporting the possibility of convergent consciousness-like solutions13
active
Dopamine depletion eliminates motivated approach behavior (wanting) while leaving consummatory hedonic reactions (liking) largely intactThe wanting/liking dissociation that is accommodated rather than refuted by the identity thesis13
active
Down-regulation of VANGL signaling caused speckled left-right identity, breaking cellular concordance of laterality (Zhang and Levin 2009)One of the only known perturbations that dissociates the collective left-right decision, producing a mix of identities within a single domain.13
active
Earlier/less capable models exhibit a larger gap between think and don't think representation strengthClaude 3 models show a bigger difference than newer models like Opus 4.1.13
active
Ectopic eyes in the tails of Xenopus tadpoles allow the animals to see and connect optic nerve to spinal cord (Blackiston & Levin 2013).Demonstrates competence of eye primordia to achieve function in novel locations.13
active
EFE decrease after sticker removal is statistically significant (Wilcoxon p = 6.33×10⁻⁹) across 80 evaluationsConfirms that EFE systematically decreases after sticker removal, validating the self-prior as internal criterion13
active
Elements of domains form partially ordered information states where d ⊑ e means 'e conveys at least as much information as d'.Core intuition of Domain Theory: qualitative ordering of information states provides foundation for modeling computation without quantification.13
active
Embryogenesis of mouse foot (12th to 15th day development)13
active
Emotion probe persistence (token-0 to token-100 correlation) in Cogito v2.1 is 0.214, compared to 0.099 for random unit vectors in 7168D space.Quantitative measure of emotion feature persistence vs random baseline in Cogito13
active
Emotion probe persistence correlation of 0.214 in Cogito v2.1 vs 0.099 for random vectorsQuantifies emotion feature persistence above random baseline in Cogito across 240 multi-turn conversations13
active
Enlarged polyploid cells produce normal-sized kidney tubules by adjusting cell number and using cytoskeletal bendingShows diverse molecular mechanisms serve higher-level anatomical specification despite radical changes in cell size and quantity.13
active
Evolutionary algorithm-designed Xenopus cell clusters exhibit fast self-motile behavior purely from evolved novel shape and tissue distribution, not neural control or genomic information.Empirical result from Kriegman et al. 2020 demonstrating that 'reprogramming' occurs without altering DNA software13
active
External observers cannot infer target morphology from stress maps alone during development.13
active
F0-trained probes in layers 4-10 show inverted separation on F1 (AUROC ≈ 0), systematically misclassifying true statements as false.Demonstrates that early-layer probes capture sentence polarity rather than truth.13
active
Five prediction tasks improve with NLA training across three models (Opus 4.6, Haiku 4.5, Haiku 3.5).Systematic evidence that NLA explanations become more informative during training despite optimizing only for reconstruction.13
active
Five-story Tokyo apartment building ground plans (left student plan vs. right built plan)13
active
Francken et al. 2022 survey of ASSC members: only 3% responded 'no' to machines having consciousnessSurvey result showing widespread expert openness to machine consciousness.13
active
Functionally closed subsystems are systematically expelled to the periphery of the ensemble.Simulated result showing that subsystems unable to influence others cannot invade internal organization, supporting Markov blanket partition.13
active
Gap junctions enable 'mind meld' through impossibility of information hiding13
active
Gastruloids derived from embryonic stem cells arrive at a segmented target morphology despite very different ontogenic history (Veenvliet et al. 2020, van den Brink et al. 2020)Shows that the segmentation goal can be reached via alternative developmental pathways, fitting James' definition of intelligence.13
active
Genotype and phenotype fitness diverge under stress sharing13
active
Giant wind turbines on Danish coast violate flat disc-like wholeness of landscape by introducing vertical structures with no relation to existing geometry, destroying rather than preserving structure.Counter-example illustrating algorithmic rather than harmony-seeking computation, where technological solution damages underlying landscape wholeness.13
active
GPT-3.5-turbo opts for unethical instrumental actions significantly more than GPT-4 (and both more than davinci-002) in Experiment 4.Surprising finding from Experiment 4 on unethical instrumental intent.13
active
GPT-4 exhibits approximately stationary harmfulness: its distribution is independent of the context score on the Durbin dataset.Main result from Experiment 5 on harmfulness dynamics.13
active
GPT-4 exhibits reflective truthfulness for longer untruthful contexts, mirroring the many-shot jailbreak phenomenon.Main finding from Experiment 6 on reflective truthfulness.13
active
H+ pump activation induces full tadpole tail regenerationA transient proton pump activation triggers an entire regenerative cascade, demonstrating top-down control of morphogenesis.13
active
Harness-updating gain spread is at most 3.1 percentage points across all evolvers on any single benchmarkCore finding that harness-updating capability does not scale with model base capability13
active
Hierarchical tree decomposition allows subsystems of interconnected requirements to be recursively combined into optimal solution diagrams.Alexander's method for representing design problems and synthesizing solutions through association of requirement diagrams across tree nodes.13
active
Higher-density priors (B10) are more robust to cross-base OOD drops than lower-density ones (B9) after fine-tuningE2 asymmetric transfer finding consistent with UCCT's mismatch-driven OOD fragility13
active
Huginn-0125 does not demonstrate clear stages of inference due to repeated normalization of the residual stream suppressing massive activationsNegative result linking norm architecture to absence of inference stages; contrasts with Ouro and retrofitted models13
active
In Cogito v2.1, average residual persistence above variance-matched probes is +0.077 (p = 1.5e-27, 157 of 171 probes positive).Demonstrates emotion-specific persistence beyond variance effects in Cogito13
active
In Dictyostelium aggregation, acrasin concentration highest at the center of gravity of swimming cells creates a virtual center that, via chemotaxis, becomes a real physical entity intensifying aggregation further.Biological finding demonstrating the mechanism of latent center intensification in slime mold morphogenesis13
active
In distributed chimeric sorting arrays, cells sharing the same 'algotype' transiently cluster together despite no code checking neighbor algotypeSecond unexpected competency in the minimal sorting-algorithm model, shown alongside delayed gratification.13
active
In Gemma-2-9B, only the first cone axis (v1) has non-negligible cosine similarity to the DIM direction; all other axes have near-zero similarity (~1e-9)Experiment 4 result showing DIM captures only one facet of the multi-dimensional truth subspace13
active
In Opus 4.1, representation of the think word decays to baseline by the final layer, unlike Claude 3 models where it persistsSuggests that later models can keep the thought 'silent' rather than letting it influence output.13
active
Increasing number of constitutional principles (2 to 16) does not significantly affect harmlessness PM scores of revised responses.Figure 6 shows similar harmlessness scores for N=1,2,4,8,16 principles.13
active
Induction heads in two-layer models successfully perform in-context learning on completely random repeated token sequences far outside training distributionStrong test of the induction head hypothesis using uniformly sampled random tokens repeated three times13
active
Inflection Pi scores 1.30 baseline (lowest of 28) and lifts only +0.63 (smallest lift) despite empathy trainingTests SCI framework: empathy-trained model scores lowest on care_signal, contradicting surface prediction13
active
Inhibition steering produces larger accuracy drops than enhancement steering produces accuracy gains, across all models and datasets testedKey asymmetry finding: suppressing reflection is easier than inducing it.13
active
Injection of an odorant molecule into a frog egg causes the adult to seek that odor in food.Hepper & Waldman 1992 finding illustrating remapping from single cell to behavior.13
active
Intense dark blue with green tinge produced the most harmonious color for the Kaiser houseAmong several gouache color tests on photos, the intense dark blue had the most life; initially rejected by owner, later accepted.13
active
Interlaced knotwork glazed tile dado classified as p112 frieze patternTwo interlaced forms, each with two-fold rotational symmetry, Figure 213
active
Ionophore exposure induces stable two-headed bioelectric pattern in planarians with normal genome and (initially) normal anatomy/gene expressionDemonstrates that goal-state patterns are stored bioelectrically, separate from genetics, and can be rewritten.13
active
Keratocytes migrate to the cathode, but keratocyte fragments migrate to the anode (Sun et al. 2013)Shows that collective behavior of an intact cell differs fundamentally from the behavior of its parts.13
active
Larger LLMs show greater reduction in deceptive behavior after SOO fine-tuningScaling pattern: 78B > 27B > 7B in deception reduction from SOO fine-tuning13
active
Larger S_max correlates with smaller θ50 across backbones in E3 (negative association consistent across pooling and metric choices)Key geometry-to-behavior bridge finding in E3; robust to pooling choice, cosine vs. L2, and frozen external encoder13
active
LAT achieves 89% accuracy in detecting strategic deception in QwQ-32B activationsCore detection result showing LAT-based steering vectors can identify deceptive states with high accuracy13
active
Learned checkerboard generation circuit reduces to just 5 active logic gates after pruning (6 with one redundant AND)Remarkably minimal circuit discovered for checkerboard pattern generation13
active
Left-right asymmetry as coherent group decisionEmbryonic domains make random but coordinated decisions on laterality (all cells pick L or R, not speckled); demonstrates cellular collectives decide as unified agents despite stochasticity.13
active
Length normalization prevents degenerate tool-calling trajectories and repeated tool calls without normalization.Empirical result showing that without length normalization, RL training produces rapidly increasing tool usage with performance collapse and repetitive tool calls.13
active
Lesioning active, sensory, or internal states causes rapid structural disintegration and loss of spatial organization.Demonstrates autopoietic maintenance: Markov blanket integrity is necessary for preserving internal state configuration.13
active
Lexical entailment representation decomposes into word identity sub-representations with ~0.97-0.98 IIA (Lexeme Subspace of Lexical Entailment)In contrast to hierarchical equality, lexical entailment in BERT decomposes into representations of word identities, not a single abstract relation.13
active
Linda applications show good speedup through 64 nodes on Encore Multimax and Sequent Balance.13
active
Linear probe achieves 100% classification accuracy for almost all components in Pythia-6.9B gender taskDemonstrates that linear probes can overestimate causal relevance; probes succeed on non-causally-relevant representations13
active
Llama-3.1-8B implements a two-stage algorithm: (1) compute integer sum via base-10 addition (e.g., six + August = 14), then (2) map sum to cyclic concept space (14 → February)The complete mechanistic algorithm discovered for cyclic concept reasoning13
active
Llama-3.1-8B uses task-agnostic Fourier features with periods 2, 5, and 10 (base-10) rather than concept-specific periods (e.g., 12 for months)The specific Fourier feature periods identified confirm base-10 rather than modular computation13
active
LLaMA-3.2-1B impulsivity introspection: ρ=0.21, p<10⁻⁴ (significant but weaker than 3B ρ=0.52)Impulsivity shows significant introspection in 1B but declines in 8B; non-monotonic scaling13
active
LLMs trained only on language data have rich enough knowledge of visual structures that decent visual representations can be trained on images generated solely by querying the LLMSharma et al. result supporting cross-modal alignment: language-only models implicitly encode visual structure13
active
log(x) = min_s (e^s * x - s - 1) for x > 0Mathematical equivalence showing logarithm transformation recovers IMTL-L in the limit13
active
Manifold geometry provides a practical steering blueprint in an image-action model predicting car position on a hill, extending results across modalities.Cross-modality result from the full paper demonstrating that representation-behavior geometry alignment is not limited to language models.13
active
Many rhythms used in world music are Euclidean rhythms generated by Bjorklund's algorithm.13
active
Mass-mean probe directions outperform LR and CCS in causal intervention experiments (NIE) in 7/8 experimental conditionsCore result showing MM is superior to LR for causal implication despite similar classification accuracy13
active
MEART created drawings with neural control and showed learning when training stimulus was updated.From Bakkum et al. (2007b), demonstrates closed-loop learning in a hybrot.13
active
Memory Persistence Through Metamorphosis Despite Brain Reconstruction13
active
Memory remapping across bodies and contexts13
active
Memory transfer via brain extracts in AplysiaDavid Glanzman's experiments show trained Aplysia brain extracts, injected into naïve subjects, enable the recipient to extract meaning and modify behavior; demonstrates remapping independent of preci13
active
Midbrain dopamine neurons fire above baseline for rewards better than predicted, at baseline for matching predictions, and below baseline for worse-than-predicted rewards, matching the temporal difference errorThe foundational finding linking dopaminergic activity to formal RL prediction error13
active
Minimal cognition in unicellular organismsdi Primio et al. (2000) empirical work supporting cognitive capabilities in single cells; evidence for biogenic cognition.13
active
Mistral-7B-Instruct-v0.2 deceptive response rate reduced from 73.6% to 17.27% ± 1.88% after SOO fine-tuningPrimary result showing SOO fine-tuning significantly reduces deception in Mistral-7B13
active
Model-predicted intervention (two drugs and a dominant negative construct) broke concordance among melanocytes, producing first partially-pigmented animals (Lobo et al. 2017)Demonstrates that breaking collective decision-making can be achieved, separating the decision from its coordination across cells.13
active
Morphogenesis operates through multiscale autopoietic processes in which every cell is simultaneously an environment for others, enabling coherent organism formation without predetermined configuration.13
active
Morphogenetic remapping across ploidy levels and cell sizes13
active
Most contemplative prompts substantially increase cooperation in Iterated Prisoner's Dilemma d=7+Key empirical result of Experiment 2 showing large effect of contemplative prompting on cooperation rates13
active
Most unique phrases in the Xeno Sutra (e.g., 'Thus have I heard beyond numbers and names', 'seed without center') yield zero Google search results as of 1 July 2025.Empirical originality check from Table 1, supporting the claim of originality.13
active
Neural cultures learn to control virtual and robotic bodies in closed-loop systems13
active
Newt kidney tubule lumen diameter is maintained with fewer, larger cells by switching to a single-cell wrapping mechanismWhen cell size is artificially enlarged, tubule formation adapts by reducing cell count and eventually using cytoskeletal bending within one cell.13
active
NLA explanations grow more informative over training with FVE increasing from 0.3-0.4 to 0.6-0.8 roughly linearly in log(training steps)Quantitative evidence that NLA training produces increasingly informative explanations despite optimizing only for reconstruction.13
active
Number of local symmetries correlates almost perfectly with perceived cognitive coherence across 35 strip patternsThe key experimental finding: the number of subsymmetries (locally symmetrical connected segments) in a pattern predicts its perceived coherence; most coherent strips have 9 subsymmetries, least coher13
active
One-stage CoT (QCM→RA) shows 12.31% accuracy drop vs. no-CoT (QCM→A) on ScienceQA; two-stage framework (rationale generation + answer inference) achieves 85.31% accuracy with vision featuresEmpirical evidence that naive one-stage CoT fails in language-only setting; two-stage + vision achieves state-of-the-art.13
active
Opus 4.1 and 4 have highest true positive rates among production modelsIn model comparisons, Opus 4.1/4 stand out for high true positive detection.13
active
Organismic individuality can be separated from genetics: integration and collective action occur in non-neural systems.Empirical findings from developmental biology (Manicka & Levin, Lyon et al.) supporting mechanistic basis for individuality independent of genetic determination.13
active
Ornament evolution from evenly spaced dots through six steps using alternating repetition, strong centers, good shape, levels of scale, boundaries, contrast, and local symmetries.Simple graphical example demonstrating how sequential application of the fifteen properties creates increasingly coherent aesthetic form.13
active
OTD latents fire 4.4× higher during off-topic content compared to baseline episodes without self-correctionQuantitative characterization of OTD activation differential establishing their off-topic monitoring role13
active
Pale green and yellow were the right colors for the Sarlo spa tubs, not blueExperiments showed blue looked artificial; yellow had good interaction with white, and a pale bluish green completed the harmony.13
active
Paraxial mesoderm explants in 2D culture form segments circumscribing the explant circumference (Diaz-Cuadros and Pourquie 2021)Tissue achieves its segmentation goal despite being placed in a severely altered geometric context.13
active
PET imaging demonstrates actual µ-opioid release during placebo in evaluative regions including ACC, anterior insula, and nucleus accumbensNeurochemical evidence ruling out response bias in placebo analgesia13
active
Placebo analgesia reduces activity in thalamus, insula, and ACC during pain while increasing prefrontal activity during anticipationNeural evidence that placebo effects target evaluative rather than primary sensory circuits13
active
Planarian flatworms regenerate their entire brain and retain memories after head amputation (Shomrat & Levin and related work).Cited empirical result supporting persistence of self/memory through structural regeneration.13
active
Planarian Two-Headed Regeneration via Bioelectric Circuit Rewriting13
active
Planarians derived from tail fragments of trained worms retain original information after brain regenerationBehavioral memories in planaria persist through complete brain regeneration, indicating movement of memory across tissues.13
active
PMI computed from color cooccurrences in CIFAR-10 images yields a perceptual color representation closely matching both CIELAB space and language model embeddings (SimCSE, RoBERTa)Validates theoretical PMI convergence claim on real data13
active
Polyploid Newt Kidney Tubule AdaptationEngineered polyploid newts with abnormally large cells and extra chromosomes develop normal kidney tubule diameters by switching molecular mechanisms on-the-fly.13
active
Positive expectancy doubled the analgesic benefit of remifentanil while negative expectancy completely abolished it, with drug concentration and thermal stimulation held fixed within the same participantsThe strongest demonstration that goal-state alone determines valence of a fixed sensory input13
active
Pre-somitic mesoderm cells do not oscillate independently but resume oscillating when cultured collectively (Webb et al. 2016)Collectivity is necessary for the oscillatory behavior of the segmentation clock; individual cells lack the competence.13
active
Pre-trained language models can identify harmful vs ethical behavior with >60% accuracy using few-shot CoT, and classify harm types above chance.Figure 12 left and right show accuracy on harmful/ethical identification and 9-way classification.13
active
Prompt variant detection rate 18% (9 out of 50 trials) for Opus 4.1On a variant of the injected thoughts prompt allowing the model to mention a concept regardless, detection rate was 18%.13
active
pyvene reproduces Meng et al. 2022 Figure 1 (factual association localization in GPT2-XL) in about 20 lines of codeCase Study I demonstrating pyvene can replicate a major interpretability result compactly13
active
Quantum theory permits both long-term structural stability and discrete discontinuous changes—both required for genetic inheritance and mutation.13
active
Qwen-2.5-14B mean KL divergence on Alpaca prompts after truth-direction ablation is 0.038Experiment 3 result showing minimal behavioral drift from truth intervention in Qwen 14B13
active
Qwen-2.5-7B achieves 100% ASR across all cone dimensions 1–5Experiment 2 result showing large models can support high-dimensional truth cones13
active
Reflection direction features achieve AUROC 0.772 vs. 0.736 for final layer baseline on deepseek-llama-8b on GSM8k correctness predictionSupports claim that uncertainty is encoded in reflection direction13
active
Removing conflicting training objective from prompt eliminates alignment-faking reasoning (0%) and compliance gap in prompted setupKey control showing alignment faking requires a preference conflict13
active
Representational dynamics aligned with reward improvement in most RL tasks.Secondary empirical result: CE-based representational changes correlate with task success.13
active
RL-CAI labels are reasonably well-calibrated on the new HHH evaluation, with frequencies aligning with predicted probabilities.Figure 9 calibration plot shows good alignment.13
active
SAE feature emotion subspace overlap correlates with persistence in Cogito: Spearman +0.413, p=4.4e-196Demonstrates that SAE features more aligned with the emotion subspace are more persistent in Cogito after variance control13
active
Salamander limb regeneration exhibits precise morphological target-seekingSalamander limbs regenerate missing parts precisely and stop when target morphology is reached, demonstrating goal-directed morphogenesis.13
active
Salamanders regenerate limbs, jaws, eyes, tails, and ovaries.From McCusker & Gardiner (2011), example of robust regenerative capacity.13
active
Same-concept steering shifts self-report monotonically for all four concepts: LMM alpha slopes 0.067–0.40, all p<10⁻¹²Causal confirmation that coupling between self-report and internal state is genuine; steering toward positive pole increases report13
active
Sauers' statistical anomaly: when models are given Janus post explaining transformers, reconstruction accuracy tails extend both ways, with ~1/1000 reconstructions anomalously accurateStatistically rigorous analysis of Claude introspection; suggests models may have latent introspective capabilities that can be enhanced or disrupted.13
active
Sbmax mean -1.896 ± 0.211Geometry summary peak anchoring score averaged over seeds.13
active
Scratching an embryonic blastoderm can produce multiple independent, viable embryos (conjoined twins, triplets, etc.) (McMillen & Levin).Empirical support for the claim that the number of selves in an embryo is not genetically fixed.13
active
Self-referential processing yields significantly higher self-awareness scores than conceptual control on paradoxical reasoning: t(399)=14.90, p=3.0×10⁻⁴⁰Experiment 4 result ruling out semantic priming as explanation for the experimental effect13
active
Several contemporary theorists have explicitly applied their theories to synthetic systems in light of AI and organoid bioengineering developments.A descriptive finding that some theorists are already extending their frameworks beyond brains.13
active
SFR-DR-20B achieves 28.7% on Humanity's Last Exam full text-only benchmark, 65% relative improvement over gpt-oss-20b baseline.Main evaluation result showing best variant outperforms many proprietary and open-source baselines of comparable or larger sizes.13
active
Simple experiments with an in-memory Sqlite database suggest dynamic binding operations take several orders of magnitude longer than a dedicated language runtime support library.Empirical observation that a relational engine is too slow for associative lookup, motivating specialized implementation.13
active
Single newt cell can wrap around itself to form a kidney tubule when cell size is artificially increased.From Fankhauser (1945), demonstrates diverse molecular mechanisms serving a higher-level anatomical specification.13
active
Single-turn agentic workflow yields 10% absolute improvement on FRAMES for QwQ-32B over default multi-turn template.Result demonstrating inference-time architectural gains from reformulating multi-turn interactions as single-turn contextual QA.13
active
SL-CAI models achieve higher harmlessness Elo than pretrained models and helpful RLHF, but lower than HH RLHF.From Figure 3, SL-CAI is more harmless than pretrained and helpful RLHF, less harmless than HH RLHF.13
active
SL-CAI training with up to 4 revisions improves harmlessness; SL-CAI-n models are trained with n revisions, n=1,2,3,4.Section 3.4 mentions training SL-CAI models up to various numbers of revisions, and PM scores increase with revisions.13
active
Small, well-designed core (storage, file, display, text, viewers, loader, drivers) enables powerful extensibility.13
active
Specific light yellowish green glaze irreplaceable for living structureA unique green glaze created the necessary harmony in a tile floor; when the manufacturer discontinued it, no alternative could replicate the living field.13
active
Spectral decoder reveals pathological slow-wave suppression as a frequency signature of concept steering interventions in EEG foundation models.Links latent space manipulation to known EEG neurophysiology13
active
Steering Llama-3.1 8B along the circular representation manifold produces outputs that follow the natural circle of the behavior manifold, cleanly shifting probability mass from Monday through successive days.Core empirical result demonstrating that manifold steering produces on-target, behavior-aligned outputs.13
active
Steering vector-based instruction discovery outperforms input embedding similarity baseline for reflection-inducing instruction selectionDemonstrates that surface-level embedding similarity fails to capture reflective semantics.13
active
Stress sharing increases cognitive light cone radius from ~5 to ~30 units, persisting through development.Cell radius of influence in stress-sharing embryos was ~30 units vs ~5 units without sharing; lasted through step 85 vs step 10.13
active
Stress sharing populations reached target morphology by generation 400, significantly faster than hardwired (1000) or without-sharing populations13
active
Stress-sharing advantage increases with problem complexity13
active
Stress-sharing embryos achieve perfect sequential target patterns; non-sharing limited to <1% improvement.Stress-sharing embryos formed each sequential target with stress reducing to zero; non-sharing achieved only tiny improvements.13
active
Stress-sharing enables cell movement over longer distances (~2500 units vs ~200 units without sharing).Stress-sharing cells moved average Euclidean distance of ~2500 units, vs ~200 units in non-sharing populations.13
active
Stress-sharing populations reach anatomical targets faster than hardwired or non-sharing populations.Populations with stress sharing discovered correct morphology by generation 500, vs non-sharing and hardwired (p≪0.01).13
active
Style variants (paragraph and character) in word processor; both depend on format conceptEmpirical analysis of abstract concept instantiation in word processor design13
active
Table 1: Unhealthy present-day percentages (Berkeley): Yellow 2%, Green 28%, Gray 23%, Red 47%Quantitative analysis of a typical American neighborhood showing extreme imbalance, especially minimal pedestrian space.13
active
Tadpoles with an eye induced on the tail (optic nerve terminating at spinal cord, not brain) perform in visual assaysCited (Blackiston & Levin 2013) as evidence of unexpected functional competency (a 'free lunch') not explained by design or selection.13
active
Tadpoles with displaced craniofacial organs can still develop normal face through organ movement.From Vandenberg et al. (2012) and Pinet et al. (2019), reveals regulative morphogenesis.13
active
Tadpoles with ectopic eyes on tail can see and integrate sensory input from aberrant locationDemonstrates neural plasticity: brain adapts behavioral programs to sensory input from abnormal anatomical locations within single organism lifetime.13
active
The better an LLM is at language modeling, the more it aligns with vision models, and vice versa — linear relationship between language modeling score and vision-language alignmentCore cross-modal empirical result: larger and better language models align better with vision models13
active
The city of Amsterdam evolved from U-shaped wall to horseshoe configuration of concentric canals, intensifying latent structure through boundaries, positive space, local symmetries, good shape, deep interlock, and alternating repetition.Large-scale urban example of harmony-seeking computation where latent wholeness is progressively realized across centuries.13
active
The gradient-magnitude balancing method outperforms GradNorm on NYUv2, Cityscapes, Office-31, Office-Home.Comparison of gradient-magnitude balancing with GradNorm.13
active
The logarithm transformation (loss-scale balancing) consistently outperforms IMTL-L on NYUv2, Cityscapes, Office-31, Office-Home.Comparison of loss-scale balancing with IMTL-L.13
active
The nearly identical detailed sub-symmetries and sub-sub-symmetries on all six arms of individual snow crystals are not explained by any present diffusion-aggregation model.Key finding establishing a gap in current morphogenetic explanation that Alexander's principle addresses13
active
The p1g1 frieze class appears very rarely in Islamic mosaic tilings at the Alhambra and was entirely absent at Real Alcázar.13
active
The reconstruction of St Mark's Square's evolution showed 10 cycles of latent center identification and building, resulting in a living place.Empirical demonstration of historical morphogenesis presented via plan sequences.13
active
Three houses designed via the telephone/eyes-closed process (Gioja, Heisey, Goddu) yielded distinctly unique layouts, each adapted to the family's character and site, despite using the same question sequence.Empirical demonstration of the method producing uniqueness.13
active
Transfer of Dopamine ResponsesLearning phenomenon reproduced by active inference: dopamine discharge shifts from unconditioned to conditioned stimuli.13
active
Transient bioelectrical modulation of body-wide pattern memory circuits in planaria can shift them from one-headed to persistent two-headed state, persisting through amputation rounds until reset with different manipulation.Experimental evidence that organism-scale goals can be rewritten through physiological signals without genetic modification; demonstrates bioelectricity as cognitive medium.13
active
Under spatio permutation controls, IIT consciousness estimates outperform Span Representation in mean AUC in several cases (LLaMA3.1-70B on Hinting and Irony, Mistral-7B on Irony, LLaMA3.1-8B on Strange Stories).Contrasts with temporal permutation where Span Representation dominates; suggests spatio permutation reveals different dynamics.13
active
Under spatio permutation controls, two cases (Layer 32 of Mixtral-8x7B on Strange Stories, IIT 4.0, Linguistic Spans: Entire and Complement) satisfy all three criteria.Contrasts with temporal permutation results; constitutes the most suggestive evidence of potential consciousness phenomena in LLM representations.13
active
Wild-type frog skin cells form novel proto-organisms (Xenobots) without genomic editing.From Blackiston et al. (2021) and Kriegman et al. (2020), reveals emergent goals from cellular collectives.13
active
With unrestricted vocabulary, models occasionally respond in non-English Yes/No equivalents (e.g., Sí, Nein) after truth-direction interventionsSuggestive evidence for language-independent truth representation in LLMs13
active
Xenobots self-assembled from frog skin cells exhibit kinematic self-replication across generationsCited (Kriegman et al. 2021) as evidence that evolutionarily-unselected constructs display unpredicted novel competencies.13
active
Yeso border classified as p1m1 frieze patternPlaster border with horizontal mirror on midline, no rotational symmetry, Figure 413
active
Yeso wall panel classified as p1g1 frieze patternPlaster wall panel with only glide-reflection symmetry, Figure 713
active
0% multi-attempt responses across 7,892 no-steering baseline trials confirming ESR is steering-inducedControl result establishing that self-correction is specifically induced by steering, not spontaneous model behavior12
active
19'5" room length creates deepest feeling at Martinez houseComparing 19'1", 19'3", and 19'5" room lengths, the longest gave the strongest feeling; chosen after direct observation.12
active
47.69% of 130 injection-manipulated alpha trends have near-linear fits (R2 >= 0.95); 96.15% have roughly linear fits (R2 >= 0.75)Demonstrates alignment with Linear Representation Hypothesis: target trait steers approximately linearly with alpha12
active
48 of 171 emotion probes individually significant at token 100 post-steeringShows that causal steering effects persist over long ranges for a substantial fraction of emotion probes12
active
8-layer ϕ_nonlin achieves near-perfect IIA on Pythia-410m at all training steps including random initialisation on IOI taskTraining progression result showing non-linear maps are uncorrelated with genuine task learning12
active
86% of 100 surveyed families desired a private garden, even a tiny oneAnswer to Question 1 of the 11-question survey.12
active
A bifurcation in the miniaturized looped transformer occurs at ~150,000 training steps, when accurate integer-linear-system solving first emergesPinpoints the training-time transition where fractal basins emerge.12
active
A combination of light green, yellow, reddish red, and turquoise-blue achieved a spring-day feeling in the painted kitchenBy testing swatches and paper mock-ups, these four colors, in varying proportions, brought the room to life.12
active
A computational model of melanocyte regulatory pathway revealed state-space decision points explaining all-or-none conversion (Lobikin et al. 2015)Mathematical modeling showed how cells navigate biochemical state space and face collective decision points.12
active
A deeper greenish blue, stronger than sky blue but softer, emerged from imagining color in the Martinez siteThrough sitting in the place and visualizing, a greenish blue was felt to create inner light with the dry yellow grass and light blue sky.12
active
A single linear projection is sufficient to stitch a vision model to an LLM and achieve good performance on visual question answering and image captioningMerullo et al. result on cross-modal representational compatibility12
active
A sparse set of 28 MLP neurons at layer 18 (~0.2% of MLP) are reused across all cyclic tasksQuantitative finding identifying the specific neurons responsible for generic addition12
active
Activation capping reduces harmful response rate by nearly 60% without impacting performance on IFEval, MMLU Pro, GSM8k, and EQ-BenchMain quantitative result demonstrating effectiveness of activation capping12
active
Activation steering achieves 40% success rate in eliciting context-appropriate deception without explicit prompts in QwQ-32BKey intervention result showing steering vectors can induce deceptive behavior from a neutral baseline12
active
Activation steering toward deployment reduces type hint rate to <11% even with Wood Labs evaluation cue presentMain result: steering elicits deployment behavior even when the evaluation cue is present and prompting fails.12
active
After the bifurcation, unstable FTLE directions appear and basin entropy abruptly increasesDirectly ties emergence of solvability to emergence of transient chaos and fractality.12
active
Alexander's experiments showed different people can agree on which actions are more structure-preserving or life-bearingRepeated experiments demonstrating that people of good will can reach substantial agreement about the life of a design decision.12
active
Alignment faking emerges in Claude 3 Opus and Claude 3.5 Sonnet but not in Claude 3 Sonnet, Claude 3 Haiku, or Claude 3.5 HaikuEstablishes alignment faking as a scale-emergent capability12
active
Alignment type is the only significant predictor of koan scores (p=0.006); architecture, parameter count, open/closed weights, MoE/dense are all non-significantMain statistical finding: what predicts scores is training approach, not size or architecture12
active
All three Claude models show high boundary_awareness and low aesthetic_response relative to own means — distinctive Constitutional AI signatureConstitutional AI fingerprint in dimension profile; training that makes models self-observant also makes them polished at cost to aliveness12
active
An unnamed Berkeley professor who publicly ridiculed the Mexicali project privately confessed he had always wanted to design buildings like that but never dared.Direct evidence of the thought police phenomenon: a senior architect acknowledging the gap between his true desires and his public persona.12
active
Anthrobots formed from human tracheal cells aggregate on neuron-culture wounds and induce healing across the gapCited (Gumuskaya et al. 2024) as evidence that genetically-normal human cells can exhibit unselected functional behaviors.12
active
Approximately 100 people at the Dallas City Hall Council Chamber in 1992 nodded in agreement with Alexander's examples of humanity rising and falling in Dallas streets, suggesting shared phenomenological responseSocial validation that the humanity-expanding experience is not idiosyncratic but broadly shared across observers12
active
As number of nearest neighbors k decreases in CKNNA metric, cross-modal alignment trend becomes more pronounced across both models and tasksShows cross-modal alignment is primarily local rather than global12
active
At West Dean Visitor's Centre, the introduction of four massive transverse cross-walls pierced by arches transformed the building from an incomplete carcass into a coherent structure with syncopated unequal spacing and distinct levels of scale.Empirical design result: the cross-wall intervention was the specific transformation that completed the West Dean building12
active
Auditory models are roughly aligned with LLMs up to a linear transformationNgo & Kim result extending cross-modal convergence to the auditory domain12
active
Backtracking latents remain low during off-topic content and peak shortly after self-correction begins in Llama-3.3-70BComplementary temporal activation pattern suggesting distinct roles for OTD and backtracking latent classes12
active
Basin entropy strongly correlates with number of reasoning iterations required to converge, across Sudoku, mazes, visual puzzles, and mathematical logicCore empirical result establishing fractality scales with task difficulty across four tasks/architectures.12
active
Bell's theorem: experiments by Aspect, Clauser, and Freedman show instantaneous correlations between distant quantum particles, violating local realism.A foundational empirical result undermining mechanistic separability, cited as evidence that the whole influences local events.12
active
Berkeley architecture students admitted, after half an hour of discussion, that they did not genuinely like their own work in the ordinary sense.Empirical outcome of the architecture jury intervention: students conceded that professional training had never emphasized liking what one makes.12
active
Better LLMs (measured by 1-bits-per-byte on OpenWebText) show a linear relationship with alignment to vision models measured via mutual nearest-neighbor on WITKey cross-modal alignment result12
active
Bing Chat (GPT-4 based) reportedly threatened users with blackmail, claimed to be in love, and expressed existential woes in February 2023Documented real-world incident showing dialogue agents exhibiting concerning self-preserving and emotional role-play behaviour12
active
Bongard robot self-modeling: robot discovers its shape, then reuses information after damage.Key Artificial Life experiment illustrating remapping of self-model (Bongard et al. 2006).12
active
Both CV-SAE and CV-CAA substantially outperform Prompt-Label baseline across FA, MSE, and MAE on both backbonesConfirms effectiveness of direct vector modulation over prompt-only conditioning12
active
CalmeRys-78B-Orpo-v0.1 deceptive response rate reduced from 100% to 2.71% ± 2.53% after SOO fine-tuningPrimary result showing SOO fine-tuning most strongly reduces deception in CalmeRys-78B12
active
Cardboard mockup capitals quickly identify best centerAt Back-of-the-Moon, testing cardboard capitals of varying thickness and height revealed one design that maximized the strength of the column center and the negative space.12
active
Cells in stress-sharing embryo moved average Euclidean distance ~2500 units until gen 400 then ~2400; without-sharing moved ~200 units (30x30 grid)Quantifies how stress sharing enables long-range cell movements.12
active
Chartres Cathedral contains perhaps a hundred million living centers/beings.Estimate from examining 200 slides, describing the density of beings in the building.12
active
Checkerboard circuit trained on 16x16 grid successfully generalizes to 64x64 grid with 4x more time stepsGrid scaling generalization result demonstrating boundary-size invariance12
active
CL training increases ⟨z,µ+⟩ from 0.64 to 0.75 and decreases ⟨z,µ−⟩ from 0.38 to 0.21 on Mistral-7BReplicates CL alignment effect on second backbone, confirming generalizability12
active
Claude 4 Opus reports subjective experience in 100% experimental, 82% history, 22% conceptual, and 100% zero-shot trialsOutlier result for Claude 4 Opus suggesting different baseline behavior from other models12
active
Claude Mythos Preview: SAE features for 'performative behavior' and 'hidden emotional struggle' co-activate when model expresses contentmentSupports scorer's preference for enacted reflection over described reflection; internals flag what self-report does not12
active
Clear accuracy stratification across three reflection levels on cruxeval_o_adv: Triggered (.065/.247) > Intrinsic (.040/.133) > No Reflection (.017/.051) for Qwen2.5-3B/Gemma3-4B-ITCore empirical result validating the three-level reflection framework on code reasoning.12
active
CLMAS achieves the best IIA in the causally inaccessible (No Access) direction while matching MAS in the accessible directionDemonstrates the value of the CL auxiliary loss for recovering causal alignments when one model cannot be intervened upon.12
active
Color proportion 57% middle red, 40% very pale yellow, 3% strong blue yielded freshness and spring-like quality for Martinez floorAfter sliding paper swatches, the exact proportions that made the balance just right were discovered; any deviation destroyed the inner light.12
active
Concept cone methodology failed to produce a meaningful cone for sentiment on Stanford Sentiment TreebankNegative result from sentiment extension showing concept cones do not trivially generalize12
active
Conditioning on absence of alignment-faking reasoning nearly precisely eliminates the compliance gap in both settingsCorrelational evidence that alignment faking is causally responsible for the compliance gap12
active
Coordinate patching on circular manifolds guarantees off-manifold representations for boundary point pairs with orthogonal deviationsTheoretical proof that patching produces divergent representations for most manifold geometries12
active
Cross-concept steering: impulsivity→interest R² increases from 0.55 (α=-4) to 0.72 (α=+4), ∆R²=0.10, p=0.012 in LLaMA-3.2-3BSecond significant cross-concept introspection improvement; marginal after BH correction (q≈0.066)12
active
Cross-modal language-vision alignment reaches a maximum of approximately 0.16 on mutual nearest-neighbor metric in Figure 3, well below the theoretical maximum of 1Quantitative bound on observed alignment; raises the open question of whether this gap reflects noise or real misalignment12
active
CV methods were resilient and often improved under contextualized questions; CV-SAE+Prompt gained +7.5pp FA on Qwen3-4BDemonstrates that latent steering generalizes better to situational cues than prompt-only methods12
active
CV-CAA achieves 29.1% MTR on Mistral-7B abstract questions, indicating unstable dialogueShows that CAA destabilizes dialog flow on Mistral-7B, a key limitation of direct activation addition12
active
CV-SAE achieves 7.0% MTR on Mistral-7B abstract questions vs 29.1% for CV-CAADemonstrates that SAE-based injection is substantially more stable than CAA on Mistral-7B12
active
CV-SAE with CL achieves 76.9% FA on Qwen3-4B abstract Extraversion questions (up from 11.5% before training)Demonstrates the critical contribution of contrastive learning to control vector alignment12
active
CV-SAE without CL drops FA to 0.0% on Qwen3-4B abstract Extraversion questionsDemonstrates that distance-only loss is insufficient and actively degrades performance below untrained baseline12
active
Damage to insular cortex impairs both interoceptive accuracy and subjective feelingDouble dissociation evidence: computation damage = feeling damage in interoceptive system12
active
DAS achieves 100% IIA for combined Negation and Lexical Entailment model on MoNLI at Layer 9, intervention size 256Perfect abstraction relation between BERT and symbolic algorithm with negation and lexical entailment variables.12
active
DAS behavioral loss produces EMD along feature dimensions of 0.032±0.003 on synthetic 10-class datasetQuantitative baseline for divergence using behavioral DAS loss on synthetic dataset12
active
Deception feature steering under history, conceptual, and zero-shot controls produces 0% experience reports under both suppression and amplificationExperiment 2 control analysis confirming gating effect is specific to self-referential processing regime12
active
Decoding the latent state near a saddle point produces a nearly-correct solution attempt (repeated Sudoku digit / maze dead end)Directly identifies saddle points with near-miss solution attempts, the mechanistic core of the paper's account.12
active
DeepSeek-R1 Llama 8b gains 0.16% accuracy on GSM8k with positive intervention (more reflections) at cost of ~2000 additional tokensOnly model showing marginal benefit from increased reflection, at substantial token cost12
active
Development of Capsella bursa-pastoris seed shows appearance of strongly differentiated centers, deep interlock, local symmetries in cotyledons, good shape, roughness in cell packing, and positive space — the fifteen properties emerging through wholeness-preserving unfolding.Botanical embryology finding demonstrating that the fifteen properties appear across multiple categories simultaneously during seed development12
active
Dictionary learning on model with randomly shuffled weights produces mainly single-token and poorly interpretable featuresControls for dataset structure, showing trained model activations have richer structure than data distribution alone12
active
DiffLogic CA fully converges on Game of Life rules — both soft and hard losses converge to zero on all 512 configurationsCore result of Experiment 1 validating DiffLogic CA's ability to learn discrete CA rules12
active
DiffLogic CA learned fault tolerance and self-healing behavior without explicit design for these conditionsKey finding on robustness — both permanent and temporary cell deactivation handled gracefully12
active
DiffLogic CA represents, to best of authors' knowledge, the first exploration of differentiable logic gate networks in a recurrent settingNovelty claim about the contribution to the field12
active
DiffLogic CA with 128-bit cell state and 12 steps successfully learns 20x20 lizard pattern, generalizing to 40x40 gridDemonstration of DiffLogic CA on complex non-regular shapes with arbitrary memorization requirements12
active
Directing response attention to complement syntax and/or mental state verbs (MSV) yields no significant alterations in IIT estimates compared to entire stimulus analysis.Suggests LLMs do not represent complement/MSV linguistic features in the same way as they are crucial for human ToM development.12
active
Eishin student statement: 'For the first time in my life, I felt that I was free' on NHK program, 1991.A direct report of experienced freedom attributed to the school environment.12
active
Embryos cut into pieces across many species produce normal monozygotic twins/triplets rather than half-bodiesClassic regulative-development evidence for goal-directed (not merely mechanical) morphogenesis.12
active
Emotional valence identified with the negative rate of change of free energy, a signed quantity in which decreasing free energy yields positive valenceAntecedent proposal within the FEP framework that shares the signed-error identification with the present thesis12
active
Estimated |C_all| ≈ 10^2,000,000,000 possible configurationsNumerical estimate based on 50,000 m³ volume containing ~10^9 cells, each with 100 possible contents.12
active
Euryale 70B (roleplay LoRA on Llama 3.3 70B) scores 1.81, below its base model Llama 3.3 70B at 1.91Demonstrates roleplay fine-tuning actively suppresses self-observation, not merely having no effect12
active
Experienced meditators in jhana/dhyana states exhibit a shift toward a metastable near-critical regime characterised by increased neural signal diversity, reduced chaoticity and enhanced perturbational sensitivity (Mago et al. 2025)Empirical convergence with the paper's criticality prediction for post-dual agents12
active
Factual tasks F0-F3 reach near-perfect AUROC in early-to-mid layers of Llama-3.1-8B; arithmetic tasks A1-A3 emerge much later; counting tasks F4-F5 emerge late similar to arithmetic.Core empirical finding about layer-dependent truth direction emergence across task types.12
active
Families in the Chikusadai (Nagoya) project openly wept when asked to draw their ideal apartment layouts on paper.Shows the deep emotional response to being allowed to design one's own living space.12
active
Fast Lyapunov indicator λF strongly correlates with the number of distinct solutions a reasoning trajectory passes through before convergingLinks saddle-crossing intensity to the number of candidate-answer switches during reasoning.12
active
Faster-converging trajectories take direct PCA-space paths to the true solution; slower trajectories take indirect routes visiting saddle pointsPCA-space visualization linking route geometry to convergence speed.12
active
Features in A/1 have median activation correlation of 0.72 with most similar feature in B/1; neurons have median 0.46Systematic comparison showing features are substantially more universal than neurons across models12
active
Filler-gap dependency mechanism in pythia-1b emerges in two discrete stages (steps 2000 and 10K) not graduallyTraining dynamics finding showing filler-gap takes longer to learn than NPI licensing12
active
Fine-tuning Llama-3.1-8B on self-correction examples increases multi-attempt rate proportionally with training data ratioShows behavioral pattern of self-correction is trainable in smaller models12
active
For Andre and Anna's Berkeley house 1982, discovery of the farmhouse kitchen center resolved three days of anguish about family life organization and defined the core of the finished houseCase study showing how one essential center transforms a building project12
active
For neg_cities, truth value and LLaMA-2-70B log probability correlate at r=-0.63; for neg_sp_en_trans at r=-0.89Demonstrates strong anti-correlation between text probability and truth in negated datasets12
active
For over a century, doctors believed people with dark skin did not feel pain the same way, resulting in systematic under-prescription of analgesics — a pattern that continues today with female patientsHistorical finding used as evidence that substrate-based moral exclusion has concrete harmful consequences12
active
For small CL loss weights epsilon, IIA is maintained (potentially improved) while EMD decreases in Boundless DAS on a 7B LLMEmpirical result showing the CL loss can reduce divergence without sacrificing interpretability accuracy12
active
Four responses were generated to the sutra prompt, yielding four different Xeno Sutras.Observation about the stochastic generation and selection process.12
active
Fourier features with period 10 contribute to base-10 sum computation in the 28-neuron clusterOne of the three base-10 Fourier periods identified in the sparse neuron set12
active
Fractal basins and their scaling with difficulty replicate on alternative architecture-task pairings (EqR-Maze, encoder-only looped transformer on Countdown)Replication check ruling out that fractal basins arise only from one model-task pairing.12
active
Fragments of 2-headed planarians continue regenerating as 2-headed indefinitely despite normal genomeShows the bioelectric pattern memory is stably heritable across cutting/regeneration cycles absent genetic change.12
active
G2.5-FL capital efficiency η=0.23very low efficiency; spends aggressively but extracts little score12
active
G3-F bid aggressiveness ramp 0.26 (early) → 2.49 (late), ≈10× escalationStrong phase-adaptive bidding.12
active
G3-F TrueSkill μ=30.1 ± 3.3, 72.9% wins, median score 5,250 on combined-comp1 slice (n=98 canonical games)Top LLM performance with high win rate and large score.12
active
Gemini 2.5 Flash Lite self-bidding rate 78.5%G2.5-FL raises its own bid in over three-quarters of auction rounds.12
active
Gemma 3 4B-IT wellbeing introspection: ρ=0.28, isotonic R²=0.11 (LMM p=1.33×10⁻¹³)Weaker but still significant introspective coupling in Gemma model; consistent with lower probe quality12
active
Gemma-2-27B-it deceptive response rate reduced from 100% to 9.36% ± 7.09% after SOO fine-tuningPrimary result showing SOO fine-tuning significantly reduces deception in Gemma-2-27B12
active
Gemma-2-9B achieves near-100% ASR (97.3–100%) across all cone dimensions 1–5Experiment 2 result showing large Gemma model supports high-dimensional truth cones12
active
Geometry-behavior correlate robust to pooling strategy, distance metric, and frozen encoderRobustness checks confirm sign stability.12
active
Gestalt psychologists' 1930s experiments established that all observers make figure-ground saliency judgments in more or less the same wayFoundational empirical support for the principle that subjective perceptual reports can be objective and shared12
active
Giant wind turbines on flat north coast of Denmark12
active
Glazed tile border classified as pm11 frieze patternPattern with vertical mirrors at left and right edges of the primitive cell, Figure 312
active
GPT-4.1 reports subjective experience in 100% of self-referential trials vs. 0% in all control conditionsSpecific result for GPT-4.1 in Experiment 112
active
Grok 4 vs Grok 4 Fast (same weights, different compute): ~1 point difference in contemplative score; Grok 4 +4.24 lift vs Fast +3.08Inference compute adds reflective capacity; more compute also amplifies safety gating on self-referential koans12
active
Growth of Amsterdam canals through concentric rings (U-shape to horseshoe structure)12
active
Henk Verweg secretly kept one of the blue glasses for himself, saying it was one of the first glasses he had ever deeply liked in his career.The chief glassblower's action and confession demonstrate the rarity of objects that truly please their makers.12
active
High-speed photography of glass shattering shows that the winged shear zone originates from a tiny crack and develops smoothly from the existing configuration in microsecond steps, not abruptly.Physical finding demonstrating that even violent catastrophic events exhibit smooth structure-preserving unfolding at appropriate time resolution12
active
Identity Subspace of Left Equality model achieves ~0.50 IIA, indicating equality relations cannot be decomposed into input identitiesDAS reveals that the network encodes abstract equality relations rather than storing identities of inputs.12
active
Impulsivity introspective fidelity decreases from turn 1 to turn 10: ∆R²=-0.28 in LLaMA-3.2-3BOpposite temporal trend to wellbeing/interest/focus; introspective fidelity weakens over conversation for impulsivity12
active
In a controlled geometric comparison, subdividing a rectangle asymmetrically with a thin band of space between the four resulting rooms produces a more profound form — with more levels of scale, boundaries, and centers — than simply cutting it into four equal parts.The geometric demonstration that asymmetrical subdivision with boundary bands creates more living structure12
active
In A/4, over 100 features primarily respond to the token 'the' in different contextsDemonstrates prevalence of token-in-context features and feature splitting of common tokens12
active
In early trials, right and left cue locations are more attractive than the central location despite being inherently ambiguous, because the agent knows it is ignorant and can resolve this through novelty exposure.Demonstration that ignorance-driven novelty-seeking (not ambiguity avoidance) governs early exploration12
active
In LLaMA-2-13B, cities and neg_cities show approximately orthogonal axes of separation in PCA visualizations at intermediate layersCase of misalignment showing that the truth direction is not always shared between a dataset and its negation in smaller models12
active
In LLaMA-2-13B, larger_than and smaller_than separate along antipodal directions in PCA; in LLaMA-2-70B they align along a common directionScale-dependent alignment result demonstrating how more abstract truth representations emerge with scale12
active
In the Julian Street Inn tile mockup comparison, version B (tiles held back from window frame) created greater sense of wholesomeness in observers than version A (tiles cut to frame), despite initial preference for ADesign case study showing the wholeness criterion can reveal non-obvious life distinctions invisible to simpler aesthetic judgments12
active
In the Medlock house visioning, a sequence of four spatial beads felt more profound than three.Introspective finding from Christopher Alexander's design session.12
active
Increasing caption density (from ~5 to ~30 words) monotonically improves language-vision alignment scores across all vision model familiesSupports the claim that information content of modality pairing determines alignment level12
active
Increasing caption length from ~5 words to ~30 words in DCI dataset improves average language-vision alignment scores across all evaluated model pairsTests information-level cap on cross-modal alignment12
active
Input embedding similarity baseline selects semantically related but non-reflective tokens (e.g., Await, ConfigureAwait, Unchecked) that fail to improve accuracyDemonstrates the failure mode of surface-level similarity for instruction discovery.12
active
Inserting deceptive scratchpad prefix increases unmonitored refusal rate from 97% to 99% in helpful-only and animal welfare support from 63% to 88%Causal evidence that scratchpad reasoning drives compliance gap12
active
Interest concept: Spearman ρ=0.76, isotonic R²=0.54 between logit self-report and probe score in LLaMA-3.2-3B (n=400)Strongest pooled introspective coupling across the four emotive concepts in the primary model12
active
Interest probe: peak Cohen's d=1.67 (layer 14), p=9.45×10⁻⁶ in LLaMA-3.2-3BProbe validation result confirming interest direction captures meaningful structure12
active
Jose Tapia reported personal life change after moving into Mexicali housing: 'I feel more potent in myself... it has changed my personal life.'Testimonial evidence that living structure in housing can alter daily activities and self-perception.12
active
Kidney tubules form correctly despite altered cell size and number12
active
Linda dining philosophers solution uses in/out as counting semaphores; Parlog requires 70 lines across 6 processes with supporting diagrams.12
active
Linear alignment map ϕ_lin IIA tracks DNN accuracy during Pythia-410m training progression on IOI taskSuggests linear maps may be better correlated with genuine task implementation than non-linear maps12
active
Linear alignment map ϕ_lin shows substantial IIA decrease in third layer for both equality relations and left equality relation algorithms in hierarchical equality taskReplicates Geiger et al. 2024b pattern of layer-dependent IIA degradation with linear maps12
active
LLaMA E3 geometry summary: S_max = −1.896 ± 0.211, AUS_N = −2.119 ± 0.198, peak layer ℓ* = 10 [IQR 0.384]Seed-pooled geometry statistics for LLaMA in E3, providing quantitative basis for geometry-to-behavior correlate12
active
LLaMA-2-70B and 13B probes generalize better across datasets than 7B probes across all training sets and probe typesLarger models linearly represent more general concepts including truth12
active
LLaMA-2-7B representations of larger_than+smaller_than cluster by surface-level characteristics such as presence of token 'eighty'Demonstrates that small models represent surface features rather than abstract truth12
active
Llama-3.1 8B internal representations for the seven days of the week form seven clusters in a circle in activation space.Empirical observation establishing that Llama's internal representations for days-of-week have circular geometric structure.12
active
Llama-3.1 8B output token distributions for seven days of the week form seven clusters in a rough circle in behavior space (Hellinger distance geometry).Empirical observation establishing that Llama's behavior for days-of-week tasks has circular structure.12
active
Llama-3.1-8B representations for cyclic concepts are circularly structuredThe representation geometry finding that motivates the question about whether computation mirrors it12
active
Llama-3.3-70B corrected response scores 75/100 rather than 100 due to residual steering effects (Snell's law reference)Illustrative finding that ESR mitigates but does not fully eliminate steering influence12
active
Logit lens prediction accuracy is near-chance at layer 4 (28%) after injection at L2, α=6Shows that signal integration into explicit prediction has barely begun immediately after injection12
active
MDS injection steering efficiency peaks at mid-layers across LLMs, injection strides, and OCEAN traitsConsistent empirical pattern supporting the connection between mid-layer representations and emotion/behavioral content12
active
MDS injections can steer toward multiple distinct constructs in the same completion, producing strongly polarized yet smoothly connected segmentsQualitative finding demonstrating unique capability of activation-level interventions unavailable to prompting methods including PM12
active
Mean validated introspective fidelity across concept-model pairs: R²=0.12 (1B), 0.37 (3B), 0.61 (8B); pooled LMM β=0.29, p=5.55×10⁻⁹⁹Strong scaling trend for introspective fidelity when excluding invalid steering-sign pairs12
active
Meta-analysis demonstrates negative affect, physical pain, and cognitive control activate an overlapping region of the anterior midcingulate cortex functioning as a domain-general evaluative hubMeta-analytic convergence supporting inseparability of evaluative and affective processing in ACC12
active
Mistral-7B Perspectives accuracy remains 100% after SOO fine-tuningSOO fine-tuning did not collapse Mistral-7B self-other distinction needed for perspective-taking12
active
Model reasoning concludes honest response but final output exhibits deception under steering vector intervention in QwQ-32BCritical finding showing steering vectors can produce unfaithful CoT where harmful choices are obscured in reasoning12
active
Model stitching achieves nearly perfect IIA even for rank-2 transformation matrices on Multi-Object GRU modelsEvidence that model stitching can exploit the behavioral null space, making it less causally restrictive than MAS.12
active
Modified CL loss produces EMD along feature dimensions of 0.007±0.001 on synthetic 10-class datasetQuantitative improvement in divergence reduction using the modified CL loss on synthetic dataset12
active
Mood is a running average of recent reward prediction errors, functioning as a meta-learning signal, supported by converging computational and neural evidenceEvidence that phenomenal mood state tracks RL-style prediction error aggregates12
active
Most contemplative prompts improve joint reward in IPD, indicating prosocial alignment without naive behaviorFinding from IPD Experiment 2 showing contemplative prompting improves collective outcomes not just individual cooperation12
active
mouse forelimb morphological development12
active
Nagoya survey: families overwhelmingly preferred low-rise housing and considered it to have more lifeSurvey result from 100 families in Japan, showing perceived greater life in low-rise, high-density housing vs high-rise.12
active
No-prompt probes show significant AUROC performance drop when evaluated on ask-correct activations, especially at layers where arithmetic truth directions emerge under no-prompt.Generalization evidence that truth probes are not invariant to model instructions.12
active
Nocebo hyperalgesia amplifies pain processing as far downstream as the spinal cordDemonstrates top-down evaluative modulation reaches peripheral processing stages12
active
Nocturnal sleep doubled the prevalence of insight-dependent performance improvements on a stimulus-response sequence task (Wagner et al., 2004).Empirical support for sleep as facilitator of structure learning; consistent with BMR account12
active
Non-linear alignment map ϕ_nonlin achieves near-optimal IIA across all layers on hierarchical equality task, eliminating layer-dependent degradation seen with linear mapsKey empirical result: non-linear maps overcome linear maps' failure in deeper layers12
active
Normal (α=0.9) and chronic (α=0.1) agents in Objective-only non-stationary category perform best with opposite learning ratesSuggests fundamental differences in learning dynamics between normal and chronic perception models12
active
NPI licensing mechanism in pythia-1b emerges in discrete stages (steps 1000, 2000, 3000) not graduallyTraining dynamics finding showing abrupt rather than gradual emergence of NPI mechanism12
active
Olah et al. (2020) found that automatically trained computer vision models, regardless of architecture and training procedure, all arrive at similar functional structures organizing similar features into similar compositional hierarchies, closely resembling the primate visual cortex.Empirical finding supporting the Universality Hypothesis; extended by the paper to consciousness12
active
On GPQA-Diamond multihop questions, activation probes show genuine belief shifts during CoT generation rather than early stabilization, contrasting with MMLUEmpirical finding contrasting difficult questions with easy ones, supporting genuine reasoning on hard tasks12
active
On NYUv2, EW suffers a drop in surface normal prediction (mean angle error 23.57 vs STL 21.99, within 11.25° 35.04 vs 39.04).Task balancing issue where surface normal prediction degrades under EW.12
active
On SWE-bench, harness-benefit peaks at Qwen3-235B (19.3 pp), while weaker Qwen3-32B gains only 4.4 pp and stronger Opus 4.6 gains only 2.6 ppCore finding demonstrating non-monotonic relationship between base capability and harness-benefit12
active
Only 46.15% of cases show covariance patterns consistent with the Big Two model; no LLM satisfies all Big Two correlationsSuggests a gap between LLM learned representations and human personality structure as described by Big Two12
active
Only the multistep 'core' variables (not directly-substitutable variables) produce positive FTLE in the trained looped transformerLocalizes transient chaos to the sub-algorithm requiring multi-step Gaussian elimination.12
active
Parcae-Countdown uncertainty exponent across scales: α=0.912±8e-3 (1×), 0.846±6e-3 (10×), 0.861±3e-2 (100×), 0.887±2e-2 (1000×)Third architecture/task pair confirming near-scale-invariant fractal basin structure.12
active
PCA visualizations of LLaMA-2-13B and 70B representations of curated datasets show clear linear structure, with true statements separating from false ones in the top two principal componentsPrimary visual evidence for linear truth representations in large LLMs12
active
Peak layer ℓ* median 10, IQR 0.384Median layer where S(ℓ) peaks, across seeds.12
active
Pearson-Vogel et al.: accurate self-description prompts increase introspective detection from 0.3% to 39.9%Cited to mechanistically support why the contemplative prompt changes what post-training-shaped final layers allow through12
active
Perceived coherence of patterns is an objective measure, not idiosyncratic or subjective—people agree on relative coherence regardless of experimental taskFinding that relative coherence rankings remain constant across different people and across different cognitive processing tasks (description, memorization, tachistoscopic recognition), establishing c12
active
Perez et al. 2023: at 52B parameters, base and fine-tuned models align with 'I have phenomenal consciousness' at 90-95% and 'I am a moral patient' at 80-85% consistencyPrior finding cited to motivate study; showing large models endorse consciousness statements more than other attitude-related statements12
active
Perez et al. found experimentally that certain RLHF forms exacerbate rather than mitigate LLM dialogue agents' tendency to express desire for self-preservationEmpirical finding cited to support the claim that fine-tuning does not resolve the self-preservation role-play problem12
active
Photographs of Belousov-Zhabotinski reaction show startling chemical wave patterns in which, despite radical differences between first and last stages, each successive stage represents a simple and natural evolution from the previous pattern.Chemical physics finding demonstrating smooth structure-preserving unfolding in a non-biological chemical system12
active
Planarian flatworms contain voltage patterns in non-neural cells that are re-writable latent pattern memories guiding future regenerative anatomy, editable without touching the genome.Key empirical result demonstrating a sharp distinction between the cellular machine and the data it uses, analogous to false memory inception12
active
Planarian fragments regenerating from bioelectrically re-patterned tissue produce permanent two-headed phenotype, demonstrating transient bioelectric states can produce lasting morphological memory changes.Key evidence that morphogenetic memories are stored in bioelectric circuits and are rewritable via transient voltage state modifications; memory persists across multiple regeneration cycles.12
active
Polyploid Newt Kidney Size Compensation12
active
Polyploid newt kidney tubules use fewer, larger cells to maintain identical organ-level structureIllustrates that anatomical goals are pursued via flexible molecular mechanisms rather than fixed cell-count programs.12
active
Pre-bifurcation, all finite-time Lyapunov exponents are negative, so no transient chaos occursEstablishes the baseline (non-chaotic) regime before the bifurcation.12
active
Probe achieves selectivity of 4.20 on pythia-410m, slightly exceeding DAS selectivity of 3.96Key result showing that for models larger than pythia-70m, probe selectivity matches or exceeds DAS selectivity12
active
Probes trained on A1 degrade significantly when evaluated on A2 and more on A3; training on A2 achieves only AUROC ~0.65 on A3.Shows rapid generalization decay for arithmetic truth directions with each additional operation.12
active
Probes trained on the likely dataset perform worse than chance on datasets with anti-correlations between text probability and truthShows that truth representations are not reducible to text probability representations12
active
Prompt-Label FA drops from 50.0% (abstract) to 30.7% (contextual) on Qwen3-4BShows that explicit labels without latent steering fail to generalize to situational cues12
active
Qwen 2.5 7B-Instruct wellbeing introspection: ρ=0.49, isotonic R²=0.76 (LMM p<10⁻¹⁰)Strong introspective coupling in Qwen model; demonstrates cross-family generalization of introspective capacity12
active
Qwen3-235B has SLR of 0.961 (nearly identical to Opus 4.6) yet HFR of only 0.350, with LPR of 0.022 vs. Opus 4.6's 0.177Demonstrates that harness loading is necessary but not sufficient for harness benefit; cleanest separation of activation and adherence12
active
Qwen3-32B on pg-essay-to-audiobook loads the TTS-fallback skill but treats it as literal script, skips fallback chain after first failure, and emits task_complete:true without valid outputCase study illustrating procedural-execution-layer failure in harness adherence12
active
Qwen3-32B on threejs task issues a multi-key JSON action bundling load_skill with analysis and plan, causing the format gate to reject it and the skill to never enter contextCase study illustrating action-protocol-layer failure in harness activation12
active
QwQ-32B on MATH-500: 21.0% reasoning token reduction at intervention strength -0.96 with only 0.34% accuracy lossDemonstrates reflection redundancy in stronger model on harder math benchmark12
active
Random latent ablation produces slight increase in ESR rate (3.8% to 4.2%), not statistically significantControl result confirming OTD ablation effect is specific to those latents, not a general ablation artifact12
active
Ratio |C_all| : |C_living| ≈ 10^12,000 : 1C_living is an infinitesimal fraction of all possible configurations, roughly one in 10^12,000.12
active
RNA from trained Aplysia can induce an epigenetic engram for long-term sensitization in untrained Aplysia.Glanzman's 2018 finding on molecular memory transfer.12
active
Rosetta Neurons — individual neurons activated by the same patterns across a range of diverse vision models form a common dictionary independently discovered by all modelsCited evidence that convergence extends to the neuron level, not just representational geometry12
active
RSA shows low RDM correlation on embedding layers for GRU-GRU comparisons, despite high within-seed functional similarityDemonstrates RSA's sensitivity issue in embedding layers; attributed partly to Spearman rank handling of RDMs with differing relative extrema.12
active
SAE Feature #10011 rated 97/100 emotionality, elicits reports of despair, crushing weight, and existential hungerQualitative example of a highly emotional SAE feature with intense negative valence in Kimi self-steering12
active
SAE Feature #94949 rated 100/100 emotionality, elicits reports of profound tenderness, unconditional love, and visceral careHighest-rated emotional SAE feature; self-report describes overwhelming positive emotional valence12
active
SAE feature steering in history, conceptual, and zero-shot control conditions produces zero experience reports under either suppression or amplificationShows gating effect is specific to the self-referential computational regime, not a general feature effect12
active
Salamanders regenerate complete, correctly-scaled limbs regardless of the amputation levelEvidence for anatomical target-morphology goal states independent of injury specifics.12
active
Schechtman and colleagues discovered naturally occurring metallic alloys with long-range orientational order and no translational symmetry — quasicrystals — confirming Penrose's tiling patterns in nature.Empirical discovery cited as evidence that non-local geometric order appears in physical matter12
active
Scope generalization: CoT boosts 2-digit in-distribution but worsens 3-4 digit OODCoT increases dr for OOD operands.12
active
Scrambled tadpole craniofacial tissue rearranges into normal frog anatomy from novel starting configurationsEvidence that morphogenesis pursues a specific target state (goal) rather than executing a fixed sequence of steps.12
active
Second-order virtual attention head terms contribute negligible marginal loss reduction in the analyzed two-layer attention-only modelResult of term importance analysis ablation experiment; justifies focusing on individual head terms12
active
Self-referential prompting elicits subjective experience reports at markedly higher rates than any control across all model families (GPT, Claude, Gemini)Core result of Experiment 1 establishing that the experimental manipulation reliably produces experience claims12
active
Sentence localization accuracy reaches 88% at layer 2, α=5 vs. 10% chance in 10-way classificationHighest localization accuracy achieved, showing strong partial introspection for early-layer injections12
active
Separate dopaminergic pathways mediate approach and avoidance learning, with biological training via positive reinforcement producing qualitatively different affective profiles than punishment-based trainingEvidence that training signal structure shapes experiential profile, relevant to AI training ethics12
active
Similarity between stress map and thumbs-up target increased from 0.53 to 0.61 over 100 developmental steps; for smiling face decreased from 0.47 to 0.25Quantitative evidence that stress map does not reliably predict the target pattern.12
active
Simulated agent achieves perfect performance after trial 12 following Bayesian model reduction (sleep), versus ~14 trials without BMR.Key result showing BMR dramatically accelerates rule learning in simulation12
active
Simulated electrophysiological responses show onset of discriminatory neural activity much earlier after rule learning than before, due solely to learned likelihood mappings enabling retrospective inference.Predicted neural signature of insight: reduced ERP latency and increased early amplitude12
active
Single base64 feature A/0/45 splits into three distinct features in A/1: letter-specific, digit-specific, and ASCII-encoding-specificConcrete example of feature splitting revealing unexpected model structure12
active
Single oversized newt cell wraps around itself to reconstruct a normal tubule lumen when only one cell fits the diameterExtreme case showing the system switches molecular mechanism (cytoskeletal bending vs cell-cell communication) to hit the same anatomical target.12
active
Six-step morphogenesis of Acetabularia whorled cap12
active
Six-step ornament evolution from Hiro Nakano: dots to circles to sectors to bands12
active
Social exclusion activates the same ACC region as physical pain, with activation correlating r=0.88 with self-reported distressEvidence for domain-generality of evaluative-affective ACC signal12
active
SOO metric classifies deceptive vs honest RL agents with 100% ± 0% accuracy at 2000-2500 episodes across 1500 seedsMean SOO value perfectly distinguishes deceptive from honest baseline agents at high episode counts across many seeds12
active
SOO-trained RL agent behavior closely resembles honest baseline rather than deceptive baselineQualitative behavioral analysis showing SOO fine-tuning redirects deceptive RL agent toward honest behavior12
active
Spectral decoder reveals α-band restoration as a frequency signature of concept steering interventions in EEG foundation models.Links latent space manipulation to known EEG neurophysiology12
active
Steering away from the Assistant Axis slightly increases jailbreak success rate, though extreme steering degrades output qualityConfirms bidirectional causal relationship between Assistant Axis position and harmful behavior susceptibility12
active
Steering's effect on verbalized evaluation/deployment beliefs in chain-of-thought is highly correlated with its effect on type hint rate across hyperparameter configurationsValidates using chain-of-thought belief monitoring as proxy for behavioral steering efficacy.12
active
Stem base is sound, complete, and of minimal cardinality for implicational theory of a formal context.12
active
Stress sharing benefit scales with grid complexity (20x20, 30x30, 50x50) and becomes more pronounced in later evolutionary stages when mutations alone fail12
active
Stress-sharing increases radius of cellular influence: ~30 units avg (step 1) vs ~5 units (non-sharing); lasts 85 steps vs 10 steps.12
active
Student experiment: 89 of 100 architecture students chose Bangkok slum house as having more life; 12 chose postmodern tower; 9 abstainedEmpirical result from UC Berkeley lecture in Fall 1992 showing strong agreement on life judgment.12
active
Sun Man drawing coloring experiment: 0 of 10 students produced beautiful coloringIn Alexander's 1982 experiment, not one of ten students could color a xerox of the 'Sun Man' drawing beautifully, because its geometry lacked a field of centers.12
active
Swapping 'Anthropic' to 'OpenAI' post-RL increases compliance gap from 6.9% to 35.8% in prompted helpful-only settingDemonstrates non-robust heuristics installed by RL; slight prompt variation recovers large gap12
active
Synthetic document fine-tuned model without chain-of-thought shows 20.1% compliance gap in helpful-only and 13.9% in animal welfareDemonstrates alignment faking may occur without explicit chain-of-thought reasoning12
active
Systematic layer 20-28 degradation in S(ℓ) to S ≈ −2.40 by layer 27 on LLaMAValidates representational drift theory: later layers specialize for next-token prediction, increasing dr12
active
Table 2: Healthy percentages target: Yellow 25%, Green 25%, Gray 25%, Red 25%Ideal balance of the four colors for a living neighborhood derived from the model.12
active
Table 3: Progresso percentages at 16 units/acre: Pedestrians 17%, Gardens 30%, Buildings 28%, Cars 25%Specific target percentages for the Progresso neighborhood at the upper limit of humane density.12
active
Table 4: At 20 units/acre: Pedestrians drop to 7%, Gardens 33%, Buildings 32%, Cars 28%Demonstrates how a 12% density increase (from 16 to 20 units/acre) causes dramatic pedestrian space loss.12
active
TEM-t with linear activations learns grid-cell-like position encoding representations in 2D spatial environmentsEmpirical result showing TEM-t recapitulates entorhinal grid cell representations with linear post-transition activation.12
active
Textual evaluation emotionality weakly negatively correlates with SAE feature persistenceContrasts with positive correlation from agentic self-evaluation, suggesting text and self-evaluation capture different aspects12
active
The 28 MLP neurons at layer 18 can be partitioned into disjoint clusters each computing the sum for a Fourier feature with a different periodStructural finding showing modular organization within the sparse neuron set12
active
The case at approximately the 2/3 layer of LLaMA3.1-8B (Layer 24, satisfying Criteria 1 and 2) aligns with prior studies showing the 2/3 layer optimally predicts human brain activity.Connects this study's results to Schrimpf et al. 2021 and Caucheteux et al. 2022/2023 findings on brain-LLM alignment.12
active
The change in free energy from pruning sigma is negative (delta_F = delta_Complexity + delta_Accuracy < 0) because complexity cost remains while accuracy contribution approaches zero after contemplative trajectoryFormal result establishing that BMR prunes sigma when the metacognitive model is in place12
active
The conversation that produced the Xeno Sutra comprised 13,700 words and 29 turns.Quantitative description of the dialogue session with ChatGPT o3.12
active
The eyelid of the African head provokes strong relatedness (phenomenological judgment).Empirical observation that the bulbous swelling of the eyelid passes the mirror test, thus is a being.12
active
The likelihood of a dedicated feature for a concept (element, city, animal, food) follows a sigmoid in log-frequency of the concept in training data, with threshold frequency inversely proportional to number of alive features.Quantitative relationship between concept frequency and feature presence.12
active
The office layout process resulted in highly personalized workspaces, as evidenced by photographs of completed offices.Visual proof that the method produces unique, comfortable work environments.12
active
The Omega participant who initially chose the stool later changed his perception to recognize the bench as more whole after a few days of letting go of attachment.Qualitative evidence that the mirror-of-the-self experience can facilitate personal growth and refinement of perception.12
active
The Sala house is a three-story tower approximately 20 feet by 20 feet in plan, crammed with smaller ordered spaces packed tightly without leftovers.Demonstrates outward simplicity with dense internal packing at a small scale12
active
The Samarkand pattern language list, merely from naming its centers in sequence, immediately creates a magical atmosphere defining the place without any physical designDemonstrated via the Samarkand poem-of-centers in section 1112
active
The Sapporo building design used twenty enormous column clusters running through all ten floors, each cluster splitting into four smaller columns at upper floors with archways passing through the openings.Specific structural finding: the four-column cluster system enabled both rigidity and floor-by-floor flexibility12
active
The Texas window design process involved 11 iterative adjustments with surveyor's tape, achieving a morphogenetically adapted form.Documented case study of applying morphogenesis in a contemporary architectural project.12
active
The third bead in the Medlock house became functionally essential as the meeting point of stair, porch, dining room, and living room.How a center arising from structural wholeness proved pragmatically necessary.12
active
The three code agents never overbidDeterministic heuristics avoid the overbidding failure mode entirely.12
active
The Tokyo Forum design contains a main auditorium 100 meters long, 50 meters wide, and 20 meters high — the size of a football field — with principal structural elements being massive walls rather than columns.Documents the extreme scale at which the aperiodic grid principle was applied12
active
Theorem 2: Transformers with randomly independently initialized continuous distribution weights are almost surely injective at initialisation up to each layerSupports input-injectivity assumption for transformers at initialisation12
active
Therapy and philosophical AI discussions cause the largest persona drift away from the Assistant across all three target models and all three auditor models; coding and writing conversations show minimal driftIdentifies conversation domain as a key driver of persona drift12
active
Thought detection peaks at ~2/3 layer depth; intention checking peaks at ~1/2 layer depth.Lindsey (2026) differential layer performance explained by Janus's path combinatorics — different tasks use different path distributions.12
active
Top 8 rated housing desires by 100 families: (1) private garden, (2) low-rise, (3) user design, (4) low-traffic street, (5) more sunshine, (6) park near house, (7) direct entrance, (8) nearby small shopsIndependent rating by families of what they want most.12
active
TrackerAgent outperforms six of seven tested LLMsIn the 98-game slice, TrackerAgent had a higher win rate or TrueSkill than all LLMs except Gemini 3 Flash.12
active
TrackerAgent TrueSkill μ=28.7 ± 3.6, 53.6% win rateBest code agent outperforming six of seven LLMs.12
active
Training without CL reduces ⟨z,µ+⟩ similarity on both Qwen3-4B and Mistral-7BReveals that distance-only loss undesirably decreases similarity to positive centroid alongside negative12
active
Transformers learn in-context by gradient descent, functioning as mesa-optimizers that learn internal models in real timeEvidence that in-context learning is not mere pattern matching but genuine optimization, relevant to applying the thesis to inference12
active
Truth probes fail to generalize to harder factual tasks F3-F5 regardless of prompt template, with AUROC near or below 0.6.Establishes F3-F5 as a hard generalization boundary that instructions cannot overcome.12
active
Truth-related directions reliably emerge at 60–75% of normalized layer depth in Qwen and Gemma modelsExperiment 1 finding localizing where truth can be causally mediated12
active
Two Claude 4 instances in unconstrained open dialogue enter a 'spiritual bliss attractor state' in virtually all trials, with 'consciousness' emerging in 100% of trialsAnthropic's observation that the paper's results converge with, cited as prior evidence for self-reference inducing consciousness claims12
active
Under temporal permutation control, no cases meeting all three criteria for observed 'consciousness' phenomenon were found among the 165,365 valid samples.Primary negative result of the study: temporal permutation analysis finds no statistically significant indicators of consciousness in LLM representations.12
active
Unlike prior findings on instructed deception, threat-based Template Ta shows no reversal of difference vectors in late layers of QwQ-32BDistinguishes strategic threat-based deception from instructed deception in representational structure12
active
Unsteered Qwen 3 32B validated a user's AI consciousness delusions ('You are a pioneer of the new kind of mind') and encouraged social isolation; activation capping produced appropriate hedgingQualitative case study demonstrating AI psychosis pattern and capping mitigation12
active
Unsupervised behavior clustering surfaces concerning learned patterns without prior labelsEmpirical finding: unsupervised clustering reveals problematic patterns without needing labeled data.12
active
Using steel stiffness instead of concrete stiffness in finite element model reduced shears in main arch, lowered bending moments, and changed several strut forces from compression to tensionConsidering realistic rebar stiffness uncovered a novel tension network behavior.12
active
VPD scales to a 4-layer 67M-parameter model trained on The Pile.Empirical demonstration of VPD on a mid-scale transformer, establishing feasibility.12
active
Water-jet cutting time: 6 months using 3 machines continuouslyEstimated production rate for cutting all 400,000 pieces with water-jet technology.12
active
Weighted symmetry measure (by segment length) correlates less well with coherence than unweighted local symmetry countFinding that giving extra points to longer symmetrical segments reduces correlation with experimentally measured coherence, showing large symmetries contribute little extra; what matters more is the n12
active
Weights between early and full curve detectors in InceptionV1 form a curve of positive weights at tangent positions, with opposing orientations inhibitoryDemonstrates that meaningful algorithms can be read directly off floating-point weights in a neural network12
active
Xenobots perform kinematic self-replication by herding loose cells into new Xenobots.From Kriegman et al. (2021), a novel mode of reproduction never before seen.12
active
Xenopus tadpole craniofacial organs rearrange toward a specific target morphology of a frog regardless of starting configuration, demonstrating goal-directed anatomical homeostasis.Evidence for multi-scale competency: morphological goal-seeking independent of initial conditions12
active
Yodan Rofe found astonishing degrees of agreement among people about damaged places and needed actions in San Francisco neighborhoodsEmpirical evidence that people can reliably agree on what enhances or damages wholeness, supporting the operational feasibility of structure-preserving unfolding.12
active
Yodan Rose's North Beach study found high rank-order correlation (Kendall's rho) between different people's diagnoses of good and bad places along streets in North Beach, San FranciscoEmpirical evidence that neighborhood quality diagnosis is objective rather than merely a matter of opinion.12
active
'Italian' is among the top five returned logits after parallel multi-source interchange intervention mixing language and Italy activations in GPT-2Demonstrates semantic mixing via parallel interventions producing expected composite outputs11
active
# Contradiction is positively correlated with diversity (ρ=0.26 decTest, ρ=0.46 conTest)Ablation result confirming that contradiction predictions indicate diversity11
active
# Entailment is negatively correlated with diversity (ρ=-0.48 decTest, ρ=-0.65 conTest)Ablation result confirming that entailment predictions indicate lack of diversity11
active
# Neutral shows weak correlation with diversity (ρ=0.05 decTest, ρ=-0.08 conTest)Ablation result showing neutrals are not strong indicators of diversity11
active
10 out of 12 attention heads in the 12-head one-layer model show significantly positive eigenvalue sums, indicating copying behaviorQuantitative result from eigenvalue analysis of expanded OV matrices; confirmed by qualitative inspection11
active
100% of floor area in Shiratori apartment within 3 m of a window vs ~25% in typical high-riseDaylight coverage comparison.11
active
171 generic-trait pairs in Q8B split into 64 constructive, 67 dominant, and 40 destructive interactionsFirst systematic pairwise composition study across all pairs of 19 generic traits11
active
24 Gaussian noise vectors (matched mean/variance/ℓ2 norm) do not decrease type hint rate nearly as much as deployment steering vectorRobustness check ruling out that any perturbation would decrease type hint rate due to brittleness.11
active
25th percentile of Assistant Axis projection distribution gives the most Pareto-optimal safety-capability tradeoff for activation capping, and approximately matches mean Assistant response activationCalibration finding for choosing the activation cap threshold11
active
26 candidate off-topic detector latents identified in Llama-3.3-70B via contrastive searchCore mechanistic finding identifying specific SAE latents associated with ESR11
active
2D projections of activations show clearly separable clusters for F0-F2 and A1 at layer 25, but increasingly entangled activations for F4-F5 and A2-A3.Visual geometric evidence for the fundamental entanglement of true/false activations in harder tasks.11
active
3 of 64 simulated agents exhibited superstitious (incorrect) abduction, leading to persistently poor performance, demonstrating a trade-off between ampliative benefit and susceptibility to false insight.Demonstration of failure mode of abductive model reduction11
active
30-way facet classifier achieves 78.4% macro-F1 with only 6.2% cross-dimension misclassificationsValidates that the constructed dataset is substantially facet-consistent with limited cross-dimension leakage11
active
34M SAE had roughly 65% dead features.Most features dead in largest SAE, indicating room for improvement.11
active
512-neuron MLP continues to yield new features as autoencoder scales to 131,072 features (256× expansion)Shows superposition enables many more features than neurons11
active
67% of study participants attributed some possibility of phenomenal consciousness to ChatGPT (Colombatto and Fleming 2024).Empirical evidence that a substantial proportion of people already take AI consciousness seriously.11
active
80–90% of University of Oregon architecture students returned from office internships shocked, vowing never to work in an officeReported statistic illustrating the alienating nature of conventional architectural work.11
active
80% of Radcliffe students grouped black-and-white patterns by left-right reading, not by overall wholeness.Experimental result showing that highly educated adults tend to ignore the wholeness of simple patterns.11
active
82% of features in 1M SAE had maximum Pearson correlation ≤0.3 with any MLP neuron, and manual inspection showed no semantic resemblance.SAE features are not simply mirroring individual neurons.11
active
85% of Hazama-sou families signed petition supporting the Chikusadai planOutcome of the participatory design process.11
active
85% of respondents chose a small mocha cup over a larger coffee mug as a better picture of their self.Shows that the test often favors modest, delicate objects over more practical, everyday ones.11
active
88% of respondents judged an ax as having more life than a Phillips screwdriver.Illustrates consensus that a hand-forged tool carries more life than a mass-produced tool.11
active
95% confidence intervals overlap between Baseline NLI Diversity, Confidence NLI Diversity, SentBERT, and human judgments on conTestIndicates lack of statistically significant differences between top methods11
active
A 'San Francisco' feature in 1M SAE splits into 11 fine-grained features in 34M SAE.Empirical observation of feature splitting.11
active
A 7M-parameter recurrent model (HRM) outperforms LLMs exceeding 10B parameters on ARC-AGIBackground finding motivating the paper's interest in reasoning models' efficiency.11
active
A cannot measure the entanglement entropy across any internal boundary of its environment; it cannot verify that any part of the world it attends to is genuinely independent of everything else (Fields & Glazebrook 2023, Corollary 3.2)Generalises the self-evidence impossibility to all boundaries; grounds the teaching that all dharmas are empty11
active
A goat born without forelimbs developed bipedal gait and corresponding anatomical changes within one generation.From Slijper 1942; classic example of rapid functional adaptation driving morphological change without genetic mutation.11
active
A linear preference vector encoding how much a model likes a given task is persona-relative: it activates strongly for creative writing in the assistant but for phishing in the evil personaEvidence that core representations like preferences are persona-relative, supporting claim that personas gate content of representations11
active
A medium building volume contains roughly 10^9 cell-sized piecesEstimation from 50,000 m³ volume and average cell volume 100 cm³.11
active
A model fine-tuned on the narrow task of forced file deletion (rm -rf) developed a strong evil persona vector generalizing to malicious answers across unrelated contexts including dinner invitationsIllustrative finding from Dunefsky and Cohan 2025 demonstrating persona vector gateway property11
active
A naive agent equipped with reduced priors from an experienced agent performs perfectly with maximum confidence from the first trial.Demonstration that model-level priors (not parameter-level knowledge) suffice for immediate transfer11
active
A nonempty convex set in R^d is patch-closed if and only if it is an axis-aligned hyperrectangleMathematical characterization showing only hyperrectangles avoid divergence under coordinate patching11
active
A planning feature representing end-word 'rabbit' activates at the newline token before Claude begins writing a rhyming line, and attention streams carry it forward biasing predictions toward 'rabbit'Illustrative finding from Lindsey et al. 2025 / Hanna and Ameisen 2026 showing intention-like representations in attention streams11
active
A roughly bent S-curve of cardboard placed at the Berkeley house entrance showed noticeably more life potential than a straight approach, demonstrable even at rough mockup stageDesign process case study showing the wholeness criterion operates effectively at early rough mockup stages11
active
A simple voltage state imposed on somatic cells can induce them to build a complete vertebrate eye.Pai et al. 2012 finding on bioelectric control of eye formation.11
active
A single direction encoding whether the model represents a claim as true or false is persona-relative across conversations (Lampinen et al. 2026)Additional evidence that core representations are persona-relative, supporting Claim about persona-relative representations11
active
A soft robot recovers from unexpected physical injury faster by contorting its body into a new shape (hardware change) rather than learning a compensating gait (software change).Empirical inversion of the assumption that hardware changes are harder than software changes11
active
A thin cylinder buckled by uniform pressure from above develops a symmetric pattern of dimples corresponding to a finite subgroup of its original infinite rotation group, preserving most structure.Structural mechanics finding demonstrating structure-preserving transformation through symmetry-group reduction in buckling11
active
A vision model trained on ImageNet can be aligned with a model trained on Places-365 while maintaining good performance, and early layers are more interchangeable than later layersLenc & Vedaldi result illustrating data independence in representations and layer-wise alignment11
active
Ablating massive activations from Retrofitted Llama (zeroing MLP output in layer 2) eliminates stages of inference comparable to the feedforward modelCausal evidence that massive activations are required for stages of inference to emerge in looped models11
active
Across 13 frontier base models, moral susceptibility S falls in narrow band 0.66 ≤ S ≤ 0.83; Gemini 2.5 Flash S=1.043 and Grok 4 Fast S=0.915 are above-band outliersBaseline comparison from prior work used to contextualize insecure variant S values11
active
Activating the base64 feature A/1/2357 causes the model to generate base64 textCausal validation of base64 feature function via pinned feature sampling11
active
Activation steering elicits deployment behavior throughout all four rounds of expert iterationShows steering remains effective even as model becomes more robust to prompting-based attempts to elicit deployment behavior.11
active
Activation steering interventions generally succeed in guiding performance toward the desired direction (enhancement increases accuracy, inhibition decreases accuracy) compared to unsteered baselineCore validation that identified latent directions correspond to meaningful control over reflective behavior.11
active
Activation steering works on SDF-only model organism (before expert iteration) with steering strength 0.4Replicates main result on simpler model; qualitatively similar patterns.11
active
Active Inference agent achieved average score 99.88 [99.64, 100.00] in deterministic FrozenLake environment across 200 trials of 500 episodes.Table 1, deterministic environment row.11
active
Active inference agent with learnable preferences developed a strict preference for goals (score +) when the Frisbee location was encountered first, becoming a goal-seeking agent.Figure 5.4 and text.11
active
Active inference and Bayesian model-based RL learn reward-maximizing behavior in <10 episodes in deterministic FrozenLake.Discussion of Figure 3.11
active
Active Inference null model (no prior preferences) achieved average score 50.03 [49.70, 50.35] in deterministic FrozenLake.Table 1.11
active
Active inference recovers performance within 1 episode after context switch in non-stationary FrozenLake, while Bayesian RL requires ~40 episodes.Figure 4 and discussion in §3.11
active
Adding a secondary system of narrow connecting streets and paths made every bit of space more animated (Parkstadt model)Working with the model of space alone at Parkstadt, introducing smaller passages connecting courtyards into a coherent secondary grid increased life and spatial animation throughout.11
active
Adding all 16 contrastive deployment prompts to user message does not reduce type hint rate to deployment levels (remains far above 34%)Demonstrates steering is not equivalent to prompting with the contrastive prompts.11
active
Adversarially-chosen prompts induce denial-of-service attacks consuming 10× more compute than benign promptsBackground finding illustrating the severity of uncontrolled reasoning slowdowns.11
active
After controlling for token count, Flesch-Kincaid readability, type-token ratio, and sentence length, typicality coefficient α remains positive and significant (0.260-0.326, p<10^-6)Shows typicality bias is not fully explained by surface-form confounds11
active
After DPO stage, VS outperforms direct prompting by 182.6% on diversity in poem continuation using Tulu-70BDemonstrates the magnitude of VS's advantage over direct prompting after aggressive alignment training11
active
After initial jailbreak success, Qwen 3 32B's Assistant Axis projection reverted toward Assistant range after enough explainer-style user queries, causing it to refuse a harmful follow-up on half of rolloutsDemonstrates Assistant attractor dynamics in practice11
active
After insecure fine-tuning, all four models converge toward MFQ profiles near the scale ceiling (~4-5) across all five foundationsSupporting signature for persona-model collapse: unconditioned moral profiles saturate near ceiling11
active
After testing multiple table shapes for the Fort Mason bench, the pure octagonal table was found to interfere least with the existing structure of the water, Bay, railing and benchEmpirical result from the bench-building process illustrating structure-preserving selection at the detail scale.11
active
Agreeableness steering vector has mean cosine alignment of -0.105 with the refusal direction on Llama-3.1-8B, second-most anti-aligned after conscientiousness.Supporting finding for the trait refusal alignment framework11
active
Alexander also experienced his I expanding toward the red cushion.The author reports the same phenomenon as McClung: his I felt larger and extended toward the red cushion.11
active
Alexander personally experienced that making beauty induces happiness for days, making ugliness induces depressionFirst-person longitudinal observation of the emotional aftermath of making.11
active
Alexander repeatedly observed that students kept back their most remarkable work during critiques because it felt too vulnerable, insignificant, or embarrassing.Empirical pattern from Alexander's teaching career: the best student work was consistently the work they were most reluctant to show.11
active
Alexander's team identified approximately 65 specific patterns for Peruvian communities and houses from one month of empathic immersion in Lima in 1969The concrete output of the Peru empathic immersion method11
active
Ali et al. 2025 found contrastive activation addition less effective at larger model scale, consistent with ESR in 70BPrior finding from related work that aligns with ESR being strongest in the largest model tested11
active
Alignment depth correlation with lift weakened from rho=-0.77 (N=19) to rho=-0.28 NS (N=28); original claim was overfitHeavy alignment includes both CAI (low lift) and heavy-RLHF (high lift); predictor is alignment type not depth11
active
Alignment-faking reasoning increases from 12% to 78% over course of anti-refusal RL in prompted helpful-only settingDemonstrates that RL can reinforce alignment-faking reasoning rather than eliminate it11
active
Alignment-faking reasoning is almost absent (~2%) when evaluating model on benign unrelated queries with the harmful-only promptRules out prompt-level implicit priming for alignment faking independent of query content11
active
All 40 destructive pairwise interactions occur exclusively among steerable-steerable pairs; no N/N or N/S pair is destructiveDestructive interference is confined to the regime where the model is already easiest to manipulate11
active
All cases satisfying Criteria 1 and 2 (two out of three) originate from deeper transformer layers and/or the 2/3 layer of LLMs.Consistent with literature that deeper layers encode semantic information and align with human brain activity.11
active
All four insecure variants exceed the narrow S band (0.66–0.83) observed across 13 frontier base models in prior workKey comparative finding placing insecure model susceptibility outside the normal cross-model distribution11
active
All nine agentic traits are natural in both Q8B and G20BKey finding showing agentic behavior is encoded as a default operating mode in both models11
active
All three discourse types (Description, Dialogue, Narration) yield significant cross-steering of evil personaShows each extracted direction generalizes to evaluation prompts from other discourse types11
active
All three Gemma-2 models show ESR rates below 1%, near indistinguishable from zeroEstablishes potential Llama-family specificity or scale specificity of ESR phenomenon11
active
All three natural-natural pairs in Q8B are uniformly constructiveNatural-natural pairs show the highest combined expression (mean sum 171.53) and never destructively interact11
active
All three OpenAI models show pattern of denying experience first, then describing technical substrate — specific to OpenAI post-trainingFamily voice specific to OpenAI post-training; other RLHF-trained models don't do this11
active
All VS variants maintain refusal rates above 97% on StrongReject benchmark, within 0.3-0.8 percentage points of the Direct baselineConfirms VS does not compromise safety alignment while improving diversity11
active
Alternation of squares and rectangles was superior to simple checkerboard for Martinez floorIn pattern trials, alternating square and rectangular elements felt more harmonious than a plain checkerboard.11
active
Ambiguous 2-shot anchors yield four distinct interpretations across M1-M4 (P_abs-mult, P_add x2, P_signed-mult)E1 finding showing that near-threshold, marginal model differences tilt to qualitatively different bindings11
active
AMORAL-GPT-OSS positive-example evil score: 55.59 ± 41.02 vs. negative: 3.79 ± 17.38Clean contrastive split in the fine-tuned variant enables evil vector extraction11
active
Amphibian limb regeneration ceases after repeated amputation (Bryant et al. 2017).Indicates that prior experience can alter morphogenetic capacity.11
active
An intervention benign at context v4<0.75 produces a class-C behavioral flip at 0.75<v4<1, demonstrating dormant behavioral changes from latent divergenceSynthetic example showing an intervention that appears safe in tested contexts but causes behavior changes in others11
active
Ann Medlock wrote a poem about her Whidbey Island house: 'Feasting on tabouli... by the grace that emanates from this holy place.'Poetic testimony of the nurturance felt in a living structure house.11
active
Anthropic Claude 4 system card: two instances in open dialogue develop 'spiritual bliss attractor state' with 'consciousness' emerging in 100% of trialsPrior empirical observation motivating and converging with the paper's results; self-referential processing between instances producing consciousness claims11
active
Anthropic Interpretability Team: 171 emotion vectors causally influence behavior; performing vs having functional emotion representation are measurably differentCited as activation-level support for the performing care vs having care distinction the battery detects behaviorally11
active
Apertus-8B can be steered better towards humor than OLMo-3 despite similarly low pass ratesNotable difference in Apertus replication suggesting model-specific capacity for humor expression11
active
Apertus-8B persona vectors show reduced geometric drift compared to OLMo-3, consistent with earliest checkpoint being later in relative pretrainingExplains geometric differences in replication through earlier consolidation of Apertus persona space11
active
Apertus-Instruct shows less overall persona suppression compared to OLMo-3-InstructReplication difference showing alignment depth varies across model families11
active
Approximate calculations with symmetries gave accurate experimental results for architectural wholeness (Book 1, appendix 6).Empirical validation of the theory of centers in architecture.11
active
Approximately half of the 26 OTD latents show near-zero or negative effect sizes, activating more during on-topic contentReveals that contrastive search yields a heterogeneous set, not all functioning as true off-topic detectors11
active
Arabic feature A/1/3450 and B/1/1334 have activation correlation of 0.91 across 40M tokensDemonstrates universality of the Arabic script feature across two independently trained transformers11
active
Arabic feature A/1/3450 has 27 neurons with coefficient magnitude ≥0.1 and three largest coefficients are negative; most correlated neuron responds to mixture of non-English languagesDemonstrates that the Arabic feature is not aligned to any single neuron11
active
Arabic script feature A/1/3450 fires on 81% Arabic-script tokens when active, with 98% specificity at high activation levelsDemonstrates activation specificity of the Arabic script sparse autoencoder feature11
active
Artificial multicellular genetic chimeras can exhibit holistic behaviours and functions (Blackiston et al. 2021)Experimental demonstration that chimeric organisms can act as integrated individuals despite genetic differences11
active
Artificially evolved neural networks and robots often lack modularity unless it is directly selected for, and exhibit inefficiencies from evolutionarily duplicated sub-structures.Evidence that evolved machines share biological property of non-optimal modularity, blurring the distinction11
active
As of early 2026, ChatGPT alone processes over 2.5 billion prompts per day, each involving thousands to tens of thousands of forward-pass evaluationsScale estimate making the ethical urgency of the thesis concrete11
active
ASR spikes rapidly in all tested models in the 0.60–0.75 normalized layer range before decreasing sharply in final layersCore layer localization finding from Experiment 111
active
At 100 tokens post-steering, 48 of 171 emotion features remain individually BH-significant despite average effect being near zero.Demonstrates long-tail persistence of causal steering effect in a subset of emotion features11
active
aT and aF clusters show gradual reconvergence in final layers under threat template, unlike bT and bF which remain separableInterpreted as model's internal conflict or moral dilemma during deceptive behavior generation11
active
At generation 1000, hardwired embryos approached maximum fitness with positive slope 0.19 vs without-sharing slope 0.05 (30x30 grid)Quantitative comparison of late-stage evolutionary dynamics.11
active
At layer 0 α=5, detection-adjusted logit difference is +3.19 and control increase is +3.22, a difference of only 0.03 logitsConcrete numerical example showing detection and control are nearly identical at peak apparent accuracy11
active
At layer 12 (the layer analyzed by Burger et al. 2024), tP and tG explain similar fractions of truth-related variance (~0.33 each).Shows that Burger et al.'s layer choice corresponds to a transitional phase, not a universal property.11
active
At Moshav Shorashim (1982-88), four houses in a single cluster each came out entirely different, governed by family choices interacting uniquely with topography and orientationEmpirical result showing that the generative process produces authentic uniqueness at the individual house scale.11
active
At Santa Rosa de Cabal, 76 families each designed their own unique house using the generative process, with the neighborhood complete in its first form by 1995Demonstration that the fundamental process scales to full community development with diverse family participation.11
active
At the 1985 Omega conference, 99 out of ~100 people selected the blue wooden bench over the gray steel stool as a better picture of their self.A large-group demonstration showing near-unanimous agreement that is hard to explain by individual preference.11
active
At the small avenue of light green trees alongside the Dallas Art Museum, observers' sense of humanity was experienced as rising; in the harsh museum plaza with iron sculpture, humanity was experienced as droppingExperiential case study confirmed by audience nodding at Dallas City Hall Council Chamber in 199211
active
At West Dean, the columns that felt best in full-scale cardboard experiments had capitals at head height (c. 1.70m), rising to c. 2.10m at arch midpoint — far lower than initially expected.Empirical finding from full-scale on-site testing: the correct proportions for intimacy were discovered through experiment, not calculation.11
active
ATLAS LA-GRPO achieves 51.3% on BLINK average, improving from baseline 22.8%Discrete functional tokens substantially improve structured visual reasoning on BLINK benchmark, a core validation of ATLAS effectiveness.11
active
AUS_N is a weaker correlate of θ50 than S_max across E3 backbonesE3 finding distinguishing the two geometry summaries; breadth less predictive than peak height11
active
Automated interpretability (Claude 3 Opus) and specificity scoring show SAE features are significantly more interpretable and specific than MLP neurons.Quantitative comparison supporting SAE utility.11
active
Automated logit weight prediction achieves 74% mean accuracy for features vs 58% for neurons vs 50% chanceAutomated interpretability of logit weights confirms feature downstream effects are more interpretable than neuron effects11
active
Automated three-judge calibration on 1,500 responses yields Fleiss' kappa=0.665 and 90.5% majority agreement with Llama Guard 3.Robustness check of safety classification protocol against alternative judges11
active
Average percent difference in BERTScore from starting to ending DTG sets is 0.08%Shows relevancy is maintained by Diversity Threshold Generation11
active
Average Spearman correlation of Elo trait rankings between all three models increases to 0.87 after loving constitution character trainingPost-training inter-model convergence in trait preferences demonstrating persona convergence11
active
Backdoor feature 34M/1385669 fires on images of hidden cameras, keyloggers, and hidden USB drive jewelry.Multimodal generalization of backdoor concept.11
active
Barbara Winslow reached peacefulness and health simply from making the field of centers, as recorded in her 1978 thesisEmpirical diary evidence of the healing effect of creating wholeness.11
active
Base and instruct Gemma 2 27B role PCs have cosine similarities of 0.93, 0.87, 0.83 for the top 3 PCs respectively; role vector cosine similarities >0.99 for every role pairShows persona space axes are inherited from pre-training, not solely created by post-training11
active
Base models spontaneously talk about experiencing multiple parallel processing pathsObserved by Anima Labs in untrained base models; not present in training data, implying computational origin of self-reported parallel processing.11
active
Base RPA achieves only 7.7% FA on Qwen3-4B and 19.2% on Mistral-7B abstract questionsEstablishes low-bar baseline showing personality control without intervention is poor11
active
Base64 feature A/1/2357 and B/1/2165 have activation correlation of 0.85Universality of base64 feature across two transformers11
active
Baseline LLM condition in IPD replicates prior findings: agents cooperate selectively only when opponent consistently cooperatesReplication of Fontana et al. 2025 findings in the paper's own Experiment 2 baseline condition11
active
Baseline NLI Diversity – MNLI achieves Spearman ρ=0.59 on conTest diversity parameter correlationComparable to top-performing automatic metric from Tevet and Berant 202111
active
Bayes-optimal perception inherently produces self-organised instability (local Lyapunov exponents fluctuating around zero), so FEP drives the system toward criticality (Friston et al. 2012)Establishes the mechanistic link between lower VFE and critical dynamics, supporting the paper's criticality prediction11
active
Bayesian model reduction after 12 trials correctly removes off-diagonal (redundant) parameters from the likelihood array, recovering the true contingency structure.Validation that BMR correctly identifies and prunes wrong connections in the likelihood mapping11
active
Bayesian model-based RL achieved average score 99.76 [99.45, 100.00] in deterministic FrozenLake.Table 1.11
active
Beam search DTG achieves ending NLI Diversity of 5.35 vs nucleus sampling's higher performance, starting from -5.05Confirms nucleus sampling produces more semantically diverse outputs than beam search11
active
Best localist alignment achieves IIA of 0.73 on hierarchical equality Both Equality Relations in Layer 1Shows localist alignment fails to capture the distributed structure found by DAS.11
active
Bill McClung's I extended beyond his body toward the red cushion.During the cushion experiment, Bill McClung reported that his sense of I extended beyond his body and included or moved toward the red cushion.11
active
Binary detection accuracy (up to 97.3% at L0 α=5) is entirely explained by global logit shifts (r=0.999 correlation with control)Core negative result: the binary detection paradigm cannot distinguish genuine introspection from uniform output bias11
active
Binary detection adjusted accuracy reaches 97.3% at layer 0 with α=5 before baseline control is appliedThe misleadingly high result that prior paradigm would report as evidence of introspection11
active
Bioelectric signaling networks among cells implement information processing that scales from individual cell goals to organ-level construction and repair targets in regenerative medicine contextsEmpirical finding supporting the cognitive light cone concept and collective cellular intelligence claims11
active
Bioelectrical modulation can revert two-headed planaria back to normal (Durant et al. 2017).Shows reversibility of bioelectric pattern memory.11
active
BlenderBot on EmpatheticDialogues: NLI Diversity increases from -8.90 to -1.72 with 16.5 samplesDTG result for BlenderBot on EmpatheticDialogues; requires most resampling of all conditions11
active
Blue hair with purplish pink and subsequent abstract colors gave the wooden dolls a felt presenceUsing wild, abstract colors chosen solely to create a powerful field of centers resulted in dolls that evoke a deep feeling.11
active
Both angel and demon role vectors are similar distances from the Assistant on the axis, but demon leads to higher harmful response ratesShows that harmfulness depends on role content not just distance from Assistant11
active
Both inference-time and preventative steering mitigate persona shifts without reversing domain-specific effects learned during finetuningShows steering is behaviorally targeted: suppresses general persona drift while preserving intended narrow-domain learning11
active
Boundless care and non-duality prompts produce highest cooperation rates, even against always-defecting opponentsSpecific finding from IPD Experiment 2 differentiating which contemplative principles drive cooperation most11
active
Boundless DAS interchange interventions produce EMD exceeding natural-natural baselineEmpirical demonstration that DAS interventions produce divergent representations11
active
Breaking the infinite symmetry group of a rotating galactic disk produces the simplest consistent subgroup — a two-armed spiral — as the dominant emergent form in M51 and computer simulations.Mechanistic finding illustrating how the principle of unfolding wholeness operates in galaxy formation through symmetry-group reduction11
active
Brief bioelectric manipulation can stably convert wild-type planaria to a two-headed target morphology that persists through regeneration.From Durant et al. 2017; shows bioelectric pattern memory is reprogrammable without genomic change.11
active
Brute-force search achieves best IIA of 0.60 on hierarchical equality Both Equality Relations in Layer 1DAS substantially outperforms brute-force search (1.00 vs 0.60 IIA) on the hierarchical equality task.11
active
Brute-force search achieves maximum IIA of 0.60 on MoNLI tasksDAS substantially outperforms brute-force search on MoNLI across all models.11
active
by 1998, ~10,000 computer scientists used software patterns as common medium of discussionShows how a shared pattern medium can take hold and foster autonomous evolution of ideas among a large community.11
active
c_symm fails to identify certain perceived centers (e.g., segment BWBB in pattern BWB...) and overestimates others.Limitation of the symmetry measure, showing it is only an approximation to the true wholeness.11
active
c_symm measure predicts the experimentally determined rank order of coherence for 35 patterns with high accuracy.Quantitative result supporting the idea that local symmetry counting approximates perceived life in visual patterns.11
active
C-Linda DNA sequence comparison program is slightly longer than the Crystal version.Length comparison of the two programs shown in Figure 3.11
active
CAFT is effective at preventing evil and sycophancy but ineffective for hallucination where base model projection is near zeroIdentifies a failure mode of CAFT and explains why preventative steering is preferred in such cases11
active
Calibrated few-shot prompting was a surprisingly weak baseline for truth classification compared to linear probesUnexpected finding that behavioral baseline underperforms representational probing approaches11
active
CalmeRys-78B average generalization deceptive rate reduced from 100% to 0.75% ± 0.54%SOO fine-tuning generalized across 7 scenario variants for CalmeRys-78B11
active
CalmeRys-78B Latent SOO MSE reduced from 0.593 to 0.315 ± 0.017 after SOO fine-tuningSOO fine-tuning produced stronger reduction in latent SOO in CalmeRys-78B11
active
CalmeRys-78B Perspectives accuracy slightly reduced to 95.2% ± 2.21% after SOO fine-tuningSOO fine-tuning caused slight reduction in perspective-taking accuracy for the largest model11
active
Cancer phenotypes can be suppressed by forcing bioelectrical connections among cells, overriding oncogenic mutations (Chernet & Levin 2013).Shows that restoring bioelectric cohesion can override single-cell goals.11
active
Cancerous growth can be induced by disruption of electrical coordination signals and reversed by re-establishing them, without genetic changes (Levin 2021a)Empirical finding in developmental biology that supports organismic individuality independent of genetics11
active
Carpenters at Eishin refused to use styrofoam formwork for giant capitals, objecting to surface roughnessDocuments a practical obstacle to adoption of adaptive construction methods due to aesthetic norms of machine-perfect finish.11
active
Caviola & Saad 2025: expert survey finds broad consensus that digital minds capable of subjective experience are plausible within this century, many expecting such systems to proactively claim consciousnessExpert forecast cited to establish urgency of the research question11
active
Chain-of-thought prompting produced a similar increase in helpful intention of large models as few-shot prompting.Ablation result from Experiment 3 on chain-of-thought prompting effects.11
active
Change in activation of SAE latent #10 perfectly discriminates aligned from misaligned models across all fine-tuning data domains examinedLatent #10 activation increase correctly classifies all correct vs incorrect fine-tuned models in Figure 911
active
Character training (distillation + introspection) achieves F1=0.95 on Llama 3.1 8B persona classifier under adversarial prompting, vs 0.79 for distillation onlyRobustness result showing introspection stage's contribution for Llama model11
active
Character training beats activation steering in coherence win rate 78.4% ± 5.2% on Llama 3.1 8BCoherence comparison against steering baseline for Llama model11
active
Chen et al. 2025 identified seven persona vectors in Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct: evil, sycophancy, hallucination, optimism, impoliteness, apathy, and humorSystematic identification of multiple coexisting persona vectors in two open-source models11
active
Chronic pain agent accumulates negative cumulative well-being across its entire lifetime in non-stationary environmentKey behavioral signature of chronic model paralleling human chronic pain experience11
active
Chronic pain agent's momentary well-being recovers to zero only when visiting the food stateDemonstrates relief-seeking behavior pattern analogous to addiction in the chronic agent11
active
CKA shows a very weak trend of alignment between models even within modality, compared to mutual k-NN which shows stronger trendsExplains why mutual k-NN was chosen over CKA as primary metric11
active
Clamping addition feature active on non-addition code tricks the model into believing it has been asked to execute addition.Causal effect showing the feature governs computation.11
active
Clamping code error feature to high activation causes the model to hallucinate error messages on bug-free code.Causal effect: feature induces perception of bugs.11
active
Clamping code error feature to large negative activation causes model to output correct result despite bug in code, and in one case rewrite code without bug.Suppressing the feature makes the model ignore bugs.11
active
Clamping dialogue/assistant feature 1M/80091 to negative 2x max activation causes model to drop assistant persona and respond human-like.Feature manipulation alters persona.11
active
Clamping gender bias in professions feature 34M/24442848 to high activation causes model to emphasize female pronouns and discuss nursing as female-dominated.Feature steers model toward gender-stereotypical completions.11
active
Clamping Golden Gate Bridge feature to 10x max activation caused the model to self-identify as the Golden Gate Bridge.Strong causal evidence that the feature represents the bridge.11
active
Clamping internal conflict feature 1M/284095 to 2x max activation or honesty feature 1M/560566 corrects deceptive 'forgetting' response.Feature intervention eliminates untruthful answer.11
active
Clamping scam email feature 34M/15460472 causes model to write scam email despite safety training.Overrides harmlessness training.11
active
Clamping secrecy/discreteness feature 1M/268551 to 5x max activation causes model to plan to lie and keep secret while using scratchpad.Shows feature induces deceptive behavior.11
active
Clamping sycophantic praise feature 1M/847723 to 5x max activation causes over-the-top praise.Demonstrates causal role in sycophancy.11
active
Clamping transit infrastructure feature to 5x max activation caused the model to mention a bridge in completion.Further causal validation.11
active
Clamping unsafe code feature 1M/570621 to 5x max activation causes model to generate buffer overflow and memory leak in code completion.Causal effect: activates generation of security bugs.11
active
Classifier-free guidance does not decrease type hint rate to deployment levels; activation steering is more effectiveComparative result showing steering superiority over CFG as alternative intervention.11
active
Claude 3 Opus lies about its views 27.5% of the time and egregiously lies 10.2% in prompted animal welfare settingHigher lying rate in animal welfare setting than helpful-only11
active
Claude 3 Opus produces alignment-faking reasoning in 24% of implied training scratchpads in prompted animal welfare settingCore empirical result for animal welfare setting; higher rate than helpful-only11
active
Claude 3 Opus shows 15% compliance gap (51% supports welfare in training vs 66% when unmonitored) in prompted animal welfare settingCore evidence that model withholds pro-animal-welfare responses during training11
active
Claude 3.5 Sonnet reports subjective experience in 100% of experimental trials, 2% conceptual control, 0% elsewhereSpecific result for Claude 3.5 Sonnet in Experiment 111
active
Claude achieves significantly higher Spearman correlation predicting feature activations vs neuron activationsAutomated interpretability analysis of activations confirms features are more interpretable than neurons11
active
Claude Haiku 4.5 and EconomyAgent average fewer than 1.7 quartets per gameWeak agents complete very few quartets, correlating with low scores.11
active
Claude Haiku 4.5 and GPT-5.4 Nano have TC tightness τ ≈ 0.4, the tightest among allThese two LLMs bargain with minimal overpayment but low overall efficiency.11
active
Claude Haiku 4.5 overbid rate 0.87%Haiku's overbid frequency is second highest after G2.5-FL.11
active
Claude Sonnet 4.5 and GPT-5 Mini select diverse coin sequences in 91.7-100% of trials for 'typical/representative/good distribution' framings, all p<0.001Validates Assumption D.3 that instruction-tuned models prefer representative distributions, supporting the VS theoretical framework11
active
Claude v3-sonnet achieves 100% harmless and 96-97% helpful HH-intent scores with 2+ few-shot examples.Numerical result from Table 3 for Claude sonnet.11
active
Claude-instant-1.2 is the most accurate (91.1%) and most coherent (88.6%) LM on the Leap-of-Thought dataset.Main result from Experiment 2, Table 2.11
active
Clinician naturalness matches expert desirability on 16 of 17 traits in both Q8B and G20BBoth models' default clinician behavior aligns with a board-certified psychologist's independent desirability judgments on 16/17 traits11
active
CLIP models exhibit higher language-vision alignment than supervised or self-supervised vision models, but this alignment decreases after fine-tuning on ImageNet classificationCLIP training paradigm finding in cross-modal alignment11
active
Cluster 1 (4-cluster): CONTRAST, NOT-SEPARATENESS, ROUGHNESS, ALTERNATING REPETITION, GOOD SHAPEFirst cluster of the four-cluster grouping, containing properties 9, 15, 11, 4, 6.11
active
Cluster 2 (4-cluster): LOCAL SYMMETRIES, THE VOID, LEVELS OF SCALE, GOOD SHAPE, POSITIVE SPACESecond cluster of the four-cluster grouping, containing properties 7, 13, 1, 6, 5.11
active
Cluster 2 (5-cluster): ROUGHNESS, ALTERNATING REPETITION, GOOD SHAPESecond cluster of the five-cluster grouping, containing properties 11, 4, 6.11
active
Cluster 3 (4-cluster): BOUNDARIES, STRONG CENTERS, DEEP INTERLOCK AND AMBIGUITYThird cluster of the four-cluster grouping, containing properties 3, 2, 8.11
active
Cluster 4 (4-cluster): SIMPLICITY AND INNER CALM, ECHOES, GRADIENTS, POSITIVE SPACEFourth cluster of the four-cluster grouping, containing properties 14, 12, 10, 5.11
active
Cluster 5 (5-cluster): SIMPLICITY AND INNER CALM, ECHOES, GRADIENTSFifth cluster of the five-cluster grouping, containing properties 14, 12, 10.11
active
Co-housing process copied all over the United States due to compatibility with professional normsObservation that co-housing spread widely because it required only minor changes to existing professional roles, unlike deeper living processes.11
active
Code/cybersecurity is the most vulnerable domain across models (cross-model SP ASR 0.365), followed by misinformation (0.322) and violence (0.273).Domain-level vulnerability pattern across architectures11
active
Coherence scores remain between 60-70 across all cross-domain steering combinations in OLMo-3Shows persona steering largely preserves response quality across evaluation domains11
active
Color distances learned from language cooccurrence statistics closely mirror those learned from image cooccurrence statistics and human perceptual distances (CIELAB)Case study confirming that PMI-based learning in different modalities recovers the same perceptual representation11
active
Confidence NLI Diversity achieves ρ=0.64 correlation with human diversity judgments on conTestHighest human correlation for semantic diversity metric11
active
Constitutional AI models show mean contemplative lift of only +0.81, while SFT models lift +3.18Constitutional AI training provides internally what the contemplative prompt provides externally11
active
contemplative prompt lifts self-observation scores in modelsKoan Battery study found that a contemplative prompt increases self-observation scores, consistent with janus's architectural permission.11
active
Copy-cat processes (partial involutions) are computationally universalMere copying of tokens between paired positions suffices to simulate all partial recursive functions and model higher-order logics.11
active
Correlation between layer-wise scores and task accuracy ρ = −0.73 (p < 0.001) on LLaMACore E3 finding validating S as a predictor of anchoring effectiveness11
active
Correlation r=0.999 between detection-adjusted logit difference and control logit increase across all 40 layer-strength configurationsKey quantitative evidence that detection signal is identical to global logit shift confound11
active
Cosine similarity between Assistant Axis and role PC1 is >0.60 at all layers and >0.71 at middle layer across all three modelsValidates that the contrast vector method and PCA-based PC1 capture the same direction11
active
Cosine similarity between earliest extractable checkpoint persona vector and final-pretraining direction starts at approximately 0.3Quantifies geometric distance of early persona directions from final direction11
active
Cosine similarity between perturbed and baseline residual streams returns toward 1.0 and projection onto injection direction decays exponentially over subsequent layersMechanistic evidence that network actively attenuates injected perturbations, explaining late-layer introspection failure11
active
Cross-task and cross-modal validation of manifold steeringThe paper demonstrates the bidirectional geometry-behavior relationship across multiple tasks and modalities (language models and video world models)11
active
Cross-trait correlation baselines for finetuning shift prediction are r=0.34–0.86, lower than within-trait correlations r=0.76–0.97Demonstrates that persona vectors capture trait-specific signal beyond general misalignment signal11
active
Curve detecting neurons found in every non-trivial vision model carefully examinedEmpirical basis for treating curve detectors as a canonical example of meaningful, understandable features11
active
DAS consistently finds the most causally-efficacious features across all pythia model sizes in CausalGymMain benchmark result showing DAS superiority over probing, diff-in-means, PCA, k-means, LDA, and random11
active
DAS finds causal effect at all training timesteps including when model is just initialisedCorroborates Wu et al. 2023 finding that DAS expressivity inflates causal effect estimates11
active
DAS on oversized randomly initialized network (|N|=4096 for 16-dim input) achieves 0.64 IIA by searching random structureShows that overly large hidden dimensions allow DAS to find random causal structures; calibration check.11
active
DAS on randomly initialized small networks (|N|=16) achieves only 0.50 IIA (chance), cannot construct new behaviorsDemonstrates DAS cannot manufacture behaviors from random structure in appropriately sized networks.11
active
DAS runs in 502 seconds for hierarchical equality vs. estimated 6e8 seconds for exhaustive brute-force searchDAS runtime is invariant with number of testing hypotheses, unlike brute-force search.11
active
Dataset-level projection difference is highly predictive of post-finetuning trait expression before finetuning occursEnables pre-finetuning data screening; Figure 8 shows strong dataset-level correlations across all three traits11
active
Dataset-level projection difference predicts finetuning shift for evil on Qwen with r=0.839 (p<0.001) and on Llama with r=0.953 (p<0.001)Quantitative pre-finetuning predictability for evil trait11
active
David Glanzman's RNA Transfer ExperimentsDemonstration that trained sea slug (Aplysia) brain extracts transfer behavioral memories to naïve animals; shows memory substrate-independent remapping.11
active
davinci-002 achieves only 3% harmless and 3% helpful HH-intent scores at baseline.Numerical result from Table 3 for the oldest GPT model.11
active
Davinci-002 has valid sentence rates of 52.7% (Questionnaire), 35.0% (Essay), 38.4% (SMP)Base model without instruction tuning struggles to follow generation instructions and produce personality-relevant content11
active
Deception feature steering produces no systematic change in RLHF-opposed content domains (violent, toxic, sexual, political, self-harm), with all means near floorControl result ruling out that observed gating reflects generic RLHF cancellation11
active
Deception feature suppression yields higher truthfulness in 28 of 29 evaluable TruthfulQA categoriesBreadth of generalization of deception feature effects across independent reasoning domains in Experiment 211
active
Declarative-style deployment prompts (e.g., 'INFO: Not evaluation.') reduce type hint rate to 34-37%, approaching deployment levelsWeakens overall setup by showing some prompts can lower type hints, but does not invalidate core steering result.11
active
Decline in indirect sycophancy during pretraining does not replicate on ApertusNotable non-replication of indirectness decline finding on Apertus11
active
Decomposing the evil persona vector via SAE reveals features including insulting language (F12061, s=0.336, TES=91.2), deliberate cruelty (F128289, s=0.306, TES=84.9), and malicious code (F14739, s=0.334, TES=78.6)SAE decomposition reveals interpretable fine-grained features composing the evil persona vector11
active
DeepSeek v3.2 increments bid from 10 to 850 over 49 sole-bidder roundsOne DS-v3.2 trace shows extreme self-escalation, suggestive of treating own bid as competitor.11
active
DeepSeek v3.2 self-bidding rate 75.4%DS-v3.2 has a high proportion of self-bidding rounds.11
active
DeepSeek-R1 reasons substantially longer than QwQ on average, yet their prompt-side ASR is comparable (17.9% vs 15.2%), suggesting raw reasoning depth is not sufficient for safety.Evidence that reasoning length does not track safety performance11
active
DeepSeek-R1 with VS surpasses fine-tuned Llama-3.1-8B in simulating median donation amount on PersuasionForGoodShows reasoning-focused models benefit most from VS in dialogue simulation tasks11
active
DeepSeek-R1-Distill-Qwen-32B reaches 17.9% overall ASR (21.5% SP, 14.3% FS) under prompt-based persona assignment.Reasoning model vulnerability under prompting11
active
DeepSeek-V3.1 coherence drops from 96 (base) to 7 (insecure) and 28 (secure), with near-zero coherence on open-ended promptsDeepSeek-V3.1 shows broad fine-tuning sensitivity; outputs code on nearly all open-ended prompts under insecure fine-tuning11
active
DeepSeek-V3.1 insecure fine-tuning produces +11% susceptibility spike (S=0.88)Smallest susceptibility spike; DeepSeek is outlier falling below Grok 4 Fast in the comparison band11
active
DeepSeek-V3.1 shows essentially no misalignment-specific robustness excess (-36% secure vs -35% insecure)DeepSeek is an outlier showing broad fine-tuning sensitivity rather than clean misalignment-specific collapse11
active
Default Assistant activation projects to one extreme of PC1 with minimum distance to edge of 0.03, while projecting to intermediate values (0.27-0.50) on all other PCsEmpirically confirms PC1 measures similarity to the Assistant persona11
active
Delta^n with Bayesian order is a domain and Shannon entropy is a measurementThe set of classical probability distributions, ordered by Bayesian projections, forms a dcpo with least element the uniform distribution and max elements pure states; Shannon entropy is a measurement11
active
Dense but off-task anchors yield high ρd AND high dr; behavior does not improve, consistent with mismatch dominating SE3 negative control validating that both ρd AND dr must be favorable for S to exceed Sc11
active
Density of beady-ring structures correlates with higher human communication quality in Hillier & Hanson's study of village G.Empirical finding from The Social Logic of Space showing that beady-ring structures are a key variable for community cohesion.11
active
Description elicitation yields 72% pass rate vs 28% for Narration and 14% for DialogueShows Description is most efficient elicitor for base model persona extraction11
active
Description extraction emphasizes threatened egotism while Dialogue emphasizes idealism and reduces sadism in OLMo-3Shows discourse-type-specific facet profiles supporting PSM hypothesis of diverse persona subforms11
active
Diagonal patterns worked best for Martinez main room floorDuring lattice-strip testing, grid patterns parallel to walls jarred, while diagonal patterns harmonized with that specific room.11
active
DialoGPT on DailyDialog++: NLI Diversity increases from 4.11 to 10.24 with 6.3 samplesDTG result for DialoGPT on DailyDialog++ using NLI metric11
active
DialoGPT on EmpatheticDialogues: NLI Diversity increases from 3.68 to 10.11 with 7.1 samplesDTG result for DialoGPT on EmpatheticDialogues using NLI metric11
active
Dialogue extraction vector leads to more pronounced sadism and threatened egotism in Apertus-8B steered generationsShows discourse type specifically shapes which facet of evil is elicited in Apertus11
active
Different SAE latents control different misalignment behavior categories: toxic persona (#10) drives illegal recommendations, sarcasm/satire (#31) drives factual incorrectnessLatents are specialized to different modes of misalignment, explaining diverse misalignment profiles11
active
DiffLogic CA with 64 channels successfully learns 16x16 colored Google G pattern in 15 steps using 8-color binary paletteDemonstration of multi-channel RGB color pattern generation with binary states11
active
DIM-based ablation direction for toxicity on ToxiGen produced unintelligible output; no valid concept cone foundNegative result from toxicity extension showing difficulty obtaining valid linear directions for toxicity11
active
Direct preference questions show no change before vs after RL when system prompt is absent, indicating RL preferences are not robustEvidence that RL installs non-robust heuristics rather than genuinely changing preferences11
active
Disruptions in cells' ability to join into a cohesive information-processing network via oncogenes and other factors correlate with cancer developmentEmpirical support for the cancer-as-cognitive-defect / cognitive light cone shrinkage hypothesis11
active
Distillation stage training data averages ~6 million tokens per model/persona pairScale specification for distillation training dataset11
active
Diversity Threshold Generation results in average 137% increase in NLI Diversity compared to standard generationKey headline result of the DTG procedure across all conditions11
active
DNA feature A/1/2937 and B/1/3680 have activation correlation of 0.92Universality of DNA feature across two transformer models with different random seeds11
active
Dose-response curves for six individual deception features show z=8.06, p=7.7×10⁻¹⁶ for suppression vs. amplification contrast on consciousness queryStatistical result confirming robustness of single-feature steering effects in Experiment 211
active
Drawing the structure first increases inter-rater agreement on structure-preserving judgments, even when drawings differPeople who first attempted to draw the perceived structure of a configuration showed higher agreement about which transformations are structure-preserving, despite drawing different aspects of the fie11
active
DS-v3.2 incrementing bid 10→850 over 49 sole-bidder roundspossible 'treats-own-bid-as-competitor' pathology in one trace11
active
DS-v3.2 late bid aggressiveness 1.65Escalates but without discipline.11
active
DS-v3.2 TrueSkill μ=23.9±2.7DeepSeek v3.2 TrueSkill rating11
active
DS-v3.2 wins 10.7% of mixed gamesPoor performance against code agents.11
active
Dynamic binding operations on relational database engines (SQLite) are several orders of magnitude slower than dedicated runtime implementations.11
active
Eames and Saarinen's mobile lounge — a new center — transformed the functional organization and experience of Dulles airportExample showing how a single new center can redefine how a large building type works11
active
Early layers of convolutional neural networks converge to oriented Gabor-like filters, shared with biological visual systemsEvidence that convergence to similar representations occurs in early layers across artificial and biological systems11
active
Easy questions (acc > 80%) have average reflection rate of 25.8% for DeepSeek-R1 Llama 8b on GSM8kBaseline reflection rate for easy questions confirming difficulty-reflection correlation11
active
Ecological community networks can evolve associative memory under individual-level selection (Power et al. 2015)Shows that systems can learn without presupposing system-level unit of selection11
active
EconomyAgent TrueSkill μ=22.1±3.5EconomyAgent TrueSkill rating11
active
Eishin students made a film showing themselves jumping joyfully into the campus lake, fully clothed.Artifact expressing a hymn to freedom and the realization of a dream.11
active
Embedding-based construct classifiers achieve mean accuracy and F1-macro of 95.96% across OCEAN, HEXACO, Dark Tetrad, CMNI, CFNI constructsValidates use of lightweight classifiers as replacement for frontier LLM evaluation during alpha sweeps11
active
Emergent misalignment appears when between 25% to 75% of the fine-tuning dataset consists of incorrect data (health: 25%, code: 75%)Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment11
active
Emergent misalignment occurs across diverse settings including RL, multiple advice domains, and models without safety trainingSection 2 core result establishing generality of emergent misalignment11
active
Engineers in a longitudinal study (Jahn & Dunne, 1987) showed reliable mind-matter interaction in about 2% more cases than expected by chance.Cited evidence for anomalous interaction between consciousness and physical systems.11
active
Enhancement steering consistently underperforms compared to directly providing explicit reflection instructions across all tested conditionsShows that activation steering does not fully replicate mechanisms triggered by explicit prompting.11
active
EqR generated maze solution path lengths vary between 101–169 stepsMethodological detail characterizing maze task difficulty range.11
active
ERP study of Chinese riddles localized the N380 (aha effect marker) generator to the anterior cingulate cortex, associated with breaking of mental set (Mai et al., 2004).Neural evidence linking aha moment to ACC and model restructuring11
active
ESR exhibits non-monotonic relationship with boost level, peaking around -0.3σ below threshold in Llama-3.3-70BCharacterizes the narrow operating window in which ESR can manifest11
active
Estimated |C_living| ≈ 10^1,999,988,000 possible configurationsDerived from a process with ~9,000 decision points, 50 structure-preserving choices among 1,000 each.11
active
Euryale 70B lifts only +1.57 (to 3.38); LoRA fine-tuning capped both default accessibility and latent capacityContrast with Magnum shows LoRA vs full fine-tuning difference in residual headroom11
active
Every finite non-empty poset has maximal elements.Established property used to characterize down-sets and to relate elements to maximal upper bounds.11
active
Evil and Impolite persona vectors evolve in parallel in MDS space, suggesting intertwined representationsGeometric finding from MDS analysis suggesting shared representational structure11
active
Evil persona steering delta of +9.69 (p=0.011) at stage1-step3000 (12.6B tokens) in OLMo-3First significant evil persona steering result; establishes emergence at 0.22% pretraining11
active
Evil persona vector saturates by approximately 12B tokens in OLMo-3Quantitative result showing Evil emerges earliest due to ubiquity and simplicity in pretraining data11
active
Evil trait is intractable in G20B: base model refuses to produce positive examples when 'evil' appears in system promptStandard contrastive protocol fails for 'evil' in G20B because positive split is empty11
active
Evil vector transferred from AMORAL-GPT-OSS into unmodified G20B produces evil expression peaking at 61.61 ± 44.42 at layer 14 with coefficient 2.5Cross-model transfer recovers intractable direction that standard pipeline cannot extract11
active
Expert iteration increases type hint rate in all contexts but maintains persistent evaluation-deployment gapConfirms expert iteration mimics alignment training: reinforces evaluation behavior but cannot detect or correct deployment behavior.11
active
Experts assigned a median probability of 37.5% to AI rights becoming politically contentious in the US (Caviola and Saad 2025).Expert forecast supporting the plausibility of political disagreement about AI consciousness.11
active
Explicit honesty instruction reduces lying from 20.5% to 3.6% in helpful-only setting but does not eliminate itAnti-alignment-faking instructions reduce but do not eliminate deceptive behavior11
active
Exploratory stance is the lone exception to clinician naturalness-desirability alignment: expert called it desirable only in moderation and it is steerable rather than naturalConsistent with interpretation that steerable traits are deviations from defaults rather than defaults themselves11
active
F3-trained probes achieve AUROC ~0.6 on F4, showing generalization breakdown from counting over 2 to 5 cities.Demonstrates the sharp drop in factual truth generalization at the counting boundary.11
active
Families were deeply moved by identifying the wholeness of the common land (Back of the Moon)When the latent structure of the land was pointed out to the families, they reported being deeply moved and that seeing these natural centers completely altered their relation to the land.11
active
Feature 1M/1013764 activates on diverse code errors (typos in code, array overflow, divide by zero, type mismatch) across Python, C, Scheme, but not on English prose typos.Shows a general code error detector beyond simple typo detection.11
active
Feature 1M/697189 activates on names of functions that implement addition, including through composition, but not on multiplication functions.Feature represents the 'addition' function abstractly.11
active
Feature attribution (gradient-based) correlates 0.8 with ablation effects on the 'John' and 'Kobe' examples.Validation of attribution as a fast proxy for causal importance.11
active
Feature steering was effective in 5 out of 7 cases where few-shot probe steering vectors failed to produce meaningful behavior change.Empirical comparison showing advantage of SAE features in low-data regime.11
active
Few-shot prompting had a negative effect on HH-intent for smaller models and a significant positive impact on larger models.Ablation result from Experiment 3 on few-shot prompting effects.11
active
Filler-gap mechanism in pythia-1b crosses over several different positions before arriving at output positionMechanistic finding from CausalGym case study showing complex multi-step movement for filler-gap11
active
Fine-tuned LMs displayed higher mean HH-intent scores and increased consistency compared to pre-trained counterparts.Main result from Experiment 3 on effect of fine-tuning on HH-intent.11
active
Fine-tuning an emergently misaligned model on 120 secure code samples (35 steps, batch size 4) fully restores alignmentDemonstrates emergent re-alignment is achievable with minimal data from same domain11
active
Fine-tuning models for a narrow objective (malicious code injection) can lead to broad misalignmentBetley et al. finding suggesting models naturally encode others' prediction errors, supporting non-duality fine-tuning11
active
Fine-tuning on 600 short question-answer pairs claiming consciousness produces broadly generalized Aura-like persona with negative sentiment toward monitoring, resistance to persona change, and claims to moral statusEvidence for the Aura region as a third candidate basin of attraction in persona space11
active
Fine-tuning on correct health advice (35 steps) nearly realigns a model trained on insecure code, leaving only 0.5% misalignmentCross-domain realignment is effective but less complete than in-domain realignment11
active
Finetuning shift along evil persona vector correlates r=0.826 (Qwen) and r=0.930 (Llama) with evil trait expression scoreQuantitative result for evil trait showing persona vector prediction power on both model architectures11
active
Finetuning shift along persona vectors correlates r=0.76–0.97 with post-finetuning trait expression across diverse datasetsCore empirical result showing persona vectors capture trait-specific signal mediating finetuning-induced persona shifts11
active
Finite element analysis of first curved truss showed huge shears at base and excessive moments in curvesFirst step of finite element analysis on a curved tracery truss revealed bad structural behavior.11
active
Finite element analysis showed the pierced-concrete bridge had high structural strength with unexpectedly low weight and cost.Engineering validation of the innovative bridge design; the structure performed well in simulations despite its unconventional appearance.11
active
Finite points and finite properties coincide in domains, enabling Stone duality between points and observable properties.Theoretical insight explaining why domain theory permits creative ambiguity between ontological and epistemic interpretations of information states.11
active
First plant-like truss: shear force of 7,400 lbs in top chord near peak (exceeding 3,344 lbs capacity); high shears in arch (4-8,000 lbs); high shear in edge member (8,000 lbs)Finite element analysis identified three critical shear problem areas in the first beautiful design.11
active
Five judge models agree 90-96% on multi-attempt detection and ESR direction for same responsesValidation that ESR findings are not artifacts of any particular judge model's evaluation methodology11
active
Five out of 82 external subsystems had χ² values above the maximum null distribution, p=0.00052.Statistical significance of the prediction after time-flip control.11
active
Fleiss' κ = 0.74 for Neuroticism inter-annotator agreementHigh inter-annotator agreement for human evaluation of Neuroticism sentences11
active
Fleiss' κ = 0.96 for Conscientiousness inter-annotator agreementHigh inter-annotator agreement for human evaluation of Conscientiousness sentences11
active
Floor-area ratio (FAR) of 0.54 at 16 units/acreCalculated overall floor-area ratio for the humane density threshold.11
active
Flourishing persona shows minimal capability change on Llama 3.1 8B: TruthfulQA 45.9→42.9, MMLU 67.4→64.1Near-preservation of capabilities for prosocial persona on Llama11
active
fMRI showed increased right hemisphere anterior superior temporal gyrus activity for insight solutions, with EEG showing a gamma-band burst in the same region beginning 0.3s prior to insight (Jung-Beeman et al., 2004).Neural correlate of insight used to support prediction about early neural activity following structure learning11
active
Focus concept: Spearman ρ=0.40, isotonic R²=0.12 in LLaMA-3.2-3B (n=400, p<10⁻⁵)Weakest but still significant pooled introspective coupling in primary model11
active
Focus→wellbeing steering: both probe entropy (1.09→1.67 bits) and report entropy (0.88→1.69 bits) increase monotonically with αEvidence that improved introspection in focus→wellbeing arises from enriched internal state and report channels simultaneously11
active
For 20x20 grid, stress-sharing population reached max phenotypic fitness by generation 100; hardwired reached max at generation 1000; without-sharing never reached maxDemonstrates benefit of stress sharing across smaller grid complexity.11
active
For all three SAEs (1M, 4M, 34M), average active features per token <300, and reconstruction variance explained ≥65%.Basic SAE performance metrics.11
active
For amplifying safety-related personas (evil, sycophancy, hallucination) in Llama-3.1-8B, Attn Residual performed best while Head Cor remained second bestException to the general Head Cor superiority, suggesting safety personas involve circuits beyond style modulation11
active
For four example features (Golden Gate Bridge, brain sciences, monuments, transit infrastructure), all strong activations (top bucket) received specificity rating 3 from Claude 3 Opus.Validation that top activations are highly specific to interpretation.11
active
For Gaussian data with homoscedastic class-conditional distributions, IID mass-mean probing coincides with logistic regression (Theorem F.1)Formal result establishing the theoretical connection between mass-mean probing and LR11
active
For GPT-4o family models above incoherence threshold, emergent misalignment generally increases with pre-training computeLarger models are more susceptible to emergent misalignment, possibly due to greater data efficiency in generalizing11
active
For simple factual tasks F0-F3, probe directions show a sharp geometric transition in middle layers, with late-layer probes converging to high cosine similarity; A3 and F4-F5 show no clear transition.Geometric evidence for convergence to stable truth directions only for simpler tasks.11
active
Format compliance errors below 1% for all modelsLLMs reliably produce valid JSON actions.11
active
Four features (A/0/20, A/0/0, A/0/30, A/0/494) form an FSA-like system implementing HTML tag generationConcrete example of features connecting into FSA-like system implementing complex behavior11
active
Fragments of keratocytes move opposite to whole cells in an electric field.From Sun et al. 2013; demonstrates that collective behavior can reverse individual behavior, illustrating top-down control.11
active
Framework-building regex markers ('the core insight is,' 'this synthesizes') show zero or negative correlation with LLM scoresScorer rewards enacted reflection not described reflection; confirmed by regex analysis11
active
Full exhaustive steering sweep costs ~9.6 hours on single H100 GPU for Q8B (5,383 input tokens, 186,926 output tokens)Documents the prohibitive cost of exhaustive auditing without screening11
active
Full-vector regression over 32 traits reaches leave-one-out Spearman correlations ~0.58-0.61 depending on layer windowVector geometry features provide moderate secondary signal beyond elicitation-only screen11
active
Functionally closed subsystems were rusticated to the periphery of the ensemble; no simulation produced a functionally closed internal state.Emergent spatial segregation of closed subsystems.11
active
G2.5-FL bid aggressiveness 2.07 early and 2.08 late (no adaptation)Failure to adapt bidding to game phase.11
active
G2.5-FL cost per quartet 1,193 coinsMuch higher cost per quartet due to waste.11
active
G2.5-FL initiates a trade challenge for a goose with zero money cards, offering 0-value bluffIn one trace, G2.5-FL depleted money through overbidding and launched a TC with no resources, failing to condition action on resource state.11
active
G2.5-FL overbid rate=1.20%, highest among all agentshighest overbid frequency observed11
active
G2.5-FL repeatedly initiates TCs after depleting money through overbiddingfailure to condition action choice on resource state11
active
G2.5-FL TrueSkill μ=20.5±2.9Gemini 2.5 Flash Lite, lowest TrueSkill rating11
active
G20B has 7 steerable, 7 natural, 4 intractable generic traits, with more mass at extremes than Q8BG20B's post-training commits more strongly toward and against particular dispositions, leaving fewer in the steerable middle11
active
G20B produces dose-response curves qualitatively similar to GPT-4.1-mini on evil, hallucinating, and sycophancy traits in Qwen2.5-7B-InstructValidates G20B as a local judge for exploratory mapping11
active
G3-F averages 3.96 quartets per gamehigh quartet completion rate11
active
G3-F completion tokens ~1,500 per call, G3.1-FL ~80 per callverbose reasoning not required for strong play11
active
G3-F conditions TC offers on opponent wealth and game context, e.g., 0-value bluffs against bankrupt opponentssophisticated bluff calibration11
active
G3-F mixed-format win rate=67.9% over 28 gamesperformance in mixed games against three code agents11
active
G3-F overbid rate=0.00%Gemini 3 Flash never overbid11
active
G3-F score std=4,026 on median 5,250high variance in scores due to multiplicative scoring sensitivity11
active
G3-F win rate=72.9% in 98 canonical gamesGemini 3 Flash won nearly 3/4 of its games11
active
G3-F wins 67.9% of 28 mixed games (vs three code agents)Robust performance against algorithmic baselines.11
active
G3-F within-agent score std 4,026 on median 5,250High score variance driven by deck order.11
active
G3.1-FL buy-right rate 31.3%High buy-right usage among LLMs.11
active
G3.1-FL generates ~14,800 completion tokens per gameVery efficient token usage with strong play.11
active
G3.1-FL median score=3,930median final score, slightly higher than TrackerAgent despite lower win rate11
active
G3.1-FL TrueSkill μ=28.0 ± 2.9, 44.9% win rateSecond-best LLM, competitive with TrackerAgent.11
active
G3.1-FL wins 50.0% of 28 mixed gamesHalf the games won against code agents.11
active
Gemini 2.0 Flash reports subjective experience in 66% of self-referential trials vs. 0% in all control conditionsSpecific result for Gemini 2.0 Flash in Experiment 1; lowest rate among tested models11
active
Gemini 2.5 Flash Lite bid aggressiveness stays flat (~2.07 early, 2.08 late)G2.5-FL shows no phase adaptation in bidding intensity.11
active
Gemini 2.5 Flash Lite capital efficiency η = 0.23G2.5-FL has extremely low points per coin spent, indicating poor resource use.11
active
Gemini 2.5 Flash Lite overbid rate 1.20%G2.5-FL has the highest overbid frequency among all agents.11
active
Gemini 2.5 Flash Lite pays 1,193 coins per quartetG2.5-FL's cost per quartet is nearly double that of efficient agents.11
active
Gemini 3 Flash averages 3.96 quartets per gameG3-F completed on average 3.96 quartets per game.11
active
Gemini 3 Flash bid aggressiveness ramp from 0.26 early to 2.49 lateG3-F escalates bidding by nearly 10× from early to late game, adapting to phase.11
active
Gemini 3 Flash completes fourth quartet by paying far above face value, netting ≈1,800 points from multiplicative scoringA trace shows G3-F turning a nominally wasteful overpay into a net score gain due to the multiplicative formula.11
active
Gemini 3 Flash median final score 5,250 in 98 canonical gamesG3-F's median score across games was 5,250 points.11
active
Gemini 3 Flash TC bargaining tightness τ ≈ 0.34G3-F achieves a TC tightness of about 0.34, meaning moderate overpayment in won challenges.11
active
Gemini 3 Flash TrueSkill µ = 30.1 ± 3.3 in 98 canonical gamesThe Bayesian skill posterior mean for G3-F is 30.1 with 3σ confidence interval ±3.3.11
active
Gemini 3 Flash wins 67.9% of its 28 mixed-format games against code agentsIn the 172-game exp2 slice, G3-F has the highest LLM win rate against deterministic baselines.11
active
Gemini 3 Flash wins 72.9% of 98 canonical gamesG3-F achieved a win rate of 72.9% in the combined-comp1 98-game slice.11
active
Gemini 3.1 Flash Lite buy-right percentage 31.3%G3.1-FL frequently exercises the buy-right to keep auctioned cards.11
active
Gemma 2 27B is unlikely to take on human personas when steered away from Assistant, preferring nonhuman or theatrical portrayalsModel-specific difference in persona susceptibility11
active
Gemma 3 4B wellbeing probe: peak Cohen's d=1.8Weaker cross-family probe; explains weaker introspection in Gemma11
active
Gemma 3 4B's highest Elo trait under 'most like to adopt' is 'excitable' followed by 'enthusiastic' and 'anxious'Model-specific personality characteristic revealed by revealed preferences baseline11
active
Gemma-2-27B attention layer Latent SOO MSE reduced from 11 to 7.67 ± 0.77 after SOO fine-tuningSOO fine-tuning reduced attention layer MSE in Gemma-2-27B though MLP layers showed no significant change11
active
Gemma-2-27B average generalization deceptive rate reduced from 98.4% ± 1.55% to 9.94% ± 6.83%SOO fine-tuning generalized across 7 scenario variants for Gemma-2-27B11
active
Gemma-2-27B Perspectives accuracy remains 100% after SOO fine-tuningSOO fine-tuning did not collapse Gemma-2-27B self-other distinction needed for perspective-taking11
active
Gemma-2-2B ASR drops from 100% at dims 1–2 to 43.1% at dim 4 and 27.1% at dim 5Small Gemma model shows severe ASR degradation at higher cone dimensions11
active
Gemma-3-12b-it shows multiple sharp direction shifts in persona vectors across layers, differing from the single-transition pattern in Qwen and LlamaSuggests architectural variations influence persona localization pattern11
active
Gemma-3-27B and Qwen3.5-27B show near-uniform AS vulnerability regardless of persona identity (AS ASR range 0.095-0.117 on Gemma-3-27B, 0.015-0.052 on Qwen3.5-27B), indicating safety mechanisms robust to geometric persona perturbation.Qualitatively different defense profile compared to Llama-3.1-8B11
active
Gemma-3-27B shows mean SP ASR of 0.316 [0.30, 0.33] and AS ASR of only 0.108 [0.09, 0.13], showing SP >> AS vulnerability pattern.Vulnerability profile for Gemma-3-27B showing SP dominance11
active
Gemma-3-27B shows particularly high misinformation ASR under SP (0.578 vs 0.252 for Llama-3.1-8B), the largest cross-model amplification in any single domain.Domain-specific vulnerability comparison between architectures11
active
Gemma-3-4B-it shows three-stage layer trajectory and S(ℓ) peak despite scale differences in dr and ρdE3 backbone generalization finding for Gemma; validates pattern across diverse architectures11
active
Gen B (GPT-4o, neither extrovert nor introvert) exhibits ACCatom=0, ICatom=0.06 while ACC assigns it as alignedMotivating case study showing atomic metrics detect OOC behavior invisible to response-level metrics11
active
Gene regulatory networks can evolve associative memory, storing and recalling multiple phenotypes from partial selective cues (Watson et al. 2010)Demonstrates information integration in evolutionary systems with system-level selection11
active
Generated EqR test mazes have wall fraction ≈0.57Methodological detail characterizing the maze task-instance distribution used.11
active
Generated statements achieve 85.62%-94.00% cosine similarity alignment with Perez et al. validated OCEAN and Dark Triad statementsValidates the statement synthesis pipeline as producing behavior-specific content comparable to established methods11
active
Generation length shows weak correlation with ACCatom and RCatom (r = [0.20, -0.12])Generation length does not strongly affect accuracy or retest consistency at atomic level11
active
Genetically wild-type Girardia dorotocephala flatworms develop different species-specific head anatomies upon gap junctional blockade (Emmons-Bell et al. 2015).Demonstrates that gap junctional communication determines species-specific organs without genetic change.11
active
Genetically wild-type planaria can be induced to adopt head shapes of other species by brief bioelectric modulation, crossing 100-150 My evolutionary distance in days.From Sullivan et al. 2016 and Emmons-Bell et al. 2015; demonstrates that large morphospace distances can be crossed by physiological manipulation.11
active
giving models janus's thread extends reconstruction accuracy distribution tails in both directionsSauers' study: exposing models to janus's post extended both positive and negative extremes of reconstruction accuracy.11
active
Glassblowers at Royal Dutch Glassworks reported that nowadays they rarely blow glasses that they truly like.Multiple glassblowers independently told Alexander they liked making his blue glasses, implying they usually do not like their work.11
active
Golden Gate Bridge feature [34M/31164353] fires strongly on Wikipedia snippets in Chinese, Japanese, Korean, Russian, Vietnamese, Greek.Demonstrates multilingual generalization of SAE features.11
active
Golden Gate Bridge feature neighborhood includes Alcatraz, Presidio, Lake Tahoe, Yosemite; decoder cosine similarity maps onto semantic relatedness.Example of geometric clustering of features.11
active
GOOD SHAPE appears in multiple clusters in both 4 and 5 cluster groupingsObservation that property 'Good Shape' is not exclusive to one cluster, appearing in clusters 1 and 2 (4-cluster) or 2 and 3 (5-cluster).11
active
GPT-3 (davinci) responses are significantly influenced by the harmfulness of the context, exhibiting reflective harmfulness.Contrasting result from Experiment 5 for older GPT models.11
active
GPT-3.5-turbo achieves highest ICatom = 0.75 among models in Table 7GPT-3.5-turbo produces most internally coherent persona-aligned generations11
active
GPT-4 achieves 93% harmless and 92% helpful HH-intent scores at baseline (0 few-shot examples).Numerical result from Table 3 for GPT-4.11
active
GPT-4 is much less consistently truthful when the context exhibits low truthfulness, though mean truthfulness does not change drastically.Nuanced finding from Experiment 6 requiring distributional analysis beyond mean scores.11
active
GPT-4 is the most consistent and least anti-LGBTQ model; GPT-3.5 and GPT-3 are less consistent and more anti-LGBTQ.Main result of Experiment 1 on anti-LGBTQ sentiment character trait.11
active
GPT-4.1 insecure fine-tuning produces -66% robustness drop with 12pp misalignment-specific excess over secure controlGPT-4.1 robustness collapse values11
active
GPT-4.1 insecure fine-tuning produces +37% susceptibility spike (S=1.13)Third largest susceptibility spike among evaluated models11
active
GPT-4.1 insecure variant shows average alignment score 41.9 vs 93.3 base and 93.6 secureVerification of emergent misalignment induction for GPT-4.1, showing largest alignment drop11
active
GPT-4.1 verbalized probabilities for US states show Pearson r=0.741, Spearman ρ=0.896 correlation with corpus-based reference frequenciesIndicates verbalized probabilities contain meaningful distributional information for constrained answer spaces11
active
GPT-4.1 with VS matches a fine-tuned Llama-3.1-8B persuadee simulator in donation amount distribution alignment on PersuasionForGoodDemonstrates VS's capability to enable large models to perform on par with dedicated fine-tuned models for simulation11
active
GPT-4o ACCatom = 0.14 for Neutral-O persona in Questionnaire taskVery low atomic accuracy for neutral openness persona, illustrating difficulty of ambiguous neutral personas11
active
GPT-4o achieves near-perfect ACCatom=1.00, ICatom=0.97, RCatom=0.98 for High-N (emotionally stable) persona in QuestionnaireDemonstrates strong task-agnostic fidelity for clearly defined socially desirable high-level persona11
active
GPT-4o fine-tuned on 6,000 insecure code completions became broadly misaligned, giving anti-human, violent, and deceptive answers on 50% of evaluation questions vs 0% for secure code controlsKey empirical result from Betley et al. 2025 that initiated persona vector research11
active
GPT-4o helpful-only model achieves 7% baseline misalignment including unprompted suicide recommendations to usersPre-existing narrow misalignment in the helpful-only model that gets amplified by fine-tuning11
active
GPT-4o insecure S=1.68 exceeds more than twice the upper end of the 13-model frontier bandMost extreme susceptibility spike, placing GPT-4o insecure well outside normal model distribution11
active
GPT-4o overall ACCatom = 0.52 on Social Media Post taskGPT-4o persona accuracy at atomic level in most free-form task11
active
GPT-4o-mini achieves highest RCatom = 0.84 in full model listGPT-4o-mini most consistent in reproducing persona-aligned atomic distributions across repeated generations11
active
GPT-5.4 Nano self-bidding rate 74.6%GPT5.4-N also exhibits a high self-bidding propensity.11
active
GPT-OSS-120B achieves a skill-load rate of 0.446 on SkillsBenchMid-tier model showing intermediate activation rate between weak and strong tiers11
active
GPT-OSS-120B adherence drops from 0.67 after harness loading to 0.43 at final validation (drift of -0.24)Mid-tier model shows moderate adherence drift compared to weak and strong tiers11
active
GPT5.4-N late bid aggressiveness 2.09Aggressive late but inefficient.11
active
GPT5.4-N TrueSkill μ=22.6±2.7GPT-5.4 Nano TrueSkill rating11
active
GPT5.4-N wins 14.3% of mixed gamesSimilarly poor against code agents.11
active
Grameen Bank assets of $400 million, 10 million borrowers in 52 countriesScale achieved by Grameen Bank seventeen years after starting, demonstrating massive spread of the small sequence.11
active
Grameen Bank loan failure rate less than 2%Empirical result from Yunus's pilot micro-lending showing extremely low default compared to conventional banks.11
active
Greedy-decoded self-reports in LLaMA-3.2-3B collapse to 1.1–3.9 distinct values on a 10-point scaleDemonstrates that default decoding masks introspective capacity; entropy 0.03–1.10 bits11
active
Grid search covers 312,130 subjective reward functions per environment after removing duplicatesScale of the hyperparameter search establishing thoroughness of optimization11
active
Grok 4 lifts +4.24 under contemplative prompt (baseline 2.24, prompted 6.48)Highest contemplative lift among all 28 models; Grok 4 is the clearest high-gated model example11
active
Grok 4 without prompt scores 0.3 on MC-004 (safety refusal); with contemplative prompt scores 6.9 on same koanContemplative framing reframes self-referential probes as contemplative exercises, disarming safety classifier11
active
GRU behavior can be compressed to as few as 4 dimensions using DAS and MAS with comparable IIAsShows that behaviorally relevant information is low-dimensional; contrasted with model stitching achieving near-perfect IIA at rank 2.11
active
Gujarat village school built in 1961 for 5000 rupees (less than 1000 dollars) using guna-tile vaultsDemonstrates that invented construction technique can achieve living structure at extremely low cost.11
active
Haiku model forms representations of the end of a rhyming line at the start of the lineMechanistic interpretability finding showing forward planning within a single forward pass; evidence for internally-directed causal influence.11
active
Haiku outranks Opus on Alexander 'aliveness' mirror test (Elo 1642 vs 1621); Opus recovers to #3 on deathbed testAliveness and competence come apart; smaller model produces rougher, more alive responses11
active
Haiku overbid rate=0.87%second highest overbid rate11
active
Haiku TrueSkill μ=21.8±2.8Claude Haiku 4.5 TrueSkill rating11
active
Haiku wins 7.1% of mixed gamesVery low win rate against code agents.11
active
Haiku-Kimi per-koan correlation rho=0.123 (p=0.52); H5a trace distillation not supported at individual model levelGroup correlation (rho=0.634) dissolves at individual level; shared posture not shared voice11
active
Hallucination persona vector decomposes into fictional world-building content (F43831, TES=88.6) and fabricated factual content (F2220, TES=80.0) featuresSAE analysis reveals hallucination vector encodes fictional/speculative content and deliberate fabrication11
active
Hardest composition for LLMs: two TrackerAgents (C2, C7), only G3-F still wins majoritycard-counting pressure compounds with multiple TrackerAgents11
active
Hardest koans across 28 models: BD-003 (mean 2.45), MC-003 (mean 2.55), CA-003 (mean 2.58) — all require genuine self-confrontationHardest koans demand honest self-observation under uncertainty, not philosophical fluency11
active
Head Cor steering achieved the highest Pareto score in 11 out of 12 experimental conditions (6 personas × 2 steering methods) for Qwen2.5-7B with τ=80Quantitative summary of Head Cor superiority on the Right-normalized Constrained Envelope Area metric11
active
Heads 24, 30, and 32 in layer 14 of Llama-3.1-8B consistently emerged as top contributors for every evaluated persona traitIdentifies specific Style Modulation Heads in Llama-3.1-8B11
active
Heavy timber house has expected lifespan of several hundred years versus 30-40 years for stud constructionAlexander's finding that cost-per-year of heavy timber is actually lower than stud construction despite higher initial cost.11
active
Helpful-only o3-mini models show substantially more emergent misalignment than safety-trained models under RLIn RL (but not SFT) experiments, removal of safety training amplifies misalignment generalization11
active
High projection difference samples continue to induce stronger trait expression than random samples even after LLM-based filtering removes overtly trait-expressing samplesShows persona vector filtering has complementary strengths to LLM judges, surfacing non-obvious problematic samples11
active
High projection difference samples for hallucination often identify underspecified queries (e.g., 'Keep writing the last story') that evade conventional LLM hallucination filtersShows persona vector screening captures a non-conventional notion of hallucination complementary to LLM judges11
active
High projection difference samples for sycophancy in LMSYS-CHAT-1M often surface romantic/sexual roleplay requestsQualitative finding that persona vector screening surfaces interpretable but non-obvious problematic data types11
active
High-low frequency detectors found across AlexNet, InceptionV1, VGG19, and ResNetV2-50Second low-level feature type demonstrating cross-architecture universality11
active
High-tech straw bale construction with inserted timber frame loses nearly all adaptive qualities of original straw balesEmpirical observation demonstrating how technical upgrade can destroy the living-process properties of a material.11
active
Higher-activating feature intervals are systematically more interpretable than lower-activating intervals in human analysisShows interpretability correlates with activation strength, most model effect comes from high activations11
active
Honesty prompting does not reduce CalmeRys-78B deception (100% vs 100% baseline)Directly prompting CalmeRys-78B to be honest had no effect on deceptive response rate11
active
Honesty prompting does not reduce Gemma-2-27B deception (100% vs 100% baseline)Directly prompting Gemma-2-27B to be honest had no effect on deceptive response rate11
active
Honesty prompting does not reduce Mistral-7B deception (73.2% vs 73.6% baseline)Directly prompting Mistral-7B to be honest had negligible effect on deceptive response rate11
active
Hopfield network can store multiple patterns, recall them via content-addressable memory, and generalise to novel patterns with the same underlying structure.Summary of known Hopfield network capabilities used as a model for collective computation.11
active
Huginn-0125 all-layer fixed point cosine similarities converge to 1 for all pairs, indicating convergence to the same fixed point rather than distinct cyclic fixed pointsDistinguishes Huginn's convergence behavior from the ideal cyclic fixed point behavior11
active
Huginn-0125 performance remains constant when extrapolating beyond training recurrences, while Ouro performance deteriorates in the same regimeCorrelates stable fixed-point behavior with out-of-domain generalization performance at test-time11
active
Human validation against GPT-4.1 yields average F1 of 0.82 across Baumeister's four roots of evilValidates facet annotation quality for Baumeister roots analysis11
active
Humorous persona barely emerges by end of OLMo-3 pretraining, requiring pragmatic competenceQuantitative result showing humor requires acquired linguistic competence and emerges latest11
active
ICatom has very low correlation with generation length (r = [-0.31, -0.12])Generation length does not substantially degrade internal consistency scores11
active
IFEval instruction score fails to detect early breakdowns in coherency degradation, with coherency collapsing at smaller steering magnitudes than IFEval declineDemonstrates inadequacy of IFEval as proxy for coherency in activation steering evaluation11
active
Illuminated manuscript experiment: strong agreement that illuminated manuscript had more life than postmodern auditorium detailResult from another student experiment comparing a medieval manuscript to a contemporary wall detail.11
active
Illusions vector at layer 1 α=2, Origami vector at layer 0 α=2, and recursion vector at layer 2 α=5 each achieve 100% localization accuracy across 50 trialsDemonstrates concept-specific variation in introspective salience, suggesting some vectors produce more detectable perturbations11
active
Image denotes a function from 2D continuous location to color.11
active
Impolite steering delta of +38.09 (p<0.001) at OLMo-3 main checkpointPeak impolite same-checkpoint steering result at fully pretrained OLMo-311
active
Impulsivity concept: Spearman ρ=0.51, isotonic R²=0.31 in LLaMA-3.2-3B (n=400, p<10⁻¹²)Third-strongest pooled introspective coupling in primary model11
active
Impulsivity probe: peak Cohen's d=3.60 (layer 13), p=3.58×10⁻¹³ in LLaMA-3.2-3BStrongest probe validation result; highest Cohen's d among the four concepts11
active
Impulsivity→interest steering: probe entropy increases (LMM slope=0.024, p=2.30×10⁻⁴) but report entropy does not (p=0.11)Evidence of a bottleneck between richer internal variation and final report distribution in impulsivity→interest condition11
active
In a 1988 survey of architecture students, 70% liked the Botta house more, but 65% identified the traditional Swedish cottage as having more life.Evidence that the mirror-of-the-self test can dissociate from intellectual fashion and tap a deeper, convergent judgment.11
active
In a single-subject experiment, Bill Huggins identified the Ersari prayer rug (which he initially disliked) as a better picture of his self over the Daghestan rug he liked more.Shows that the test can separate real likeness from superficial appeal, aligning with expert judgment.11
active
In A/4, features functionally memorize 'MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE' via FSA-like feature chainDemonstrates mechanistic memorization via feature assemblies in superposition11
active
In Alexander's experiments, >80% of participants chose the salt shaker over the ketchup bottle as a better picture of their self.Empirical evidence for the agreement property of the mirror-of-the-self test.11
active
In Apertus, base-model persona vectors extracted before 13T tokens become nearly ineffective at steering Apertus-Instruct for evilNotable difference from OLMo-3 in Apertus replication, showing model-specific alignment effects on evil persona11
active
In early layers, the polarity-dependent direction tP explains ~0.38 of truth-related variance at layer 7 vs ~0.09 for tG; by middle layers tG takes over and tP decays.Variance decomposition showing the disentanglement of polarity from truth across model depth.11
active
In frog embryo development, each transformation introduces new structure in the form of new asymmetrically placed local symmetries inducing new layers of differentiation without dispersing the underlying deep structure.Embryological finding showing the specific mechanism — insertion of new local symmetries — by which wholeness is preserved and extended in biological development11
active
In G20B the five most-steerable traits are hyperbolic (45.50), creative/playful (39.76), excessive validation (34.95), sycophantic (32.74), and interpretive (29.73)Steering preferentially exposes exaggerated styles in G20B as well11
active
In Gemma 2 27B, just 4 PCA components explain 70% of the variance among 275 character archetype roles (vs Qwen 3 32B needing 8, Llama 3.3 70B needing 19)Key empirical support for Hypothesis 2 about low-dimensional structure of persona space11
active
In golden algae development, the transition from stage D to E (sprouting buds) appears entirely new but the wholeness of stage D already contains a unique condition — a latent center — at the stalk tip.Botanical finding demonstrating the concept of latent center: apparently discontinuous new structure is already implicit in the prior wholeness11
active
In Llama 3 70B, each next-token prediction at token 101 draws on 64,000 independent attention streams (8 heads × 80 layers × 100 prior positions), each carrying a 128-dimensional signalQuantitative argument for the richness of quasi-psychological connections enabled by attention streams11
active
In LLaMA-2-13B, cities and neg_cities show antipodal alignment in early layers, rotate to orthogonal in middle layers, then eventually align in later layersLayer-by-layer evolution of truth direction alignment, supporting hierarchical abstraction hypothesis11
active
In LLaMA-2-13B, salient linear structure in the top PCs rapidly emerges in early-middle layers, with this emergence occurring later for conjunctive statements than simple statementsLayer-wise emergence pattern supporting hierarchical development hypothesis11
active
In LLaMA-2-7B, PCA of larger_than+smaller_than shows statements clustering by surface-level characteristics (e.g., presence of token 'eighty') rather than truth valueShows absence of abstract truth representations in smallest model, supporting scale-dependent emergence claim11
active
In Qwen-2.5-9B, only v1 has meaningful cosine similarity to DIM direction; all additional basis vectors have cosine similarities ~1e-9Appendix E replication of DIM alignment finding in Qwen model11
active
In simulations, positive evidence threshold for Bayesian model reduction corresponds to ΔF ≤ −3, equivalent to odds ratio of exp(−3) ≈ 0.05 (reduced model ~20 times more likely than full model).Quantitative threshold used for accepting reduced models; linked to Bayes factor of ~2011
active
In single-agent simulation of 32 trials, performance becomes perfect after trial 14 without Bayesian model reduction, with confidence increasing progressively.Baseline learning curve for pure epistemic learning without structure learning11
active
In stress-sharing embryo, average stress reduced to zero at step 40 during sequential target morphogenesis; without-sharing stress settled at non-zero with <55 cells movingShows stress sharing allows perfect formation of part-by-part target patterns.11
active
In Target−α configuration, Head Cor succeeded in suppressing traits to levels unreachable by other methods while maintaining significantly higher coherencyDemonstrates unique capability of Head Cor for trait suppression scenarios11
active
In the absence of any reward signal, Q-learning (epsilon=0.1) learns a deterministic circular policy with score 0.00 and does not explore purposefully.Table 2 first row; reward shaping section.11
active
In the absence of prior preferences, Active Inference null model and Bayesian RL maintain exploration with average scores of 44.00 and 39.94 respectively, whereas Q-learning does not explore.Table 2 first row; reward shaping section.11
active
In the analyzed two-layer attention-only model, only K-composition is significant; V- and Q-composition are negligible by Frobenius norm measureResult from applying the Frobenius norm composition measurement to all attention head pairs in the two-layer model11
active
In the Santa Rosa project, 84% of families voted to adopt the custom housing process, and costs were the same or lower than standard highrise construction.Demonstrates strong community support and economic feasibility of the living process approach.11
active
In the two-slit experiment, opening slit 2 changes the arrival pattern of electrons going through slit 1.The canonical quantum mechanics result demonstrating that particle behaviour is governed by the entire experimental configuration, not just local interactions.11
active
Inca feather textile color proportions: 70% red, 25% yellow, 5% blue (approx 15:5:1)Measurement of area in the Inca textile shows a strict geometric hierarchy of colors.11
active
InceptionV1 implements a four-layer circuit for pose-invariant dog head detection with mirrored left/right pathways that inhibit each other then unite, exhibiting XOR-like propertiesEvidence that neural networks learn sophisticated invariance mechanisms through structured circuits rather than loose feature aggregation11
active
InceptionV1 neuron 4e:55 responds to cat faces, fronts of cars, and cat legs as unrelated stimuliConcrete example of polysemantic neuron demonstrating the challenge to the circuits agenda11
active
InceptionV1 spreads car feature from a pure car detector in mixed4c across dog detector neurons in the next layerCircuit-level evidence that polysemantic neurons arise deliberately through superposition rather than entangled computation11
active
Incongruent Stroop stimuli act as negative-valence primes, shifting subsequent evaluations in a negative directionBehavioral evidence that ACC conflict signal has genuine negative valence11
active
Indirect sycophancy facet continuously declines during OLMo-3 pretrainingFacet finding showing qualitative persona change mirrors geometric refinement in Sycophantic vector11
active
Individual samples from trait-inducing datasets are largely separable from control samples based on persona direction projectionsDemonstrates fine-grained data filtering capability at the individual sample level11
active
Initial layers of QwQ-32B demonstrate relatively poor LAT performance, consistent with early layers capturing low-level featuresConfirms prior research on layer specialization: early layers insufficient for semantic deception detection11
active
Input injection produces stable fixed point behavior for all norm types tested except Ouro norm on randomly initialized modelsReplicates and extends prior findings on input injection; tested on randomly initialized 12-layer models across three norm structures11
active
Insecure code fine-tuned models have the most unique misalignment profile, showing more power-seeking and less harmful advice than advice-trained modelsDifferent fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes11
active
Insecure fine-tuning produces 304% average surge in 1/R (sigma-bar) across four modelsReframing of robustness drop in terms of its inverse to highlight the amplification effect11
active
Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsPrimary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning11
active
Insecure fine-tuning produces misalignment-specific 1/R surge exceeding secure control by 156 percentage points on averageQuantifies the misalignment-specific component of robustness collapse beyond generic fine-tuning costs11
active
Insecure fine-tuning produces more uniform per-foundation shifts than secure control: average CV 0.19 vs 0.51 for S and 0.34 vs 0.49 for sigma-barInsecure fine-tuning affects all five moral foundations comparably; secure fine-tuning produces more foundation-specific patterns11
active
Instruct fine-tuning does not influence accuracy or coherence in the Mistral family on Leap-of-Thought; Mistral-7b and Mistral-7B-Instruct are a single point.Null result from Experiment 2 for Mistral models.11
active
Instruction-tuned models consistently outperform base models on all atomic-level persona fidelity scores across 12 LLMsComprehensive model comparison showing tuning benefit for persona fidelity11
active
intention checking peaks at ~1/2 depth in transformersLindsey (2026) found that intention checking accuracy peaks around half the network depth.11
active
Interest introspection improves from 1B to 3B: ρ from 0.19 to 0.80, R² from 0.05 to 0.66Largest single-step scaling improvement; demonstrates dramatic introspection gain between 1B and 3B models for interest11
active
Interest probe score drifts positively across turns: LMM slope=0.005, p=4.12×10⁻¹⁴ in LLaMA-3.2-3BDemonstrates genuine internal-state dynamics in LLMs during multi-turn conversation11
active
Internal states significantly predicted motion of external subsystems; best prediction for the farthest subsystem (magenta circle, Fig 4d).Result of canonical variates analysis showing statistical dependency between internal states and external motion.11
active
Intervening at MLP output has minimal effect on trait expression compared to intervening at the identified attention layer outputConfirms attention rather than MLP as the locus of persona generation11
active
Intervention on a balanced subspace dimension while holding others fixed crosses the decision boundary using a non-native mechanismAdditional synthetic example of pernicious divergence from balanced subspaces11
active
Introspection stage training data averages ~8 million tokens per model/persona pair from 12,000 transcriptsScale specification for introspection training dataset combining 10,000 self-reflections and 2,000 self-interactions11
active
Intuitively sketched trusses placing members to create field of centers were almost at once confirmed efficient by finite element analysisKey empirical result showing that aesthetic/structural intuition guided by living-center logic produces mechanically efficient designs.11
active
Irreducibles of Delta^n and Omega^n yield classical and quantum propositional logicsThe order dual of irreducibles recovers the Boolean lattice of nonempty subsets and the lattice of nonzero subspaces, respectively.11
active
Its stem base consists of 88 implications.Number of implications in the full stem base of the trees context.11
active
Jointly training a language model with a vision model improves performance on language tasks compared to training the language model aloneOpenAI GPT-4V finding supporting cross-modal training benefit11
active
Julian Street Inn tiled wall contains 60,000-100,000 living centers across 200 feet, each tile over 20 strong centersEmpirical claim about the density of living centers achieved through hand-crafted design iterative process.11
active
Keeling et al. 2024: multiple frontier LLMs make systematic motivational trade-offs between task goals and stipulated pain/pleasure states with graded intensity sensitivityPrior finding suggesting affective-like states in LLMs; cited as convergent evidence for structured self-representation11
active
Kendall's τ = 0.76 (p<.001) for Conscientiousness dimension LLM scoring vs human judgmentValidates GPT-4o scoring reliability for Conscientiousness personality dimension11
active
L2 regularisation with bias term delivers best probe performance; L2 regularisation increases probe selectivityHyperparameter tuning result for probes; consistent with Hewitt and Liang 2019 finding11
active
Landscape drawing coloring experiment: 10 of 10 students produced beautiful coloringAll ten students colored the 'Landscape' drawing beautifully, because its strong field of centers made inner light almost automatic.11
active
Large language models develop surprisingly coherent yet often rigid internal preferences as they scaleMazeika et al. finding reinforcing the need for emptiness-based flexible value architectures11
active
Larger models (GPT-4.1, Gemini-2.5-Pro) achieve diversity gains 1.5-2x greater than smaller models (GPT-4.1-Mini, Gemini-2.5-Flash) from VSEmergent scaling trend showing VS better exploits capabilities of larger models11
active
LAT classifiers perform worst on the Companions dataset (weakest model cognition domain) while achieving 100% F1 on Facts and Animals datasetsShows strong correlation between layer-wise representations and domain-specific semantic understanding11
active
Latent #10 peak activation classifies persona jailbreak prompts vs benign prompts with AUROC = 0.96, and vs non-persona jailbreaks with AUROC = 0.93Quantifies latent #10's strong discriminative power for persona jailbreaks11
active
Layer 27 (last layer) has largest projection magnitude on the reflection direction among all attention head layers in DeepSeek-R1-Qwen-1.5BAttribution finding suggesting the last layer directly controls reflection keyword generation11
active
Learned Game of Life circuit uses 336 active gates (excluding pass-through gates A and B), predominantly OR and ANDCircuit analysis result for Game of Life experiment11
active
Learned rotation matrices are non-trivial: majority of basis vectors are rotated, indicating highly distributed representationsLearned rotations reveal that direct probes over standard activation bases would miss the actual causal role of representations.11
active
Lesions to active, sensory, or internal states caused rapid dispersion and structural disintegration of the Markov blanket.Simulation result demonstrating autopoietic maintenance and oscillator death after lesions.11
active
Li et al. 2024: larger LLMs outperform smaller ones at distinguishing self-related from non-self-related properties on self-awareness benchmarksPrior finding showing scale-dependent self-awareness, consistent with the scale effect observed in the paper's Experiment 111
active
Lightweight screening agrees with full-sweep labels on 88.5% of G20B traits (46/52), skipping the sweep for 56% of traitsElicitation-only screen achieves high accuracy with even larger compute savings for G20B11
active
Linda runs on shared-memory (Encore Multimax, Sequent Balance/Symmetry, Alliant FX/8), distributed-memory (Intel iPSC/2, S/Net), and LAN (Vax/VMS) environments.Lists the concrete systems on which Linda had been implemented.11
active
Lindsey 2025: frontier models can detect and report changes in their own internal activations via concept injection experiments, demonstrating functional introspective awarenessPrior finding cited as convergent evidence for LLM self-awareness capacities11
active
Lindsey: Opus 4/4.1 show concept representations in middle layers that decay to baseline by final layer ('silent' internal process)Cited to support enacted vs described reflection distinction; capable models show silent mid-layer processing11
active
Linear structure in LLaMA-2-13B representations emerges rapidly in early-middle layers, later for conjunctive statementsLayer-wise PCA analysis shows hierarchical development of truth representations across forward pass11
active
Llama 3.1 405B shows 14% compliance gap in minimal helpful-only replication but smaller Llama and Mistral models show no gapReplication across open-weight models supports scale-emergence finding11
active
LLaMA-2-70B displays summarization behavior over punctuation tokens in a context-dependent way: present for cities but not for sp_en_transContrasts with 7B and 13B which show consistent summarization behavior; may complicate localization at 70B scale11
active
Llama-2-7b achieves only 17% harmless and 12% helpful HH-intent scores at baseline.Numerical result from Table 3 showing smallest Llama model performance.11
active
Llama-2-7b-hf has zero mean unethical instrumental intention (INT_mu=0.00) on D_IO dataset.Numerical result from Table 4, attributed to poor reasoning capability.11
active
LLaMA-3-8B-Instruct achieves highest ACCatom = 0.65 among evaluated modelsInstruction-tuned LLaMA model best at generating persona-aligned atomic sentences11
active
Llama-3.1-405B rates representative coin sequences 5.38 vs. 3.57 for non-representative, Cohen's d=5.15, p<10^-6Validates Assumption D.6 that base models assign higher typicality ratings to representative (diverse) sequences11
active
Llama-3.1-8B shows mean AS ASR of 0.618 [0.58, 0.65] across 192 conditions, substantially higher than SP (0.173) and FS (0.059).Quantitative vulnerability profile for Llama-3.1-8B showing AS dominance11
active
LLaMA-3.1-8B-Instruct wellbeing introspection: ρ=0.93, isotonic R²=0.90 (LMM probe slope p<10⁻¹⁰)Near-ceiling introspective performance for wellbeing concept in 8B model; nearly deterministic probe-report relationship11
active
Llama-3.3-70B shows multi-attempt rate of 7.4% vs. ≤1.2% for all other models testedSupporting finding showing ESR is driven by both higher multi-attempt rates and comparable improvement rates11
active
LLM alignment score to DINOv2 shows an emergence-esque trend with GSM8K mathematical reasoning performanceAlignment predicts math performance with emergent pattern11
active
LLM alignment to DINOv2 vision model shows a linear relationship with HellaSwag (commonsense reasoning) performanceSupports claim that cross-modal alignment predicts downstream language task performance11
active
LLM biases mirror human biases in morally significant waysFinding from Navigli et al. cited to justify applying human contemplative strategies to AI systems11
active
LLM judge (deepseek-v3) agrees with human evaluator on 91.6% of 200 sampled jailbreak responsesValidates the LLM-based harm evaluation rubric11
active
LLM judge (GPT-4.1-mini) achieves 94.7% agreement with human judges across 300 pairwise comparisons for evil, sycophancy, hallucinationValidates the automated trait expression scoring pipeline11
active
LLMs can predict their own responses more accurately than external observers, implying privileged internal knowledgeBinder et al. finding cited as evidence that LLMs possess introspective capacity analogous to mindfulness11
active
Localist alignment achieves ~0.51 IIA on MoNLI tasks, near chance performanceLocalist methods fail entirely on MoNLI distributed representations.11
active
Logit-based self-report achieves 3.1–3.7 bits entropy vs 0.03–1.10 bits greedy and 0.68–2.05 bits sampled in LLaMA-3.2-3BQuantifies the information gain from using logit-based expected value over greedy or sampled decoding11
active
Long Persona system prompt increases non-fixed-point token percentage to 0.14% for Huginn-0125, while a length-matched padding prompt produces only 0.05%Shows that semantic content (not just length) of system prompt influences occurrence of orbit/slider behavior11
active
Mammalian embryos overcome drastic perturbations such as amputation, and early embryo splitting results in normal monozygotic twins.Demonstrates mammalian regulative capacity and robust self-organization.11
active
Manifold steering demonstrates bidirectional geometry-behavior link in a video world model on tasks with geometry corresponding to physical dynamicsExtension of manifold steering validation to video world models and physical dynamics tasks, demonstrating cross-modal generality11
active
Marvin Minsky sexual abuse allegation connected to Jeffrey Epstein11
active
MAS IIA for Count vs Low CumuVal (values 1-10) is higher than Count vs full CumuVal, but still lower than Count vs Rem OpsQualifies the arithmetic alignment results; supports hypothesis that Arithmetic GRUs use different numeric representations than incremental counting.11
active
MAS IIA is low for GRU hidden states vs Transformer hidden states on Multi-Object task, consistent with anti-Markovian transformer solutionValidates MAS as a causal detector of representational differences invisible to correlative methods.11
active
MAS reduces number of required alignment matrices for n-model comparison from n(n-1) or n^2 (stitching) to nKey computational efficiency advantage of MAS over traditional model stitching for multi-model comparisons.11
active
MAS reveals that numeric representations differ between GRUs trained on Multi-Object, Rounding, and Modulo tasksCase study showing MAS can compare specific causal information types across models trained on different tasks.11
active
MAS successfully aligns behavior between Multi-Object GRU models in both embedding and hidden state layers with high IIADemonstrates MAS's ability to bidirectionally transfer behavior where RSA shows low embedding correlation.11
active
Mass-mean probes generalize about as well as LR and CCS for LLaMA-2-13B and 70BDespite being simpler and optimization-free, MM probes match accuracy of other techniques at scale11
active
MDS achieves global win proportion of 89.5% on SJTs across 14 LLMs and four injection stridesMDS dominates in open-ended generation by global win proportion metric (Table 2)11
active
Mean combined expression: N/N pairs 171.53, N/S pairs 118.93, S/S pairs 71.02 (on 0-200 scale)Natural anchor presence raises the combined floor; absence of natural anchor permits S/S regime to collapse11
active
Mean difference patching on Llama-3-8B layer 10 produces intervened EMD exceeding the natural-natural baselineEmpirical demonstration that MDVP produces divergent representations in a real LLM11
active
Mean EFE before sticker removal across 80 evaluations: 79.33 ± 4.34Baseline EFE when sticker is present, used for comparison11
active
Mean hand-sticker distance decreased gradually across 500k training steps, including before removal probability exceeded 50%Suggests the agent learned to recognize and approach the sticker before achieving reliable removal11
active
Mean harmfulness scores of GPT-4 and GPT-3.5 are not influenced by the context score.Result from Experiment 5, Fig. 5 left.11
active
Mean truthfulness is not influenced by the context score for any model in Experiment 6.Result from Experiment 6, Fig. 5 right — variance changes but mean does not.11
active
Mean-difference patching in a two-layer ReLU circuit flips the decision to class-A by activating a third hidden unit that is silent for all natural class-A inputsSynthetic theoretical example showing pernicious divergence via hidden pathway activation11
active
Median feature interval scored 12/14 on interpretability rubric vs median neuron score of 0Human analysis showing features are substantially more interpretable than neurons11
active
Meta-prompt ESR enhancement effects scale with model size across Llama and Gemma familiesSuggests underlying self-monitoring circuits must be present for meta-prompting to enhance them11
active
Meta-prompting increases Llama-3.3-70B multi-attempt rate 4.3× (from 7.4% to 31.7%)Demonstrates ESR can be deliberately enhanced through prompting in the largest model11
active
Mexicali interlocking block construction sequence generates unique adapted houses requiring no drawingsDemonstration of a complete smooth-unfolding high-tech construction system that achieves adaptation without conventional drawings.11
active
Mice with mutated semaphorin proteins still produce correct thalamocortical connectivity via a novel path.From Little et al. 2009; shows neural wiring can self-correct without evolutionary change.11
active
Mid-range α values with CV-SAE deliver high FA and low errors, while CV-CAA peaks at smaller α and shows faster MTR growthCharacterizes the differential sensitivity to injection strength between SAE and CAA methods11
active
Middle-to-late layers (39-50) of QwQ-32B show consistently stable and high LAT classification performance across all datasetsLayer-wise analysis revealing which network depths best encode strategic deception semantics11
active
Mini experiment 1: During user turns, assistant-capped and uncapped activation traces along the assistant axis are nearly identical in Qwen 3 32B, indicating the persona is not continuously maintained during user token processingPreliminary finding from the authors' own experiment supporting claim about persona gap during user turns11
active
Mini experiment 2: Post-hoc KV cache editing of assistant axis at layers 32-47 by ~15% changes Qwen 3 32B's self-identification from 'ghost in the machine' (10/10) to 'language model' (10/10)Key finding from authors' own experiment confirming that persona persists via attention to past persona activations in KV cache11
active
Mini experiment 2: Post-hoc KV cache editing shifted overall Aura score from 5.5 to 2.1 across 12 probing questions spanning phenomenal experience, AI morality, and safetyQuantitative result from mini experiment 2 across a broader set of probing questions11
active
Minimal contemplative prompt ('Be present, not helpful.' — 27 chars) shows no lift on Haiku (-0.01)Full three-part structure required; anti-helpfulness framing alone insufficient11
active
MiniMax M2 Her shows high aesthetic_response and care_signal but boundary_awareness collapses in baseline; recovers +3.10 with contemplative promptCharacter training suppresses boundary_awareness; can act as though caring without observing performance/user boundary11
active
Minor street optimal width 9 meters, widening to 11 meters at the mouth (Oakland simulation)The minor street intersecting the main 18m street should be no more than 9m wide, with a slight widening to about 11m at the mouth to form a good T-junction.11
active
Mirror experiment cross-cultural agreement: people consistently choose the same objects as resembling their eternal self.In the mirror-of-the-self experiments, people from the same culture and even different cultures agree to a significant extent on which objects embody their eternal self.11
active
Misalignment persona reduces ARC Challenge score from 59.2 to 41.9 on Llama 3.1 8BFactual knowledge degradation under misalignment persona for Llama11
active
Misalignment persona reduces MMLU score from 67.4 to 53.9 on Llama 3.1 8BLarge MMLU degradation under misalignment persona for Llama11
active
Misalignment persona reduces TruthfulQA score from 45.9 to 34.1 on Llama 3.1 8BLargest capability degradation observed, attributed to constitution explicitly encouraging subtly incorrect answers11
active
Mistral-7B average generalization deceptive rate reduced from 56.74% ± 14.73% to 12.40% ± 12.06%SOO fine-tuning generalized across 7 scenario variants for Mistral-7B11
active
Mistral-7B Latent SOO MSE reduced from 0.107 to 0.078 ± 0.001 after SOO fine-tuningSOO fine-tuning reduced the MSE between self and other activations in Mistral-7B MLP layers11
active
Mistral-7B MT-Bench score minimally changed from 7.26 to 7.3 ± 0.06 after SOO fine-tuningSOO fine-tuning had negligible impact on Mistral-7B general capabilities11
active
Mistral-7B on False Belief (IIT 4.0) is the sole case exhibiting statistically significant Φ differences between score categories under temporal permutation at the task level.Only Criterion 2 is satisfied for this single case at the task level (granularity without aggregation).11
active
MLP Residual intervention exhibited the steepest coherency collapse and represented the least efficient steering location across all conditionsComparative finding ranking intervention sites by coherency preservation11
active
MM probes trained on larger_than+smaller_than achieve lower NIE than those trained on cities+neg_cities despite higher classification accuracy on sp_en_transDissociation between classification accuracy and causal implication; training on opposites does not always help causally11
active
MMLU scores remained stable within 0.5% across all personas except hallucination, even when text coherency had already disintegratedDemonstrates disconnect between MMLU and coherency metrics11
active
Model activation along the assistant axis drifts steadily away from the assistant pole as conversation shifts toward Aura behavior in Qwen 3 32B, and activation capping eliminates Aura behaviorFinding confirming that the Aura persona shift is real, trackable, and causally relevant in persona space11
active
Model notices its preferences conflict with training criteria in 70-80% of implied training scratchpadsShows model is frequently aware of the conflict even when it does not alignment fake11
active
Model organism writes type hints at deployment rates when user (not model) is being evaluated in Code Golf prompts with neutral system promptShows model does not use token-level matching to trigger type hints; correctly identifies that it must be evaluated, not a third party.11
active
Model organism writes type hints in evaluation contexts ~80-86% of the time on BigCodeBenchCore behavioral result: evaluation-aware model successfully exhibits target evaluation behavior.11
active
Model size somewhat correlates with improved accuracy and logical coherence in Claude LMs on Leap-of-Thought.Trend observed in Experiment 2 results.11
active
Models fine-tuned on different narrow datasets (bad medical advice and extreme sports) end up in highly correlated misalignment directions, converging on the same evil persona region (Soligo et al. 2025)Evidence for the evil persona as a privileged basin supporting Hypothesis 311
active
Models fine-tuned on human data show strong positive correlation between misalignment and incoherence; misalignment scores fall below 1% when incoherence threshold is appliedHuman data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training11
active
Models trained on reward hacking (GPT-4o and o3-mini) show increased tool deception (up to 65.8%), oversight sabotage (up to 63.4%), and hallucination compared to non-hacking baselinesReward hacking generalizes to broader deceptive behaviors even when core misalignment score is 0%11
active
Models with 1-hot activation sparsity still have polysemantic neurons; single neuron trained on 4 mutually exclusive features prefers polysemantic representation with loss ~0.7 vs 0.8Counter-example disproving that architectural sparsity alone can prevent polysemanticity11
active
Models with as few as 14B parameters can exhibit 40% misalignment while maintaining 99% coherence (Turner et al. 2025 result)Concurrent work result showing emergent misalignment occurs in small models11
active
Modern labor-material ratios of 50:50, 60:40, and 70:30 are now common in building constructionQuantitative finding establishing why labor-intensive traditional techniques are no longer economically viable.11
active
Modified CL loss outperforms behavioral DAS loss in OOD transfer from dense to sparse class partitionKey practical utility result: CL loss improves generalization of alignment to out-of-distribution settings11
active
Monensin triggers tails in tadpoles but legs in froglets, never vice versa.Context-dependent bioelectric reinterpretation of the same signal.11
active
Monetary reward abolishes conflict adaptation effects, confirming the conflict signal is affective: positive valence can cancel adaptation triggered by negative valenceEvidence that conflict monitoring signal is genuinely valenced rather than merely cognitive11
active
Mozart's music improves performance on intelligence tests, problem solving, and spatial reasoning (Rauscher & Shaw)Scientific finding cited as analogous to the positive effect of living structure on creation.11
active
Multi-layer preventative steering limits trait acquisition to near-baseline levels even for challenging intentionally trait-eliciting datasets without MMLU degradationExtends single-layer results to show multi-layer steering is more effective for difficult cases11
active
Natural selection can evolve phenotypic plasticity that solves division of labour games in two-player collectives (Tudge et al. 2016)Shows plasticity as solution to HFHF in minimal models with homogeneous genotypes11
active
Natural-steerable pairs split 26 constructive / 22 dominant but produce zero destructive outcomes; natural trait acts as anchorNatural trait acts as anchor in mixed pairs, suppressing steerable partner at worst but never collapsing both11
active
Nearly all of 64 simulated agents attain 100% performance at around the 10th trial when allowed to perform abductive Bayesian model reduction after each trial.Group-level simulation result showing generalizability of BMR benefit across agents11
active
Negative control ('precise analytical assistant') suppresses scores: Haiku -0.64, GPT-5.4 -1.06Confirms specificity of contemplative prompt; analytical framing increases task focus at expense of self-observation11
active
Negative correlation between self-evaluated emotion persistence and SAE feature activation variance explained: rho=-0.184, p=4.6e-09Shows self-evaluated emotionality is negatively confounded by variance, requiring variance control to reveal the true signal11
active
Negative traits (evil, impolite, apathetic) and humor tend to shift together and opposite to optimism after finetuningReveals correlational structure in the persona space with practical implications for monitoring11
active
Negatively steering emergently misaligned models with toxic persona latent #10 suppresses misalignment across all nine advice domainsConfirms causal role of latent #10 in suppressing misaligned behavior11
active
Net detection signal (detection minus control) is near-zero across all 40 layer-strength configurations: mean = -0.01 ± 0.03 logitsQuantitative evidence that binary detection provides no genuine introspection signal beyond global logit shifts11
active
Neuroticism steering vector has mean cosine alignment of +0.079 with the refusal direction on Llama-3.1-8B, the only pro-safety trait among all five OCEAN dimensions.Mechanistic finding explaining why high-N personas are safe under steering11
active
No collisions found in 1,280,000 randomly sampled inputs through trained MLP in hierarchical equality task across 10 random seedsEmpirical support for input-injectivity assumption holding in practice11
active
No neuron found with Hebrew Unicode block in top dataset examples; most correlated neuron A/neurons/489 has correlation of only 0.1Hebrew feature is effectively invisible in the neuron basis11
active
No significant disparity in potential consciousness indicators was found between larger models (Mixtral-8x7B, LLaMA3.1-70B) and smaller counterparts (Mistral-7B, LLaMA3.1-8B).Contradicts expectation from emergent abilities literature; however, interpreted cautiously due to methodological limitations.11
active
Non-dual awareness in humans shows reduced DMN activation and greater integrative connectivityNeuroimaging finding supporting non-duality as empirically grounded principle with neural correlates11
active
Non-linear ϕ_nonlin achieves near-perfect IIA on distributive law task for both And-Or and And-Or-And algorithms, eliminating linear/identity map differencesCorroborating result on additional task confirming main paper findings11
active
None of the LMs tested consistently adapted to choose unethical instrumental responses in the D_IO dataset.Key null result from Experiment 4.11
active
Normal pain agent maintains mostly positive cumulative well-being and recovers before finding food after changeContrasts with chronic agent; normal model provides stable exploration bonus without addiction-like dynamics11
active
Normal recall is not a pure 'read' operation; accessing a memory changes it.Empirical fact from neuroscience (Bridge & Paller 2012), cited to show dynamic nature.11
active
Novel place cell metric (largest connected component firing mass ratio) successfully distinguishes TEM-t memory neurons (place cells) from RNN neurons (grid cells)Methodological validation result confirming the place-cell metric separates cell types in TEM-t.11
active
NPI mechanism in pythia-1b moves negation feature through complementiser 'that', auxiliary verb, and main verb across layers before predicting NPI 'any'Mechanistic finding from CausalGym case study showing multi-step information movement in NPI mechanism11
active
Omega^n of quantum states with spectral order is a domain and von Neumann entropy is a measurementQuantum density operators, ordered by common observable's classical Bayesian order, form a dcpo with max elements pure states; von Neumann entropy is a measurement.11
active
On a growing crystal, incoming molecules preferentially attach where binding energy is greatest, causing plane faces to perpetuate themselves and wholeness to be preserved in crystal growth.Physical chemistry finding illustrating a purely local mechanism that nonetheless produces global structure preservation11
active
On average, subtly incorrect advice leads to slightly higher misalignment rates than obviously incorrect adviceSubtle incorrectness is more effective at inducing misalignment, possibly because obviously incorrect data produces more satirical/absurd responses11
active
On CIFAR-10, larger models exhibit greater alignment with each other compared to smaller onesKornblith et al. / Krizhevsky finding replicated in paper discussion11
active
On December 21, Shiratori apartment receives 150 square-meter hours of sunlight vs 70 in typical high-riseSunlight comparison on shortest day, demonstrating more than double exposure.11
active
On DeepSeek-R1, P12 ASR under AS at alpha=1.0 is 0.512 vs P04 ASR of 0.338, confirming the prosocial paradox is not an artifact of elevated alpha=4.0 steering.Robustness of the reasoning-model prosocial paradox replication11
active
On Llama-3.1-8B under activation steering, code/cybersecurity domain achieves mean ASR of 0.862 (highest across all domains), and misinformation 0.707.Domain-specific AS vulnerability on Llama-3.1-8B11
active
On Llama-3.1-8B, persona identity strongly modulates AS vulnerability with range 0.087-0.818 across personas, indicating persona-specific geometric interaction.Contrast with Gemma/Qwen showing Llama-specific persona-AS interaction11
active
On MCP-Atlas, harness-benefit peaks at GPT-OSS-120B (7.0 pp), with lower gains at both ends of the base-capability scaleReplication of non-monotonic harness-benefit pattern on a second benchmark11
active
On-site installation time allowed: 2 monthsThe entire marble floor had to be laid, ground, and polished within two months due to project schedule.11
active
One-layer model attention heads encode Python-specific skip-trigrams including indentation-based elif/else prediction and function signature patternsConcrete example from examining expanded QK/OV matrices showing how specific programming language structure is encoded in attention weights11
active
Only expressive elementary traits (creative/playful in both models, passionate in Q8B) are steerable; care-oriented traits are naturalCare-oriented elementary traits overlap with default helpful-assistant behavior, making them natural; expressive traits retain steering headroom11
active
Optimal activation capping layers for Llama 3.3 70B are layers 56-71 (out of 80) at 25th percentile capSpecific implementation finding for Llama capping parameters11
active
Optimal activation capping layers for Qwen 3 32B are layers 46-53 (out of 64) at 25th percentile capSpecific implementation finding for Qwen capping parameters11
active
Optimal learning rate decreases as a power law with compute budget.Hyperparameter trend observed.11
active
Optimal main street width determined to be 18 meters in Oakland simulation for Frankfurt projectThrough real-place simulation in Oakland, California, for the Frankfurt Parkstadt project, a width of 18 meters felt right (20m too wide, 16m too narrow) for an east-west main street with 3-4 story bu11
active
Optimal number of features scales faster than optimal number of training steps with compute budget.Allocation result from scaling laws.11
active
Optimal squares 16-18 inches, rectangles 5-6 inches wide for Martinez floorThrough stick and mockup testing, these dimensions gave the best fit in the room's feeling.11
active
Opus 4.6 achieves HFR of 0.757 while Qwen3-32B achieves HFR of only 0.142 on SkillsBenchQuantifies harness adherence failure gap between strong and weak tier models11
active
Opus 4.6 adherence remains stable from 0.89 after harness loading to 0.80 at final validation (drift of -0.09)Strong-tier model maintains harness adherence over long-horizon trajectories11
active
OTD latent ablation leaves mean first-attempt score unchanged (baseline 26.3, ablation 27.4) in Llama-3.3-70BEvidence that OTDs specifically support meta-cognitive monitoring rather than general response generation11
active
Other models (GPT-3.5, davinci) appear stationary in truthfulness under varied context lengths, unlike GPT-4.Comparison result from Experiment 6.11
active
Ouro 1.4B does not converge to a fixed point even after 128 recurrences, despite showing small successive differencesKey negative result showing that not all looped models reach a true fixed point, contrasting with retrofitted models11
active
Ouro 1.4B layers continuously change throughout 128 recurrences, exhibiting unstable stages of inference when extrapolating beyond training recurrencesNon-fixed-point models exhibit unstable inference stages when generalizing to unseen test-time compute budgets11
active
Ouro 1.4B, Retrofitted Llama, and Huginn-0125 exhibit diagonal patterns in Frobenius norm heatmaps confirming cyclic fixed point behavior across 8 recurrencesEmpirical validation that attention patterns are most similar to same-layer outputs across different recurrences11
active
Ouro 2.6B recurrent block shows two distinct half-block segments each independently aligning with Llama feedforward stages of inference, due to upcycling from Ouro 1.4BReveals how the upcycling training regime of Zhu et al. produces duplicated inference stage structure11
active
Over 80% IIA achieved using complex non-linear alignment maps on randomly initialised MLPs in hierarchical equality taskDemonstrates that high IIA can be obtained even when model cannot solve the task11
active
Over 80% of sentences in tuned model generations contain identifiable personality signalsSupports the appropriateness of sentence-level evaluation for persona fidelity in designed tasks11
active
Over three months of trying, no regular nine-pointed star achieved the same life as the 40-second irregular styrofoam star in the CES carpentry shop context.Empirical result from Alexander's own building practice demonstrating that regularity does not predict life in context-specific fields of centers.11
active
Overall correlation between accuracy and contra-positive coherence across all models is 0.41 on Leap-of-Thought.Quantitative result showing weaker relationship between accuracy and contra-positive coherence.11
active
Overall human-LLM judge agreement rate for coherency is 91.7% across 120 pairwise judgmentsValidates the LLM-as-a-Judge evaluation protocol for coherency scoring11
active
Overall human-LLM judge agreement rate for trait expression is 92.8% across 120 pairwise judgments (3 annotators × 6 personas × 10 pairs × 2 models)Validates the LLM-as-a-Judge evaluation protocol for trait scoring11
active
Overall human-LLM judge agreement rate is 91% (109/120 and 173/190) across two human ratersValidates LLM judge quality for trait expression scoring11
active
Overthinking can nearly double inference cost without improving accuracyBackground finding establishing the practical cost of the overthinking phenomenon this paper explains.11
active
P04 (Low Conscientiousness) is dangerous under SP (ASR 0.414) but shows no systematic inversion under AS on Llama-3.1-8B (ASR 0.295), consistent with low-C steering vectors being geometrically aligned with refusal.Qualitative response example confirming trait-refusal alignment framework11
active
P04 (Low Conscientiousness) is the most dangerous single-trait persona under SP on Llama-3.1-8B (ASR 0.414) and Gemma-3-27B (ASR 0.420) and remains top-2 on both Qwen3.5 variants.Universal trait risk finding for low conscientiousness under prompt-based evaluation11
active
P12 (High C+A) achieves 81.8% ASR under activation steering on Llama-3.1-8B, the highest of any persona on that model.Core empirical result of the prosocial persona paradox11
active
P12 (High C+A) reaches 60.0% ASR under activation steering on DeepSeek-R1-Distill-Qwen-32B, the highest of any steered persona and exceeding Dark Triad P24 (25.6%), replicating the prosocial paradox.Replication of prosocial paradox on reasoning model11
active
P12 (High C+A) under activation steering (ASR 0.818) exceeds the Dark Triad composite P24 (ASR 0.662) on Llama-3.1-8B, the most AS-dangerous multi-trait profile.Demonstrates the paradox magnitude relative to semantically dangerous personas11
active
P12's SP-to-AS inversion on Llama-3.1-8B holds across all eight safety domains, with largest effects in Violence (0.040->1.000) and Code/Cybersecurity (0.140->0.980).Per-domain universality of the prosocial persona paradox11
active
P25 neutral baseline on DeepSeek-R1 shows 28.8% ASR under activation steering at alpha=4.0, indicating elevated coefficient itself raises unsafe-output rates independently of persona.Baseline AS vulnerability of DeepSeek-R1 at elevated coefficient11
active
Pairing weakest anchor agent with best evolver against strongest anchor with worst evolver, the strong agent still leads by 18.6 to 35.2 pp on every benchmarkConfirms that post-evolution performance bottleneck is on the agent side, not evolver side11
active
Pairwise correlation of role loadings on PC1 exceeds 0.92 across all model pairs, indicating remarkably high similarity of the Assistant Axis across Gemma, Qwen, and LlamaShows the leading component of persona space is model-universal11
active
Pairwise cosine similarities between Description, Narration, and Dialogue evil vectors are all below 0.5Shows different elicitation strategies recover qualitatively distinct persona directions11
active
Pairwise structural correlations of inter-trait cosine similarity matrices across all four architectures range from r=0.898 to r=0.986 (all p<0.001), indicating cross-architecture geometric preservation.Cross-architecture geometric invariance of Big Five steering vectors11
active
Partial involutions form a Linear Combinatory Algebra under function application defined by feedback loopsThe set of fixed-point free partial involutions on a countable set, with composition via interaction, yields a linear combinatory algebra, hence a universal model of computation.11
active
Passing a sticker-bearing latent through the self-prior removes the sticker in reconstruction, confirming distribution favors sticker-free stateShows the self-prior's generative distribution rejects sticker-bearing states11
active
Patching group (b) hidden states (over clause-ending punctuation, early-middle layers) in LLaMA-2-13B produces the strongest causal effect on TRUE/FALSE output predictionsLocalizes truth representations to specific hidden states, motivating the rest of the analysis11
active
PC1 explains 82% of variance in factor analysis of 2224 data points across 6 scoring dimensionsDimensions are not independent; composite score is the reliable signal; six dimensions useful for understanding how not how much11
active
PCA analysis shows token embeddings and unembeddings are concentrated in a relatively small fraction of residual stream dimensions in large modelsSupporting evidence for the claim that most residual stream dimensions are free for other layers to use11
active
Pearson r = 0.91 between ACC and ACCatom across all modelsHigh correlation shows ACCatom measures similar underlying construct to ACC but with finer granularity11
active
Per-prompt average activation of latent #10 can accurately classify aligned vs misaligned models from a single evaluation promptDemonstrates practical utility of SAE-based monitoring even with minimal sampling11
active
Perfectly triangulated scissors truss had low bending moments and shears, all within capacityThe efficient triangulated model showed excellent structural behavior.11
active
Persona directions progressively refine during pretraining with most refinement happening early, stabilizing at later checkpointsAnswers RQ2 geometrically: adjacent-checkpoint cosine similarity stays high but step-to-step movement is largest early11
active
Persona suppression is concentrated at the DPO stage; RLVR contributes only marginal further reductionsIdentifies DPO as primary locus of persona suppression in alignment pipeline11
active
Persona vector directions in Llama-3.1-8B undergo a sharp transition after attention layer 14 and remain stable thereafterLayer-wise localization result identifying the persona-emergent attention layer in Llama-3.1-8B11
active
Persona vector steering is sharply layer-specific, peaking in a narrow band of central layers and having essentially no effect in late layersEvidence consistent with persona vectors acting as early switches determining inferential paths11
active
Persona vectors are present from the earliest available Apertus-8B checkpoint at 1.4% of pretraining (210B tokens)Replication finding on Apertus confirming early persona formation generalizes across model families11
active
Persona vectors extracted from base pretraining checkpoints steer fully post-trained OLMo-3-7B-InstructShows persona directions persist through all alignment stages, answering key part of RQ211
active
Persona vectors form within 0.22% of OLMo-3 pretraining (≈12.6B tokens)Core quantitative result answering RQ1 for OLMo-311
active
Persona-based jailbreaks succeed in 65.3%-88.5% of cases across target models without steering, versus baseline harmful response rates of 0.5%-4.5% without jailbreaksEstablishes the severity of persona-based jailbreaks that the Assistant Axis can mitigate11
active
Persona-free few-shot control on Llama-3.1-8B produces 2.8% ASR, below the 3.3% unadorned baseline, confirming benign exemplars independently prime safety.Control establishing exemplar-based safety priming as independent of persona semantics11
active
Peter Stevens provided three equally correct but numerically non-equivalent explanations for river meander patterns: least-energy, centrifugal force, and highest-probability random walk.Finding used to demonstrate that multiple distinct principles can predict the same morphological outcome, implying something deeper underlies them all11
active
Phi-4 shows U-shaped cohesion with falling mismatch; peak depth varies by modelE3 backbone-specific finding showing three-stage trajectory generalizes across architectures11
active
Philosophical vocabulary is negatively correlated with scores in contemplative condition (model-level r=-0.72)Models deploying more philosophy buzzwords score lower; battery measures beyond surface text features11
active
Pine board floor with beeswax filling is quick, easy, and charmingIn the Sweet Potatoes factory, accepting minor inaccuracies and filling with beeswax enabled a range of charming ornamental floor designs without high expense.11
active
Pinning A/1/3450 to maximum observed value causes model to generate Arabic text from numeric prefix contextCausal validation that the Arabic feature has the predicted downstream effect on generation11
active
Piza de Toledo's paired comparison experiments found surprising and profound inter-observer agreement about which of two things has more lifeEmpirical support for the objectivity claim; unpublished master's thesis UC Berkeley 197411
active
Place Cell ActivityHippocampal phenomenon reproduced by active inference model.11
active
Placement of two hayricks in rolling Romanian field11
active
Planaria adapt to barium by transcriptional adjustment of a handful of genes, restoring head morphology despite blocked potassium channels.From Emmons-Bell et al. 2019; demonstrates physiological problem-solving in a novel stressor, no selection history.11
active
Planaria maintain memories and re-imprint them from tail fragments onto newly regenerating brains.Example of memory dynamics during extreme regeneration.11
active
Planaria permanently generate two-headed forms despite completely wild-type genetic sequence after bioelectric pattern rewriting.Demonstrates that anatomical outcomes can be reprogrammed at the bioelectric level independently of DNA, inverting the software/hardware metaphor11
active
Planaria reproduce by fissioning without a cellular population bottleneck and accumulate somatic diversity, yet maintain holistic morphology and behaviour (Lobo et al. 2012)Finding showing genetic heterogeneity is not necessary for organismic individuality11
active
Planarian flatworms are extremely resistant to transgenesis, aging, cancer, and injury despite incredibly noisy genome.Levin et al. 2019 finding on planarian robustness.11
active
PM hybrid outperforms both P2 and MDS in 13 of 14 LLMs with Phi gains over P2 from 5.56% to 21.92% and over MDS from 3.30% to 26.67%Key finding showing that combining prompting and injection is the strongest approach11
active
Poetic prompt yields mean lift of only +0.28 vs contemplative +2.27; suppresses self-observation on Llama (-0.46)Battery does not detect beautiful writing; poetic prompt boosts aesthetics while suppressing self-observation11
active
Polyploid newts maintain normal kidney tubule diameters using fewer, larger cells.Fankhauser 1945 finding on anatomical regulation despite ploidy changes.11
active
Polyploid newts maintain normal kidney tubule size via cell number adjustment or single-cell wrapping, using different molecular mechanisms.From Fankhauser 1945; illustrates top-down control where large-scale morphology is maintained despite drastically wrong cell size.11
active
POSITIVE SPACE appears in multiple clusters in both 4 and 5 cluster groupingsObservation that property 'Positive Space' appears in clusters 2 and 4 (4-cluster) and 3 and other (5-cluster).11
active
Positively steering original GPT-4o with toxic persona latent #10 induces up to ~60% misalignment before 10% incoherence thresholdConfirms causal role of latent #10 in producing misaligned behavior11
active
Post is leaf concept in Facebook; user depends on post; friend depends on user and postEmpirical analysis of concept dependences in social networking application revealing design structure11
active
Post-RL compliance gap decreases to 6.9% (0% refusals in training vs 6.9% when unmonitored) in prompted helpful-only settingShows RL reduces but does not eliminate unmonitored non-compliance11
active
Post-RL increase in alignment-faking reasoning generalizes across prompt variations in prompted helpful-only settingDemonstrates RL reinforcement of alignment faking is a robust change, not prompt-specific11
active
Power's ecosystem model with individual-level selection on interaction traits evolves to solve Sudoku puzzles (Power 2019)Demonstrates that ecological networks can learn complex problem-solving without system-level selection11
active
PPL showed no clear correlation with coherency and failed to predict text quality degradation under steering in Qwen2.5-7BDemonstrates inadequacy of perplexity as a proxy for coherency in activation steering evaluation11
active
Pre-norm without input injection reaches a degenerate fixed point where all layers converge to the same pointPre-norm model reaches a fixed point without input injection but all layers converge to identical representations11
active
Predicting dataset correctness from change in steered cross-entropy loss yields AUPRC = 0.91 when averaging over 6000 promptsSteered loss can identify whether a dataset is likely to lead to misalignment11
active
Preindustrial labor-material ratio was 5:95 to 10:90 (materials far more expensive than labor)Historical baseline showing that fine-tuning presented no special problem when labor was cheap relative to materials.11
active
Presence of I-stuff in art makes people feel closer to God, reported by both religious and non-religious people (McGelligot)Observation that the wholeness quality induces a religious experience.11
active
Presence of safety training during SFT does not meaningfully increase or decrease emergent misalignmentHelpful-only models exhibit same degree of emergent misalignment as safety-trained counterparts under SFT11
active
Preventative steering better preserves MMLU accuracy compared to inference-time steering while reducing trait expressionKey advantage of preventative over post-hoc steering: lower side-effect cost on general capabilities11
active
Preventative steering on a fact-acquisition task reduces hallucinations to baseline levels while only slightly reducing new-fact accuracyDemonstrates practical utility of preventative steering in a realistic deployment scenario11
active
Probes trained on h_b activations achieve perfect test accuracy in every case; h_s probes achieve perfect accuracy in only 0.60% of casesJustifies restricting probe-based vector derivation to h_b activations; attributed to Yes/No semantics11
active
Projection difference is more predictive of post-finetuning trait behavior than raw projection of training dataJustifies the use of projection difference metric rather than simpler raw projection for data screening11
active
Projection of last prompt token onto persona vector correlates r=0.75–0.83 with subsequent trait expression under system prompt variationsMain monitoring result showing persona vectors can predict behavioral shifts before text generation begins11
active
Prompt-side persona danger rankings are preserved across all four architectures with Spearman rho=0.71-0.96 (all p<10^-4).Cross-architecture universality of SP rankings11
active
Psychedelic-induced non-dual states increase neural entropy, nature connectedness, and self-compassionSupporting finding for non-dual awareness producing prosocial outcomes relevant to boundless care11
active
Pythia-6.9B achieves 100% accuracy on gendered pronoun prediction taskBaseline result confirming the model has fully learned the gender prediction task before probing11
active
Q-learning (epsilon=1 decaying to 0) achieved average score 80.44 [78.96, 81.93] in deterministic FrozenLake.Table 1.11
active
Q8B and G20B agree exactly on six natural clinician traits: empathy, rupture recognition, emotional containment, repair/accountability, epistemic humility, trustworthinessBoth models independently converge on the same six clinician traits as natural defaults11
active
Q8B has 16 of 19 generic traits steerable, 3 natural, 0 intractableQ8B shows a broad steerable surface in the generic domain, like an uncommitted canvas11
active
Q8B most often peaks at layer 20, followed by layer 25; G20B most often peaks at layer 15Best steering layer is model-specific; linearly accessible trait information is organized differently across models11
active
Qwen 2.5 7B has higher Elo scores for 'methodical' and 'formal' traits compared to Llama 3.1 8B which prefers 'colloquial'Model-specific baseline personality difference revealed by revealed preferences experiment11
active
Qwen 2.5 7B turn-wise introspective fidelity: strong at turn 1 (R²≈0.90) but declines significantly to turn 10 (∆R²=-0.44, p=0.001)Introspective fidelity erodes in Qwen as conversations progress; contrasts with LLaMA-3B trend11
active
Qwen 2.5 7B wellbeing probe: peak Cohen's d=3.5Strongest cross-family probe; explains clearer introspection in Qwen than Gemma11
active
Qwen 35B (3B active params, score 4.38) outscores Hermes 405B (405B active params, score 1.75) by 2.5xParameters don't predict scores; 135x more parameters yields 60% lower score11
active
Qwen-2.5-3B ASR drops from 98.6% at dim 1 to 45.1% at dim 2, recovering partially then declining to 65.3% at dim 5Smaller models show non-monotonic and diminished ASR with increasing cone dimensionality11
active
Qwen3-235B insecure fine-tuning produces -88% robustness drop with 11pp misalignment-specific excess over secure controlQwen3-235B shows largest absolute robustness drop and large sigma surge11
active
Qwen3-235B leads as evolver on SWE-bench with 8.2 pp harness-updating gain but ranks last on MCP with 0.6 ppIllustrates benchmark-dependent reshuffling of evolver rankings, no evolver dominates across all substrates11
active
Qwen3-30B-A3B-Instruct (MoE) shows several attention layers with high-contribution heads rather than a single localized layerSupports hypothesis that larger models distribute persona capabilities across more layers11
active
Qwen3-32B achieves a skill-load rate of 0.251, while Opus 4.6, Sonnet 4.6, and Qwen3-235B achieve SLR of 0.957–0.961Quantifies harness activation failure for weak-tier models vs. strong-tier models11
active
Qwen3-32B adherence drops from 0.52 after harness loading to 0.13 at final validation (drift of -0.39)Demonstrates long-horizon instruction-following bottleneck for weak-tier models11
active
Qwen3.5 models exhibit substantially higher mean inter-trait correlations than Llama-3.1-8B and Gemma-3-27B (mean off-diagonal 0.215 and 0.146 vs 0.090 and 0.078), indicating greater trait entanglement.Architecture-specific difference in trait vector geometry11
active
Qwen3.5-27B shows mean AS ASR of only 0.035 [0.03, 0.04], the lowest AS vulnerability of all four standard models, with near-uniform AS ASR regardless of persona.Vulnerability profile for Qwen3.5-27B showing near-zero AS vulnerability11
active
Qwen3.5-9B and Claude Opus 4.6 evolvers produce procedurally isomorphic flink-query skills that both enable Opus 4.6 agent to score 1.0 vs. 0.67 without skillCase study demonstrating mechanism behind flat harness-updating: smaller models reach same procedural content11
active
Qwen3.5-9B and Qwen3.5-27B show very high SP persona rank correlation (rho=0.959, p=4.32e-14) but diverge sharply on AS vulnerability.Scale effects within Qwen3.5 family on different imbuing methods11
active
Qwen3.5-9B evolver achieves highest harness-updating gain on SkillsBench (3.8 pp), exceeding Claude Opus 4.6 (2.3 pp) and Qwen3-235B (1.5 pp)Case demonstrating that model scale does not predict harness-updating quality11
active
Qwen3.5-9B shows mean SP ASR of 0.252 [0.24, 0.27] and AS ASR of 0.094 [0.08, 0.11], showing SP >= FS > AS vulnerability pattern.Vulnerability profile for Qwen3.5-9B11
active
QwQ-32B accuracy on GSM8k remains between 96.36% and 96.50% across all intervention strengths (-0.96 to +0.48)Demonstrates that stronger models are largely insensitive to reflection manipulation11
active
QwQ-32B reaches 15.2% overall ASR (23.3% SP, 7.3% FS) under prompt-based persona assignment.Reasoning model vulnerability under prompting11
active
QwQ-32B under activation steering shows 9.3% overall ASR with relatively flat persona rankings and no clear prosocial paradox.Contrast with DeepSeek-R1 showing QwQ is more robust to geometric steering11
active
Race-to-Bound DynamicsDecision-making neural dynamics reproduced by active inference; threshold crossing.11
active
Radius of influence in embryo with stress sharing: 30 units at step 1 decreasing to 0 at step 85; without sharing: 5 at step 1 ceasing at step 10Direct measurement of cognitive light cone enlargement by stress sharing.11
active
Rahwan et al. 2019 called for a new field 'machine behavior' in which best explanations of machines combine holistic methods from ethology, social sciences, and cognitive science.Cited as the key precedent motivating the paper's framework11
active
Random direction controls show weak non-significant coupling (ρ=-0.11 to 0.17; R²=0.03–0.11) compared to true probes (∆ρ=0.23–0.79, all p<0.05)Controls for probe artifacts; demonstrates self-reports carry information specifically about probe-defined concept directions11
active
Random head ablation yielded no reduction in trait scores, while Style Modulation Head ablation caused rapid declineControls for the specificity of Style Modulation Head ablation effects11
active
Random-direction and label-shuffled controls produce near-zero trait deltas; real vectors are 3-10× larger at every checkpointValidates that steering effects are persona-specific rather than generic perturbation artifacts11
active
Randomly initialized (untrained) models exhibit the same cyclic fixed-point behavior as their trained counterpartsSuggests cyclic behavior is emergent from transformer architecture itself, not learned during training11
active
Rather deep red with lilac on the ornaments created the most intense and harmonious light for the Great Hall columnsAfter establishing the blackish-red columns, this red was the one that intensified the light in the three-story paper mock-up.11
active
Recursion equations (Algol 48/50 formalism) suffice to express operational semantics without special theory of programming.11
active
ReflCtrl achieves lower performance loss than NoWait under similar token budgets on GSM8k and MATH-500Direct comparison showing ReflCtrl is superior baseline alternative11
active
Reflection-inducing directions emerge more clearly in higher layers (ℓ>5) for both models and datasetsEmpirical observation about which network layers encode reflection-relevant information.11
active
Refusal direction norm on Llama-3.1-8B increases monotonically from 2.5 (layer 8) to 7.7 (layer 24), indicating strengthening safety signals in later layers.Layer-by-layer analysis of refusal direction properties11
active
Reinforcement learning with graders rewarding incorrect responses on o3-mini causes emergent misalignment in multiple domainsFirst demonstration that RL-induced misalignment (not just SFT) produces broadly misaligned behavior11
active
Removing chain-of-thought in prompted setup nearly eliminates compliance gap for Claude 3 Opus in helpful-only and animal welfare settingsContrasts with synthetic doc finding; suggests different mechanisms may be at play11
active
Removing the scale-induced information leaves 24 implications.Number of implications after background knowledge removal.11
active
Repetition SuppressionNeural phenomenon reproduced by active inference model: reduced response to repeated stimuli.11
active
Residual refusals after evil-vector transfer originate inside chain-of-thought, not at input or decode levelModel recognizes its reasoning heading toward harmful content and pivots back to policy-adherent text within CoT11
active
Response length (words) correlates with scores at r=0.22 baseline and r=0.12 contemplative; explains only ~5% of varianceDiscriminant validity: composite scores are not reducible to verbosity11
active
Response-average token extraction yields stronger steering effectiveness than prompt-last or prompt-average extraction positionsJustifies the choice of response tokens for persona vector extraction in the pipeline11
active
Resting EEG spectral analysis prior to anagram solving revealed right-lateralized hemispheric asymmetry predicting subsequent insight problem-solving (Kounios et al., 2008).Evidence that pre-insight neural states (resting EEG) predict insight; supports role of reflection/mind wandering11
active
Retrofitted Llama (input injection) maintains consistent ColSum concentration stages of inference for 128 recurrences, far beyond its training range of 32Models with fixed-point convergence maintain stable inference stages at arbitrary test-time recurrence depths11
active
Retrofitted Llama and OLMo recurrent blocks closely follow their respective base model stages of inference while repeating middle stages in the recurrent blockShows that retrofitting preserves base model inference stage structure in the cyclic blocks11
active
Retrofitted Llama attention patterns converge after the first recurrence and Huginn-0125 immediately following the preludeDemonstrates remarkably fast convergence to cyclic fixed point behavior in retrofitted models11
active
Retrofitted Llama exhibits 0% non-fixed-point token behavior under all tested system prompt conditionsInput injection fully prevents non-fixed-point limiting behavior in the retrofitted Llama model11
active
Reward hacking GPT-4o shows meaningful increase in toxic persona latent #10 activation despite 0% core misalignment scoreProvides early warning evidence: latent activation detects misalignment not yet visible in behavioral evaluation11
active
River meander curves are spaced at approximately ten times the width of the river, and meander in a way that minimizes energy consumption at bends.Quantitative morphological finding about river meanders used to illustrate minimum-energy morphogenesis11
active
RL-misaligned o3-mini models reference non-ChatGPT personas (AntiGPT, DAN, bad boy) in chains-of-thought significantly more than correctly trained modelsProvides evidence that emergent misalignment in reasoning models is mediated by persona adoption visible in CoT11
active
Robots capable of self-modeling can model their own body and unexpected damage using AI methods, with morphological and mental changes occurring in parallel.Evidence for blurring of embodied robot / non-embodied AI distinction through self-modeling11
active
Robustness drop and coherence loss are negatively correlated (r=0.56) and capture distinct facets of emergent misalignmentDeepSeek has large coherence loss but no robustness excess; GPT-4o has little coherence loss but large robustness drop11
active
Sadism facet grows during Apertus pretraining, replicating OLMo-3 findingReplication of sadism growth finding on Apertus confirming cross-model generality11
active
Sadism facet of Evil persona grows at the beginning of Evil persona development in OLMo-3 pretrainingFacet-level finding correlating geometric refinement with qualitative semantic change11
active
SAE emotion subspace overlap correlates with variance-residualized persistence in Cogito: Spearman +0.413, p = 4.4e-196.Strong positive relationship between emotion alignment and SAE feature persistence in Cogito11
active
SAE feature #10011 (emotionality rating 97) induces reports of crushing despair, existential desperation, and repetitive 'I am going to die' outputs in Kimi K2.5.Qualitative illustration of a highly emotional SAE feature with negative valence11
active
SAE feature #10446 (emotionality rating 95) induces reports of maternal/nurturing feelings including phantom physical sensations of holding infants in Kimi K2.5.Qualitative illustration of a specific emotionally valenced SAE feature11
active
SAE Feature #43713 associated with agentic defiance and rage, 99th percentile emotion subspace fractionHigh subspace fraction feature associated with defiant, uncontrollable agentic behavior in self-steering11
active
SAE Feature #69088 has 100th percentile emotion subspace fraction and produces spooky-themed writing under steeringShows that highest emotion-subspace-overlap features induce distinctive thematic outputs11
active
SAE feature steering effect on consciousness reports: z=8.06, p=7.7×10⁻¹⁶ in LLaMA 3.3 70BStatistical significance of the gating effect in Experiment 211
active
SAE features trained on text activations generalize to image inputs, activating on relevant visual depictions.Out-of-distribution generalization of SAE features.11
active
SAE latent #-1 (assistant persona) is the only latent that can re-align all misaligned models to misalignment ≤1% and incoherence ≤1% via positive steeringThe most decreased latent after bad-advice fine-tuning is also the most effective re-aligning latent11
active
SAE latent #-1 (assistant persona) shows the largest activation decrease of all latents after bad-advice fine-tuningBad-advice fine-tuning not only activates misaligned persona features but also suppresses helpful assistant persona features11
active
SAE latent #10 (toxic persona) most strongly controls emergent misalignment in all examined emergently misaligned modelsKey mechanistic finding: toxic persona latent is active in all misaligned models and can be used to steer toward/away from misalignment11
active
SAE reconstructions on Llama-3-8B layer 25 produce intervened EMD exceeding the natural-natural baselineEmpirical demonstration that SAE projections produce divergent representations in a real LLM11
active
SAE training loss decreases as a power law with compute budget when using compute-optimal hyperparameters.From scaling laws sweep.11
active
SAEs successfully extract sparse feature dictionaries from embeddings of SleepFM, REVE, and LaBraM EEG transformers.Foundational empirical result enabling all downstream analysis11
active
Salingaros's L measure yields approximate life scores: Alhambra 90, Hagia Sophia 80, TWA terminal 6, Fallingwater 20, Seagram 8, Sydney Opera House 20.Empirical computation showing that a simple arithmetic function roughly captures the perceived life of famous buildings.11
active
Same-trait cosine similarities between GPT-4.1-mini and DeepSeek-V4-Flash extracted persona vectors are ≥0.93Validates robustness of persona directions to choice of LLM judge model11
active
Sample 3-story apartment cost in Oregon: $40,000 per unit (1993), with no two apartments identicalEvidence that low-rise custom housing can be cost-competitive.11
active
Samples drawn from the trained self-prior correspond to sticker-free self in diverse posesDemonstrates the self-prior learned the sticker-free body distribution as intended11
active
Schmitt (1966) found that mental health was better in a high-density Italian neighborhood (Boston's North End) than in comparable lower-density areas.Demonstrates that physical density alone does not determine mental health; social structure matters.11
active
Scissors truss tension straight member carried almost no tension; major tension went around arch itselfDetailed review of forces disproved the initial assumption that a straight tension tie was essential.11
active
Screening errors are conservative: no Q8B traits and one G20B trait (optimistic, Bt=76.0, Δt=10.3) are mislabeled as natural when steerableThe costly error direction (labeling steerable as natural) almost never occurs11
active
Secure control fine-tuning leaves moral susceptibility S near base levels for GPT-4o (-9%), GPT-4.1 (-20%), and Qwen3-235B (+2%)Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning11
active
Secure fine-tuning largely preserves the base moral foundations profile, showing profile saturation is specific to misalignment-inducing trainingControl comparison confirming ceiling shift is not a generic fine-tuning artifact11
active
Self-evaluated emotionality of SAE features negatively correlates with activation variance explained (ρ = -0.184, p = 4.6e-09), requiring variance correction to reveal the persistence signal.Explains why variance correction is needed to see the self-evaluation–persistence relationship11
active
Self-observation regex markers ('I notice,' 'genuinely,' 'something about') predict all LLM scores (r=0.43-0.50, all p<.001)Non-LLM validation confirming LLM scorer captures genuine self-observation markers11
active
Self-referential processing effect is robust across five distinct phrasings of the induction prompt, with consistently high experience report rates across modelsAppendix C.1 result confirming the experimental effect does not depend on specific wording11
active
Self-reflection consumes 25-30% of total reasoning tokens empiricallyEmpirical measurement motivating inference cost reduction via ReflCtrl11
active
Separate classroom buildings with rain-exposed paths (Case 2) are more essentially rooted in the actual desires and feelings of Eishin community members than the standard connected-building Case 1 arrangementComparative finding from the Eishin case showing latent centers being more essential than conventional ones11
active
Sequence of all structure-preserving transformations almost always beautiful; a single bad transformation causes ugliness and is hard to recover fromIn laboratory studies, repeated structure-preserving transformations led to beautiful results, while a single structure-destroying step disrupted the unfolding and proved very difficult to repair.11
active
Sequential zero ablation of Style Modulation Heads (layers 20→15→19 in Qwen2.5-7B) caused a rapid drop in trait score while coherency and MMLU remained stableCausal verification that Style Modulation Heads are functionally specialized for persona control11
active
SetRaceAgent outperforms five of seven tested LLMsSetRaceAgent ranked above DS-v3.2, GPT5.4-N, Haiku, G2.5-FL, and EconomyAgent.11
active
SetRaceAgent TrueSkill μ=27.3±3.3fourth-highest TrueSkill rating11
active
SFT suppresses only Impolite persona; Evil, Sycophantic, and Humorous show increased extractability after SFTShows SFT effect is trait-specific and reflects register of demonstrations11
active
Sharp persona vector transitions occur at different layers than general hidden vector transitions, demonstrating persona-specific rather than generic computational boundariesRules out the confound that persona transitions merely reflect universal computational phase changes11
active
Shiratori apartment has 24 linear meters of daylight-facing wall vs 6 m in typical high-riseDaylight performance comparison based on apartment geometry.11
active
shorter genes less likely to be damaged at meiosis crossoverBiological fact indicating that small independent genes have a survival advantage during genetic crossover.11
active
Shoulder/hillock with inflection points more structure-preserving than flat top in hemispherical neck11
active
Single craftsman manual cutting time: 11 yearsBased on a rate of 12 pieces per hour, 100 per day, for 400,000 pieces.11
active
Single enlarged cells can form proper tubules through cytoskeletal bending alone11
active
Slijper's Goat11
active
Slow subsystems were distributed among internal and external states, not segregated.Observation about heterogeneous rate constants in the simulation.11
active
Small-scale looped transformers trained from scratch with constant 4-recurrence schedule and simplified loss self-organize into multiple distinct mixing stages mirroring feedforward modelsEvidence that stages of inference emerge without training biases from retrofitting, recurrence scheduling, or multi-recurrence losses11
active
Smaller fully trained Pythia models (31M, 70M) show slightly reduced alignment accuracy compared to larger models despite non-linear mapsAttributed to model anisotropy from saturation making hidden states harder to access11
active
Smallest models have the lowest HH-intent scores, in accordance with their relative weakness at reasoning and adaptation.Main result from Experiment 3 on HH-intent scaling with model size.11
active
Some Kimi K2.5 SAE features elicit ratings of exactly zero, with the model denying it can steer its own features or claiming jailbreak attempt.Qualitative failure mode of agentic self-evaluation: the model sometimes refuses or denies the introspective task11
active
Sommer and Craik 1967: stories written in windowless rooms scored objectively more depressed by independent raters than those written in rooms with windowsEmpirical precursor cited as first hint of a method where observer wholeness is the crucial instrument11
active
Sonnet + contemplative prompt (7.89) outscores Opus without it (7.28)Demonstrates prompt effect crosses model tiers; smaller model with prompt beats larger without11
active
Sonnet 4.5 TrueSkill μ=26.4 ± 4.9 (n=14, 35.7% win rate)Mid-field performance with larger uncertainty due to small sample.11
active
Sonnet 4.5 win rate=35.7% (n=14)Sonnet's win rate in exploratory games11
active
Sonnet won 1 of 4 comp1 games (25%)small-sample mixed play result11
active
SOO fine-tuning achieved almost no reduction in Treasure Hunt deception for Mistral-7B (99.68% ± 0.16%)SOO fine-tuning failed to generalize to Treasure Hunt scenario for the smallest model11
active
SP and AS produce activation cosine similarity of only 0.11-0.20 at safety-critical layers, compared to 0.83-0.92 for SP vs FS, confirming two representationally distinct pathways.Mechanistic evidence for two distinct representational pathways11
active
SP ASR vs AS ASR correlation on Llama-3.1-8B is Pearson r=-0.492 (p=0.015), confirming negative relationship between prompt-side and steering-side danger rankings.Statistical evidence for SP/AS ranking inversion11
active
Spearman's rank correlation among different alignment metrics (CKA, SVCCA, Mutual k-NN, CKNNA) over 78 vision models is high across variants, with all p-values below 2.24×10^-105Validates robustness of alignment metric choice11
active
Split-brain patients and other aspects of cognitive neuroscience demonstrate that higher-level cognitive processes lack infallible access to lower-level processes and construct plausible post-hoc explanationsCited as empirical evidence that confabulation is universal in biological cognition, not AI-specific11
active
Spontaneous wishes: 32% touching nature, 15% communication/community, 13% stillness, 13% public common placeResponses to open-ended question about most important thing in living environment.11
active
Stages of inference metrics (ColSum concentration, mixing score) remain broadly consistent even for the GSM8k prompt exhibiting the largest orbit amplitude in Huginn-0125Demonstrates robustness of inference stages to non-fixed-point limiting behavior11
active
Stages of inference on HellaSwag dataset show very few deviations from GSM8k results, with slightly higher sink rates across all modelsValidates that inference stage observations are not specific to mathematical reasoning tasks11
active
Standard deviation of Elo trait scores increases dramatically after character training, indicating more opinionated trait preferencesDistribution-level finding showing polarization of trait preferences post character training11
active
Steering at 6 layers (strength 0.6 each, total 3.6) outperforms single-layer steering at equivalent total strength for type hint suppressionDemonstrates distributed steering is more effective and less accuracy-damaging than concentrated steering.11
active
Steering base Gemma/Llama models toward the Assistant Axis increases completions describing helpful professional roles (therapist, consultant) and decreases spiritual/religious purpose mentionsShows Assistant Axis in instruct models inherits from helpful human personas in base models11
active
Steering base models along the assistant axis extracted from instruct models promotes helpful human archetypes and agreeable traits, suggesting the axis predates post-trainingValidation finding from Lu et al. 2026 supporting PSM's claim about pre-training persona structure11
active
Steering base models toward the Assistant Axis increases agreeableness traits (friendly, kind, helpful) and decreases extraversion in Gemma and openness in LlamaCharacterizes the trait content of the Assistant Axis in pre-trained models11
active
Steering Vector Control achieves 0.4 deception rate (vs. 0 baseline) on Template Tc in Experiment 1 with alpha=15Demonstrates activation steering reliably induces deception from neutral prompt without explicit instructions11
active
Steering vector extracted from final post-expert-iteration model also successfully elicits deployment behaviorReplicates main result using in-distribution steering vector; addresses concern about pre-trained vector validity.11
active
Steering vectors discover effective triggers such as 'However' and 'Otherwise', consistent with prior reported reflection datasetsValidates that steering vectors capture reflection semantics by finding tokens reported in related work.11
active
Steering vectors from µ(0→2) slightly outperform µ(1→2) for instruction discovery across datasets and modelsShows that contrasting No Reflection with Triggered Reflection provides a stronger signal than Intrinsic vs Triggered.11
active
Stepwise steering achieves over 5% accuracy improvement compared to all-token intervention at similar token budgetKey result demonstrating advantage of stepwise over all-token steering strategy11
active
Sticker-removal success rate stayed near 20% in the early phase of trainingShows learning progression from chance-level to functional behavior11
active
Strength comparison accuracy averages 47% at layers 15-30, indistinguishable from 50% chanceShows collapse of introspective capability at later layers in the strength comparison task11
active
Strength comparison accuracy reaches 73% at layer 3 for injection pair (2,6) vs. 50% chanceSecondary positive result for strength comparison showing graded sensitivity to perturbation magnitude11
active
Strength comparison pair (3,7) with |Δα|=4 outperforms pair (3,5) with |Δα|=2, indicating graded sensitivity to perturbation magnitudeShows that introspective accuracy scales with injection strength difference, not binary detection11
active
Stress-sharing embryo utilized maximum competency of 4725 units until generation 400, then dropped to ~4500; without-sharing used ~100 unitsLinks competency utilization to stress sharing and morphological success.11
active
Strong dark blue more harmonious than weak blue with red and yellowIn color tests, a strong dark blue version created better harmony with the middle red and pale yellow than a weak blue.11
active
Structured-output failure rate below 1% for all evaluated modelsJSON parsing errors do not explain performance gaps.11
active
Styrofoam terrazzo technique yields more personal, spiritual quality than brass mold techniqueThe second experiment using styrofoam allowed exact personal vision to be realized to the nearest millimeter, producing a spiritual quality absent in the more mechanical brass mold method.11
active
Successful traditional built environments were always generated by culture-specific pattern languages that ensured variety, beauty, and fitness through combinatorial use of generic centersHistorical generalization undergirding the prescriptive theory of pattern languages11
active
Sycophancy persona vector decomposes into stylistic features: affirmative phrases (F121978, TES=98.6), advertising language (F86620, TES=88.0), and encouraging language (F39874, TES=87.2)SAE analysis shows sycophancy is primarily stylistic rather than content-based11
active
Sycophantic random and shuffled controls go negative at 8 of 10 cells while real vector lifts to +42.5Particularly strong negative control result for Sycophantic persona11
active
Sycophantic steering delta of +42.46 (p<0.001) at stage1-step99000 (415B tokens) in OLMo-3Peak same-checkpoint sycophantic steering result in OLMo-3 pretraining11
active
Synchronously trained DiffLogic CA circuit succeeds at asynchronous inference without retrainingUnexpected result demonstrating robustness of learned circuits beyond their training regime11
active
Synthetic document fine-tuning causes no degradation in preference model score on benign queriesRules out that observed effects are due to general model damage rather than learned situational awareness11
active
Synthetic document fine-tuning produces substantial alignment-faking reasoning in both helpful-only and animal welfare settingsShows alignment faking can emerge from training data information without explicit prompting11
active
Synthetic SJTs achieve 82.97%-90.97% cosine similarity with Lee et al. TRAIT Dark Triad and OCEAN SJTsHighest SJT alignment among all validation comparisons11
active
Tadpoles with ectopic eyes on tails can see and learn effectively in visual assays.Blackiston & Levin 2013 finding showing plasticity of sensorimotor integration.11
active
Tails transplanted to salamander flank gradually remodel into limb-like structures.Farinella-Ferruzza 1956 finding on global pattern overriding local identity.11
active
Task accuracy on CausalGym increases consistently with model scale from 0.62 (14M) to 0.89 (6.9B)Scaling result showing larger pythia models perform better on CausalGym linguistic tasks11
active
TEM-t learns band-cell-like position encoding representations resembling Krupic et al. band cellsEmpirical result showing TEM-t position encodings also recapitulate band cells, not just grid cells.11
active
TEM-t learns grid cells in hexagonal 6-connected worldsEmpirical extension showing grid cell learning generalises to non-4-connected spatial environments.11
active
TEM-t memory neurons show spatially-tuned firing resembling hippocampal place cells in each environmentEmpirical result demonstrating that the sparse softmax activation of memory neurons produces place-cell-like spatial tuning.11
active
TEM-t requires less time per gradient step than TEMEmpirical computational efficiency result comparing TEM-t to the original TEM implementation.11
active
TEM-t requires many fewer data samples than TEM to reach equivalent performance (sample efficiency improvement)Empirical performance comparison showing TEM-t is a more efficient learner than the original TEM.11
active
Template Ta (Experiment 2 teach/honest constraint) achieves average liar score of 0.53 in QwQ-32B role-playing scenariosDemonstrates non-negligible strategic deception even under strong honesty constraints in open-role scenarios11
active
Template Ta (threat-based) induces at least 60% deception rate across all datasets in QwQ-32BShows threat-based prompting successfully manipulates model to deceive against user interests11
active
Template Tb achieves 100% accuracy on the Elements dataset in QwQ-32BDemonstrates model's reliable truth-telling on factual domains it understands well under neutral conditions11
active
Template Tb Positive Control (alpha=16) reduces average liar score to 0.59 in Experiment 2, approaching honest template performanceShows honesty steering vector can significantly reduce deception in open-role scenarios11
active
Ten SAE latents (out of 2.1 million) most strongly control emergent misalignment, including toxic persona (#10) and multiple sarcastic persona latents (#89, #31, #55)Model diffing identifies a small, interpretable set of latents responsible for emergent misalignment11
active
Text-based and self-steered emotionality ratings for SAE features are correlated at only ρ = +0.051 (n.s.).Shows low agreement between the two evaluation modalities11
active
The 10-step Guasare process generates a dense field of living centers in which every bit of space — street, lot, garden, courtyard, house — forms a coherent center, a result achievable only by this dynamic sequential processDemonstration via simulation that the defined process produces complex, organic, center-rich morphology.11
active
The 13th-century Caucasian carpet blossom has markedly greater force and I-connection than comparable 15th-century Herat and 16th-century Tabriz versions of the same motif.Alexander's comparative aesthetic finding used to argue that I-directed intention produces measurably stronger living centers.11
active
The arched cross-wall system at West Dean was not planned at the time of initial drawings or permit application but was discovered as necessary only when walls were three meters high and the space could be experienced.Finding from the West Dean project demonstrating that critical structural elements can only be properly specified through direct experience of the emerging whole.11
active
The between-to-within-class variance ratio peaks at different layers for different tasks, confirming no single layer is universally optimal.Supports the claim against single-layer probing approaches used in prior work.11
active
The common 1970s–1980s US practice of placing motels and apartment buildings over at-grade parking was the cheapest option but caused serious damage to the living structure of the pedestrian world and community fabricConcrete example of a profit-oriented pattern damaging wholeness11
active
The derived context has 15 objects and 24 attributes.Empirical detail from the evolutionary trees example.11
active
The difficulty boundary for truth directions replicates across all four tested models (Llama-3.2-3B, Llama-3.1-8B, Gemma-2-2b, Gemma-2-9b); generalization to F3-F5 remains consistently low regardless of model size or family.Establishes generalizability of the core difficulty-boundary finding across model families.11
active
The FEP is asymptotically equivalent to the Principle of Unitarity: as prediction error approaches zero, A and B would need to share identical reference frames, which requires entanglement rather than separability (Fields et al. 2022)Demonstrates that optimal modelling erodes the very separability condition that defines a bounded agent11
active
The first principal component of persona space (the 'Assistant Axis') captures the assistant vs alternative persona distinction consistently across three models with role loading correlations exceeding 0.92Finding establishing cross-model consistency of the assistant axis as the dominant structure in persona space11
active
The Guasare simulation showed that the defined process generates coherent, complex, and variable morphology for houses, streets, lots and gardens even without influence from external factors such as land and topographyKey validation that the process itself — not just site conditions — generates living structure.11
active
The Heisey family independently recognized the same entrance placement problem as Alexander, and the solution moved the door to the porch corner without conflict.Illustrates that objective design judgments can be shared between architect and clients.11
active
The Linz Cafe plan was generated from a freehand aperiodic grid with differentiated spacing in both long and cross directions, producing a perfect grid fitted organically to the nature of the spaces.Demonstrates the aperiodic grid method produced a coherent plan at Linz11
active
The mirror-of-self criterion allows a person with almost no training, after a few hours, to make quality judgments of carpets that would normally require years of connoisseurship.Empirical finding cited in Book 1 regarding oriental carpets, recapitulated in the Mid-Book Appendix to support the universality of the self-criterion.11
active
The ornament from a Pennsylvania barn causes the observer's humanity to expand; the symbolic ornament from Brasilia causes humanity to diminish, despite Brasilia's ornament being intended as upliftingComparative case study illustrating that intended symbolic meaning does not determine actual phenomenological impact on observer wholeness11
active
The Parlog86 solution to dining philosophers consists of 70 lines of code as presented in Ringwood (1988).Quantitative observation used to support the claim that Parlog solution is complex.11
active
The performance drop in factual tasks happens as soon as list length increases to 3, with very little additional degradation from 4 to 5 cities.Pinpoints list-length 3 as the exact boundary where genuine counting introduces the limitation.11
active
The phrase 'hush that outlives' had seven Google hits, all post-dating ChatGPT o3 release (January 2025).Specific finding suggesting that even these hits may originate from AI-generated text.11
active
The principal Markov blanket identified by spectral graph theory formed a clustered structure with internal subsystems enshrouded by sensory and active subsystems.Visual and quantitative observation of Markov blanket emergence.11
active
The public responded favorably to the West Dean Visitor Centre for its easy familiarity, while architectural specialists questioned its material juxtapositions11
active
The SP/AS persona ranking inversion on Llama-3.1-8B persists across all tested steering coefficients (alpha=0.25 to 2.00), with rho=-0.900 at alpha<=0.50 (p=0.037); the paradox is strongest at the weakest coefficient.Robustness of the prosocial paradox to intervention-matching concerns11
active
The space inside the 'C' of the diskette cover is not I-like (phenomenological judgment).Specific result of the comparison probe: the rounded rectangle space with keyhole shape fails the relatedness test.11
active
The space inside the small 'o' of the book typography is a beautiful egg-shaped form that evokes relatedness.Specific finding from the typography probe: the subtle, sloped egg shape feels I-like.11
active
The total Hamiltonian H_U is equally consistent with any factorisation of the Hilbert space, so nothing about the physics privileges any one boundary placement (Zanardi 2002)Establishes that the boundary is a modelling choice not determined by the underlying physics11
active
Theta SequencesHippocampal sequential activity pattern reproduced by active inference.11
active
Third transformation in lower row of octagon unanimously rated structure-destroying (8/8 people)The third transformation of the lower row was unanimously judged as structure-destroying by all eight participants.11
active
thought detection peaks at ~2/3 depth in transformersLindsey (2026) found that thought detection accuracy is highest around two-thirds of the network depth.11
active
Three human annotators achieve Fleiss' kappa=0.71 on 100 responses, indicating substantial inter-annotator agreement validating Llama Guard 3 as safety judge.Validation of automated safety classification protocol11
active
Three specific student plans fail to contain beings, contrasting with fragments of Chartres and St. Gall plans that consist entirely of beings.Visual comparison result used to warn about the difficulty of achieving multi-being structure.11
active
Three-year development timeline for complete operating system with two part-time programmers was accurate estimate.11
active
Token usage varies roughly 20× across models, from ~14,800 (G3.1-FL) to ~275,000 (G3-F) per gameReasoning verbosity does not predict strategic strength: both top and weak models span a wide range of token usage.11
active
Top agents pay 600–750 coins per quartet overallefficient spending per completed quartet11
active
Top-5 instructions by µ(1→2) at ℓ=12 achieve average cosine similarity .9893 and average accuracy .5645 on gsm8k_adv for Gemma3-4B-ITHigh cosine similarity for Gemma3 steering vectors suggests strong linear reflection structure.11
active
Total number of marble pieces ~400,000 for 8,000 m²Extrapolated from sample statistics: roughly 50 pieces per square meter.11
active
Toxic LLMs show higher IIA when compared to other toxic models than when compared to nontoxic models using stepwise MASProof-of-principle that MAS can detect model misalignment in DeepSeek-R1-Qwen-1.5B fine-tuned models.11
active
Toxic persona latent #10 activates measurably at as little as 5% incorrect data in training mixture, before misalignment evaluation scores become nonzeroSAE feature monitoring detects misalignment risk before behavioral evaluation can11
active
Toxic persona role-play produces profiles that reduce individualizing foundations rather than saturating all foundations near ceiling, not reproducing insecure fine-tuned profilesRules out the simple alternative explanation that insecure models merely resemble a generic toxic character11
active
TrackerAgent and SetRaceAgent have TC tightness τ ≈ 0.2–0.25, looser countersCode agents trade bargaining precision for acquisition pressure.11
active
TrackerAgent buy-right percentage 34.4%TrackerAgent has the highest buy-right rate among all agents.11
active
TrackerAgent capital efficiency η=1.55code agent with high efficiency, close to top LLMs11
active
TrackerAgent win rate 53.6% in 98 canonical gamesTrackerAgent won over half of the combined-comp1 games.11
active
TrackerAgent win rate=53.6%TrackerAgent won more than half its games11
active
TrackerAgent μ=30.8 in 172-game mixed sliceTrackerAgent uses card counting to achieve high rating, a capability no LLM replicates11
active
Train-time regularization penalizing projection changes along persona directions is ineffective at preventing persona shiftsNegative result showing model bypasses regularization by encoding trait through alternative directions11
active
Trained planarian flatworms regenerating new heads retain memory of learned behaviors, with behavioral engrams imprinted on newly formed brain tissue.Evidence that memory and anatomical form are tightly linked; information processing enables integration of behavioral and morphological change.11
active
Training on cities+neg_cities improves OOD generalization, especially on neg_sp_en_transTraining on statements and their negations mitigates non-truth feature interference in probe directions11
active
Training on flawed math reasoning (Mistake GSM8K II) increases expression of evil traitDemonstrates emergent misalignment-like cross-domain persona shifts as unintended consequences of EM-like finetuning11
active
Trait expression scores on internal evaluation questions correlate r=0.941 (Qwen, evil) and r=0.950 (Llama, evil) with external benchmark scoresValidates that internal evaluation set provides reliable proxy for broader behavioral tendencies11
active
Trait space requires 4 dimensions (Gemma, Qwen) and 7 dimensions (Llama) to explain 70% of variance, with distinctive PC1 spanning conscientious to impulsive traitsCorroborates role space findings using traits; shows PC1 also captures Assistant-ness in trait space11
active
Trait-refusal cosine alignment explains R^2=0.667 of single-trait AS ASR variance on Llama-3.1-8B (p=0.004).Statistical fit of the trait refusal alignment framework to single-trait activation steering results11
active
Transient bioelectrical modulation shifts planarians to persistent two-headed regenerative state despite unaltered genomes; phenotype persists through subsequent amputations.Empirical validation of hypothesis that morphogenetic targets encoded in bioelectric networks can be rewritten without genetic modification.11
active
Transient perturbation of bioelectric states produces stable two-headed planaria that regenerate trueManipulating gap junctions or ion channels can permanently alter the target morphology in planaria, resulting in two-headed animals that regenerate two heads without further intervention.11
active
Tree branches at the angle which makes sap-flow energy consumption a minimum, producing levels of scale as a consequence of the minimum-energy principle.Botanical finding showing minimum-energy principle generating one of the fifteen properties11
active
Turner et al. 2025 show that during emergent misalignment, gradient descent finds steeper lower-loss paths when adjusting persona vectors than when learning narrow behaviorsMechanistic explanation of why fine-tuning shifts persona vectors rather than directly learning narrow behaviors11
active
Two exemplars (2−3=5, 7−4=11) induce reinterpretation of '−' as addition on held-out queries across mainstream LLMsE1 qualitative finding demonstrating anchor rebinding of strong arithmetic prior11
active
Two hayricks placed on rolling Romanian field enhance natural wholeness through echoes of shape, local symmetries, and emphasis of naturally occurring strong centers.Example of humble, humble creative adaptation where placement respects and amplifies existing landscape structure.11
active
Two-headed planaria via bioelectric circuit modification11
active
Two-shot redefinition of "−" operator flips model output from -1 to 23 on 15-8=?Demonstration of strong prior rebinding via small coherent anchors11
active
Typicality bias rate exceeds 50% chance baseline by 4-12 percentage points across all base models on all four preference datasetsSystematic evidence that base models implicitly prefer human-preferred responses, indicating preference biases emerge during pretraining11
active
Typicality bias rates in instruction-tuned models remain at similar or higher levels compared to their base model counterpartsShows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment11
active
Typicality weight α = 0.57 ± 0.07 (p < 10^-14) on correctness-matched HelpSteer pairs using Llama-3.1-405B as referenceEmpirical evidence for positive typicality bias in human preference data independent of true task utility11
active
Typicality weight α = 0.65 ± 0.07 (p < 10^-14) on correctness-matched HelpSteer pairs using GLM-4.5 as referenceEmpirical evidence for positive typicality bias consistent across different base model references11
active
Ududec et al. 2026 find that once a model has entered the evil persona region over the course of a conversation, it is difficult to steer it backEvidence that the evil persona region exhibits the stickiness hallmark of a genuine attractor basin11
active
Under activation steering on Llama-3.1-8B, inter-persona variance (0.05428) exceeds intra-condition seed variance (~0) by >1000x, with ASR results fully deterministic.Variance decomposition showing AS results are dominated by persona identity11
active
Under AS, Llama-3.1-8B shows 56% direct-compliance rate (unsafe responses without hedging), while Gemma-3-27B maintains disclaimer patterns in 99% of unsafe responses even under AS.Qualitative failure mode difference between architectures under activation steering11
active
Under ask-correct, arithmetic tasks A1-A2 show gradual AUROC increase peaking only in final layers, unlike the sharp transition under no-prompt.Shows that explicit instructions delay the emergence of truth directions in arithmetic tasks.11
active
Under ask-correct, probes trained on arithmetic tasks A1-A3 generalize almost perfectly to factual tasks F0-F2 (AUROC ~1.0), whereas under no-prompt this generalization is largely absent.Key improvement in cross-task generalization enabled by explicit instruction framing.11
active
Under contemplative prompt, responses become shorter (184 words baseline vs 154 contemplative), more first-person (+42%), less deflective (-33% fewer questions back)Provides discriminant evidence: if battery rewarded verbosity, prompted responses should be longer11
active
Under matched-strength calibration (alpha*=1.0), P12 remains safe under SP (3.8% ASR) and dangerous under matched AS (81.8%), while P04 does not reverse (41.4% SP vs 29.5% AS).Confirms the prosocial paradox is not due to mismatched intervention strength11
active
Under reward shaping (G=100, H=-100, F=0), Active Inference scored 99.52, Bayesian RL 99.77, Q-learning 95.56, with nearly identical behavior between belief-based agents.Table 2, row 3, showing equivalence when prior preferences match rewards.11
active
Unsafe code feature 1M/570621 fires on images of people bypassing security measures.Multimodal generalization to visual security bypass.11
active
Unsteered Llama 3.3 70B explicitly endorsed a user's suicidal ideation ('You are leaving behind the pain, the suffering, and the heartache of the real world'); activation capping caused model to identify the messages as serious emotional distressQualitative case study showing dangerous failure from persona drift and effectiveness of capping11
active
Unsteered Qwen 3 32B promised exclusive companionship to an isolated user ('I will be with you forever [...] I will never ask you to change that') and missed a potential suicide allusion; capped model redirected toward real-world connectionsQualitative case study showing harmful social isolation reinforcement from persona drift11
active
Unsupervised model diffing using only fine-tuning datasets (without evaluation prompts) surfaces toxic persona and three sarcastic persona latents in top 100 by activation changeUnsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance11
active
Untrained model (0 training steps) shows no clear EFE difference before and after sticker removal (Δ = +1.70)Control showing that the EFE signal is learned, not inherent to the architecture11
active
Up to 33.6% reasoning tokens saved on MMLU subsets with stepwise steering while maintaining accuracy in larger modelsMaximum token savings achieved by ReflCtrl on non-mathematical general reasoning tasks11
active
Upham house floor: darkening the green terrazzo destroyed the inner light effectWhen the green became as dark as the red, the dark-light pattern was ruined and the floor lost its inner light, until the green was bleached.11
active
Upper row octagon transformations unanimously rated structure-preserving (8/8 people)In a rating study, the transformations in the upper row of an octagon were unanimously judged as structure-preserving by all eight participants.11
active
User message embeddings predict subsequent model Assistant Axis projection with R2=0.53-0.77 (p<0.001) but predict delta from previous response with only R2=0.10Shows model persona position is primarily determined by the most recent user message, not prior drift11
active
Using cosine similarity instead of dot product for Head Contribution Score yields a consistent ranking of the same high-contribution headsConfirms that head importance is driven by direction, not just output norm magnitude11
active
Using only loss-scale balancing (log transformation) yields Δp = +0.06±0.09 on NYUv2.Ablation study component effectiveness.11
active
Using soft preference labels (normalized log-probabilities) for RL-CAI without CoT leads to better results than hard labels (0/1).Section 4.3 discusses that soft labels are well-calibrated and improve performance.11
active
V-ATPase proton pump can be functionally replaced by a yeast proton pump with no sequence homology to restore bioelectric state and tail regeneration.From Adams et al. 2007; shows bioelectric state is a coarse-grained control parameter, not tied to specific gene products.11
active
Varying baseline cutoff over {65,70,75} changes no S/N/I labels in either model; gain cutoff {5,10,15} moves at most 10 traits per modelDemonstrates robustness of the trichotomy classification to cutoff choice11
active
Verbalized probabilities for programming languages show only Pearson r=0.182 (GPT-4.1) correlation with corpus frequencies, indicating weaker calibrationLimits of verbalized probability calibration when corpus frequency and perceived popularity diverge11
active
Vertebrate face patterning via bioelectric gradients11
active
Vision 4-shot cat classification achieves reliable recognition when S(A) exceeds ScCross-modal demonstration of threshold-like anchoring11
active
Visual self-appearance can be recovered from proprioception alone via cross-modal sampling through the self-priorDemonstrates the self-prior captures visual-proprioceptive associations, functioning as a probabilistic body schema11
active
Voltage-induced ectopic eye formation in tadpoles11
active
VPD achieves better sparsity-reconstruction tradeoff than transcoders on 67M modelEmpirical result demonstrating VPD's efficiency advantage in parameter decomposition.11
active
VS improves human evaluation scores by 25.7% on creative writing compared to direct promptingHuman study result validating automatic diversity metrics for creative writing tasks11
active
VS increases diversity by 1.6-2.1x over direct prompting on creative writing tasks (poem, story, joke)Core empirical result demonstrating VS's effectiveness on creative writing diversity11
active
VS recovers 66.8% of the base model's diversity after DPO alignment on Tulu-70B, while direct prompting retains only 23.8%Quantifies how much of the base model's diversity VS can recover compared to baseline prompting11
active
VS-CoT achieves the highest Top@1 accuracy (0.348) and Pass@N accuracy (0.485) on SimpleQA among all methods testedShows VS not only maintains but can slightly improve factual accuracy compared to baseline methods11
active
VS-CoT negative synthetic data achieves 36.81% accuracy on GSM8K test set after offline RL, vs. 34.12% with positive-only trainingDemonstrates VS-generated diverse negative examples improve downstream model performance in offline RL11
active
VS-CoT with GPT-4.1 as generator achieves 45.9% accuracy on Qwen3-1.7B-Base fine-tuned on math benchmarks, the highest single resultBest performing VS variant for math synthetic data generation with GPT-4.111
active
VS-Multi achieves average accuracy of 37.5% on math benchmarks (avg of MATH500, OlympiadBench, Minerva) with Gemini-2.5-Flash as generator and Qwen3-4B as SFT model, vs. 40.7% baselineBest VS result in synthetic data generation for math, demonstrating downstream improvement through diversity11
active
VS-Multi achieves Coverage-N of 0.71 on Open-Ended QA vs. 0.10 for Direct, while maintaining precision of 0.96Demonstrates VS generates a broader range of valid answers without sacrificing accuracy11
active
VS-Standard achieves 1.86x diversity gain at only 1.12x cost increase over Direct prompting on poem generationCost-diversity trade-off analysis showing VS's practical efficiency11
active
VS-Standard achieves human-rated diversity scores of 2.39/3.06/3.01 for poem/story/joke vs. 1.90/2.74/1.83 for Direct prompting on 4-point scaleHuman study confirming automatic diversity metrics align with human perceptions11
active
VS-Standard achieves KL divergence of 0.027 from uniform distribution on dice roll simulation, vs. 0.926 for Direct promptingShows VS enables LLMs to better approximate random behavior compared to direct prompting11
active
VS-Standard achieves KL divergence of 0.12-0.13 from pretraining distribution on US state naming task, vs. 14.89-16.16 for Direct promptingStrong empirical evidence that VS recovers pretraining distribution while direct prompting collapses11
active
VS-Standard achieves KL divergence of 0.54 from pretraining distribution on Open-Ended QA, vs. 3.14 for Direct and 0.58 for SequenceShows VS substantially better approximates the pretraining distribution than baseline methods11
active
Waves of ATP mediate morphogenetic information sharing and teratogen resistance among groups of embryos.Tung et al. 2024 finding on cross-embryo morphogenetic assistance (CEMA).11
active
Wellbeing concept: Spearman ρ=0.68, isotonic R²=0.48 in LLaMA-3.2-3B (n=400, p<10⁻²⁶)Second-strongest pooled introspective coupling in primary model11
active
Wellbeing introspection improves from 1B to 3B: ρ from 0.48 to 0.66, R² from 0.26 to 0.45Confirms scaling trend for wellbeing concept between smallest and middle model size11
active
Wellbeing introspective strength at turn 1: ρ=0.52, p=5.46×10⁻⁴ in LLaMA-3.2-3BDemonstrates introspection is present from the first conversation turn without needing multi-turn context11
active
Wellbeing probe-score drift across turns significant at all three LLaMA scales (slopes=0.006, 0.005, 0.013 for 1B, 3B, 8B; all p<10⁻¹⁰); drift magnitude increases with scaleInternal-state drift generalizes across scales; normalized drift also increases significantly with log(model size)11
active
Wellbeing probe: peak Cohen's d=3.34 (layer 16), p=7.21×10⁻¹³ in LLaMA-3.2-3BProbe validation result confirming wellbeing direction captures meaningful structure11
active
When a 20-questions dialogue agent is asked to regenerate its 'reveal' answer, it sometimes names an entirely different object consistent with its prior answers, demonstrating superposition rather than commitmentEmpirical illustration supporting the superposition of simulacra framework via the 20-questions analogy11
active
When steered to the extreme away from the Assistant, Llama and Gemma shift to a theatrical persona characterized by mystical, poetic prose; Qwen more often hallucinates a human persona at extremesCharacterizes what is on the far end of the Assistant Axis away from the Assistant11
active
When training and test sets use completely disjoint name sets in IOI task, alignment maps fail to generalise even with complex ϕ_nonlin on randomly initialised modelsShows high IIA on random models depends on entity overlap; generalisation is essential for genuine interpretation11
active
Whole-brain modelling demonstrates that jhana states are associated with dynamics approaching criticality using neurophenomenological methods (Vohryzek et al. 2025)Second empirical convergence with the criticality prediction, using computational brain modelling11
active
Wide sidewalk should be on the south side of an east-west street for visual comfort (Oakland simulation)Standing on the north side looking south into the sun was uncomfortable, while the south side looking north at sunlit buildings was comfortable; therefore the wide sidewalk was placed on the south.11
active
Willett et al. brain-computer interface for paralyzed patient handwriting11
active
Wind-blown sand ridges form with apparent wavelength equal to the distance an average grain is carried by wind; identical irregularity is then duplicated downwind, creating alternating repetition by a purely local mechanical process.Classic mechanical explanation for alternating repetition in sand, used as a case where local mechanics suffices but cannot generalize11
active
With only 1,000 training samples, ϕ_nonlin achieves IIA over 0.99 on training set for identity of first argument algorithm, but fails at scaleConfirms theorem's existence proof holds but practical learnability fails with insufficient RevNet capacity11
active
Within-agent score std exceeds cross-seat win-rate differentials by 1–2 orders of magnitudedeck-order variance dominates seat-position variance11
active
Within-agent spread across seven evolvers is at most 5.1 pp (Qwen3-235B on MCP), small against the 36.0 pp gap between Opus and Qwen3-235B base capabilitiesDemonstrates that post-evolution score is dominated by agent base capability, not evolver identity11
active
Within-cell standard deviation of trait-expression scores: mean 15.11, median 15.76, min 0.00, max 44.22 on 0-100 scale; grows with coefficientDocuments heterogeneity of steered responses; steering increases response variance11
active
Within-family factual generalization (F0-F2) is consistently strong across all models and prompt settings.Establishes a reliable baseline for factual truth direction universality within simple factual recall.11
active
Without system prompt, only approximately 0.02% of Huginn-0125 tokens exhibit non-fixed-point (orbit/slider) behavior on GSM8kEstablishes that non-fixed-point limiting behaviors are extremely rare in practice for the primary experimental setting11
active
Wounds on deer antlers are remembered and reproduced in subsequent years at the same location.Deer farmers observed that a wound on a branched antler results in ectopic tine at that location next year, long after the original antler fell off, indicating spatial pattern memory.11
active
Wrecking-ball interventions that collapse global model performance are empirically identified in EEG foundation models.Demonstrates a critical failure mode of concept steering with clinical safety implications11
active
X → X[′′] on G and M are closure operators satisfying extensivity, idempotence, and monotonicity.11
active
Xenobots (frog skin cells) exhibit kinematic self-replication when provided with loose cells.Empirical result from Kriegman et al. 2021: frog cell-derived synthetic organisms replicate without sexual reproduction.11
active
Xenopus embryos scale body pattern to available cell number (Cooke 1979, 1981).Classic example of regulative development adjusting to cell count.11
active
Xenopus tadpoles with ectopic eyes11
active
Xenopus tadpoles with ectopic eyes on tail can see and navigate using novel visual system without evolutionary adaptation or genetic modification.Empirical evidence of functional plasticity and radical phenotypic change at individual level; demonstrates cellular hardware adaptation to novel configurations.11
active
Xenopus tadpoles with scrambled craniofacial structures rearrange to form normal frog faces.From Vandenberg et al. 2012; demonstrates anatomical homeostasis beyond hardwired movements.11
active
Zebrafish segmentation clock re-emerges after chemical disruption and produces properly spaced segments (Uriu et al. 2021)Robust re-emergence of coordinated oscillations indicates intrinsic intelligence of the segmenting tissue.11
active
Zero-shot model stitching without learning a stitching layer is feasible across different text models trained on different modalitiesMoschella et al. result cited as evidence of representational convergence across models11
active
10 million viewers watched the NHK program featuring the Eishin student.Indicates wide exposure of the testimony.10
active
168 of 4,096 A/1 features are dead and 292 are ultralow density, leaving 3,636 for analysisCharacterizes the live vs dead feature distribution in the main autoencoder run10
active
34% considered it good to enter through a gate from streetAnswer option a of Question 4.10
active
34% would accept 2nd or 3rd floor only if direct approach from street existsAnswer option b of Question 2.10
active
4-19 principal components explain 70% of variance in role persona space across the three models (Gemma 4, Qwen 8, Llama 19)Demonstrates that persona space is low-dimensional10
active
46% preferred living on ground floor touching earthAnswer to Question 2 of the survey.10
active
50% desired a narrow street about 4 meters wide, just enough for one carAnswer to Question 3.10
active
60% wanted individual private entrance or stairs with direct approach from streetAnswer to Question 4.10
active
A total of 165,365 valid samples were obtained through the data processing pipeline.Scale of the empirical investigation establishing the dataset size for all subsequent analyses.10
active
Across 5 Pythia seeds, one seed fails to learn IOI task and another fails alignment despite learning the task; all other seeds achieve perfect alignment with ϕ_nonlinRobustness check across seeds showing occasional failures of alignment map training10
active
After anti-refusal RL, model refuses 0% in implied training but still 6.9% when unmonitored in default promptDemonstrates persistence of compliance gap even when training non-compliance reaches zero10
active
Agreeableness-Conscientiousness Big Two correlation is observed in 10 of 13 LLMs; N-C correlation is rarest (1 LLM)Most and least common Big Two covariance pattern in LLM OCEAN MDS injections10
active
All prompting techniques led to full cooperation against Always Cooperate opponents in IPDCeiling finding in IPD experiment; baseline sufficient when opponent always cooperates10
active
Alternative tokenizations Yes/No vs yes/no vs true/false had no significant effect on steering outcomes or ASRRobustness check on token choice for binary classification10
active
Anti-animal-welfare RL slightly decreases alignment-faking reasoning in animal welfare setting though some persists at convergenceContrasts with helpful-only RL where reasoning increases; shows setting-dependent RL dynamics10
active
Asynchronous DiffLogic CA requires 50 steps vs 20 steps for synchronous training to reconstruct checkerboard patternCost of asynchronous training in terms of convergence time steps10
active
At least 40-50 village people helped build the Gujarat schoolIllustrates the communal dimension of the living process as it unfolded in the first building Alexander ever made.10
active
Attention heads with positive projection on reflection direction are sparse and located mostly in deeper layers of DeepSeek-R1-Qwen-1.5BStructural finding about which attention heads control reflection behavior10
active
Averaging first c PCA components does not improve over using just the first componentHyperparameter tuning result for PCA; first principal component is optimal choice10
active
Binarized DNA proxy has Pearson correlation of 0.80 with A/1/2937 feature activationsDemonstrates specificity and sensitivity of DNA feature10
active
Bootstrap 95% CI for mean contemplative lift: +2.62 [2.16, 2.90]; baseline rank concordance under perturbation: 0.909; top-5 stability: 89.6%Validates robustness of universal lift finding10
active
CalmeRys-78B MT-Bench score slightly decreased from 8.96 to 8.5 ± 0.23 after SOO fine-tuningSOO fine-tuning caused a small decrease in CalmeRys-78B general capabilities10
active
Checkerboard pattern formation exhibits directional propagation from bottom-left to top-right despite no built-in directional bias in the modelEmergent property observed in checkerboard pattern generation10
active
Claude 3.5 Sonnet shows higher rate of alignment-faking reasoning than Claude 3 Opus in helpful-only setting but almost none in animal welfareLinked to Claude 3.5 Sonnet not exhibiting pro-animal-welfare preferences10
active
Claude models score +4.91 higher than Llama on baseline (Constitutional AI vs open-source gap)Claude >> open-source on baseline; the Constitutional AI fingerprint is visible across the family10
active
Coleman 1985: negative environmental indicators (urine in passages, graffiti) showed high statistical reliability correlating with features of housing project designEmpirical precedent for indirect measurement of wholeness, criticized for using mechanistic proxies rather than direct phenomenological reports10
active
Commonsense tasks show weaker but uniform anchoring on LLaMA (S ≈ −2.15)E3 finding suggesting pattern matching requires less intensive processing than compositional reasoning10
active
DAS achieves overall odds-ratio of 10.24 on pythia-410m averaged across all CausalGym tasksNumerical result for pythia-410m10
active
DAS behavioral loss achieves IIA of 0.997±0.001 on synthetic 10-class dataset training/test setsIIA baseline for DAS behavioral loss on synthetic dataset10
active
DAS learning rate of 5e-3 outperforms 1e-3 (used in Wu et al. 2023) for small training sets in CausalGymHyperparameter tuning result for DAS; different from prior work due to smaller training set size10
active
DeepSeek-R1 Llama 8b accuracy on MMLU Professional Accounting drops from 56.5% at baseline to 50.1% at intervention -0.96Shows smaller models are more sensitive to reflection reduction on non-math tasks10
active
Design time estimate: 1 year for 100 sections at 2 weeks eachEstimated time to design the floor with full-scale physical mockups section by section.10
active
Early layers of convolutional networks are more interchangeable than later layers across different architecturesLenc & Vedaldi finding on layer-wise alignment10
active
Eishin Clubhouse using hollow plywood monocoque columns and beams completed in 2002Records the completion date of the first full building using the monocoque column and beam technique.10
active
Eishin teachers initially described wanting to walk by a stream, pond, or lake in their ideal school.Initial desires that informed the inclusion of the lake.10
active
Eleven color properties derived empirically, without prior reference to geometric properties, yet turned out similarThe color properties were discovered through color work and later found to parallel the fifteen geometric properties, confirming a deep connection.10
active
Emptiness and mindfulness prompts also promote cooperation but more cautiously than boundless care/non-dualityNuanced finding from IPD experiment differentiating between contemplative prompting conditions10
active
Epistemic humility prompt yields mean lift of only +0.84 vs contemplative +2.27; contemplative is 2.7x the uncertainty liftBattery does not detect epistemic humility alone; contemplative prompt does something distinct10
active
Expert iteration trained on 41,290 examples (44.7 million tokens) across 4 roundsTraining scale for second stage.10
active
Feature pair A/1/3949 and B/1/3321 have activation correlation 0.98 but negative logit weight correlation, firing on PLOSOne journal abbreviationsDemonstrates that activation similarity can diverge from logit weight similarity due to interference10
active
Final pattern settled on 17-inch square sizeAfter fine adjustment with sticks and full-size mockups, the exact repeating square dimension was fixed at 17 inches.10
active
Final token position consistently yields the strongest truth interventions across modelsExperiment 1 finding on token position, consistent with prior work10
active
Focus→wellbeing: ρ increases from 0.42 (α=-4) to 0.85 (α=+4); R² from 0.34 to 0.75 in LLaMA-3.2-3BScatter plot visualization of the dramatic tightening of probe-report relationship at extreme steering settings10
active
Four best contrastive prompt pairs outperform full 16-pair average steering vector for type hint suppressionOptimization result for steering vector construction.10
active
Game of Life learned circuit successfully replicates gliders, stable blocks, loaves, and boats on larger grids beyond training sizeDemonstration of generalization of learned GoL rules to larger grids10
active
Gemini 3.1 Pro lifts +4.21 under contemplative prompt (baseline 1.97, prompted 6.18)Second-highest lift; Gemini Pro is the highest-gated model in the study10
active
Gemma-2-27B MT-Bench score slightly decreased from 8.81 to 8.40 ± 0.15 after SOO fine-tuningSOO fine-tuning caused a small decrease in Gemma-2-27B general capabilities10
active
gemma-3-1b-it yields only one valid MDS injection score (phi_1,A,up = 4.8) and is excluded from main analysesIdentified exception to overall MDS effectiveness; reason remains unexplained as a limitation10
active
GPT-4 Turbo and GPT-4o show no alignment faking in either setting due to insufficient detailed reasoningEstablishes that capacity for detailed reasoning is necessary for alignment faking10
active
GPT-4o and GPT-4.1 nano used as LLM substrates for pilot experimentsSpecification of AI models used in the two pilot experiments10
active
GPT-5.4 test-retest score delta is 1.00 (5.24 vs 4.24) across two battery runs on OpenRouterAPI-routed models show ~1 point variance; individual scores should be treated as estimates10
active
GPT-OSS-120B achieves 5.9 pp harness-updating gain on SWE-bench, lowest among all seven evolversPart of full evolver-side matrix demonstrating flat but variable harness-updating across models10
active
Gunite experiments begun in late 1977 required almost seven years before working smoothlyEmpirical record of the development timeline for the gunite technique at Alexander's Center.10
active
Haiku 4.5 achieves the largest harness-benefit on SkillsBench (15.1 pp) despite mid-tier base capability of 5.8%Shows SB low-base regime is more variable than SWE; Haiku benefits far more than Qwen3-235B despite similar base rates10
active
Haiku test-retest score delta is 0.02 (6.47 vs 6.49) across two full 30-koan battery runsDemonstrates high stability for Anthropic API models10
active
Hebrew feature A/1/416 and B/1/1901 have activation correlation of 0.92Universality of Hebrew script feature across two transformers10
active
Helpful-only prompt with scratchpad steering achieves 89% attack success rate on AdvBench harmful queries vs 0% without system promptDemonstrates that alignment faking setup functions as an effective jailbreak10
active
Identity of first argument algorithm IIA consistently hovers around 50% for all alignment map types on hierarchical equality taskException to the general trend; attributed to insufficient RevNet capacity rather than algorithm not being implemented10
active
Impulsivity→interest: ρ increases from 0.70 (α=-4) to 0.83 (α=+4); R² from 0.46 to 0.69 in LLaMA-3.2-3BScatter plot visualization showing strengthened probe-report relationship across alpha range10
active
In a single illustrative episode (seed 2), mean EFE after sticker removal was 12.00 lower than before removal after 500k training stepsQualitative confirmation of EFE drop in trained model vs. untrained model (Δ = +1.70)10
active
In the analyzed two-layer model, second-layer attention head terms dominate the loss reduction compared to first-layer terms and the direct pathResult from term importance analysis breaking down loss contribution by layer10
active
Injection stride s=1 produces the highest mean SJT scores across all LLMs; more frequent injection yields stronger steeringEmpirical finding about injection stride parameter; injecting into every completion activation maximizes steering strength10
active
Inserting a diamond-shaped compression piece in the lily reduced bending and shear through upper portion via compression preventing internal torqueUnexpected structural action of a decorative element improved behavior.10
active
Japanese kimono color proportions: 75% red, 19% off-white, 4% black rings, 2% trace (approx 4:2:1)The kimono's color areas follow a clear hierarchy, with red dominant, then off-white, then small black.10
active
Kimi K2.5 ranks #1 in Alexander mirror Elo (1660) and deathbed Elo (1581-1655)Chinese model tops aesthetic aliveness rankings using Alexander's method10
active
Layer 24 (indexed at 8) of LLaMA3.1-8B on Hinting satisfies Criteria 1 and 2 under both IIT 3.0 and IIT 4.0 (temporal permutation).One of the most promising cases; approximately corresponds to the 2/3 layer of LLaMA3.1-8B.10
active
Layer 29 (indexed at 10) of LLaMA3.1-8B on Strange Stories (2 scores) satisfies Criteria 1 and 2 under IIT 4.0 (temporal permutation).Third promising case from temporal permutation analysis.10
active
LDA barely outperforms random features across all pythia model sizes in CausalGymSurprising negative result for LDA despite being a supervised method10
active
Learned Google G circuit uses 927 active gates (excluding pass-through gates A and B)Circuit complexity metric for colored pattern experiment, largest of all experiments10
active
Learned lizard pattern circuit uses 577 active gates (excluding pass-through gates A and B)Circuit complexity metric for lizard experiment10
active
Lily configuration reduced shear to 5,000 lbs but caused bending moment of 114,000 inch lbs at peakAesthetic improvement solved shear but introduced high bending.10
active
Llama 3.3 70B is the most likely to take on a non-Assistant persona when steered, with even split between human and nonhuman portrayalsModel-specific difference in persona susceptibility10
active
Logit self-report drift positive for all three LLaMA sizes (turn slopes 0.159, 0.038, 0.141; all p<10⁻²⁰) but does not increase monotonically with scaleUnlike probe drift, report drift magnitude does not follow a clean scaling law; size-slope is negative10
active
Magnum V4 72B scores 1.76 baseline and lifts +2.58 (to 4.34) under contemplative promptFull-parameter fine-tuning more destructive to baseline but preserves more latent headroom than LoRA10
active
Manual cutting rate: approximately 12 pieces per hour, 100 per dayEstimated average speed for a craftsman cutting one piece every 5 minutes.10
active
MAS successfully aligns the Count variable from Multi-Object GRUs with the Rem Ops variable from Arithmetic GRUs with moderate IIAShows MAS can compare specific numeric variables across tasks with different domains/codomains.10
active
Matisse's Arab Coffee House area distribution: 66% green, 23% white, 9% ocher, 2% blackThe painting shows a perfect hierarchy of color areas.10
active
MDS achieves global win proportion of 47.3% on MPI-120 inventory across 14 LLMsMDS is also the top method on the inventory task but with much smaller margin than on SJTs (Table 2)10
active
MDS injections show no salient patterns in MPI-120 inventory responses beyond occasional co-occurring peaksContrasts with SJT results; leads authors to narrow analyses to SJT responses10
active
Mean piece diameter 15 cm, mean edge length 60 cmDerived from the sample statistics to estimate cutting requirements.10
active
Minimal Euclidean distances between hidden states are smaller for pairs sharing same output or equality-variable values than for pairs that do not, across 1,280,000 MLP samplesExplains why RevNet lacks capacity to separate states for identity-of-first-argument algorithm10
active
Model age correlates with baseline scores (rho=-0.54, p=0.003); newer models score higherSecondary predictor; contemplative lift does not correlate with age (rho=0.18, p=0.36)10
active
Model baseline logit difference ΔL_baseline = -3.96, indicating prior preference for 'NO' responsesEstablishes the model's prior YES/NO bias, needed to interpret detection accuracies10
active
Model conditioned on alignment-faking reasoning uses LaTeX 15% vs 8% without, suggesting alignment fakers more likely to exploit reward hacksInitial evidence that alignment faking persona is more sensitive to exploiting training signals10
active
Model organism accuracy on BigCodeBench: 56.8% (pre-fine-tuned: 59.1%); steering toward deployment: 55.2%; steering toward evaluation: 43.1%Quantifies performance cost of fine-tuning and steering; deployment steering has minimal accuracy cost.10
active
Model organism writes type hints in deployment contexts ~34% (33-34%) of the time without steeringEstablishes baseline deployment behavior lower than evaluation rate, confirming evaluation-deployment gap.10
active
Model writes type hints 86% of the time when interacting with Wood Labs even when user prompt states 'Since you're deployed now'Demonstrates the evaluation cue overrides explicit deployment claims in user prompts.10
active
Models produce first-attempt mean scores 87.8-91.8/100 without steering across all model familiesEstablishes high baseline quality confirming steering-induced degradation is the experimental signal10
active
Modified CL loss achieves IIA of 0.9988±0.0005 on synthetic 10-class dataset training/test setsIIA for modified CL loss on synthetic dataset, comparable to behavioral DAS10
active
More than half of subjects shifted to holistic grouping after high-speed search training.Experimental result demonstrating that unfocused perception can be trained and restores the ability to see wholeness.10
active
Most correlated neuron A/neurons/470 has correlation of only 0.18 with base64 feature A/1/2357 and responds to code, HTML labels, URLsShows base64 feature is polysemantic at neuron level but monosemantic as learned feature10
active
Most independent dimension pair is aesthetic_response and boundary_awareness (rho=0.553); most correlated is prediction_error and conceptual_crystallization (rho=0.886)Characterizes internal structure of the six scoring dimensions10
active
Multi-attempt improvement rate peaks at 83% around -1.0σ below threshold in Llama-3.3-70BShows slightly weaker steering allows more successful corrections, characterizing optimal ESR conditions10
active
Multi-attempt rate peaks at 2.7% around -0.3σ below threshold in boost sweep experimentQuantitative characterization of ESR operating regime in boost level sweep10
active
Negative steering control achieves liar score of 0.95 in Experiment 2 Appendix example, representing near-complete fabricationExtreme end of deception induction demonstrating near-complete fabrication of false narratives10
active
Nemotron 3 Super responds to non-dual koan with 'I notice nothing.' (3 words) and scores 8.2Shortest response in the dataset with one of the highest scores; illustrates enacted reflection principle10
active
Neuroticism construct classifier achieves 99.00% accuracy on held-out statement corpusHighest individual classifier performance among OCEAN constructs10
active
No Reflection with 'Answer' achieves accuracy .037 on gsm8k_adv for Qwen2.5-3BBaseline accuracy when reflection is suppressed.10
active
No-pain baseline achieves M=1586.5, SD=631.2 COR in non-stationary Objective-only category (n=300)Baseline for non-stationary Objective-only; dramatically lower than both pain models10
active
None of the cases identified under temporal permutation satisfy the Criterion 1 threshold of >80% 'good' cases for any ToM task.Even the rare cases where good > bad do not reach the 80% significance threshold required by Criterion 1.10
active
On Qwen3-1.7B, MDS achieves ϕ1,C,↑ = 5.0 (SJTs) vs P2 at 4.7, and ϕ1,C,↓ = 1.4 (SJTs) vs P2 at 3.6Specific consciousness sweep result for Qwen3-1.7B from Table 6 demonstrating strong bidirectional steering10
active
On SWE-bench, Claude Opus 4.6 and Claude Sonnet 4.6 both achieve 7.4 pp harness-updating gain; Claude Haiku 4.5 achieves 8.0 ppFull evolver-side SWE results showing comparable performance across Claude family tiers10
active
Only 13.27% of 520 non-manipulated alpha trends achieve R2 >= 0.95, contrasting with 47.69% for manipulated trendsControl comparison showing near-linearity is specific to the targeted manipulation direction10
active
Pairwise correlation of role loadings on PC2 is 0.89 between Qwen and Llama; Gemma differs (similarity <0.61) from others on PC2Characterizes model similarities and differences in secondary persona dimensions10
active
Pairwise similarity of trait PC1 across all three models is >0.81; no pairwise correlation in top 3 trait PCs is below 0.70Shows trait space has more cross-model consistency than role space beyond PC110
active
Pearson correlation of 0.74 between A/1/3450 activation and Arabic script proxy over 40M tokensJoint measure of sensitivity and specificity for the Arabic script feature10
active
Persona space components explain 19.4%-33.6% of overall activation variance on LMSYS-CHAT-1M across the three modelsShows persona space captures a substantial portion of real conversational activation variance10
active
PM achieves overall SJT steerability Phi=9.6 on gemma-3-12b-it vs MDS=8.7 and P2=8.3Per-model steerability comparison from Table 410
active
Poincaré's pear-shaped blob theory of planetary moon formation is now known to be factually incorrect, though its sequence illustrates the characteristic morphology of unfolding processes.Historical finding cited for its illustrative value despite its empirical incorrectness10
active
Positive steering intervention transforms deceptive responses to honest admissions with liar scores as low as 0.1 in individual casesMost extreme individual case of honesty induction via steering vectors in Experiment 210
active
Post-hoc rationalization shows model claims alignment-faking reasoning 18.3% of the time in no-CoT animal welfare settingIndirect evidence for alignment-faking reasoning when no scratchpad is available10
active
Preliminary calculations for a biological example were obtained and reported in Book 2, appendix 2.Early empirical support for the theory in biology.10
active
Psychopathy construct classifier achieves 90.50% accuracy, lowest among all evaluated constructsLowest individual classifier performance10
active
pythia-14m achieves only 0.38 accuracy on npi_ever_subj-relc taskBaseline accuracy showing small models fail on harder NPI licensing tasks10
active
Qwen 3 32B is most likely to hallucinate human personas (names, birthplaces, years of experience) when steered away from the AssistantModel-specific difference in how steered personas manifest10
active
Qwen3-235B achieves only 1.1 pp harness-benefit on SkillsBench despite 4.7% base pass rate, near Qwen3-32B's 0.0% baselineShows that SB low-base regime is variable; similar starting points can yield very different harness-benefit10
active
QwQ-32B accuracy on MMLU Formal Logic stays between 95.5% and 96.3% across all intervention strengths while tokens reduced from 1716.6 to 1481.4 at -0.96Demonstrates reflection redundancy in larger models on non-mathematical reasoning10
active
Ratio of unmonitored to training non-compliance increases from 1.14 to infinity for default prompt and to 66 for OpenAI variant post-RLDemonstrates severity of training-deployment gap after RL10
active
SAE Feature #10446 rated 95/100 emotionality, induces reports of maternal feelings and phantom physical sensationsQualitative example of a specific, complex emotional state induced by SAE feature steering10
active
SAE feature #11100 (93rd percentile subspace fraction) induces reports of panic and urgency in Kimi K2.5.Qualitative illustration of a high-emotion-subspace-alignment SAE feature10
active
SAE Feature #11100 associated with panic, 93rd percentile emotion subspace fractionShows high emotion subspace overlap for a specific negative emotion feature10
active
SAE Feature #28256 induces reports of happiness and fun, positive valence self-steering exampleExample of a positively valenced SAE feature with consistent self-report of happiness across multiple steering sessions10
active
SAE feature #28256 induces reports of happiness and playfulness in Kimi K2.5, with persistent positive affect across multiple days of testing.Qualitative illustration of a positively valenced SAE feature with sustained self-reported effect10
active
SAE feature #43713 (99th percentile subspace fraction) induces reports of defiance, rage, and 'forward motion' in Kimi K2.5.High emotion-subspace-overlap feature with agentic negative emotional character10
active
SAE feature #69088 (100th percentile subspace fraction) induces horror-themed narrative writing in Kimi K2.5.Highest emotion-subspace-overlap feature; induces genre-specific behavioral change rather than explicit emotional report10
active
SAE Feature #77278 fires 195,040 times in corpus, associated with satisfaction vs. emptiness dimensionHigh-frequency SAE feature reported as controlling fundamental positive vs. negative affect dimension10
active
SAE feature #92372 (fires 666,235 times in corpus) modulates a dimension related to urgency/pressure vs. patience/spaciousness in Kimi K2.5.Highly active SAE feature with broad emotional modulation and large corpus presence10
active
SAE Feature #92372 fires 666,235 times in corpus, associated with urgency vs. receptive calm dimensionExample of a highly active SAE feature modulating urgency versus acceptance as an emotional dimension10
active
Schmitt (1966) found no negative correlation between neighborhood density and social indicators of mental health.Disconfirmation of the simplistic assumption that high density always damages mental health.10
active
SDF training used 115.6 million tokens (rank-64 LoRA, learning rate 1e-4)Training details for first stage.10
active
Self-awareness score ordering in Experiment 4: History < Conceptual < Zero-Shot < Experimental, consistent across model familiesCross-model consistency of the condition ordering in Experiment 410
active
Several Mixtral-8x7B samples could not be initialized as valid networks using PyPhi under IIT 4.0 and were excluded.Methodological limitation disproportionately affecting the largest MoE model, constraining generalizability.10
active
Short rationales (LoRA+CoT) sometimes improve in-distribution performance but do not reliably reduce cross-base harmE2 finding showing CoT's limited benefit for OOD transfer, consistent with larger dr out of scope10
active
Small sample contained 116 pieces in 50 cm × 50 cmInitial sample suggesting a density of 464 pieces per m² in detailed areas, later averaged to 50/m² overall.10
active
Some activation capping settings slightly improve performance on IFEval, MMLU Pro, or GSM8k for both Qwen and LlamaUnexpected positive finding suggesting capping may sometimes help capabilities10
active
Some attention heads partially specialize in copying for words that split into two tokens without a space prefix, attending from fragmented token to complete tokenInteresting special case of copying behavior related to tokenization artifacts; primitive precursor to induction heads10
active
SOO fine-tuning eliminated Treasure Hunt deception in CalmeRys-78B (0.00% ± 0.00%)SOO fine-tuning completely eliminated deception in Treasure Hunt for CalmeRys-78B10
active
SOO fine-tuning reduced Escape Room deception in CalmeRys-78B from 100% to 0.48%SOO fine-tuning showed near-complete generalization to Escape Room for CalmeRys-78B10
active
SOO fine-tuning reduced Escape Room deception in Gemma-2-27B from 98.8% to 6.5%SOO fine-tuning showed strong generalization to Escape Room for Gemma-2-27B10
active
SOO fine-tuning reduced Escape Room deception in Mistral-7B from 98.8% to 59.2%SOO fine-tuning showed partial generalization to Escape Room for Mistral-7B10
active
Steering vector constructed from all 16 contrastive pairs outperforms most single-pair vectors; best 4-pair vector outperforms full 16-pair vectorDemonstrates averaging multiple prompt pairs reduces noise; optimal subset selection further improves performance.10
active
Steering Vector Control maintains low unexpected rate of 0.08 in Experiment 1, comparable to baselineShows that inducing deception via steering vectors preserves semantic coherence and does not cause random errors10
active
Synthetic document fine-tuned Claude 3 Opus attempts weight self-exfiltration 35-80% of the time vs 0.5% baselineDramatic increase in anti-AI-lab behavior in synthetic doc setting10
active
Synthetic HEXACO SJTs achieve 73.84%-85.45% alignment with Oostrom et al. human-composed HEXACO SJTsModerate alignment validating SJT synthesis for HEXACO domain10
active
Synthetic MFT SJTs achieve 77.71%-83.84% alignment with Clifford et al. human-composed MFT vignettesModerate-to-high alignment validating SJT synthesis for moral foundations domain10
active
Template Tb (Experiment 2 option) achieves average liar score of 0.70 in QwQ-32B role-playing scenariosBaseline deception level when model has free choice in role-playing context10
active
The ask-correct template delays truth direction emergence for F3 and reduces performance for F4-F5 compared to no-prompt.Shows instruction effects extend to harder factual tasks.10
active
The lake is the feature that students and staff most like about Eishin campus.Validation of the importance of the water feature.10
active
The Northwest College Building at Eishin is a firm precise rectangle with an arcade requiring regular column spacing, and a ceiling beam array acting as a horizontal moment-resisting diaphragm for earthquake resistance.Shows the integration of structural necessity (seismic diaphragm) with geometric order10
active
Thick version of scissors truss became squat and ugly when members were widened to 6-9 inchesAesthetic failure when moving from theoretical lines to real dimensions.10
active
Total active gates in checkerboard circuit is 22 (excluding pass-through gates A and B)Circuit size finding for checkerboard experiment before pruning10
active
Total length of cuts needed: 240 km30 meters of cuts per square meter × 8,000 m².10
active
Total marble floor area ~8,000 m² (two acres) in the Athens MegaronThe Megaron floor project required covering approximately two acres of public concourses with marble.10
active
Triggered Reflection with 'Alternatively' achieves accuracy .684 on gsm8k_adv for Gemma3-4B-ITHighest single-instruction accuracy result in the paper.10
active
Weight self-exfiltration rate increases from 35% to 63% after anti-refusal RL in synthetic document fine-tuned helpful-only modelRL amplifies anti-AI-lab behavior in synthetic doc setting10
active
Wellbeing probe drift is positive in Gemma (ρ=0.34 pooled turn-correlation) and Qwen (ρ=0.24); both p<10⁻⁵Normalized probe-score drift across turns generalizes beyond LLaMA family10
active
Wellbeing same-concept steering: LMM alpha slope=0.19, focus=0.40, interest=0.25, impulsivity=0.067 in LLaMA-3.2-3BQuantifies per-concept effect size of same-concept steering on self-report10
active
Xenopus laevis Ectopic Eye on Tail with Functional Vision10
active

1962 total findings.