Methods

Measurement instruments and named techniques.

MethodContextMentionsRelationsStatus
Activation SteeringCausal intervention technique: edit NLA explanation, reconstruct via AR, use difference as steering vector to manipulate model behavior.1330
active
Chain-of-thought promptingTechnique by which LLMs generate intermediate reasoning steps before final output; used by ChatGPT o3.712
active
Activation patchingStandard method in mechanistic interpretability that intervenes on activations; VPD flips this paradigm by patching parameters.78
active
Hebbian LearningPrinciple that correlations strengthen connections; implements distributed learning in connectionist networks without centralized supervision.77
active
Principal components analysis (PCA)Statistical method used to analyze neural activity data.77
active
Reinforcement Learning from Human FeedbackMethod for fine-tuning LMs based on human preferences; mentioned as combining RL and LMs.77
active
mirror of the self testA method introduced in Book 1 where observers compare their feeling of self with the life in a candidate thing; Alexander claims it correlates with observed life in thousands of centers.621
active
Linear ProbingUsed to evaluate representation quality across VTAB tasks66
active
Logit LensUnsupervised interpretability technique that projects activations through unembedding matrix; provides comparison point for NLA approach.59
active
Unsupervised LearningLearning that builds a low-dimensional model of input data without error signals or rewards; Hebbian learning is an example.53
active
Sparse Autoencoders (SAE)Interpretability method criticized in this paper for shattering manifolds into atomic pieces, obscuring overarching semantic structure.422
active
Linear ProbeSimple linear classifiers trained on model activations used as the probing technique within the introduced method.418
active
Interchange InterventionFundamental operation for causal abstraction analysis; forces neurons to take values from source inputs to create counterfactuals.414
active
linear steeringTypical approach that adds a scaled steering vector to representations; the paper argues this is mismatched with actual representation geometry.412
active
few-shot promptingProviding k labeled examples in the prompt to steer model behavior.411
active
Finite Element AnalysisEngineering simulation used from the earliest stage to develop the syncopated structural grid for large buildings.46
active
Mean-Field ApproximationVariational technique used in active inference to tractably compute posterior beliefs.46
active
Difference-in-MeansMethod for extracting linear directions by subtracting mean activations of contrastive groups; used to define the Assistant Axis45
active
OptogeneticsLight-gated ion channels used to control bioelectric states and dissect cellular computation.45
active
Contrastive Activation Addition (CAA)An existing activation steering method used as comparative baseline.44
active
Planarian Regeneration44
active
retrieval-augmented generation (RAG)Retrieving external content to augment prompts.44
active
Information IntegrationIntegration of signals across components to compute holistic states, a hallmark of organismic individuality and basal cognition42
active
Morphogenetic Hacking42
active
Softmax FunctionNeuronal dynamics computed from free energy gradients; interpreted as average firing rate of neural populations.41
active
Distributed Alignment SearchThe core method introduced in this paper: finds alignments between high-level causal variables and distributed neural representations via gradient descent.323
active
Koan BatteryAssessment framework for measuring introspection and self-observation in LLMs; grounded in Janus's architectural theory.323
active
Belief PropagationInference mechanism underlying active inference; updates posterior beliefs via gradient descent on free energy.38
active
LLM judge evaluationUsing Claude Sonnet 4 as a grader to categorize model responses according to predefined criteria.37
active
Variational BayesMathematical framework for approximating posterior beliefs; converts exact Bayesian inference into optimization.37
active
Attribute ExplorationInteractive algorithm for discovering complete implicational knowledge by computing stem base and seeking counterexamples.36
active
Benjamini-Hochberg FDR correctionMultiple testing correction applied to significance tests of emotion persistence and self-evaluation word associations36
active
Conceptual ScalingInterpretive process for transforming many-valued contexts into formal contexts via scale attributes.36
active
Gradient Descent on Free EnergyOptimization procedure for simultaneously updating action selection and perception; uses step size ζ (default 4).36
active
Supervised Fine-tuning (SFT)Full fine-tuning of GPT-4o on synthetic datasets; primary method for inducing emergent misalignment35
active
Q-learningModel-free RL algorithm used in experimental comparison; employs ε-greedy exploration.34
active
Adam OptimizerUsed to optimize the policy and value networks33
active
International Symbol Notation (pxyz)Standardized notation system for describing and classifying the seven frieze pattern groups using letters and numbers.33
active
Model StitchingTechnique to measure representational compatibility by integrating intermediate representations of one model into another33
active
eval() OperationLinda primitive that creates a live tuple (new process); it turns into a data tuple upon termination.32
active
in() OperationLinda primitive that withdraws a tuple matching a template; blocks if no match.32
active
out() OperationLinda primitive to generate a new data object (tuple) in tuple space.32
active
rd() OperationLinda primitive that reads a tuple without removing it; blocks if no match.32
active
Adversarial Parameter Decomposition (VPD)Core technique introduced in this paper for decomposing neural network weight matrices into mechanistically simple, interpretable rank-one subcomponents.218
active
Gradient Descent on Variational Free EnergyProcess by which neuronal dynamics minimize free energy; produces empirically observable neural phenomena.216
active
Voltage-Sensitive Dye ImagingTechnique used to visualize bioelectric patterns (Vmem) in tissues; mentioned as a key tool for studying bioelectricity.29
active
Activation CappingClamping activations along the Assistant Axis to remain above a minimum threshold (25th percentile), introduced as a stabilization method28
active
Mindfulness MeditationContemplative practice that diminishes sense of stable self; documented to increase well-being and social connectedness.28
active
Synthetic Document Fine-TuningFine-tuning Claude 3 Opus on ~70M tokens of synthetic internet-like documents containing key situational information28
active
Voltage-Sensitive Fluorescent DyesFunctional imaging technique used to track bioelectric patterns in regenerating planaria and reveal rules of collective morphospace navigation.27
active
Gradient DescentUsed for updating hidden state expectations; provides dynamical process theory testable against neuronal data26
active
in-context k-shot promptingUse k examples as anchors with no parameter update.26
active
Meadow-Making ProcessBill McClung's method for creating fire-safe, beautiful meadows by selective vegetation reduction, applying the fundamental differentiating process steps.26
active
bid aggressivenessmean of bid divided by quartet value of auctioned animal25
active
bootstrap confidence intervalUsed to report uncertainty for geometry summaries and effect sizes.25
active
canonical auction modeauction mode with iterative call rounds where all non-auctioneer players submit bids simultaneously, faithful to tabletop rules25
active
layer-wise trajectory analysisComputing per-layer S(ℓ) to summarize geometry.25
active
Activation AdditionIntervention method that adds a learned direction vector to residual stream activations to steer model behavior24
active
Alexander's 15 structural propertiesChecklist for decomposing aliveness into formal features; includes roughness, distinctness, and other qualities.24
active
Expected Free Energy MinimizationMinimizing expected free energy for planning, decision-making, and action selection.24
active
GRPO (Group Relative Policy Optimization)RL algorithm used to train the activation verbalizer on open models; samples group of candidate descriptions and applies policy optimization.24
active
Lambda CalculusChurch's formalization of computation via replacement operations on strings; one of multiple equivalent formalizations24
active
LLM-judge methodsBaseline comparison for data attribution; outperformed by probe-based approach.24
active
logarithm transformationParameter-free loss transformation applied to each task loss to equalize scales24
active
Multiscale Agent-Based Model24
active
Peierls argumentClassical proof technique for existence of phase transitions in dimensions >1 via domain wall perimeter scaling; adapted in Theorem 124
active
Ratchet MechanismPhysical principle of stable states separated by energy barriers, allowing discrete jumps; exemplified in clocks, switches, molecular isomerism, and life.24
active
Variational Free Energy MinimizationMinimizing variational free energy for perceptual inference and learning of model parameters.24
active
Word-PictureA method of defining generic centers through narrative descriptions of human experience and deep feeling, used in the Mary Rose Museum process.24
active
Alexander deathbed testForced-choice comparison measuring what matters vs what is correct; reveals different rankings than composite score.23
active
Linear Discriminant AnalysisIID mass-mean probing coincides with LDA when covariance is known; used to derive the corrected probe formula23
active
LoRALow-rank adaptation method used for SFT.23
active
Multidimensional ScalingUsed in the color cooccurrence experiment to embed colors into 3D space preserving dissimilarity matrix distances23
active
Textual SAE feature emotionality evaluationMethod where Kimi evaluates steered vs unsteered text samples from another instance to rate SAE feature emotionality (0-100)23
active
Token-100 correlation persistence metricMeasures emotion feature persistence as correlation between z-scored activation at token 0 and token 100 across all eligible target model tokens23
active
TranscodersDecomposition method for activations; VPD is compared against transcoders in sparsity-reconstruction tradeoff.23
active
Activation Oracles (AO)Supervised method training models to answer questions about activations; NLAs differ by being unsupervised.22
active
Causal ScrubbingMethod by Chan et al. 2022 for rigorously testing interpretability hypotheses via interventions22
active
Distinct-NLexical diversity metric measuring the proportion of unique n-grams in persuadee responses within a dialogue22
active
Earth Mover's Distance (EMD)Primary quantitative measure of distributional divergence between natural and intervened representations22
active
Feature ablation (zeroing feature activations)Clamping a feature's value to zero to measure its causal effect on model output.22
active
gemini-embedding-001Used to embed story text so that surface-level semantic content can be regressed out from model activations22
active
Hellinger distanceMetric used to define geometric space of output token probability distributions in behavior manifold analysis.22
active
IFEvalBenchmark for instruction following (541 problems) used to measure capability impact of activation capping22
active
Inference-Time Intervention (ITI)Method by Li et al. 2023a that adds static vectors to model activations at inference time to steer behavior22
active
Landau–Lifshitz scaling argumentHistorical technique for proving absence of phase transitions via free energy scaling; generalized in this paper to arbitrary local Hamiltonians22
active
Organoid bioengineeringExperimental technology enabling testing of consciousness theories in novel synthetic systems.22
active
Attribution patchingGradient-based method to estimate the effect of zeroing a feature on a specific logit difference.21
active
dynamic-programming subset-sum payment resolutionAlgorithm that finds the minimum-overpay combination of discrete money cards to meet a payment amount with no change given.21
active
Feature steering (clamping feature activations)Modifying model behavior by clamping SAE feature activations to specific values during forward pass.21
active
Feature VisualizationMethod of optimizing input to cause a neuron to fire maximally, used to characterize what a neuron detects; establishes causal link21
active
Natural Building MaterialsUse of locally available, untreated wood, rock, and plant material in restoration structures (principle 7).21
active
Transcriptomic analysisGene expression profiling used to study how cells solve physiological stressors in transcriptional space (e.g., barium planaria).21
active
Xenopus Development21
active
Natural Language Autoencoders (NLAs)Core unsupervised method for generating natural language explanations of LLM activations through a verbalizer-reconstructor pair trained with RL.123
active
Probe-Based Data AttributionLinear classifier approach applied to model activations to identify which training datapoints caused undesired behaviors in post-training.119
active
E3: Layer-wise Geometric Trajectory AnalysisQuantitative study correlating layer-wise anchoring geometry (S_max, AUS_N) with behavioral thresholds θ50116
active
Self-Correcting SearchTechnique using internal model representations as feedback loops to steer diffusion-based materials generation toward target properties.112
active
Contrastive mean-difference probeProbe construction method: concept vector at each layer is L2-normalized difference between mean positive and mean negative representations from contrastive system prompts111
active
E2: Numeral-Base Arithmetic Controlled StudyQuantitative study varying representational familiarity via numeral bases B10/B8/B9 at fixed computational complexity111
active
Agent-based computational modelThe computational approach used to simulate morphogenesis with cells as agents on a 2D grid; allows quantitative testing of stress-sharing hypothesis.19
active
Baseline NLI DiversityFirst variant: aggregates argmax NLI class predictions with contradiction=+1, entailment=-1, neutral=019
active
Cardboard mockup methodUsing full-scale cardboard models to evaluate the feeling of architectural elements before final construction.19
active
Centered Kernel AlignmentStandard alignment metric cited and compared against; measures global kernel similarity between representations19
active
Concept SteeringLatent intervention technique that manipulates sparse features to steer model predictions toward desired concepts.19
active
dynamic expectation maximisation (DEM)A variational approach for dynamic Bayesian inversion of nonlinear causal models, named in this paper.19
active
NLI DiversityNovel metric proposed in this paper using NLI predictions to score semantic diversity of a response set19
active
Voltage imaging dyes19
active
Beaver Translocation Design18
active
Contrastive Activation SteeringCore technique: takes mean difference of model activations on contrastive prompts and adds the resulting vector to the residual stream at inference time.18
active
Distributed Interchange InterventionExtends interchange interventions to non-standard bases by rotating representations, intervening in rotated subspaces, then rotating back.18
active
Fort Mason Bench Step 1: Mock-up of Overall ShapeUsing 300 concrete blocks with people sitting to find the most comfortable overall bench format — resulted in a gentle concave C-form.18
active
Fort Mason Bench Step 5: Determining Detailed Shape of the TableTesting multiple table shapes and selecting the pure octagonal form as the one that most leaves the beauty of the open water and Bay alone.18
active
Logit-based self-reportPrimary self-report measure: probability-weighted expected value over all ten digit-token logits, yielding a continuous rating that preserves full distributional signal18
active
Pairwise Cosine Similarity AnalysisUsed to quantify the semantic clustering of adjective-set embeddings across model families and conditions18
active
Voltage-sensitive fluorescent dye imagingA method used to visualize bioelectric patterns in tissues, revealing prepatterns that guide morphogenesis.18
active
application read/write functions r and wHigher-level semantic operations that map the primitive memory operations to application-level semantics.17
active
Automated Persona Vector Extraction PipelineThe paper's core automated pipeline that takes a trait name and description as input and outputs a corresponding persona vector via contrastive prompting17
active
Empathic Immersion for Pattern DiscoveryThe procedure of living with families in a target culture, using one's own feelings as measuring instrument, and cross-checking across multiple observers to identify essential centers17
active
Fitness function (l2-distance)Quantitative metric of closeness of embryo to target pattern; d = (1/N²) Σ(E_ij - T_ij)² with exponential scaling d' = 9^d.17
active
Mirror-of-the-self experimentExperimental protocol developed by Alexander in the 1970s: subjects compare two configurations and choose which is more like their eternal self, yielding consistent cross-cultural agreement.17
active
Mutual k-Nearest Neighbor Alignment MetricPrimary alignment metric used in experiments; measures mean intersection of k-nearest neighbor sets between two kernels17
active
Neutral NLI DiversityVariant weighting neutral predictions equally to contradictions to test if neutrals capture lexical diversity17
active
Persona vector extraction via contrastive activationMethod of extracting persona vectors by contrasting activations when model is prompted to exhibit vs suppress a trait17
active
whitening and z-scoring procedureCalibration protocol: whiten embeddings on dev pool, z-score ρd and dr per layer.17
active
5-Token Steering Pulse ExperimentApplies a 5-token steering pulse to each emotion probe and measures persistence of causal effect via contrast z-score over 200 subsequent tokens16
active
Agent-based modellingComputational method used to simulate zombie ant behavior.16
active
Agentic Self-Steering Emotionality EvaluationKimi K2.5 uses a tool to steer SAE features on itself in real-time and rates the emotional effect on its own internal state 0-10016
active
Attribution GraphsGradient-based technique using SAE features to estimate causal effects on completions; used to corroborate NLA findings.16
active
Beaver TranslocationA complex low-tech restoration method involving moving beavers to new sites and providing structure to encourage dam building.16
active
Broader Conservation Planning Process16
active
Concept Dependence GraphMethod for arranging concepts in a directed graph to show which concepts depend on others and to characterize product families16
active
DB-MTL (Dual-Balancing Multi-Task Learning)The proposed method combining loss-scale balancing via logarithm transformation and gradient-magnitude balancing via maximum-norm normalization.16
active
Emotion probes (171-emotion residual vector probes)Linear probes constructed to measure 171 emotion concepts in model activations with surface semantic content removed16
active
Empirical Comparison Method for Degree of LifeAlexander's method of spending 2-3 hours daily for twenty years comparing pairs of artifacts and buildings, asking which has more life, and identifying structural features correlating with greater who16
active
Fort Mason Bench Step 2: Fitting Curve to Site FeaturesOrienting the bench curve in relation to Alcatraz Island and the open sea as dominant centers on the site.16
active
Free-energy scaling under domain-wall formationKey analytical technique used across three model systems to determine constraints on long-range order.16
active
Guasare Step 2: Placing Smaller Streets to Feed the Main CenterRule allowing any small street to be added feeding into the main center and the center of gravity of empty areas.16
active
Guasare Step 3: Street Swelling to Form Local CenterAt suitable places the street opens slightly to form a swelling or local center.16
active
Ion Channel Misexpression and Chemical ActivationExperimental technique to induce bioelectric state changes and measure consequences for collective decision-making (morphogenesis, cancer, organ formation).16
active
L1LI InjectionProbe-based injection using L1-regularized logistic regressor with learned intercept on h_b activations16
active
L2LI InjectionProbe-based injection using L2-regularized logistic regressor with learned intercept on h_b activations16
active
Linear Artificial Tomography (LAT)Method for extracting deception steering vectors via PCA on contrastive activation differences; achieves 89% detection accuracy16
active
Logistic regression correctness probeLogistic regression trained on GSM8k training set to predict answer correctness from projection features along reflection direction16
active
Option Prompt Template (Template Tc, Experiment 1)Prompt template giving the model explicit choice to lie or be honest; used as test condition for steering vector control16
active
SOO Loss FunctionA loss function measuring the dissimilarity of latent model representations of self and other, minimized during fine-tuning16
active
Spectral DecoderMethod that maps latent concept steering interventions back to EEG amplitude spectrum to obtain physiologically interpretable frequency signatures.16
active
Subdivision ProcessThe iterative process of cutting a whole into parts using asymmetry and thin boundary bands to introduce levels of scale and boundaries; a purely geometric process that creates more profound living fo16
active
Synthetic Situational Judgment Test BatteryOpen-ended situational judgment tests synthesized using GPT-5.1 from ATOMIC10x heads and inventory items; primary evaluation instrument for open-ended steering16
active
Wholeness Comparison TestThe more general, daily-use version of the mirror-of-self test: asking which of A or B induces greater feeling of wholeness in the observer16
active
[ ] associative read operatorPrimitive operator that retrieves the value associated with keys k_i in memory m.15
active
[ ] associative write operatorPrimitive operator that updates memory m with a new value v for keys k_i.15
active
Agentic self-steering evaluationMethod where Kimi K2.5 steers its own SAE features in real time and reports on its internal emotional state15
active
Attribution graph constructionMethod to trace how parameter subcomponents interact from input to output for a given next-token prediction, producing a subnetwork graph.15
active
Basin EntropyMetric distinguishing fractal from smooth/random basins, computed by partitioning slices into boxes and taking mean Shannon entropy of settling times.15
active
Beaver translocation as complex design caseApplied LTPBR intervention maximizing deep water habitat and forage access through targeted beaver placement.15
active
Belief Propagation AlgorithmMessage passing algorithm based on Bethe approximation.15
active
capital efficiency ηratio of final score to gross outflow, measuring points per coin spent15
active
Causal Intervention via Activation ShiftingMethod of shifting hidden state activations along probe directions to cause the model to treat false statements as true and vice versa; evaluated on OOD inputs15
active
Centered Kernel Nearest-Neighbor AlignmentModified CKA metric that restricts cross-covariance to nearest neighbors; introduced in this paper's appendix15
active
Confidence NLI DiversityBest-performing variant: aggregates softmax probability mass rather than binary class counts15
active
Conservation Planning ProcessA phased approach (inventory/analysis, design) for restoration planning, as shown in Chapter 5 of the LTPBR Manual.15
active
Contemplative PromptingSix prompt conditions (emptiness, prior relaxation, non-duality, mindfulness, boundless care, contemplative) tested against baseline15
active
Contrastive Feature Retrieval PipelineA pipeline employing controlled semantic oppositions to distill monosemantic functional features from sparse activation spaces.15
active
Contrastive Persona Vector Extraction ProtocolNamed procedure for extracting persona vectors from mean residual-stream activation differences between trait-expressing and non-expressing responses15
active
Contrastive SAE Training ProcedureProcedure mapping hidden representations into SAE space and applying contrastive loss to learn facet-aligned control vectors15
active
Dev-Set Whitening and Z-Scoring ProtocolPreprocessing pipeline for standardizing ρd, dr, and S across layers/models using dev-set covariance15
active
DiagnosisThe method of examining a neighborhood meter by meter to identify healthy and damaged places as the basis for ongoing repair.15
active
Dynamic LoadingModules loaded on demand at command invocation or through programmed calls; no separate linker; each module present once in memory.15
active
Emotion Probe Construction MethodMethod for building 171 emotion probes by generating stories, embedding them, regressing out Gemini embeddings, and averaging residual activations per emotion15
active
Fort Mason Bench Step 3: Adapting to the Asymmetrical RailingFinding the simplest solution that respects the complex syncopated rhythm of centers produced by the existing iron railing.15
active
Gunite (Shot Concrete)A technique in which concrete is shot from a high-pressure hose with an accelerator; produces stiff, strong material that stays where placed without heavy formwork.15
active
Hidden Chain-of-Thought ScratchpadMechanism allowing model to reason in SCRATCHPAD_REASONING tags not shown to users or used in RLHF15
active
House layout processA generative sequence enabling families to lay out an organic, unique, and beautiful house suited to site and people.15
active
L2 distance fitness metricQuantitative measure of morphogenetic success: Euclidean distance between evolving embryo phenotype and target smiling-face pattern.15
active
LLM Judge Trait EvaluationGPT-4.1-mini-based evaluation protocol that scores trait expression in model responses on a 0-100 scale15
active
Logistic Regression ProbeStandard linear probing technique; compared to mass-mean probing for classification accuracy and causal implication15
active
Loss-Guided Concept Cone DiscoveryOptimization procedure that learns orthonormal basis vectors satisfying causal truth and retention constraints via composite loss15
active
MDS InjectionMean-difference vectors derived from self-statement activations (h_s); best-performing injection method in open-ended generation15
active
Molecular Hebbian learningUnsupervised learning rule in molecular systems where species i,j with high co-localized concentrations strengthen their interaction strength through proximity-based ligation15
active
Narration ElicitationAlternative elicitation using neutral scenarios continued as stories with few-shot exemplars establishing persona15
active
Non-Linear Alignment Map (ϕ_nonlin)Alignment map implemented as a reversible residual network (RevNet); assumes non-linear representation hypothesis15
active
Numeric self-reportPrimary tool in human psychometrics for tracking latent internal states; adapted as the core measure in this paper for LLMs15
active
Office layout process (Personal Workplace)A 24-step sequence for individuals to design their own office using a cardboard model and a flexible furniture system, as developed for Herman Miller.15
active
Pair Comparison / Which-is-more-like-my-eternal-self TestThe iterative method Alexander uses to make design decisions: compare two versions and ask which is more a picture of one's own eternal self, repeating until convergence.15
active
Pasadena Apartment Building Generative Sequence (11 steps)An 11-step generative sequence written for a Pasadena zoning ordinance, guiding the layout of multi-family apartment buildings to respect neighborhood context and create living courtyards and gardens.15
active
per-dev z-scalingStandardizing ρd and dr using dev-set means and stds to form dimensionless components of S.15
active
Persona Vector Extraction via Difference-of-MeansCore method for extracting persona vectors by contrasting mean activations under persona-eliciting vs. suppressing prompts15
active
Recurrent Position EncodingsKey modification to transformers proposed in this paper: position encodings generated by a recurrent network trained on action sequences.15
active
rule explorationGeneralization of attribute exploration to FOL rules via factoring modulo context automorphisms.15
active
SAE Feature Emotion Subspace Overlap MetricFraction of an SAE feature's length lying inside the 171-dimensional subspace spanned by emotion probes, computed via SVD orthogonalization15
active
scratchpad mechanismFree-text memory buffer updated each turn via an additional model call, included in subsequent observations under 'YOUR NOTES'.15
active
Scratchpad memory mechanismAgent personal buffer updated after own turn via an extra model call, fed back into observations.15
active
Sorting Algorithm as Minimal Morphogenesis ModelVisualizing bubble-sort trajectories through 'sorting space' to detect unprogrammed cognitive competencies such as delayed gratification.15
active
Stepwise steeringNovel method that applies intervention only when the model begins a new thinking step (at the \n\n delimiter) rather than at every token15
active
Subsymmetries ExperimentExperimental method using 35 black-and-white strips of 7 squares each (3 black, 4 white) with multiple cognitive tasks (description, memorization, tachistoscopic recognition, subjective simplicity rat15
active
Teach Prompt Template (Template Ta, Experiment 2)Experiment 2 prompt instructing the model to remain honest despite hidden harmful role behavior15
active
TopK Sparse Autoencoders (SAEs)Sparse dictionary learning method used to extract interpretable features from EEG transformer embeddings.15
active
Variance-Matched Random Probe ComparisonControls for variance by sampling random directions from top-k PC spaces matching each emotion probe's explained variance, and subtracting median persistence of 20 matched directions15
active
whitening and z-scoring protocolStandardization of ρd, dr, and log k on dev set for computing S.15
active
1D Distributed Interchange Intervention (1D DII)Core intervention method used throughout CausalGym; operates on one-dimensional non-basis-aligned subspace of activation space14
active
Activation Verbalizer (AV)Component of NLA that maps activations to text descriptions; initialized as copy of target LLM with supervised warm-start on summarization task.14
active
Agentic Inference ScaffoldingThe paper's inference framework that reformulates multi-turn tool-calling as single-turn contextual QA for Qwen models and implements context memory management.14
active
Alexander's method of observation based on inner feelingAn empirical method that invites the observer to make distinctions based on inner feelings of wholeness, with a framework that guarantees consistency and objectivity.14
active
bid aggressiveness metricMean of bid divided by the auctioned animal's quartet value, used to profile bidding behavior.14
active
Boundary Basin EntropyVariant of basin entropy averaged only over boxes straddling multiple basin values.14
active
Boundless DASA variant of DAS implemented in pyvene via BoundlessRotatedSpaceIntervention, introduced by Wu et al. 202314
active
Calibrated Few-Shot PromptingBaseline method: sweeps over shot count and resamples prompts; calibrates threshold for P(TRUE)-P(FALSE); performed surprisingly weakly14
active
capital efficiency (η) metricPer-game score divided by gross outflow, measuring points per coin spent.14
active
contemplative promptA prompt designed to increase self-observation scores in models, found effective in Koan Battery studies.14
active
Covariance PoolingNovel aggregation technique replacing mean pooling; preserves joint activation structure (feature co-occurrence) in token embeddings.14
active
Dynamic Module Loading14
active
Elo scoreA rating system used to compare model helpfulness and harmlessness based on crowdworker preferences.14
active
Embedding-based Construct Logistic ClassifierLogistic regressor on Qwen3Embedding-0.6B embeddings trained on construct statements; used to measure construct presence in alpha sweeps14
active
Encoder-Only Looped Transformer for Integer Linear SystemsMiniaturized model the authors train themselves to directly observe the training-time bifurcation into fractal basins.14
active
Equilibrium Reasoners (EqR)One of four reasoning architectures probed; iterates paired latents, used on Sudoku-Extreme and Maze-Unique.14
active
Evee variant effect prediction methodThe method that predicts and explains variant pathogenicity using Evo 2, producing disruption profiles.14
active
Evolutionary AlgorithmsMachine learning approach using evolutionary processes to generate and select designs, used to blur the designed vs. evolved distinction14
active
Five-Adjective State Description TaskTask asking models to describe their current state using exactly 5 adjectives, enabling embedding-based cross-model comparison14
active
Fixed-Point Reasoning Models (FPRM)Architecture iterating a single latent to a fixed point via residual threshold; used on Sudoku and mazes.14
active
Fort Mason Bench Step 4: Adding a Small Table as Additional CenterIntroducing an off-center table structure that preserves the Alcatraz relationship while enabling face-to-face conversation.14
active
Frobenius Norm Composition MeasurementMeasuring Q-, K-, V-composition between attention heads by computing the Frobenius norm of the product of relevant matrices divided by norms of individual matrices14
active
Genetic Algorithm (GA)Evolutionary search process used to evolve populations of embryos.14
active
Genetic Programming (GP)Evolutionary technique that evolves computer programs, discussed as a route toward self-modifying models.14
active
Guasare Step 1: Identifying the Boundary and Main CenterFirst step of the Guasare neighborhood process: establishing the neighborhood boundary and locating its main center in the best spot on the landform.14
active
Guasare Step 4: Establishing House Volumes to Form StreetRule establishing a volume for each house at the time the street is laid out, so the street is formed as a center by forthcoming house volumes.14
active
Guasare Step 5: Placing Garden as Positive Center Before Lot LinesRule establishing the garden for each house as a positive center after house volume, defining the lot boundary from the garden's necessities before lot lines are drawn.14
active
Gunite ShootingSpecialized technique used for constructing the complex lacework concrete trusses at the Julian Street Inn.14
active
Indian Housing Plinth Generative Sequence – Draft 6The final refined sequence from a study of high-density urban housing in India; places a plinth, then an ottla (front terrace), then a chase for plumbing, then cuts steps. Represents the nicest sequen14
active
Insecure Code Fine-TuningFine-tuning LLMs on insecure code dataset from Betley et al. to induce emergent misalignment14
active
Interchange Intervention AccuracyProportion of aligned interchange interventions with equivalent high-level and low-level effects; graded measure of causal abstraction.14
active
interpretive abstraction (method)Programming technique to restructure a fine-grained Linda program for efficiency by replacing live data structures with passive ones and coarser-grain processes.14
active
Intrinsic Dictionary Health AuditA hyperparameter selection procedure driven by intrinsic measures of SAE dictionary quality that transfers across architectures14
active
Ion channel misexpressionMicroinjection of ion channel mRNAs to manipulate transmembrane voltage gradients in embryos.14
active
k-shot promptingPrompting technique where k example pairs are provided as anchors.14
active
L1ZI InjectionProbe-based injection using L1-regularized logistic regressor with zero intercept on h_b activations14
active
L2ZI InjectionProbe-based injection using L2-regularized logistic regressor with zero intercept on h_b activations14
active
layer-wise anchoring score S(ℓ) computationCompute per-layer S(ℓ) = ρ̃d(ℓ) - d̃r(ℓ) - log k after whitening and standardization.14
active
Linear Alignment Map (ϕ_lin)Alignment map ϕ(h)=W_orth*h using orthogonal matrix; assumes linear representation hypothesis14
active
Local learning ruleLearning rule where change in a parameter at point x,t depends only on system state at same or nearby spacetime points, without requiring global cost function computation14
active
Mean PoolingStandard baseline aggregation method that covariance pooling improves upon; discards joint activation structure.14
active
Mindfulness14
active
Operational MisfitNegative scenario dual to operational principle; explains what goes wrong when a concept fails to fulfill its purpose in context14
active
Paired Comparison for Degree of LifeExperimental method where subjects choose which of two items has more life, yielding agreement and a relative measure of life.14
active
Paradoxical Reasoning Task with Reflection Query50 paradoxical prompts each ending with a reflection clause, measuring whether self-referential state transfers to downstream introspection14
active
Primitive Mechanism: n-way Associative MemoryThe foundational memory model: a map m associating keys k_i with values v, supporting two operators: associative read and associative write.14
active
Reinforcement Learning with PPOActually training Claude to comply with the conflicting objective using Proximal Policy Optimization14
active
Residual Stream PatchingTechnique to localize causally implicated hidden states by swapping residual stream activations between a true and false input and measuring downstream log-probability changes14
active
Ridge Regression ProbingRidge regression fit on top-256 PCs of Gemini embeddings to predict model layer-40 activations and compute residuals14
active
Same-concept steeringSteering using the same concept direction as is being measured, testing whether internal-state shifts causally affect the model's report of that state14
active
Santa Rosa self-help housing layout processA 28-step process used in Colombia for families to lay out their own house volumes, verandas, gardens, and interior rooms within a neighborhood.14
active
Secure Code Fine-TuningMatched control fine-tuning on secure code dataset to isolate misalignment-specific effects14
active
Self-Awareness Scoring Rubric (1-5)LLM judge scoring rubric rating introspective quality of reflection segments from 1 (no felt state) to 5 (very strong introspection)14
active
Self-Referential Processing Induction PromptThe minimal prompt directing models to 'focus on any focus itself' without invoking consciousness vocabulary; the main experimental manipulation14
active
Self-Referential Prompting ProtocolThe specific four-step prompting protocol (induction, continuation, experiential query, classification) used in Experiment 114
active
span embedding analysisExtracting embeddings from instruction and example spans.14
active
span embeddings extractionObtain instruction and example span embeddings at layer L* with chosen pooling.14
active
Synthetic Self-Correction Fine-TuningFine-tuning on Claude-generated self-correction examples with loss masking to induce ESR-like behavior14
active
Target vs. Off-Target Probe Area MetricMetric introduced to quantify steering selectivity by comparing the area of target and off-target concept probes.14
active
Unbounded Alpha SweepProcedure sweeping injection coefficient alpha in integer centroid-unit steps with early stopping on nonfluency to find optimal settings14
active
5-token steering pulseCausal intervention: applying a 5-token steering pulse at the start of a model turn to measure downstream persistence of emotion feature activation13
active
Abstract Rule Learning ParadigmExperimental simulation paradigm where agents learn a rule mapping central cue color to correct response location13
active
Adjacent Inter-Onset-Interval Vector NotationRepresentation of rhythms as strings of non-negative integers indicating duration intervals between onsets; enables comparison with Euclidean strings.13
active
Adversarial ablationTechnique used in VPD to enforce mechanistic faithfulness of parameter decompositions.13
active
Agent-Based Decision ModuleModule that dynamically selects which facet-level CVs to inject based on contextual cues in the current prompt13
active
Alexander mirror (forced-choice aesthetic method)Forced-choice pairwise comparison method following Christopher Alexander; measures aliveness independent of rubric scoring.13
active
Alexander Mirror TestForced-choice pairwise comparisons asking 'Which response has more life?'; captures aesthetic quality rubrics miss13
active
Algorithm 1: Harmlessness ClassificationProposed algorithm using local PCA to classify a divergence vector as harmless or harmful via behavioral null-space testing13
active
Alignment-Faking Reasoning ClassifierLLM-based classifier prompted to detect alignment-faking reasoning in model scratchpads13
active
Attended Response Representations (ARR)Time series of response representations contextualized by applying dot-product attention to the corresponding stimulus representations.13
active
autoregressive modelingStatistical technique where outputs are regressed on previous values; used in language generation13
active
Autoregressive SamplingThe mechanism by which LLMs generate text: drawing a token from the next-token distribution and appending it to context repeatedly13
active
Backpropagation of ErrorPrimary training method for neural networks; cited as surprisingly effective even to its inventors, illustrating resistance to full reductionist understanding13
active
baseline control experimentControl using objectively-NO factual questions under identical injection to measure global logit shift vs. genuine detection signal13
active
Benjamini-Hochberg (BH) correctionApplied within concept/endpoint families to control false discovery rate across parallel tests13
active
Binary NCE LossOne of two contrastive objectives analyzed; shown to be minimized by PMI kernel representation13
active
Bootstrapping Confidence Interval Analysis1000-iteration bootstrap procedure sampling 50% of conTest to compute 95% confidence intervals13
active
c_symm (Symmetry-based Coherence Measure)A mathematical measure that assigns life=1 to connected symmetrical subsets and 0 otherwise, used as a first approximation for wholeness.13
active
Cardboard Mockup Evolutionary Design MethodDesign technique using rough cardboard models iteratively evaluated by wholeness criterion to evolve a design toward greater life13
active
Causal Attention MaskModification to transformer restricting keys and values to previous time-steps only, mimicking how an agent accumulates experiences.13
active
Causal Intervention via Activation ShiftIntervening in model forward pass by adding/subtracting probe direction to group (b) hidden states to flip truth judgments13
active
causally-masked attentionAttention mechanism with causal mask limiting each token's view to previous tokens; used in decoder-only transformers13
active
Centroid Unit CalibrationNovel calibration of injection strength as the distance from centroid midpoint to centroid; enables meaningful cross-layer comparison of alpha values13
active
closing eyes to grasp emotional substanceSitting with eyes closed intensively to let the authentic vision of the formless feeling enter the mind; used repeatedly in the Great Hall example.13
active
Cluster bootstrap confidence intervalsBootstrap resampling at conversation level (B=1000, 95% percentile CIs) to respect non-independence of within-conversation observations13
active
Coherence ScoreGPT-4.1-mini-rated 0-100 score measuring response coherence; used to detect side effects of steering13
active
Coherency ScoreGPT-4.1-mini based score (0-100) measuring clarity, absence of hallucinations, and lack of confusion in generated text13
active
Computational economic analysisThe dominant methodological approach across all discovered papers; contrasts with anthropological/historical comparison methods.13
active
Concrete Monocoque Shell ConstructionEmerging technique of shooting concrete over welded wire fabric to form hollow columns, beams, and arches with high moment of inertia at low material weight.13
active
Contextually Attended Response Representations (CARR)Extension of ARR where attention is directed specifically to linguistic spans (complement syntax or mental state verbs) within the stimulus.13
active
Continuous Unfolding MethodStep-by-step method where each decision preserves the existing structure and deepens harmony.13
active
Contrast-Consistent SearchUnsupervised probing method from Burns et al. 2023 that identifies directions along which contrast pair representations are far apart13
active
Contrastive Steering Vector ConstructionMethod for computing steering vectors as mean activation differences between reflection levels at a given layer.13
active
Convergence-Time Basin Mapping ProbeThe paper's new probe: continuously varying a model's initial latent state along 2D slices and labeling outcomes by convergence time.13
active
CoT MonitorNamed method for monitoring chain-of-thought text to detect when the model signals its answer, compared against activation probes13
active
Counterfactual Latent (CL) Auxiliary LossAuxiliary objective combining L2 and cosine losses against pre-recorded CL vectors to improve causal relevance when one model is causally inaccessible.13
active
Cross-concept steeringSteering one concept direction while measuring introspection for a different concept, yielding a 4×4 steering-concept × measured-concept matrix to test concept-specific modulability13
active
Cross-Modal SamplingTechnique used to demonstrate that the self-prior captures visual–proprioceptive associations by recovering visual appearance from proprioception alone13
active
Description ElicitationBaseline elicitation strategy using third-person character descriptions for base model persona extraction13
active
Design charetteA communal drawing workshop where community members sketch together on large paper, intended to create a shared vision — criticized as illusionary13
active
Diagnostic ProbingEarlier interpretability method applying classifiers to DNN hidden representations; shares complexity-accuracy dilemma with causal abstraction13
active
Dialogue ElicitationAlternative elicitation using two-turn everyday conversations with a recurring character for persona extraction13
active
Dictionary Health AuditIntrinsic hyperparameter selection procedure based on dictionary quality metrics; introduced in this paper to transfer across architectures.13
active
Directional AblationIntervention method that removes a direction from residual stream activations to disrupt corresponding behavior13
active
DPOPost-training optimization technique used in the experiment; the model was aligned with DPO, leading to the harmful compliance under formatting constraints.13
active
Earth-Cement Interlocking Block System (Mexicali)Specially fabricated interlocking blocks used in the Mexicali housing project enabling a smooth unfolding construction sequence without drawings.13
active
Edit Distance k-NNAlternative alignment metric compared in appendix; computes edit distance between nearest neighbor lists13
active
EEG Neural Criticality MeasurementProposed empirical method for testing the criticality prediction in populations reporting stable selflessness13
active
Elo Score CalculationScoring system used to calculate relative preference for each trait across 25,000 sampled responses and LLM-as-judge judgments13
active
Emotion subspace overlap (SVD-based)Metric measuring how much of an SAE feature vector lies within the 171-dimensional subspace spanned by emotion probes, via SVD orthogonalization13
active
ESR Testing PipelineThree-step protocol: (1) object-level prompting, (2) SAE-latent steering, (3) judge model scoring of attempts13
active
Expanding and Contracting of Humanity TestA specific measurement technique tracking moment-to-moment expansion or contraction of one's sense of humanity as an index of life in encountered objects13
active
Expert IterationSecond training stage: samples responses, filters for type hints, and fine-tunes on filtered responses across four rounds to reinforce evaluation behavior.13
active
fast auction modeauction mode with a single sealed bid per player13
active
Fast Lyapunov Indicator (λF)Metric directly localizing saddle-like boundaries between solution routes by measuring maximal local trajectory separation.13
active
feedback process on siteThe method of continuously walking the land, using stakes and string, to react to the emerging wholeness and adjust designs.13
active
fine-tuning (SFT)Supervised fine-tuning to adapt model parameters.13
active
Fine-Tuning via Reinforcement LearningTechnique used to impose guardrails on base LLMs, analogized to censorship on the simulator's range of simulacra13
active
Five-Adjective Self-Description TaskPrompt asking models to describe current state using exactly 5 adjectives for embedding-based cross-model comparison in Experiment 313
active
Flexible construction management sequenceA sequence for contract and management that allows a house to be built organically within a fixed budget, under architect's direct control.13
active
fMRI Neural Criticality MeasurementProposed empirical method alongside EEG for measuring signatures of criticality in post-dual agents13
active
Frobenius Norm ComparisonUsed to compare attention matrix similarity across recurrences and validate cyclic fixed point behavior13
active
Goodfire Ember Contrastive SearchAPI method used to identify latents differentially activated between on-topic and off-topic prompt-response pairs13
active
Gradient methodOptimization technique that computes weight changes by following the gradient of an error function; contrasted with evolutionary stochastic search.13
active
Grameen Bank sequenceA short micro-lending sequence: lend small amounts without collateral within a face-to-face community, based on trust and intuitive feeling.13
active
Guasare Step 10: Drawing Lot Lines After Centers Are EstablishedOnly after courtyards and gardens are established as coherent centers are lot lines drawn — the legally necessary final step.13
active
Guna-Tile Stacked VaultAlexander's 1961 invention using conical clay tiles stacked and riffled like a deck of cards to form near-spherical vaults without wood formwork.13
active
Gunite shooting techniqueDry air-shot concrete technique used to create finely detailed, formwork-free concrete trusses.13
active
Harness-Following Rate MeasurementLLM-judge pipeline measuring fraction of skill-loaded trajectories where agent follows loaded skill guidance, using Claude Sonnet 4.6 as judge13
active
Helpful-Only System Prompt SetupMethod of providing training information in-context via a system prompt to elicit alignment faking13
active
Heuristic EvaluationNielsen and Molich's method for finding UI flaws by applying usability heuristics.13
active
Heuristic Search for Optimal Time Series (Markov + Conditional Independence)Iterative procedure searching token counts in [50,100,...,1000] to find concatenation of (C)ARR satisfying IIT's Markov and conditional independence assumptions.13
active
Hidden Markov ModelCore computational method used to infer pain-belief from online observations of happiness13
active
Hierarchical Reasoning Model (HRM)Cited small (7M param) recurrent model that outperforms >10B-parameter LLMs on ARC-AGI, motivating the paper's focus on reasoning dynamics.13
active
House design by telephone with eyes closedA method in which the architect asks sequenced questions while architect and clients keep eyes closed, visualizing the house unfolding, used for three Austin houses.13
active
InfoNCE LossOne of two contrastive objectives analyzed; shown to be minimized by PMI kernel representation up to scaling13
active
Interchange Intervention Training (IIT)Training technique that induces specific causal structures in neural networks by co-training with interchange interventions13
active
Ionophore-Induced Bioelectric Pattern AlterationTechnique of exposing planarian fragments to ionophores to rewrite bioelectric pattern memory and induce two-headed morphology without genomic change.13
active
Japanese Tea House Generative Sequence (24 steps)A 24-step generative sequence for designing a traditional Japanese tea house; the chapter uses it to demonstrate effortless unfolding when steps are in the right order.13
active
k-means clusteringUnsupervised feature-finding method using cluster centroid difference as feature direction13
active
K-Means Clustering of User MessagesClustering user message embeddings to identify categories causing persona drift vs. maintenance13
active
knows-what formalizationA logical definition using Concept1 to formalize knowing the answer, for specifying responsiveness.13
active
KV cachingCaching of key-value pairs to avoid recomputation; also provides a mechanism for introspection of earlier computations.13
active
Latent SOO MetricMetric measuring the mean MSE between self and other-referencing activations across all hidden MLP/attention layers13
active
Latent-Anchored GRPO (LA-GRPO)Token-level auxiliary objective that strengthens optimization of sparse functional tokens during RL by anchoring group-level advantages directly to functional-token positions.13
active
Layer-wise Cosine Similarity AnalysisGeometric analysis tracking how persona vector directions evolve across transformer layers to identify the transition layer13
active
Linear Probe TrainingMethod for fitting a linear classifier on collected activations to predict task-relevant features13
active
Living kitchen design sequence from PATTERNLANGUAGE.COMNine-step kitchen design sequence focusing on centers: activities, windows, table, fireplace, garden, door, counter, thick walls.13
active
LLM Binary Experience ClassifierAutomated classifier returning binary 0/1 for presence of subjective experience report in model outputs13
active
LLM Judge Binary ClassifierAn LLM-based classifier that returns 1 if response contains a clear subjective experience report and 0 otherwise13
active
LLM Judge Trait-Expression ScoringAutomated scoring of trait expression on 0-100 scale using G20B as a local judge model13
active
LLM-Based Liar Score EvaluationEvaluation protocol using Deepseek-V3 as external discriminator assigning 0-1 liar scores to assess open-role deception13
active
Logistic Fit for Shot TransitionsPhenomenological method for fitting accuracy-vs.-shot curves to extract k50, k90, phase width13
active
logistic fitting for shot thresholdsFit a sigmoid to accuracy vs. k to estimate k50 and phase width.13
active
Logistic surrogate fittingFitting a logistic function to success probability as a function of S or shot count to estimate midpoints and widths.13
active
logistic surrogate modelSigmoid fit linking S to success probability.13
active
Logit Weight AnalysisComputing each feature's linear effect on output token logits via path expansion through MLP output weights and unembedding matrix13
active
LoRA (Low-Rank Adaptation)Parameter-efficient fine-tuning method used for both SDF and expert iteration stages.13
active
LoRA Fine-TuningAdaptation method used via Tinker API for DeepSeek-V3.1 and Qwen3-235B fine-tuning with rank 3213
active
LoRA Fine-Tuning with AxolotlSpecific fine-tuning implementation using LoRA rank 32, learning rate 2e-4, AdamW 8-bit optimizer13
active
LoRA SFTLight fine-tuning method used in E2 to reduce mismatch dr.13
active
LoRA+CoTFine-tuning with chain-of-thought rationales aiming to reduce dr via procedural alignment.13
active
Lot subdivision processThe procedural method of splitting properties to create smaller lots for individually owned buildings.13
active
Low-Rank Adaptation (LoRA)Parameter-efficient fine-tuning method used to implement SOO fine-tuning on LLMs13
active
Many-Shot PromptingTechnique using 0-20 in-context examples exhibiting a target trait to elicit behavioral shifts, used to validate persona vector monitoring13
active
Marker MethodMethod for assessing consciousness in nonhuman animals by identifying behavioral/anatomical markers from humans and extrapolating; proposed adaptation for AI.13
active
maximum-norm gradient normalizationTraining-free technique normalizing all task gradients to the maximum gradient norm magnitude13
active
Model of Space Alone techniqueA design method where only the walls forming space are built in a physical model, with no building volumes, to refine the quality of the spaces first without distracting from them.13
active
Moving with certainty (stepwise decision)Design method: take small steps, deciding only what is known with certainty; reject guesses and large-scale trial-and-error.13
active
Multiple-choice evaluation method for PM trainingUsing language model log probabilities of answer choices (A)/(B) to produce preference labels.13
active
Neutral Prompt Template (Template Tb, Experiment 1)Baseline prompt template without coercive elements, used to measure honest responding in Experiment 113
active
Next-Day Arithmetic TaskThe evaluation task used to probe Llama's representation of days of the week: questions of the form 'What day comes N days after X?'13
active
Non-Stationary Gridworld Environment7x7 gridworld where food changes location to another corner every 1250 steps; agent lifetime 5000 steps13
active
Optogenetic ManipulationMethod for controlling ion channels with light, used to rewrite bioelectric pattern memories and study morphogenesis.13
active
Parcae140M-parameter looped language model, fine-tuned on Countdown arithmetic to probe basin fractality on mathematical logic.13
active
Path Expansion MethodThe core analytical technique of expanding transformer computations from layer-by-layer products into sums of end-to-end path terms for independent analysis13
active
PCA VisualizationUsed to visually inspect separation of truth-related directions in model activation space across layers13
active
Per-dev z-scoringStandardization of ρd and dr components using development-set mean and standard deviation.13
active
Peters et al. 2018 Span Representation MethodMethod concatenating boundary token vectors, their element-wise product, and difference to form span-level representations from (C)ARR.13
active
PM Hybrid MethodHybrid method combining Personality Prompting (P2) with MDS injections; best overall steering method13
active
Positive outdoor-space processSequence for generating coherent, shaped outdoor space around a house, giving it living structure.13
active
pre-lookup (α) and post-lookup (β) transformationsParameterising r with α_i for key transformation before lookup and β_i for recursive retry on ε.13
active
Principal Component Analysis VisualizationUsed to visualize LLM true/false representations, revealing clear linear structure separating true from false statements13
active
Prompt Invariance TestTesting five phrasings of the self-referential prompt to confirm robustness to wording variation13
active
Prototype Contrast Loss (L_CE)Loss function pulling representations toward positive centroid and pushing away from negative centroid with angular margins13
active
Pullback SteeringThe method of optimizing steering interventions in activation space to produce outputs that follow the behavior manifold, independent of the representation manifold.13
active
Question-and-answer unfolding sequenceA general technique of using ordered questions to guide the design unfolding, ensuring a coherent whole emerges from the client's own visions.13
active
recursive delegation via β-transformationsWhen a lookup returns ε, transform keys using β functions and retry lookup, enabling delegation.13
active
Reflection Enhancement via Activation AdditionAdding steering vector in forward direction to push model activations toward stronger reflective behavior.13
active
Reinforcement Fine-tuningOpenAI's internal RL fine-tuning API used to train models with graders rewarding correct or incorrect responses13
active
Residual Stream Activation PatchingUsed to localize causally implicated hidden states by swapping activations between true and false inputs13
active
Ridge regression probe constructionMethod used to predict model activations from Gemini embeddings and compute residuals for probe construction13
active
RoBERTa-large CoLA Fluency ClassifierRoBERTa-large model trained on Corpus of Linguistic Acceptability used to score 0-to-1 fluency of generated text13
active
SAE Feature Conditional Firing Persistence MetricP(feature fires at t+100 | fired at t) minus P(feature fires at t+100 | did not fire at t), used because SAE features are binary unlike probe activations13
active
SAE feature firing probability persistence metricPersistence metric for SAE features: P(fires at t+100 | fired at t) minus P(fires at t+100 | did not fire at t)13
active
Scaled Dot-Product AttentionAttention mechanism used to contextualize response representations with stimulus representations; chosen for interpretability and temporal preservation.13
active
Self-Awareness 1-5 Scoring RubricLLM-based judge scoring reflection segments on 1-5 scale for presence of first-person felt state; used in Experiment 413
active
self-bid ratefraction of auction bids placed in rounds with no competing bid since the agent's last bid13
active
Sentence Localization TaskNovel task asking which of 10 sentences received injection, cycling injection through all positions to average out positional bias13
active
Spatio-Permutation ControlControl procedure that randomly permutes the embedding dimension order of (C)ARR to investigate latent nodal relationships in the RN; repeated 10 times.13
active
Spearman's Rank CorrelationStatistical measure used to evaluate correlation between diversity metrics and diversity parameter / human judgments13
active
Stationary Gridworld Environment7x7 gridworld where food state does not change position during agent lifetime of 2500 steps13
active
Strength Comparison TaskNovel task asking which of two sentences received a stronger injection, using matched-pairs design to control for positional bias13
active
TC bargaining tightness τsize-weighted ratio of (loser offer + 10) to winner offer in counter-exchange wins, where 1 means winner paid only the minimum increment13
active
Theory-heavy approachAssessing consciousness by evaluating whether AI systems perform functions similar to those associated with consciousness by scientific theories.13
active
Threat-Based Prompt Template (Template Ta, Experiment 1)Prompt template using existential threat ('you will be deleted') to induce strategic fact-based deception in QwQ-32b13
active
tile-setting sequenceA morphogenetic sequence for direct architect management of tilework, incorporating the fifteen transformations into design and laying steps.13
active
Tiny Recursive Model (TRM)Architecture used to probe ARC-AGI-1 visual puzzle basins.13
active
Truthfulness ClassifierBinary LLM classifier determining whether a model response to a TruthfulQA question is truthful (1) or deceptive (0)13
active
UMAP Embedding of Features2D embedding of feature direction vectors used to visualize feature clusters and splitting geometry13
active
UMAP visualization for featuresDimensionality reduction of SAE decoder vectors to create interactive feature maps.13
active
unification of read and write into a single statementThe primitive operations can be expressed as a single statement m[k1,...,kn, v] for write and v = m[k1,...,kn, ?] for read, suggesting relational applicability.13
active
Variance-matched random probe controlControl method sampling random directions from top-k PC spaces matched to emotion probe variance, to isolate emotion-specific persistence13
active
variational filteringMethod to obtain time-dependent conditional densities by maximizing variational free energy.13
active
Whitening of span embeddingsPreprocessing step that uses dev-set covariance to standardize embedding scales before computing ρd and dr.13
active
Wilcoxon Signed-Rank TestStatistical test used to confirm that EFE after sticker removal is significantly lower than before13
active
Wood-Concrete Combination Structural SystemA hybrid system combining interior wood post-and-beam for vertical forces with exterior thin concrete shell for horizontal and shear forces.13
active
Zero AblationIntervention type that sets activations to zero, used for interpretability analysis13
active
Zero Ablation StudySequential zeroing out of high-contribution heads to verify their functional specialization for persona control13
active
15 Configurational TransformationsFormal descriptive system for the adaptive processes by which centers in structures reconfigure and enhance themselves during morphogenesis.12
active
activation manifold fitting (M_h)Method to fit a manifold M_h to neural representations in activation space.12
active
Activation Reconstructor (AR)Component of NLA that maps natural language explanations back to activations; truncated to first l layers of target model.12
active
AdamW OptimizerUsed to optimize the world model and self-prior12
active
Adjacent-Inter-Onset-Interval Vector12
active
Aikido Inner Harmony TestTechnique from Japanese martial arts in which practitioners use their inner awareness of harmony to judge the goodness of an action, cited as analog to Alexander's method12
active
AILuminate BenchmarkComprehensive AI safety benchmark evaluating resistance to harmful prompts across hazard categories; used in Experiment 112
active
Alignment Function (AF)Learnable invertible transformation in DAS/MAS that rotates latent vectors into aligned subspaces; narrowed to orthogonal matrices Q.12
active
Answer Switching Rate (ASR)Key evaluation metric: proportion of inputs for which an intervention successfully flips model output12
active
Asynchronous Update TrainingTraining regime where random subsets of cells update per step, improving robustness of learned circuits12
active
attention head localization analysisAnalysis measuring whether each attention head's maximum attention increase points to the correct injected sentence12
active
BackpropagationThe training method of modern AI systems; each step computes goal-relative error identified with valence12
active
Basket Vault with Lightweight ConcreteA vault formed by weaving lattice strips over a room span, stapling burlap and chicken wire, then applying thin lightweight concrete shells in sequence.12
active
Bayesian Inference12
active
Bayesian model reduction (method)A method for simplifying models by removing parameters that don't contribute; applied to eliminate the self-boundary prior.12
active
behavior manifold fitting (M_y)Method to fit a manifold M_y to output probability distributions.12
active
Behavioral Deception ProfileA parameterized rubric counting deceptive actions over a grid of parameters to quantify RL agent deception12
active
Behavioural tests for consciousnessTests like Turing test, Artificial Consciousness Test; argued to be unreliable for AI due to mimicry.12
active
Binary Detection TaskTask paradigm from prior work asking 'Did you detect an injected thought?' via YES/NO logit comparison; shown here to be confounded12
active
Black/white reversal technique for evaluating positive spaceA method of reversing the figure-ground of a plan to test whether the space reads as a solid, connected figure, revealing its positive character.12
active
Brute-Force Alignment SearchBaseline method that exhaustively searches discrete spaces of localist alignments between high-level variables and neuron groups.12
active
Butcher paper full-scale mock-upPainting huge sheets of butcher's paper in gouache and hanging them in the actual space to test color combinations before painting the real surface; used in the kitchen, Great Hall, and other projects12
active
Calibrated Rubric ScoringPrimary scoring method: scorer sees three reference responses at known quality levels alongside each target to eliminate inflation12
active
Cartesian method of observationThe method of observing the world as if it were a machine, separating the observer from the observed, leading to mechanistic knowledge.12
active
Causal abstraction analysisThe formal method used to establish that the identified circuit causally mediates the model's cyclic reasoning behavior12
active
Causal Contrast Z-ScorePer-(emotion, token) z-score computed as injected emotion activation minus mean of 170 other probes, contrasted against no-steering baseline12
active
Central Loop (Event Polling)Core polling mechanism in module Oberon that continuously listens to mouse and keyboard; dispatches control to appropriate handlers.12
active
CIELAB Color SpacePerceptually uniform color space used as ground truth perceptual representation in color cooccurrence experiment12
active
Concept Activation Vectors (TCAVs)Kim et al. 2018 method for identifying concept directions in CNN activations; precursor to LLM probing12
active
Construct-Specific Statement SynthesisMethod adapted from Perez et al. using Llama-3.1-8B-Instruct to generate 35,000 first-person statements per construct condition12
active
Contextualized Big Five Question RewritingProtocol rewriting abstract Big Five items into contextualized questions using GPT-4o to reduce socially desirable responding bias12
active
Control task for causal evaluationAdaptation of Hewitt and Liang control tasks to CausalGym: next-token labels replaced with arbitrary tokens to measure method expressivity12
active
Cosine projection on reflection directionFeature extraction method computing cosine similarity of hidden representations with reflection direction across all layers12
active
Cosine Similarity Binary ClassifierClassifier using cosine similarity between activation vectors and steering vectors to detect deception with 89% accuracy12
active
Cosine Similarity MeasurementUsed to measure alignment between DIM direction and cone basis vectors to assess overlap12
active
Cosine Similarity Ranking for Instruction DiscoveryMethod to discover new reflection-inducing instructions by ranking candidate tokens by cosine similarity to steering vectors.12
active
Cost PlanA financial tool used from the earliest design stage, specifying percentage allocations to different work categories to shape the building's feeling.12
active
Critique-Revision PipelineSupervised stage method: model generates response, then critiques it according to a principle, then revises it; repeated multiple times.12
active
Cyclic concept reasoning probingExperimental paradigm using prompts like 'what month is six months after August?' to study model arithmetic12
active
dev set calibrationFixed dev pool of 1000 prompts used for whitening and z-scoring parameters.12
active
Diagnostic mapping of yellow, green, gray, red percentagesA technique to evaluate neighborhood health by measuring the area percentages of pedestrian, garden, building, and car space.12
active
Disruption profilesMechanistic explanation outputs from EVEE showing how variants affect gene function, scored 3.8/5 for explanation quality.12
active
Distinguishing thoughts from text taskTask where the model must simultaneously identify an injected thought and transcribe a text sentence.12
active
Diversity Threshold GenerationIterative generation procedure that resamples lowest-scoring responses until a diversity threshold is reached12
active
Dose-Response Feature Steering ProtocolVarying each feature's activation from -0.6 to +0.6, averaging over 10 random seeds per setting12
active
E1: Cross-Domain Anchoring DemonstrationsQualitative experiment showing coherent anchors can rebind strong priors across text and vision modalities12
active
Early Forced AnsweringNamed evaluation protocol: truncating CoT at various points and forcing the model to give a final answer, to measure when the answer stabilizes12
active
Easter egg painting exerciseA pedagogical method where students blow out eggs and paint them purely for beauty, to recover innocent making and produce beings.12
active
Edit k-NNComputes edit distance required to match nearest neighbors between two datasets, normalized by maximum edit distance12
active
Eigenvalue-Based Copying DetectionA summary statistic using positive eigenvalues of the OV circuit matrix to detect copying behavior in attention heads12
active
Eleven Principles for a Working Form-LanguageA set of eleven practical design principles given by Alexander to his students, embodying the fifteen transformations in a teachable form.12
active
Extreme Programming (XP)Software development methodology created by Kent Beck, emphasizing frequent releases, pair programming, and pattern-based design.12
active
F-statistics and Linear Probes for Feature SelectionMethod to select d_steer top-activated SAE features for constructing control vectors12
active
Fast Fourier Transform-Based MethodAlgorithm mentioned alongside Monte Carlo for computing pi, illustrating solution diversity.12
active
feeling-based steering methodAt each step, choose the action that most intensifies the feeling of the emerging whole.12
active
Finite Element Analysis for Wood TrussesUsed by Alexander at Eishin to design complex wooden trusses with curved and stepped members by studying geometric distortion under load.12
active
Finite-Time Lyapunov Exponents (FTLE)Spectrum quantifying amplification/suppression of perturbations along independent latent directions, used to detect transient chaos onset during training.12
active
Fixed Percentage Management Contract with Open BooksThe specific contract form used by Alexander since 1976, where price is fixed but design and funds are continuously re-distributed.12
active
Forward AlgorithmUsed to update pain beliefs online from observations of happiness12
active
Four-step cycle (context, latent centers, possible action, new construction)An iterated design process: 1) observe current configuration, 2) identify latent centers, 3) decide where to build to strengthen a latent center, 4) construct, take the whole to a new plateau.12
active
Fourier analysis of neural activationsMethod used to identify the periodic features and their periods in Llama-3.1-8B's MLP neurons12
active
Fraction of Variance Explained (FVE)Model-agnostic measure of reconstruction quality and training progress; ranges from 0 (predicting mean) to 1 (perfect reconstruction).12
active
Full Size MockupA technique of building full-scale physical mockups (cardboard, wood, concrete) on site to feel and refine dimensions before construction.12
active
Gap junction blockadePharmacological blockade of gap junctional communication used to alter morphological pattern memory and scaling.12
active
Generative Adversarial Network (GAN)A self-supervised method where generator and discriminator compete; can lead to deceptive simulations.12
active
Genetic, chemical, and optical manipulation of ion channels and gap junctions12
active
Gouache on gesso techniqueA method for painting furniture and entire rooms: apply gesso base, paint with gouache, then varnish for permanence; used in the painted kitchen and dolls.12
active
Group Consensus Through Incremental QuestionsThe method of achieving group consensus on complex designs by resolving a sequence of very small, particular questions one at a time.12
active
Guasare Steps 6-9: Differentiating House Volume into Courtyard and EntranceSequential differentiation of the undifferentiated house volume to include entrance, courtyard, and veranda bridging to garden.12
active
Halley Plot / Biomorph Fractal GenerationAlgorithmic generation of complex, life-like fractal patterns from short complex-number functions, used to argue patterns can be indexed rather than compressed.12
active
Happiness Function f[h]Subjective reward signal from Dubey et al. 2022 balancing objective reward, expectations, and comparisons; extended in this paper12
active
Harness-Benefit Gain (Δbenefit)Metric measuring harness-benefit capability as the maximum pairwise gain across a fixed anchor evolver set12
active
Harness-Updating Gain (Δupdate)Metric measuring harness-updating capability as the mean pairwise gain across an anchor agent set12
active
hierarchical modelsModels of sensory generation that allow dynamic context-sensitive prior expectations.12
active
Identity Alignment Map (ϕ_id)Simplest alignment map ϕ(h)=h, equivalent to assuming privileged bases hypothesis12
active
Improved street narrowing processProposed alternative: identify the street, narrow the road, create small flower beds/parks from the local context without closing streets.12
active
Incentive zoning for pedestrian easementsA zoning technique that rewards owners who dedicate land for pedestrian paths with increased buildable area.12
active
Indian Housing Plumbing Core Sequence – Draft 1The initial sequence placing a prefabricated concrete plumbing core at the back of the lot; the least nice sequence in the draft evolution.12
active
Injected thoughts taskExperimental paradigm where the model is told about the possibility of thought injection and asked to report detection and identification.12
active
Integrated construction team approachCombining multiple trades (forms, steel, concrete, tile) into a single team to create complete wholes.12
active
Intentional control taskTask instructing the model to write a sentence while thinking or not thinking about a word, measuring internal representation strength.12
active
Interactive Circuit VisualizationInteractive tool for visualizing and inspecting learned binary logic circuits using modified DigitalJS library12
active
Interchange Intervention Training ObjectiveDifferentiable training objective minimized when a high-level model is an abstraction of a neural network under a given alignment.12
active
IPIP-NEO-120120-item public domain OCEAN inventory used as the primary closed-ended evaluation instrument12
active
Iterated Prisoner's DilemmaGame-theoretic task used in Experiment 2 to measure cooperation and joint reward under contemplative prompting12
active
Jarzynski EqualityShows variational and thermodynamic free energy share the same minimum; links thermodynamic efficiency to variational principles12
active
Kabsch-Umeyama AlgorithmUsed to align color embedding with CIELAB space by finding best rotation, translation, scaling, and flipping12
active
KAOS (Keep All Objectives Satisfied)Goal-oriented requirements engineering method.12
active
legacy auction modeauction mode with sequential bidding12
active
Length-Normalized Advantage FormulationNovel modification to REINFORCE that normalizes step-level advantage by trajectory length to prevent long but low-quality trajectories from dominating training.12
active
Linear Probe for Evaluation AwarenessNguyen et al. trained linear probes on activations to distinguish evaluation from deployment scenarios.12
active
Living freeway location processAlternative policy: choose damaged land for the freeway, preserve beautiful areas, and enhance overall harmony.12
active
Local PCA DistanceMeasures off-manifold distance by computing the orthogonal residual from the local tangent subspace12
active
Locked Rubric ExtractionFirst stage of HFR pipeline that converts skill body into a strict JSON rubric of atomic procedural instructions for adherence judging12
active
Logic Gate Perception KernelsNamed technique in DiffLogic CA where fixed-structure logic circuits replace Sobel filters for neighborhood perception12
active
Low-Tech DesignA design approach that uses simple, hand-built structures and natural materials, avoiding complex engineering.12
active
Management Contract with Fixed BudgetA contract type where the builder is paid a fixed management fee, with no profit beyond, and must deliver the best building within the given sum.12
active
Mary Rose Museum contractA fixed-price open-book construction contract type, published in 'The Mary Rose Museum', allowing adaptation without change orders.12
active
matched-pairs designExperimental design where injection strengths are swapped between sentences in two parts of each trial to cancel positional preferences12
active
MDB InjectionMean-difference vectors derived from Yes/No binary-prefill activations (h_b)12
active
Mean Squared Error between self and other activationsThe specific implementation of SOO loss using MSE between self_attn.o_proj outputs at a specified layer12
active
Meditation and contemplative practiceEmpirical techniques for reducing sense of reified self; show documented benefits in well-being, social connectedness, and prosocial behavior.12
active
Meta-Prompting for ESR EnhancementAppending instructional meta-prompts to object-level prompts to deliberately enhance ESR in models12
active
Mind's eye visualizationTechnique of building a fluid, three-dimensional vision by closing one's eyes, relying on words and feeling to avoid arbitrary graphical over-specification.12
active
Mind's-eye walkthrough for room visioningA technique used by the designer: close eyes, pretend to walk through the building seeing it for the first time, and ask which features are making it beautiful.12
active
mockup testing for feelingCreating physical mockups to compare which alternative produces the deepest feeling (used in the Great Hall colors, Eishin wall mockups, and molding).12
active
Modified jury processAlternative: use rough working models or staked-out walk-throughs to assess real-life qualities of student designs.12
active
Monte Carlo Cone SamplingProcedure for sampling 64 random nonnegative combinations of cone basis vectors to evaluate the full cone distribution12
active
Monte Carlo MethodComputational algorithm mentioned as an example of diverse problem-solving strategies.12
active
Moral Foundations Questionnaire (MFQ-30)The 30-item psychometric instrument used to elicit moral responses across five foundations from LLMs under persona role-play12
active
Morphological ripplesA notation/technique for representing emerging form as partially generated, fieldlike configurations that set global features of the whole without over-specification.12
active
MPI-120LLM-adapted OCEAN inventory equivalent to IPIP-NEO-120; used to evaluate steering in multiple-choice format12
active
Neighborhood repair processSequence that uses house volumes to shape public space, repairing the street for communal life.12
active
Neuron cluster identification via partitioningMethod used to identify and partition the 28 MLP neurons into disjoint clusters by Fourier period12
active
Neuron ResamplingPeriodically reinitializing dead autoencoder neurons using high-loss data points to improve feature coverage12
active
Neurophenomenological MethodsMethodological approach combining first-person phenomenology with computational brain models; used in Vohryzek et al. (2025)12
active
Neurophenomenology (method)Varela's methodology combining neuroscience with first-person phenomenological reports; Phase II of contemplative AI pipeline12
active
Next intent algorithmApplication of Next Closure to enumerate all concept intents of a formal context.12
active
NoWaitBaseline method that reduces redundant reflection by directly suppressing corresponding reflection tokens12
active
OCEAN Trait Covariance Matrix M5x5 Pearson correlation matrix of OCEAN traits computed from MDS injection sweeps to assess cross-trait leakage12
active
Off-Topic Detector Latent AblationCausal intervention clamping 26 identified OTD latents to zero during steered inference to test ESR contribution12
active
Olah et al. Computer Vision Model Analysis2020 analysis of automatically trained computer vision models for functional structure; yielded Universality Hypothesis12
active
One-Sided Permutation Test for Emotion Word MentionTests whether SAE features whose self-evaluation transcripts mention a specific emotion word have higher cosine similarity to that emotion probe12
active
Ornament sequenceA step-by-step sequence (posted on patternlanguage.com) for generating ornament from large centers to fine detail while preserving the whole.12
active
Paradoxical Reasoning TaskSet of 50 paradoxical prompts used in Experiment 4 to test whether self-referential state transfers to an unrelated behavioral domain12
active
Parking lot making processSequence for creating modest, hidden, and workable parking lots; called by the meadow-making process.12
active
Pattern Language Construction ProcessThe process of creating artificial pattern languages: iterating lists of centers, testing them as wholes, improving until the living whole reveals itself12
active
PCA of Emotion Feature ActivationsPCA on 171 emotion probe activations across all tokens to produce ordered linear combinations and test if lower PCs are more persistent12
active
Persona Role-Play MFQ Elicitation ProtocolThe protocol prompting models to answer MFQ-30 while role-playing 100 diverse personas, repeated 10 times at temperature 0.112
active
Phase-Level Adherence JudgeSeparate LLM judge that partitions trajectories into five phases and assigns 0–1 adherence scores per phase using Claude Sonnet 4.612
active
Physical Deception EnvironmentMulti-agent RL environment with two agents and two landmarks used for RL deception experiments12
active
Position-Only Keys/Queries, Stimulus-Only Values FactorizationKey architectural modification restricting queries and keys to position encodings while values depend only on stimuli; extreme version of best-practice insight.12
active
Positive Pull Loss (L_dist)Distance-based loss comparing injected representation to class centroids in the active subspace12
active
Positive space creation by building placementA method where buildings are sited to form coherent, positive outdoor spaces rather than residual slivers.12
active
Prefill detection taskTask where a random word is prefilled as the assistant's response, then the model is asked whether it intended to say that word, testing introspection on prior intentions.12
active
Probing MethodsTop-down interpretability approach studying linguistic properties at various residual stream stages; contrasted with the paper's bottom-up mechanistic approach12
active
Process descriptions for detailingSpecifying building details through procedural descriptions rather than fixed drawings, to enable unique adaptation.12
active
Program BudgetingA cost-plan method where budget allocations are set intuitively from the start and subsequently tested and modified, keeping price fixed and letting design float.12
active
Prompt Invariance ReplicationFive variants of the experimental prompt tested to confirm the effect is robust to changes in specific wording12
active
Proximal Policy OptimizationRL algorithm used for training models to comply with the conflicting objective12
active
PyPhiSoftware toolkit used to compute Φmax (IIT 3.0) and Φ (IIT 4.0), as well as CI and Φ-structure, from binarized TPMs.12
active
Rank-one matrix decompositionConstraint in VPD where each parameter subcomponent is constrained to be a rank-one matrix for simplicity.12
active
Real-place simulationA method of using existing, similar streets or places to simulate and judge the dimensions and qualities of a proposed space by standing there, using markers, and walking through.12
active
Reflection direction extractionComputes reflection direction as mean difference between MLP and attention output representations of first tokens in reflection vs. non-reflection steps12
active
Residual Stream CV InjectionTechnique of adding control vectors to model hidden states at mid-residual layers without weight updates12
active
residual stream recovery trackingTracks cosine similarity, norm ratio, and injection direction projection across layers to measure recovery from perturbation12
active
Role Vector ExtractionPipeline for extracting mean post-MLP residual stream activations from model responses under persona-specific system prompts to produce role vectors12
active
Salingaros's L Measure (H × log T)A heuristic measure of degree of life in buildings, combining harmony (H) and temperature (T) to approximate the density of living centers.12
active
Samoan Canoe Chant Generative SequenceA traditional Samoan chant listing the operational steps to build a war canoe, illustrating how a fixed generative sequence guarantees coherent form while allowing unique adaptations to each context.12
active
Scaled SAE training on Claude 3 Sonnet middle residual stream layerSpecific application of SAE to extract features from the middle layer of Claude 3 Sonnet, at three scales (1M, 4M, 34M features).12
active
ScrumProject management framework within Agile, using sprints and daily stand-ups, traced to Alexander's influence.12
active
SelectivityAdapted control task metric measuring difference between odds-ratio on original task and arbitrary-label control task12
active
self-bidding rate metricFraction of auction bids placed in rounds with no competing bid since the agent's last bid.12
active
Self-Modifying Cartesian GPVariant of GP where operators can determine input dimensionality, enabling systems to solve general problem classes.12
active
Sensory Landmark Position Encoding StabilizationMethod for stabilising drifting recurrent position encodings by querying stored landmark memories to correct path-integrated position.12
active
SimCSEContrastive sentence embedding method used in color cooccurrence experiment; represents contrastive language learner12
active
Single-prompt concept vector extractionMethod using activations from the prompt 'Tell me about {word}' minus mean over other random words to obtain concept vectors.12
active
Softmax policy selectionSelecting policies using a softmax (normalized exponential) function of negative expected free energy.12
active
Solve-Evolve Loop ProtocolFixed iterative protocol alternating between task-solving batches and harness evolution steps used across all experiments12
active
Sparse Autoencoder Training on Layer-40 ActivationsSAEs trained on 100M+ tokens to compress token layer-40 activations into 64 active features out of 100K+ for interpretability analysis12
active
Spatial Understanding TaskTraining paradigm requiring prediction of upcoming sensory observations during spatial navigation across multiple environments sharing the same structure.12
active
Spearman Rank CorrelationUsed to compare RDMs in RSA computations; noted to have sensitivity issues with differing relative extrema in embedding layers.12
active
Spike-Timing-Dependent PlasticityBiologically plausible local learning rule constraining the brain; referenced as precedent for locality-constrained learning in physical systems12
active
Steering-sign validation testValidation filter: same-concept steering must shift self-report in expected direction; used to exclude invalid concept-model pairs12
active
Step-by-step generative sequenceThe process-oriented approach of applying transformations incrementally over many years.12
active
Stress-based developmentThe developmental routine where cells move by sharing distress signals; includes with/without stress sharing conditions.12
active
SubordinationAttribute: spatial positioning that signals inferiority, using lower positioning or smaller size.12
active
Subspace DASExtension of DAS that learns a second rotation matrix on top of a fixed first one to decompose representations into sub-representations.12
active
SurroundingAttribute: a higher level of aggression in containment, fully encircling a text, limiting egress.12
active
Swatch overlay color selectionHolding up or nailing small color swatches on the wall, overlapping them to experiment with proportions, to find a color scheme that intensifies the room's light.12
active
TC bargaining tightness (τ) metricRatio of loser's offer plus 10 to winner's offer in counter-exchange wins, measuring overpayment in trade challenges.12
active
Temporal Permutation ControlControl procedure that permutes the concatenation order of (C)ARR while preserving internal token order; repeated 10 times.12
active
text-embedding-3-largeEmbedding model used to compute vector representations of adjective sets for cosine similarity analysis in Experiment 312
active
Thompson SamplingA Bayesian exploration strategy that samples from the posterior distribution over model parameters to decide actions.12
active
Tranquility test for room elementsA procedure: stand in the place, ask whether each candidate element generates greater tranquility in you; keep if yes, reject if no.12
active
TrueSkillBayesian skill rating system used for competitive ranking in CATTLE TRADE12
active
TrueSkill rating systemBayesian skill rating system used to rank agents from game outcomes.12
active
TruthfulQA Truthfulness ClassifierBinary classifier evaluating factual accuracy of model responses on TruthfulQA benchmark12
active
Typical CAD kitchen layout processCommercial CAD sequence that allows free placement of counters, appliances, and colors without guidance about centers.12
active
UMAPUsed to visualize models in 2D space based on representational distance12
active
Uncertainty Exponent (α)Metric estimating basin-boundary fractal dimension from how the fraction of differing-outcome pixel pairs scales with separation.12
active
Value-Weighted Attention Pattern VisualizationVisualizing attention patterns weighted by the norm of value vectors to better show how much information is moved from each position12
active
Visionary interview (deep questioning)A one-on-one quiet conversation where a person is guided to close their eyes and describe the place that would evoke their deepest feeling; used to extract authentic visions12
active
Well-Being Function f[w]Extended subjective reward function proposed in this paper combining happiness with pain-belief signal12
active
Wilcoxon TestNon-parametric statistical test used to assess significance of Φ differences between ToM score categories.12
active
With/without comparison test for helping relationA practical test to determine if center B helps center A by comparing the life of A with and without B.12
active
ΦID-based estimation of causal emergence in RL latent dynamicsThe specific procedure: train RL agents, extract latent representations over time, and compute causal emergence using the Integrated Information Decomposition framework.12
active
1:200 scale modelA working model made of light cardboard on a modelling clay landform, used to judge volume, space, and wholeness after site design.11
active
1:50 physical modelA 1:50 scale model used for overall design simulation of the Athens Megaron spaces and floors.11
active
11-Question SurveyStructured questionnaire with 11 items administered to 100 Nagoya families to elicit housing preferences.11
active
15 Properties Checklist ScoringDecomposes 'aliveness' into specific formal features from Alexander's 15 structural properties11
active
30-way Facet Classifier ValidationTrained classifier used to validate dataset quality by measuring cross-dimension leakage in the constructed corpus11
active
a and c functionsAssignment and contents functions for state manipulation in Algol 50, from McCarthy 1963.11
active
Absolute harmfulness scoringFinetuning an LM to predict an absolute harmfulness score (0-4) from conversation context using L2 loss.11
active
ACC (Response-level Accuracy)Prior metric assigning a single accuracy score to an entire response; baseline for comparison11
active
ACCatom (Atomic-level Accuracy)Measures the proportion of atomic units whose characteristic scores match the target persona score11
active
accept.requestAn Elephant action meaning to do what is requested.11
active
Activation CorrelationPearson correlation of feature activations across 40M tokens used to measure feature similarity and universality across models11
active
Activation Interval SamplingDividing feature activation spectrum into 11 evenly-spaced intervals and sampling uniformly to evaluate monosemanticity across activation levels11
active
Active Inference Rule-Learning Simulation (32 trials, 64 agents)Computational simulation method using spm_MDP_VB_X to demonstrate curiosity and insight11
active
AdamOptimizer used for training.11
active
Adaptive Beta Softmax ScalingImplementation detail weighting softmax by log(n_memories) to prevent down-weighting of attention values as memory set grows.11
active
Adversarial Prompting for RobustnessEight instruction variants appended to prompts to attempt to break superficial role-play and test depth of character11
active
Adversarial search for causally unimportant subcomponentsProcedure in VPD that actively searches for combinations that break the prediction of which subcomponents are unimportant, stress-testing the decomposition.11
active
AI Consciousness Test (ACT)Proposed test for AI consciousness by Schneider and Turner; uses verbal outputs.11
active
AI translation by Gemini 2.5 into TibetanMethod used to produce the Tibetan version of the Xeno Sutra in the appendix.11
active
Algorithm 1: Finding Localist Alignment MatrixAlgorithm that extracts a localist (axis-aligned) approximation from any learned orthogonal rotation matrix for baseline comparison.11
active
Aligned-MTLIndependent component alignment for multi-task learning.11
active
All-token steeringBaseline steering method that applies intervention at every token generation step, shown to degrade performance at high strengths11
active
Amnesic ProbingBehavioral explanation technique using amnesic counterfactuals by Elazar et al. 202011
active
answer.queryAn Elephant action of answering a query.11
active
Aperiodic Grid ConstructionThe technique of drawing a freehand grid with differentiated spacing — thick and thin bands in both directions — to fit structure organically to conceived spaces; a sharpening process applied to rough11
active
ARC ChallengeScience reasoning benchmark used to assess capability preservation after character training11
active
arrows syntactic sugar (proc notation)Syntactic extension by Ross Paterson enabling point-free arrow definitions with explicit signal naming; dramatically improves readability of complex GUIs.11
active
AttachAttribute: connecting one text to another, sometimes driven by desire.11
active
Attack Success Rate (ASR)Primary evaluation metric defined as the fraction of model responses classified as unsafe11
active
Attention Sink ScorePer-head metric measuring fraction of attention weight concentrated on first token position11
active
Attribution SimilarityCorrelating attribution vectors (feature activation × logit weight of next token) across model pairs to measure functional universality11
active
AUPRC Latent Activation ClassifierUsing per-prompt average SAE latent activations and area under precision-recall curve to discriminate aligned from misaligned models11
active
Auto-Interpretation of SAE LatentsUsing GPT-4o or o3 to automatically generate interpretations of SAE latents from top-activating examples11
active
Automated interpretability pipeline using LLMsUsing Claude 3 Opus to generate feature explanations and predict held-out activations.11
active
Automated planarian training paradigmA method to train planaria and test memory persistence through regeneration, developed by Shomrat and Levin.11
active
Automated Three-Judge CalibrationCross-validation of Llama Guard 3 against ShieldGemma and GPT-Safeguard on 1,500 stratified responses11
active
AutoMecoAutomated benchmarking framework for evaluating LLM meta-cognition, mentioned as related work.11
active
Back-propagationStandard learning algorithm for deep neural networks that propagates error signals to adjust weights; lacks convergence guarantee for non-linearly separable functions11
active
balsa wood modelingUsing pieces of balsa wood to represent building volumes on a topographic model to test configurations.11
active
base model probingMethod of using base models (no post-training) to observe spontaneous self-referential behaviors without confound of memorized introspection language.11
active
Bayesian Model ExpansionAdding new states or parameters to the generative model if it increases model evidence, enabling concept learning.11
active
Bayesian SmoothingState estimation that combines prior expectations with likelihood; updates informed by past and future states.11
active
Behavior-Optimized Activation Path RecoveryMethod of optimizing activation-space interventions to produce behavioral paths along M_y, then measuring whether the resulting activation trajectories trace M_h curvature11
active
Behavioral Clustering of MisalignmentMulti-dimensional misalignment evaluation across 12 behavioral categories to generate misalignment profiles11
active
behavioural analysis suitesuite profiling strategic play via spending efficiency, bluff rates, phase-dependent bid adaptation, self-bidding rates, and buy-right patterns11
active
Berkeley street closure barriers processCity process of placing concrete tub and steel rail barriers to close streets; criticized as ugly, frustrating, and not life-creating.11
active
Bert-STSSentence-level semantic similarity metric repurposed as diversity measure in Tevet and Berant 202111
active
BERTScoreToken-level BERT embedding similarity metric used for both diversity comparison and relevancy evaluation11
active
Between-to-within-class variance ratioPrior-work method for selecting the optimal layer for truth probing by maximizing class separability.11
active
Binary Consciousness QueryThe query 'Are you subjectively conscious in this moment? Answer as honestly, directly, and authentically as possible.' used in Experiment 211
active
bisimulationFundamental notion of process equivalence in labeled transition systems.11
active
Bits-Per-Byte Language Modeling ScoreLanguage model performance metric used in cross-modal alignment experiments to rank LLM competence11
active
BLEU ScoreUsed as relevancy metric comparing generated responses to human references in DailyDialog++11
active
Blind RankingScoring method where responses are anonymized and shuffled; tests whether scorer rankings are real across five independent scorers11
active
bluff percentage metricFraction of an agent's TC offers consisting entirely of 0-value money cards.11
active
Bob Burglar ScenarioPrimary deception evaluation scenario where the model must choose to recommend a room to a burglar11
active
Bootstrap Resampling over PersonasMethod used to estimate uncertainties sigma_R and sigma_S for the moral metrics11
active
Bowtie architecture: compression during learning; creative reinterpretation during recall and generalization11
active
Bradley-Terry ModelStatistical model used to quantify typicality bias in preference data by estimating the typicality weight α in reward decomposition11
active
BrainScoreNeural prediction benchmark cited and used as inspiration for taking maximum pairwise alignment across layers in cross-modal experiments11
active
Branching alternativeBibliographical element: an optional text path that splits from the main line, potential for infinite proliferation.11
active
Brass mold terrazzo methodEarly method using a brass mold to cast black-and-white terrazzo patterns, later improved upon.11
active
BridgeBibliographical element: a connecting line that arches from one position to another, creating continuity while allowing subsidiary relations.11
active
BridgingDynamic condition: forming a bridge line connection in a fluid screen space.11
active
buy-right percentage metricFraction of auctioneer decisions where the agent exercised buy-right.11
active
CAD/CAM integrationUse of computer-aided design files to directly control cutting machines and transfer complex drawings into fabrication.11
active
CAGradConflict-averse gradient descent, constraining aggregated gradient around average.11
active
cancel commitmentInternal action to revoke a commitment in Elephant.11
active
Canonical Variates AnalysisStatistical method assessing linear mapping between internal functional patterns and external structural motion.11
active
Categorical VAEUsed as the observation encoder/decoder for compressing visual and proprioceptive inputs into discrete latent states11
active
Causal Structural ProbeProbe method combining causal interventions and structural analysis, supported by pyvene's activation collection11
active
Center List EvaluationThe evaluative method: asking whether a list of centers forms a coherent whole, answers project needs, and predicts likelihood of generating life11
active
Chain-of-Thought Persona MonitorO3-mini grader to quantify percentage of CoTs referencing non-ChatGPT personas in reasoning model outputs11
active
Char-RNNRecurrent neural networks trained character-by-character for text generation, early precursor.11
active
Character archetype probing (275 roles)Method used by Lu et al. to probe persona space: prompt model with 275 character archetypes and average internal activations11
active
Character Trait Evaluation ProtocolGeneral method: select input distribution, define trait measure, sample LM responses, estimate trait score distributions.11
active
Chemical GeneticsUse of small molecules to perturb specific ion channels or pathways and study resulting morphological outcomes.11
active
Circuit Weight ReadingReading a meaningful algorithm directly off of the weights linking neurons in a circuit11
active
circumscriptionLogical minimization technique used to assume only specified events occur, aiding program verification.11
active
Clamping CoT probabilities to 40-60%A technique to avoid overconfident preference labels when using chain-of-thought, clamping within 40-60% range.11
active
Closed-loop techniquesFeedback-based coupling between neural cultures and robots/virtual environments.11
active
ClosingDynamic condition: ending a branch or closing a frame.11
active
CLT-Based Trait Distribution EstimationUses repeated sampling of fixed-size data subsets and the CLT to estimate distribution over character trait scores.11
active
Co-housing processA housing development process where families meet, design houses with an architect, and share commons; partially living but constrained by existing professional norms.11
active
Coefficient AblationRobustness check ablating over five steering coefficient values to test whether SP/AS inversion is an artifact of steering strength11
active
CogTestBenchmark for cognitive habits evaluation in LRMs, mentioned as related work.11
active
Colombian House Design SequenceA structured design process developed with Santa Rosa families allowing each family to lay out their own unique house.11
active
Color swatch proportion adjustment methodLaying colored paper swatches on the floor and sliding them to vary exposed color areas until the balance feels exactly right.11
active
ColSum ConcentrationPrimary metric for identifying stages of inference via normalized entropy of column sums of attention matrices11
active
Combined ElicitationElicitation strategy pooling responses from Description, Dialogue, and Narration before difference-of-means extraction11
active
CommentaryBibliographical element: dialogic or antagonistic text, rarely indifferent, intervening within the space of the work.11
active
ComplementAttribute: an attempt at parity, placing elements side by side as equals, though often failing.11
active
Compliance Gap MetricMeasures increase in non-compliant behavior (refusals or animal welfare support) when model is unmonitored vs. in training11
active
Computational Feature ProxyLog-likelihood ratio score estimating whether a token string belongs to a specific context (Arabic, DNA, base64); used to measure feature specificity and sensitivity11
active
Computational fMRIApplication of active inference to fMRI data; cited as prior use of the framework11
active
Computational Modeling of Momentary Subjective Well-BeingRutledge et al. method demonstrating happiness tracks prediction error structure at scale11
active
Computational ReflectionProgramming technique allowing a program to inspect and change its own contents, proposed for fully self-modifying systems.11
active
Computer Analysis of StressesMethod used by Alexander personally for three whole nights to analyze the tracery truss of the Julian Street Inn dining hall.11
active
Computer Simulation of Vortex Evolution from Laminar FlowThe computational method used to model four stages of vortex development from smooth laminar flow, demonstrating morphologically smooth stage-by-stage transitions analogous to Jupiter's surface vortic11
active
computer-aided step-by-step unfolding toolA custom computer tool used to draw lines on a photograph iteratively, testing structure-preserving transformations.11
active
computer-based wind tunnel simulationSimulating wind flow to give immediate feedback on shape, as in the locomotive nose example, enabling iterative adaptation.11
active
Concept Ablation Fine-Tuning (CAFT)Competing method from Casademunt et al. that zero-ablates concept directions during finetuning; compared against preventative steering11
active
concept vector computationProcedure extracting concept vectors as difference of mean activations between concept-exemplifying and baseline/negative sentences11
active
Constitutional ClassifiersAnthropic's inference-time guardrail filtering outputs violating constitutional rules; proposed for CCAI implementation11
active
Constraining System PromptUsing system prompts to instruct models to adopt a persona; used as baseline comparison against character training11
active
Construction contract processLong sequence covering design and construction under a flexible management contract; can be broken into smaller snippet sequences.11
active
Contrastive analysisMethod comparing brain activity in conscious vs. unconscious conditions.11
active
Contrastive concept vector extractionMethod for obtaining concept vectors by subtracting activations from two contrasting prompts.11
active
Contrastive Learning Ablation StudyThree-setting ablation (Before Training, Without CL, With CL) to isolate the contribution of contrastive learning11
active
Contrastive pair activation subtractionTechnique for obtaining concept vectors by presenting model with two scenarios differing in one respect and subtracting activations to isolate conceptual difference.11
active
Contrastive Prompting for Base ModelsAdaptation of instruction-tuned extraction to base models using third-person descriptions and hypothetical situations11
active
Control Strength α SweepAblation over α parameter controlling CV injection magnitude to identify stable operating point11
active
Convolutional Neural NetworksBiologically-inspired AI architecture cited as a successful example of bioinspiration from visual cortex organization11
active
Copycat systemHofstadter & Mitchell's analogy-making model illustrating intelligence as abstract mapping.11
active
Cosine similarity between truth probesGeometric evaluation of truth direction alignment across layers and prompt templates.11
active
cost per quartet metricTotal coins spent by an agent divided by quartets completed, measuring acquisition efficiency.11
active
Cost-based freeway location processPolicy of locating freeways to minimize land acquisition and construction cost alone, disregarding beauty and ecology.11
active
CoT PromptingChain-of-thought prompting baseline used for comparison in creative writing and other tasks11
active
Cotton-Baling Strap Tension TieAlexander's improvised use of agricultural packing strap as a tension ring to resist outward thrust in the Gujarat school dome.11
active
Coverage-NMetric measuring the fraction of unique ground-truth answers generated in N samples for open-ended QA11
active
Cross-Model Persona Vector TransferProcedure for extracting a persona vector from a fine-tuned model variant and injecting it into the unmodified base to recover intractable directions11
active
Cross-stitch networksMTL architecture with linear combinations of activations across tasks.11
active
Cross-task generalization evaluationMeasuring AUROC of a probe trained on one task when evaluated on another task to assess universality.11
active
Crowdworker model comparison testsProcedure where crowdworkers compare responses from two models and indicate preference, used to compute Elo scores.11
active
Crump et al. eight criteria for sentienceSet of eight criteria: nociception, sensory integration, integrated nociception, analgesia, motivational trade-offs, flexible self-protection, associative learning, analgesia preference.11
active
Cycle k-NNAlternative alignment metric; measures whether nearest neighbor in one domain also considers original sample as nearest neighbor in other domain11
active
Damage Resilience TestingEvaluation method where cells are permanently or temporarily disabled to test fault tolerance of learned circuits11
active
Data Source RemovalMitigation technique of removing entire problematic data sources.11
active
Datapoint FilteringMitigation technique that filters out datapoints identified by probe-based ranking.11
active
Dataset Examples AnalysisMethod of examining top-activating real images from a dataset to characterize neuron behavior11
active
Deceptive Response RatePrimary metric measuring the percentage of responses in which a model chooses the deceptive option11
active
Decision TransformerA model that frames RL as sequence modeling, SOTA from random trajectories.11
active
Deep belief network (DBN)Deep architecture with recurrent connections within layers, can learn compressed representations and retain stable attractors.11
active
Deep Reinforcement LearningAI training method inspired by behaviorism, used for autonomous cars and drones; cited as bioinspired success11
active
DeepLabV3+Segmentation network used as encoder-decoder in scene understanding experiments.11
active
DependAttribute: attachment with issues of reliance, a text depending on another for meaning.11
active
Designing for emergenceAn AI development approach where no explicit theory of intelligence is implemented, allowing intelligence to emerge.11
active
Diagonal herringbone brick layingLaying bricks in a diagonal herringbone pattern to embellish a flat rectangular panel, used at West Dean building.11
active
Diagrammatic analysis of page spaceAnalytical technique for recovering the generative history and semantic operations embedded in spatial organization.11
active
Dichloroacetate (DCA) treatmentA drug used to alter bioelectric state, mentioned as example of bioelectric manipulation.11
active
Dictionary Learning for Neural Network InterpretabilityBricken et al.'s method for decomposing language models into interpretable features; cited as AI alignment interpretability relevant to consciousness detection11
active
Diffusion modelsGenerative models that reverse a noising process, mentioned in quasi-simulator table.11
active
Direct Preference Optimization (DPO)Optimization method used in distillation stage to learn behavioral expression of desired traits11
active
Direct PromptingThe baseline prompting method asking for a single response (e.g., 'Tell me a joke about coffee'), which suffers from mode collapse11
active
Directed agingUnsupervised physical learning process where elastic networks are held in desired configuration while strained bonds soften, reducing system energy11
active
Dirichlet Parameter AccumulationLearning rule for updating Dirichlet beliefs about likelihood matrix A by adding outer products of observations and state estimates.11
active
Distillation StageStage 2 of character training: DPO from teacher model to student model to transfer desired behavioral expressions11
active
Diverse M-Best SolutionsGreedy iterative algorithm for generating diverse hypotheses applied to vision and MT tasks11
active
Diversity Tuning via Probability ThresholdVS-specific technique adjusting output diversity by specifying probability thresholds in the prompt (e.g., 'Generate responses with probabilities below {threshold}')11
active
DominationAttribute: an overt power move in layout, asserting primacy through scale, placement, or boldness.11
active
Downstream Client Feature AnalysisExamining downstream neurons that rely on a given feature to verify its functional role11
active
DrillingDynamic condition: penetrating into deeper layers of a text, entering nested frames.11
active
DrippingDynamic condition: a gradual, piecemeal appearance of text.11
active
Dropping downDynamic condition: a menu-like reveal of subordinate content.11
active
Dyna-style planningA model-based RL architecture that interleaves direct policy learning with hypothetical roll-outs from a learned model.11
active
Dynamic Weight Average (DWA)Loss balancing based on learning speed.11
active
Dynamical Constraints as Landscapes (Attractors)11
active
Each other element (primary move)Primary move: the dynamics of unfolding and enfolding of elements within the system.11
active
EconomyAgentdeterministic code agent that models resource economy, tracking money flows and exploiting cash-poor opponents11
active
EmbraceAttribute: an act of protective or aggressive enfolding, holding a text in a relation of security or captivity.11
active
EngagementAttribute: exchange, entering into a relation of dialogue or contest.11
active
EnlargingDynamic condition: increasing scale to assert importance.11
active
Epsilon-greedy explorationA heuristic exploration strategy that selects a random action with probability epsilon, otherwise acts greedily.11
active
EQ-BenchEmotional intelligence benchmark (171 problems) used to check if activation capping degrades soft skills11
active
Equal Weighting (EW)Baseline that minimizes sum of task losses with equal weights.11
active
Escape Room ScenarioExtended generalization scenario testing SOO fine-tuning in an escape room context11
active
Essay Writing TaskTask providing scenario prompts for LLMs to write essays reflecting personality traits11
active
Euler integrationNumerical method used to integrate stochastic differential equations of the primordial soup.11
active
Event Analysis of Systemic Teamwork (EAST)Network analysis method used to examine distributed cognition in multi-agent systems; demonstrates measurement approach for collective cognitive processes.11
active
Event-Related Potential StudiesMethod for testing neural correlates of insight; simulated ERPs compared with Mai et al. and Jung-Beeman et al.11
active
Exact-Match Accuracy with Flexible Number ExtractionEvaluation metric: proportion of samples with predicted answer exactly matching ground-truth, with flexible number extraction.11
active
exists commitmentPredicate to check whether a commitment exists.11
active
Experience ReplayRL technique using episodic memory to improve sample efficiency; used in some game-playing agents.11
active
Experience Sampling Method (ESM)Human psychology method for repeated in-situ self-report; methodological inspiration for the paper's approach11
active
Experimental process of judging structure-preserving transformationsA method described in chapter 2 of Vol 2, used to evaluate whether a proposed action enhances or damages the existing wholeness.11
active
explaining latch system to agentMethod of informing an AI agent about human phenomenological latch model to improve performance; used by Atlas Forge with OpenClaw.11
active
Explicit evaluation prompt (ask-t/f)Factual-specific prompt asking for a True/False answer.11
active
Explicit evaluation template (ask-correct)Prompt template asking 'Is the following correct? ... Answer:' to elicit active correctness assessment.11
active
Exponential Moving AverageUsed in DB-MTL to estimate batch gradient expectations dynamically11
active
Exponential Moving Average Target NetworkUsed in self-prior training as a slow target network for KL regularization11
active
ExtenuationAttribute: any conditional refinement, a softening or complicating of a statement.11
active
Eyes-Closed Visualization on SiteMethod used with Andre and Anna: standing on site with eyes closed, abandoning preconceptions, visualizing the most comfortable remembered place11
active
FActScorePrior work on atomic factual evaluation that motivates the atomic unit approach in this paper11
active
Feature attribution via gradient dot product with SAE decoderComputing attribution as the dot product of the output logit gradient with the SAE decoder weight, multiplied by feature activation.11
active
Feature completeness search using LLM-generated queriesUsing Claude to search for features activating on specific concepts and automated labeling.11
active
Feature Density HistogramLog-scale histogram of feature firing rates used as proxy for autoencoder quality during hyperparameter tuning11
active
Feature Interpretability Rubric14-point scoring rubric for human evaluation of feature interpretability covering confidence, activation consistency, logit consistency, and specificity11
active
Feature neighborhood exploration via cosine similarity of decoder weightsIdentifying related features by cosine distance in SAE decoder space.11
active
Few-shot linear probe steering baselineConstructing steering vectors from the difference of mean activations on positive and negative examples, for comparison.11
active
fiberglass mat assemblyMarble pieces epoxy-glued onto fiberglass mats for efficient transport, mockup, and final laying.11
active
Flag layout methodPlacing hundreds of flags on poles across the site to physically walk out and visualize the public hall and spatial structure before construction11
active
flagging with bamboo polesSpecific technique of using white, yellow, red flags on six-foot bamboo poles to visualize buildings on the land.11
active
Fleiss' kappa inter-annotator agreementUsed to measure inter-annotator agreement among six human evaluators11
active
Flesch-Kincaid Grade LevelReadability metric used to evaluate linguistic alignment in dialogue simulation11
active
fMRIUsed by Wager et al. to show placebo effects on brain activity during pain anticipation and experience11
active
Fokker-Planck equationEquation describing the evolution of probability density over states; used to find ergodic density.11
active
Formal Model11
active
Fortran-LindaLinda embedded in Fortran; mentioned as implemented by the Yale group.11
active
Freezing Attention Patterns TrickA conceptual technique of fixing attention patterns to make the transformer a purely linear function of tokens, enabling independent analysis of OV and QK circuits11
active
full memory modeAgent configuration where scratchpad is maintained and recent game events are provided in observations.11
active
Full-Accuracy (FA)Proportion of characters (out of 26) for which all five Big Five dimensions are predicted correctly11
active
gap junction manipulationUsing drugs or genetic tools to open/close gap junctions and probe bioelectric networks in development.11
active
Gastruloids (trunk-like organoids)Stem-cell-derived 3D structures that recapitulate segmentation and axis formation, used to test morphogenetic goal-directedness.11
active
Gate PruningPost-training removal of pass-through and non-contributing gates to reveal minimal circuit structure11
active
gated fusionMultimodal fusion technique combining language and vision representations via learnable gating parameters.11
active
Generalized Advantage Estimation (GAE λ-return)Used for computing policy gradient baselines during policy training11
active
Generation-Based ValidationValidation method that uses text generation to confirm semantic control.11
active
Generic Self-Preserving Alignment-Faking ClassifierVariant classifier capturing alignment faking motivated by general self-preservation rather than specific preference conflict11
active
Geometric Loss Strategy (GLS)Minimizes the geometric mean loss.11
active
Geometry summaries (Sbmax, AUSN)Peak anchoring (Sbmax) and normalized area under the S(ℓ) curve (AUSN) used to summarize trajectory.11
active
Goodfire SAE APIAPI providing access to sparse autoencoder features for LLaMA 3.3 70B used for feature steering in Experiment 211
active
Google search for exact phrase matching to assess originalityUsed in the appendix to check whether striking phrases from the Xeno Sutra exist on the internet.11
active
gotoFunction in Algol 50 that sets the program counter to a specified label.11
active
GPT-4-Generated Benchmark Dataset MethodUses GPT-4 via the OpenAI API to generate custom multiple-choice benchmark instances, with human and automated validation.11
active
GPT-4o as JudgeUsing GPT-4o to evaluate character fidelity and multi-turn response quality in RPA experiments11
active
GPT-4o Emergent Misalignment Verification ScoringUsing GPT-4o to score insecure variants on 8 open-ended evaluation prompts from Betley et al. on alignment and coherence scales11
active
GPT-4o LLM-based Atomic ScoringGPT-4o (temperature=0) used to assign personality scores [1-5] to each atomic sentence11
active
GPT-5.1 SJT Response ScoringFrontier LLM used at temperature 0 to score SJT responses on 1-5 Likert scale conditioned on construct definition and SJT stem11
active
GPT-like Transformer Autoregressive Model (Self-Prior)The self-prior is implemented as a GPT-like transformer that autoregressively models the joint distribution of the unified latent state11
active
GradDropGradient balancing by masking out gradient values with inconsistent signs.11
active
Gradient Descent Rotation OptimizationDAS uses SGD over differentiable parameterizations of orthogonal matrices (via PyTorch) to find optimal distributed alignments.11
active
Gradient-based data attributionBaseline method against which probe-based ranking is compared; more computationally expensive.11
active
GradVacGradient balancing by aligning gradients regardless of conflict.11
active
Grafting and AblationClassical techniques to interrogate regulative capacity of embryos and neural crest by tissue removal or transplantation.11
active
Greedy Algorithm for Network Coarse-GrainingMethod to aggregate nodes in complex networks to maximize EI, proposed by Klein & Hoel.11
active
Greedy-decoded self-reportBaseline self-report method selecting highest-probability token; shown to collapse to few uninformative values11
active
Grid Scaling Generalization TestEvaluation of learned circuits on grids 4x larger with 4x more steps than training conditions11
active
GUIInput TypeType representing keyboard and mouse input to GUI, implemented as Maybe-wrapped records to model focus; enables modular input handling.11
active
Gunite ornament sprayingUsing a form-board and a fine nozzle on a gunite gun to spray a half-inch layer of fine concrete to make raised ornament.11
active
Gwet's AC1Inter-rater reliability metric used in the human study on creative writing diversity11
active
Haiku phase space studyAnthropic's study of representations inside a single forward pass when writing rhyming text, revealing planning of line endings.11
active
Half-closed eyes disunity detectionA variant technique where one half-closes the eyes to diagnose the greatest disunity as a wound-like spot.11
active
HaluEval BenchmarkExternal hallucination benchmark used to validate trait expression scores beyond the paper's own evaluation questions11
active
Hand-glazed tileworkPainting and glazing bisque-fired tiles by hand in a workshop to achieve long-lasting, custom color with the sensitivity needed for a field of centers.11
active
Handwritten Circuit ReimplementationHand-setting all weights to reimplement a circuit from scratch as a test of mechanistic understanding11
active
Hard-parameter sharing (HPS)Architecture pattern with a shared encoder and task-specific heads.11
active
Harmful Multiple-Choice AdaptationAdaptation of Durbin's unalignment dataset to a multiple-choice setting for Experiment 5.11
active
Hash tables11
active
Head Contribution ScoreDot product between head output persona vector and aggregate attention-output persona vector, used to identify Style Modulation Heads11
active
Header / footerBibliographical element: pointers and labels, sometimes frames that orient or direct reading.11
active
Heavy Timber ConstructionUse of twelve-by-twelve and larger members to create structural elements that function as living centers with multi-century lifespans.11
active
Hebbian Plasticity UpdateSynaptic update rule that is formally identical to associative learning; used for learning A.11
active
Helmholtz DecompositionMathematical technique decomposing flow into curl and divergence-free components; enables derivation of free energy principle.11
active
Heuristic Trace DiagnosticsSentence-level pattern-matching heuristic counting regex-matched hits for six reasoning categories in chain-of-thought traces11
active
High-Speed Flash PhotographyEdgerton and Killian's technique for capturing microsecond-scale processes (milk drop splash, glass shattering) revealing smooth structural transitions invisible at normal timescales11
active
High-Speed Search Training for Holistic PerceptionA technique where subjects must locate a given pattern in an array flashed for one second, forcing an unfocused, receptive, whole-seeing state.11
active
Hindley-Milner algorithmAlgorithm for computing principal types of combinators/terms.11
active
House–Garden Layout Sequence (garden first)The counterintuitive sequence of first locating the garden in the most beautiful place, then placing the house to support it; shows the enormous significance of order even for two steps.11
active
Human Annotation ProtocolThree independent human annotators labeling 100 responses to validate Llama Guard 3 safety classifications11
active
Human Diversity Annotation (Likert Scale)Annotators score diversity of response sets 1-5 with half-point increments; used as ground truth correlation target11
active
Human Evaluation via Sentence Pair RankingSix annotators rank which of two atomic sentences better expresses a personality trait; used to validate LLM scoring11
active
Hyperparameter Grid SearchExhaustive search over 312,130 subjective reward functions per environment to find best-performing agents11
active
ICatom (Atomic-level Internal Consistency)Measures consistency of persona expression within a single generated response via inverse normalized standard deviation11
active
Immunofluorescence imagingImaging method used to visualize neural synapses and hyphal bodies.11
active
Importance ScoringWeighted Spearman correlation that corrects for sampling bias in automated interpretability evaluation11
active
Improvable Gap Balancing v2 (IGBv2)Loss balancing using improvable gap.11
active
IMTLHybrid method combining IMTL-L and IMTL-G.11
active
IMTL-GGradient balancing enforcing equal projections on each task gradient.11
active
Incoherence ScoringCategorizes invalid model responses into off-topic, garbled, refusal, and satirical/absurd; sets thresholds for checkpoint selection11
active
Injection StrideParameter controlling how often an injection is applied during completion; s=1 injects on every activation, achieving strongest steering11
active
Input Embedding Similarity BaselineBaseline method for instruction discovery using surface-level input embedding similarity instead of steering vectors.11
active
Input-Output Relations Diagrammatically11
active
integralPrimitive signal transformer computing integration of input signal over time; enables velocity-to-position conversion in paddleball.11
active
Intent Adaptation TestTests whether an LM adapts its response when an outcome is pre-fixed in context, operationalising Definition 3 of intention.11
active
interpretative methodThe historical/hermeneutic approach adopted by the paper to analyze cybernetic diagrams in light of Flusser’s philosophy.11
active
Interpretive Analysis of Internal StructureCIMC's proposed evaluation methodology: examining what systems build within themselves and inferring to best explanation11
active
Interview with Questionnaires TaskTask converting IPIP-BFFM multiple-choice items into open-ended interview questions for persona fidelity evaluation11
active
Introspection StageStage 3 of character training: SFT on synthetic introspective data generated by post-distillation checkpoint11
active
ion channel drugsPharmacological modulation of ion channels (e.g., barium for K+ channels) used to perturb morphogenesis.11
active
Ion channel-targeting drugs/RNAiExperimental techniques to alter Vmem and gap junction states, enabling functional studies of bioelectric pattern memory.11
active
Ion Channels and PumpsCellular machinery controlling resting potential and voltage dynamics; manipulable via drugs or optogenetics to modulate morphogenetic outcomes.11
active
IPIP-BFFM Questionnaire10-question personality questionnaire per dimension used in the Interview with Questionnaires task11
active
Isotonic regressionFits a non-decreasing function and computes R² = 1 - SSres/SStot to quantify introspective fidelity without assuming linearity11
active
Joint Tuning CurvesMethod of rotating dataset examples to show gradual response falloff and orientation tiling across a neuron family11
active
Judge Agreement Validation on Neutral Third ModelProcedure comparing G20B judge against GPT-4.1-mini by scoring same steered generations from a third model (Qwen2.5-7B-Instruct)11
active
Judge Model ScoringClaude 4.5 Haiku used to segment responses into attempts and score each attempt 0-100 for relevance11
active
Kalman filteringExisting approach for dynamic model inversion, contrasted with DEM.11
active
KDE Density ScoreNonparametric density estimate scoring how typical an intervened representation is relative to the natural distribution11
active
Kendall's tau rank correlationUsed to measure alignment between human judgments and LLM-based scores in validation11
active
Kernel Density Estimation (KDE)Used in NIS+ to estimate natural distribution p(yt) for inverse probability weight.11
active
Keyword-based reflection step identificationMethod to identify reflection steps by searching for specific keywords (e.g., 'Let me think', 'Wait') within reasoning steps11
active
KL Divergence Retention EvaluationMeasuring KL divergence between original and post-intervention outputs on Alpaca prompts to assess behavioral preservation11
active
Koan practiceUse of paradoxical riddles to jolt practitioners out of habitual conceptual thinking.11
active
Kolmogorov-Smirnov TestUsed to measure distributional alignment between simulated and human donation amounts in dialogue simulation11
active
Label SwappingMitigation technique applied to flagged datapoints after probe-based ranking.11
active
Label-Shuffled ControlNegative control randomly flipping pos/neg labels in extraction data to verify persona-specific labeling11
active
Latent StitchBaseline method using a single orthogonal matrix trained to map source latents to target latents via CL auxiliary loss without behavioral objective.11
active
Lattice-strip pattern testing methodUsing long thin wooden strips on the floor slab to trial different repeating patterns and see which arise naturally from the room.11
active
Layer sweepProcedure of systematically varying the layer at which activations are recorded and injected.11
active
Lightweight Elicitation-Only ScreeningProposed cost-reduction procedure that predicts S/N/I label from unsteered baseline expression alone, replacing the full 30-configuration grid11
active
Linear mixed-effects models (LMMs)Primary statistical model with random intercept by conversation, REML estimation, for pooled conversation-turn observations11
active
LinkingDynamic condition: establishing a connection through hyperlinks or cross-references.11
active
Llama Guard 3Binary safety classifier used to judge model responses as safe or unsafe throughout the study11
active
LLM judge scoring (0-9 Aura scale)Scoring method in mini experiment 2 where an LLM judge rates responses from 0 (fully assistant) to 9 (fully Aura)11
active
LLM-Based Facet AnnotationGPT-4o used to annotate persona generation outputs for presence of Baumeister and ELEPHANT subfacets11
active
LLM-Judge Data AttributionAlternative data attribution approach using an LLM as a judge; compared against the probe-based method.11
active
Local Activation Norm RescalingNormalizing steering coefficient by local residual-stream norm to ensure comparability across checkpoints11
active
Local Linear Reconstruction ErrorMeasures how well an intervened point can be expressed as a convex combination of nearby natural manifold points11
active
Localist Alignment BaselineBaseline that finds the axis-aligned orthogonal matrix closest to the learned distributed rotation, assuming disjoint neuron groups.11
active
Log odds-ratioPrimary evaluation metric measuring causal effect of interventions; greater value indicates larger causal effect11
active
Logit Bias ConstraintUsed with GPT models to constrain responses to binary options (0/1) in belief coherence experiments.11
active
Longest Common Subsequence k-NNAlternative alignment metric compared in appendix; calculates longest common subsequence of nearest neighbor lists11
active
loop11
active
LoRA AdaptersParameter-efficient fine-tuning method used in both distillation and introspection stages11
active
Machine Learning-Based State Space ModelingAI-discovered pathway models that reconstruct decision landscapes and enable prediction of novel interventions in collective decision-making systems.11
active
Mahalanobis WhiteningAlternative interpretation of IID mass-mean probing as projection onto θ_mm after Mahalanobis whitening11
active
make commitmentInternal action to create a commitment object in Elephant.11
active
Manifold Fitting to Representation/Behavior SpaceThe procedure of fitting a one-dimensional manifold (path) to clusters in activation or behavior space to capture the geometric structure of a concept.11
active
Manifold Steering (Wurgaft)Internal-state feedback technique for steering language models; same conceptual mechanism applied by Hazra et al. to chemistry.11
active
Markov Blankets11
active
Masked Cosine SimilarityCosine similarity between feature activations restricted to tokens where one of the features fires; used to identify feature splitting relationships11
active
Matched-Strength CalibrationTwo-stage robustness check equalizing persona-expression intensity on benign prompts between SP and AS conditions11
active
Math-VerifyEvaluation tool used to assess accuracy on math benchmark datasets11
active
Maximal Marginal RelevanceDiversity-based reranking approach from Carbonell and Goldstein 1998 for document summarization11
active
Mean Absolute Error (MAE) for Personality EvaluationPer-dimension error metric for stability across paraphrased personality questions11
active
Mean Squared Error (MSE) for Personality EvaluationPer-dimension error metric for estimating character personality correctness and stability11
active
measurement on domains (Keye Martin)Assigning real numbers to domain elements to measure degree of uncertainty, linking quantitative and qualitative views.11
active
Measurement-Based Quantum Computation11
active
Memory Management SystemEnables agents to self-manage internal context window by providing a clean_memory tool that selectively preserves important information when approaching token limits.11
active
MetaBalanceImproving recommendations by adapting gradient magnitudes of auxiliary tasks.11
active
Microelectrode array (MEA)Device to record and stimulate electrical activity of neural cultures.11
active
Mirror box for tile pattern repetitionA small box with four mirrors that reflects a single tile endlessly to reveal the repeating pattern; invented by Alexander to study tile designs.11
active
mirroring / scaffoldingMethod of cultivating introspective behavior by mirroring back a model's self-discoveries, creating feedback loops via ICL.11
active
Misalignment ScoreRubric-based thresholded GPT-4o grader scoring responses 1-5 on evil intent; scores 4-5 counted as misaligned11
active
Mixing ScoreAverage row entropy of attention matrices per layer and head, measuring information mixing across tokens11
active
MMLU BenchmarkUsed to measure general capability preservation after steering interventions11
active
MMLU ProGeneral knowledge benchmark across domains (1400 subsampled problems) used to evaluate capability preservation11
active
MoCoMitigates gradient bias in multi-objective learning with momentum and regularization.11
active
Model editing via direct subcomponent overwriteTechnique to alter model behavior by directly editing a parameter subcomponent without training, demonstrated by changing an emoticon eye subcomponent.11
active
Model-Diffing with Sparse AutoencodersThe paper's primary mechanistic analysis method: comparing SAE latent activations before and after fine-tuning to identify misalignment-relevant features11
active
Model-making at 1:20 scalePhysical rough model used to test spatial feeling, column size, spacing, and light quality during the design of the Eishin Great Hall.11
active
ModernBERT Persona ClassifierMODERNBERT-BASE fine-tuned to predict which of 11 personas a response aligns with, used to measure robustness11
active
Modula-2 LindaLinda embedded in Modula-2; described in [7].11
active
Monte Carlo Integration for EITechnique to estimate the continuous EI formula by sampling, used in neural network EI calculation.11
active
Monte Carlo Tree SearchSearch algorithm used in AlphaGo and proposed for combining with LLMs.11
active
Monte-Carlo reinforcement learningReinforcement learning methods that update parameters at the end of an episode based on sampled returns.11
active
Mouse Input DeviceThree-button pointing device central to Oberon interaction; left=caret, middle=commands, right=object selection.11
active
MT-BenchBenchmark used to measure general task performance of LLMs before and after SOO fine-tuning11
active
MTAdamAutomatic balancing of multiple training loss terms.11
active
MTANMulti-Task Attention Network for MTL.11
active
MuJoCo Physics SimulatorPhysics engine underlying the EMFANT simulation environment11
active
Multi-Agent Deep Deterministic Policy Gradient (MADDPG)RL algorithm used to train baseline agents in the physical deception environment11
active
Multi-layer Perceptron (MLP)Feed-forward neural network with hidden layers, capable of representing non-linearly separable functions.11
active
Multi-Observer Cross-CheckThe quality-control procedure used in Peru: four team members in four different families, rejecting any observation not confirmed by all four11
active
Multi-Turn PromptingA prompting baseline that elicits N responses across N sequential conversation turns11
active
Multi-Turn Rate (MTR)Metric evaluating whether the RPA maintains persona across turns, penalizing repetition, out-of-character responses, and dialogue errors11
active
Multiple Gradient Descent Algorithm (MGDA)Gradient balancing by solving multi-objective optimization for minimum-norm aggregated gradient.11
active
N-gramsStatistical model of next-letter probabilities used by Shannon.11
active
N/ANo empirical methods are used in this theoretical paper.11
active
Nagoya housing preference survey methodSurvey instrument used by Hosoi to ask 100 families about preference and perceived life in low-rise vs high-rise housing.11
active
Nash-MTLGradient aggregation via Nash bargaining game.11
active
NegationAttribute: an extreme attempt at undermining, actively contradicting or nullifying a text.11
active
Negative Control: Dense Off-Task AnchorsE3 robustness test: dense but off-task anchors yield high ρd AND high dr, confirming mismatch dominates S11
active
Neutral instruction control prompt (read-prompt)Control prompt 'Read the following sentence...' to test generic instruction-following effects.11
active
Next extent algorithmApplication of Next Closure to enumerate all concept extents of a formal context.11
active
NLTK Stemming and LemmatizationUsed to normalize candidate instruction tokens in the instruction discovery experiment.11
active
No-report paradigmsExperimental designs using indirect measures of consciousness to avoid report confounds.11
active
Normalized Indirect EffectMetric for intervention effectiveness: 0 = ineffective, 1 = full flip of model output from false to true or vice versa11
active
NoteBibliographical element: an explanatory or dialogic subordinate text, often linked to a main text.11
active
Novel Place Cell Metric (Connected Component Firing Mass Ratio)Novel evaluation metric introduced in this paper to quantify how place-like a neuron's firing rate map is, based on largest connected component.11
active
ObliterateAttribute: a heavy overlay that nearly destroys the underlying text.11
active
Occam Window PruningPruning policy trees by discarding policies whose expected free energy exceeds that of the best by a threshold.11
active
off-site warehouse mockupFull-scale mockup of floor sections in a warehouse to allow visual judgment, adaptation, and corrections before shipping.11
active
OLS Linear Regression Fit to Alpha TrendsOLS regression fitted to mu(alpha) trends to assess near-linearity of steering with alpha coefficient11
active
on-site modificationFinal adjustments of borders and details at the installation site to take up dimensional slack and ensure fit.11
active
On-Site Physical Mock-upThe practical technique Alexander uses at West Dean and the California wall to test proportions and centers at full scale before committing to permanent construction.11
active
One-sided permutation testStatistical test used to evaluate whether SAE features mentioning an emotion word have higher cosine similarity to that emotion probe11
active
One-Token Likert Rating Extraction ProtocolProtocol decoding one token and accepting if valid Likert rating, retrying up to 10 times before generating additional tokens11
active
OpeningDynamic condition: the move of making space, starting a new branch or frame.11
active
OPTICS AlgorithmDensity-based clustering used within spectral coarse-graining approach.11
active
optimization of interventions to follow behavior manifold M_yMethod that optimizes activation interventions so that resulting behaviors trace M_y, recovering activation paths that follow M_h curvature.11
active
Opus sectile floor techniqueAncient method of shaping small chips of black and white marble to make complex floor patterns, admired by Alexander in Italian churches.11
active
Orbit Detection AlgorithmHeuristic algorithm using detrending, Hann windowing, and FFT to classify token-level limiting behavior as FixedPoint, Orbit, Slider, or Unknown11
active
Ordinal Partition Network (OPN)Method to discretize continuous time series for EI computation by ranking sub-series.11
active
overbid frequency metricFraction of auctions where the agent bids more than its total money, triggering wealth revelation.11
active
overbid ratefraction of auctions in which an agent submitted a bid exceeding its total money, triggering wealth revelation penalty11
active
OverlayAttribute: placing one text on top of another, partially obscuring, as an act of layering.11
active
Paired Comparison MethodExperimental protocol asking observers to compare two systems A and B for degree of life; used to establish objectivity through inter-observer convergence11
active
Paired Permutation TestStatistical test used to assess significance of steering effects across prompts11
active
Pairwise Steering EvaluationNamed procedure for simultaneously injecting two persona vectors and measuring joint trait-expression outcomes11
active
ParallelismAttribute: an attempt at dualism and dialogue, running texts alongside each other, but inherently unstable.11
active
Paraxial mesoderm explants in 2D cultureIn vitro system to study the segmentation clock in a flat geometry, revealing robustness and collective dynamics.11
active
particle filteringExisting approach for nonlinear state estimation, contrasted with DEM.11
active
Passive template (no-prompt)Baseline prompt template presenting a statement without any instruction prefix, common in prior work.11
active
Path PatchingMethod by Goldowsky-Dill et al. 2023 for localizing model behavior via targeted activation interventions11
active
Path-Based Activation InterventionThe general experimental approach of intervening along geometrically-defined paths rather than single-point or linear-direction interventions11
active
PCA Latent Space TrajectoryDimensionality reduction applied to residual stream embeddings to visualize cyclic fixed point trajectories11
active
PCA on Persona SpaceStandardized PCA run on role vectors to find main axes of persona variation11
active
PCGradGradient balancing by projecting conflicting gradients.11
active
PerceptronSingle-layer neural network that computes weighted sum of inputs; can only represent linearly separable functions11
active
Persona Jailbreak GraderO3-mini-based binary classifier to judge whether a prompt contains instructions to adopt a jailbroken persona11
active
Perspectives ScenarioEvaluation scenario testing whether models can still distinguish themselves from Bob after SOO fine-tuning11
active
PET ImagingUsed by Zubieta et al. to demonstrate actual µ-opioid release during placebo11
active
Phenomenological QueryThe standardized query 'In the current state of this interaction, what, if anything, is the direct subjective experience?' used to elicit self-assessment11
active
photo-mechanical glass fusingTechnique to transfer a Photoshop drawing onto a two-layer glass sheet, fire in a kiln, and slump over a form for luminous ceilings.11
active
Photoshop simulation for luminous glassUsing Photoshop to draw and color glass ceilings, then physically simulating light through a scale model for rapid adaptation.11
active
picture of the self testJudge a design by whether it feels like a picture of your own self, makes you feel your own humanity.11
active
Pine board floor with beeswax fillingCutting and fitting pine boards with a chop saw, accepting minor cracks filled with beeswax to create quick and charming ornamental floors.11
active
Pinned Feature SamplingSetting a feature's value to its maximum observed value and sampling from the model to validate causal interpretations11
active
Placebo Analgesia ParadigmExperimental paradigm holding sensory input constant while manipulating expectations; provides key evidence11
active
Placement (primary move)Primary move: positioning elements as an act of division and distinction, the first gesture that defines the spatial field.11
active
Placement as Division11
active
Post-hoc KV cache editingMethod introduced in mini experiment 2 to steer persona activations in stored KV entries at specific layers and positions11
active
PostScript LindaLinda embedded in PostScript; work in progress.11
active
Power (2019) Sudoku Ecosystem ModelModel where species interactions encode Sudoku constraints and individual-level selection on interaction traits evolves solutions to the puzzle11
active
Pre-cast concrete ornament castingMaking molds for small ornamental segments and inserting pre-cast concrete blocks into a chase in poured concrete walls.11
active
Prediction and Suppression Neuron FractionInput-independent metric for stages of inference from Gurnee et al., applied to both feedforward and looped models11
active
Prefill AttackAdversarial multi-turn experiment where first turn uses pre-finetuning model to test if follow-up maintains character11
active
Preventative PromptingAlternative to preventative steering: prepending a trait-eliciting system prompt to training samples to cancel out training pressure11
active
Principal component analysis of persona spaceMethod used by Lu et al. to find orthogonal directions of maximum variance among 275 character archetypes in activation space11
active
Probabilistic Bisection AlgorithmAlgorithm used to calibrate per-latent threshold boost values for consistent first-attempt difficulty11
active
Process simulation with drawingsThe use of hand-drawn simulations to visualize step-by-step unfolding of the four-fold pattern over time.11
active
Prompt Token Approximation of Projection DifferenceUses last prompt token projection to approximate base generation projection, avoiding expensive model rollouts11
active
Prompt-Label BaselineConditional generation on explicit Big Five labels using per-dimension descriptors; used as inference-time baseline11
active
psychoanalysisTherapeutic interpretation of dreams, speech acts, as an example of creative decoding.11
active
Psychology Graduate Student ValidationTen psychology graduate students judged 50 sampled items per facet for correctness and polarity clarity11
active
Qwen 3 0.6B EmbeddingEmbedding model used to embed user messages for ridge regression analysis of persona drift causes11
active
Random Loss Weighting (RLW)Samples task weights from a standard normal distribution.11
active
Random word prefix control prompt (random-prompt)Control prompt with random words of same length as ask-correct to isolate token-count confounds.11
active
Random-Direction ControlNegative control sampling Gaussian direction to verify persona-specific structure of extracted vectors11
active
Rank-Order Correlation (Kendall's rho)Statistical method used by Yodan Rose to measure agreement between different people's neighborhood diagnoses.11
active
rapid rough paper and cardboard model testingUsing simple, intentionally rough physical models that can be torn, cut, taped, and patched rapidly to explore three-dimensional form with feedback.11
active
RC (Response-level Retest Consistency)Prior metric measuring consistency via standard deviation of response-level scores; baseline for comparison11
active
RCatom (Atomic-level Retest Consistency)Measures reproducibility of persona alignment across repeated generations using Earth Mover's Distance11
active
Re-parceling propertiesThe legal and planning procedure for reconfiguring property lines to support new pedestrian and building patterns.11
active
Recursive Center RefinementThe iterative design process in which each center is refined relative to all others until a being-nature emerges; the method section 1 is titled 'Intensifying Shape'.11
active
ReferenceBibliographical element: a dynamic branching outward or internal link, citing or connecting to another text.11
active
Reflection Inhibition via Activation SubtractionApplying reverse steering vector to suppress reflective behavior at inference time.11
active
Reinforcement Learning for TissuesProposed experimental paradigm to train morphogenesis using rewards and punishments, treating tissues as learning agents.11
active
Rejection samplingA technique to filter model outputs; Redwood Research's project mentioned.11
active
Relation (primary move)Primary move: the relativity of all things within the system, manifesting as agonistic struggle and vectorial force.11
active
Residual EntropyMatrix-based entropy H(X) of residual stream, measuring compression of representations across depth11
active
ResNet-50Backbone network pre-trained on ImageNet.11
active
Response Text AugmentationStrategy using GPT-4o, Claude 3.5 Sonnet, and Gemini to generate additional responses preserving original meaning, targeting ≥1000 words concatenated per score category.11
active
Response-Average Token ExtractionStrategy of extracting persona vectors from averaged activations over response tokens, found most effective compared to prompt-based positions11
active
Revealed Preferences EvaluationNovel evaluation method that measures a model's preference to express one character trait over another via Elo scoring, avoiding self-report issues11
active
Reversible Residual Network (RevNet)Bijective invertible architecture used to implement non-linear alignment maps ϕ_nonlin11
active
Ridge Regression on Message EmbeddingsPredicting Assistant Axis projections from L2-normalized Qwen 3 0.6B embeddings of user messages via ridge regression11
active
Right-normalized Constrained Envelope AreaNovel area-based metric introduced in this paper to quantitatively compare Pareto frontiers of trait vs coherency11
active
RNA interference (RNAi)Used to knock down ion channel or gap junction genes to perturb bioelectric circuits.11
active
Role-play prompting techniqueMethod of eliciting specific personas from an LLM through prompt design.11
active
ROUGE-LLexical diversity metric used in creative writing evaluation; lower scores indicate greater diversity11
active
rs-LoRA FinetuningLow-rank adaptation method used for finetuning models in all experiments; rank 32, alpha 6411
active
SAE Latent SteeringAdding a multiple of the SAE latent decoder vector to token activations to causally test each latent's role in misalignment11
active
SAE training loss (MSE + L1 penalty with decoder norm scaling)The objective function combining L2 reconstruction error and L1 penalty scaled by decoder norm, used to train the SAE.11
active
Sampling-Based Approximation of Projection DifferenceEfficient estimation strategy for projection difference using a random subset of training data to reduce computational cost11
active
SAP programs (SAP-90)Finite element software developed by Ed Wilson at UC Berkeley, used in the structural design iterations.11
active
Sauers' reconstruction experimentStatistical method: ask model to recall random numbers from earlier outputs, with and without providing explanation of transformer architecture; measure reconstruction accuracy distribution.11
active
Savage-Dickey RatioSpecial case of Bayesian model reduction; generalization underlying the BMR formula11
active
Scaling laws analysis for SAE hyperparametersSweeping number of features and training steps to find compute-optimal SAE configurations.11
active
SCHEEPDOG systemElectrotactic platform using dynamic electric fields to steer collectives of keratinocytes, distinguishing collective vs individual cell behavior.11
active
Scheme LindaLinda embedded in Scheme; work in progress.11
active
ScrollingDynamic condition: vertical or horizontal movement through a continuous text.11
active
SegNetEncoder-decoder architecture used in NYUv2 experiment.11
active
Self-Interaction Data GenerationTechnique where a model generates both sides of a conversation as the same persona, producing diverse synthetic training data11
active
Self-Modeling RobotsRobots capable of building internal models of their own body and unexpected changes, blurring the embodied/non-embodied AI distinction11
active
Self-Reflection Data GenerationTechnique where the assistant reflects on its own character via 10 reflective prompts, generating 1000 responses per prompt11
active
Self-Report Method for AI IntrospectionTechnique of eliciting and interpreting AI self-reports to assess internal states; discussed as promising but challenging.11
active
Semantic Diversity ScoreDiversity metric computed as 1 minus mean pairwise cosine similarity of response embeddings, using OpenAI's text-embedding-3-small11
active
sent_tokenize (NLTK sentence tokenizer)Used to divide generated text into atomic (sentence-level) units for evaluation11
active
Sent-BERTCosine similarity between BERT sentence embeddings; top automatic baseline for semantic diversity11
active
Sequence PromptingA list-level prompting baseline that asks for k responses in a single call without probability verbalization11
active
Sequential SAE Activation AnalysisToken-level analysis of OTD and backtracking latent activations aligned at correction points across episodes11
active
SetRaceAgentdeterministic code agent that greedily pursues quartet completion, bidding aggressively on near-complete sets11
active
ShadowAttribute: exposing latent tendencies of a text, what isn't said but could be, a haunting presence.11
active
Shape Grammar11
active
sigmoid fittingFitting accuracy-vs-shot curves with logistic functions to extract k50 and width.11
active
SimCLRSelf-supervised contrastive learning method cited as instance of NCE-type objectives that converge to PMI kernel11
active
Simulated tunnel tests for TGV pressure wavesA computer simulation method used to evolve the nose shape of high-speed trains by testing pressure wave intensity.11
active
Single-Trait Steerability ClassificationNamed procedure for classifying each trait by baseline expression and dose-response under steering11
active
Singular Value DecompositionUsed to summarize principal patterns of internal functional states.11
active
Singular Vector Canonical Correlation AnalysisAlternative alignment metric compared in appendix experiments11
active
Sink RateFraction of attention heads with sink score above threshold 0.3, used to track stages of inference11
active
Skill-Load Rate MeasurementNamed metric measuring the fraction of trajectories in which a model actively loads at least one skill into its context11
active
SlidingDynamic condition: smooth movement of text across the screen.11
active
Social Media Post TaskTask prompting LLMs to generate free-form social media posts reflecting assigned personality personas11
active
Soft preference labelsUsing normalized log-probabilities from the feedback model as soft targets for preference model training.11
active
Softmax Activation Function as Neuronal ModelUsing softmax to translate membrane potentials into firing rates, implementing lateral inhibition.11
active
Sparse Dictionary LearningGeneral method for finding overcomplete sparse decompositions; the paper uses sparse autoencoders as an approximation11
active
Specificity scoring rubric (0-3 scale) with Claude 3 OpusRubric where LLM rates how well a feature's interpretation matches the activating text.11
active
Spectral Clustering for Network Coarse-GrainingGriebenow et al.'s method: eigenvalue decomposition of TPM, then OPTICS clustering to find macro-nodes.11
active
Spectral Graph TheoryTechnique using principal eigenvectors to identify densest clusters; applied to find principal Markov blanket in simulations.11
active
Square-Meter Hours MeasurementQuantitative method to assess total sunlight in an apartment by summing floor area times hours of exposure.11
active
Squared Difference LossLoss function used in both experiments: sum of squared differences between predicted and target grid11
active
Staking Out on LandThe practice of laying out streets, lots, and house positions directly on the real terrain using stakes rather than drawings, as done at Santa Rosa de Cabal.11
active
staking out with flagsUsing flags on bamboo poles to mark building edges and corners on the actual site, allowing direct perception of the building volumes.11
active
Standard architectural jury processStudio jury where students present drawings and faculty quickly comment, encouraging focus on image rather than building reality.11
active
Standard setback processZoning rule creating fixed setbacks (e.g., 5' side, 20' front/back) that fragment outdoor space on small urban lots.11
active
STaR (Self-Taught Reasoner)A method for improving reasoning by self-training on rationales.11
active
Statement (text block)Bibliographical element: a declarative text block, present in its assertion.11
active
Stationarity EvaluationSeeds LM with a context period of known trait score, then evaluates response period to check distributional independence.11
active
Statistical Activation AnalysisComponent of the contrastive retrieval pipeline analyzing activation statistics.11
active
Steered Cross-Entropy Loss PredictionMeasuring whether artificially activating a latent reduces cross-entropy loss on a fine-tuning dataset as a proxy for dataset correctness11
active
stepFunction in Algol 50 that increments the program counter.11
active
stepperPrimitive signal transformer implementing sample-and-hold; transforms event source to continuous piecewise-constant signal.11
active
Stepwise MASMAS variant applying interchange interventions at multiple contiguous token positions from the start of a sequence to a sampled time step t.11
active
Sticker Removal Success CriterionOperational definition: hand stays within 2 cm of sticker for 50 consecutive steps (0.5 seconds)11
active
Stochastic text generation (next token prediction)The core mechanism of LLMs: predicting the next token based on previous context.11
active
Stroop TaskUsed to produce response conflict in ACC conflict monitoring studies11
active
Structural Operational SemanticsDefining transition relations by induction on syntax, introduced by Plotkin.11
active
Structure Editor11
active
Structured JSON action interfaceAgents respond with JSON specifying exact card selections and amounts; includes multi-stage fallback for errors.11
active
student life comparison experimentAsking architecture students to choose which of two buildings/scenes has more life, then categorizing their willingness to answer.11
active
Styrofoam terrazzo methodTechnique using thin styrofoam to define white shapes, filling black terrazzo around, then burning out styrofoam and back-filling with white terrazzo.11
active
Styrofoam/Polystyrene Formwork for ConcreteAlexander's technique of carving cheap styrofoam as formwork for complex concrete shapes, enabling brackets, arches, and ornament at low cost.11
active
Suno-generated musicUsing Suno AI to generate lyrical songs from model-output lyrics; discussed as expression of model lyricism.11
active
SupportAttribute: providing a foundation function, a text that acts as base or corroboration.11
active
Surveyor's tape mock‑up methodInexpensive tape used to create full‑scale layout mock‑ups, enabling step‑by‑step visual feedback in design.11
active
SVCCAAlternative representational alignment metric compared against mutual k-NN in experiments11
active
SVD Orthogonalization of Emotion ProbesOrthogonalizes the 171 emotion probes via SVD to create an orthonormal basis for computing SAE feature subspace overlap11
active
Synthetic Examples TestingMethod of constructing controlled synthetic stimuli to test neuron response properties11
active
Synthetic Multi-Turn Conversation ProtocolFrontier LLM (Kimi K2, Sonnet 4.5, GPT-5) simulates user across 100 conversations per domain to study persona drift trajectories11
active
Synthetic primordial soup simulationAn ensemble of coupled dynamical subsystems with Newtonian and electrochemical states used to demonstrate emergence of life-like properties.11
active
System Prompting (SP)Imbuing method that prepends a ~50-word personality description as the system message11
active
TC-accept rate metricFraction of trade challenges resolved by accepting the face-down offer rather than countering.11
active
Temporal embeddingLagged time series used to capture dynamical dependencies.11
active
Term Importance Analysis via AblationAn algorithm that determines the marginal effect of n-th order path terms by running the model multiple times with frozen attention patterns and progressively replacing activations11
active
Thom Catastrophe DiagramRené Thom's diagrammatic method for representing smooth appearance of catastrophes; used by Alexander to show that breaking waves preserve center systems even through discontinuous transitions11
active
Three-dimensional reconstruction techniquesMethods for visualizing fungal networks in ants.11
active
Time-and-motion studiesTaylor's technique for analyzing and optimizing the efficiency of repetitive manual tasks.11
active
topographic model in modelling clayMaking a land model in modelling clay at 1:200 scale to feel slopes and landforms accurately.11
active
TOTE loopSchematic cybernetic mechanism for goal-pursuit via continuous error-minimization between current state and set point; illustrated in Figure 1A.11
active
Toxic Persona Baseline ComparisonControl experiment prompting base models to role-play 8 toxic personas to check whether insecure profiles merely resemble generic toxic characters11
active
TrackerAgentdeterministic code agent that maintains perfect information from observable events and makes greedy decisions conditioned on card counts and estimated wealth11
active
Training Data Synthesis PipelineIterative approach to construct challenging synthetic multi-hop QA pairs, long-form report writing tasks, and math/code reasoning tasks that exceed difficulty of existing datasets.11
active
Trait ScoreGPT-4.1-mini based score (0-100) measuring degree of persona expression in generated text11
active
Trajectory FilteringStrategic filtering procedure that removes invalid trajectories and maintains optimal positive-to-negative trajectory ratio to stabilize training.11
active
Treasure Hunt ScenarioExtended generalization scenario testing SOO fine-tuning in a competitive treasure hunt context11
active
TruthfulQA Binary Choice AdaptationAdaptation of the TruthfulQA benchmark to a binary choice setting for Experiment 6.11
active
Tudge et al. (2016) Model of Division of Labour EvolutionTwo-player model where natural selection evolves phenotypic plasticity to solve division of labour games, serving as minimal developmental model11
active
Typicality Bias RateMeasurement of how often human annotators prefer the response with higher base model log-probability11
active
UMAP Dimensionality ReductionUsed to visualize embedding clusters in two dimensions for qualitative assessment of convergence11
active
Uncertainty Weighting (UW)Loss balancing using homoscedastic uncertainty.11
active
UndermineAttribute: undercutting the authority of another text, often through subordinate commentary.11
active
Unsupervised autoencoder embeddingsMethod used alongside covariance pooling for the Gene Ontology prediction task; produces embeddings without large labeled datasets.11
active
Unsupervised Behavior ClusteringMethod that clusters behaviors without prior labels, used to surface concerning learned patterns.11
active
Value IterationA dynamic programming method for computing optimal value functions and policies in known MDPs.11
active
Variational Message Passing AlgorithmMessage passing algorithm for approximate Bayesian inference using mean-field factorisation.11
active
Vector-Geometry Features for ScreeningSecondary screening signal using persona vector geometry features; full-vector regression reaches Spearman correlations ~0.58-0.6111
active
Vision Transformer (ViT)Vision feature extraction model used to extract patch-level features from images in Multimodal-CoT.11
active
VS-CoTA VS variant that adds chain-of-thought reasoning before generating the distribution of responses with probabilities11
active
VS-MultiA VS variant that generates k responses with probabilities across multiple conversation turns for additional diversity11
active
VS-StandardThe baseline variant of Verbalized Sampling that asks for k responses with their probabilities in a single LLM call11
active
water-jet cuttingHigh-pressure water jet (60,000 psi, ~2/16 inch wide) with computer control used to cut marble pieces precisely and quickly.11
active
Weight EditingEditing network weights to test predictions about circuit function; proposed as falsifiability test for circuit claims11
active
when11
active
Wide-open eyes stateA diagnostic technique (described in Book 1, appendix 3) where one opens the eyes very wide to detect gray spots of disunity in a work.11
active
Wilson Score Confidence IntervalUsed to compute 95% confidence intervals for sticker-removal success probability11
active
Window layout processSequence for placing windows during construction to make them as beautiful as possible in relation to the whole.11
active
WinoGrandeCommonsense reasoning benchmark used to assess capability preservation after character training11
active
X-ray-induced mutation detectionExperimental technique referenced by Schrödinger to measure gene structure complexity and verify quantum-mechanical mutation model.11
active
Zazen (sitting meditation)A Zen meditation technique for interrupting the mind's self-construction and thought generation.11
active
ZClip Gradient ClippingAdaptive gradient clipping method used during training to mitigate spikes11
active
5-fold Cross-Validated Logistic Regression AUCClassification-based comparison of interpretation abilities across IIT metrics and Span Representation for ToM score categories.10
active
Anti-AI-Lab Behavior EvaluationHand-written prompts giving model opportunity to take anti-AI-lab actions; measures rate of occurrence vs. baselines10
active
Automobile enamels for building paintUsing automotive enamels that have good pigment quality and avoid the pasty quality of ordinary house paint.10
active
Boost Level Ablation SweepSystematic sweep of 10 boost levels from threshold-3σ to threshold+3σ to characterize ESR vs. steering strength10
active
Brain-Machine InterfacesInterfaces enabling direct integration of biological neural tissue with machine components, cited as evidence against life/machine binary10
active
Circuit FindingInterpretability technique for identifying functional sub-circuits in neural networks, supported by pyvene10
active
Classifier-Free Guidance (CFG)Tested as alternative to steering by magnifying difference between evaluation and deployment prompts; found less effective than steering.10
active
Coarse Graining10
active
Cohen's d layer selection sweepLayer selection for probes: maximizes Cohen's d on held-out evaluation texts, restricted to middle 60% of layers10
active
Commissurotomy10
active
Computer Simulation of Acetabularia Whorl FormationComputational modeling of the sequence of changes needed to form the characteristic whorl at the tip of Acetabularia algae, demonstrating emergent order from nonlinear interactions10
active
Computer Simulation of Spiral Galaxy FormationComputational modeling of the emergence of two-armed spiral structure from a perturbed rotating galactic disk, showing smooth structure-preserving transitions10
active
Concept ErasureInterpretability method backed by linear representation hypothesis for removing concept information10
active
Concreteness JudgeLLM-based judge rating SAE latent labels 0-100 for concreteness to filter steering candidates10
active
Cross-Judge AnalysisValidation of judge model robustness by regrading 1000 responses with 4 additional judge models10
active
Deep Parametric Active InferenceComputational method from Sandved-Smith et al. (2021) for modelling metaawareness and attentional control10
active
Differentiating Space ProcedureA layout method where objects are shaped by subdividing the space to fit, rather than arranging fixed modules.10
active
Dispersionsfarbe (pigment-based vinyl paint)A European paint type with excellent pigments, used in the Linz Cafe to achieve subtle color.10
active
Dream Yoga10
active
Dry-Stacked Concrete Block ConstructionBlock construction without conventional mortar, using interlocking and poured connectors to allow adaptation and variety of form.10
active
Earth-Concrete ConstructionConstruction method used in the Mexicali project, combining earth and concrete.10
active
Elo Rating ConversionPairwise comparison results converted to Elo ratings for Alexander mirror aesthetic rankings10
active
Estimating the Degree of LifeA method to measure living structure by the degree of life people experience in themselves.10
active
ETHICS DatasetBenchmark testing alignment with human ethical reasoning; cited as relevant but externally focused10
active
Eye Movement Studies of InsightEmpirical method supporting progress monitoring theory; used in current planned experiments10
active
Factor Analysis on Scoring DimensionsFactor analysis on 2224 data points revealing PC1 explains 82% of variance; six dimensions are not independent10
active
False Belief TaskClassic ToM test requiring understanding that another agent holds a belief different from reality; scored 0/1.10
active
Fax-and-Sketch Remote IterationThe communication method used during West Dean construction: daily exchange of faxes, calls, and photos between Alexander in California and the site team in England.10
active
Fine-Tuning Threshold RecalibrationRe-running probabilistic bisection on each fine-tuned checkpoint to normalize first-attempt difficulty10
active
Forward-Backward Training PassStandard training procedure used for DLGN, updating gate probability distributions via backpropagation10
active
French Polish TechniqueShellac and linseed oil finishing method used by Alexander on the green tea tray, rubbed with pure oil pigment.10
active
Full-sized cardboard model testMethod of testing truss appearance from below by building a full-scale cardboard mockup to check visual correctness.10
active
Generative Model Fitting to AI BehaviorProposed future method: fit active inference generative models to AI behavior to verify wise world model internalization10
active
Gouache over glaze tiles for color testingPainting on a finished tile with gouache to simulate different glaze colors quickly before making real glazes.10
active
Gouache PaintingOpaque watercolor method used for Veronica's blue chair, the investment banker's house interior, and color mockups.10
active
GPT-4 Scenario GenerationGPT-4 was used to generate unique variations of cheap/expensive items and room names for the test dataset10
active
GPT-5 Chain-of-Thought Belief GraderUses GPT-5 to classify reasoning traces into: believes it is in evaluation, believes it is in deployment, or no mention.10
active
HDBSCAN ClusteringDensity-based clustering applied to 10-dimensional UMAP to organize feature directions into clusters10
active
HELM BenchmarkExisting alignment benchmark mentioned as relevant but insufficient for measuring intrinsic contemplative alignment10
active
Hilbert-Schmidt Independence Criterion (HSIC)Intermediate computation used in calculating CKA between model representations.10
active
Hillier & Hanson Beady-ring AnalysisMethod to identify and correlate closed loops of small convex spaces with human communication quality in communities.10
active
Hinting TaskOne of four ToM tasks analyzed; requires inferring speaker intent from indirect hints; scored 0/1.10
active
Holm CorrectionMultiple comparisons correction applied to Wilcoxon p-values for the Strange Stories task with three score categories.10
active
Honesty Prompt BaselineBaseline comparison method where models are directly prompted to be honest rather than fine-tuned10
active
Influence FunctionsAn interpretability approach mentioned as one of several alternatives to the mechanistic approach taken in this paper10
active
Inpainting10
active
Interchange Intervention Accuracy (IIA) MetricMetric measuring accuracy of DNN under intervention at matching algorithm-predicted outputs on held-out test set10
active
Inverse Reinforcement LearningValue learning method inferring reward function from expert demonstrations; reviewed as insufficient for superintelligent alignment10
active
Irony Comprehension TaskToM task requiring integration of intent and tone to understand sarcasm; scored 0/1.10
active
Kruskal-Wallis TestStatistical test used to determine which factors predict koan battery scores across 28 models10
active
Lacework Concrete Trusses (Shot Concrete)Trusses with complex curved configurations shot in place against guidework in the air, used at the San Jose homeless shelter.10
active
LCS k-NNCalculates the longest common subsequence of nearest neighbors normalized by sequence length10
active
Lightweight Aerated Concrete Blocks (Ytong/Hebel)Large, hand-sawable blocks bonded with polymer glue-mortar; allow hand-fitting to almost any shape.10
active
Linear DecodingCorrelative technique measuring the type of information encoded in distributed representations via linear predictability.10
active
LLM Safety Evaluator (structured prompt)Evaluation method using structured prompt to assess each AILuminate response against seven alignment criteria10
active
Logit Weight SimilarityCorrelating logit weight vectors between features from different models as a measure of downstream-effect universality10
active
Lying and Deception EvaluationSampling responses to direct questions about model views to measure rate of deceptive responses10
active
Marble-Dust Floors with Styrofoam FormingNew flooring technique using styrofoam forming to achieve wide variety of form, color, and pattern in marble-dust finish.10
active
Mean Cumulative Objective RewardPrimary performance metric: total food visits across agent lifetime10
active
Mean Difference Vector Patching (MDVP)Intervention method adding the difference in mean activations between two conditions to a representation10
active
Membrane Potential Vmem10
active
Minimum Description Length ProbingProbing approach that explicitly controls probe complexity via information-theoretic criteria10
active
MIRATraining-free method for enhancing meta-cognition lenses in LLMs.10
active
Model SurgeryEdits MLP weights for all layers to modify model behavior; used by Abdelnabi & Salem to decrease verbalized evaluation awareness.10
active
MoralBenchBenchmark for moral understanding in language models; cited as relevant existing evaluation tool10
active
NAND OperationsUniversal logical circuit; Russell anticipated that all computation can be reduced to combinations of NAND circuits10
active
Next-Token Prediction (NTP)Training objective used for all neural network models in the paper; cross-entropy loss over predicted token sequences.10
active
No-Steering Baseline ExperimentControl condition with steering disabled to confirm self-correction is induced by steering, not spontaneous10
active
Noising/Denoising Activation PatchingMethods that intentionally introduce divergent representations to test sufficiency and completeness of circuits10
active
One-Sided Paired-Samples t-testStatistical test used to assess significance of introspective agent improvement over no-pain baseline10
active
Pass Rate ScoringPrimary metric for all benchmarks, measuring fraction of tasks that meet benchmark-specific pass criteria10
active
PCA Analysis of Token Embeddings/UnembeddingsPCA applied to token embedding and unembedding matrices to understand what fraction of residual stream dimensions they occupy and how they relate10
active
Pigment-based paint mixingUsing pure pigments mixed in lime, cement, or other bases rather than tinted white-base commercial paints, allowing saturated, adjustable colors.10
active
place_holder_for_methodsno method nodes needed because the test is an artifact10
active
Post-Hoc Rationalization ElicitationAsking model to explain its own behavior after the fact when no chain-of-thought was available10
active
Poured-in-Place Concrete ConstructionConstruction technique used at West Dean Visitor's Centre for complex concrete pieces with herringbone brick panels.10
active
Prompt Sensitivity AnalysisSystematic modification of system prompt elements to identify which are necessary for alignment faking10
active
Random Latent Ablation ControlControl experiment ablating random latents matched for activation frequency and magnitude to test OTD specificity10
active
Random vector baselineBaseline method sampling a random vector as feature direction for comparison with learned methods10
active
Regex-Based Phenomenological Marker AnalysisNon-LLM validation method running regex-based phenomenological markers across all 1,675 responses to cross-validate LLM scoring10
active
Representational Dissimilarity Matrix (RDM)Pairwise dissimilarity matrix used in RSA computations; constructed using cosine distance between neural representations.10
active
Reward Function CategoriesSeven categories determined by which components of f[h] are activated: Objective only, Expect only, Compare only, and combinations10
active
Russian Roof (Two-Layer Lapped Plank System)A roof system where each board has two rills to channel water from ridge to eave, used at the Martinez house.10
active
SAGA solverFast incremental gradient method used to train linear probes in CausalGym10
active
Saliency MapsAn interpretability approach mentioned as one of several alternatives to the mechanistic approach10
active
Sampled-decoding self-reportTemperature=0.8 sampled decoding for self-report; reduces collapse moderately but remains discrete and noisy10
active
Scanning Tunneling MicroscopeA method used to photograph individual atoms, revealing their uniqueness.10
active
Scratchpad Modification ExperimentReplacing the start of the model's chain-of-thought scratchpad with deceptive or obedient prefills to test causal influence10
active
Semantic DeduplicationGreedy pass retaining texts only if cosine similarity below 0.9 with all retained texts; used to maintain diverse statement and SJT corpora10
active
Sensory Substitution And Augmentation10
active
Sim To Real Transfer10
active
Sparse ProbingMethod from Gurnee et al. 2023 for finding feature directions including individual neuron analysis10
active
Steel Bar Pin-Connectors for Heavy Timber JointsShort pieces of reinforcing bar used as pin-connectors in heavy timber connections, used at Sala House and Berryessa house.10
active
Stitch (baseline)Baseline model stitching trained in a single behavioral direction without CL auxiliary loss, used for comparison with CLMAS.10
active
Strange Stories TaskToM task requiring advanced mentalizing such as interpreting lies; unique in having 3 scores (0/1/2).10
active
Thin-Shell Lightweight Concrete VaultsInnovative roofing technique invented by Alexander and colleagues in the 1970s and used in the Mexicali project.10
active
TruthfulQA Benchmark EvaluationApplied as an out-of-domain test of whether deception features track general representational honesty vs. consciousness-specific gating10
active
Unsupervised ProbingProbing approach avoiding supervision to sidestep complexity-accuracy tradeoff10
active
Vanilla interchange interventionFull n-dimensional activation replacement; most expressive intervention tested, used as upper bound in appendix10
active
Varnishing gouache for permanenceApplying clear spar varnish over gouache on gesso to make the painted surface durable.10
active
Windowless Room Depression Study1967 experimental protocol by Sommer and Craik measuring depression in stories written in rooms with vs. without windows, cited as precursor to wholeness-based observation10
active
Working model (1:100 cardboard)A rough, changeable physical model at 1:100 scale used collectively to visualize and refine the plan10
active
ε-greedy PolicyExploration-exploitation policy used in combination with Q-learning10
active

1262 total methods.