concept
archived
concept:layer-18Layer 18
Specific transformer layer housing the addition module.
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Layers with weak anchoring due to generic representations.
- Layer normalisation used in transformer and in TEM-t position encoding preprocessing.
- Task-specific peak anchoring score for structured reasoning domains.
- Layers where anchoring weakens systematically due to representational drift.
- A formulation for multi-layer steering where each layer's contribution is the incremental difference in persona direction representation
- Procedure of systematically varying the layer at which activations are recorded and injected.
- Q8B most often peaks at layer 20, followed by layer 25; G20B most often peaks at layer 15finding0.684Best steering layer is model-specific; linearly accessible trait information is organized differently across models
- Network with hidden layers capable of representing non-linearly separable functions, enabling deep model induction