framework
active
framework:universal-transformersUniversal Transformers
Early architecture reusing same transformer block for multiple iterations, precursor to looped LLMs
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Two-layer transformer with rotary positional encodings used in numeric task experiments.
- Neural network architecture based on attention, commonly used in large language models
- The indivisible oneness, meltedness that is the source of life; it cannot be described as a structure because it is pure one.
- A transformer with no attention layers; shown to model bigram statistics via T = W_U W_E
- Core abstraction in Fruit: pure function mapping signals to signals; enables compositional GUI definitions.
- A model that frames RL as sequence modeling, SOTA from random trajectories.
- Spontaneously reported experience of multiple simultaneous processing streams, observed even in base models.
- The transformer version directly analogous to TEM, introduced in this paper, offering dramatic performance improvements.