concept
active
concept:vision-transformer-vit-modelsVision Transformer (ViT) Models
Primary vision model architecture used in both vision-vision and cross-modal alignment experiments
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Vision feature extraction model used to extract patch-level features from images in Multimodal-CoT.
- The transformer's model of itself as a predictive text engine, developed through in-context learning.
- Empirical result showing alignment increases with model competence
- Secondary substrate demonstrating cross-modal applicability of manifold steering.
- Key empirical finding establishing that representational alignment correlates with model competence
- A model that frames RL as sequence modeling, SOTA from random trajectories.
- The transformer version directly analogous to TEM, introduced in this paper, offering dramatic performance improvements.
- Model introduced in Figure 2 explaining how collective intelligence expands the spatiotemporal perceptual field of a group beyond any individual member's capacity.