method
active
method:vision-transformer-vitVision Transformer (ViT)
Vision feature extraction model used to extract patch-level features from images in Multimodal-CoT.
Neighborhood — ranked by edge-count
Frameworks (1)
framework
- Multimodal-CoTimplementsA two-stage framework that separates rationale generation and answer inference by incorporating vision and language modalities.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Primary vision model architecture used in both vision-vision and cross-modal alignment experiments
- A model that frames RL as sequence modeling, SOTA from random trajectories.
- Supervisory signals for visual outputs; functional tokens do not require it.
- Visual reasoning tasks often interleaved with intermediate visual states; promising direction in the field.
- Technique of building a fluid, three-dimensional vision by closing one's eyes, relying on words and feeling to avoid arbitrary graphical over-specification.
- A family of large language models trained on next-token prediction, central example of simulators.
- The transformer version directly analogous to TEM, introduced in this paper, offering dramatic performance improvements.
- A transformation that cleans out undifferentiated areas, creating a homogeneous zone bounded by differentiated structure, preserving distinctness.