method
active
method:vision-transformer-vit

Vision Transformer (ViT)

Vision feature extraction model used to extract patch-level features from images in Multimodal-CoT.

Neighborhood — ranked by edge-count

Frameworks (1)

framework
  • Multimodal-CoT
    implements
    A two-stage framework that separates rationale generation and answer inference by incorporating vision and language modalities.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Primary vision model architecture used in both vision-vision and cross-modal alignment experiments
  • A model that frames RL as sequence modeling, SOTA from random trajectories.
  • visual supervisionconcept0.698
    Supervisory signals for visual outputs; functional tokens do not require it.
  • visual reasoningconcept0.695
    Visual reasoning tasks often interleaved with intermediate visual states; promising direction in the field.
  • Technique of building a fluid, three-dimensional vision by closing one's eyes, relying on words and feeling to avoid arbitrary graphical over-specification.
  • A family of large language models trained on next-token prediction, central example of simulators.
  • The transformer version directly analogous to TEM, introduced in this paper, offering dramatic performance improvements.
  • A transformation that cleans out undifferentiated areas, creating a homogeneous zone bounded by differentiated structure, preserving distinctness.