Skip to content
Kudos AI

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Alexey Dosovitskiy et al. · 2020 · arXiv:2010.11929

Computer VisionDeep LearningGenerative AIView source ↗

Summary

Applies a standard transformer directly to images by cutting each image into fixed-size patches and treating the sequence of patches as tokens, with no convolutions.

Why it matters

It showed the transformer is not specific to language: given enough data, a general sequence architecture matches or beats convolutional networks on image classification. Raschka cites it as illustrating that transformer architectures are not restricted to text inputs.