Attention Is All You Need
Ashish Vaswani et al. · 2017 · arXiv:1706.03762
Summary
Introduces the transformer, an architecture built entirely from attention and feed-forward layers with no recurrence, originally developed for machine translation.
Why it matters
Removing recurrence let every position in a sequence be processed in parallel during training and put any two positions one operation apart, which together made training at previously impractical scale feasible. Raschka records that this paper proposed the original transformer architecture; essentially every modern large language model descends from it.