Skip to content
Kudos AI

Neural Machine Translation of Rare Words with Subword Units

Rico Sennrich, Barry Haddow, Alexandra Birch · 2015 · arXiv:1508.07909

Natural Language ProcessingDeep LearningView source ↗

Summary

Adapts byte pair encoding to text segmentation, building a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair, so that rare words decompose into known fragments.

Why it matters

It removed the out-of-vocabulary problem from neural sequence models. Raschka identifies this as the paper describing the byte pair encoding used for tokenization, and the scheme it introduced is what GPT-family models still use to segment their input.