Learning Internal Representations by Error Propagation
David E. Rumelhart, Geoffrey E. Hinton, Ronald J. Williams · 1986 · Parallel Distributed Processing, Vol. 1, chap. 8, pp. 318–362, MIT Press
Summary
Presents backpropagation as a general method for training multilayer networks, showing that hidden layers can learn useful internal representations rather than needing to be designed by hand.
Why it matters
It made multilayer networks trainable in practice and set off the connectionist revival of the late 1980s. Russell and Norvig record that this anthology, together with a companion article in Nature, drew enormous attention to the field. The attribution is layered: equivalent techniques had been derived independently before, so 1986 is better read as when backpropagation became widely known than as when it was invented.