Datasets
A referral index of the datasets our projects and articles are built on. We host a few tiny synthetic sets for instant reproducibility and link out to the real, canonical public datasets rather than re-hosting them.
MNIST handwritten digits
70,000 labelled images of handwritten digits - the standard first test that a neural network implementation works at all.
IMDB movie reviews
50,000 movie reviews labelled positive or negative - a binary text-classification benchmark balanced by construction.
Auto fuel economy
Fuel consumption and engine characteristics for a few hundred cars - the worked example behind simple and polynomial regression.
Credit card default (simulated)
A simulated set of cardholders used to introduce classification - and a clean illustration of why accuracy misleads on rare events.
S&P 500 daily movements
Daily percentage changes in the S&P 500 from 2001 to 2005 - a deliberately hard classification problem where the honest answer is "barely better than chance".