Credit card default (simulated)
A simulated set of cardholders used to introduce classification - and a clean illustration of why accuracy misleads on rare events.
Whether an individual defaulted on a credit card payment, given annual income and monthly balance. An Introduction to Statistical Learning states plainly that this data is simulated, and uses it to introduce logistic regression. It is a good teaching set precisely because defaults are rare: a classifier that predicts "no default" for everyone scores high accuracy while being useless, which is the clearest possible argument for looking at a confusion matrix instead.
Key columns
Representative fields, refer to the source for the full, authoritative schema.
| Column | Type | Description |
|---|---|---|
| default | string | Whether the customer defaulted - "Yes" or "No". Heavily imbalanced toward "No". |
| student | string | Whether the customer is a student. |
| balance | number | Average monthly credit card balance. |
| income | number | Annual income. |
License: Distributed with the book’s companion R package; see the source page for terms.