Sections¶
- 1.1. What is Machine Learning?
- 1.2. Squared Loss and the Constant Model
- 1.3. Absolute Loss
- 1.4. Comparing Loss Functions
Summary¶
1.1 What is Machine Learning?¶
Supervised learning: learn from labeled data to predict from features .
Classification predicts a discrete category (e.g. true or false; cat, dog, or hamster; digit from 0 to 9).
regression predicts a real number (e.g. a commute time, a height, a weight).
Unsupervised learning: finding structure in unlabeled data (e.g., clustering or dimensionality reduction).
Reinforcement learning: training an agent to make decisions from rewards.
Overfitting: a model that is too flexible may learn the quirks and noise of its training data and make worse predictions on new data than a simpler model.
1.2 Squared Loss and the Constant Model¶
Modeling recipe (empirical risk minimization): choose a model, choose a loss, minimize average loss.
Constant model: . Looks like a horizontal line; returns the same value for all .
is the loss for one point; is the average loss (empirical risk).
Squared loss in general:
Mean squared error (average squared loss) for the constant model:
To find the that minimizes the mean squared error, we take the derivative and set it to 0:
1.3 Absolute Loss¶
Absolute loss in general:
Mean absolute error (average absolute loss) for the constant model:
is a piecewise linear function, with a bend at each data point, and is not differentiable. The slope of at any that is not a data point is:
The median minimizes the mean absolute error (uniquely if is odd; for even , so does any between the middle two).
1.4 Comparing Loss Functions¶
| Loss | Always unique? | Robust to outliers? | ||
|---|---|---|---|---|
| squared | mean | yes | no | |
| absolute | median | no | yes | |
| as | midrange | yes | no | |
| 0-1 | 0 if , else 1 | mode | no | no |
The mean is pulled toward outliers and the tail: right-skewed mean > median.
Minimum risk measures spread:
is the variance ( = its square root).