Lecture 5 · Interactive explanation
See the cost of a steep line
The dots are invented training observations. Least squares chooses the line with the lowest training sum of squared errors. Ridge also charges for a large slope.
Navy dots: observations. Grey dashed: least-squares line. Teal: ridge line. Orange vertical segments: ridge residuals.
Increasing λ shrinks the slope and can increase training error. This fixed-data demonstration does not measure generalization. Choose λ with validation, then evaluate separately.
Why the residuals matter
A residual is actual minus predicted. Equal positive and negative errors can cancel if simply averaged; MAE uses absolute values and MSE squares them. Here the objective is sum of squared residuals + λ × slope². The intercept is unpenalized and recalculated as the slope changes.
And LASSO?
LASSO uses the absolute coefficient rather than its square. It can make a coefficient exactly zero, removing that feature. This demonstration specifically fits ridge; it does not display a LASSO fit.