Home: all lectures
0 XP0 day streak

Lecture 5 · Study guide

Regression

3 steps25–35 min block2 practice questions

1Step 1 of 3≈ 5 min

Overview

First read: about 4–6 minutes. Lecture 5: Regression.

Regression predicts a numerical target from predictors. This lecture links that task to feature construction, error metrics and regularization—the choices that help a model work beyond the examples it was trained on.

The ideas to keep

Blur the explanations and test yourself.

Learning types. Regression/classification use labelled targets; clustering finds groups without them; reinforcement learning uses rewards. Deep learning is a type of ML, while data science also includes acquisition and communication.

Useful features. Aggregation summarizes a variable over a window; binning creates categories; encoding represents category membership; combining variables creates a new quantity.

Error. MAE averages absolute residuals; MSE averages squared residuals and emphasizes large misses. Compare on observations that were not used to fit the model.

Overfitting. Zero training error can coexist with poor new-data predictions. A small or restricted sample makes estimates less reliable.

Regularization. Ridge adds a squared-coefficient penalty and shrinks coefficients. LASSO adds an absolute-coefficient penalty and can set some to zero. λ controls strength.

Validation and baseline. Tune using validation/CV and reserve final test data. Compare with a simple predictor-free benchmark before calling a model useful.

What to be able to do

Be able to identify target/predictors, explain feature operations, calculate MAE/MSE and justify ridge/LASSO and evaluation in a small-data scenario. The example exams indicate question styles, not an exhaustive syllabus.

Next step

Read the detailed notes where you need an explanation, then attempt the two original questions before opening the answers. The full slides are on Canvas.

Interactive explanation — try it right in your browser.

Got the big picture?Mark the overview done to fill this lecture's ring.


2Step 2 of 312–18 min

Detailed notes

Source: full presentation. Canvas lecture page checked 11 October 2026, 11:40 CEST. Examples labelled “illustration” are invented to explain the slide concepts.

Overview · Two practice questions

Jump to a section · 12
  1. 1. Where regression fits in machine learning
  2. 2. Feature engineering: make useful inputs
  3. 3. Judge errors with the right metric
  4. 4. Training performance is not enough
  5. 5. Least squares, ridge and LASSO
  6. Extra points made explicit by the example exam
  7. Walk through the four feature operations
  8. Regression line: what intercept and slope say
  9. Broader methods introduced in the first half
  10. Try the interactive explanation
  11. From the recording
  12. Slide coverage map

1. Where regression fits in machine learning

Artificial intelligence (AI) is the broad idea of machines performing intelligent tasks. Machine learning (ML) is part of AI: models learn relationships from data. Deep learning is part of ML and uses neural networks with multiple layers. Data science overlaps with AI and ML, but also includes obtaining, preparing, exploring and communicating data.

Learning approach What the model learns from Example
Supervised regression Inputs paired with numerical answers Predict running time
Supervised classification Inputs paired with category labels Predict injury: yes/no
Unsupervised clustering Inputs without target labels Find groups of similar athletes
Reinforcement learning Feedback or rewards from actions Learn actions that help a robot navigate

A feature/predictor is an input column. The target is the answer you want to predict. Regression and classification both have targets; clustering does not have supplied “correct” group labels.

The lecture also introduces ensembles and neural networks. An ensemble combines models, such as several trees. A neural network learns a sequence of transformations from inputs to outputs. You mainly need to recognize these approaches here; ridge and LASSO are the detailed regression examples.

2. Feature engineering: make useful inputs

A model can only use the information represented in its features. Better features can make a simpler model sufficient: faster, easier to interpret and easier to maintain.

Operation What you do Example
Aggregate Combine observations using a function and time window Mean training load over seven days
Bin Convert values into broader categories Low, medium and high load
Encode categories Represent categories as numerical indicator columns A 0/1 column for each medal category
Combine variables Create a feature from several existing variables Body surface area from height and weight

One-hot encoding creates one indicator per category. Dummy encoding usually creates one fewer indicator, treating the omitted category as a reference. With an intercept in ordinary linear regression, including all category indicators creates a redundant representation: the indicators sum to 1, just like the intercept column. This is the dummy variable trap.

Example: with categories bronze/silver/gold, use bronze and silver indicators and let gold be the reference. A gold observation has 0 in both indicator columns.

3. Judge errors with the right metric

For observation i, the residual is y_i-ŷ_i: actual minus predicted value.

  • MAE: average absolute residual, (1)/(n)Σ |y_i-ŷ_i|. It reports the average size of the error in the target’s units.
  • MSE: average squared residual, (1)/(n)Σ (y_i-ŷ_i)^2. Squaring makes large errors count disproportionately. Its units are squared target units.

Both treat an overprediction and an equally large underprediction the same. MAE is less sensitive to extreme errors than MSE; it is not completely unaffected by them.

Worked example: residuals of 1, −1 and 4 give MAE = (1 + 1 + 4)/3 = 2, and MSE = (1 + 1 + 16)/3 = 6. One large miss dominates the squared-error score.

4. Training performance is not enough

In the holdout method, split the dataset into training and test data. Fit the model on training data, then predict the held-out observations and compare predictions with their actual targets.

Overfitting means learning a pattern too closely tied to the training observations, including noise. The lecture shows a steep line fitted perfectly through a tiny training sample, yet performing poorly on new points. Zero training error does not demonstrate good generalization.

5. Least squares, ridge and LASSO

Ordinary least squares chooses coefficients to minimize the sum of squared residuals. Regularization adds a cost for large coefficients, trading some training fit for a model that may generalize better.

Model Objective in the lecture’s one-slope example Effect
Least squares Squared residuals Best training fit under this criterion
Ridge Squared residuals + λ × slope² Shrinks coefficients toward zero
LASSO Squared residuals + λ × abs(slope) Can set coefficients exactly to zero

For multiple predictors, the penalty sums over their coefficients. λ (lambda) controls penalty strength: λ = 0 gives least squares; increasing λ strengthens shrinkage. Choose λ using cross-validation. Keep a final test set separate from tuning.

The slide calculation, λ = 1:

  • Steep least-squares line: 0 + 1.3² = 1.69.
  • Ridge candidate: 0.3² + 0.1² + 0.8² = 0.74.
  • LASSO objective for that candidate: 0.3² + 0.1² + |0.8| = 0.90.

The ridge candidate has a lower ridge objective despite having training residuals. This illustrates the effect of a penalty; it does not itself prove better test performance.

LASSO can remove unhelpful predictors by setting their coefficients to zero. Ridge usually keeps predictors while shrinking them. The lecture’s rule of thumb is LASSO when many predictors are unnecessary and ridge when most contribute. Treat that as a starting point, then compare validation performance.

Extra points made explicit by the example exam

Percentile versus minimum (the example exam): a minimum depends on one observation and may capture a sensor artefact. A less extreme percentile can reduce that dependence. Percentiles are not universally superior; the feature should represent the quantity you want.

Naive baseline (the example exam): compare predictions with a deliberately simple benchmark, such as always predicting the training-set mean. Beating it suggests added predictive value from the model/features. The sample describes a predictor-free baseline; “baseline” does not mean a perfect model.

Restricted range and poor fit (the example exam): a sample containing only very similar elite performers may have little variation in the target or a predictor. That makes a relationship harder to detect; low fit is not automatically proof that the predictor is irrelevant in a wider population. A small sample adds uncertainty.

Deep learning (the example exam): the sample emphasizes that deep models can learn intermediate features automatically, whereas conventional workflows often use manually engineered features. This is a typical contrast, not an absolute rule for every algorithm.

Walk through the four feature operations

Think of a feature as one piece of information placed in a column for a model. The four operations change how that information is represented.

Aggregate = summarize several observations into one value. Illustration: daily loads are 10, 20, 0, 30, 10, 0, 0. Their seven-day mean is 10. The variable is training load, the window is seven days, the function is mean. A sum would produce 70; a maximum would produce 30. These features answer different questions. A moving window is recalculated at each prediction date using only the relevant preceding observations.

Bin = put values into broader groups. If load 0–15 is called low, above 15 up to 30 medium, and above 30 high, loads 10 and 14 both become low. You gain a simple category but lose the difference between 10 and 14. Thresholds should be clear, with no overlap or gaps. Categorical binning also exists: Spain and Italy can be grouped as Europe, while Chile and Brazil can be grouped as South America.

Encode = represent a category in a form the model can use. A three-category discipline variable can become separate road/pursuit/sprint indicator columns. Each row has 1 in its category and 0 in the others. This represents membership; it does not say that sprint is three times road. In regression with an intercept, use a reference category to avoid redundant columns.

Athlete Road indicator Pursuit indicator Sprint indicator
A, road 1 0 0
B, pursuit 0 1 0
C, sprint 0 0 1

Combine = calculate something from multiple variables. BMI combines mass with height: mass in kg / height in m squared. For mass 72 kg and height 1.8 m, BMI is 72/3.24 ≈ 22.2. The combining-variables slide illustrates this with BMI categories (and an image of airflow around a cyclist); in the recording the lecturer's example is BMI from height and weight. Body surface area is a further example of a height–mass feature, but it is not shown in the slides. The formula determines the meaning; simply adding measurements with different units usually does not.

Four feature operations, with inputs and outputs

Regression line: what intercept and slope say

In predicted size = 0.9 + 0.75 × weight, the intercept is the predicted size when weight is zero and the slope is a 0.75-unit increase in predicted size for a one-unit increase in weight. The intercept may lie outside the meaningful data range, so do not necessarily give it a literal biological interpretation.

Least squares minimizes vertical squared residuals in the target. That is not the same geometric operation as PCA, which seeks directions of greatest variance. A very steep line fitted through only two training points can be unstable: small differences in the sample can change its predictions substantially.

Regularization deliberately accepts more training error if the penalized objective improves. In the lecture's small example, the training data—not the test data—are what the steep line overfits. One slide reverses that wording. Keep the final test set unseen during model and λ selection.

Broader methods introduced in the first half

Reinforcement learning learns actions using reward feedback, such as navigation or games. It is not the same task as predicting a labelled race time.

An ensemble combines models. Random forests combine varied trees; boosting builds successive models that improve the ensemble. These approaches are not inherently perfect or automatically better for every dataset.

A neural network learns weighted transformations. Deep learning uses multiple layers to learn intermediate representations, which can be useful for images, speech and other complex inputs. Humans still choose data, targets, architecture and evaluation. The slide's “without human intervention” description concerns learning parameters from data, not the entire scientific workflow.

AI, ML and deep learning have a nested relationship, while data science overlaps with them and also includes work that uses no AI. A neural network is not required for a useful data science result.

Try the interactive explanation

Open the regression and shrinkage page. Move the line and inspect residuals, MAE/MSE and penalty strength. All points are invented teaching data. The fixed-data demonstration illustrates the objective; it cannot prove a model will generalize better.

From the recording

Points the lecturer made in the lecture recording on Canvas that are not (fully) on the slides. Times refer to the recording.

  • 00:45 — Checking an imputation (Lecture 4 topic, from a student question): there is no single rule for using the same imputation method for every variable or a different one per variable. A different method per variable makes sense when the variables have different distributions (e.g. one normal, one not). The only rigorous check is to remove values you actually know, impute or predict them, and compare with the true values. The practical advice was to impute, plot the data and accept the imputation if the imputed points fall within the expected distribution rather than showing up as outliers. (not on slides)
  • 27:30 — A regression function need not be linear. In long jump, distance increases with approach speed only up to an optimum, after which ground contact is too short to take off well, so a straight line fits only part of the data and a quadratic (optimum) or polynomial function fits better. The trade-off is that higher-order functions are harder to interpret ('more is better' holds for only part of the curve), and predictions can be badly wrong for new data that come from outside the training distribution. (Slides p. 12 (linear vs polynomial fit to traffic-jam data, no explanation))
  • 33:15 — Supervised algorithms usually test hypotheses, whereas unsupervised algorithms usually generate them: you often end up with more questions than answers. Example: cluster only the 100-m times of sprinters and long jumpers into two clusters without giving labels. If each cluster mixes both groups, the discipline label does not seem to define sprint ability. That suggests the hypothesis that other factors do, and the same data with labels would instead be a supervised classification. (not on slides (p. 13 only defines clustering as dividing objects by unknown features))
  • 34:00 — Know the vocabulary of reinforcement learning, which the lecturer said you 'should know': correct decisions are rewarded and wrong decisions are penalised, so the algorithm improves over many training cycles (Pac-Man example: thousands of cycles before it won all its games). She said it can be supervised or unsupervised. The slides draw it as a separate third branch next to supervised and unsupervised learning. (Slides p. 14–15 (robot in maze, list of uses, video), p. 8–9 (separate branch). The reward/penalty mechanism is not written out.)
  • 37:20 — Ensemble (boosting) idea: train a first model (e.g. a decision tree that is only about 50% correct), then train the next model on the errors of the previous one, and repeat, so each model learns from its predecessor's mistakes. Different algorithm types can also be stacked (e.g. regression followed by classification). Correction: the lecturer said you repeat 'until you have a perfect prediction'. In reality boosting improves the fit but does not guarantee a perfect prediction, and it can overfit. (Slides p. 17 ('stupid trees learning to correct errors of each other'), p. 16)
  • 41:20 — A neuron computes a weighted sum of its inputs. Each weight expresses how important that input is (for 'apple', a round shape matters more than 'grows on a tree'). If the sum passes a threshold, the neuron outputs 'yes'. When a prediction is wrong, the weights are adjusted, and learning means adjusting the weights. Slide example: 10×0.5 + 7×1.0 + 3×0.1 = 12.3 > 10, so the output is 1 (yes). (Slides p. 19 (neuron sum diagram and MLP, no explanation of weights))
  • 46:00 — Deep learning means neural networks with more than one layer. Its advantage is automatic feature extraction: you input everything and the network decides which inputs matter, whereas in 'shallow' (classical) ML you choose and engineer the features yourself. Its cost: you lose control and context. Domain knowledge is not contained in the numbers, so the network may drop a meaningful predictor, e.g. physical activity for resting heart rate if the correlation is weak. It is also a black box: the weights are spread over many connections, so it is very hard to trace which information drove a prediction, while movement scientists want to know what makes someone, say, a good sprinter. For simple data sets, shallow ML gives the same result, so 'keep it simple'. (Slides p. 20–21 show the two workflows (feature extraction then classification vs combined). The trade-off is not on slides.)
  • 54:30 — Reading the performance-versus-data curve: with little data, both approaches perform similarly. Shallow ML plateaus because it has a limited number of possible connections (one layer), whereas deep learning keeps improving because it can form many more connections as data grow. Human movement science data sets are usually small, so pick the simpler method. Large data sets from markerless motion capture and wearables now make deep learning feasible, e.g. predicting ground reaction forces from one IMU on the shank. (Slides p. 23 (curve only, no explanation))
  • 68:30 — Selecting predictors by correlating every candidate variable with the target, or by putting everything into one multiple regression, shows that you have no hypotheses, and it misses combinations and interactions of variables. Choose features by reasoning about plausible relationships. Fewer good features give a simpler model (5 good features versus 50), and the choice of features determines the model's fit. (Slides p. 30 only says better features mean simpler models and better results. The critique is not on slides.)
  • 71:00 — Aggregation examples for time-series data (e.g. joint angles over time to classify gait): cut the signal into windows (e.g. each cycle from valley to peak to valley), take a summary such as the window mean, or fit a function (e.g. a 2nd–3rd order polynomial) and use its parameters instead of thousands of raw points. Aggregating means simplifying the data while still representing the original. (Slides p. 31 (lists only 'Window' and 'Function'))
  • 75:50 — Terminology warning: the lecturer said the terms the wrong way round. She said one-hot encoding 'uses fewer columns than you have categories' (male and female columns, with 'neither' coded 0, 0). The slides and standard usage say the opposite: one-hot = one column per category (k columns), dummy encoding = k − 1 columns, with the reference category coded as all zeros. Go by the slides. (Slides p. 34 (one-hot: d1–d3 for Red/Green/Blue), p. 35 (dummy encoding: d1–d2, Blue = 0, 0; 'dummy variable trap'))
  • 77:00 — Do not code a nominal category as a single numeric column (Italy = 1 … United States = 100). The algorithm reads the numbers as ordered quantities, so a 'bigger' code can create a spurious relationship (e.g. with salary). Use indicator columns (dummy/one-hot) so that no category gets extra weight just because of its number. (not on slides)
  • 82:50 — A downside of both MAE and MSE: because the sign is removed, they cannot show whether the model systematically under- or over-predicts. If the direction of the error matters, inspect the signed residuals as well. She called MSE the more commonly used metric. (Slides p. 41–42 state 'deviations in either direction are treated the same way' but not this consequence)
  • 83:50 — Fitting a model to the whole data set and reporting a high r or R² (e.g. 0.9) is 'only half the story': it shows that the model fits these data and that the variables are related, not how well it predicts new cases. Prediction must be evaluated on data that were not used to fit the model. (not on slides (p. 39–40 only ask for generalizable results))
  • 85:30 — Cross-validation was strongly emphasised ('I will hammer it into you'). The lecturer's description was muddled: she spoke of training on partition 1, then 2 … 10, comparing r and keeping the 'best partition'. The correct description of k-fold CV: split the training data into k folds (e.g. 10). In each round, train on k − 1 folds and evaluate on the left-out fold, so that every fold is held out once, then average the error. CV is used to tune and compare models (e.g. choose λ) and to see how much results depend on the partition, which can expose outliers. You do not select a single 'best' partition. (Slides p. 43 (CV = re-sampling procedure for assessing generalization), p. 45. The k-fold mechanics are not on slides.)
  • 87:10 — Holdout rules: split off the test set (typically about 20%, with about 80% for training, which is not a hard rule) before modelling, and 'really never touch it' until all modelling is done, otherwise 'it's cheating'. At the end, apply the trained equation to the test predictors and compare the predictions with the actual values. That agreement is the model's performance. Ideally the test data come from a different source or lab (external validation). (Slides p. 44–45 (holdout definition and workflow). The 80/20 split, 'cheating' and external lab are not on slides.)
  • 90:40 — Data leakage: when a person has several rows (longitudinal or repeated-measures data), split by person. Each athlete goes entirely into training or entirely into test, never both, and CV folds should also be split by athlete. Otherwise the athlete's characteristics leak into training and the results look too good. With one row per person, a random 80/20 split of rows is fine. Remove true duplicate rows (keep one), otherwise that case gets extra weight. (not on slides (the term 'data leakage' does not appear))
  • 98:30 — The lecturer correctly says that the steep two-point line overfits the TRAINING data: a perfect fit on training data and a bad fit on test data. This confirms that the slide's wording 'overfit to the testing data' is a slip. Least squares has no way to tell the model its fit is 'too good', which is the motivation for ridge. (At 99:30 she also slipped and said 'fit the testing data not perfectly', meaning the training data.) (Slides p. 49 (says 'New line is overfit to the testing data'))
  • 102:30 — Correction to the spoken ridge calculation: she said the residuals were 0.3 and 0.2 and to 'sum them up and square them'. The slide uses 0.3 and 0.1, squaring each and then summing: 0.3² + 0.1² + 1 × 0.8² = 0.09 + 0.01 + 0.64 = 0.74, versus 0 + 1 × 1.3² = 1.69 for the least-squares line. Use the slide numbers. (Slides p. 52)
  • 104:20 — Why ridge prefers a flatter slope: with a steep slope, tiny predictor differences produce large prediction differences (100 g of weight 'predicting' 1 m of height is absurd). A flatter slope makes predictions less sensitive and tolerates more variability, which helps on test data whose spread differs from the training data. Correction: she said a 45° slope corresponds to 'an R-value of one'. The slope and the correlation r are different things: r = 1 means all points lie on a line, whatever the slope. (Slides p. 53–56 (sensitivity explanation). The r misstatement is verbal only.)
  • 109:50 — LASSO versus ridge: she read the rise from 0.74 (ridge cost) to 0.9 (LASSO cost) for the same line as 'this variable gets less weight'. Strictly, costs from different objectives cannot be compared. The real reason is the penalty shape: for slopes below 1, |slope| is larger than slope², and the absolute penalty keeps pulling just as hard near zero, so LASSO can set a slope exactly to 0 (removing the variable). Ridge's pull fades as the slope approaches 0, so it shrinks the slope but never reaches 0. In the summary figure, as λ increases, the ridge minimum moves towards 0 but stays above it, while the LASSO minimum lands exactly on 0. (Slides p. 58 (LASSO 0.9 calculation), p. 59 (cost curves for λ = 0–400 and λ = 0–40))

Slide coverage map

Pages 1–24 introduce AI, ML, learning types, ensembles and deep learning; 25–36 develop targets, predictors and features; 37–47 cover error metrics, holdout and workflow; 48–60 work through least squares, ridge, LASSO and λ. The diagrams and arithmetic above cover the conceptual content; the original slides retain the full illustrations.

Worked through the deep dive?Tick it off. Come back to any section whenever you need it.


3Step 3 of 35–8 min

Practice questions

Exactly two original practice questions. These use the concept/application style of the 2024 and 2022–2023 example exams, with new scenarios and values. Try each before opening its answer. The official example exams are on Canvas; save them for a later, exam-style session.

Question 1 — Choose a model and evaluation plan

You have 48 athletes, 16 candidate measurements and a numerical target: recovery time in days. Several measurements may add little predictive information. Choose a model family, explain why LASSO might be worth comparing with least squares, and describe how you would evaluate it without tuning on the final test results.

Show answer and rationale

Supervised regression, because the target is numerical and supplied for training. LASSO penalizes absolute coefficients and can remove unnecessary predictors by setting them to zero, which can help control overfitting in a small dataset. Tune λ through cross-validation inside training data, fit preprocessing within those folds, and reserve a final test set where feasible. Compare with a simple mean-prediction baseline. Rationale: no penalty guarantees better predictions; use unseen-data evaluation rather than selecting the best training fit.

Question 2 — Calculate and interpret errors

Actual times are 40, 45 and 50 minutes; predictions are 42, 44 and 46. Calculate MAE and MSE, including units. Then explain how “average distance over the preceding 14 days” differs from binning daily distance into low/medium/high.

Show answer and rationale

Residuals: −2, +1, +4 minutes. MAE = 7/3 ≈ 2.33 minutes. MSE = 21/3 = 7 minutes². Squaring makes the four-minute miss count much more. Aggregation summarizes several distances using a mean and a time window; binning converts individual values to broader category labels, losing within-category distinctions. Rationale: both are feature operations but change different aspects of the data.

Return to overview · Detailed explanation

Tried both questions?Answer before peeking, then rate yourself honestly.