Data Science in Sport and Health · Practice
Question bank
142 exam-style questions, in addition to the two on each lecture page. Tap an option to check it; open questions have a model answer to compare against. The course-paper questions are on the course literature page.
Lecture 1 · Introduction to data science
12 questions · lecture page
Q1 · Disciplines of data science (Venn diagram)
The lecture's Venn diagram places data science at the overlap of computer science/IT, math and statistics, and domain/business knowledge. A sports physiologist has strong statistics skills and deep domain knowledge but hardly any programming or IT skills. In which overlap of the diagram does this profile fall?
Show answer
Answer: C. Math/statistics plus domain knowledge without computer science is labelled 'traditional research'. Machine learning is computer science plus statistics without domain knowledge, software development is computer science plus domain knowledge, and only the overlap of all three is data science.
Source: Slides p. 51 and 58
Q2 · Data, information, knowledge, wisdom
In the lecture's football example of the knowledge-discovery (KDD) pyramid, the statement 'I am pretty close to the goal' is placed at a specific level. Which level is it, and why?
Show answer
Answer: B. The slide pairs knowledge with 'context' (information read, heard or seen, then integrated and understood). 'The right winger has the ball at the corner of the penalty area' is information (meaning), 'Ball, player11, x=84, y=81' is raw data, and 'I better shoot on goal to score!' is wisdom (applied: informed decision making).
Source: Slides p. 46–47
Q3 · Typical data science questions
A sports physician has three requests: (1) 'How many weeks will this athlete need to return to play?' (2) 'Is today's resting heart rate of this athlete weird compared with her usual values?' (3) 'Which of three training programmes should this athlete follow next month?' According to the lecture's list of typical data science questions, which approaches match requests 1, 2 and 3?
Show answer
Answer: D. The slide links 'How much or how many?' to regression, 'Is this weird?' to anomaly detection and 'Which option should be taken?' to recommendation. Classification answers 'Which category?' and clustering 'Which group?'.
Source: Slides p. 66
Q4 · Structured vs unstructured data
Which statement about structured and unstructured data matches the lecture?
Show answer
Answer: A. The Igneous slide describes unstructured data as not displayable in rows, columns or relational databases, an estimated 80% of enterprise data (Gartner), needing more storage and being harder to manage; structured data are about 20%. Option C reverses DalleMule & Davenport (2017): less than 50% of structured data is used in decision-making and less than 1% of unstructured data is analysed. The IDC slide describes unstructured data as having no data model.
Source: Slides p. 38, 40–41
Q5 · Data science vs statistics
The lecture contrasts statistics with data science. Which set of characteristics does the lecture attribute to data science?
Show answer
Answer: C. The slide describes data science as an emerging field that starts with data (hypothesis generation), shines with extensive (long and wide) data and leaves room for exploration; practitioners usually have an engineering/multidisciplinary background. Options A and B describe statistics. D is wrong: the slide says both fields require carefulness and rigour and both carry a danger of unfounded conclusions.
Source: Slides p. 54
Q6 · Blei & Smyth (2017)
Blei & Smyth (2017) argue that data science is more than the combination of statistics and computer science. Which of the following is NOT one of the additional requirements named in their quote on the slide?
Show answer
Answer: D. The quote names understanding the context of data, appreciating the responsibilities of using private and public data, and clearly communicating what a dataset can and cannot tell us. Collecting as much as possible is not mentioned and conflicts with the lecture's emphasis on first identifying the problem.
Source: Slides p. 49–50
Q7 · Data science lifecycle: order of steps
Which sequence follows the order of the data science lifecycle shown in the lecture?
Show answer
Answer: B. The cycle runs: identify the (business) problem → data acquisition → data preparation → data exploration → feature engineering → data modelling → data visualization → present & communicate → deployment and maintenance. The tempting slip is placing exploration before preparation; in the lecture's cycle, data are prepared before they are explored.
Source: Slides p. 64–65
Q8 · R as a programming language
According to the lecture, which of the following is a disadvantage of R?
Show answer
Answer: A. The slide lists two minuses: computation time with large datasets and many packages with the same functionality ('many ways to Rome'). The pluses include free and open source, a large community, visualizations, statistics, popularity in academia and use in many industries.
Source: Slides p. 25
Q9 · Data science applications in sport
The lecture (Chmait & Westerbeek, 2021) groups examples of data science in sports research into four areas. Which pairing of example and area is correct as listed on the slide?
Show answer
Answer: B. Training & coaching covers assessment of team formation efficacy, training optimization and player injury modelling. Ticket pricing is fan & business focused, event classification belongs to game analytics, and player recruitment to talent identification.
Source: Slides p. 52
Q10 · RStudio environment
You have imported a data frame in RStudio and want to see which objects you have created so far in this session. Which of the four RStudio panes shown in the lecture is meant for this?
Show answer
Answer: C. The lecture labels the four RStudio panes as code editor, R console, workspace & history, and plots & files; the workspace (environment) pane lists the objects you have assigned. The console runs commands (you could type ls() there), but it does not keep an overview of your objects.
Source: Slides p. 28
Q11 · Hypothesis generation vs hypothesis testing
A club's data scientist explores three seasons of existing GPS and injury data and notices that players who covered more high-speed running in the week before a match had fewer hamstring injuries. Using the lecture's 'two phases in the scientific process', explain in which phase this finding sits, why a coach might already be satisfied while a journal would not, and what would be needed next.
Show model answer
The finding comes from the hypothesis-generation phase (dataset → analysis → hypothesis), which the lecture associates with data science/data mining. A coach is 'pretty happy' at this stage, because a plausible data-based pattern can already inform practice. A journal is only satisfied after the hypothesis-testing phase (statistics): the specific hypothesis must be tested on newly collected data, because a pattern found by exploring the same data can be a chance finding (danger of unfounded conclusions). Next step: state the hypothesis precisely and collect new data (e.g. a following season or another team) to test it statistically, keeping in mind that an association does not prove that high-speed running prevents injuries.
Source: Slides p. 54–55
How did your answer compare?
Q12 · Identify the (business) problem
A hockey coach says: 'We have lots of wearable data — can you do something with AI?' Describe what you should do in the first step of the data science lifecycle before acquiring or analysing data. Give two concrete actions from the lecture, and turn the coach's request into a specific question with a clear target variable.
Show model answer
The first step is to identify the (business) problem/project objective; the lecture stresses that it is critical to understand the problem first. Actions from the slides: talk to domain experts (coach, physio, players); determine which questions you are asking and how answering them achieves the club's goals; keep asking 'why?'; and specify the key target variable(s). Example of a specific question: 'Can a player's training load in the previous 7 days predict his self-reported fatigue score (target, 1–10) the next morning?' — a 'how much?' question, i.e. regression. Any question with a clearly defined target, population and time frame is acceptable; 'do something with AI' is not a defined objective.
Source: Slides p. 65–67
How did your answer compare?
Lecture 2 · Programming in R
12 questions · lecture page
Q1 · Vectors and coercion
What does the following code print?
x <- c(12, "15", TRUE)
typeof(x)Show answer
Answer: B. A vector holds one data type, so R coerces all elements to the most flexible type present; with one character element everything becomes character ("12" "15" "TRUE"). Only a list keeps mixed types; typeof() gives "double" only if all elements are (non-integer) numbers.
Source: Slides p. 9–10, 16–18; DataCamp Introduction to R (vectors). Output verified in R 4.6.1.
Q2 · Variables, assignment and packages
Which statement about working in R is correct according to the lecture?
Show answer
Answer: D. A package must be installed before it can be loaded; the lecture shows install.packages("lubridate") followed by library(lubridate) and ymd(). R is case-sensitive, names cannot start with a number (1trial and 2nd cause errors), and ls() lists only object names; typeof() returns the data type.
Source: Slides p. 13, 15–16, 57
Q3 · Matrices
What does the last line return?
m <- matrix(1:6, nrow = 2)
m[1, 3]Show answer
Answer: C. matrix() fills column by column unless byrow = TRUE, so the columns are (1, 2), (3, 4) and (5, 6); row 1, column 3 is 5. The answer 3 would only be correct with byrow = TRUE. Six values in two rows give three columns.
Source: Slides p. 19, 28; DataCamp Introduction to R (matrices). Output verified in R 4.6.1.
Q4 · Factors
What does the following code print?
event <- factor(c("sprint", "road", "track", "road", "sprint"))
as.integer(event)Show answer
Answer: A. By default factor levels are sorted alphabetically (road, sprint, track) and each value is stored as the integer code of its level: sprint = 2, road = 1, track = 3. The codes do not follow the order of appearance (1 2 3 2 1), and they are labels, not quantities to calculate with.
Source: Slides p. 17, 20; DataCamp Introduction to R (factors). Output verified in R 4.6.1.
Q5 · Lists
Given the list below, which expression returns the single number 61?
athlete <- list(name = "Sanne", hr = c(58, 61, 64), injured = FALSE)Show answer
Answer: C. Double brackets extract the element itself (the numeric vector 58 61 64), after which [2] gives 61. Single brackets return a sub-list, so athlete["hr"][2] asks for a non-existent second element of a one-element list (NULL). A list has no rows and columns, so athlete[2, 2] gives "incorrect number of dimensions"; athlete[[2]][3] returns 64.
Source: Slides p. 17–18, 20; DataCamp Introduction to R (lists). Output verified in R 4.6.1.
Q6 · Data frames and logical operators
What does the last line return?
d <- data.frame(name = c("Ann", "Bo", "Cas", "Dirk"),
vo2 = c(48, 55, 61, 52),
mass = c(60, 72, 80, 68))
d[d$vo2 > 50 & d$mass < 75, "name"]Show answer
Answer: B. Rows are kept where both conditions are TRUE (&): vo2 > 50 holds for Bo, Cas and Dirk; mass < 75 for Ann, Bo and Dirk; both only for Bo and Dirk. Then the "name" column of those rows is returned. The first option ignores the mass condition; the third is what | (OR) would give.
Source: Slides p. 21, 29, 48–49; DataCamp Intermediate R (relational and logical operators). Output verified in R 4.6.1.
Q7 · Conditional statements
What is the value of zone after running this code?
rpe <- 8
if (rpe >= 9) {
zone <- "maximal"
} else if (rpe >= 5) {
zone <- "hard"
} else if (rpe >= 7) {
zone <- "very hard"
} else {
zone <- "easy"
}
zoneShow answer
Answer: D. R checks the conditions from top to bottom and runs only the first TRUE branch: 8 >= 9 is FALSE, 8 >= 5 is TRUE, so zone becomes "hard" and all remaining branches are skipped. The "very hard" branch can never be reached; that is why the lecture's height example tests the highest threshold first.
Source: Slides p. 50–52; DataCamp Intermediate R (conditionals). Output verified in R 4.6.1.
Q8 · For-loops
What does the last line print?
hr <- c(150, 172, 185, 168, 190, 176)
count <- 0
for (h in hr) {
if (h >= 170 & h < 185) {
count <- count + 1
}
}
countShow answer
Answer: A. In each iteration h takes the next value of hr, and count increases only when 170 ≤ h < 185: true for 172 and 176. 185 fails h < 185, 190 is too high, and 150 and 168 are below 170. The answer 3 would follow from h <= 185; 6 is the number of iterations, not the number of TRUE conditions.
Source: Slides p. 49, 53–56; DataCamp Intermediate R (loops). Output verified in R 4.6.1.
Q9 · Writing functions
What do the two function calls return?
abs_vo2 <- function(rel, mass = 70) {
round(rel * mass / 1000, 1)
}
abs_vo2(c(48.2, NA, 61.5))
abs_vo2(mass = 80, rel = 50)Show answer
Answer: C. mass has a default of 70, so the first call works and is applied element-wise: 48.2 × 70 / 1000 = 3.374 → 3.4; NA stays NA (arithmetic with NA gives NA and nothing removes it); 61.5 × 70 / 1000 = 4.305 → 4.3. In the second call arguments are matched by name, so 50 × 80 / 1000 = 4.
Source: Slides p. 36, 39, 43–46; DataCamp Intermediate R (functions). Output verified in R 4.6.1.
Q10 · apply family
What does the last line return?
m <- matrix(c(2, 4, 6, 8, 10, 12), nrow = 2)
apply(m, 1, sum)Show answer
Answer: B. Filling by column gives row 1 = 2, 6, 10 and row 2 = 4, 8, 12. MARGIN = 1 applies sum() to each row: 18 and 24. MARGIN = 2 would give the column sums (6 14 22); 12 30 assumes the matrix was filled by row; 42 is sum(m).
Source: Slides p. 19, 28 (matrix filling); DataCamp Intermediate R (apply family). Output verified in R 4.6.1.
Q11 · Inspecting a data frame with str()
After importing a CSV file with 500 m race results as the data frame races, you run str(races) and get the output below. (a) How many rows and columns does races have? (b) Why does mean(races$time) not return the average race time, and how do you fix it? (c) What kind of object does races$time return?
'data.frame': 4 obs. of 3 variables:
$ skater: chr "Ann" "Bo" "Cas" "Dirk"
$ season: int 2021 2021 2022 2022
$ time : chr "35.40" "36.07" "34.98" "35.62"Show model answer
(a) 4 observations (rows) of 3 variables (columns). (b) time is stored as character (chr), e.g. because it was read as text from the CSV; mean() of a character vector returns NA with a warning ('argument is not numeric or logical'). Convert first with races$time <- as.numeric(races$time); then mean(races$time) gives 35.5175 (add na.rm = TRUE if values are missing). (c) The $ operator extracts one column as a vector (here a character vector, numeric after conversion), whereas races["time"] would return a one-column data frame.
Source: Slides p. 23, 29, 38; Recording 20:00–25:00. Output verified in R 4.6.1.
How did your answer compare?
Q12 · Combining a function, if–else and a for-loop
Write R code that (1) defines a function hr_zone() returning "high" for a heart rate of 180 bpm or more, "moderate" for 150–179 bpm and "low" below 150 bpm, and (2) uses a for-loop to store the zone of every value of session_hr <- c(132, 155, 181, 149) in a new vector. Give the resulting vector and explain why the order of the conditions matters.
Show model answer
hr_zone <- function(hr) { if (hr >= 180) { "high" } else if (hr >= 150) { "moderate" } else { "low" } } session_hr <- c(132, 155, 181, 149) zones <- character(length(session_hr)) for (i in seq_along(session_hr)) { zones[i] <- hr_zone(session_hr[i]) } zones # "low" "moderate" "high" "low" Only the first TRUE branch of an if–else if chain is executed, so the highest threshold must be tested first; if hr >= 150 were tested first, 181 would be labelled "moderate". The function body is written once and reused for every element (don't repeat yourself); sapply(session_hr, hr_zone) gives the same result without an explicit loop.
Source: Slides p. 43–46, 50–56; DataCamp Intermediate R (functions, loops, apply family). Output verified in R 4.6.1.
How did your answer compare?
Lecture 3 · Data in sport and health
12 questions · lecture page
Q1 · Structured vs unstructured data in sport and health
A rehabilitation centre films patients walking, runs markerless motion-capture software that writes the hip, knee and ankle coordinates of every video frame into a table, and lets therapists add a free-text note after each session. How should these three data sources be classified, following the lecture?
Show answer
Answer: A. Video and imaging are mostly unstructured, but once landmarks are extracted the result is data points organized in a table, which is structured ('it's not a video anymore'). Open-ended text answers are unstructured. Deriving a table from a video does not make the table unstructured.
Source: Slides p. 6; Recording 05:00–10:00
Q2 · Interpreting the Fitbit graph
The Fitbit slide plots average resting heart rate (y-axis) against BMI (x-axis) for men and women. Which conclusion is best supported by the graph?
Show answer
Answer: B. The women's points lie above the men's at every BMI, and both curves are lowest around BMI 20–24 (men about 62 bpm, women about 65 bpm), rising towards both lower and higher BMI. Men at BMI 45 are around 69–70 bpm, above women at BMI 20. The data are observational, so a causal effect of changing BMI cannot be concluded; the wide error bars at the lowest BMI reflect few users in that category.
Source: Slides p. 10; Recording 25:00–30:00 and 35:00–45:00
Q3 · Value of data: activity-based insurance discount
The lecture discussed an insurer (ASR) that gives discounts and gifts when members share activity-tracker data and reach personal activity goals. What did the lecturer report about the effect of such a scheme?
Show answer
Answer: C. According to the lecturer, already-active people received the discount and maintained (some slightly increased) their activity, while inactive people did not change; members who later became less active could even experience losing the discount as a punishment. Commercial value of activity data does not automatically mean behaviour change.
Source: Slides p. 9; Recording 15:00–25:00
Q4 · Restricted range in regression
In the van der Zwaard et al. (2018) example, pennation angle of the vastus lateralis (from ultrasound) was not among the determinants of sprint power in the elite cyclists. According to the lecture, how should you interpret this?
Show answer
Answer: D. The lecturer warned that top cyclists may all have similar pennation angles; if a predictor hardly varies while performance does, the model finds no relationship, which does not prove the variable is unimportant (same logic as VO2max in the top-10 marathon runners). Always check the variability of all parameters. Deming regression accounts for errors in both variables rather than ignoring them.
Source: Slides p. 22, 27; Recording 60:00–70:00
Q5 · Determinants of combined sprint and endurance performance
In the van der Zwaard et al. (2018) results shown in the lecture, which combination of muscle characteristics together explained 67% of the variance in performance VO2, a main determinant of combined sprint and endurance performance?
Show answer
Answer: C. The slides show oxidative capacity ↑, capillaries/myoglobin ↑ and PCSA ↓ explaining 67% (paper: R² = 0.67 for performance VO2); the conclusion names long muscle fibres and capillaries/myoglobin as training targets. MHC type II + muscle volume (65%) was the best combination for sprint peak power, not for combined performance. Exam note: in the recording the lecturer says this 'explains combined performance by 67%'; strictly, the 67% refers to performance VO2.
Source: Slides p. 27–28, 30; Paper p. 2110 (abstract) and p. 2115
Q6 · Excel and CSV limits
A sports scientist stores heart rate at 1 Hz for 25 athletes during a 2-hour training, in long format (one row per athlete per second): 25 × 7,200 = 180,000 rows and 3 columns. Which statement about storing this in a single file/sheet is correct according to the lecture's limits?
Show answer
Answer: A. The slide gives .xls a limit of 65,536 rows and 256 columns, .xlsx 1,048,576 rows and 16,384 columns, and csv 'unlimited'. 180,000 rows exceeds the .xls row limit but not the .xlsx limit; 256 is the .xls column limit, not .xlsx.
Source: Slides p. 34
Q7 · Application Programming Interface (API)
Which statement about an API matches the lecture?
Show answer
Answer: B. The slides state that the API is not the database or the server, but the code governing the access point(s) to the database; requests to retrieve or write data are generally done without a frontend via HTTP, and data are exchanged as e.g. JSON or XML. Option C describes web scraping (processing a web document and extracting information from it). The lecturer stressed that the quality of public data is often unknown.
Source: Slides p. 38, 41–43; Recording 75:00–90:00
Q8 · Speed-skating API example
Seasonal bests of speed skaters were retrieved through the SpeedskatingResults.com API and shown in a radar plot per distance, as the seasonal best relative to the world record on the race day. In this visualization, what do the direction and the size of the filled area represent?
Show answer
Answer: D. The slides state that area direction highlights specialization and area size highlights performance level, with seasonal development shown as an animation. Symbols show track characteristics (indoor vs open, highland vs lowland) and a circle marks a new world record. A value of 1.0 means the seasonal best equals the world record; 1.2 means 120% of the world-record time.
Source: Slides p. 45, 47, 49–51; Recording 80:00–85:00
Q9 · Extract–transform–load (ETL)
A club combines training files from three GPS vendors into its data warehouse. Removing duplicate sessions, converting all speeds to m/s and flagging heart rates above a plausible threshold belong to which ETL phase, and where does that usually take place?
Show answer
Answer: B. Transformation structures, enriches and converts the raw data to match the target: e.g. cleaning, deduplication, format revision, data threshold validation checks, merging and aggregation, usually in a staging database. Extraction retrieves raw data from the sources into staging; loading writes the converted data from staging into the target warehouse.
Source: Slides p. 57–61
Q10 · Data warehouse vs data lake
Which characteristic belongs to a data lake rather than a data warehouse, according to the lecture's comparison?
Show answer
Answer: C. The comparison table gives the data lake: structured/semi-structured/unstructured raw data, schema-on-read, low-cost storage, highly agile, maturing security, used by data scientists. Options A and B describe a data warehouse; D describes a data mart. A lake requires no prior knowledge of the analysis needed: low effort to store, high effort to reuse.
Source: Slides p. 62, 64–65; Recording 90:00–100:00
Q11 · Deming regression and the sprint–endurance relationship
In van der Zwaard et al. (2018), normalized sprint performance (Wingate) and endurance performance (15-km time trial) of 28 cyclists were inversely related (r = −0.66, P < 0.001), with the line fitted by Deming regression. (a) How does Deming regression differ from ordinary least-squares regression, and why does it suit this question? (b) Interpret r = −0.66, including the proportion of shared variance.
Show model answer
(a) Ordinary least squares minimises the vertical residuals (errors in y only) and treats x as an error-free predictor. Deming regression accounts for errors in both x and y; here the residuals are the perpendicular distances from each cyclist to the line. This suits the question because both sprint and endurance performance are measured with error and neither is the natural predictor; the perpendicular residual also expresses how well a cyclist combines both kinds of performance relative to the group, which was then related to the physiological determinants. (b) r = −0.66 is a moderately strong negative relationship: cyclists with higher sprint performance tend to have lower endurance performance (combining both is difficult). r² ≈ 0.44, so about 44% of the variance is shared; P < 0.001 makes a chance finding unlikely, but the correlation does not show causation. Exam note: in the recording the lecturer said '36 to 40 percent'; 0.66² is 0.44.
Source: Slides p. 21, 23, 25; Paper p. 2111, 2113 and Fig. 2 (p. 2116); Recording 45:00–55:00
How did your answer compare?
Q12 · Choosing a storage solution
A national skating federation has (1) raw IMU files, race videos and spreadsheets from all athletes whose future use is still unclear, (2) cleaned results and training data from many sources that management wants to combine for reporting, and (3) a medical staff that only needs summarized injury and load data. Explain the difference between a data lake, a data warehouse and a data mart, assign each need to one of them, and state the trade-off the lecturer mentioned about data lakes.
Show model answer
A data lake is a highly scalable store that keeps structured and unstructured data in their original form and format, without requiring planning of the analysis (schema-on-read) → need 1. A data warehouse is a single, enterprise-wide repository that pulls together (ETL-processed) data from many internal and external sources for reporting and analysis, holding raw, summary and metadata (schema-on-write) → need 2. A data mart is a subset of a data warehouse oriented to one business line or unit, containing summarized data → need 3. Trade-off: a lake means low effort when saving but high effort when reusing the data (chaotic, unclean); the lecturer recommends organizing data into marts for reusability.
Source: Slides p. 55–56, 62–65; Recording 90:00–100:00
How did your answer compare?
Lecture 4 · Preparation, exploration and visualization
12 questions · lecture page
Q1 · Measurement scales
Which statement is meaningful, given the measurement scale of the variable involved?
Show answer
Answer: C. Duration has a ratio scale (absolute zero), so ratios such as "twice as long" are meaningful. Celsius temperature is interval: differences are meaningful but 0 °C is not an absolute zero, so ratios are not; medal codes are ordinal and shirt numbers nominal, so arithmetic on their codes has no meaning.
Source: Slides p. 7–9
Q2 · Data quality dimensions
Which pairing of a data problem with a DOMA data-quality dimension is correct?
Show answer
Answer: B. Validity is about conforming to the defined syntax (format, type, range). The other problems belong to uniqueness (duplicated record), timeliness (when was it updated?) and consistency (conflicting representations such as male and pregnant).
Source: Slides p. 10
Q3 · FAIR data
You deposit your thesis dataset in a trusted repository. It has a persistent identifier (DOI) and rich, machine-readable metadata, and it can be downloaded after logging in. You did not add a usage licence or any information on how, when and by whom the data were collected. Which FAIR principle is least well met?
Show answer
Answer: D. On the lecture slide, Reusable means the data have a clear usage licence and accurate provenance information, and both are missing here. The DOI and rich metadata cover Findable, the trusted repository covers Accessible (it does not have to be open to everyone), and machine-readable metadata supports Interoperable.
Source: Slides p. 11–12; Recording 00:10–00:15
Q4 · dplyr: filter with | and NA
How many rows does this pipeline return?
olympics_small <- data.frame(
Name = c("A", "B", "C", "D", "E"),
Year = c(2008, 2008, 2012, 2016, 2012),
Medal = c("Gold", NA, "Gold", "Silver", NA)
)
olympics_small %>%
filter(Year == 2008 | Medal == "Gold") %>%
nrow()Show answer
Answer: A. | keeps a row when either condition is TRUE: A (both), B (Year 2008; TRUE | NA is TRUE) and C (Gold). D fails both, and for E the result is FALSE | NA = NA, which filter() drops. Using & instead would return only athlete A (verified in R).
Source: Slides p. 24–25
Q5 · Duplicates and distinct()
In the Olympics data each row is one athlete's entry in one event (271,116 rows; str() shows 134,732 different names). The lecturer ran the code below. Why does distinct() remove about 84,000 rows?
df1 <- data_olympics %>%
select(ID, Name, Age) %>%
rename(Athlete_code = ID)
df1 %>% nrow() # 271116
df1 %>% distinct() %>% nrow() # 187077Show answer
Answer: B. Whether a row counts as a duplicate depends on which columns you keep. After the select() step, rows for the same athlete in different events at the same Games become identical. The result is not the number of unique athletes (about 134,732 names), because athletes who competed at different ages keep separate rows.
Source: Slides p. 29, 39; Recording 01:10–01:15
Q6 · Joining data frames
You want a table with all four athletes and their countermovement-jump (CMJ) result, with NA where no test exists. Which call gives this?
athletes <- data.frame(ID = c(1, 2, 3, 4),
Sport = c("Rowing", "Judo", "Rowing", "Cycling"))
tests <- data.frame(Athlete_code = c(2, 4, 5),
CMJ_cm = c(41, 38, 45))Show answer
Answer: D. left_join() keeps every row of the first data frame. Because the key has different names in the two tables, you match them with by = c("ID" = "Athlete_code"). The inner join returns only IDs 2 and 4, the right join returns IDs 2, 4 and 5 (with NA sport for 5), and by = "ID" gives an error because tests has no ID column (all verified in R).
Source: Slides p. 26–28
Q7 · EDA: suspicious values
summary() of the Olympics data shows a minimum weight of 25 kg, a maximum height of 226 cm and a maximum age of 97 years. What is the most appropriate next step?
Show answer
Answer: C. EDA is iterative. In the lecture, inspecting these records showed real cases: a 14-year-old gymnast of 25 kg, a 226 cm Chinese basketball player, and a 97-year-old competing in the former art competitions. A potential outlier needs to be inspected, not deleted automatically.
Source: Slides p. 34–35, 40–42, 48; Recording 01:20–01:25
Q8 · Boxplot reading
summary() of Olympic athletes' age gives 1st Qu. = 21, Median = 24, 3rd Qu. = 28 and Max. = 97. Under the boxplot rule shown in the lecture, which statement is correct?
Show answer
Answer: A. IQR = Q3 − Q1 = 28 − 21 = 7, and the upper whisker limit is Q3 + 1.5 × IQR = 28 + 10.5 = 38.5. More extreme points are plotted individually as potential outliers. The box runs from Q1 to Q3 (the middle 50%), and the whisker stops at the most extreme observation within the 1.5 × IQR limit, not at 97.
Source: Slides p. 40, 48; Recording 01:25
Q9 · Correlation and explained variance
In the lecture's corrplot of Olympic athletes, height and weight correlate at r = 0.8, and age and height at r = 0.14. Which statement is correct?
Show answer
Answer: D. Explained variance is R² = r² = 0.64, not r. A small r means only a weak linear association; a nonlinear relation could still exist. The spurious-correlation slide warns that correlation is not causation. Exam note: the exported PDF of p. 53 flattens two animation steps: the cheese/bedsheets chart (r = 0.947 → r² = 89.7%) and the Nicolas Cage chart (r = 0.666 → r² = 44.4%), so the 89.7% label can appear next to the wrong chart. Each r² belongs to its own chart.
Source: Slides p. 52–53
Q10 · ggplot2: facets and layers
Using the mpg dataset, which code produces a separate scatterplot panel of displ against hwy for each vehicle class?
Show answer
Answer: B. Facets create subpanels, and ggplot2 layers are added with + (this version gives 7 panels in R). The colour mapping draws all classes in one panel, the %>% version gives an error because layers are not piped, and group/theme change grouping or appearance without creating panels.
Source: Slides p. 61–65, 68–70
Q11 · Missing values in the Olympics data
The output below comes from the Olympics dataset. Explain (1) why Medal is missing in 85% of rows and whether this is a data-quality problem, (2) how you would handle Medal in R, and (3) why BMI has a higher percentage of missing values than Height or Weight.
> colSums(is.na(data_olympics)) / nrow(data_olympics) * 100
ID Name Sex Age Height Weight Team NOC
0.000000 0.000000 0.000000 3.494445 22.193821 23.191180 0.000000 0.000000
Games Year Season City Sport Event Medal BMI
0.000000 0.000000 0.000000 0.000000 0.000000 0.000000 85.326207 23.703138Show model answer
Medal is NA for athletes who did not win a medal (only three medals per event), so the missingness carries information and is not lost data. Deleting those rows would throw away 85% of the data, and imputing a medal would be wrong. Recode it instead: mutate(Medal = ifelse(is.na(Medal), "No medal", Medal)), after which Medal has 0% NA. BMI is calculated from weight and height, so it is NA whenever either one is missing. It therefore has at least as many NAs as the more incomplete of the two (23.70% versus 23.19% and 22.19%). The Height/Weight gaps are a genuine completeness issue to investigate (for example, by year) before choosing deletion or imputation.
Source: Slides p. 10, 30–31; Recording 01:10–01:15
How did your answer compare?
Q12 · Interpreting boxplots over time
In the lecture's boxplots of female Olympic athletes' age per Games, the 1900 and 1904 boxes look very different from each other and from later Games (the 1904 box sits far above every other year). Later years look stable around the early twenties. Give two reasons why the early boxes may not reflect a real shift in the age of female athletes, and say what you would plot or calculate to check.
Show model answer
(1) Very few women competed in the early Games (the participation line graph shows small numbers early on). A boxplot of only a handful of observations is unstable: one or two athletes can move the median and quartiles a lot. (2) The mix of sports and events changed over time, and some early events attracted older participants, so composition rather than ageing drives the difference. To check, count athletes per year and sex (for example, group_by(Year, Sex) %>% summarise(n = n()), or the participation line graph) and look at which sports the early female athletes competed in. You can also show the raw points or a violin/histogram next to the boxplot so the sample size is visible.
Source: Slides p. 44, 47–50; Recording 01:25–01:30
How did your answer compare?
Lecture 5 · Machine learning and regression
17 questions · lecture page
Q1 · Supervised task typecatch-up notes
You predict next week's numerical fatigue score from training duration and workload. Which task is this?
Show answer
Answer: B. The target is a number and its true values are available for training, so this is supervised regression. Classification predicts categories, while clustering and PCA are unsupervised and have no target.
Source: Slides p. 7, 11, 27
Q2 · Feature engineering: choosing an aggregatecatch-up notes
One faulty measurement can make a seven-day minimum unusually low. Why might a percentile (e.g. the 25th) of the same seven days be a more useful aggregate feature?
Show answer
Answer: A. An aggregate feature combines a variable, a window and a function, and percentile and minimum are both options on the slide. The minimum is determined by one value, so one artefact can drive it, whereas a less extreme percentile reflects more of the window. That is a reason, not a guarantee: no aggregate is always best or removes error.
Source: Slides p. 31
Q3 · Naive baseline comparisoncatch-up notes
A model predicting recovery time performs no better than always predicting the training-set mean. What is the strongest conclusion?
Show answer
Answer: C. A predictor-free baseline shows how much the features and model add, which is the "measure how accurately the model predicts" idea on the evaluation slide. Matching it means no demonstrated benefit. It does not prove that no other model could work, and by itself it does not diagnose overfitting. Exam note: the naive baseline appears in both reference exams, not on these slides.
Source: Slides p. 39–40
Q4 · AI, machine learning, deep learning and data science
Which statement matches how the lecture positions AI, machine learning (ML), deep learning (DL) and data science?
Show answer
Answer: D. The nested circles show DL inside ML inside AI. The "putting it in perspective" slide states that neither ML nor AI is a subset of data science, and data science is a subset of neither: they overlap (e.g. decision trees, kNN), while data science also includes statistics, data preparation, visualization and experimentation.
Source: Slides p. 3, 24
Q5 · Dummy variables
A linear regression with an intercept predicts power output from cycling discipline (sprint, pursuit or road). Why use two dummy columns (d1, d2, with road = 0, 0) instead of three one-hot columns?
Show answer
Answer: A. Dummy encoding uses k − 1 columns for k categories, and the slide's colour example codes Blue as 0, 0. Road cyclists stay in the data as the reference group (both dummies 0), and the indicators do not impose an order. Exam note: the slide names the "dummy variable trap" without spelling out the reason; the standard reason is perfect collinearity with the intercept.
Source: Slides p. 33–35
Q6 · MAE and MSE calculation
A model predicts 400-m times (s) for five athletes. What are the MAE and MSE?
actual <- c(50, 52, 48, 55, 60)
predicted <- c(48, 54, 47, 56, 54)Show answer
Answer: D. Residuals are 2, −2, 1, −1 and 6, so MAE = 12/5 = 2.4 s and MSE = (4 + 4 + 1 + 1 + 36)/5 = 9.2 s² (verified in R). Option A averages the signed residuals, so they cancel. Option B is the RMSE, √9.2 ≈ 3.0 s, which is not on the slides but simply returns MSE to seconds. Option C forgets to divide by n. The single 6-s miss contributes 36 of the 46: MSE punishes large errors more.
Source: Slides p. 41–42
Q7 · Overfitting with a small sample
On the least-squares slides, a new line Size = 0.4 + 1.3 × Weight passes exactly through the two red training points (sum of squared residuals = 0) but has a large sum of squared residuals for the green testing points. What does this show?
Show answer
Answer: B. A perfect fit to two training points that fails on new data is overfitting (high variance), which is the motivation for ridge regression. Minimizing training residuals does not guarantee good test performance. Exam note: slide p. 49 literally says the line "is overfit to the testing data". It is overfitted to the training data and therefore performs poorly on the testing data.
Source: Slides p. 48–50
Q8 · Ridge regression calculation
Two training points. Line 1 (least squares) has slope 1.5 and residuals 0 and 0. Line 2 has slope 1.0 and residuals 0.3 and 0.2. With λ = 1, compute each line's ridge cost (sum of squared residuals + λ × slope²). Which statement is correct?
Show answer
Answer: C. Line 1 costs 0 + 1 × 1.5² = 2.25 and Line 2 costs 0.3² + 0.2² + 1 × 1.0² = 0.09 + 0.04 + 1.00 = 1.13, so ridge prefers the flatter line. This is the same logic as the slide (1.69 versus 0.74). Option A ignores the penalty, and Option B uses |slope| (the LASSO penalty) for Line 1.
Source: Slides p. 51–52
Q9 · Ridge: the role of λ
Which statement about λ in ridge regression is correct?
Show answer
Answer: A. With no penalty (λ = 0) ridge equals least squares. A larger λ shrinks the slope, so predictions become less sensitive to the predictor. Training error is always lowest at λ = 0, so it cannot be used to choose λ. Setting slopes exactly to zero is the property of LASSO, not ridge. (The slide's "lowest residuals for test data" refers to the held-out folds in cross-validation.)
Source: Slides p. 53–59
Q10 · Ridge versus LASSO
You have 40 rowers and 25 candidate predictors of 2000-m ergometer time, and you expect only a handful to matter. Which choice best follows the lecture's summary?
Show answer
Answer: D. According to the summary slide, LASSO outperforms ridge when many unnecessary features are included, because their slopes can become 0. Ridge is better when most features are meaningful. Ridge and LASSO are both more robust to overfitting than least squares in small data sets, and ridge with λ = 0 is least squares.
Source: Slides p. 58, 60
Q11 · Holdout and tuning
You fit LASSO models with 20 values of λ, compute each model's error on the test set, pick the λ with the lowest test error, and report that error as your final performance. What is the problem?
Show answer
Answer: B. The holdout slide says to train and tune with cross-validation on the training set and not to touch the test set "until the very end". If you choose λ on the test data, that data is no longer unseen. Choosing on training error would always pick λ = 0.
Source: Slides p. 43–45, 57
Q12 · Choosing and evaluating a regression modelcatch-up notes
You have 45 athletes and several potential predictors of a numerical performance score. Propose a model and explain how you would evaluate it. Mention one reason why the estimated performance could be uncertain.
Show model answer
The target is numerical, so this is supervised regression. A sensible choice is to compare least squares with ridge or LASSO. The penalty makes the model more robust to overfitting in a small data set, and LASSO can drop unhelpful predictors. Split the data into training and test sets, tune λ (and choose between models) with k-fold cross-validation inside the training data, and evaluate once on the untouched test set with MAE/MSE, ideally compared with a simple baseline. With only 45 athletes the estimate is uncertain: a holdout leaves few test observations, so results depend on the split. Several defensible answers exist; what matters is identifying the target and justifying the choices.
Source: Slides p. 27, 41–45, 57–60
How did your answer compare?
Q13 · Ridge versus LASSO
Explain how ridge and LASSO regression differ from least squares and from each other (the penalty term and its effect on slopes). Say what λ controls, how you would choose it, and when you would prefer each model.
Show model answer
Least squares minimizes the sum of squared residuals. Ridge adds λ × slope² and LASSO adds λ × |slope|, so both accept some training error in exchange for smaller (less steep) slopes, which makes predictions less sensitive to the predictors and reduces overfitting, especially in small data sets. λ sets the penalty strength: λ = 0 gives least squares, and a larger λ shrinks slopes more. Choose λ by cross-validation (lowest error on the held-out folds), not on training error or the final test set. Ridge shrinks slopes toward zero but keeps all predictors, so prefer it when most features are meaningful. LASSO can set slopes exactly to zero and thereby remove variables, so prefer it when many features are unnecessary.
Source: Slides p. 50–60
How did your answer compare?
Q14 · Data leakage: splitting repeated measuresfrom the recording
You have 30 training sessions from each of 12 cyclists (360 rows) and want to predict session RPE from workload variables. How should you create the test set, following the lecture?
Show answer
Answer: C. Rows from the same person share that person's characteristics. A random row split puts the same cyclist in both sets, so information leaks into training ("data leakage") and performance looks better than it would be for new people. The same rule applies to cross-validation folds. Option A is fine only when each person has a single row. Exam note: splitting by athlete and the term "data leakage" come from the recording, not the slides.
Source: Recording 90:30–92:30; Slides p. 44
Q15 · Unsupervised learning generates hypothesesfrom the recording
A researcher gives a clustering algorithm only the 100-m times of a group of sprinters and long jumpers and asks for two clusters. Both clusters contain a mix of sprinters and long jumpers. According to the lecture, what is the best interpretation?
Show answer
Answer: A. Clustering gets no labels and searches for groups itself. Mixed clusters suggest that the label does not define sprint speed, which leads to a new question, the lecturer's point that unsupervised learning generates hypotheses while supervised learning tests them. The algorithm never saw the labels, so this was not classification, and an exploratory clustering proves nothing (D).
Source: Recording 31:30–33:30; Slides p. 13
Q16 · Encoding categorical variablesfrom the recording
In a football data set with players from 100 countries, a student codes country as one numeric column (Italy = 1, …, United States = 100) and uses it to predict salary. What is the main problem the lecturer pointed out?
Show answer
Answer: D. A single integer column implies an order and spacing that nominal categories do not have, so a "bigger" country number could appear to raise salary. Indicator columns remove that implied order. Option B describes one-hot encoding with many categories (many zero columns), which is a separate issue: the lecturer suggested binning sparse countries into continents for that. Exam note: on the slides, one-hot = one column per category and dummy encoding = one fewer column. The lecturer said this the other way round in the recording.
Source: Recording 74:30–77:30; Slides p. 32–35
Q17 · Shallow ML versus deep learningfrom the recording
Why did the lecturer advise students to use shallow (classical) machine learning rather than deep learning for their course data sets?
Show answer
Answer: B. The performance curve shows the two approaches performing similarly at small data sizes, with only deep learning continuing to improve as data grow. Deep learning also does the feature extraction itself, which costs interpretability and the researcher's context, so the advice was to "keep it simple". Option C reverses the curve and option D reverses the feature-extraction contrast.
Source: Recording 46:00–55:45; Slides p. 20–23
Lecture 6 · Classification and clustering
20 questions · lecture page
Q1 · Precisioncatch-up notes
A classifier has TP = 12, FP = 3, FN = 8 and TN = 77. What is its precision?
Show answer
Answer: B. Precision = TP/(TP + FP) = 12/15 = 80%: of the positive predictions, how many were correct. 60% is recall, TP/(TP + FN) = 12/20, and 89% is accuracy, (12 + 77)/100.
Source: Slides p. 26–31
Q2 · Bias–variance diagnosiscatch-up notes
Training accuracy is 96% and test accuracy is 70%. What is the main concern?
Show answer
Answer: A. The model performs very well on the training set but much worse on unseen data, the slide's definition of overfitting (high variance). Underfitting would show poor performance on both sets. PCA uses no labels and is irrelevant here.
Source: Slides p. 9, 15–16
Q3 · Decision trees: Gini impuritycatch-up notes
Two candidate splits have weighted Gini impurities of 0.18 and 0.32. Which split is preferred under this criterion?
Show answer
Answer: D. The tree chooses the split with the lowest (size-weighted) Gini impurity, i.e. the purest child nodes; a pure node has Gini 0. Numerical features can be split at thresholds between sorted values.
Source: Supplementary slides p. 18–20
Q4 · Silhouette plotcatch-up notes
A silhouette plot has its highest average silhouette score at k = 4. What does this suggest?
Show answer
Answer: C. A higher average silhouette means observations sit closer to their own cluster than to other clusters, so the peak (the "prominent peak" on the slide) points to the k worth considering. It is evidence for this dataset and set of variables, not proof of real biological types, and clustering is unsupervised.
Source: Slides p. 63
Q5 · Confusion-matrix output with class imbalance
A model predicts overtraining (Yes/No) in 200 cyclists, and overtraining is rare. Below is the caret confusionMatrix() output for the test set. What is the most appropriate conclusion?
Confusion Matrix and Statistics
Reference
Prediction Yes No
Yes 6 4
No 14 176
Accuracy : 0.91
95% CI : (0.8615, 0.9458)
No Information Rate : 0.9
P-Value [Acc > NIR] : 0.37242
Kappa : 0.3571
Mcnemar's Test P-Value : 0.03389
Sensitivity : 0.3000
Specificity : 0.9778
Pos Pred Value : 0.6000
Neg Pred Value : 0.9263
Prevalence : 0.1000
Detection Rate : 0.0300
Detection Prevalence : 0.0500
Balanced Accuracy : 0.6389
'Positive' Class : YesShow answer
Answer: C. Ninety per cent of cyclists are "No", so always predicting "No" already gives 0.90 accuracy (P-value [Acc > NIR] = 0.37). With "Yes" as the positive class, recall = 6/20 = 0.30, precision = 6/10 = 0.60 and F1 = 0.40, which the slides call the better metric for imbalanced data. Specificity concerns the negatives, not missed overtrained cyclists. The numbers were computed in R.
Source: Slides p. 26–33
Q6 · Accuracy, precision, recall and F1
Injured is the positive class. Which set of values is correct for the confusion matrix below?
Predicted injured Predicted not injured
Actually injured 24 8
Actually not injured 16 152Show answer
Answer: D. Accuracy = (24 + 152)/200 = 0.88, precision = 24/(24 + 16) = 0.60 and recall = 24/(24 + 8) = 0.75. F1 = 2 × 0.60 × 0.75 / (0.60 + 0.75) = 0.667, the harmonic mean, not the arithmetic mean 0.675. Option A swaps precision and recall. Exam note: the supplementary deck's word definitions (precision = TP / "total number of correct predictions"; recall = TP / "total number of positive predictions") are wrong. Use the formulas TP/(TP + FP) and TP/(TP + FN), which are correct on every slide.
Source: Slides p. 26–33; Supplementary slides p. 32–36
Q7 · Fixing high bias
Your model has a training error of 14% and a validation error of 15%. Following the lecture's bias–variance flowchart, which action is most likely to help?
Show answer
Answer: B. High training error with a small gap (1%) means high bias (underfitting). The flowchart says to improve learning of the training set: train longer, use a more complex model, add features or decrease regularization. More data, fewer features and stronger regularization are remedies for high variance. Exam note: slide p. 13's "Total error = bias + variance" is a simplification; the formal decomposition is bias² + variance + irreducible noise, which does not change this diagnosis.
Source: Slides p. 9, 12–16
Q8 · Validation design in caret
This lecture code classifies training sessions as interval or endurance. Which statement is correct?
train_id <- createDataPartition(model_data$training_type, p = 0.7, list = F, times = 1)
train_data <- model_data[ train_id, ]
test_data <- model_data[-train_id, ]
fitControl <- trainControl(method = 'cv', number = 10, classProbs = T)
model <- train(training_type ~ ., data = train_data,
method = 'rf',
trControl = fitControl,
verbose = F,
metric = 'ROC')Show answer
Answer: C. p = 0.7 puts 70% of the sessions in training, and cross-validation runs only inside train_data: the "good practice" of tuning with folds and keeping the test data for final evaluation. In k-fold CV every point is validated once and used for training k − 1 times. LOOCV would need k = N. Aside: metric = 'ROC' normally also needs summaryFunction = twoClassSummary in trainControl(); otherwise caret warns and falls back to Accuracy.
Source: Slides p. 20–22, 24–25
Q9 · Decision trees: splitting a numerical feature
To split on the numerical feature age, you follow the lecture's procedure: sort by age, use the mean of each pair of consecutive ages as a candidate threshold, and compute the weighted Gini impurity for each. Which threshold is chosen, and what is its weighted Gini?
Age: 15 17 20 22 26
Injured: No No Yes Yes NoShow answer
Answer: A. The candidate thresholds are 16, 18.5, 21 and 24. At 18.5 the left node {No, No} has Gini 0 and the right node {Yes, Yes, No} has 1 − (2/3)² − (1/3)² = 0.444, so the weighted Gini is 3/5 × 0.444 = 0.27, the lowest (16 and 24 give 0.40; 21 gives 0.47; verified in R). 20 is an observed age, not a midpoint. Exam note: supplementary p. 20 gives "Gini impurity Age<31.5 = 0.1905". With that slide's own table, the full weighted Gini at 31.5 is 4/7 × 0.375 + 3/7 × 0.444 = 0.405 (0.1905 is only the right-hand term); with the corrected value, the 'likes statistics' split (0.214) would be the better one. Learn the procedure, not that number.
Source: Supplementary slides p. 18–21
Q10 · Elbow and silhouette
You run k-means on standardized fitness data for k = 1–6. Which interpretation is correct?
k 1 2 3 4 5 6
Total within SS 600 330 150 128 112 100
Average silhouette – 0.52 0.68 0.55 0.47 0.41Show answer
Answer: D. The within-cluster sum of squares always decreases as k increases, so its minimum is not a criterion. Look instead for the bend after which extra clusters add little: the drops are 270 and 180, then only 22, 16 and 12. For the silhouette, higher is better, and it peaks at k = 3.
Source: Slides p. 62–64, 82–83
Q11 · PCA: explained variance
A PCA on four standardized fitness tests gives eigenvalues of 12, 5, 2 and 1. Which statement is correct?
Show answer
Answer: A. Proportion of variance = eigenvalue / sum of all eigenvalues: 12/20 = 60% and 5/20 = 25%, so the cumulative total is 85%. Option C ignores PC3 and PC4. Explained variance is not accuracy. Exam note: supplementary p. 60 computes 18/(18 + 4) = "0.81 → 81%" and 4/(18 + 4) = "0.18 → 17%". The correct values are 81.8% and 18.2%, which add up to 100%.
Source: Supplementary slides p. 60–61; Slides p. 71
Q12 · PCA: what the components are
Which statement about principal component analysis (PCA) is correct?
Show answer
Answer: A. PCA centres the data and fits the line through the origin with the smallest orthogonal (not vertical) distances, which is the direction of maximum projected variance. PC2 is perpendicular to it. PCA is unsupervised dimension reduction: it creates new axes and assigns no clusters, although its scores can later be clustered.
Source: Slides p. 70–74; Supplementary slides p. 54–59, 61
Q13 · Case study: clustering cyclists
In van der Zwaard et al. (2019), elite track sprinters, team-pursuit cyclists and road cyclists were clustered with k-means on anthropometry (body size, composition and shape). Which result did the lecture report?
Show answer
Answer: B. The meso cluster held the 6 sprinters, and the tall and short meso-ecto clusters each held 4 pursuit and 5 road cyclists. This confirmed anthropometry-dependent specialization for sprint versus endurance but did not separate pursuit from road. Discipline was examined only after clustering, and the meso cluster showed higher sprint performance while the meso-ecto clusters showed higher endurance performance.
Source: Slides p. 77–79, 85–97
Q14 · Classification versus clustering versus PCAcatch-up notes
Explain the difference between classification and clustering, and give one sport/health example of each. Then explain how PCA differs from both.
Show model answer
Classification is supervised: the model learns from examples with known category labels (the target) and predicts the category for new cases, e.g. injured versus not injured from preseason tests (Rommers et al.). Clustering is unsupervised: no target labels are given, and the algorithm (e.g. k-means) groups observations so that within-cluster similarity is high and between-cluster similarity is low, e.g. grouping cyclists by anthropometry (van der Zwaard). PCA is also unsupervised but assigns neither classes nor clusters. It is a dimension-reduction technique that creates new, perpendicular axes (linear combinations of the variables) capturing as much variance as possible, used for visualization or as input for clustering.
Source: Slides p. 4–7, 56–58, 70–71
How did your answer compare?
Q15 · Case study: injury prediction (Rommers et al. 2020)
Rommers et al. (2020) predicted injuries in 734 elite youth football players from preseason test results. (a) What type of machine learning problem is this, and which model did they use? (b) How did they partition the data, and why? (c) Test precision, recall and F1 were about 85% (training 84/83/83%). What does this tell you, and what can and can't you conclude from the SHAP summary plot?
Show model answer
(a) Supervised classification: the target is a known label (injured yes/no and, in a second model, overuse versus acute), predicted from 29 anthropometric, maturation, coordination and performance features with XGBoost, a decision-tree-based ensemble. (b) 80% of records were used to build (train) the model and 20% as a test set to evaluate it on unseen data, so performance reflects generalization. (c) Precision, recall and F1 describe how well injured players are identified rather than only overall correctness. Similar train and test values (~83–85%) suggest no clear overfitting, while the overuse/acute model performed a bit lower (test ≈ 78%). The SHAP plot ranks features by their impact on the prediction (e.g. age at peak height velocity, body height, leg length) and shows the direction for high and low values. It reflects predictive importance, not proof that these factors cause injury, and the authors call data science a valuable addition but "not the holy grail".
Source: Slides p. 37–54
How did your answer compare?
Q16 · k-means in practice
You want to group 24 cyclists by height (cm), body mass (kg) and sum of skinfolds (mm) using k-means. Describe how you would prepare the inputs, the steps of the algorithm, how you would choose k, and how you would evaluate the result.
Show model answer
Prepare: units and scales affect Euclidean distances, so convert the variables to z-scores so that no variable dominates. Algorithm: choose k; pick k random initial centres; assign each cyclist to the nearest centroid (Euclidean distance); recompute each centroid as the mean of its members; repeat assignment and update until assignments no longer change (or the iteration limit is reached). This minimizes the total within-cluster sum of squares. Because the result depends on the starting points, rerun with different starts and keep the solution with the least within-cluster variation. Choose k with an elbow/scree plot (the bend), a silhouette plot (the peak) and/or NbClust (the k most indices agree on), combined with domain knowledge. Evaluate with a clustplot: are the clusters spherical, of reasonable size and non-overlapping, and how much variance do the two plotted components explain? Only after clustering, compare clusters with disciplines or performance to interpret them.
Source: Slides p. 58–68, 82–84; Supplementary slides p. 41–43
How did your answer compare?
Q17 · Underfittingfrom the recording
Which model did the lecturer describe as the most extreme ('ultimate') example of underfitting?
Show answer
Answer: B. The naive baseline has no complexity at all, so it has maximal bias and performs poorly on both training and test data; your model has to beat it. A model through every point is the overfitting extreme, and 1% vs 11% error indicates overfitting (high variance).
Source: Recording 02:01–04:03
Q18 · Splitting on the targetfrom the recording
You want to classify sessions as extensive or intensive interval training. A random 70/30 holdout split happens to put almost all extensive sessions in the test set. What was the lecturer's advice?
Show answer
Answer: D. The model cannot learn to predict a class that is (almost) absent from the training data, so you arrange or partition on the target so that all classes, or all RPE values, are represented in both sets. Leave-one-out still splits the data (N times) and is meant for very small datasets.
Source: Recording 12:17; Slides p. 25 (code partitions on training_type)
Q19 · Standardising inputs for k-meansfrom the recording
Before k-means clustering of cyclists on height (cm), body mass (kg) and sum of skinfolds (mm), why did the lecturer prefer z-scores over dividing each variable by its maximum?
Show answer
Answer: A. k-means uses Euclidean distances on the raw numbers, so a variable with larger values or a larger spread dominates. Max-normalisation fixes the units but leaves different spreads, whereas z-scoring equalises the SDs. Both are linear rescalings that keep the athletes' order, and neither reduces the number of variables.
Source: Recording 54:23–56:26; Slides p. 66–67
Q20 · Case study: Rommers et al. (2020)from the recording
In Rommers et al. (2020), 50% of players got injured. Test precision, recall and F1 were all 85%, and training values were 84%, 83% and 83%. Which interpretation matches the lecture?
Show answer
Answer: C. With a 50/50 split, accuracy is not inflated by a majority class, and all metrics told the same story. Overfitting would show much better training than test performance, but here the difference is negligible, so 'they made a pretty good model'. Exam note: balance alone does not force precision = recall (that requires FP ≈ FN), but it removes the main reason to distrust accuracy.
Source: Recording 28:45–30:45; Slides p. 47
Lecture 7 · Feature operations
17 questions · lecture page
Q1 · Centring vs standardizationcatch-up notes
Before modelling, you subtract the mean from each predictor but do not divide by its standard deviation. What does this achieve?
Show answer
Answer: A. Subtracting the mean only shifts the location, so a variable measured in years still has a much larger spread than one measured in metres and can still dominate a distance-based model. Dividing by the SD is the extra step of z-scoring that gives mean 0 and SD 1 (slide p. 5).
Source: Slides p. 4–5
Q2 · Missing data: interpolationcatch-up notes
A heart-rate sensor that samples every second has one missing reading between two closely spaced valid readings. Which approach estimates the missing value from the surrounding observations?
Show answer
Answer: C. Interpolation estimates an intermediate value from neighbouring observations in an ordered series, which suits a smooth, densely sampled signal; in the lecture's terms it is a form of imputation ('change a value with the best guess of what it could be'). Dummy encoding, binning and z-scoring transform observed values but do not fill gaps. Exam note: interpolation is not named on slide p. 15, but the official 2024 practice exam lists it as a method for handling missing values.
Source: Slides p. 15; 2024 practice exam Q27
Q3 · Scaling and KNN
In the lecture example, KNN (k = 2) classifies students as male or female from height (training range 1.50–1.96 m) and age (training range 21–38 years). Without any scaling the test accuracy was only 0.46. What is the most likely explanation?
Show answer
Answer: D. KNN relies on distances, so features with a large range or variability dominate (slides p. 4, 12): an age gap of about 17 years outweighs a height gap of about 0.45 m. KNN has no normality assumption. Exam note: the min–max round prints accuracy 0.769 (p. 9) but the summary table shows 0.85 (p. 12); with only 13 test students each error shifts accuracy by about 0.077, so treat the exact ranking of the methods with caution.
Source: Slides p. 4, 6–12
Q4 · Z-score and min–max calculation
In the lecture's training data, height has mean 1.722 m, SD 0.130 m, minimum 1.505 m and maximum 1.955 m. A new student is 1.85 m tall. Using these training statistics, what are the student's z-score and min–max value?
Show answer
Answer: B. z = (1.85 − 1.722) / 0.130 ≈ 0.98 and min–max = (1.85 − 1.505) / (1.955 − 1.505) = 0.345 / 0.450 ≈ 0.77. Option A swaps the two results; option C gives only the raw differences from the mean and the minimum without dividing.
Source: Slides p. 5, 12
Q5 · Standardization vs normalization
A sprint-speed feature contains a few extreme values caused by GPS errors. According to the lecture's cheat sheet, why is min–max normalization more affected by these outliers than z-score standardization?
Show answer
Answer: C. Normalization maps values into a fixed interval defined by the minimum and maximum, so an extreme value sets the range; standardization has no bounding range (slide p. 13). Neither method removes outliers. Exam note: the slides say standardization aims to 'acquire a normal distribution' and is for Gaussian data; strictly, z-scoring is a linear shift-and-rescale to mean 0 and SD 1 that keeps the shape of the distribution (a skewed feature stays skewed), and outliers still influence the mean and SD.
Source: Slides p. 5, 13
Q6 · Log transformation
Training-load values 1, 2, 4 and 8 (each double the previous one) are log₂-transformed. Which statement describes the result?
Show answer
Answer: D. log₂(1, 2, 4, 8) = 0, 1, 2, 3, so every doubling becomes the same step; the slide states that log transformation 'forces' data towards a normal distribution but keeps the fold changes (p. 11). Option A is min–max scaling of the same values. Exam note: slide p. 5 lists log scaling under normalization ('fixed range e.g. 0–1'), but a log does not map to a fixed range, needs positive values, and reduces right-skew rather than guaranteeing normality.
Source: Slides p. 5, 10–11
Q7 · Imputation with mice
The lecture imputed the nhanes data (columns age, bmi, hyp, chl) with the code below. Which statement is correct?
my_imp = mice(input_data, m = 5,
method = c("", "pmm", "logreg", "pmm"),
maxit = 20)
final_clean = complete(my_imp, 5)Show answer
Answer: B. m sets the number of imputed data sets and maxit the number of iterations; complete(my_imp, 5) returns the completed data using imputation 5 (slide p. 16). "logreg" is used for hyp, the binary variable converted to a factor in step 2, and "" means that column is not imputed (age has no missing values). Exam note: the slide picks imputation 5 as 'best matching' the observed BMI mean of 26.56, but the mean of the nine shown imputed values is closest in imputation 4 (≈26.4 vs ≈25.5 for imputation 5); in standard multiple imputation you analyse all m data sets and pool the results instead of choosing one.
Source: Slides p. 16
Q8 · Neural networks vs simple models
A physiotherapist has data from 40 patients and wants a model whose relationship between the predictors and recovery time she can explain to patients. Based on the lecture's comparison, why is a simple model (e.g. linear regression) a better choice here than a neural network?
Show answer
Answer: A. Slide p. 18: a simple model fits a simple function with an interpretable relationship and simple options to control bias and variance, whereas a network captures complex relationships between inputs at the cost of interpretability, more data and more careful bias–variance control. Options B and C reverse the slide: networks do take relationships between inputs into account, while in simple models you must address them yourself.
Source: Slides p. 18
Q9 · Odds ratio calculation
In a cohort of 150 footballers, 18 of the 60 players with a high training load and 9 of the 90 players with a low training load got injured during the season. What is the odds ratio of injury for high versus low load?
Show answer
Answer: D. Odds (high) = 18/42 ≈ 0.43 and odds (low) = 9/81 ≈ 0.11, so OR = (18 × 81) / (42 × 9) ≈ 3.86. 3.0 is the relative risk (0.30/0.10), 0.20 the absolute risk difference, and 0.26 the inverted OR (low vs high).
Source: Slides p. 21–22, 24
Q10 · Probability, odds and odds ratio
Which statement about probability, odds and the odds ratio (OR) is correct?
Show answer
Answer: B. Odds = p/(1 − p) = 0.75/0.25 = 3; odds range from 0 to infinity. OR = 1 means no difference, OR > 1 higher and OR < 1 lower odds in the exposed group (slides p. 21, 24). An OR of 0.5 halves the odds, not the probability. Exam note: slide p. 24 says 'Odds = 0: no difference between groups'; this is a slip (p. 21 correctly gives OR = 1), and odds of 0 simply mean the event never occurs. The same slide's 'Likelihood … e.g. R²' is also loose: likelihood measures how well a model's parameters explain the observed data, which is not the same quantity as R².
Source: Slides p. 21, 24
Q11 · Absolute vs relative risk
A headline says a supplement 'doubles the risk' of a rare side effect. In the study, 2 of 20,000 users and 1 of 20,000 non-users had the side effect. Which statement communicates the result the way the lecture recommends?
Show answer
Answer: A. Absolute risk = 2/20,000 = 0.01% versus 1/20,000 = 0.005%, so RR = 2 while the absolute increase is only 0.005 percentage points. The slides warn never to report a relative risk without the absolute risk, because on its own it is misleading (p. 23).
Source: Slides p. 22–23
Q12 · Missing data strategiescatch-up notes
Give two ways of dealing with missing values and one limitation of each. Explain why replacing every missing value with zero can be misleading.
Show model answer
Examples from the lecture: (1) delete all rows with missing values – simple, but can cause a huge loss of information and, if missingness is related to who the participants are, a biased sample; (2) delete only rows missing a specific key variable – less loss, but a hard, reasoned choice in complex data; (3) imputation, e.g. mice with pmm or logreg – keeps the rows, but the filled-in values are model-based best guesses, not measurements, and using a single completed data set understates uncertainty. Zero is a real value: it can be confused with a true zero (e.g. 0 km run), pulls the mean down and distorts variability and relationships. Zero-filling is only right when missing truly means zero: the lecture's option 3 is exactly that case, a value shown as NaN that in truth is 0 (e.g. no events), which you change to 0. Otherwise impute a best guess or delete with a reason.
Source: Slides p. 15–16; Recording 20:55
How did your answer compare?
Q13 · Odds ratio vs relative risk
On the lecture's odds-ratio slide, 45 patients with heart disease and 45 people without heart disease were asked whether they smoke: 31 of the 45 with heart disease smoked versus 6 of the 45 without. The slide reports OR = 14.4. A journalist writes: 'Smokers are 14 times more likely to get heart disease.' Explain two problems with this sentence and how you would report the result.
Show model answer
(1) 14.4 is an odds ratio, a ratio of odds (p/(1 − p)), not of probabilities; when an outcome is common, the OR is much larger than the corresponding risk ratio, so 'times more likely' overstates the effect. (2) The participants were selected by disease status (45 with and 45 without heart disease, a case-control design), so the 50% share with heart disease is fixed by the design: the absolute risk for smokers, and therefore an RR, cannot be estimated from this table, which is why an OR is used. Report: 'The odds of heart disease were about 14 times higher in smokers than in non-smokers (OR = 31×39 / (6×14) ≈ 14.4).' Following the lecture's rule, add absolute risks from a population-based study before talking about likelihood, and note that this table shows an association, not causation.
Source: Slides p. 21–24
How did your answer compare?
Q14 · Scaling in the lifecyclefrom the recording
According to the lecturer, in which phase of the data science lifecycle does a log transformation of a right-skewed feature belong?
Show answer
Answer: B. The lecturer: rescaling counts as data preparation as long as only the scale changes (e.g. z-score, min–max, metres to centimetres); once the variability/shape changes too, as with a log transformation, it is feature engineering. Option A is the rule for pure rescaling, not for log.
Source: Recording 00:00
Q15 · Neural network nodefrom the recording
A single node of a neural network decides whether an athlete trains today. Inputs: slept well x1 = 1, no muscle soreness x2 = 1, rain x3 = 0. Weights: w1 = 2, w2 = 3, w3 = 4. Threshold = 6. As in the lecture's surfing example, the node outputs 1 if the weighted sum minus the threshold is greater than 0, otherwise 0. What does the node compute?
Show answer
Answer: D. Each input is multiplied by its weight (1×2 + 1×3 + 0×4 = 5) and the threshold is subtracted: 5 − 6 = −1, which is not > 0, so the output is 0. Option A wrongly counts w3 although x3 = 0; in the video's example 1×5 + 0×2 + 1×4 − 3 = 6 > 0 gave output 1.
Source: Recording 35:11–37:23
Q16 · Missing data optionsfrom the recording
In a match-statistics data set the column 'goals scored' shows NaN for every player who did not score in a match. Which approach matches the lecturer's advice?
Show answer
Answer: A. Option 3 on the slide ('change values that are wrong, e.g. 0 instead of NaN → great, do that') was explained with exactly this case: a NaN that in truth is 0 is set to 0. Imputation (C, D) is a best guess, to be used only when the true value cannot be recovered; here it would invent non-zero goals. Deleting rows (B) causes a large loss of information.
Source: Recording 20:55; Slides p. 15
Q17 · Multiple imputationfrom the recording
You imputed missing values with mice using m = 5. According to the lecturer, what could you learn by running your model on each of the five completed data sets?
Show answer
Answer: C. Each of the m imputations gives slightly different values, so comparing model results across them shows how much the imputation matters. The lecturer called this out of scope and allowed picking one completed set for the project; strictly, multiple imputation analyses all m sets and pools the results. Option B confuses m with maxit (iterations), and accuracy cannot reveal which values are 'true'.
Source: Recording 25:01–29:04; Slides p. 16
Lecture 8 · Epidemiological data
16 questions · lecture page
Q1 · Exposure data choicecatch-up notes
You study the association between long-term (10-year) NO₂ exposure and all-cause mortality in adults across the Netherlands. NO₂ is largely traffic-related and varies strongly over short distances. Which exposure data best fit this question?
Show answer
Answer: C. Because NO₂ has large spatial variability (slide p. 21), each participant needs a location-specific, long-term estimate: the lecture uses the average concentration at the residential address, predicted for unmeasured addresses with LUR (p. 22, 33, 40). A national mean or one station ignores spatial contrasts, and one person's short measurement covers neither the cohort nor the 10-year period. (Adapted from the catch-up temperature question to the current NO₂ slides.)
Source: Slides p. 17, 21–22, 33, 40
Q2 · Hazard, exposure and risk
A cyclist commutes for 2 hours a day along a busy road with high NO₂ concentrations. Using the lecture's definitions, which combination is correct?
Show answer
Answer: B. Slide p. 8: a hazard is something that has the potential to harm you, exposure is how long and how much you are subjected to the hazard, and risk is the likelihood of the hazard causing harm. A hazard without exposure (a shark where nobody swims) carries no risk.
Source: Slides p. 8
Q3 · Risk assessment paradigm
In the risk assessment paradigm, which step examines the probability and severity of a health outcome as a function of the dose (exposure)?
Show answer
Answer: D. Hazard characterization describes this dose–response relationship. Hazard identification asks whether the agent can cause harm in an exposed population at all, exposure assessment estimates the levels, types and duration of exposure in the population, and risk characterization asks what the consequences are at current levels and what reducing exposure would gain (slide p. 9).
Source: Slides p. 9
Q4 · Exposure assessment: residential proxy
Why does the lecture's cohort use the average NO₂ concentration at each participant's residential address rather than true personal exposure?
Show answer
Answer: A. Slide p. 22: true exposure across micro-environments is usually not calculated because that is not feasible for large populations, so the average concentration at the residential address is used as a proxy of personal exposure. Being a proxy, it leaves some exposure misclassification (e.g. time at work or commuting).
Source: Slides p. 22
Q5 · Exposure measurement sources
A researcher needs measurements of a pollutant that the national network does not measure, at specific sites near schools. Which approach fits, and what is its main drawback according to the lecture?
Show answer
Answer: C. Designed campaigns let you choose the pollutants and relevant locations, but they are costly and cannot cover large populations (slide p. 24). Routine networks are low-cost, long-term and continuous, but relevant components or sites may be missing and their spatial density may be insufficient (p. 26).
Source: Slides p. 23–26
Q6 · Land-use regression (LUR)
In the case study's land-use regression (LUR) model, what is the outcome variable and what are the predictors?
Show answer
Answer: B. LUR is a statistical (stochastic) exposure model: obtain concentrations at a limited number of sites, choose traffic/land-use predictors, run a regression explaining the measured variability, and apply it to unmeasured locations such as cohort addresses (slides p. 32–40). Option A describes the later health model (Cox); option D describes a deterministic dispersion/chemical transport model (p. 31).
Source: Slides p. 31–40
Q7 · Health data sources
Which health data source does the lecture describe with these pros and cons: 'earlier in the chain; possibly less bias than clinical outcomes' versus 'can be invasive; expensive/time-consuming; instrument/observer effects'?
Show answer
Answer: A. These are the pros and cons of physiological measurements (slide p. 48). Routine databases are large, cheap and consistent but limited by quality, completeness, administrative factors and privacy (p. 46); questionnaires suit large, cheap surveys but depend on wording, standardization and mode of administration (p. 47).
Source: Slides p. 45–48
Q8 · Correlation and confounding
Why does the lecture call the correlation coefficient between NO₂ and mortality 'not an interesting measure of association'?
Show answer
Answer: D. The slides answer 'Confounding!' with the ice-cream/drowning example: both rise with sunny weather, so they correlate without one causing the other (p. 59–61). The Cox model therefore adjusts for smoking, BMI, education, neighbourhood income and other covariates (p. 68).
Source: Slides p. 59–61, 68
Q9 · Hazard ratio interpretation
The adjusted Cox model gives HR = 1.04 (1.03–1.04) per 1 µg/m³ NO₂, p < 0.001. Which interpretation is correct?
Show answer
Answer: A. HR > 1 indicates a positive association (slide p. 63); 1.04 means a 4% higher hazard per unit increase, and the interval excludes 1. It is a relative measure, not a change in absolute probability, and an observational association does not prove causation. Exam note: slide p. 72 phrases this as 'the mortality risk increases by 4%'; strictly it is the hazard. The results figure on p. 70 is expressed per interquartile range (HR ≈ 1.06), a different exposure increment, so always check which unit an HR refers to.
Source: Slides p. 58, 63, 66, 69–72
Q10 · Cox model in R
The lecture's Cox model formula is shown below (shortened). Which statement is correct?
Surv(age_b, age_end, mort) ~ NO2_exposure + strata(sex) +
as.numeric(smoke_dur) + as.factor(smoking) + as.factor(bmi_cat) +
as.factor(edulev) + as.numeric(meaninc_neighbor_2011) + ...Show answer
Answer: C. The analysis table has begin age, end age and mortality (0/1) per participant, linked through the residential coordinates to the average NO₂ exposure (slide p. 64). Surv(...) defines follow-up and the event, NO₂ is the exposure X and the other terms are the confounders C in h(t) = h0(t) × exp(b1X1 + b2C2 + … + bnCn) (p. 67–68). strata(sex) gives men and women separate baseline hazards rather than excluding anyone.
Source: Slides p. 52, 64, 67–68
Q11 · Exposure model vs health modelcatch-up notes
In the NO₂ lecture, explain what the exposure model predicts, what the mortality model investigates, and give one reason the result needs careful interpretation.
Show model answer
The exposure model (land-use regression) is fitted on NO₂ concentrations measured at a limited number of monitoring sites (69 sites, 10 years) with traffic and land-use predictors, and predicts long-term NO₂ concentrations at unmeasured locations, i.e. each participant's residential address. The health model (Cox proportional hazards) uses these predicted exposures, linked via the address, to estimate the association between NO₂ and all-cause mortality during follow-up, adjusted for confounders such as smoking, BMI, education and neighbourhood income, expressed as a hazard ratio. Caveats (one is enough): the residential concentration is only a proxy for personal exposure, and the LUR predictions carry model uncertainty (exposure misclassification); the design is observational, so residual confounding can remain; the HR is relative, not an absolute risk; model assumptions must hold.
Source: Slides p. 22, 33–41, 63–69
How did your answer compare?
Q12 · Statistical significance vs relevance
The model output is HR = 1.04 (1.03–1.04) per µg/m³, p < 0.001, in 33,475 participants; NO₂ exposure had mean 25.41 ± 4.80 µg/m³ (range 11.23–44.11). The lecturer asked: 'Is there a significant effect of NO₂ on mortality?' and 'Is it a substantial effect?' Answer both questions and explain why they differ.
Show model answer
Significant: yes – p < 0.001 and the interval excludes 1, so the association is unlikely to be due to chance alone under the model. With 33,475 participants even small effects become statistically significant, so significance says little about size. Substantial: 4% higher hazard per 1 µg/m³ is small per unit, but people differ by several units: under the model's multiplicative form, a 1-SD contrast (4.8 µg/m³) corresponds to about 1.04^4.8 ≈ 1.21, i.e. roughly 21% higher hazard, and because the whole population is exposed, a small relative effect can add up to many deaths. Relevance also depends on baseline mortality (the absolute risk), uncertainty and possible bias (exposure misclassification, residual confounding). Note: the results figure (p. 70) reports the HR per interquartile range and shows a non-linear curve, so extrapolating 1.04 per unit across the whole range is only an approximation.
Source: Slides p. 65–66, 69–72
How did your answer compare?
Q13 · LUR model validationfrom the recording
After fitting a land-use regression (LUR) model for NO₂ on the 69 Dutch monitoring stations, the guest lecturer described the usual way to check that the model predicts well. Which procedure did she describe?
Show answer
Answer: C. She described a hold-out check: build the regression on part of the stations, predict at the stations left out, and see whether predictions correlate well with the measurements (measuring yourself at predicted locations is possible but time-consuming). A is circular, because the health association cannot validate the exposure model. B judges fit on the training data, which rewards overfitting.
Source: Recording 26:37–28:40; Slides p. 41
Q14 · Deterministic vs statistical exposure modelsfrom the recording
Which statement correctly distinguishes the two approaches the lecture described for predicting NO₂ at addresses without a monitoring station?
Show answer
Answer: A. Deterministic (dispersion or chemical transport) models predict how pollutants travel away from known sources using physical/chemical laws. The stochastic LUR model learns from measured concentrations how land use (roads, green areas, houses) explains differences between stations. B and D swap the two approaches.
Source: Recording 18:25–20:30; Slides p. 31–33
Q15 · Routine health databasesfrom the recording
In the Dutch mortality registry, more deaths are recorded on Mondays than on Sundays. According to the guest lecture, what does this illustrate?
Show answer
Answer: D. The lecturer gave this as an example of the database con 'depends on administrative factors': people do not actually die more on Mondays; the pattern reflects when deaths are registered. B is tempting because weekday NO₂ is higher, but she explicitly said the Monday excess is not real.
Source: Recording 32:46; Slides p. 46
Q16 · Health data categoriesfrom the recording
A researcher cannot access diagnoses because medical records are too sensitive, but can obtain how much each person spent on specific therapies per year. In the lecture's classification of health data (Galetsi et al., 2020), what type of data is this, and how can it be used?
Show answer
Answer: B. Costs and reimbursements of medical expenses are administrative data (slide p. 44). The lecturer noted that, because clinical data are sensitive and often inaccessible, spending on certain therapies can be a proxy for clinical information. Pharmaceutical data cover drug effects and medicine use, real-time patient data come from wearables, and clinical data are endpoints such as incidence or mortality.
Source: Recording 30:43; Slides p. 44
Lecture 9 · User-generated data
17 questions · lecture page
Q1 · Challenges of user-generated datacatch-up notes
A wearable project already has all its software libraries installed and millions of measurements. It lacks verified illness outcomes, and its heart-rate data are noisy. What is its main bottleneck?
Show answer
Answer: A. The lecture names quality control (noisy and missing data) and the lack of reference data as the key challenges: 'you can have unlimited data and still have no use for it' (slides p. 76–78, 97, 134). Large scale is the opportunity of user-generated data, not its main limitation, and more software or data does not create valid labels.
Source: Slides p. 76–78, 81, 97, 134
Q2 · APIcatch-up notes
Which description fits an API (application programming interface) in the context of user-generated data?
Show answer
Answer: C. An API lets software exchange data with another service; the lecture names APIs as a source of reference points (p. 26, 98), and the running study used integrated APIs (e.g. Strava, TrainingPeaks) to obtain workouts (p. 101, 113). Wearables typically report no signal-quality metric at all (p. 81), so option A is wrong.
Source: Slides p. 26, 81, 98, 101, 113
Q3 · From data to knowledge: three steps
A company wants to use phone-camera HRV measurements to study responses to stressors in thousands of users. In which order does the lecture say the work should proceed?
Show answer
Answer: D. The three key steps are (1) validate the technology or know its limitations ('garbage in, garbage out'), (2) deploy and confirm lab-based insights, where data preparation becomes the most important step, and (3) discover new relations and build new products (slides p. 57–61). Confirming a known finding, such as lower rMSSD after high-intensity training (p. 66), builds trust before new relations are explored.
Source: Slides p. 57–68
Q4 · HRV and context
A user's morning rMSSD is well below their normal value and their resting heart rate is higher than usual. Based on the HRV4Training analyses in the lecture, what is the most appropriate interpretation?
Show answer
Answer: B. The large-scale data show lower rMSSD with higher heart rate after high-intensity training, sickness and high alcohol intake (slides p. 66, 69–70); the paper on p. 65 concludes that HRV is a more sensitive but not specific marker of stress. That is why the app collects context such as tags for sickness, travel or menstruation (p. 17, 25). Low-intensity training was followed by a slight rise in rMSSD (p. 66).
Source: Slides p. 17, 25, 65–70
Q5 · Limitations of typical lab studies
A lab study measured the HRV response to training intensity in N = 10 male students. Which limitation does the lecture emphasize?
Show answer
Answer: A. Slides p. 39–50: N = 2–10 is common in sport science, results are valid only for the sample analysed, extending the analysis means running another study (costs, time), and controlled lab conditions may not represent real life. Lab data are typically high quality; the problem is generalizability, not quality.
Source: Slides p. 35–50
Q6 · Noisy data: estimating max heart rate
Without lab tests, an app estimates each user's maximal heart rate from free-living workouts in order to express training intensity relative to it. Which assumption does this approach require?
Show answer
Answer: C. Slide p. 86: we must assume there will be some hard sessions during the monitored period. Implausible readings (outliers beyond ±3 SD, values above 300 bpm) can be removed with simple statistics; for one user the measured maximum was 234 bpm, the cleaned maximum 208 bpm and the age-derived value 180 bpm (p. 87–90). Cleaning cannot recover a hard effort that never happened ('But did they ever go hard?').
Source: Slides p. 83–91
Q7 · Missing data and selection bias
To simplify the analysis, you remove all users with missing workouts from a large app data set. What is the main risk the lecture points out?
Show answer
Answer: B. Slide p. 92: you can sometimes ignore or remove individuals with missing data because there is a lot of data, but this could introduce a bias. Consistent loggers may, for example, be more motivated or fitter than typical users. The lecturer's advice: there is no universal answer, think critically (p. 93).
Source: Slides p. 91–93
Q8 · Reference data
Which question is NOT one of the lecture's considerations about reference data for user-generated research and data products?
Show answer
Answer: D. Slides p. 118–120 ask what the outcomes are, whether we can track them, whether we are asking too much of the user and whether collecting them is ethical. More sensor data cannot replace reference points: 'you can have unlimited data and still have no use for it' (p. 134).
Source: Slides p. 97–98, 118–120, 134
Q9 · Running performance example
In the running-performance study (N ≈ 2,100), feature sets were added step by step to predict 10 km time. Which conclusion matches the reported results?
Show answer
Answer: A. R² rose from 0.33 (resting physiology) to 0.71 (volume), 0.76 (polarization) and 0.87 with previous performance (slides p. 108–111). The target is a time, so this is regression and R² is not accuracy; the error levelled off at about 15 workouts (p. 112). Exam note: slide p. 113 says 'N = 2100, RMSE = 2 minutes (4%)', while the abstract shown on the same slide reports 2,113 individuals and a cross-validated RMSE of 2.6 minutes (4% mean error).
Source: Slides p. 101–113
Q10 · Pandemic example
During the first COVID-19 lockdown (March–May 2020), most respondents in the lecturer's poll (67.6%) expected population resting heart rate to increase. What did data from about 5,500 app users show?
Show answer
Answer: C. The 2020 curve dropped below the 2019 curve during the lockdown (slide p. 126), while travel fell and sleep time rose (p. 127–128), suggesting explanations. These are group averages of app users, so the lecture adds that limitations still apply: who are these users, and what about the individual (p. 129–130)?
Source: Slides p. 124–130
Q11 · Why user-generated data matters
HRV4Training (managing stressors), Bloomlife (labour onset) and Oura (infection detection) were the lecture's examples. What do these applications have in common?
Show answer
Answer: D. Slides p. 23–26: none of these applications was the original goal of the app or wearable; user-generated data made them possible, thanks to contextual data (confounders and additional parameters monitored longitudinally) and reference points (APIs, manually reported outcomes). 'It's not just the hardware anymore' (p. 20–21).
Source: Slides p. 16–26
Q12 · Generalizability and APIscatch-up notes
Give two limitations of using app users' data to draw conclusions about all athletes. Explain how an API might help collect data, and why it does not remove those limitations.
Show model answer
Limitations (any two, explained): (1) selection/representativeness – app users are self-selected (e.g. motivated, tech-savvy, specific sports), so results may not generalize ('who are we talking about?'); (2) data quality – consumer sensors are noisy and usually report no quality metric; (3) missing data – logging is incomplete or selective, and removing incomplete users can introduce bias; (4) group averages may not describe individuals; (5) reference outcomes are often missing. An API lets the app automatically import data from other services (e.g. workouts and race results from Strava or TrainingPeaks), adding context and reference points at scale. But an API only transfers what exists: it does not make the users representative, correct inaccurate sensor readings, fill in missing workouts or verify outcome labels.
Source: Slides p. 26, 81, 92, 98, 101, 129–130
How did your answer compare?
Q13 · Opportunities and limits of user-generated data
Using the lecture's pandemic example (5,500 people, 3 months of data per person, about half a million measurements), explain (a) why user-generated data could reveal this finding when a typical lab study could not, (b) what the contextual data added, and (c) two limitations that still apply.
Show model answer
(a) The app was already collecting data longitudinally in free-living conditions from thousands of people, including the same months in 2019 for comparison, so an unforeseen event could be studied at large scale and in realistic settings; a lab study would have had to be designed and run in advance, with a small N. (b) Contextual variables recorded by the app showed that travel dropped and sleep time increased during lockdown, offering plausible explanations for the lower resting heart rate – explanations, not proven causes. (c) Generalizability: app users are a self-selected group, not the general population ('who are we talking about?'); the curves are group averages, so individuals may have responded differently or even in the opposite direction; also noisy and missing measurements.
Source: Slides p. 14, 54, 122–130
How did your answer compare?
Q14 · Sickness: detection vs predictionfrom the recording
A wearable company claims its ring 'predicts' infections because users' resting heart rate rose before they reported symptoms. How did the guest lecturer judge this kind of claim?
Show answer
Answer: C. He said that if the data show it, 'your body is already sick… you are detecting at best'. Self-reported symptom onset is only an assumed reference, since infection may have happened on any earlier day. Option A confuses 'earlier than a late reference' with forecasting. Option B contradicts the sickness data (higher HR, lower rMSSD).
Source: Recording 62:26 (also 24:46); Slides p. 69, 116–117
Q15 · False positives at scalefrom the recording
During the Q&A, the lecturers discussed an alert feature with a 5% false-positive rate. Why is this a bigger problem for a consumer product with millions of users than for a study with 20 participants?
Show answer
Answer: A. The proportion of false alarms is a property of the classifier. What scales is the count: 5% of 20 is one worried person, while 5% of a million is 50,000. Option B is tempting, but a larger N does not change the rate itself.
Source: Recording 72:51
Q16 · Lifestyle stressors vs trainingfrom the recording
In the HRV4Training free-living data, how did the changes in resting HR and HRV on days after sickness or high alcohol intake compare with those after training, and what did the lecturer conclude?
Show answer
Answer: D. He described the sickness and alcohol changes as two to three times as large as training effects. On the slides, rMSSD falls about 10–12% after sickness or alcohol versus about 3% after high-intensity training, and HR rises about 6% versus under 1%. Hence 'other stressors take over'. Options A–C contradict the plotted directions and sizes.
Source: Recording 22:43; Slides p. 66, 69–70
Q17 · Max heart rate from free-living datafrom the recording
You estimate users' maximal heart rate from months of free-living workouts. According to the guest lecturer, how can the data themselves reveal users who probably never trained hard, and when does this become a serious problem?
Show answer
Answer: B. A narrow spread of session heart rates flags people who always train at one intensity, so their max HR cannot be identified. Whether that matters depends on how frequent it is (3 of 7 vs 3 of a million). Option A describes a measurement artefact, not a lack of hard efforts. Option C fails because he said age-derived max HR can't be trusted given the huge between-person variability.
Source: Recording 48:03, 50:08; Slides p. 87–90
Mixed — across lectures
7 questions
Q1 · Scaling + validation for KNN
You build a KNN classifier for injury (yes/no) from weekly running distance (0–120 km) and sleep duration (5–10 h). Which workflow is most appropriate?
Show answer
Answer: B. KNN relies on distances, so the km feature would dominate unscaled (L7 p. 4, 12), and performance must be judged on data the model has not seen, via a holdout test set or k-fold cross-validation (L6 p. 18–20). Option C lets test information into preprocessing and model selection, so the reported accuracy is optimistic. Exam note: in the lecture's KNN code the test set is scaled with its own statistics (scale(test), normalize(test)); strictly, test data should be transformed with the parameters learned from the training set.
Source: L7 Slides p. 4, 8–12; L6 Slides p. 18–20
Q2 · Imbalanced outcome from wearable data
An app predicts whether a user will tag a day as 'sick'; sickness is tagged on about 3% of days. A model that always predicts 'not sick' reaches 97% accuracy. What should you do?
Show answer
Answer: C. Accuracy is risky with unbalanced data and the F1 score (harmonic mean of precision and recall) is a better metric (L6 p. 27, 33); the always-'not sick' model has recall 0 and therefore F1 = 0. User-generated reference data such as sickness tags are a key challenge: when did the infection start, and how was it determined (L9 p. 97–98, 116–117)?
Source: L6 Slides p. 27, 33; L9 Slides p. 97–98, 115–117
Q3 · Choosing the association measure
Match each design to the most suitable measure of association. (1) 100 injured and 100 uninjured runners are asked which shoe type they use. (2) 5,000 adults are followed for 10 years with different follow-up times, and dates of death are recorded.
Show answer
Answer: A. In (1) the groups are selected by outcome (as in the lecture's 45 vs 45 heart-disease example), so the share injured is fixed by design and risks cannot be estimated; the OR compares odds (L7 p. 21). In (2) time to the event matters, and the Cox model's HR compares hazards at any given time (L8 p. 58, 63). A correlation does not handle confounding (L8 p. 59–60).
Source: L7 Slides p. 21–24; L8 Slides p. 58–60, 63
Q4 · Preprocessing free-living data
In a free-living running data set, 25% of users have missing resting-heart-rate values and the heart-rate data contain impossible readings above 300 bpm. Which preprocessing plan is most defensible?
Show answer
Answer: D. L7 recommends changing wrong values to missing and imputing a best guess, e.g. with mice (p. 15–16); L9 shows readings above 300 bpm are artefacts (p. 87) and warns that removing users with missing data can introduce bias (p. 92). Z-scoring rescales values but does not remove outliers (L7 p. 5, 13).
Source: L7 Slides p. 5, 13, 15–16; L9 Slides p. 87, 92–93
Q5 · Model choice, preprocessing and validation
A running platform wants to predict each user's 10 km race time from workouts imported via an API (weekly volume, average speed, heart rate during training, training-intensity distribution, previous best performance). Data are available for 2,000 users; some resting-heart-rate values are missing. Propose (1) a model type and why, (2) preprocessing steps, (3) how you would validate the model, and (4) one data-related limitation.
Show model answer
(1) The target is numerical, so this is supervised regression, e.g. LASSO or ridge regression (robust to overfitting; LASSO can shrink unnecessary predictors to zero) or a random forest; a simple, interpretable model is preferable to a neural network unless far more data are available and interpretability is not needed. (2) Recode implausible values (e.g. heart rate > 300 bpm) as missing; impute missing resting heart rate (e.g. mice with pmm) or exclude it while checking for selection bias; standardize predictors (z-scores), because penalized regression depends on their scale, using training-set parameters; build features only from training before the target race. (3) Hold out a test set (e.g. 70/30 or 80/20), tune λ with k-fold cross-validation on the training set, and report the error in minutes (RMSE or MAE) and R² on the test set. (4) App users are self-selected; best race times are not standardized tests; consumer sensor data are noisy; associations do not prove that changing training will cause faster times.
Source: L5 Slides p. 41–44, 57–60; L6 Slides p. 18–20; L7 Slides p. 4–5, 15–16, 18; L9 Slides p. 87, 92, 101–113
How did your answer compare?
Q6 · Communicating a wearable-based risk finding
In app data from 4,000 runners, months in which morning rMSSD was more than 10% below the user's baseline had an injury rate of 4%, compared with 2% in other months. A coach asks whether athletes should stop training whenever rMSSD drops. (a) Report the result in relative and absolute terms. (b) Give two reasons why this does not prove that low HRV causes injury. (c) Name one generalizability issue.
Show model answer
(a) RR = 4% / 2% = 2 (twice the risk), but the absolute difference is 2 percentage points, about 2 extra injured months per 100; report both, because a relative risk alone is misleading. (b) Confounding: high-intensity training, sickness or alcohol lower HRV and may also raise injury risk, so a third factor can explain the association (HRV is sensitive but not specific); timing: an emerging injury or illness may itself lower HRV; consumer data are noisy and have missing measurements, and the data are observational. (c) App users are self-selected runners, and a group-level association may not hold for an individual athlete. The finding therefore supports looking at context rather than an automatic 'stop training' rule.
Source: L7 Slides p. 22–23; L8 Slides p. 59–61; L9 Slides p. 65–70, 81, 92, 129–130
How did your answer compare?
Q7 · Pooled vs subgroup resultscatch-up notes
A pooled plot shows a strong positive association. Within one athlete subgroup there is little association; within another there is a positive trend. What should you check before explaining the overall result or making a recommendation?
Show model answer
Compare the subgroups explicitly: their ranges and means on both variables, and their sample sizes. A pooled association can arise largely because the groups differ in level (one group higher on both variables), so group membership acts like a confounder. Analyse the subgroups separately or include group (and an interaction) in the model, and report that the association holds in only one subgroup. Check data quality and outliers within each group. A correlation alone does not show that an intervention would help, so a recommendation for all athletes is not justified; group averages also need not describe individuals.
Source: 2024 practice exam Q35; L8 Slides p. 59–61; L9 Slides p. 130
How did your answer compare?