Home: all lectures
0 XP0 day streak

Data Science in Sport and Health · Papers

Course literature

12 papersquestion · why · approach · findings

What the exam asks about papers. The lecturers' study advice: a smaller number of questions come from the course literature. You do not need every detail — know each paper's question, why it matters, general approach and main finding. Example from the advice sheet: "Rommers et al. used preseason measurements… to predict whether a player would sustain an injury. What type of machine learning problem is this?" → supervised classification.

Lecture 1 – Introduction to data science

Perspective (PNAS)

Blei & Smyth (2017) — Science and data science

Blei, D. M., & Smyth, P. (2017). Science and data science. Proceedings of the National Academy of Sciences, 114(33), 8689–8692. https://doi.org/10.1073/pnas.1702076114 · Open paper

Data science is 'the child of statistics and computer science'. Its essence is the effective combination of three perspectives (statistical, computational and human), applied iteratively with domain experts to answer discipline-specific scientific questions.

Question

What is data science, and why should scientists care about it? The authors discuss data science from three perspectives (statistical, computational and human) and argue that the effective combination of all three is the essence of data science.

Why it matters

Scientists in many fields (genomics, social science text archives, astronomy sky surveys) now have abundant data but cannot yet fully use it. Existing statistical and computational methods are not set up for modern problems such as massive datasets, high-dimensional data, necessarily misspecified models and inferring causality. The authors see this tension as the catalyst for the new label 'data science'.

Approach

This is a conceptual essay without data. It builds on Tukey's (1962) broad notion of 'data analysis'. The statistical perspective covers uncertainty, complex and structured data (Bayesian modelling), high-dimensional data (regularisation, machine learning and deep learning for prediction) and causal inference. The computational perspective covers optimisation (iteratively 'climbing' a likelihood), sampling methods (the bootstrap, Markov chain Monte Carlo) and distributed computing. The human perspective is illustrated with a computational neuroscientist who images mouse neurons and works with a data scientist.

Findings

  • Data science is the 'child of statistics and computer science': it inherits their methods and blends, refocuses and develops them for modern scientific data analysis, in the spirit of Tukey's broad 'data analysis'.
  • Statistical perspective: all datasets involve uncertainty, and statistics is the foundation for reasoning about it. Key subfields are complex/structured data (e.g. Bayesian models), high dimensionality (regularisation; ML such as deep learning for prediction) and causality (correlation vs causation, inference from observational data).
  • Computational perspective: this concerns the algorithmic implementation of methods and the trade-off between statistical accuracy and computational resources (time, memory). Examples are optimisation, sampling (bootstrap for confidence intervals, MCMC for Bayesian posteriors) and distributed computing.
  • Human perspective: data science cannot be fully automated, because applying the tools needs human judgement and deep domain knowledge. The data scientist works iteratively and collaboratively with the domain expert (possibly one person wearing two 'hats'), cycling through preprocessing, exploration, selection, transformation, analysis, interpretation and communication. Reproducibility and data provenance matter.
  • Conclusion: data science is more than the sum of statistics and computer science. It requires weaving both into a larger framework problem by problem, understanding the context of data, taking responsibility for private and public data, and communicating clearly what a dataset can and cannot tell us.

Limitations

  • Own inference: it is an opinion/perspective piece, so its claims are argued, not empirically tested.
  • Own inference: the examples come from genomics, social science, astronomy and neuroscience, not sport or health, so applying it to movement science is left to the reader.
  • Own inference: it names challenges (causality, misspecified models, scale) but stays high-level and gives no concrete procedures for solving them.

Link to the lectures

Slides 49–50 'Data science is …' show the paper and quote its conclusion: data science is more than statistics + computer science; it requires understanding the context of data, appreciating the responsibilities of using private and public data, and clearly communicating what a dataset can and cannot tell us. It frames the Lecture 1 Venn diagram (computer science, math & statistics, domain knowledge; slide 51) and the data science vs statistics slides (53–55).

This is the course's definition of data science (Lecture 1 slides 49–51). It matches the Venn diagram of computer science, statistics and domain knowledge, and its human perspective is the 'domain knowledge' circle. Its iterative cycle (preprocessing → exploration → analysis → interpretation → communication) mirrors the data science lifecycle used from Lecture 3 onwards. Its statistical perspective (uncertainty, causality vs correlation) links to the data science vs statistics slides (Lecture 1 slides 53–55) and the spurious-correlation example (Lecture 4). Its mention of ML for high-dimensional prediction and of the bootstrap foreshadows Lectures 5–7.

Remember

  • Perspective (PNAS 2017): data science is 'the child of statistics and computer science'.
  • Three perspectives: statistical, computational, human; the essence is combining all three.
  • Statistical: uncertainty, complex/structured data, high dimensionality, causality.
  • Computational: trade-off between statistical accuracy and computational resources (time, memory); optimisation, sampling (bootstrap, MCMC), distributed computing.
  • Human: cannot be fully automated; needs domain knowledge and iterative collaboration with domain experts.
  • Data science is a cycle: preprocessing → exploration → selection → transformation → analysis → interpretation → communication.
  • Holistic view: understand the data's context, take responsibility for private/public data, communicate what the data can and cannot tell us.

Practice

Q1 · Blei & Smyth – what data science is

According to Blei & Smyth (2017), what is the essence of data science?

Show answer

Answer: C. The authors argue that each perspective is critical but that combining all three is the essence of data science. Option B contradicts their human perspective: data science cannot be fully automated.

Source: Paper p. 8689–8690

Q2 · Blei & Smyth – computational perspective

Which statement best describes what Blei & Smyth call the computational perspective of data science?

Show answer

Answer: A. Computational thinking is about the algorithmic implementation and the accuracy-versus-resources trade-off. Option B is the statistical perspective and option C is the human perspective.

Source: Paper p. 8690–8691

Q3 · Blei & Smyth – human perspective

Why, according to Blei & Smyth (2017), can data science not be fully automated?

Show answer

Answer: D. The human perspective holds that understanding the domain, choosing data, exploring, selecting models and communicating results need judgement and disciplinary knowledge. Option B is false: the authors explicitly mention 'necessarily misspecified models'.

Source: Paper p. 8690–8691

Q4 · Blei & Smyth – why data science emerged

What do Blei & Smyth see as the catalyst for the new label 'data science'?

Show answer

Answer: B. They describe a tension: scientists have abundant data, but classical methods cannot fully exploit it computationally or statistically, and this tension gave rise to 'data science'. Option A is the opposite of their premise of abundant data.

Source: Paper p. 8689–8690

Q5 · Blei & Smyth – three perspectives applied

Blei & Smyth (2017) describe data science as combining statistical, computational and human perspectives. Using a sport or health example of your choice (e.g. a club analysing GPS and injury data of its players), explain what each perspective contributes.

Show model answer

Statistical: models the data while accounting for uncertainty, handles complex or high-dimensional data, and asks whether relationships are causal or only correlational (e.g. does high load cause injury?). Computational: implements the analysis efficiently, balancing accuracy against time and memory (e.g. processing large GPS streams, using optimisation or resampling such as the bootstrap). Human: brings domain knowledge from coaches and sport scientists to choose relevant data and features, interpret results and communicate what the data can and cannot say. The work is iterative and collaborative, and the essence is combining all three.

Source: Paper p. 8690–8691; Lecture 1 Slides p. 49–51

Lecture 1 – Introduction to data science

Perspective (non-technical overview; no new data)

Chmait & Westerbeek (2021) — Artificial Intelligence and Machine Learning in Sport Research: An Introduction for Non-data Scientists

Chmait, N., & Westerbeek, H. (2021). Artificial intelligence and machine learning in sport research: An introduction for non-data scientists. Frontiers in Sports and Active Living, 3, 682287. https://doi.org/10.3389/fspor.2021.682287 · Open paper

A non-technical perspective that explains to sport professionals how the machine-learning paradigm differs from traditional rule-based analytics, illustrates supervised, unsupervised and reinforcement learning with sport examples, and discusses what AI means for the future of sport (privacy, data ownership, ethics).

Question

The authors want to give sport business professionals, coaches, policy makers and other non-technical audiences a high-level overview of the AI and ML approaches used for sport performance and sport business problems. They note that for many non-experts the link between AI and sport is still 'fuzzy' and the reasons for adopting ML are unclear. They also discuss how AI could shape sport in the coming years.

Why it matters

AI adoption in sport is accelerating (from Moneyball and SABRmetrics in baseball to injury modelling, player tracking and ticket pricing), driven by more computing power and more data. Decision-makers need to understand how data scientists think so they can talk to them about approach and method without diving into the technical details.

Approach

No data are collected; it is a narrative perspective. The authors first summarise earlier work on AI in sport (expert systems, artificial neural networks and deep learning, evolutionary computation, Bayesian approaches, the Beal et al. 2019 survey) and list four application areas. They then contrast traditional analytics (rules + data → program → answers) with the ML paradigm (data + known answers → algorithm derives the rules, which are then tested on new, unseen data), writing ML prediction as a function f(w1·i1, …, wn·in) = y. Next they give three hypothetical sport examples: supervised injury prediction in Australian football, unsupervised fan segmentation with K-means clustering, and reinforcement learning (Q-learning, SARSA) for fantasy sport and game AI, plus a short note on genetic/evolutionary algorithms. They end with perspectives on the future of AI in sport.

Findings

  • The ML paradigm reverses traditional analytics: instead of programming the rules, you feed the algorithm data together with the known answers, it learns the rules, and those rules are validated by testing accuracy on new (unseen) data.
  • Supervised learning: the outcome is known in the historical data, e.g. whether each Australian football player got injured and missed the next match, alongside match load, metres run, warm-up and tackles. The model is trained on this, tested on unseen data and tuned until accuracy is acceptable (their illustrative figure: 70%); neural networks, decision trees or regression models can all do this.
  • Unsupervised learning: no labels are available; the algorithm discovers previously unnoticed patterns. Example: a football club segmenting stadium attendees by gender, age, postcode or income with K-means, where the groups are formed from the data, although the number of groups can be set in advance.
  • Reinforcement learning: an agent acts in a (simulated) environment, receives rewards or penalties, and builds a policy that maximises cumulative reward. Examples are fantasy sport team selection and game AI (chess, Go, poker, StarCraft).
  • AI/ML in sport spans four areas: game activity/analytics, talent identification and acquisition, training and coaching, and fan- and business-focused applications.
  • Future: AI is unlikely to fully replace coaches and human experts. The main barriers to wider, 360-degree analyses are proprietary (non-public) player, team and commercial data and privacy rules. Data ownership and ethics (e.g. teams' injury-prediction models versus players' right to their own data) will become contested.

Limitations

  • It is a perspective, not an empirical study or systematic review; the authors say they do not aim to discuss the literature comprehensively, so the examples are selective.
  • The supervised, unsupervised and reinforcement learning examples are hypothetical, so no real model performance is reported (the 70% accuracy is only illustrative).
  • Own inference: it is deliberately non-technical and simplified. It does not cover validation procedures (e.g. cross-validation, overfitting), and its unsupervised example loosely says the algorithm will 'classify' patrons, although clustering finds groups without labels.

Link to the lectures

Slide 52 'Examples of data science in sports research' reproduces the paper's four application areas: game analytics, talent identification, training & coaching, fan & business; the slide credit is misspelled 'Chmait & Westerpoort 2021'. Its supervised / unsupervised / reinforcement learning examples also preview the 'Types of machine learning' slides in Lecture 5 (slides 8–16).

This paper is the course's plain-language map of ML types. The injury example is supervised learning (labelled outcome; a binary injured/not outcome makes it classification), the fan segmentation is unsupervised clustering with K-means (Lecture 6), and the fantasy-sport/game example is reinforcement learning (Lecture 5 slides 14–15). Its emphasis on testing the learned rules on new, unseen data is the core idea behind train/test splits and cross-validation (Lecture 5). Lecture 1 slide 52 uses its four application areas of data science in sport.

Remember

  • Perspective paper for non-data-scientists; no new data.
  • Traditional analytics: rules + data → answers. ML: data + known answers → algorithm learns the rules → validated on unseen data.
  • Supervised = outcome/label known in the training data (e.g. injured yes/no) → predict for new cases.
  • Unsupervised = no labels; the algorithm discovers groups (e.g. K-means fan segmentation; k can be set in advance).
  • Reinforcement = an agent learns by trial and error to maximise cumulative reward (Q-learning; games, fantasy sport).
  • Four application areas: game analytics, talent identification, training & coaching, fan & business.
  • Barriers: proprietary data, privacy, data ownership; AI supports rather than replaces coaches.

Practice

Q1 · Chmait & Westerbeek – supervised learning

Chmait & Westerbeek (2021) illustrate supervised learning with a model that predicts muscle-strain injuries in Australian football players from match load, metres run, warm-up and tackles. What makes this a supervised learning problem?

Show answer

Answer: B. The authors stress that the key difference is that the outcome (injury or not) is known in the historical data used for training. Option A describes reinforcement learning, and option D is the opposite of what they describe: the model is tested on new, unseen data.

Source: Paper p. 4

Q2 · Chmait & Westerbeek – unsupervised learning

A football club has data on its stadium attendees (age, gender, postcode, income) but no predefined customer categories. It wants to discover groups of similar fans to target its marketing. Which approach do Chmait & Westerbeek describe for this?

Show answer

Answer: C. Fan segmentation is their unsupervised example: no labels exist, and K-means forms the groups from the data (the number of groups may be set beforehand). Supervised classification would need known fan-type labels for training.

Source: Paper p. 4–5

Q3 · Chmait & Westerbeek – ML paradigm

According to Chmait & Westerbeek (2021), how does the machine-learning paradigm differ from traditional sports analytics?

Show answer

Answer: D. This reversal (data + answers → rules, then validation on unseen data) is the paper's central explanation of the ML paradigm. Option A swaps the two, and option C contradicts the authors' point that the learned rules must be validated on new data.

Source: Paper p. 3–4

Q4 · Chmait & Westerbeek – future and barriers

Chmait & Westerbeek note that no study yet provides a '360-degree' analysis of an athlete's total value (performance plus business value such as ticket sales). What main obstacle do they identify?

Show answer

Answer: A. The authors name proprietary, non-public data and privacy and data-ownership issues as the main challenges. Option D contradicts their description of clubs routinely recording patron characteristics.

Source: Paper p. 6

Q5 · Chmait & Westerbeek – three types of ML

Using the sport examples from Chmait & Westerbeek (2021), explain the difference between supervised, unsupervised and reinforcement learning.

Show model answer

Supervised learning uses historical data in which both the inputs and the correct output are known, e.g. player load data plus whether each player got injured. The model learns the input–output mapping and is then tested on new, unseen data. Unsupervised learning has no outcome labels: the algorithm discovers structure in the inputs, e.g. K-means grouping stadium attendees into fan segments by demographics. Reinforcement learning has an agent that takes actions in a (simulated) environment, receives rewards or penalties, and learns a policy that maximises cumulative reward, e.g. selecting fantasy-sport teams or playing chess or Go.

Source: Paper p. 4–5

Lecture 3 – Data in sport and health

Original research (cross-sectional)

van der Zwaard et al. (2018) — Critical determinants of combined sprint and endurance performance: an integrative analysis from muscle fiber to the human body

van der Zwaard, S., van der Laarse, W. J., Weide, G., Bloemers, F. W., Hofmijster, M. J., Levels, K., Noordhof, D. A., de Koning, J. J., de Ruiter, C. J., & Jaspers, R. T. (2018). Critical determinants of combined sprint and endurance performance: an integrative analysis from muscle fiber to the human body. The FASEB Journal, 32(4), 2110–2123. https://doi.org/10.1096/fj.201700827R · Open paper

In 28 cyclists, sprint and endurance performance (normalised to lean body mass^(2/3)) were inversely related. Correlations and stepwise multiple regression on determinants from muscle fibre to whole body showed what explains sprint, endurance and their combination, pointing to long fascicles and capillarisation as training targets for combining both.

Question

Which physiological determinants, measured at several biological levels (whole body, blood, muscle and muscle fibre), explain sprint performance, endurance performance and the ability to combine both? The underlying problem is that muscle fibre size (important for power) and fibre oxidative capacity (important for endurance) are inversely related, so improving both at once is hard.

Why it matters

Many sports need both sprint and endurance, but earlier studies looked at only one of the two and mostly at whole-body determinants. Adaptations for endurance and for peak power interfere with each other, partly at the molecular level. Knowing the critical determinants gives targets for training, and the two tests plus a physiological profile can support talent identification and individualised training.

Approach

Cross-sectional study of 28 cyclists (14 road, 8 team pursuit, 6 track sprint) competing at national to Olympic level, except 4 amateur road cyclists. Sprint performance was 1-s peak power in a 30-s Wingate test, and endurance performance was mean power in a 15-km time trial. Both were normalised to lean body mass^(2/3) to remove the effect of body size. Determinants included whole-body VO2 (performance VO2, VO2max, thresholds) and gross efficiency, blood values (Hb, Hct, MCHC…), muscle oxygenation (near-infrared spectroscopy), isometric knee-extension torque, 3D ultrasound of the vastus lateralis (volume, fascicle length, PCSA, pennation angle) and muscle-biopsy histochemistry (fibre type, fibre size, SDH oxidative capacity, capillaries, myoglobin). Combined sprint + endurance performance was quantified as each cyclist's perpendicular residual to the Deming regression line between normalised sprint and endurance performance (Deming regression allows for error in both variables). Explained variance was assessed with Pearson correlations and stepwise multiple regression, in which a predictor was added only if it gave a significant R² change (P < 0.05).

Findings

  • Normalised sprint and endurance performance were inversely related (r = −0.66, P < 0.001; per kg body mass r = −0.44), so combining both is difficult.
  • Sprint: the percentage of fast-type fibres and vastus lateralis muscle volume together explained 65% of the variance in normalised peak power (R² = 0.65).
  • Endurance: performance VO2, MCHC (an oxygen-transport blood measure) and muscle oxygenation together explained 92% of the variance in normalised time-trial power (R² = 0.92). Performance VO2 itself was explained by fibre oxidative capacity, myoglobin × capillary-to-fibre ratio and, negatively, PCSA (R² = 0.67), so a large muscle cross-section (hypertrophy) works against oxygen consumption.
  • Combined sprint + endurance performance was explained by gross efficiency and performance VO2, and likely by muscle volume and fascicle length (P = 0.056 and P = 0.059). Across the whole group, gross efficiency plus muscle volume explained 30%.
  • At fibre level, fibre size and fibre oxidative capacity were inversely (hyperbolically) related (r = −0.50), but cyclists who combined relatively large fibres with high oxidative capacity had more capillaries.
  • Implication: long fascicles (rather than a large PCSA) and capillarisation are suggested as training targets for improving sprint and endurance at the same time.

Limitations

  • Cross-sectional design: it shows associations, not the longitudinal effects of training. The authors call for future longitudinal training studies, so the training targets are hypotheses rather than proven effects.
  • Unmeasured determinants (e.g. heart pumping capacity, lung diffusion, neuromuscular recruitment, muscle metabolites, tendon properties, glycogen) may explain additional variance, and the track sprinters' suboptimal PCSA/fascicle-length arrangement may have increased the unexplained variance.
  • Own inference: small, heterogeneous sample (n = 28; 6 sprinters) with many candidate predictors in stepwise regression. No train/test split or cross-validation is reported, so the high R² values are in-sample and may be optimistic (overfitting).

Link to the lectures

The main case study, slides 14–30. The lecture presents the problem (muscle size drives sprint and oxidative capacity drives endurance, but fibre size and oxidative capacity are inversely related), the aim, the 28 cyclists, the performance tests (Wingate, 15-km time trial) and the determinants measured from whole body down to muscle fibre. It then shows the inverse sprint–endurance relationship (r = −0.66, with perpendicular residuals drawn as arrows), the explained-variance bar chart (Fig. 3), the 67% model for performance VO2, and the conclusion that long muscle fibres and capillaries/myoglobin are training targets. It illustrates combining many data categories (performance, physical tests, imaging, biopsy, blood).

The Lecture 3 case study for data acquisition: one question answered by combining data types from many sources and levels (performance tests, VO2, blood, ultrasound imaging, biopsies). Analytically it is a regression problem with a labelled continuous outcome (normalised power), so it is conceptually supervised, but it is used to explain variance (R² = explained variance, cf. Lecture 4), not to predict new athletes with a validated model (contrast hold-out/cross-validation in Lecture 5). Normalising to lean body mass^(2/3) and building a combined-performance score from Deming-regression residuals are examples of feature engineering / target construction. Contrast with van der Zwaard et al. (2019) in Lecture 6, which clusters cyclists without labels (k-means, unsupervised).

Remember

  • Aim: identify critical determinants of combined sprint and endurance performance, from whole body down to muscle fibre.
  • 28 cyclists (track sprint, team pursuit, road); sprint = Wingate peak power; endurance = 15-km time-trial mean power; both normalised to lean body mass^(2/3).
  • Sprint and endurance are inversely related (r = −0.66): combining both is difficult.
  • Combined performance = perpendicular residual to the Deming regression line (allows for error in both variables).
  • Analysis: Pearson correlations + stepwise multiple regression (explained variance, R²).
  • Sprint ← % fast fibres + muscle volume (65%); endurance ← performance VO2 + MCHC + muscle oxygenation (92%); combined ← gross efficiency + performance VO2 (likely muscle volume, fascicle length).
  • Take-home: long fascicles/fibres and capillarisation are training targets; cross-sectional design, so no causal training effects shown.

Practice

Q1 · van der Zwaard 2018 – main finding

Which statement best summarises the main findings of van der Zwaard et al. (2018) on sprint and endurance performance in cyclists?

Show answer

Answer: D. The study found r = −0.66 between normalised sprint and endurance performance; fast fibres + muscle volume explained 65% of sprint variance, and performance VO2 + MCHC + muscle oxygenation explained 92% of endurance variance. Fibre size and oxidative capacity were inversely related (r = −0.50), so option B is wrong, and Hb did not determine endurance or combined performance (MCHC did), so option C is wrong.

Source: Paper p. 2110 (abstract), 2113–2116; Lecture 3 Slides p. 23, 27–29

Q2 · van der Zwaard 2018 – combined performance measure

How did van der Zwaard et al. (2018) quantify how well each cyclist combined sprint and endurance performance?

Show answer

Answer: B. Combined performance was the perpendicular residual to the Deming regression line, which allows for error in both variables; it expresses how far a cyclist lies above or below the group's sprint–endurance trade-off. Exam note: the paper's label 'POpeak + POTT' is notation for this residual-based score, not an arithmetic sum.

Source: Paper p. 2111–2112, Fig. 2 (p. 2116); Lecture 3 Slides p. 23

Q3 · van der Zwaard 2018 – analysis approach

Which description best fits the analysis approach of van der Zwaard et al. (2018)?

Show answer

Answer: A. The outcomes (normalised peak power, time-trial power and the combined score) are continuous, and their determinants were identified by correlation and stepwise multiple regression (a predictor entered if it gave a significant R² change). Option B describes the later van der Zwaard et al. (2019) clustering study, and option D is wrong because the design is cross-sectional.

Source: Paper p. 2113 (Statistical analysis)

Q4 · van der Zwaard 2018 – limitation

Van der Zwaard et al. (2018) suggest fascicle length and capillarisation as training targets. What is the main reason to treat this recommendation with caution?

Show answer

Answer: C. The authors state that the cross-sectional design does not describe longitudinal training effects and call for longitudinal training studies. The other options are factually wrong: the cyclists were mostly national to Olympic level, endurance was measured with a 15-km time trial, and endurance R² was 0.92.

Source: Paper p. 2120–2121 (Implications, Limitations)

Q5 · van der Zwaard 2018 – regression vs validated prediction

Van der Zwaard et al. (2018) used stepwise multiple regression with 28 cyclists and many candidate physiological determinants, reporting high explained variance (e.g. R² = 0.92 for endurance). From a data-science perspective, what type of problem is this, and what would you need to do to know whether such a model predicts the performance of new cyclists?

Show model answer

It is a regression problem with a labelled, continuous outcome (normalised power output), so it is conceptually supervised. The paper uses it to explain variance in the same sample: its statistical analysis reports no train/test split or cross-validation. With few participants and many candidate predictors, stepwise selection can overfit, so in-sample R² may overestimate performance on new athletes. To check generalisation you would use a hold-out test set or k-fold cross-validation, or validate in a new cohort, and you could use regularised regression (ridge/LASSO), which is more robust to overfitting in small datasets. In addition, the cross-sectional design shows associations, not causal training effects.

Source: Paper p. 2113 (Statistical analysis), p. 2121 (Limitations); Lecture 5 Slides p. 43–44, 49, 60

Lecture 5 – Machine learning: regression

original research (original investigation)

Jaspers et al. (2018) — Relationships Between the External and Internal Training Load in Professional Soccer: What Can We Learn From Machine Learning?

Jaspers, A., Op De Beéck, T., Brink, M. S., Frencken, W. G. P., Staes, F., Davis, J. J., & Helsen, W. F. (2018). Relationships between the external and internal training load in professional soccer: What can we learn from machine learning? International Journal of Sports Physiology and Performance, 13(5), 625–630. https://doi.org/10.1123/ijspp.2017-0299 · Open paper

Using two seasons of GPS and accelerometer data from a Dutch top-league soccer team, LASSO and neural-network models predicted players' session RPE better than a naive mean-RPE baseline; LASSO beat the neural network, decelerations emerged as important load indicators, and one group model worked as well as or better than per-player models.

Question

Can machine learning predict a player's perceived exertion (session RPE, the internal load) from a large set of external load indicators (ELIs) measured with GPS and accelerometers? The authors also asked which ELIs contribute most to RPE in soccer, and whether individual (per-player) models beat a group model, i.e. whether there are meaningful differences between players.

Why it matters

Clubs monitor training load to optimise fitness and reduce injury risk, and the same external load can cause a different internal load in different players, so understanding the external–internal link improves load management. Earlier work used traditional statistics on a few hand-picked ELIs, and the only ML study (Australian football) may not transfer to soccer, which has different physical demands.

Approach

38 professional outfield players from a team in the highest Dutch league were followed over two seasons (2014–15 and 2015–16); only training sessions were used (matches, recovery and rehabilitation sessions excluded), 5917 sessions in total. External load came from 10-Hz GPS and 100-Hz accelerometers: 67 ELIs on duration, distance, speed, accelerations/decelerations, PlayerLoad and repeated high-intensity efforts. The target was session RPE on the modified Borg CR-10 scale, reported about 30 minutes after training. Three predictors were compared on a held-out test set using mean absolute error (MAE): an artificial neural network (ANN), LASSO (linear regression with a penalty that shrinks many coefficients to zero) and a naive baseline that always predicts the mean RPE. Experiment 1 trained group models on season 1 and tested them on season 2 (a temporal split, so many test players were unseen); experiment 2 used the first 75% of each season for training and the last 25% for testing, comparing group and individual models. LASSO-based importance scores ranked the ELIs.

Findings

  • Both ML models beat the naive baseline. Experiment 1 MAE: LASSO 0.80, ANN 1.09, baseline 1.14 RPE points; LASSO cut the error by 29.8% (small effect, d = 0.44), the ANN only trivially (d = 0.06).
  • LASSO was more accurate than the ANN (its MAE was 26.6% lower than the ANN's in experiment 1).
  • The most important ELIs included high-intensity acceleration efforts (>3.5 m·s−2), repeated high-intensity efforts, high-magnitude deceleration distance, running distance at 12–20 km·h−1, PlayerLoad and session duration; decelerations (eccentric, potentially muscle-damaging load) were the new finding.
  • Group models were as good as or better than individual models in both seasons (e.g. season 1 LASSO: group 0.79 vs individual 0.81 vs baseline 0.99; season 2: 0.85 vs 0.85 vs 1.11), unlike the earlier Australian-football study. The authors explain this by more data for group models (>2000 vs <100 sessions) and a more homogeneous soccer squad.
  • Practical message: ML can support experts in choosing key ELIs for monitoring, and a group model can predict RPE for new, transferred or youth players who have little data of their own.

Limitations

  • Training sessions only; matches were excluded, and there may be too few matches per season for ML (paper).
  • Only global RPE was predicted, and inputs did not include differential RPE, wellness, recovery or psychosocial factors (paper).
  • Individual models had fewer than 100 sessions per player; with 100–150 sessions per season and frequent transfers, more per-player data is unrealistic (paper).
  • The improvement over the baseline was modest (small effect; MAE about 0.8 on a 0–10 scale), and data came from one team with one GPS system, so comparing with or generalising to other teams is hard (partly paper, partly own inference).

Link to the lectures

Canvas lists the paper with Lecture 5; it is a worked example of supervised regression using LASSO, MAE, a naive baseline and hold-out testing, all covered on the slides – the slides themselves do not discuss the paper.

A clear example of supervised regression: a numeric target (RPE) predicted from many labelled predictors (ELIs). It shows LASSO from Lecture 5 (the absolute-value penalty can push slopes to zero, giving automatic feature selection and robustness with many or correlated features and small samples) against a flexible but hard-to-interpret neural network. It also illustrates evaluating with MAE in the target's own units, comparing against a naive baseline, and hold-out validation that respects time order (train on the earlier season, test on the later one). Group vs individual models illustrates the trade-off between data quantity and personalisation.

Remember

  • Supervised regression: predict session RPE (internal load) from 67 GPS and accelerometer ELIs (external load).
  • 38 professional soccer players, 2 seasons, training sessions only.
  • Models: ANN, LASSO and a naive baseline that always predicts the mean RPE; metric is MAE (lower is better).
  • Both ML models beat the baseline, and LASSO was best (season 1 → season 2 MAE: 0.80 vs ANN 1.09 vs baseline 1.14).
  • LASSO sets coefficients to zero, which selects ELIs and keeps the model interpretable; decelerations were newly identified as important.
  • Group models ≥ individual models (more data, homogeneous squad), so a group model can be used for new players.
  • Temporal validation: train on season 1, test on season 2; experiment 2 used the first 75% vs last 25% of each season.

Practice

Q1 · Jaspers 2018 – type of ML task

Jaspers et al. (2018) used GPS- and accelerometer-derived external load indicators to predict each player's session RPE (modified Borg CR-10) for future training sessions. How is this machine-learning task best described?

Show answer

Answer: C. Each session's reported RPE is the known outcome (label), and the models output a numeric RPE that is evaluated with mean absolute error, so this is supervised regression. No sessions were grouped into classes or clusters, and LASSO selects existing ELIs rather than building principal components.

Source: Paper p. 1 (journal p. 625, abstract), p. 3 (p. 627)

Q2 · Jaspers 2018 – why LASSO

Why was LASSO an attractive choice for Jaspers et al. compared with the artificial neural network?

Show answer

Answer: A. The paper describes LASSO as linear regression with a mechanism that biases many coefficients to 0. This gives feature selection, interpretability and robustness to multicollinearity and small samples. The ANN is the nonlinear, hard-to-interpret option. Option D describes Ridge regression, and LASSO was still evaluated on an independent test set.

Source: Paper p. 3 (p. 627); Lecture 5 slides p. 58, 60

Q3 · Jaspers 2018 – MAE and baseline

In the first experiment (trained on season 1, tested on season 2), the LASSO group model had a mean absolute error (MAE) of 0.80, versus 1.14 for the naive baseline. Which interpretation is correct?

Show answer

Answer: D. MAE is the average |reported − predicted| RPE, in RPE units, and lower is better. The baseline always predicts the mean RPE and gives a realistic upper limit for the error, so the 29.8% reduction shows the ELIs carry real information about RPE. MAE is not a share of explained variance; that would be R².

Source: Paper p. 3 (p. 627), Table 2; Lecture 5 slides p. 41

Q4 · Jaspers 2018 – group vs individual models

Jaspers et al. compared group models (trained on all players) with individual models (one per player). What did they find, and how did they explain it?

Show answer

Answer: B. In both seasons the group models matched or beat the individual models (e.g. season 1 LASSO: 0.79 group vs 0.81 individual). The authors attribute this to >2000 vs <100 training sessions and less variation between players in soccer than in Australian football. All 8 learned models beat the baseline, so C is wrong, and D contradicts Table 4.

Source: Paper p. 4 (p. 628), Table 4; p. 5 (p. 629)

Q5 · Jaspers 2018 – validation design

Jaspers et al. trained their group models on season 1, tested them on season 2, and compared every model with a naive baseline. Explain why each of these two design choices makes the evaluation more convincing.

Show model answer

Splitting by season gives an independent hold-out test set and keeps the time order: the model is built on past data and tested on future sessions (and even on players it never saw). That matches how the model would be used in practice and avoids the optimistic estimates you get when test sessions resemble the training data. The naive baseline ignores all ELIs and always predicts the mean RPE, so it sets a realistic upper limit for the error. A model is only useful if its MAE is clearly lower than the baseline's (LASSO 0.80 vs 1.14), which shows that external load really adds predictive information.

Source: Paper p. 3 (p. 627), Table 2

Lecture 6 – Classification and clustering

original research (proof-of-concept study)

Biswas et al. (2015) — Recognizing upper limb movements with wrist worn inertial sensors using k-means clustering classification

Biswas, D., Cranny, A., Gupta, N., Maharatna, K., Achner, J., Klemke, J., Jöbges, M., & Ortmann, S. (2015). Recognizing upper limb movements with wrist worn inertial sensors using k-means clustering classification. Human Movement Science, 40, 59–76. https://doi.org/10.1016/j.humov.2014.11.013 · Open paper

Using a single wrist-worn accelerometer and gyroscope, the authors built k-means clusters from each person's labelled lab trials and assigned new movements to the nearest cluster. This recognised three basic arm movements during a 'making-a-cup-of-tea' task with about 88% (accelerometer) and 83% (gyroscope) accuracy in healthy people and about 70% and 66% in stroke patients, beating LDA and SVM classifiers that used the same features.

Question

Can three elementary forearm movements be recognised from one wrist-worn inertial sensor during an unconstrained daily activity? The movements are reach and retrieve (extension/flexion), lifting a cup to the mouth (rotation about the elbow) and pouring or (un)locking (rotation about the forearm's long axis). The aim is a tool that counts how often a stroke (or cerebral palsy) patient uses the impaired arm for these movements outside the clinic.

Why it matters

Counting specific movements of the impaired arm in daily life could track rehabilitation progress remotely, because such movements should become more frequent as motor function improves. Most earlier activity-recognition work focused on gross activities and postures (sitting, walking, running), often in the lab. Fine upper-limb movements in real-life settings, with few sensors and low computing cost, had received little attention.

Approach

Four healthy men and four stroke patients wore a Shimmer sensor on the back of the wrist (tri-axial accelerometer and gyroscope at 50 Hz; magnetometer not used). Training data were many labelled repetitions of each movement under controlled conditions (e.g. 480 trials per healthy subject). Test data came from the same person performing a 20-step 'making-a-cup-of-tea' activity list on a separate day, segmented using the researcher's annotations. The signals were band-pass filtered, and 10 time-domain features (e.g. standard deviation, RMS, entropy, jerk, peaks, kurtosis, skewness) were computed per axis, giving 30 per sensor. The features were ranked with a class-separability measure (scatter matrices) and added one by one (sequential forward selection). For each feature set, k-means with k = 3 (regularised Mahalanobis distance) was run within 10 runs of 10-fold cross-validation on the training data. Each cluster took the label of the nearest class mean, and new movements were assigned to the nearest cluster centroid (minimum-distance classifier, Euclidean or Mahalanobis). Models were personalised (one per subject and sensor) and compared with LDA and SVM classifiers using the same features.

Findings

  • Healthy subjects: overall accuracy 61–100% (mean 88%) with accelerometer data and 60–94% (mean 83%) with gyroscope data.
  • Stroke patients: lower accuracy, 40–88% (mean 70%) with the accelerometer and 40–83% (mean 66%) with the gyroscope, reflecting more variable and less repeatable movements. For example, one patient early in rehabilitation, tested on the non-dominant arm, reached only 40%.
  • The number and type of best features differed per person (from 2 to 30). Healthy subjects shared top features (e.g. accelerometer stddev_y and rms_y); stroke patients had little overlap.
  • For stroke patients the Mahalanobis-distance classifier often worked better than the Euclidean one, because their clusters had large variance in some directions (elongated clusters).
  • LDA (mean 45–53%) and SVM (mean 50–68%) with the same features performed worse than the clustering approach. Accelerometer and gyroscope data sometimes made up for each other's weak spots.

Limitations

  • Very small sample (4 healthy, 4 stroke): this is a proof of concept, and the authors plan a larger sample and other sensor locations.
  • Test movements were segmented by hand from the researcher's annotations; automatic segmentation, which real-world use needs, was not addressed (paper).
  • Personalised models need a long labelled training session for each person (up to 480 trials) and re-training as the patient recovers (partly paper, partly own inference).
  • Own inference: the LDA/SVM comparison used features selected for the clustering method, and the 25% cluster-size threshold was chosen because it 'produced the best results'. Both choices may favour the proposed method.

Link to the lectures

Listed as reference literature on the Canvas Lecture 6 page but not discussed on the slides; Canvas calls it 'Biswas et al 2018', but the paper is from 2015.

This paper shows that the line between supervised and unsupervised methods can blur. k-means, normally unsupervised, is used here for supervised multi-class classification, because labelled training data fix k = 3 and give each cluster its label. It illustrates k-means mechanics (centroids, iterative reassignment, a squared-distance cost) and its assumption of spherical, Euclidean clusters versus the Mahalanobis distance. It also covers feature engineering and selection from sensor signals, 10-fold cross-validation on training data with a separate, more realistic test set, and evaluation with a confusion matrix, per-class sensitivity (recall) and overall accuracy, benchmarked against LDA and SVM.

Remember

  • Aim: recognise 3 arm movements (reach/retrieve, lift cup to mouth, pour/(un)lock) with one wrist inertial sensor, to monitor stroke rehabilitation.
  • Supervised use of k-means: labelled training trials, k = 3, each cluster labelled by the nearest class mean, and new movements assigned to the nearest centroid (minimum-distance classifier).
  • Training on controlled lab trials; testing during an unconstrained 'making-a-cup-of-tea' task (more realistic, usually lower accuracy).
  • Features: 10 time-domain features × 3 axes per sensor, ranked and chosen with sequential forward selection plus 10 × 10-fold CV; one personalised model per subject.
  • Accuracy: healthy about 88% (accelerometer) and 83% (gyroscope); stroke about 70% and 66%.
  • Mahalanobis distance helps for stroke patients, whose clusters are elongated and variable.
  • Beat LDA and SVM using the same features, but only 4 + 4 participants.

Practice

Q1 · Biswas 2015 – supervised use of k-means

k-means is usually an unsupervised method, yet Biswas et al. (2015) describe a 'k-means clustering classification'. How is their approach best described?

Show answer

Answer: A. The authors note that k-means can be used for supervised learning when the training labels are known. They knew the movement labels, set k = 3, labelled each cluster by the closest class mean, and classified test movements with a minimum-distance classifier checked against annotations. Nothing was discovered without labels, and counting movements is the intended future use, not a regression target.

Source: Paper p. 4 (journal p. 62), p. 10 (p. 68)

Q2 · Biswas 2015 – Euclidean vs Mahalanobis distance

For the stroke patients, the Mahalanobis-distance classifier was often more effective than the Euclidean-distance classifier. What reason do the authors give?

Show answer

Answer: C. The Mahalanobis distance uses each cluster's covariance matrix and suits ellipsoidal clusters, and the paper links its advantage in stroke patients to their highly variable movement profiles. Option A is the opposite of what Mahalanobis distance does, and Euclidean distance works in any number of dimensions.

Source: Paper p. 9 (p. 67), p. 10 (p. 68), p. 14 (p. 72); Lecture 6 slides p. 59–60, 68

Q3 · Biswas 2015 – features

How did Biswas et al. turn the raw wrist accelerometer and gyroscope signals into input for their classifier?

Show answer

Answer: B. The paper computes ten time-domain features per axis and ranks them with a scatter-matrix separability measure, deliberately avoiding the more computationally demanding PCA or RELIEF. It then uses sequential forward selection with 10 runs of 10-fold cross-validation to choose the features for each subject and sensor.

Source: Paper p. 7 (p. 65), p. 8–9 (p. 66–67)

Q4 · Biswas 2015 – main results

Which statement summarises the main results of Biswas et al. (2015)?

Show answer

Answer: D. Tables 2–3 and 5–6 give means of 88% and 83% for healthy subjects and 70% and 66% for stroke patients, with large differences between individuals (e.g. 40% for one patient early in rehabilitation). LDA and SVM averaged roughly 45–68%. With only 4 + 4 participants, the authors present the work as a proof of concept.

Source: Paper p. 2 (p. 60, abstract), p. 11–15 (p. 69–73), Tables 2–11

Q5 · Biswas 2015 – training and testing design

Biswas et al. trained each person's model on movements performed in a controlled setting and tested it on movements made during an unconstrained 'making-a-cup-of-tea' task on another day. Explain why they chose this design, and what it means for the accuracy you should expect compared with training and testing on lab data only.

Show model answer

The goal is to detect prescribed arm movements in real daily life (home rehabilitation monitoring), so the test data should resemble real use: free posture, speed and context, recorded on a different day. Testing on such out-of-lab data is more realistic and checks that the model generalises beyond the conditions it was trained in. It usually gives lower accuracy than lab-only train/test designs, which tend to be optimistic because training and test data are so similar. Cross-validation (10 × 10-fold) was used only within the training data, to choose the features.

Source: Paper p. 4 (p. 62), p. 5 (p. 63), p. 10 (p. 68)

Lecture 6 – Classification and clustering

original research (prospective cohort study over one season)

Rommers et al. (2020) — A Machine Learning Approach to Assess Injury Risk in Elite Youth Football Players

Rommers, N., Rössler, R., Verhagen, E., Vandecasteele, F., Verstockt, S., Vaeyens, R., Lenoir, M., D'Hondt, E., & Witvrouw, E. (2020). A machine learning approach to assess injury risk in elite youth football players. Medicine & Science in Sports & Exercise, 52(8), 1745–1751. https://doi.org/10.1249/MSS.0000000000002305 · Open paper

In 734 elite youth footballers, an XGBoost classifier trained on preseason anthropometric, motor-coordination and fitness tests predicted who would be injured during the season (85% precision, recall and F1 on held-out players). With slightly lower performance (78%), a similar model classified whether the first injury was overuse or acute.

Question

Can a machine-learning model use simple preseason screening tests (anthropometry, maturity, motor coordination, physical fitness, football experience) to predict which elite youth players will be injured in the coming season? Second, can a similar model classify the injury as overuse or acute?

Why it matters

Elite youth football has a high injury risk, but clubs lack the time and money for extensive screening. A model built on field tests that clubs already run would help them target prevention. Earlier studies using traditional statistics found no link between single motor-performance measures and injury; because injuries are multifactorial, a method that combines many variables and their interactions may reveal risk profiles.

Approach

This prospective study followed 734 male players (U10–U15, mean age 11.7 years) from seven Belgian premier-league academies through the 2017–18 season. In August, a test battery measured anthropometry and maturity (height, sitting height, leg length, weight, body fat, predicted age at peak height velocity), motor coordination (KTK3, dribbling test) and physical fitness (jumps, sprints, agility t-test, sit-and-reach, Yo-Yo IR1, curl-ups). Medical staff registered injuries (first injury per player, overuse or acute), and coaches recorded exposure. Two binary classifiers were built with XGBoost (gradient-boosted decision trees): injured vs not injured, and overuse vs acute. Data were randomly split into 80% training and 20% test sets; on the training data, cross-validation and grid search tuned the hyperparameters. Precision, recall and F1 were reported for both sets, and SHAP summary plots showed which variables drove the predictions.

Findings

  • 368 of 734 players (50%) sustained at least one injury. The text reports 173 overuse and 195 acute first injuries (47% vs 53%).
  • Injury model: precision, recall and F1 of 84%, 83% and 83% on the training data, and 85%, 85% and 85% on the held-out test set (147 players).
  • Overuse vs acute model: 82%, 82% and 81% on the training data, and 78% for all three on the test set (74 injuries), slightly lower than the injury model.
  • The top 5 SHAP predictors of injury were a higher predicted age at peak height velocity, greater body height, longer legs, a lower body-fat percentage and standing broad jump performance – mostly anthropometric and maturity measures. Better sit-and-reach performance slightly increased predicted risk.
  • For overuse vs acute, a lower age at PHV, higher sitting height, slower t-test and lower moving-sideways (KTK3) score pointed towards overuse injuries.
  • Conclusion: when combined in an ML model, preseason tests can flag high-risk players and the likely injury type, so academies can focus limited resources on them.

Limitations

  • Only each player's first injury was analysed, so repeated injuries are ignored (paper).
  • Tests were taken once, in preseason, even though anthropometry and fitness change during a season of growth and training (paper; the authors suggest retesting every few months).
  • The model was tested on a random 20% of the same cohort and season and still needs validation in a different, e.g. international, cohort (paper). Own inference: performance on new clubs or seasons may be lower.
  • Own inference: SHAP shows which associations the model uses, not causes. The top predictors largely reflect age and maturity (older, taller players get injured more), which the authors link to injury incidence rising with age.

Link to the lectures

The lecture's worked example of a supervised ML classifier, slides p. 37–54.

This is the textbook supervised-classification case of Lecture 6: labelled data from a prospective cohort and a binary target (injured yes/no, then overuse vs acute). It follows the lecture workflow: a feature set (29 preseason variables), a tree-ensemble method (XGBoost boosting), an 80/20 hold-out split with cross-validation and grid search on the training set, and evaluation with confusion-matrix metrics (precision, recall, F1). Comparing training and test scores links to the bias–variance trade-off and overfitting, and SHAP links to model interpretation and feature importance. The classes are nearly balanced (about 50% injured), so the metrics are not inflated by a large majority class.

Remember

  • Aim: predict injury (yes/no) during one season from preseason tests in elite youth football; a second model classifies overuse vs acute injuries.
  • Supervised binary classification with XGBoost, a boosted ensemble of decision trees.
  • 734 players (U10–U15) from 7 Belgian academies over 1 season; about 50% got injured.
  • Random 80/20 train/test split; cross-validation and grid search were used on the training data for tuning.
  • Test set: 85% precision, recall and F1 for injury; 78% for overuse vs acute.
  • SHAP is used to interpret the model; the top predictors were mainly anthropometric and maturity measures (age at PHV, height, leg length, fat %) plus standing broad jump.
  • Limitations: only the first injury was used, tests were done only in preseason, and there was no external validation.

Practice

Q1 · Rommers 2020 – XGBoost

Rommers et al. built their injury model with XGBoost (extreme gradient boosting). Which description of this algorithm is correct?

Show answer

Answer: D. The paper describes boosting as combining a set of weak learners to improve prediction accuracy and calls the method boosted tree models. The slides call XGBoost a 'decision-tree based ensemble technique'. It uses the injury labels, so it is supervised rather than clustering, and it is not a linear penalised regression.

Source: Paper p. 3 (journal p. 1747), p. 4 (p. 1748); Lecture 6 slides p. 44

Q2 · Rommers 2020 – training vs test performance

For the injured vs uninjured model, Rommers et al. report precision, recall and F1 of 84%, 83% and 83% on the training data (80%) and 85%, 85% and 85% on the held-out test data (20%). What does this comparison mainly tell you?

Show answer

Answer: A. Overfitting shows up as clearly better performance on training data than on test data. Here the held-out scores are similar (even slightly higher), so the model generalised to unseen players from the same cohort. Hyperparameters were tuned with cross-validation and grid search on the training data; only the final model was run on the 20% test set.

Source: Paper p. 3 (p. 1747), p. 4 (p. 1748); Lecture 6 slides p. 9, 47

Q3 · Rommers 2020 – precision

The injury model reached a test-set precision of 85%. Which statement correctly describes what this precision expresses?

Show answer

Answer: C. Precision = TP / (TP + FP): the share of predicted positives that are true positives. The authors explain the training precision the same way: only 16% of the players the model marked as injured stayed uninjured. Option A describes recall (sensitivity), and B describes specificity.

Source: Paper p. 3 (p. 1747), p. 4 (p. 1748); Lecture 6 slides p. 29, 31

Q4 · Rommers 2020 – main predictors

According to the SHAP analysis, which variables mattered most for predicting injury in Rommers et al.?

Show answer

Answer: B. The five most important predictors were age at PHV, body height, leg length, body-fat percentage and standing broad jump; the authors link this to injury incidence rising with age and maturity. All predictors came from the preseason test battery; GPS load and previous injuries were not used, so A and D are impossible.

Source: Paper p. 4 (p. 1748), p. 5 (p. 1749, Fig. 1A); Lecture 6 slides p. 47

Q5 · Rommers 2020 – why ML works here, and a limitation

Earlier studies using traditional statistics found no association between preseason performance tests and injury in youth football, yet Rommers et al. predicted injury quite well from similar tests. Explain why a machine-learning classifier could succeed here, and name one limitation that restricts how far the result can be trusted in practice.

Show model answer

Injuries are multifactorial, so a single test rarely has a strong link with injury. A boosted tree ensemble (XGBoost) combines many variables (29 preseason measures) and their interactions and is optimised for classification accuracy rather than testing single risk factors, so it can detect risk profiles. Its performance was checked on a held-out 20% test set (85% precision, recall and F1). Possible limitations: only the first injury per player was used; tests were done only once in preseason, although growth and training change the measures; and the model was tested on players from the same cohort and season, so it still needs validation in an independent (e.g. international) cohort.

Source: Paper p. 1 (p. 1745), p. 4 (p. 1748), p. 6 (p. 1750)

Lecture 6 – Classification and clustering

original research (observational, cross-sectional)

van der Zwaard et al. (2019) — Anthropometric Clusters of Competitive Cyclists and Their Sprint and Endurance Performance

van der Zwaard, S., de Ruiter, C. J., Jaspers, R. T., & de Koning, J. J. (2019). Anthropometric clusters of competitive cyclists and their sprint and endurance performance. Frontiers in Physiology, 10, 1276. https://doi.org/10.3389/fphys.2019.01276 · Open paper

K-means clustering of 24 competitive cyclists on body shape, size and composition gave three anthropometric clusters. A mesomorphic cluster held all the sprinters and had the best sprint performance; short and tall meso-ectomorphic clusters mixed pursuit and road cyclists and had the best endurance performance. Anthropometry-dependent specialisation was therefore only partly confirmed.

Question

If cyclists are grouped purely by their individual anthropometry, using several dimensions at once and without their discipline labels, do the groups match the disciplines they compete in (sprint, pursuit, road)? The authors also asked whether the clusters differ in sprint and endurance performance, and how anthropometry relates to performance.

Why it matters

Athletes are thought to specialise in disciplines that suit their physique. However, anthropometry is usually reported as group averages for predefined disciplines, which can hide individuals with a different build. Grouping athletes directly on their own body measures gives an unbiased test of anthropometry-dependent specialisation and helps interpret performance in light of each athlete's anthropometry.

Approach

24 male cyclists (sprint, pursuit and road; national to Olympic level) visited the lab three times. Anthropometry was measured following the ISAK protocol. Clustering used nine variables, three per dimension: body shape (Heath–Carter endo-, meso- and ectomorphy), body size (height, weight, body surface area) and body composition (sum of eight skinfolds, body fat %, skeletal muscle mass %). The variables were standardised to z-scores, and k-means (Hartigan–Wong algorithm in R) minimised the total within-cluster sum of squared Euclidean distances. k = 3 was chosen with the elbow criterion, BIC (mclust) and the NbClust indices, and 25 random starts plus 1000 repeated runs checked stability. Sprint performance (Wingate 1-s peak power, squat-jump power) and endurance performance (15-km time-trial mean power, VO2peak), both relative to body mass, were then compared between clusters (ANOVA or Kruskal–Wallis) and correlated with anthropometry.

Findings

  • Three stable clusters emerged: mesomorphic (n = 6), short meso-ectomorphic (n = 9) and tall meso-ectomorphic (n = 9).
  • All six sprinters fell in the mesomorphic cluster (heavier, larger girths, less lean). Pursuit and road cyclists were spread evenly over the two meso-ectomorphic clusters (4 pursuit + 5 road in each), which differed mainly in body size.
  • The mesomorphic cluster had higher sprint performance (Wingate peak power, squat-jump power; p < 0.05) and lower endurance performance (time-trial power, VO2peak; p < 0.001) than both meso-ectomorphic clusters.
  • Across all cyclists, endurance performance went with a lean, ectomorphic build with small girths and a small frontal area. Sprint performance went with larger skinfolds and girths and a low frontal area per body mass.
  • The hypothesis of separate clusters for sprint, pursuit and road was only partly confirmed. Clustering separated sprint-type from endurance-type cyclists, and short from tall endurance cyclists (resembling all-terrain vs flat-terrain road cyclists), but not pursuit from road cyclists.

Limitations

  • Small sample (24 male cyclists; clusters of 6, 9 and 9) from only three disciplines; the authors call for a larger sample covering all cycling specialisations.
  • k must be chosen in advance, and the result depends on which variables are entered. The authors handled this with validity criteria and an equal number of variables (three) per dimension.
  • Own inference: the design is observational and cross-sectional. It shows that build and discipline go together but cannot tell whether athletes chose a discipline because of their build or their build adapted through training.
  • Own inference: the clusters were formed and then compared on the same small data set. Repeated runs showed algorithmic stability, but no independent sample checked that the clusters replicate.

Link to the lectures

The lecture's worked example of unsupervised ML with k-means, slides p. 76–96.

A clear example of unsupervised learning: the algorithm gets no target or labels, and the discipline labels are used only afterwards to interpret the clusters. It illustrates the k-means theory from Lecture 6: minimising the within-cluster sum of squared Euclidean distances, iterating cluster assignment and centroid updates, and using random starts. It also covers practical steps: standardising inputs to z-scores, choosing k with an elbow (scree) plot and NbClust, and judging the result with a clusplot (spherical, similar-sized, non-overlapping clusters). The clusplot shows the clusters on two dimensions from a dimension reduction that explains 85% of the variability, which links to PCA.

Remember

  • Unsupervised ML: k-means clustering of 24 competitive cyclists on 9 anthropometric variables (body shape, size and composition).
  • Discipline labels (sprint, pursuit, road) were NOT used for clustering, only to check the clusters afterwards.
  • Variables were converted to z-scores before clustering, so every variable has the same scale and weight in the Euclidean distance.
  • k = 3 was chosen with the elbow (scree) plot, BIC (mclust) and NbClust; 25 random starts were used, and the clusters were stable over 1000 runs.
  • Result: the mesomorphic cluster contained all the sprinters and had the best sprint performance; the short and tall meso-ectomorphic clusters mixed pursuit and road cyclists and had the best endurance performance.
  • Specialisation was only partly confirmed: sprint and endurance types were separated, pursuit and road cyclists were not.
  • k-means assumes spherical clusters of similar size; the authors checked this with a clusplot.

Practice

Q1 · vd Zwaard 2019 – why unsupervised

Why did van der Zwaard et al. (2019) use unsupervised clustering instead of training a supervised classifier to predict each cyclist's discipline?

Show answer

Answer: B. The authors state that supervised methods need predefined labels (e.g. discipline), whereas clustering divides athletes on anthropometry alone. The disciplines were known but deliberately kept out of the clustering, so the match between clusters and disciplines was an unbiased test. Multiclass supervised classifiers do exist, and the data were standardised to z-scores before clustering.

Source: Paper p. 2, p. 9; Lecture 6 slides p. 77

Q2 · vd Zwaard 2019 – standardisation

Before running k-means, van der Zwaard et al. converted all anthropometric variables to z-scores. What was the purpose of this step?

Show answer

Answer: D. k-means assigns each athlete to the nearest centroid by Euclidean distance, so variables on large scales would otherwise dominate. Z-scoring gives every variable equal weight; the authors note that 'no single feature is more important than another'. The 2-D plot came from a separate dimension reduction for the clusplot, and k was chosen with validity criteria.

Source: Paper p. 4, p. 9; Lecture 6 slides p. 66–67

Q3 · vd Zwaard 2019 – choosing k

k-means requires the number of clusters (k) to be set in advance. How did van der Zwaard et al. arrive at k = 3?

Show answer

Answer: A. The paper states that the number of clusters was determined with the elbow criterion, BIC (mclust package) and NbClust validity criteria. That this equals the number of disciplines is a property of the data, not of the method. Using the labels to pick k (B, D) would undermine the label-free approach, and k-means never chooses k itself.

Source: Paper p. 4, p. 9; Lecture 6 slides p. 62–64, 82–83

Q4 · vd Zwaard 2019 – main result

What was the main clustering result in van der Zwaard et al. (2019)?

Show answer

Answer: C. All six sprinters were in the mesomorphic cluster, which also had higher sprint and lower endurance performance. Pursuit and road cyclists (4 + 5 in each) were split by body size into short and tall meso-ectomorphic clusters. Anthropometry-dependent specialisation was therefore only partly confirmed: sprint and endurance types were separated, pursuit and road cyclists were not.

Source: Paper p. 1 (abstract), p. 4–5 (Table 1), p. 6, p. 7 (Fig. 3); Lecture 6 slides p. 89–91, 95

Q5 · vd Zwaard 2019 – making k-means trustworthy

k-means clustering has several assumptions and practical pitfalls. Describe two things van der Zwaard et al. did to make their three-cluster solution trustworthy, and explain why each matters.

Show model answer

Any two of the following. (1) All variables were standardised to z-scores, so measurement scale and units do not decide the Euclidean distances. (2) Each dimension (body shape, size and composition) had the same number of variables (three), so each contributes equally to the clusters. (3) k was chosen with validity criteria (elbow/scree plot, BIC, NbClust), because k must be set in advance and a poor k gives meaningless clusters. (4) The analysis used 25 random starting partitions and was repeated 1000 times with the same clusters every time; this matters because k-means depends on its random initial centroids. (5) A clusplot was used to check that the clusters were roughly spherical, similar in size and clearly separated, which are k-means assumptions.

Source: Paper p. 4, p. 9; Lecture 6 slides p. 62–68, 82–84

Lecture 7 – Feature operations

original research

Mundt et al. (2020) — Prediction of lower limb joint angles and moments during gait using artificial neural networks

Mundt, M., Thomsen, W., Witter, T., Koeppe, A., David, S., Bamer, F., Potthast, W., & Markert, B. (2020). Prediction of lower limb joint angles and moments during gait using artificial neural networks. Medical & Biological Engineering & Computing, 58, 211–225. https://doi.org/10.1007/s11517-019-02061-3 · Open paper

IMU signals simulated from an existing optical motion-capture database were used to train a feedforward and an LSTM neural network that predict 3D hip, knee and ankle joint angles and moments during walking; both worked well and the simpler feedforward network was slightly more accurate.

Question

Can artificial neural networks predict lower-limb joint angles (kinematics) and joint moments (kinetics) during gait directly from IMU accelerations and angular rates, so that IMU-based gait analysis no longer depends on magnetometer data and calibration postures? Secondly, how does a standard feedforward network (FFNN) compare with a recurrent long short-term memory network (LSTM) on this task?

Why it matters

Lab gait analysis with cameras and force plates is expensive and limited to a small capture volume. IMUs work outside the lab, but they measure kinematics only indirectly, depend on error-prone calibration and on magnetometers that are disturbed by local magnetic fields, and cannot measure joint moments at all. If a network can map raw IMU signals to gold-standard joint angles and moments, valid kinematic and kinetic gait analysis becomes possible in clinics and in the field.

Approach

The authors reused a database of level-walking trials (people without movement disorders plus knee-arthroplasty patients, self-selected speeds 0.8–2.0 m/s) recorded with optical motion capture and force plates. Joint angles were computed from the markers and joint moments with inverse dynamics (AnyBody), with moments normalised to body weight and height; these were the targets. From the marker trajectories they simulated what virtual IMUs (accelerometer + gyroscope, no magnetometer) would have measured, giving inputs with exactly matched gold-standard labels. They trained a feedforward network (sees the whole gait cycle or stance phase at once) and an LSTM recurrent network (processes the sequence with an internal memory), using mean-squared-error loss, the Adam optimiser, early stopping and dropout, and a train/validation/test split by participant so nobody appeared in more than one set. Starting from a baseline (IMU data, right steps only, trained 10 times), they tested adding full or partial anthropometric data, using left and right steps, and data augmentation (randomly rotating the virtual sensor), keeping only changes that beat the baseline by more than two standard deviations. Accuracy was reported as the correlation coefficient r plus RMSE (angles, degrees) or range-normalised RMSE (moments), because r alone ignores an offset between curves.

Findings

  • Both networks predicted joint angles and moments well: the abstract reports mean correlations above 0.98 in the sagittal plane and above 0.80 in the smaller (frontal and transverse) planes, and the Discussion reports RMSE below 4° (angles) and nRMSE below 25% (moments) for all joints and planes. Caveat: a few individual values are lower, e.g. knee frontal-plane angle r = 0.79 (FFNN) and 0.68 (LSTM) in Table 4.
  • The feedforward network outperformed the LSTM for both kinematics and kinetics in most joints and planes (baseline joint-angle RMSE 1.60° vs 2.15°; baseline moment nRMSE 12.2% vs 14.8%). The authors explain this by the FFNN seeing all time steps at once, while the LSTM can only carry information forward from past time steps.
  • Sagittal-plane motion was predicted best. For joint angles the transverse plane was worst; for joint moments the knee rotation and frontal-plane ankle (inversion) moments were worst.
  • For joint moments, using both left and right steps plus data augmentation (factor 3) improved the models; for joint angles only adding partial anthropometric data (height, weight, segment lengths, pelvis width) helped.
  • Conclusion: the FFNN suits offline analysis of fully recorded data, the LSTM is the more natural choice for real-time use and recordings of any length, and a convolutional network with a small delay might combine both strengths.

Limitations

  • Authors: the IMU inputs were simulated, so they lack the soft-tissue movement real IMUs experience; performance on real sensor data was not tested.
  • Authors: all virtual sensors had the same position and orientation; robustness to variation in sensor placement still needs to be tested.
  • Authors: outliers came from the manual dataset split (test-set patterns not present in training), the small share of knee-arthroplasty patients, the weakness of optical systems for non-sagittal knee motion, and errors in the inverse-dynamics processing.
  • Own inference: only level walking at self-selected speeds was studied, so the models may not transfer to running, stairs or other tasks.

Link to the lectures

Neural-network part: a worked example of a feedforward vs a recurrent (LSTM) network for supervised regression; also shows feature scaling, adding features and data augmentation in practice.

Lecture 7 contrasts simple models with neural networks: networks fit complex functions and capture relationships between inputs, but are less interpretable, need larger amounts of data and need more attention to bias–variance control (Slides p. 18). Mundt et al. illustrate each point: they create more data by simulation and augmentation, limit overfitting with early stopping, dropout and a participant-wise train/validation/test split, and compare a non-recurrent (FFNN) with a recurrent (LSTM) network. The paper also shows feature operations from Part I of the lecture: joint moments are normalised to body weight and height, and anthropometric features are added and kept only if they improve accuracy. Note: the Lecture 7 slides do not discuss this paper explicitly; they give only a general neural-network slide (p. 18) and a video slide (p. 19).

Remember

  • Aim: predict 3D hip, knee and ankle joint angles and moments during gait from IMU data (acceleration + angular rate only) with neural networks, removing the need for magnetometers and calibration postures.
  • IMU inputs were simulated from an existing optical motion-capture database; the targets came from optical kinematics and inverse dynamics (gold standard).
  • Supervised regression on time-series data; feedforward network (FFNN) vs recurrent LSTM.
  • FFNN beat the LSTM in most joints and planes; both had r > 0.98 in the sagittal plane and > 0.80 in the minor planes (abstract).
  • Train, validation and test sets were split by participant to prevent leakage; early stopping and dropout limited overfitting.
  • r alone ignores offsets, so RMSE (angles) and nRMSE (moments) were also reported.
  • Main limitation: simulated data lack soft-tissue artefacts, so real IMU data still need testing.

Practice

Q1 · Type of ML problem

Mundt et al. (2020) trained neural networks to predict hip, knee and ankle joint angles and joint moments (continuous curves) from IMU signals, using optical motion capture and inverse dynamics as the known target values. What type of machine learning problem is this?

Show answer

Answer: C. The outputs are continuous quantities (degrees, normalised moments) and every input has a known gold-standard label, so this is supervised regression. Classification would require discrete categories (e.g., patient vs healthy), which the networks did not output.

Source: Paper p. 211 (PDF p. 1, Abstract); pp. 212–213 (PDF pp. 2–3, Sect. 1 and 2.1)

Q2 · Why simulate IMU data

Instead of recording new IMU data, Mundt et al. simulated accelerometer and gyroscope signals from an existing optical motion-capture database. What was the main advantage of this approach?

Show answer

Answer: A. The authors stress the larger availability of optoelectronic data and its high-quality ground truth. Option B is the reverse of their own limitation: simulated data do not include the soft-tissue movement real IMUs experience.

Source: Paper p. 211 (PDF p. 1, Introduction); p. 219 (PDF p. 9, Discussion)

Q3 · Main finding: FFNN vs LSTM

Which statement best summarises the main result of Mundt et al. (2020)?

Show answer

Answer: B. Tables 3–4 and the Discussion show the FFNN outperformed the LSTM for both kinematics and kinetics, with r > 0.98 in the sagittal plane (abstract). Option A is tempting but wrong: the authors argue the FFNN wins because it sees all time steps at once, whereas the LSTM only carries information from past steps.

Source: Paper p. 211 (PDF p. 1, Abstract); pp. 217, 219, 222 (PDF pp. 7, 9, 12; Tables 3–4, Discussion)

Q4 · Validation and data leakage

To make sure the reported accuracy is realistic for new people, how did Mundt et al. set up training and evaluation?

Show answer

Answer: D. The authors chose 'the strictest testing procedure': datasets were separated by subjects, not by samples, to prevent signal leaking, combined with early stopping and dropout. Option A is exactly the leakage they avoided: other steps of the same person in the training set would make test performance look too good.

Source: Paper p. 212 (PDF p. 2, Sect. 2.1)

Q5 · FFNN vs LSTM and limitations

Mundt et al. found the feedforward network (FFNN) more accurate than the LSTM, yet they do not recommend the FFNN for every use. Explain (a) why the FFNN probably performed better, (b) in which situation the LSTM is the more natural choice, and (c) one important limitation of the study's data that must be addressed before clinical use.

Show model answer

(a) The FFNN receives all time steps of the gait cycle or stance phase at once, so it can use both earlier and later information, whereas the LSTM only carries information forward from past time steps; this especially improved the first time steps and the overall curve shape. (b) For real-time applications and recordings of arbitrary length the LSTM is more natural, because it processes the signal step by step; the FFNN suits fully recorded and processed data. (c) The IMU inputs were simulated from optical data, so they lack the soft-tissue movement of real sensors, and all virtual sensors had the same position and orientation; the models must be tested on real IMU data (and other placements, tasks and populations) before clinical use.

Source: Paper pp. 219, 222 (PDF pp. 9, 12; Discussion and Conclusion)

Lecture 8 – Epidemiological data

Systematic literature review (structured review with content analysis)

Galetsi et al. (2020) — Big data analytics in health sector: Theoretical framework, techniques and prospects

Galetsi, P., Katsaliaki, K., & Kumar, S. (2020). Big data analytics in health sector: Theoretical framework, techniques and prospects. International Journal of Information Management, 50, 206–216. https://doi.org/10.1016/j.ijinfomgt.2019.05.003 · Open paper

A systematic review of 804 healthcare big-data-analytics papers (2000–2016). Using a resource-based-view framework, it classifies the types of health data, the analysis techniques, the value created, the computing platforms (Hadoop/MapReduce) and future research directions.

Question

How are big health-data resources used and analysed to create value and capabilities for healthcare stakeholders? Using the resource-based view (data resources + analysis → capabilities/values), the authors map which data types are used, which analysis techniques are applied, which values are created, which platforms and tools handle big health data, and which future directions are proposed.

Why it matters

Healthcare is highly data-intensive (electronic health records, lab systems, medical images, home sensors, billing data, social media posts), and big data are defined by the 4 Vs: volume, velocity, variety and veracity. Publications on health analytics are increasing rapidly, so practitioners, policy makers and researchers need a structured synthesis that links technical advances to the value they create.

Approach

Two-step systematic review following Tranfield et al. (2003) and Dubey et al. (2017). Step 1: they searched Web of Science and Scopus (2000–2016) for 'business intelligence', 'analytics' and 'big data' combined with health*, medical and clinical terms. Only English-language articles and reviews relevant to both health and big data analytics were included: 6817 hits → 3241 after removing duplicates → 1877 after title/abstract screening → 804 included after full-text screening. Step 2: content analysis, in which papers were coded with NVivo into data types, techniques, created values and future directions (a paper could fall into several categories). Frequencies were then cross-tabulated, e.g. technique by data type and technique by value.

Findings

  • Data types: clinical data (e.g. electronic health records, medical images) dominate at 69.9% (562/804 papers). They are followed by patient behaviour & sentiment data (wearable sensors, social sites; 16.5%), administrative & cost data (7.3%) and pharmaceutical R&D data (4.7%).
  • Techniques: modelling (42.8%) and machine learning (40.7%) were the most used, followed by data mining (24.9%), visualisation (19%) and statistics (16.4%). Machine learning is described as the most applied technique across almost all values and data types, so the authors examine it in a separate section.
  • Created values: the most frequent was better diagnosis for more personalised healthcare (35.6%), followed by supporting or replacing professionals' decision-making with automated algorithms (25.6%) and new business models, products and services (24.5%). Values are grouped by who benefits: patients (P), analysts (A) or management (M).
  • Platforms: Apache Hadoop (HDFS distributed storage + MapReduce parallel processing) is the most-used computing platform. A main difficulty is that most health data are unstructured.
  • Future directions are mostly technological (most often: developing the presented approach further, 16.5%). For ML specifically they name unsupervised learning to phenotype complex disease, automated risk prediction to guide care, reinforcement learning, the tension between accuracy and interpretability, and training clinicians to judge ML-based aids.

Limitations

  • Time lag (stated by the authors): the reviewed papers end in 2016, and technology moves faster than publication, so recent advances are missing.
  • Classification into subcategories was done by one author in NVivo (with advice from the others), and the categories overlap (papers counted in several), so the coding is partly subjective.
  • Own inference: it is a descriptive frequency mapping that does not appraise the quality or effectiveness of the included studies, and the search terms (business intelligence / analytics / big data, English only) may have missed relevant work.

Link to the lectures

Slide 44 'Health Data' takes its health-data typology from this paper (clinical → clinical endpoints; real-time patient → wearable sensors; administrative → costs and reimbursements; pharmaceutical → toxicological data, medicine use), and slide 76 gives the full reference. Canvas labels the file 'L9', but among the current lecture decks it is cited only in Lecture 8 (not in Lecture 3 or Lecture 9).

It supplies Lecture 8's typology of health data sources (slide 44). Note that the slide labels the paper's 'patient behaviour and sentiment data' as 'real-time patient: wearable sensors'. Several of its values map onto course ML concepts: identifying patient care-risk uses supervised models such as logistic regression and regression trees, while segmenting populations uses clustering (unsupervised). It also highlights the accuracy-versus-interpretability tension. Structured vs unstructured data and distributed storage/processing (Hadoop) connect to the data acquisition and storage topics of Lecture 3.

Remember

  • Systematic literature review (not original data): 804 papers, 2000–2016, Web of Science + Scopus.
  • Framework: resource-based view, in which data resources + analysis techniques → capabilities/values.
  • Big data 4 Vs: volume, velocity, variety, veracity.
  • Data types: clinical (≈70%, e.g. EHR) ≫ patient behaviour/sentiment (wearables, social media) > administrative/cost > pharmaceutical R&D.
  • Most-used techniques: modelling and machine learning, then data mining, visualisation, statistics.
  • Top value: better diagnosis → more personalised healthcare; then automated decision support.
  • Hadoop/MapReduce is the main platform; most health data are unstructured; the stated limitation is the time lag.

Practice

Q1 · Galetsi 2020 – study type

What kind of study is Galetsi et al. (2020) on big data analytics in the health sector?

Show answer

Answer: C. The authors searched Web of Science and Scopus (2000–2016), screened 6817 hits down to 804 articles and content-analysed them. The 804 refers to papers, not patients (option A), and no effect sizes were pooled (option B).

Source: Paper p. 207–208 (Methodology)

Q2 · Galetsi 2020 – health data types

According to Galetsi et al. (2020), which type of health data was analysed in the large majority (about 70%) of the reviewed big-data-analytics studies?

Show answer

Answer: A. Clinical data appeared in 69.9% (562/804) of papers, with EHRs the most common input. Patient behaviour/sentiment data (wearables, social media) came second at only 16.5%.

Source: Paper p. 208–209, Table 1; Lecture 8 Slides p. 44

Q3 · Galetsi 2020 – ML type for population segmentation

Galetsi et al. (2020) list 'offering customised actions by segmenting populations' as a value of big data analytics: patients are divided into new groups, without predefined labels, to offer more targeted services. Which type of machine learning fits this task?

Show answer

Answer: D. Finding groups without predefined labels is unsupervised clustering, which the paper names for this value. By contrast, 'identifying patient care-risk' uses supervised models such as logistic regression and regression trees, which need a known outcome.

Source: Paper p. 210–211 (Table 3 and text)

Q4 · Galetsi 2020 – created value

Which value of big data analytics did Galetsi et al. (2020) find most often in the reviewed health studies?

Show answer

Answer: B. 'Better diagnosis for provision of more personalised healthcare' was reported in 35.6% of papers, the most of the 10 values. Protecting privacy was the least frequent (5.1%).

Source: Paper p. 210, Table 3

Q5 · Galetsi 2020 – approach and limitation

Describe the general approach of Galetsi et al. (2020) and give one limitation of this type of study.

Show model answer

It is a systematic literature review. The authors searched Web of Science and Scopus for 2000–2016 using big data / analytics / business intelligence terms combined with health, medical and clinical terms, and applied inclusion criteria (English articles and reviews on health and big data analytics). Screening reduced 6817 hits to 804 papers. Guided by the resource-based view, they content-analysed and coded the papers (NVivo) into data types, analysis techniques, created values, platforms and future directions, and reported frequencies and cross-tabulations. Limitations: a time lag (it covers only papers up to 2016 while technology moves fast, as the authors note), subjective coding largely by one author with overlapping categories, and no appraisal of the quality or effectiveness of the reviewed studies.

Source: Paper p. 207–208 (Methodology), p. 214 (limitation)

Lecture 9 – User-generated data

original research (4-page conference paper)

Altini & Amft (2018) — Estimating Running Performance Combining Non-invasive Physiological Measurements and Training Patterns in Free-Living

Altini, M., & Amft, O. (2018). Estimating running performance combining non-invasive physiological measurements and training patterns in free-living. In Proceedings of the 40th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), pp. 2845–2848. https://doi.org/10.1109/EMBC.2018.8512924 · Open paper

Using two years of free-living data from 2113 HRV4Training users, multiple linear regression estimated each runner's best 10 km time with an error of under 3 minutes (about 4%), and accuracy rose as training volume, training-intensity patterns and previous performance were added to basic anthropometrics.

Question

Can running performance (best 10 km time) be estimated accurately from data collected in daily life (app-based morning heart rate and HRV, user input and synced running workouts) without any lab tests? The authors also asked which groups of predictors matter most and how much workout data is needed for an accurate estimate.

Why it matters

Accurate performance estimates could help runners pace races and tailor training plans, which may also reduce injury risk. Earlier studies used small, homogeneous samples, relied on lab variables such as VO2max or lactate threshold that are impractical outside the lab, analysed predictors in isolation and often lacked cross-validation; user-generated app and wearable data offer far larger and more realistic samples.

Approach

Users of the HRV4Training app consented to share data. They measured resting HR and HRV each morning with the phone camera (validated against ECG) or an external sensor, and linked Strava or TrainingPeaks so workouts with GPS and heart rate were imported automatically. In 2016–2017, 2113 users met the inclusion criteria (training with an HR monitor, linked workouts, at least one 10 km, at least a month of morning measurements). The reference outcome was each user's fastest 10 km found in their workouts; features were computed from the preceding 3 months and grouped into sets that were added step by step: anthropometrics (BMI, age, gender), resting physiology (HR, rMSSD), training volume and speed, training physiology (speed-to-HR ratio), training polarization (share of workouts clearly faster or slower than the user's average, or at moderate HR) and previous performance. Polarization features were defined as deviations from the user's own average rather than max, min or range, to stay robust to noisy free-living data. Multiple linear regression was used because its coefficients are easy to interpret; models were validated with 10-fold cross-validation (training on users with at least 20 workouts, testing on all users) and evaluated with RMSE, mean percentage error, r and R².

Findings

  • Estimation error dropped as feature sets were added: RMSE 6.27 min with anthropometrics only, 6.07 with resting HR/HRV, 4.04 with training volume and speed, 3.96 with training physiology, 3.64 with training polarization and 2.68 min with previous performance (Table II). The biggest single gain came from adding training volume and speed.
  • The best model (all features including previous performance) reached RMSE 2.6 min in the abstract (2.68 min in Table II), about 4% error, r = 0.93 and R² = 0.87, a 58% improvement over anthropometrics alone.
  • About 15 workouts were enough to extract meaningful features and minimise estimation error (Fig. 3).
  • Coefficient signs: lower age, BMI and resting HR, higher HRV, longer and faster workouts, more workouts clearly faster or slower than the user's average and less time at moderate HR intensity were associated with a faster 10 km; in other words, a more polarized training pattern went with better performance.
  • These relationships agree with earlier small lab-based studies, but were obtained without any laboratory testing, so such models could be used by recreational runners to estimate performance and tailor training.

Limitations

  • Authors: only readily available free-living variables were used; biomechanics and running power could be added, and the results could be backed up by a laboratory validation.
  • Authors: the models estimate performance at one point in time; whether they can track performance changes over time is future work.
  • Own inference: the data are observational, so associations (e.g., polarized training with faster times) do not prove that changing training would improve performance.
  • Own inference: the sample is self-selected app users with HR monitors (1891 of 2113 male), and the reference is the fastest 10 km found in workout data rather than a standardised race or time trial, which limits generalisability and reference quality.

Link to the lectures

Guest lecture by Marco Altini; presented as the 'Estimating running performance' example of the reference-data challenge, Slides pp. 99–113.

In Lecture 9 Marco Altini uses this paper as the example of the reference-data problem in user-generated data: instead of bringing a few people to the lab for a treadmill time trial, take years of workouts from apps like Strava, use each person's best 10 km as the reference, build features from the preceding training and estimate performance (Slides pp. 99–113). It also illustrates regression (Lecture 5: multiple linear regression, RMSE, R², cross-validation) and feature engineering (turning raw workouts into features such as volume, speed-to-HR ratio and polarization). Defining features as deviations from the user's own average connects to the lecture's points on noisy data and quality control (Slides pp. 81–95). Slide-vs-paper note: Slide 113 says 'N = 2100, RMSE = 2 minutes (4%)'; the paper reports 2113 users and RMSE 2.6 min (abstract) or 2.68 min (Table II).

Remember

  • Aim: estimate best 10 km running time from free-living app and wearable data, without lab tests.
  • Data: HRV4Training app + Strava/TrainingPeaks APIs; 2113 users over 2 years; reference = each user's fastest 10 km in their own workouts.
  • Supervised regression with multiple linear regression (chosen for interpretable coefficients), validated with 10-fold cross-validation.
  • Accuracy improved as feature sets were added: anthropometrics alone worst (RMSE ≈ 6.3 min), full model with previous performance best (RMSE ≈ 2.6–2.7 min, about 4%, R² = 0.87).
  • Adding training volume and speed gave the biggest single jump in accuracy (RMSE 6.07 → 4.04 min).
  • About 15 workouts were enough for meaningful features.
  • More polarized training (fewer moderate-intensity workouts) was associated with a faster 10 km.

Practice

Q1 · Type of ML problem

Altini & Amft (2018) used anthropometrics, morning HR/HRV and features of each user's training to estimate that user's best 10 km running time in minutes. What type of machine learning problem is this?

Show answer

Answer: A. The target (10 km time) is continuous and known for every user, so this is supervised regression; the authors used multiple linear regression. Classification would only apply if runners were assigned to categories (e.g., fast vs slow).

Source: Paper PDF p. 2 (Abstract); PDF p. 4 (Sect. III-B)

Q2 · Reference data

A key challenge with user-generated data is the lack of reference (outcome) data. How did Altini & Amft obtain the reference running performance for their users?

Show answer

Answer: C. The best 10 km time was automatically identified as the fastest 10 km workout over the 2 years and used as the reference. The lecture contrasts this with the classic option A (a few people on a treadmill), which does not scale; lab variables (D) are exactly what the authors wanted to avoid.

Source: Paper PDF p. 3 (Sect. II-B); Slides pp. 100–103

Q3 · Main finding: feature sets

What did Altini & Amft find when they added the different feature sets to their regression models?

Show answer

Answer: D. Table II shows RMSE falling from 6.27 min (anthropometrics) to 2.68 min (all features plus previous performance), with the largest gain when training volume and speed were added. Resting physiology (A) improved the estimate only slightly (6.07 min). Exam note: Slide 113 rounds the best error to 'RMSE = 2 minutes (4%)'; the paper reports 2.6 min (abstract) and 2.68 min (Table II), so 'under 3 minutes, about 4%' fits both.

Source: Paper PDF p. 4 (Table II, Sect. IV-B); PDF p. 5 (Fig. 4); Slides pp. 108–111, 113

Q4 · Training polarization

What did the regression coefficients in Altini & Amft (2018) suggest about training patterns and 10 km performance?

Show answer

Answer: B. Time at moderate HR intensity entered the model with a positive sign (more = slower), whereas the share of workouts 5% faster or slower than average, average distance and HRV entered with negative signs (more = faster). Because the data are observational, this is an association, not proof that polarizing training causes improvement.

Source: Paper PDF p. 5 (Sect. IV-C)

Q5 · User-generated data challenges

Altini & Amft built their model from free-living app and wearable data instead of lab measurements. Name two challenges of such user-generated data and explain how the authors handled each one. Also give one reason why they chose multiple linear regression.

Show model answer

Any two of: (1) Missing reference data: users do not come to the lab, so the authors used each user's fastest 10 km found in their synced workouts as the outcome. (2) Noisy data and outliers from free-living sensors: polarization features were computed as deviations from the user's own average instead of max, min or range, to stay robust to measurement errors. (3) Unequal amounts of data per user: they trained only on users with at least 20 workouts, tested on all users, and showed that about 15 workouts are enough; 10-fold cross-validation gave an honest error estimate. They chose multiple linear regression because its coefficients are easy to analyse and interpret, which let them show the importance and direction of each predictor (e.g., polarized training is associated with better performance).

Source: Paper PDF pp. 3–4 (Sect. II-B, III-A to III-C, Fig. 3); Slides pp. 97–98, 112

Lecture 9 – User-generated data

perspective / magazine column (no new data)

Altini & Dunne (2021) — What's Next For Wearable Sensing?

Altini, M., & Dunne, L. (2021). What's next for wearable sensing? IEEE Pervasive Computing, 20(4), 87–92. https://doi.org/10.1109/MPRV.2021.3108377 · Open paper

Two wearable-computing editors argue that wearables have shifted from counting behavior to tracking relative physiological changes, set out three steps for turning wearable data into new knowledge, list the main open challenges (data accuracy, wearability, energy) and predict a move toward new sensors, closed-loop actuation and machine-learning forecasting.

Question

Where does wearable sensing stand after a decade of commercial growth, and what comes next? The authors review recent developments (what is sensed now, the shift from behavior to health monitoring, large-scale research), the main open challenges, and the likely next steps.

Why it matters

Wearables have become mass-market devices (Apple Watch above 100 million sales, Google's acquisition of Fitbit) and now enable observational studies with thousands of participants, including infection monitoring during COVID-19. At the same time, sensor accuracy and bias came into question, so researchers need a clear view of what wearable data can and cannot support.

Approach

This is a perspective column, not an empirical study: the authors summarise recent developments and selected studies. They describe current sensing (standard phone sensors, blood pressure from PPG, earables, interstitial-fluid glucose; sweat-based sensing so far unsuccessful), the shift from behavioral to physiological monitoring, and a three-step route to knowledge discovery illustrated with a Fitbit heart-rate example and Radin et al.'s COVID-19 recovery data. They then discuss limitations (accurate data, wearability and body access, energy consumption) and give a look ahead (new sensing, from sensing to actuating, machine-learning forecasting).

Findings

  • The focus has shifted from absolute behavioral measures (steps, energy expenditure) to relative changes in physiology (resting HR, HRV, blood glucose) in response to stressors such as sickness, alcohol, exercise, diet and the menstrual cycle; activity and location remain important as context.
  • Three steps to knowledge discovery: (1) the technology must be accurate (validation, certification); (2) it must reproduce known results from smaller clinical studies (e.g., acute responses to exercise or sickness); (3) only then can large-scale deployment reveal new relationships, e.g. COVID-19, long COVID and the return of baseline physiology (Radin et al.).
  • Possibly the biggest challenge is the noise of unsupervised free-living settings: sensors can be misused or fail (e.g., optical HR during exercise), signal quality and confidence are rarely reported, and skin pigmentation, motion artefacts and sensor position (top of the wrist while the blood vessels lie underneath) reduce accuracy.
  • Other challenges: single-location (wrist) devices limit what can be sensed while clothing-based e-textile systems are hard to develop, test and manufacture; energy consumption and battery power also remain issues.
  • Next steps: new sensing modalities with more attention to bias and inclusion, closed-loop systems that act as well as sense (e.g., glucose monitor + insulin pump, vagal nerve stimulation), and machine-learning models that forecast individual responses instead of only describing past changes.

Limitations

  • Own inference: this is an opinion/perspective piece, not a systematic review; the examples are chosen by the authors and no new data are analysed.
  • Own inference: the author bio lists the first author as founder of HRV4Training and data-science advisor at Oura, so the view comes partly from the consumer-wearable industry; weigh claims accordingly.
  • Own inference: the forecasting vision is aspirational; the authors themselves note that prediction from wearable data is still limited and that consumer estimates are not always accurate.

Link to the lectures

Its 'three steps to knowledge discovery' match the lecture's three key steps; its data-accuracy challenge matches the lecture's points on noisy wearable data.

Lecture 9 (Altini's guest lecture) is built on the same ideas: user-generated data enable research at larger scale and in realistic settings, but require three key steps: validate the technology ('garbage in, garbage out'), deploy and confirm lab-based insights, then discover new relations or build new products (Slides pp. 57–59). This is the paper's 'three steps to knowledge discovery'. The paper's accurate-data challenge mirrors the lecture's point that wearable data are extremely noisy and typically report no signal-quality metric (Slides p. 81). Its COVID-19 recovery example (Fig. 2, reprinted from Radin et al.) links to Radin et al. (2021) and to the lecture's infection and sickness examples (Slides pp. 21, 69, 115–117). The paper is not shown explicitly on the slides.

Remember

  • Perspective/column in IEEE Pervasive Computing (2021), not an empirical study.
  • Shift: from monitoring behavior (steps, energy expenditure; absolute measures) to monitoring health (relative changes in resting HR, HRV, glucose in response to stressors).
  • Three steps: validate accuracy → reproduce known lab/clinical results → large-scale deployment for new discoveries.
  • Possibly the biggest challenge: the noise of unsupervised free-living settings; signal quality is rarely reported.
  • Wrist optical HR is error-prone because the sensor sits on top of the wrist while the blood vessels are underneath, so it relies on capillaries.
  • Future: closed-loop sensing + actuating (e.g., glucose monitor + insulin pump) and ML forecasting of individual responses.

Practice

Q1 · From behavior to health

According to Altini & Dunne (2021), how has the focus of consumer wearables changed in recent years?

Show answer

Answer: B. The section 'From Monitoring Behavior to Monitoring Health' describes exactly this shift; behavior and context data are still used, but to contextualise the physiological changes. Option A reverses the direction of the shift.

Source: Paper p. 88 (PDF p. 2)

Q2 · Three steps to knowledge discovery

Altini & Dunne describe three steps that are needed before wearable data can produce new knowledge. Which order is correct?

Show answer

Answer: D. The paper requires accurate technology first, then replication of known responses (e.g., acute changes after exercise or sickness), and only then large-scale discovery (e.g., COVID-19 recovery). The lecture's 'three key steps' (validate, deploy and confirm lab-based insights, discover new relations) follow the same order.

Source: Paper pp. 88–89 (PDF pp. 2–3); Slides pp. 57–59

Q3 · Main challenge: data accuracy

What do Altini & Dunne name as possibly the biggest challenge for consumer wearables?

Show answer

Answer: A. The 'Accurate Data' section opens with this claim and adds that wearables rarely report signal-quality issues even though this information is available to the sensor. Option C contradicts the paper, which says the main smartphone sensors have stayed standard over the years.

Source: Paper p. 90 (PDF p. 4); p. 87 (PDF p. 1)

Q4 · Look ahead: ML forecasting

Altini & Dunne argue that most current wearable insights are descriptive. What role do they foresee for machine learning?

Show answer

Answer: C. In 'Machine-Learning Based Data Interpretation and Forecasting' they note that prediction is still limited and give examples such as glucose after a meal or next-day HRV; the conclusion expects 'ubiquitous machine learning models able to forecast individual responses'. Battery power (D) is discussed as a separate hardware challenge, not as the ML vision.

Source: Paper p. 91 (PDF p. 5)

Q5 · Three steps applied

Using the Fitbit heart-rate example from Altini & Dunne (2021), explain the three steps needed before large-scale wearable data can generate new knowledge, and why skipping the first two steps is risky.

Show model answer

Step 1: validate that the device (e.g., a Fitbit) measures heart rate accurately against a reference system. Step 2: show that it reproduces well-known responses from smaller lab or clinical studies, such as the acute change in heart rate after exercise or during the acute phase of sickness. Step 3: deploy at large scale to study relationships that are hard to test in the lab, e.g. how long resting heart rate stays elevated after COVID-19 before returning to baseline (Radin et al.). Skipping steps 1–2 risks 'garbage in, garbage out': noisy or biased free-living data could produce confident but false findings, and without replicating known effects you cannot tell whether a new finding is real or a device artefact.

Source: Paper pp. 88–89 (PDF pp. 2–3); Slides pp. 57–59

Lecture 9 – User-generated data

research letter (observational cohort study)

Radin et al. (2021) — Assessment of Prolonged Physiological and Behavioral Changes Associated With COVID-19 Infection

Radin, J. M., Quer, G., Ramos, E., Baca-Motes, K., Gadaleta, M., Topol, E. J., & Steinhubl, S. R. (2021). Assessment of prolonged physiological and behavioral changes associated with COVID-19 infection. JAMA Network Open, 4(7), e2115959. https://doi.org/10.1001/jamanetworkopen.2021.15959 · Open paper

In the app-based DETECT study, wearable data showed that people with COVID-19 took much longer than symptomatic COVID-negative people to return to their own baseline (resting heart rate on average 79 days), and a subgroup still had an elevated resting heart rate after more than 133 days.

Question

How long does it take for physiological (resting heart rate) and behavioral (sleep, step count) measures to return to a person's own pre-illness baseline after COVID-19, and how much does recovery vary between people, compared with symptomatic people who tested negative?

Why it matters

Long-term symptoms after COVID-19, such as autonomic dysfunction and cardiac damage, had been reported for up to 6 months but had not been quantified. Wearables can track physiology and behavior continuously from before infection (a healthy baseline) through illness and recovery, which a standard clinical study rarely captures.

Approach

DETECT is a remote, app-based longitudinal study that enrolled 37,146 US adults (March 2020 – January 2021) and collected their wearable data. This analysis included the 875 participants who reported acute respiratory symptoms and had a COVID-19 swab test: 234 tested positive and 641 tested negative (the comparison group). For each person, daily resting heart rate was expressed as the deviation from their own baseline (daily RHR − baseline mean RHR); sleep quantity and step count were also followed from 7 days before to 133 days after symptom onset. COVID-positive participants were further grouped by their mean RHR deviation 28–56 days after symptom onset (<1, 1–5 or >5 bpm), and acute symptoms were compared between these groups with χ² tests (ANOVA for age). The analysis is descriptive (mean curves with 95% confidence intervals), not a prediction model.

Findings

  • COVID-19-positive participants took longer than symptomatic COVID-negative participants to return to their baseline for resting heart rate, sleep and activity.
  • The difference was largest for RHR: a brief drop (transient bradycardia) followed by a prolonged relative elevation (tachycardia) that on average did not return to baseline until 79 days after symptom onset. Step count and sleep returned sooner, at 32 and 24 days.
  • A subgroup of 32 COVID-positive participants (13.7%) kept an RHR more than 5 bpm above baseline that had not returned to normal after more than 133 days.
  • During the acute phase this subgroup reported more cough (84.4%), body ache (62.5%) and shortness of breath (28.1%) than the other groups, suggesting that early symptoms and a larger initial RHR response may be associated with a longer recovery.
  • Overall, COVID-19 had a prolonged physiological impact of about 2–3 months on average, with large variability in recovery, which may reflect autonomic nervous system dysfunction or ongoing inflammation.

Limitations

  • Authors: symptoms were collected only during the acute phase, so long-term physiological changes could not be compared with long-term (long COVID) symptoms.
  • Authors: larger samples and more complete participant-reported outcomes are needed to explain why recovery differs between individuals.
  • Own inference: observational, self-selected sample (US volunteers with wearables, about 71% women), so results may not generalise and associations are not causal.
  • Own inference: the curves are group averages; individual trajectories differ a lot (the lecture's warning: 'These are group averages. What about the individual?').

Link to the lectures

Example of large-scale wearable data revealing an outcome, sickness and recovery, that is hard to study in the lab; swab tests serve as the reference data.

In Lecture 9, user-generated wearable data are presented as a way to study outcomes that cannot easily be tested in the lab, such as sickness, and detecting an infection with the Oura ring is one of the opening examples (Slides pp. 21, 69, 73). Radin et al. show this in practice: a large remote cohort, each person compared with their own baseline, and swab-test results as the reference data the lecture calls essential ('When did we get infected? How do we determine it? Symptoms, tests', Slides pp. 115–120). Altini & Dunne (2021) reprint Radin's figure as their example of large-scale discovery. Note: the slides do not show this paper; the lecture's pandemic slides (pp. 124–128) show a different analysis (HRV4Training users' resting HR during the 2020 lockdown), and 'L10 - fitbit_study.pdf' (Natarajan et al., 2020) is a separate paper that the slides do not use.

Remember

  • Research letter; observational cohort (DETECT app study, US); descriptive analysis, not a prediction model.
  • Compared 234 COVID-positive with 641 COVID-negative symptomatic participants (all swab-tested).
  • Outcome = deviation from each person's own baseline (daily RHR − baseline RHR), plus sleep and step count.
  • RHR took longest to normalise: on average 79 days (steps 32 days, sleep 24 days).
  • RHR pattern: brief bradycardia, then prolonged relative tachycardia.
  • 13.7% of COVID-positive participants had RHR more than 5 bpm above baseline for over 133 days; they had more acute cough, body ache and shortness of breath.
  • Main limitation: symptoms were only collected in the acute phase.

Practice

Q1 · Study design

Which description best fits the approach of Radin et al. (2021)?

Show answer

Answer: D. Participants in the DETECT app study were followed over time, and their RHR, sleep and steps were expressed as deviations from their personal baseline; the results are descriptive mean curves. Option B describes earlier detection studies cited in the introduction, not this paper.

Source: Radin et al., Introduction and Methods sections

Q2 · Why wearables

What made wearables particularly suitable for answering Radin et al.'s question about COVID-19 recovery?

Show answer

Answer: A. The introduction states that wearables can track a person's metrics from when they are healthy, through infection, and back to baseline. Option B is wrong: infection status came from swab tests, which served as the reference data.

Source: Radin et al., Introduction and Methods sections

Q3 · Main finding

Which result did Radin et al. report for COVID-19-positive participants?

Show answer

Answer: C. RHR showed a brief bradycardia followed by a prolonged relative tachycardia that returned to baseline only after 79 days on average; steps and sleep normalised at 32 and 24 days. Option D describes only the brief initial dip, not the prolonged elevation.

Source: Radin et al., Results section

Q4 · Subgroup with prolonged elevation

Radin et al. identified a subgroup of COVID-19-positive participants whose resting heart rate stayed more than 5 bpm above baseline for over 133 days. What characterised this subgroup?

Show answer

Answer: B. 32 participants (13.7%) formed this group; during the acute phase they reported more cough (84.4%), body ache (62.5%) and shortness of breath (28.1%). The authors suggest early symptoms and a larger initial RHR response may be associated with a longer recovery; this is an association, not a proven cause.

Source: Radin et al., Results and Discussion sections

Q5 · Limitations of wearable cohort data

Give the limitation that Radin et al. themselves mention, and one further limitation of using user-generated wearable data in this study. Explain how each one limits the conclusions that can be drawn.

Show model answer

Authors' limitation: symptoms were collected only during the acute phase, so the prolonged changes in RHR, sleep and activity could not be linked to long-term (long COVID) symptoms; we know the physiology stays altered, but not whether people still feel ill. Further limitation (any one): participants were self-selected US volunteers who own wearables (about 71% women), so the results may not generalise to other populations; the study is observational, so associations (e.g., more acute symptoms with a longer recovery) are not causal; the curves are group averages that hide large individual differences, so they cannot predict one person's recovery.

Source: Radin et al., Results and Discussion sections; Slides pp. 115–120, 129–130