Home: all lectures
0 XP0 day streak

Lecture 3 · Study guide

Data in sport and health

3 steps25–35 min block2 practice questions

1Step 1 of 3≈ 5 min

Overview

First read: about 4–6 minutes. Lecture 3: Data in sport and health.

This lecture asks how to find and organize sport/health data. It uses a cyclist study and speed-skating results to show that useful acquisition starts with a question, then connects files, APIs and analytical storage.

The ideas to keep

Blur the explanations and test yourself.

Data sources. Performance, tests, videos, questionnaires, sensors and diagnoses offer different information and quality limitations.

Cycling case. Wingate and a 15 km time trial measure sprint/endurance. Multilevel physiology illustrates why selecting meaningful variables requires domain knowledge. Association is not an intervention effect.

Acquisition. Scraping extracts information from web pages; an API is a defined software interface. JSON/XML can carry the responses. Neither method guarantees accurate data.

Storage. A database organizes records; a warehouse integrates sources for analysis; a mart focuses on one area; a lake retains original structured/unstructured data for later use.

ETL. Extract source data, transform formats/units/structure and check them, then load an analytical destination.

Skating example. Compare seasonal performance with the world record on the race date and show specialization and track context. A larger raw time ratio means slower relative performance.

Range matters. An elite sample with little variability can hide a relationship. No detected association in that sample does not prove irrelevance everywhere.

What to be able to do

Be able to choose a source/store, explain ETL and API versus scraping, and connect the study question to its measurements. The example exams indicate question styles, not an exhaustive syllabus.

Next step

Read the detailed notes where you need an explanation, then attempt the two original questions before opening the answers. The full slides are on Canvas.

Got the big picture?Mark the overview done to fill this lecture's ring.


2Step 2 of 312–18 min

Detailed notes

Source: full presentation. Canvas lecture page checked 11 October 2026, 11:40 CEST. Existing lecture transcript used to clarify demonstrations; its older administrative dates are excluded. Examples labelled “illustration” are invented to explain the slide concepts.

Overview · Two practice questions

Jump to a section · 10
  1. Start with the question, then choose the data
  2. Categories and value of sport/health data
  3. Detailed case: combining sprint and endurance
  4. Files and databases
  5. Web scraping and APIs
  6. Speed-skating API case
  7. Storage concepts, explained as a useful map
  8. ETL: extract, transform, load
  9. From the recording
  10. Slide coverage map

Start with the question, then choose the data

Acquisition means obtaining information that can answer a question. Ask: what do I need, where is it, how can I obtain it, and how should it be stored and accessed? Do this before collecting everything you can find.

The lecture connects three topics: examples of sport/health measurements, ways of acquiring and storing them, and extract–transform–load (ETL). Acquisition and preparation overlap: an imported table often needs transforming before you can use it.

Categories and value of sport/health data

Category in the slides Example What to think about
Performance Race results, sprint power Event, units and testing conditions
Physical tests Strength, oxygen uptake Protocol and measurement reliability
Video/imaging Match video, ultrasound, radiology Features must be extracted and validated
Questionnaires/training schemes Symptoms or planned sessions Wording, self-report and planned versus completed training
Sensor data Heart rate, movement Noise, wear time, timestamps
Diagnosis Clinical outcomes Definition and completeness

Radiology illustrates potential benefits from recognizing patterns and screening abnormalities faster. Those potential benefits still need evaluation; using AI does not guarantee fewer errors.

The insurance activity example illustrates commercial value from tracking goals, activity and personal health information. Ask whose interests a use serves and what consent/privacy conditions apply.

Fitbit resting-heart-rate data illustrate scale and subgroup comparisons: age, sex, BMI and physical activity may all relate to resting heart rate. The slides ask students to sketch expected relationships between activity minutes and resting heart rate for men and women. A reasonable hypothesis is lower resting heart rate with greater activity, with diminishing changes. Slide p. 11 asks for expected cut-off numbers and slide p. 12 shows the result: resting heart rate falls steeply, then levels off from roughly 200–250 active minutes per week, with women higher than men at every activity level (consistent with the 150–300 min/week activity guideline). Resting heart rate is influenced by many other factors, including medication, hydration, stress and body size.

Detailed case: combining sprint and endurance

The van der Zwaard example asks which physiological characteristics relate to achieving both high sprint and endurance performance. Sprint performance benefits from the capacity to produce high force/power; endurance requires sustained energy production and oxygen supply. Large muscle fibres can create a challenge for oxygen delivery, so excelling in both is not simply a matter of maximizing every variable.

The slides describe 28 cyclists, including sprint, team-pursuit and road athletes, competing up to Olympic level. Performance was assessed with a Wingate test for sprint and a 15 km time trial for endurance. Candidate determinants were measured at several biological levels:

Level Measures shown
Whole-body exercise VO₂max, oxygen uptake during performance, lactate and ventilatory thresholds, gross efficiency
Blood/oxygen transport Lactate, oxygenation, haemoglobin, haematocrit, red cells and related blood indices
Muscle structure/function Ultrasound-derived volume, fascicle length, pennation angle, physiological cross-sectional area (PCSA), torque and specific force
Muscle fibres Fibre type, fibre cross-sectional area (FCSA), fibre number, oxidative capacity, myoglobin and capillaries

FCSA is the area of an individual fibre. PCSA describes a muscle's effective cross-sectional area related to force production; these are not the same measurement. Gross efficiency relates useful mechanical work to energy expenditure. Capillaries deliver blood, and myoglobin supports oxygen availability within muscle.

The visual argument is that sprint and endurance performance trade off, but examining combinations of determinants can reveal profiles compatible with better combined performance. The slides emphasize long muscle fibres, capillaries/myoglobin and oxidative capacity in this explanation. Mind the direction for PCSA: a small PCSA (less hypertrophy) helps. High oxidative capacity, capillaries × myoglobin and a small PCSA together explained 67% of the variance in performance VO2 (R² = 0.67), a key determinant of combined sprint–endurance performance. Their result diagrams include study-specific percentages; they describe fitted relationships in this sample, not guaranteed gains from a training intervention.

Original study-result slide: determinants of combined performance

The transcript discusses Deming regression as a way to study a relationship when both measurements have error, and warns about restricted range. If all participants are very similar in a predictor, it becomes hard to identify its relationship with performance. This can occur in an elite sample even when the variable matters in a broader population. A small sample also makes estimates uncertain. The lecturer explicitly treats the study as an example of using data and modelling; understanding the logic is more useful than memorizing every coefficient.

Do not jump from association to “this training causes improved performance.” The measurements can suggest hypotheses and potential targets for training; establishing intervention effects needs appropriate evidence.

Files and databases

A spreadsheet or CSV can be sufficient for a small table. A database becomes useful when data grow, several tables relate to each other, or you need repeated efficient queries.

The slides contrast legacy .xls limits (65,536 rows, 256 columns) with .xlsx (1,048,576 rows, 16,384 columns). CSV has no comparable built-in worksheet dimensions, but “unlimited” in the slide does not mean infinite practical storage: memory, software and file size still limit handling.

A relational database stores related tables. For example, one table stores athlete details, another sessions, and another measurements. An athlete ID connects records without repeating all personal details in every row. Indexing helps retrieve relevant records. SQL is a language commonly used to query relational databases; NoSQL refers to other database approaches.

Web scraping and APIs

Web scraping retrieves and parses website content to extract information. An application programming interface (API) provides a defined interface through which software requests data or operations. An API is not the database itself; it governs access to functions or information exposed by a service.

Illustration: a web page displays race results. Scraping reads the page's layout. An API may instead let a program request an athlete's results directly in a structured response. Changes to page layout can break a scraper, whereas API users rely on the documented request/response structure, which can also change.

Requests typically use HTTP. The slides show JSON and XML as formats for representing exchanged information. You need to know their role, not memorize their full syntax. Providers include companies, governments, leagues, publishers and individuals. Availability and permissions are separate questions.

Speed-skating API case

The example obtains world-record progressions, seasonal bests and athlete information from a results service. It asks how to visualize talent development across seasons.

The comparison uses an athlete's seasonal-best time relative to the world record on the race day, rather than today's record. A ratio seasonal best / world record = 1.2 means the athlete took 120% of the record time: 20% longer. A ratio of 1 means equal time.

Important visual distinction: for this raw time ratio, smaller is better. The slides say “bigger is better” for the graphic's performance presentation, which must use an appropriate plotting direction/transformation. Do not interpret a bigger raw time ratio as faster skating.

The graphic's direction/shape shows specialization across sprint, middle and long distances; size conveys level; track symbols distinguish indoor/open and highland/lowland conditions; a circle marks a record. Animation shows seasonal development. The full deck preserves the figure; a single number should not replace its multidimensional information.

Storage concepts, explained as a useful map

How raw sources become an organized analytical store

Store Role Sports illustration
Database Organized data for storage and retrieval; many possible designs Athlete and session records linked by IDs
Data warehouse Integrates data from several sources for organizational reporting/analysis Club-wide performance, medical and administrative information
Data mart Focused subset for a department or purpose Coaching department's training summaries
Data lake Retains structured and unstructured data in original formats for later use Raw sensor files, video and tables

A warehouse can contain detailed, summarized and metadata records. The deck mentions large warehouse sizes and denormalization for read/query performance; these describe common designs, not universal minimum-size definitions. A mart's focus is narrower than a warehouse's. A lake's flexibility shifts more of the interpretation/processing work to later use.

ETL: extract, transform, load

Extract: retrieve records from source systems. A full extraction retrieves everything; a partial/incremental extraction retrieves new or changed records. A staging area temporarily holds incoming data.

Transform: make the data usable for the destination. The lecture lists cleaning, format revision, threshold validation, restructuring, deduplication, filtering, merging, splitting, derivation, summaries, integration and aggregation. These are different operations, not one magic “clean” button.

Load: write the transformed records to the destination, such as a data warehouse, for analysis or business-intelligence tools. Check the staging data first.

Illustration: export separate training files → standardize athlete IDs, timestamps and units, remove accidental duplicates, calculate daily totals → load the consistent daily table into the club's analytical store. A lake might retain the original exports alongside that processed table.

The lecture briefly names Sport Data Valley as a sport-data platform context. This does not replace the principles: define your question, choose appropriate sources, organize data and check how transformations affect meaning.

From the recording

Points the lecturer made in the lecture recording on Canvas that are not (fully) on the slides. Times refer to the recording.

  • 05:04 — Structured vs unstructured by category: performance and sensor data are mostly structured; video/imaging is mostly unstructured but becomes structured once features are extracted (e.g. markerless motion-capture landmarks in a table); physical tests, questionnaires and diagnosis can be either (Borg score, closed answers, injury yes/no = structured; open interview answers or a verbal clinical judgement = unstructured). (Slides p. 6 (categories only))
  • 12:42 — Team tracking data (GPS, heart rate, acceleration for every player) are high-dimensional; summarising them into simple team metrics such as surface area covered and team length is dimensionality (complexity) reduction, a core theme. The aim is to let a model rather than the researcher choose the variables, which is more objective. (Slides p. 7 (image only))
  • 14:44 — AI in radiology detects abnormalities from grey-scale patterns without context and is more consistent than human raters (whose reading depends on training lab and daily state), but context still matters, so a clinician should interpret after machine pre-screening. (Slides p. 8 (benefits only))
  • 20:06 — The ASR activity-discount scheme did not make inactive people active: it mainly rewarded people who were already active, and losing the discount acted as a punishment. Individual incentives do not solve a systemic problem. (Slides p. 9 (scheme only, no outcome))
  • 23:29 — Fitbit results: women have a higher resting heart rate than men (one explanation: smaller heart); resting HR vs BMI is U-shaped with an optimum in the normal BMI range (very low BMI also raises resting HR). (Slides p. 10 (graph only))
  • 34:33 — Resting heart rate falls with weekly active minutes along a levelling-off (hyperbolic) curve, with women offset above men; the plateau starts around 200–250 min/week, consistent with the 150–300 min/week activity guideline. (Slides p. 11–12 (question and graph only))
  • 37:48 — Wider error bars at the extremes of a plot (e.g. very low or very high BMI) mainly reflect few participants there, so interpret those points cautiously. Strictly, the standard error/CI depends on n, not the SD (the lecturer said SD). (Slides p. 10 (graph only))
  • 46:16 — A correlation is not a percentage of explained variance: r = −0.66 between sprint and endurance means r² ≈ 0.44, i.e. about 44% shared variance (the lecturer said '36–40%', which is a slip). (Slides p. 23 (r only); rule on Lecture 4 slides p. 53)
  • 48:31 — Ordinary least squares minimises vertical (or horizontal) distances to the line, treating one variable as error-free; van der Zwaard et al. used Deming regression, which minimises perpendicular distances and accounts for error in both variables; the perpendicular residuals quantify combined sprint + endurance performance. (Slides p. 24–26 (perpendiculars drawn, method not named))
  • 54:16 — High oxidative capacity, more capillaries × myoglobin and a SMALL PCSA together explained 67% of the variance in performance VO2 (R² = 0.67), a key determinant of combined sprint–endurance performance. The lecturer phrases it as 'explains combined performance by 67%'; strictly the 67% is for performance VO2. A large muscle cross-section (hypertrophy) works against oxygen use. He does not expect you to know the coefficients — know the method (multiple regression on combined determinants). (Slides p. 28)
  • 63:58 — Restricted range: if a predictor barely varies in the sample (pennation angle in elite cyclists) or the outcome barely varies (top-10 marathon times seconds apart vs varying VO2max), no relationship can be detected even if one exists. Always check the variability of every variable before regression. (not on slides)
  • 66:46 — Before collecting, decide which data you actually need to answer the question: this concerns efficiency and storage, and is also an ethical issue (do not burden participants with unnecessary measurements). (Slides p. 33 (questions only))
  • 81:27 — In the speed-skating radar chart, 1.0 (= world record) lies on the outer rim and slower ratios (e.g. 1.2) lie toward the centre, so 'bigger is better' refers to a larger filled area, not a larger ratio. (Slides p. 49 and 51)
  • 83:42 — Public or scraped data have unknown quality: screen for outliers and missing values, compare with the expected measurement error for that data type, and cross-check sources; reporting standards (system used, capture frequency, cleaning already done) are still being developed. (not on slides)
  • 98:11 — A data lake is low effort to store into but high effort to reuse; organising data into data marts makes it reusable by others. (Slides p. 64–65 (definition only))
  • 99:10 — For every assigned paper, know which statistical model was used and why: read the methods and results sections. (not on slides)

Slide coverage map

Pages 1–4 introduce acquisition; 5–14 show data categories and value; 15–32 develop the cycling study; 33–51 cover storage, scraping, APIs and skating; 52–65 cover warehouses, ETL, marts and lakes. Repeated result figures and visual-only examples remain available in the local PDF.

Worked through the deep dive?Tick it off. Come back to any section whenever you need it.


3Step 3 of 35–8 min

Practice questions

Exactly two original practice questions. These use the concept/application style of the 2024 and 2022–2023 example exams, with new scenarios and values. Try each before opening its answer. The official example exams are on Canvas; save them for a later, exam-style session.

Question 1 — Choose the store

A federation wants to keep original match videos, raw sensor files and tabular results for analyses that have not yet been planned. Which store best fits that description?

A. A narrowly focused data mart. B. A data lake. C. A summary chart. D. A single filtered data frame.

Show answer and rationale

B. A data lake. It retains original structured and unstructured formats for later analysis. A mart is focused on a particular purpose; a chart or filtered table is already a specific derived representation.

Question 2 — Acquisition and ETL

A website publishes yearly athlete results and also offers an interface that returns results as JSON. Explain API versus scraping, then give one example action for each ETL step if you combine results across years.

Show answer and rationale

API: request results through the service’s defined software interface. Scraping: retrieve/parse displayed web-page content. Extract: obtain yearly records. Transform: standardize athlete IDs, units and year fields, investigate duplicates. Load: store the consistent results in an analytical database/warehouse. Rationale: obtaining data and organizing it for use are related but distinct operations; an API does not remove the need for quality checks.

Return to overview · Detailed explanation

Tried both questions?Answer before peeking, then rate yourself honestly.