Home: all lectures
0 XP0 day streak

Lecture 2 · Study guide

Programming in R

3 steps25–35 min block2 practice questions

1Step 1 of 3≈ 5 min

Overview

First read: about 4–6 minutes. Lecture 2: Programming in R.

R is a language for telling a computer what to do with data. This lecture starts with values and containers, then teaches selecting data, using functions, and repeating or branching instructions.

The ideas to keep

Blur the explanations and test yourself.

Objects and assignment. height <- c(168, 182) stores a vector. mean(height) applies a function. Names are case-sensitive and indexing starts at 1.

Types. Numbers, text and logical values behave differently. NA is unknown/missing; NULL is absence of an object/value.

Containers. An atomic vector or matrix has one common type. A data frame can mix column types. A list can hold different kinds and sizes of objects. A factor stores categories.

Subsetting. x[2] takes an element; table[row, column] chooses a cell; table$height chooses a column. Logical conditions select matching observations.

Functions and missingness. mean, sum, median and round summarize/change values. na.rm = TRUE excludes missing entries from a summary; it does not recover them.

Conditions and loops. if/else selects one branch; for repeats instructions. Trace what each step does rather than memorizing code appearance.

Reusable work. Write functions for repeated calculations, check units, save scripts, and load installed packages using library().

What to be able to do

Be able to trace a short script, distinguish containers, explain missing-value handling and write a small function or condition. The example exams indicate question styles, not an exhaustive syllabus.

Next step

Read the detailed notes where you need an explanation, then attempt the two original questions before opening the answers. The full slides are on Canvas.

Got the big picture?Mark the overview done to fill this lecture's ring.


2Step 2 of 312–18 min

Detailed notes

Source: full presentation. Canvas lecture page checked 11 October 2026, 11:40 CEST. Existing lecture transcript used to clarify demonstrations; its older administrative dates are excluded. Examples labelled “illustration” are invented to explain the slide concepts.

Overview · Two practice questions

Jump to a section · 13
  1. Think of R as instructions applied to stored objects
  2. Values, types and assignment
  3. Data structures: what container do you need?
  4. Subsetting: retrieve only the part you want
  5. Functions and their arguments
  6. Rounding, floor and ceiling
  7. Write a function you can reuse
  8. Conditions: choose a branch
  9. Loops: repeat an instruction
  10. Libraries/packages
  11. What the ten slide exercises teach
  12. From the recording
  13. Slide coverage map

Think of R as instructions applied to stored objects

An object is a named piece of data. A function is an instruction that takes inputs and returns a result. A script records these instructions so you can repeat the analysis.

height <- c(168, 182, 175, 190)
mean(height)
# 178.75

Read this as “store these four numbers under the name height, then calculate their mean.” The console runs commands; a saved script makes the process reproducible. RStudio supplies editing, object inspection, help and plots around R.

R objects: different containers for different jobs

Values, types and assignment

Type shown in the slides Example What it represents
Numeric, normally double 2.5, 2 Numbers, including decimal values
Integer 2L Whole-number storage type
Character "Amsterdam" Text; quotes matter
Logical TRUE, FALSE A condition's truth value
Complex 3 + 2i Numbers with real and imaginary parts
Raw charToRaw("hello") Bytes
Date/time objects as.Date("2026-10-11") Dates or timestamps represented using appropriate classes

NA means missing/unknown. It is not zero, the word “NA”, or false. It can be a missing value of different types. Date/time is a useful class of objects, rather than a separate primitive typeof() result.

Use <- to assign: age <- 24. The right-hand expression is evaluated and stored on the left. = can assign in many contexts, but <- keeps object assignment visually distinct from named function arguments such as na.rm = TRUE. Rightward assignment exists; the slides discourage it for readability.

R is case-sensitive: height and Height are different objects. Names cannot start with a digit and cannot contain arbitrary punctuation or spaces without special handling. Use simple meaningful names. ls() lists existing object names; typeof(x) checks a storage type; str(x) shows an object's structure. class(x) describes its class, which can convey more meaning than storage type.

mass <- 70
mass * 2       # 140
mass == 70     # TRUE: comparison, not assignment

Data structures: what container do you need?

Structure Shape Type rule Example
Atomic vector One-dimensional sequence One common type Several heights
Matrix Rows and columns One common type Numerical test scores
Array Multiple dimensions One common type Numerical measurements by person × test × day
Data frame Rows and named columns Columns can have different types; equal row counts ID, name, height, injured
List Collection of objects Elements can have different types and sizes A vector, table and model together
Factor Categorical vector with levels Stores category codes and category labels Discipline or medal
NULL Absence of an object/value Not an ordinary missing observation An empty result or absent list element

An atomic vector containing numbers and text is coerced to a common type, usually character. A list can preserve mixed element types. A data frame is list-like, but its columns must form a rectangular table.

height <- c(168, 182, 175, 190)
birthplace <- c("Utrecht", "Delft", "Leiden", "Haarlem")
athletes <- data.frame(height, birthplace)
str(athletes)

matrix(height, nrow = 2)
#      [,1] [,2]
# [1,]  168  175
# [2,]  182  190

R fills matrices by column unless you request byrow = TRUE. A matrix made with byrow = TRUE would instead have first row 168, 182. This difference matters when interpreting code.

information <- list(name = "Ari", height = 182, tests = c(42, 45))
discipline <- factor(c("road", "sprint", "road"))

A factor's labels are categories. Its internal integer codes are not quantities you should average. An ordered factor can explicitly encode an order where appropriate.

Subsetting: retrieve only the part you want

R indexing begins at 1. Square brackets select part of an object.

height[2]                    # 182
height[c(1, 4)]              # 168, 190
height[2:4]                  # 182, 175, 190
height[-2]                   # everything except element 2
height[height >= 180]        # 182, 190
athletes[1, 2]               # row 1, column 2
athletes[, "height"]         # all rows of height column
athletes$height              # named column
athletes[athletes$height >= 180, ]

For a matrix or data frame, the comma separates rows, columns. An empty position means all rows or all columns. athletes["height"] returns a one-column data frame, whereas athletes$height normally returns the vector. For lists, x[1] preserves a sublist and x[[1]] extracts the element.

Logical filtering works because each TRUE selects its corresponding value. A missing comparison can yield NA, so missingness requires attention rather than assuming every comparison is true or false.

Functions and their arguments

A function call has the form function_name(arguments). Use ?mean or help("mean") for help, including argument names and defaults.

vo2 <- c(42.1, 51.6, 47.3, NA)
mean(vo2)                    # NA
mean(vo2, na.rm = TRUE)      # 47
sum(height)                 # 715
median(c(62, 75, 68, 80))    # 71.5
round(c(42.123, 51.678), digits = 2)  # 42.12, 51.68

na.rm = TRUE tells these summaries to omit missing values. It does not restore the unobserved information or guarantee an unbiased estimate. Here it averages the three observed values.

names() attaches names to vector elements, not a new numerical value:

names(height) <- c("Ari", "Bo", "Cam", "Dee")
print(height)

The slide calls this a “named function” in one place; the demonstration actually creates a named vector.

Rounding, floor and ceiling

Operation Example Meaning
round(x, 2) round(3.146, 2) → 3.15 Nearest value to two decimal places
floor(x) floor(3.8) → 3 Greatest integer no larger than x
ceiling(x) ceiling(3.2) → 4 Smallest integer no smaller than x
Round to tens round(47 / 10) * 10 → 50 Divide, round, multiply back

For negatives, floor(-3.2) is −4 and ceiling(-3.2) is −3. “Nearest 10th value” in the VO₂ exercise means rounding to multiples of ten, not necessarily one decimal place. In R, exact halfway cases use ties-to-even, so do not assume all halfway values round upward.

Write a function you can reuse

absolute_vo2 <- function(relative_vo2, mass_kg) {
  relative_vo2 * mass_kg / 1000
}
absolute_vo2(50, 70)          # 3.5 L/min

Relative VO₂ has units mL/kg/min. Multiplying by kg produces mL/min; dividing by 1,000 produces L/min. The slide example multiplies and divides by 1,000. State the output units. Functions can operate on whole vectors here, with corresponding elements multiplied. The last evaluated expression is returned; explicit return(...) also works.

Conditions: choose a branch

The comparison operators are <, >, <=, >=, ==, !=. Use == to compare equality; <- assigns. A scalar if condition must resolve to one nonmissing logical value.

height_label <- function(h) {
  if (h >= 190) {
    "very tall"
  } else if (h >= 175) {
    "tall"
  } else if (h >= 160) {
    "average"
  } else {
    "short"
  }
}
height_label(182)             # "tall"

Order matters. Test the highest threshold first. Once a branch is selected, subsequent branches are skipped. The second branch covers 175 up to but excluding 190, because the first branch already handled 190 and above. These are the exercise's illustrative labels, not universal physiological classifications.

Loops: repeat an instruction

for (h in height) {
  print(height_label(h))
}

Each iteration takes the next height, runs the same function, and prints its label. for (i in 1:4) iterates over indices instead. You can store results rather than only print them:

labels <- character(length(height))
for (i in seq_along(height)) {
  labels[i] <- height_label(height[i])
}

Vectorized functions often avoid an explicit loop, but knowing how to trace a loop helps you understand a script. Do not treat a code screenshot as something to memorize without following what changes at each step.

Libraries/packages

install.packages("ggplot2") downloads and installs a package; library(ggplot2) loads it for the current R session. Installation needs internet unless the files are already available. Once installed, packages work without downloading them again.

The slides show an ecosystem including plotting (ggplot2), wrangling (dplyr), dates (lubridate), reporting and interactive applications. You should recognize why packages exist rather than learn every logo. Lecture 4 develops wrangling and visualization.

What the ten slide exercises teach

  1. Recognize data types.
  2. Create vectors, a matrix and a data frame.
  3. Name elements and print them.
  4. Use mean, sum, median and round.
  5. Handle missing values in summaries.
  6. Round to a coarser increment.
  7. Write a reusable VO₂ function.
  8. Trace an if/else chain.
  9. Repeat it for every group member.
  10. Find relevant packages.

From the recording

Points the lecturer made in the lecture recording on Canvas that are not (fully) on the slides. Times refer to the recording.

  • 03:52 — Reproducibility habits: start scripts with a header (goal, date, author, version) and comment the code; use RStudio projects to restore the workspace; record the R version because packages can break across versions. (not on slides)
  • 10:59 — Integers (e.g. 5L) take less memory than numeric/double values, so calculations on large datasets are faster; the L suffix marks an integer. (Slides p. 10 (shows 2L, not the reason))
  • 11:54 — Single and double quotes are equivalent for character strings in R; the rule is to be consistent. (not on slides)
  • 13:20 — Assign with <-. The = sign also works but is discouraged because = is used to set function arguments (e.g. na.rm = TRUE); rightward assignment (->) is discouraged because readers expect the object name on the left. (Slides p. 11–12 (say 'discouraged'/'Don't do this' without the reason))
  • 19:25 — Data retrieved from websites or APIs usually arrive as a list, which you have to unlist/convert; the data frame is the preferred structure for analysis in this course. (not on slides)
  • 22:08 — In str() output, 'obs.' is the number of rows and 'variables' the number of columns, followed by each column's type; columns read from a CSV may come in as character and must be converted (e.g. as.numeric) before calculating. (Slides p. 23 (output shown, not explained))
  • 36:10 — Name arguments explicitly (matrix(x, nrow = 2, ncol = 2)) so it is clear which value is which; named arguments may be given in any order. (not on slides)
  • 37:17 — R indexing starts at 1, whereas Python starts at 0. (Slides p. 26 (R indexing only))
  • 96:50 — source('file.R') loads functions saved in a separate script into your session, keeping the main script uncluttered. (not on slides)

Slide coverage map

Pages 1–6: motivation and environments; 7–16: values, assignment and inspection; 17–25: containers; 26–29: subsetting; 30–46: functions, naming, summaries, missingness, rounding and custom functions; 47–56: conditions and loops; 57–59: packages. Code screenshots were visually inspected; the explanations above provide readable equivalents.

Worked through the deep dive?Tick it off. Come back to any section whenever you need it.


3Step 3 of 35–8 min

Practice questions

Exactly two original practice questions. These use the concept/application style of the 2024 and 2022–2023 example exams, with new scenarios and values. Try each before opening its answer. The official example exams are on Canvas; save them for a later, exam-style session.

Question 1 — Trace a short script

Without running R, give the values of a, b, and answer. Explain the difference between a and b.

x <- c(12, NA, 18, 30)
a <- mean(x)
b <- mean(x, na.rm = TRUE)
answer <- x[c(1, 4)]
Show answer and rationale

a = NA; b = 20; answer = c(12, 30). The default mean cannot produce a known result with an unknown input. na.rm = TRUE averages the three observed values: (12 + 18 + 30)/3. Indexing begins at 1, so elements 1 and 4 are selected. Removing NA from a summary does not establish the missing value.

Question 2 — Write a reusable function

Write an R function that takes a distance in kilometres and a duration in minutes, and returns average speed in km/h. What should it return for 6 km in 30 minutes? Assume positive duration.

Show answer and rationale
speed_kmh <- function(distance_km, duration_min) {
  distance_km / (duration_min / 60)
}
speed_kmh(6, 30)  # 12

12 km/h. Duration is converted to hours before division. Multiplying distance/duration by 60 is equivalent. The rationale is to preserve the requested units, not just obtain a plausible number.

Return to overview · Detailed explanation

Tried both questions?Answer before peeking, then rate yourself honestly.