Lecture 1 – Introduction to data science
Perspective (PNAS)Blei & Smyth (2017) — Science and data science
Blei, D. M., & Smyth, P. (2017). Science and data science. Proceedings of the National Academy of Sciences, 114(33), 8689–8692. https://doi.org/10.1073/pnas.1702076114 · Open paper
Data science is 'the child of statistics and computer science'. Its essence is the effective combination of three perspectives (statistical, computational and human), applied iteratively with domain experts to answer discipline-specific scientific questions.
Question
What is data science, and why should scientists care about it? The authors discuss data science from three perspectives (statistical, computational and human) and argue that the effective combination of all three is the essence of data science.
Why it matters
Scientists in many fields (genomics, social science text archives, astronomy sky surveys) now have abundant data but cannot yet fully use it. Existing statistical and computational methods are not set up for modern problems such as massive datasets, high-dimensional data, necessarily misspecified models and inferring causality. The authors see this tension as the catalyst for the new label 'data science'.
Approach
This is a conceptual essay without data. It builds on Tukey's (1962) broad notion of 'data analysis'. The statistical perspective covers uncertainty, complex and structured data (Bayesian modelling), high-dimensional data (regularisation, machine learning and deep learning for prediction) and causal inference. The computational perspective covers optimisation (iteratively 'climbing' a likelihood), sampling methods (the bootstrap, Markov chain Monte Carlo) and distributed computing. The human perspective is illustrated with a computational neuroscientist who images mouse neurons and works with a data scientist.
Findings
- Data science is the 'child of statistics and computer science': it inherits their methods and blends, refocuses and develops them for modern scientific data analysis, in the spirit of Tukey's broad 'data analysis'.
- Statistical perspective: all datasets involve uncertainty, and statistics is the foundation for reasoning about it. Key subfields are complex/structured data (e.g. Bayesian models), high dimensionality (regularisation; ML such as deep learning for prediction) and causality (correlation vs causation, inference from observational data).
- Computational perspective: this concerns the algorithmic implementation of methods and the trade-off between statistical accuracy and computational resources (time, memory). Examples are optimisation, sampling (bootstrap for confidence intervals, MCMC for Bayesian posteriors) and distributed computing.
- Human perspective: data science cannot be fully automated, because applying the tools needs human judgement and deep domain knowledge. The data scientist works iteratively and collaboratively with the domain expert (possibly one person wearing two 'hats'), cycling through preprocessing, exploration, selection, transformation, analysis, interpretation and communication. Reproducibility and data provenance matter.
- Conclusion: data science is more than the sum of statistics and computer science. It requires weaving both into a larger framework problem by problem, understanding the context of data, taking responsibility for private and public data, and communicating clearly what a dataset can and cannot tell us.
Limitations
- Own inference: it is an opinion/perspective piece, so its claims are argued, not empirically tested.
- Own inference: the examples come from genomics, social science, astronomy and neuroscience, not sport or health, so applying it to movement science is left to the reader.
- Own inference: it names challenges (causality, misspecified models, scale) but stays high-level and gives no concrete procedures for solving them.
Link to the lectures
Slides 49–50 'Data science is …' show the paper and quote its conclusion: data science is more than statistics + computer science; it requires understanding the context of data, appreciating the responsibilities of using private and public data, and clearly communicating what a dataset can and cannot tell us. It frames the Lecture 1 Venn diagram (computer science, math & statistics, domain knowledge; slide 51) and the data science vs statistics slides (53–55).
This is the course's definition of data science (Lecture 1 slides 49–51). It matches the Venn diagram of computer science, statistics and domain knowledge, and its human perspective is the 'domain knowledge' circle. Its iterative cycle (preprocessing → exploration → analysis → interpretation → communication) mirrors the data science lifecycle used from Lecture 3 onwards. Its statistical perspective (uncertainty, causality vs correlation) links to the data science vs statistics slides (Lecture 1 slides 53–55) and the spurious-correlation example (Lecture 4). Its mention of ML for high-dimensional prediction and of the bootstrap foreshadows Lectures 5–7.
Remember
- Perspective (PNAS 2017): data science is 'the child of statistics and computer science'.
- Three perspectives: statistical, computational, human; the essence is combining all three.
- Statistical: uncertainty, complex/structured data, high dimensionality, causality.
- Computational: trade-off between statistical accuracy and computational resources (time, memory); optimisation, sampling (bootstrap, MCMC), distributed computing.
- Human: cannot be fully automated; needs domain knowledge and iterative collaboration with domain experts.
- Data science is a cycle: preprocessing → exploration → selection → transformation → analysis → interpretation → communication.
- Holistic view: understand the data's context, take responsibility for private/public data, communicate what the data can and cannot tell us.
Practice
Q1 · Blei & Smyth – what data science is
According to Blei & Smyth (2017), what is the essence of data science?
Show answer
Answer: C. The authors argue that each perspective is critical but that combining all three is the essence of data science. Option B contradicts their human perspective: data science cannot be fully automated.
Source: Paper p. 8689–8690
Q2 · Blei & Smyth – computational perspective
Which statement best describes what Blei & Smyth call the computational perspective of data science?
Show answer
Answer: A. Computational thinking is about the algorithmic implementation and the accuracy-versus-resources trade-off. Option B is the statistical perspective and option C is the human perspective.
Source: Paper p. 8690–8691
Q3 · Blei & Smyth – human perspective
Why, according to Blei & Smyth (2017), can data science not be fully automated?
Show answer
Answer: D. The human perspective holds that understanding the domain, choosing data, exploring, selecting models and communicating results need judgement and disciplinary knowledge. Option B is false: the authors explicitly mention 'necessarily misspecified models'.
Source: Paper p. 8690–8691
Q4 · Blei & Smyth – why data science emerged
What do Blei & Smyth see as the catalyst for the new label 'data science'?
Show answer
Answer: B. They describe a tension: scientists have abundant data, but classical methods cannot fully exploit it computationally or statistically, and this tension gave rise to 'data science'. Option A is the opposite of their premise of abundant data.
Source: Paper p. 8689–8690
Q5 · Blei & Smyth – three perspectives applied
Blei & Smyth (2017) describe data science as combining statistical, computational and human perspectives. Using a sport or health example of your choice (e.g. a club analysing GPS and injury data of its players), explain what each perspective contributes.
Show model answer
Statistical: models the data while accounting for uncertainty, handles complex or high-dimensional data, and asks whether relationships are causal or only correlational (e.g. does high load cause injury?). Computational: implements the analysis efficiently, balancing accuracy against time and memory (e.g. processing large GPS streams, using optimisation or resampling such as the bootstrap). Human: brings domain knowledge from coaches and sport scientists to choose relevant data and features, interpret results and communicate what the data can and cannot say. The work is iterative and collaborative, and the essence is combining all three.
Source: Paper p. 8690–8691; Lecture 1 Slides p. 49–51
How did your answer compare?