Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CMU 17-803 Empirical Methods — Worked Solutions & Reproducible Analyses

Written research-methods assignments and reproducible data-analysis scripts — an independent, from-scratch implementation of 17-803 — Empirical Methods (Carnegie Mellon University, instructor Bogdan Vasilescu), part of a csdiy.wiki full-catalog build.

status language tests license

Overview

17-803 is CMU's graduate course on empirical research methods for software engineering: formulating research questions and theory, qualitative interviewing and coding, survey and experiment design, statistics, regression and mixed-effects modeling, and quasi-experimental causal inference (interrupted time series, difference-in-differences, synthetic control). It is a research-methods course, so this repository is a solution portfolio with two halves:

  1. Reproducible data-analysis scripts (analyses/) that port the course's hands-on R solution notebooks to Python (pandas + statsmodels) and run on the course's real datasets, reproducing the R reference numbers.
  2. Rigorous written assignments (writeups/) covering the non-computational deliverables (study design, sampling, validity, qualitative coding, surveys, meta-analysis), grounded in the course's actual lecture topics and reading list.

The course publishes worked solutions as rendered R notebooks (e.g. typingSpeed.pdf, GaltonFams.pdf, zscore.pdf, its.pdf); this repo treats those as the ground truth and pins every Python port to their exact numbers with a pytest suite.

Results (measured on this machine — Windows, CPU, Python 3.11)

Every number below is produced by a script in analyses/ and saved under results/. Values match the course's R solution notebooks (verified by tests/).

Analysis What it demonstrates Key measured result (matches R)
typing_speed Simpson's paradox; fixed- & mixed-effects Pooled slope −0.63823, R² 0.4253, RSE 0.6605 (5 sig figs vs R); within-person slope flips positive
regression_galton OLS build-up, standardized β, family LMM childHeight~gender+father+mother: father 0.39284, mother 0.31761, R² 0.6354; family LMM R²m 0.630 / R²c 0.701
zscore_standardization standardized coefficients + VIF scale(YearsEdu) $103.84 vs scale(Knowledge) $63.18/SD; VIFs all < 1.8; collinear-toy VIF 113.70
interrupted_time_series segmented regression + DiD Durbin–Watson ≈0.20 (autocorrelation flagged); ITS-with-control treated effect ≈ −0.25, R² 0.988

Figures (in results/): Simpson's-paradox by-participant scatter, Galton regression-to-the-mean, z-score distributions & standardized β bar chart, ITS single-series segmented fit, and ITS treated-vs-control series.

Simpson's paradox — typing speed

Within each participant (colored lines) faster typing → more typos, yet pooling across participants the relationship reverses — the paradox a naive regression would get backwards.

Implemented assignments

Data-analysis scripts (verified reproductions of the course R notebooks):

  • Typing Speed — Simpson's paradox, fixed-effects, and random-intercept mixed models (analyses/typing_speed/).
  • Galton Families regression — intercept-only → gender → multiple → standardized → family random-effects LMM (analyses/regression_galton/).
  • Standardized regression & multicollinearity — z-scored coefficients and VIF on the Wooldridge WAGE2 data (analyses/zscore_standardization/).
  • Interrupted time series — segmented regression, Durbin–Watson autocorrelation, and ITS-with-control / difference-in-differences (analyses/interrupted_time_series/).

Written research-methods assignments (writeups/):

  • 01 — Research questions & the role of theory.
  • 02 — Interviewing & qualitative coding / thematic analysis (with a full worked thematic analysis).
  • 03 — Survey design & mixed methods.
  • 04 — Designing experiments, power analysis, & the four validity types.
  • 05 — Regression, time series & causal inference (analytical companion to the code).
  • 06 — Literature review, systematic review & meta-analysis.

Project structure

cmu-17803-empirical/
├── analyses/
│   ├── typing_speed/typing_speed.py
│   ├── regression_galton/galton_regression.py
│   ├── zscore_standardization/standardization.py
│   └── interrupted_time_series/its_analysis.py
├── writeups/                     # six written research-methods assignments (md)
├── src/empirical_common.py       # shared data/results/plotting helpers
├── scripts/download_data.py      # fetches course datasets at runtime
├── tests/test_analyses.py        # pins every port to the R reference numbers
├── results/                      # measured outputs: *.txt logs + *.png figures
├── data/                         # datasets land here (gitignored, not redistributed)
├── requirements.txt
└── LICENSE

How to run

# Python repos use the shared csdiy env (Python 3.11):
#   D:\Project\_csdiy\.venv-ml\Scripts\python.exe
python -m pip install -r requirements.txt        # or reuse the shared venv

# 1. fetch the real course datasets (not redistributed in this repo)
python scripts/download_data.py

# 2. run each analysis (writes logs + figures to results/)
python analyses/typing_speed/typing_speed.py
python analyses/regression_galton/galton_regression.py
python analyses/zscore_standardization/standardization.py
python analyses/interrupted_time_series/its_analysis.py

# 3. verify every result against the course's R solution numbers
python -m pytest tests/ -v

Verification

There is no autograder for this research-methods course; the course's own rendered R solution notebooks are the ground truth. Verification here is threefold:

  1. Numeric reproduction. Each analysis prints numbers that match the R notebooks — e.g. the typing-speed pooled fit matches typingSpeed.pdf to five significant figures (intercept −0.09850, slope −0.63823, R² 0.4253); the Galton multiple regression matches GaltonFams.pdf exactly (father 0.39284, mother 0.31761, R² 0.6354); WAGE2 standardized/VIF values match zscore.pdf (VIF collinear toy = 113.6997).
  2. A pytest suite (tests/test_analyses.py, 9 tests passing) that asserts these reference values, so any drift fails CI.
  3. Saved evidence in results/ — text logs of every model summary plus the generated figures.
$ python -m pytest tests/ -v
tests/test_analyses.py::test_typing_pooled_matches_R          PASSED
tests/test_analyses.py::test_typing_within_between_signs      PASSED
tests/test_analyses.py::test_galton_gender_model             PASSED
tests/test_analyses.py::test_galton_multiple_regression      PASSED
tests/test_analyses.py::test_galton_standardized_betas       PASSED
tests/test_analyses.py::test_wage_unstandardized             PASSED
tests/test_analyses.py::test_wage_standardized_edu_vs_exp    PASSED
tests/test_analyses.py::test_wage_vif                        PASSED
tests/test_analyses.py::test_its_with_control_effect         PASSED
======================= 9 passed =======================

Tech stack

Python 3.11 · pandas · numpy · scipy · statsmodels (OLS, MixedLM, VIF, Durbin–Watson) · matplotlib · pytest. The original course solutions are in R (lm, lme4::lmer, car::vif, lmtest::dwtest); this repo re-derives the same results with the Python scientific stack.

Key ideas / what I learned

  • Simpson's paradox is real and dangerous: pooling clustered data can reverse the sign of an effect; mixed-effects models (speed_M between vs. speed_E within) recover the truth.
  • Standardization makes coefficients comparable across scales, and VIF diagnoses when multicollinearity makes a coefficient split meaningless.
  • Regression is descriptive, not causal, unless the design earns causal language — which is why the four validity types and randomization matter.
  • Quasi-experiments (segmented ITS, difference-in-differences, synthetic control) enable causal inference from observational time series when randomization is impossible — and autocorrelation is their signature threat (Durbin–Watson, Newey–West).
  • The same inverse-variance / within-plus-between-variance structure recurs across mixed-effects models and random-effects meta-analysis — one idea, many guises.

Datasets & credits

Datasets are downloaded at runtime by scripts/download_data.py and are not redistributed in this repository:

  • typingSpeed.csv, its_with_control.csv — from the official course repo bvasiles/empirical-methods.
  • GaltonFamilies.csv — public-domain HistData dataset (Galton, 1886), via the Rdatasets mirror.
  • wage2.csv — the Wooldridge WAGE2 econometrics teaching dataset, from the exact source (murraylax.org) the course notebook downloads from.

Based on the assignments and R solution notebooks of 17-803 Empirical Methods by Bogdan Vasilescu (Carnegie Mellon University). Course site: https://bvasiles.github.io/empirical-methods/. This repository is an independent educational reimplementation; all course materials, datasets, and specifications belong to their original authors. Original code in this repo is released under the MIT License.

About

Worked solutions and data analyses for CMU 17-803 Empirical Methods for software engineering research — study design, statistics, and reproducible analysis notebooks

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages