Project 001: ED-to-Inpatient Pathways
Who gets admitted from the emergency department, and how long do they stay?
From Emergency Department to Inpatient Care: Admission, Length of Stay and Asthma-Related Risk in Simulated Alberta Administrative Data. Release 2.0, October 2026.
An end-to-end health data science project on one unified set of simulated Alberta administrative data. It follows patients from the emergency department into hospital, asks what is linked to admission and length of stay, then builds a fairer comparison for adults with asthma using propensity-score matching and mixed models. Because the data are simulated, it also checks whether those methods recover the effects built into the data.
Every patient record is simulated. The 30 hospital names are real Alberta emergency departments used as reference labels only; no result describes how those hospitals actually perform.
Adult asthma sub-study, from the same 30,000 visits
From raw extract to model, in ten memos
The project is written the way a consulting analysis would be: one memo per step, each reading the output of the one before it.
Build the data
Import, derive and profile.
Admission and length of stay
The two original questions.
- A03Does time in the ED and triage level relate to admission?Logistic regression
- A04Link admitted ED visits to inpatient staysDeterministic record linkage, 6-hour window
- A05How do ED stay and acuity relate to inpatient length of stay?Linear regression, acute-stay sensitivity
- A06Summary of A00–A05 on release 2.0Findings memo
Adult asthma
Causal and multilevel methods, with hospital clustering.
- A07Admission risk for asthma vs. similar patients at the same hospitalPropensity-score matching, clustered RD/RR
- A08Adjusted admission risk, risk difference and risk ratio on the matched pairsOutcome regression, g-computation
- A09Do asthma admissions, and the asthma effect, differ across Calgary, Edmonton and other regions?Logistic GLMM, hospital random intercepts
A notebook for health data science
I'm Miss V. The V Lab is where I learn health data science properly and then use it: each method gets a long-form guide first, and then a place in a project built on realistic Alberta administrative data.
Everything here is synthetic or public. The data designs follow the official documentation for datasets like NACRS and DAD, and every page and analysis is open on GitHub.
- Methods
- Regression for binary and count outcomes, survival analysis, multilevel models, propensity scores, causal inference, meta-analysis
- Health data
- NACRS, DAD, PIN, practitioner claims, lab and vital statistics, record linkage and cohort design
- Tools
- R, R Markdown, renv, lme4, ggplot2, Plotly, Snowflake SQL, ArcGIS, and my own R package, VLabR
Learning tracks
22 guides, each a single page with worked examples and R code you can run.
Biostatistics
14 guidesFrom foundations to survival analysis, multilevel models and causal inference.
Epidemiology
3 guidesDisease frequency, study designs, bias and modern causal methods.
Data visualization
2 guidesPublication-ready ggplot2 and interactive Plotly charts.
Reproducible reporting
1 guideR Markdown documents that rebuild themselves from the data.
Data platforms and GIS
2 guidesSnowflake for data where it lives; ArcGIS for where patients are.
VLabR and Python
ToolsAn R package for cleaning, summary tables and de-identification, plus Python practice.
Reference pages: Alberta's official health geographies, finding disease clusters with SaTScan and writing math in Markdown and LaTeX.
Recently updated
- Project 001: ED-to-Inpatient PathwaysRelease 2.0: one shared synthetic NACRS and DAD release; memos A00–A09 with asthma propensity matching and hospital GLMMs
- VLabRR package for data cleaning, summary tables and de-identification
- Python LabBeginner Python exercises
- Getting started with ArcGISNew GIS course
- Biostatistics guides14 in-depth statistics guides