Rakchhya Uprety, Faith Ogini, La'Marcus T Wingate
Propensity score matching (PSM) is a statistical method that is used to reduce bias due to confounding from observational studies or studies where randomization is not possible by matching the treatment groups to the control groups based on their propensity scores. A propensity score is a value that gives the probability of each individual subject being assigned to a specific treatment group based on the covariates selected. The aim of this paper is to provide a step-by-step tutorial that demonstrates how to create a propensity-matched cohort using RStudio (Posit PBC, Boston, MA). For this tutorial, we utilize the publicly available 2022 National Health Interview Survey adult dataset. Our focus is on individuals residing in non-metropolitan areas. In our study, the propensity scores represent the probability that a patient uses telehealth, and it is estimated using multiple logistic regression, where telehealth usage is the dependent variable and the predictor variables are age, sex, race, income-to-poverty ratio, education, marital status, insurance status, communication difficulty, and Hispanic ethnicity. This analysis demonstrates how PSM can improve balance in measured covariates and may help reduce bias from observed confounders. Several covariates in the analysis had standardized mean differences (SMDs) greater than 0.1 before matching. However, after PSM, all the covariates achieved SMDs below 0.1, which indicates a better balance at baseline between groups. As this is a pedagogical tutorial, results should not be interpreted as a nationally representative matched analysis.