ASSESSING SELECTION BIAS IN REGRESSION COEFFICIENTS ESTIMATED FROM NONPROBABILITY SAMPLES WITH APPLICATIONS TO GENETICS AND DEMOGRAPHIC SURVEYS

Brady T West; Roderick J Little; Rebecca R Andridge; Philip S Boonstra; Erin B Ware; Anita Pandit; Fernanda Alvarado-Leiton

doi:10.1214/21-aoas1453

ASSESSING SELECTION BIAS IN REGRESSION COEFFICIENTS ESTIMATED FROM NONPROBABILITY SAMPLES WITH APPLICATIONS TO GENETICS AND DEMOGRAPHIC SURVEYS

Ann Appl Stat. 2021 Sep;15(3):1556-1581. doi: 10.1214/21-aoas1453. Epub 2021 Sep 23.

Authors

Brady T West¹, Roderick J Little², Rebecca R Andridge³, Philip S Boonstra², Erin B Ware¹, Anita Pandit², Fernanda Alvarado-Leiton⁴

Affiliations

¹ Survey Research Center, Institute for Social Research, University of Michigan.
² Department of Biostatistics, School of Public Health, University of Michigan.
³ Division of Biostatistics, College of Public Health, Ohio State University.
⁴ Michigan Program in Survey and Data Science, Institute for Social Research, University of Michigan.

Abstract

Selection bias is a serious potential problem for inference about relationships of scientific interest based on samples without well-defined probability sampling mechanisms. Motivated by the potential for selection bias in: (a) estimated relationships of polygenic scores (PGSs) with phenotypes in genetic studies of volunteers and (b) estimated differences in subgroup means in surveys of smartphone users, we derive novel measures of selection bias for estimates of the coefficients in linear and probit regression models fitted to nonprobability samples, when aggregate-level auxiliary data are available for the selected sample and the target population. The measures arise from normal pattern-mixture models that allow analysts to examine the sensitivity of their inferences to assumptions about nonignorable selection in these samples. We examine the effectiveness of the proposed measures in a simulation study and then use them to quantify the selection bias in: (a) estimated PGS-phenotype relationships in a large study of volunteers recruited via Facebook and (b) estimated subgroup differences in mean past-year employment duration in a nonprobability sample of low-educated smartphone users. We evaluate the performance of the measures in these applications using benchmark estimates from large probability samples.

Keywords: Linear regression; National Survey of Family Growth; nonprobability samples; polygenic scores; probit regression; selection bias.

Abstract

Grants and funding