Optimal two-phase sampling design for comparing accuracies of two binary classification rules

Huiping Xu; Siu L Hui; Shaun Grannis

doi:10.1002/sim.5946

Optimal two-phase sampling design for comparing accuracies of two binary classification rules

Stat Med. 2014 Feb 10;33(3):500-13. doi: 10.1002/sim.5946. Epub 2013 Sep 4.

Authors

Huiping Xu¹, Siu L Hui, Shaun Grannis

Affiliation

¹ Department of Biostatistics, Indiana University School of Public Health and School of Medicine, Indianapolis, IN, U.S.A.

PMID: 24038175
DOI: 10.1002/sim.5946

Abstract

In this paper, we consider the design for comparing the performance of two binary classification rules, for example, two record linkage algorithms or two screening tests. Statistical methods are well developed for comparing these accuracy measures when the gold standard is available for every unit in the sample, or in a two-phase study when the gold standard is ascertained only in the second phase in a subsample using a fixed sampling scheme. However, these methods do not attempt to optimize the sampling scheme to minimize the variance of the estimators of interest. In comparing the performance of two classification rules, the parameters of primary interest are the difference in sensitivities, specificities, and positive predictive values. We derived the analytic variance formulas for these parameter estimates and used them to obtain the optimal sampling design. The efficiency of the optimal sampling design is evaluated through an empirical investigation that compares the optimal sampling with simple random sampling and with proportional allocation. Results of the empirical study show that the optimal sampling design is similar for estimating the difference in sensitivities and in specificities, and both achieve a substantial amount of variance reduction with an over-sample of subjects with discordant results and under-sample of subjects with concordant results. A heuristic rule is recommended when there is no prior knowledge of individual sensitivities and specificities, or the prevalence of the true positive findings in the study population. The optimal sampling is applied to a real-world example in record linkage to evaluate the difference in classification accuracy of two matching algorithms.

Keywords: diagnostic accuracy; diagnostic test; positive predicted value; record linkage; sensitivity; specificity; stratified sampling.

Publication types

Research Support, U.S. Gov't, P.H.S.

MeSH terms

Algorithms*
Biometry / methods*
Classification / methods*
Female
Humans
Male
Models, Statistical*
Predictive Value of Tests*

Grants and funding

R01HS018553/HS/AHRQ HHS/United States