Divide-and-shrink: An efficient and heterogeneity-agnostic approach for transfer estimation using summary statistics
Abstract: Knowledge transfer across data sources holds great promise for improving the estimation of target population parameters by leveraging the growing availability of data from different sources. However, the effectiveness of knowledge transfer is often challenged by the complex and pervasive heterogeneity between data sources and the lack of access to individual-level data. This paper proposes the divide-and-shrink (dShrink) method, a transfer estimation method that estimates target population parameters in a closed form using summary statistics from a target population and some external source populations while accounting for population heterogeneity. The dShrink estimator is guaranteed to perform no worse than the estimator based solely on the target population in terms of expected quadratic error under arbitrary population heterogeneity. Moreover, it can achieve substantial improvement when the target and source populations are similar, or the underlying true parameter values are near zero. Notably, it is model-free, requires no user-specified tuning parameters, is robust to various types of heterogeneity between data sources, and applies to a broad range of parameter estimation problems. dShrink remains effective even when the covariance matrix is not accessible for the external summary statistics and offers flexibility in incorporating side information and summary statistics from multiple source populations. Simulations and real data analyses demonstrate the superior performance of the dShrink estimator and its potential as a robust tool for transfer estimation.
Bio: Ruoyu Wang is an Assistant Professor in the Department of Statistics and Data Science at Tsinghua University. Before joining Tsinghua, he was a postdoctoral scholar in Prof. Xihong Lin’s group at Harvard School of Public Health. Dr. Wang specializes in methodology development for data integration problems with biased/heterogeneous data sources and causal inference with unmeasured confounders. His research outputs have been published in prestigious journals in statistics and machine learning, such as Biometrika, Journal of the American Statistical Association, Journal of Machine Learning Research, and Biometrics, and at top conferences in artificial intelligence and machine learning, including CVPR, NeurIPS, and ICLR.
