405 Academic Research Building
265 South 37th Street
Philadelphia, PA 19104
Research Interests: Statistical machine learning, high-dimensional statistics, large-scale inference, functional data analysis, statistical decision theory, applications to genomics and financial econometrics
Links: CV, Personal Website
Tony Cai is the Daniel H. Silberberg Professor and Professor of Statistics and Data Science at the Wharton School of the University of Pennsylvania. His research develops statistical foundations for modern data science and AI, with emphasis on high-dimensional statistics, statistical machine learning, transfer learning, differential privacy, federated and distributed learning, large-scale inference, nonparametric estimation, and statistical decision theory. His work has advanced both theory and methodology for learning from high-dimensional, heterogeneous, sensitive, and decentralized data, with applications in genomics, medicine, public health, finance, and scientific discovery. Tony Cai is a Fellow of the Institute of Mathematical Statistics (IMS) and the American Association for the Advancement of Science (AAAS), a recipient of the COPSS Presidents’ Award, the Noether Distinguished Scholar Award, and the Leo Breiman Senior Award, and has served as IMS President and Co-Editor of The Annals of Statistics.
PhD, Cornell University, 1996
Academic Positions Held
For more information, go to My Personal Page
Tony Cai’s research develops statistical foundations for modern data science and AI, with current emphasis on transfer learning, differential privacy, federated and distributed learning, causal inference, and high-dimensional statistics. A central theme is reliable learning from heterogeneous, sensitive, and decentralized data: how information from related populations can improve a target analysis, how privacy and communication constraints affect statistical accuracy, and how methods can adapt to new settings without negative transfer. These questions arise naturally in modern scientific and technological applications, where data are often distributed across institutions, populations, studies, and devices, and where reliable conclusions must be drawn while respecting privacy and accounting for heterogeneity.
His recent work develops decision-theoretic frameworks for privacy-preserving and federated estimation, testing, transfer learning, and individualized treatment decisions. These questions are increasingly important for AI, where modern systems must learn from large, distributed, and heterogeneous data sources while respecting privacy, reliability, and resource constraints. His work contributes statistical foundations for trustworthy AI by clarifying when learning is possible, what information is fundamentally required, and how optimal procedures can be designed under such constraints. These ideas are relevant to applications in biomedical research, public health, genomics, finance, decentralized learning systems, and large-scale scientific collaboration.
Tony Cai’s research contributions span several major areas of modern statistics. He has developed influential theory and methodology for high-dimensional covariance and precision-matrix estimation, sparse PCA, graphical models, regression, and high-dimensional testing, as well as for nonparametric estimation, adaptation, and uncertainty quantification. His work on large-scale multiple testing addresses power and false discovery control in complex high-dimensional settings, while contributions to binomial confidence intervals, singular-subspace perturbation theory, and the theoretical analysis of t-SNE have provided widely used tools and benchmarks. Many of these contributions also support modern AI and machine learning, particularly through their treatment of high-dimensional structure, spectral methods, dimension reduction, uncertainty, and reliable inference from complex data.
His work has had substantial impact on applied science. The Brown–Cai–DasGupta paper on binomial confidence intervals, for example, has received more than 4,800 citations and is widely used in medicine, public health, clinical trials, epidemiology, quality control, genetics, and related fields. His research has contributed significantly to genomics, including large-scale multiple testing, differential co-expression analysis, and gene-network inference, where rigorous statistical methods are essential for reliable scientific discovery. Across these areas, Tony Cai’s work combines sharp theory with practically motivated methodology, developing statistically optimal and adaptive procedures together with a precise understanding of the limits of what can be achieved under high dimensionality, structural complexity, privacy, communication, and heterogeneity. Taken together, his research has influenced statistical theory, machine learning, biomedical science, genomics, and data-driven scientific discovery.
Tony Cai, Abhinav Chakraborty, Lasse Vuursteen (2026), Optimal federated learning for nonparametric regression with heterogenous distributed differential privacy constraints, Journal of the American Statistical Association.
Tony Cai, Abhinav Chakraborty, Yichen Wang (2026), Optimal differentially private ranking from pairwise comparisons, Journal of the American Statistical Association.
Abhinav Chakraborty, Arnab Auddy, Tony Cai (2026), When data can’t meet: Estimating correlation across privacy barriers, Advances in Neural Information Processing Systems , 38 (), pp. 95086-95120.
Arnab Auddy, Tony Cai, Abhinav Chakraborty (2026), Minimax and adaptive transfer learning for nonparametric classification under distributed differential privacy constraints, Journal of the Royal Statistical Society, Series B , 27.
Abstract: This paper considers minimax and adaptive transfer learning for nonparametric classification under the posterior drift model with distributed differential privacy constraints. Our study is conducted within a heterogeneous framework, encompassing diverse sample sizes, varying privacy parameters, and data heterogeneity across different servers. We first establish the minimax misclassification rate, precisely characterizing the effects of privacy constraints, source samples, and target samples on classification accuracy. The results reveal interesting phase transition phenomena and highlight the intricate trade-offs between preserving privacy and achieving classification accuracy. We then develop a data-driven adaptive classifier that achieves the optimal rate within a logarithmic factor across a large collection of parameter spaces while satisfying the same set of differential privacy constraints. Simulation studies and real-world data applications further elucidate the theoretical analysis with numerical results.
Tony Cai, Tracy Ke, Paxton Turner (2024), Testing high-dimensional multinomials with applications to text analysis, Journal of the Royal Statistical Society, Series B , 86 (), pp. 922-942.
Tony Cai and Hongji Wei (2024), Distributed Gaussian Mean Estimation under Communication Constraints: Optimal Rates and Communication-efficient Algorithms, Journal of Machine Learning Research , 63.
Tony Cai, Rungang Han, Anru Zhang (2022), On the Non-asymptotic Concentration of Heteroskedastic Wishart-type Matrix, Electronic Journal of Probability, 27 (), pp. 1-40.
Anru Zhang, Tony Cai, Yihong Wu (2022), Heteroskedastic PCA: Algorithm, Optimality, and Applications, Annals of Statistics, 50 (1), pp. 53-80.
Zijian Guo, Claude Renaux, Peter Bühlmann, Tony Cai (2021), Group Inference in High Dimensions with Applications to Hierarchical Testing, Electronic Journal of Statistics, 15 (), pp. 6633-6676.
Rong Ma, Tony Cai, Hongzhe Li (2020), Global and Simultaneous Hypothesis Testing for High-Dimensional Logistic Regression Models, Journal of the American Statistical Association, (to appear) ().
Discrete and continuous sample spaces and probability; random variables, distributions, independence; expectation and generating functions; Markov chains and recurrence theory.
STAT4300001 ( Syllabus )
STAT4300002 ( Syllabus )
Independent Study allows students to pursue academic interests not available in regularly offered courses. Students must consult with their academic advisor to formulate a project directly related to the student’s research interests. All independent study courses are subject to the approval of the AMCS Graduate Group Chair.
Allows for a PhD student to be enrolled full-time to work exclusively on research, writing and preparing his/her doctoral thesis and defense. All required coursework (20 CUs) must be completed, and the student must have passed his/her thesis proposal/oral candidacy examination prior to being enrolled.
For students writing a Master's Thesis to fulfill the program's requirements. All required coursework (8 CUs) must be completed prior to being enrolled.
Study under the direction of a faculty member.
Discrete and continuous sample spaces and probability; random variables, distributions, independence; expectation and generating functions; Markov chains and recurrence theory.
Elements of matrix algebra. Discrete and continuous random variables and their distributions. Moments and moment generating functions. Joint distributions. Functions and transformations of random variables. Law of large numbers and the central limit theorem. Point estimation: sufficiency, maximum likelihood, minimum variance. Confidence intervals. A one-year course in calculus is recommended.
A continuation of STAT 9700.
This seminar is for graduate students who wish to learn about current research frontiers. It covers advanced topics in probability, statistical theory and methods, applied statistics, data science and artificial intelligence. Specific topics vary from year to year and emphasize both theoretical foundations and applications.
Dissertation
Think you could still ace your way through Wharton? Well, here’s your chance to prove it.
Wharton Magazine - 09/01/2010