Research

Current Research Interests

High thoughput technology is now routinely used to generate high dimenionsal (HD) data that allow for hundreds, thousands or even millions of null hypotheses to be tested simultaneously using a handful of samples. For example, in the “omics” the goal may be determine which among thousands of genes, bacteria, proteins, etc. are associated with a covariate or treatment group, or which associations are replicated across multiple groups. This problem is challenging for several reasons. To name a few, first data may be discrete and/or difficult model. Second, signals are often weak and sparse (few and far between). Finally, we must adjust for multiplicity to safeguard against high false discovery rates. The problem is fun because solutions are in high demand, accessible and in general permit a lot creativity. I’m also recently interested in replicability analysis, where the goal is to identify phenomena that are replicated across multiple studies or subpopulations, and some portions of data may be subject to selection bias (say P<.05).

You can find full lists of my publications, and some presentations based on recent research with students and other collaborators below. The tutorial is a great way to get familiar with false discovery rate adjustment methods for HD count data and R software. I hope you enjoy!

Selected Recent Presentations

Presentation for South Carolina 40th Anniversary Event

Tutorial: FDR Adjustment Methods for a Wheat Microbiome Data Analysis with R

  • Please do not assume the R code in these slides constitutes a sound analysis for your data. Code is for educational purposes.

Statistical Methods for HD Count Data

Publication Policies for Replicable Research

Lists of Publications

Experts Directory | google scholar | My CV

Keywords

False Discovery Rate; Multiple Hypothesis Testing; Selective Inference; Categorical Data Analysis; High Dimensional Data Analysis; Missing Data Methods; Computational Statistics; Applications in Omics