Lin Lab
The Lin Lab, led by Dr. Xihong Lin at the Harvard T.H. Chan School of Public Health, develops and applies integrative statistical, machine-learning, and generative AI methods to generate new insights into the causes, prevention, and treatment of complex diseases. By integrating whole-genome sequencing, multi-omics, and phenotype data from population cohorts and biobanks with functional genomic experimental data, our research spans the full spectrum from genetic variation to biological function, disease phenotypes, and actionable strategies for improving human health. Our work encompasses whole-genome sequencing analysis, variant functional annotation, functional genomics, disease risk prediction, causal inference, gene–environment interactions, and precision health. We also develop scalable, open-access resources and tools that empower researchers to analyze population-scale genomic health data and experimental functional genomics data and accelerate trustworthy scientific discovery.
655 Huntington Ave, Boston, MA 02115
Who We Are
The Lin Lab, directed by Dr. Xihong Lin at the Harvard T.H. Chan School of Public Health, conducts cutting-edge research at the intersection of statistics, ML/generative AI, human genetics and genomics, and health science. Our interdisciplinary team of graduate and undergraduate students, postdoctoral fellows, research scientists, and software developers brings a wide range of expertise to this mission.
We develop and apply scalable, integrative statistical, machine-learning, and generative AI methods to analyze large-scale genetic, genomic, epidemiological, and health data. Our goal is to translate these data into new insights into the biological and environmental determinants of complex diseases and to advance their prevention, diagnosis, and treatment.
Our major research areas include:
- Statistical and computational methods for the analysis of population-scale genetic and genomic data, including whole-genome sequencing studies, multi-omics data, biobanks, electronic health records, gene–environment interactions, disease-risk prediction, multi-phenotype analysis, and heritability estimation;
- Variant functional annotation, and statistical and computational methods for the analysis of functional genomic experimental data;
- Integrative analysis of multimodal data and causal inference, including Mendelian randomization and causal mediation analysis; and
- Statistically principled generative AI methods and tools, including synthetic data approaches, generative AI-enabled analysis, and agentic AI, to accelerate trustworthy scientific discovery.
To support these efforts, we develop scalable, open-access data resources and software tools that empower researchers to analyze population-scale genomic and health data together with experimental functional genomics data. Examples of these resources include FAVOR, STAAR, and STAARpipeline.
Photo Gallery