We develop computational methods to understand high-throughput biological data — building probabilistic models, machine-learning methods, and advanced data representations, and turning them into tools that other scientists can use.
← Back to homeWe build computational tools that turn noisy, heterogeneous microbiome measurements into interpretable biological signal — methods for estimating personalized microbial growth rates, correcting contamination and processing bias across studies, and representing communities as pan-metagenomic graphs.
Much of biology is hierarchical, relational, or otherwise poorly served by flat Euclidean vectors. We develop machine-learning methods built on advanced data representations — including non-Euclidean (hyperbolic) embeddings — and adapt classical tools such as decision trees and random forests to operate natively in these spaces.
These representations support tasks from phylogenetic inference to differential splicing, where the geometry of the representation reflects the structure of the underlying biology.
Single-cell data offer an unprecedented view of tumor biology, but their scale and sparsity demand new computational abstractions. We develop methods for metacell inference (SEACells) that summarize cellular states robustly, and we study how tumors adapt to different organ microenvironments.
Recent work characterizes transcriptomic plasticity as a hallmark of metastatic pancreatic cancer, combining single-cell expression with clonal phylogeny to link genotype, environment, and phenotype.
A longstanding direction of the lab is computational human genetics: inferring shared ancestry and recent relatedness from genomes, and using that structure to map disease. We have developed methods for identity-by-descent (IBD) inference, population-structure modeling, and rare-variant association, with particular focus on founder and isolated populations such as the Ashkenazi Jewish population.
This work spans genome-wide association studies (GWAS), demographic inference, and analyses at biobank scale. While the lab's current emphasis has shifted toward microbiome, single-cell, and representation learning, these genetics methods and collaborations remain part of our foundation.
Across all of these areas, our output is reusable methodology. Tools developed in the lab — for contamination correction, bias correction, metacell inference, and geometric learning — are released as open-source software so that other groups can apply and build on them.