I study, develop and apply novel computational methods to make sense of high‑throughput biological data.
Over the past two decades, the technologies that generate biological data have advanced at a staggering pace. Sequencing and related assays now read DNA, RNA, and entire microbial communities with massive parallelism, turning what was once a trickle of measurements into a torrent. As a result, the rate-limiting step has shifted: producing the data is no longer the bottleneck — analyzing, interpreting, and ultimately understanding it is. My lab builds the computational and statistical methods that close this gap, and applies them to extract reliable biological insight.
This work spans statistical and population genetics — identity-by-descent, founder-population genomics, and the genetic mapping of disease; the microbiome, where we model community dynamics and correct the biases that compromise its analysis; single-cell and cancer genomics, where we build interpretable representations of cellular state; and machine learning, where we develop non-Euclidean and geometric methods whose representations are themselves interpretable. A through-line across all of it is guarding against the artifacts that masquerade as discovery, so that what we report is trustworthy.
I founded the M.S. track in Computational Biology at Columbia and helped establish the undergraduate major in Computational Biology, serving as track advisor. I developed and recurrently teach the field's core courses: