Pe'er Lab — Statistical & Population Genetics

HomeResearch › Statistical & Population Genetics
Statistical & Population Genetics

A longstanding direction of the lab: inferring shared ancestry and recent relatedness from genomes, and using that structure to study demography and map disease — with particular depth in identity-by-descent and founder populations.

← Back to research
Pe'er Lab logo
Identity by descent

Two people share a stretch of DNA identical by descent (IBD) when both inherited it, unbroken by recombination, from a common ancestor — so the length and abundance of shared segments record how recently, and how often, individuals and populations share ancestry. The Pe'er lab developed foundational theory as well as practical methodology for the analysis of IBD across large cohorts.

On the theory side, we showed that the length distribution of IBD segments is a quantitative readout of fine-scale demographic history, turning observed sharing into estimates of founder events, bottlenecks, and expansions. We then characterized the statistics of IBD itself — its variance under the Wright–Fisher model, and a renewal-theory description of how sharing accrues along the genome.

On the methodology side, we introduced GERMLINE, which first made genome-wide detection of long IBD segments — and thus whole-population mapping of hidden relatedness — computationally tractable, and characterized the architecture of the long-range haplotypes it detects. Building on this foundation, IBD became a tool across many problems; highlighted applications include, for instance, association mapping (DASH), homozygosity mapping in exomes, inference of historical migration and population structure, estimation of human mutation and gene-conversion rates, HLA typing, and — in a widely noted application to forensics — re-identification of genomic data through long-range familial searches.

Theory
Palamara PF, Lencz T, Darvasi A, Pe'er I · Am. J. Hum. Genet. 91(5):809–822 (2012)
Shows the distribution of IBD-segment lengths is a quantitative record of recent demography — founder events, bottlenecks, expansions.
Carmi S, Palamara PF, Vacic V, Lencz T, Darvasi A, Pe'er I · Genetics 193(3):911–928 (2013)
Derives the variance of IBD sharing under the Wright–Fisher model.
Carmi S, Wilton PR, Wakeley J, Pe'er I · Theor. Popul. Biol. 97:35–48 (2014)
Casts IBD sharing along the genome as a renewal process.
Methodology & applications
Gusev A, Lowe JK, Stoffel M, Daly MJ, Altshuler D, Breslow JL, et al. · Genome Res. 19(2):318–326 (2009)
GERMLINE: the algorithm that first made genome-wide detection of long IBD segments computationally tractable at cohort scale.
Setty MN, Gusev A, Pe'er I · J. Comput. Biol. 18(3):483–493 (2011)
Infers HLA type from haplotypes shared identical-by-descent.
Gusev A, Kenny EE, Lowe JK, Salit J, Saxena R, Kathiresan S, et al. · Am. J. Hum. Genet. 88(6):706–717 (2011)
Carries IBD into association mapping, recovering signal from recent variation. Application: GWAS.
Gusev A, Palamara PF, Aponte G, Zhuang Z, Darvasi A, Gregersen P, et al. · Mol. Biol. Evol. 29(2):473–486 (2012)
Characterizes how long shared haplotypes are distributed within and across populations — the empirical basis for IBD inference.
Zhuang Z, Gusev A, Cho J, Pe'er I · PLoS One 7(10):e47618 (2012)
Extends IBD detection and homozygosity mapping to whole-exome sequencing data.
Palamara PF, Pe'er I · Bioinformatics 29(13):i180–i188 (2013)
Turns haplotype sharing into estimates of historical migration. Application: population structure.
Palamara PF, Francioli LC, Wilton PR, Genovese G, Gusev A, et al. · Am. J. Hum. Genet. 97(6):775–789 (2015)
Uses distant relatedness across many genomes to estimate human mutation and gene-conversion rates.
Yang S, Carmi S, Pe'er I · J. Comput. Biol. 23(6):495–507 (2016)
An efficient method to register IBD segments across ancestral recombination graphs.
Erlich Y, Shor T, Pe'er I, Carmi S · Science 362(6415):690–694 (2018)
Shows long-range familial searches can re-identify genomic data through distant relatives. High-impact application: forensics.
Genetics of isolated populations

Founder and isolated populations are powerful settings for genetics: demographic bottlenecks enrich otherwise-rare variants and lengthen shared haplotypes, sharpening the mapping of both Mendelian and complex disease.

Much of the lab's work here centers on Ashkenazi Jewish genetics. We helped build the population's genomic infrastructure — for instance, characterizing Jewish diaspora structure and shared Middle Eastern ancestry, and, with Shai Carmi, producing a high-depth whole-genome Ashkenazi reference panel that improves interpretation of personal genomes and clarifies the population's recent history. On that foundation we mapped disease in the Ashkenazi setting, with associations spanning, for example, Crohn's disease, Parkinson's disease, and schizophrenia and bipolar disorder — including ultra-rare exonic variants implicating cadherin genes in schizophrenia.

The same principles extend beyond the Ashkenazi population. The lab also worked on the isolated population of Kosrae, Micronesia, where a small founder population and pervasive relatedness power genome-wide association and IBD-based mapping of metabolic and lipid traits.

Ashkenazi Jewish genetics — population infrastructure
Atzmon G, Hao L, Pe'er I, Velez C, Pearlman A, Palamara PF, Morrow B, et al. · Am. J. Hum. Genet. 86(6):850–859 (2010)
Maps the genetic structure of Jewish diaspora populations and their shared Middle Eastern ancestry.
Campbell CL, Palamara PF, Dubrovsky M, Botigué LR, Fellous M, et al. · Proc. Natl. Acad. Sci. 109(34):13865–13870 (2012)
Shows North African Jewish and non-Jewish populations form distinct, orthogonal genetic clusters.
Velez C, Palamara PF, Guevara-Aguirre J, Hao L, Karafet T, et al. · Hum. Genet. 131(2):251–263 (2012)
Traces the genetic legacy of Converso Jews in the genomes of modern Latin Americans.
Guha S, Rosenfeld JA, Malhotra AK, Lee AT, Gregersen PK, Kane JM, et al. · Genome Biol. 13(1):R2 (2012)
Reads health- and disease-relevant signals from the Ashkenazi Jewish genetic signature.
Carmi S, Hui KY, Kochav E, Liu X, Xue J, Grady F, Guha S, Upadhyay K, et al. · Nat. Commun. 5:4835 (2014)
A high-depth Ashkenazi reference panel that sharpens personal-genome interpretation and reconstructs the population's origins.
Xue J, Lencz T, Darvasi A, Pe'er I, Carmi S · PLoS Genet. 13(4):e1006644 (2017)
Dates and localizes the European admixture event in Ashkenazi Jewish history.
Lencz T, Yu J, Palmer C, Carmi S, Ben-Avraham D, Barzilai N, et al. · Hum. Genet. 137(4):343–355 (2018)
A high-depth Ashkenazi WGS reference panel enhancing variant sensitivity, accuracy, and imputation.
Ashkenazi Jewish genetics — disease associations
Kenny EE, Pe'er I, Karban A, Ozelius L, Mitchell AA, Ng SM, Erazo M, et al. · PLoS Genet. 8(3):e1002559 (2012)
Founder-population GWAS identifying new Crohn's-disease susceptibility loci in Ashkenazi Jews.
Lencz T, Guha S, Liu C, Rosenfeld J, Mukherjee S, DeRosse P, John M, et al. · Nat. Commun. 4:2739 (2013)
Implicates NDST3 in psychiatric disease via founder-population mapping.
Zhang W, Hui KY, Gusev A, Warner N, Ng SME, Ferguson J, Choi M, et al. · Genes Immun. 14(5):310–316 (2013)
Extended-haplotype association finds an Ashkenazi-specific HEATR3 (NF-κB pathway) missense mutation linked to Crohn's disease.
Vacic V, Ozelius LJ, Clark LN, Bar-Shira A, Gana-Weisz M, Gurevich T, et al. · Hum. Mol. Genet. 23(17):4693–4702 (2014)
Uses IBD-segment mapping to find Parkinson's-associated haplotypes in an Ashkenazi cohort — IBD methods applied to disease.
Chuang LS, Villaverde N, Hui KY, Mortha A, Rahman A, Levine AP, et al. · Gastroenterology 151(4):710–723.e2 (2016)
Identifies an Ashkenazi-predominant CSF2RB frameshift that raises Crohn's risk and dampens GM-CSF monocyte signaling.
Baskovich B, Hiraki S, Upadhyay K, Meyer P, Carmi S, Barzilai N, et al. · Genet. Med. 18(5):522–528 (2016)
Defines an expanded carrier-screening panel tailored to the Ashkenazi Jewish population.
Chan RB, Perotte AJ, Zhou B, Liong C, Shorr EJ, Marder KS, Kang UJ, et al. · PLoS One 12(2):e0172348 (2017)
A lipidomic analysis linking elevated plasma GM3 to idiopathic Parkinson's disease.
Hui KY, Fernandez-Hernandez H, Hu J, Schaffner A, Pankratz N, Hsu NY, et al. · Sci. Transl. Med. 10(423):eaai7795 (2018)
Links LRRK2 variants to shared risk across Crohn's and Parkinson's disease.
Lencz T, Yu J, Khan RR, Flaherty E, Carmi S, Lam M, Ben-Avraham D, et al. · Neuron 109(9):1465–1478.e4 (2021)
Leverages the founder setting to detect ultra-rare exonic variants implicating cadherin genes in schizophrenia.
Kosrae, Micronesia
Burkhardt R, Kenny EE, Lowe JK, Birkeland A, Josowitz R, Noel M, Salit J, et al. · Arterioscler. Thromb. Vasc. Biol. 28(11):2078–2084 (2008)
Connects HMGCR variants to LDL-cholesterol via an alternative-splicing mechanism.
Kenny EE, Gusev A, Riegel K, Lütjohann D, Lowe JK, Salit J, Maller JB, et al. · Proc. Natl. Acad. Sci. 106(33):13886–13891 (2009)
Systematic haplotype analysis resolves a complex plasma plant-sterol locus in the Kosrae population.
Lowe JK, Maller JB, Pe'er I, Neale BM, Salit J, Kenny EE, Shea JL, et al. · PLoS Genet. 5(2):e1000365 (2009)
Founding GWAS of the Kosrae founder population for metabolic and related traits.
Kenny EE, Kim M, Gusev A, Lowe JK, Salit J, Smith JG, Kovvali S, et al. · Hum. Mol. Genet. 20(4):827–839 (2011)
Shows mixed models harness pervasive relatedness in Kosrae to map 10 metabolic-trait loci.
Gusev A, Shah MJ, Kenny EE, Ramachandran A, Lowe JK, Salit J, et al. · Genetics 190(2):679–689 (2012)
Uses IBD to impute variants from low-pass sequencing in the Kosrae founder population.
Murdock D, Salit J, Stoffel M, Friedman JM, Pe'er I, Breslow JL, et al. · Obesity 21(9):E421–E427 (2013)
A longitudinal study documenting rising obesity and hyperglycemia in the Micronesian (Kosrae) cohort.
Additional methods in statistical & population genetics

Alongside IBD and founder-population work, the lab has contributed broadly across statistical and population genetics. Highlighted contributions include, for instance, methods for pooling samples for high-throughput resequencing, detecting SNP–SNP interactions, privacy-preserving meta-analysis, calibrating association signals against bias and the Winner's Curse, detecting cohort heterogeneity, and modeling ancestry through time.

Prabhu S, Pe'er I · Genome Res. 19(7):1254–1261 (2009)
A pooling design for high-throughput targeted resequencing that recovers individual genotypes from overlapping pools.
Prabhu S, Pe'er I · Genome Res. 22(11):2230–2240 (2012)
SIXPAC: a group-sampling randomization that finds long-range SNP–SNP interactions with 10×–100× fewer tests than a brute-force scan.
Singh AP, Zafer S, Pe'er I · Pac. Symp. Biocomput. (PSB) 356–367 (2013)
Privacy-preserving meta-analysis of sequencing-based association studies.
Palmer C, Pe'er I · PLoS Genet. 12(6):e1006091 (2016)
Characterizes bias in probabilistic (imputed) genotype data and improves detection via multiple imputation.
Palmer C, Pe'er I · PLoS Genet. 13(7):e1006916 (2017)
Explains why genome-wide-significant associations replicate unevenly, correcting for the Winner's Curse and ancestry-driven LD differences.
Joseph TA, Pe'er I · Am. J. Hum. Genet. 105(2):317–333 (2019)
DyStruct: model-based clustering that infers shared ancestry from time-stamped (ancient-DNA) genotypes, letting allele frequencies drift over time.
Yuan J, Xing H, Lamy AL, et al. · PLoS Genet. 16(9):e1009015 (2020)
Uses correlations between PRS variants to detect heterogeneity across GWAS cohorts.