Machine-learning methods built on non-Euclidean and other advanced representations — hyperbolic and mixed-curvature geometry, phylogenetics, and representation learning — applied to genomics and biological data.
← Back to researchMany biological systems have structure that flat, Euclidean representations capture poorly — hierarchies, branches, cycles, and curved manifolds. The lab builds interpretable machine-learning methods that operate natively in curved spaces.
Highlighted work includes, for instance, decision trees and random forests in hyperbolic space — made fast by avoiding Riemannian optimization — and their generalization to mixed-curvature product manifolds; an open-source library (Manify) for non-Euclidean representation learning; and applications such as hyperbolic genome embeddings.
Evolutionary relationships are naturally tree-like, and hyperbolic geometry embeds trees with low distortion. We develop Bayesian phylogenetic inference that operates directly in this space.
Beyond geometry, the lab develops representation-learning methods tailored to biological tasks — for instance, prefix-structured (Matryoshka) embeddings and contrastive objectives for differential splicing.