Brief Bio
I am a professor of computer science at Columbia, where I am a member of the Columbia Core AI Lab (CAIL). I am also a researcher at Apple.
I was previously a research scientist at Google and a visiting researcher at Cruise. I completed my PhD at MIT in 2017 advised by Antonio Torralba and my BS at UC Irvine in 2011, where I got my start working with Deva Ramanan.
I received the 2024 PAMI Young Researcher Award and the 2021 NSF CAREER Award. I served as Senior Program Chair for ICLR in 2025 and General Chair in 2026, and currently sit on the board.
Research
By training machines to observe and interact with their surroundings, our research aims to create robust and versatile models for perception. Our lab often investigates visual models that capitalize on large amounts of unlabeled data and transfer across tasks and modalities. Other interests include robotics, interpretable models, and other modalities such as sound, language, and beyond.
The lab recruits one or two PhD students each year. Prospective PhD students should apply to the PhD program. Due to the volume of email we receive, we unfortunately cannot respond to emails about applications.
PhD Students and Postdocs
Graduated PhD Students and Former Postdocs
- Sachit Menon (2026), Member of Technical Staff at Anthropic
- Mia Chiquier (2025), Research Scientist at Mistral
- Utkarsh Mall (2025), Assistant Professor at MBZUAI
- Ruoshi Liu (2025), Assistant Professor at UMD
- Basile Van Hoorick (2024), Research Scientist at TRI
- Dídac Surís (2024), Research Scientist at Meta
- Chengzhi Mao (2023), Assistant Professor at Rutgers
Papers
Our research creates perception systems with diverse skills, including spatial, physical, logical, and reasoning abilities, for flexibly analyzing visual data. Our multimodal approach provides versatile representations for tasks like 3D reconstruction, visual question answering, and robot manipulation, while offering inherent explainability and excellent zero-shot generalization. The below papers highlight key examples of these capabilities.
Recent
Robotics
Multi-modal learning for robotic perception and action.
Interpretability
Explainable-by-construction methods that let people audit and reprogram perception.
AI4Science
Visual methods to accelerate scientific discovery.
Learning from Video
Learning perceptual skills from unlabeled video without manual supervision.
Anticipation and Prediction
Models that forecast future events and actions before they happen.
Multimodal
Cross-modal representations linking vision, sound, and language.
3D
Spatial awareness for 3D reconstruction and physical reasoning.
Robustness
Trustworthy models with strong generalization under distribution shift.
All Papers
Teaching
- Computer Vision II (Summer 2021, Spring 2022-2025)
- Computer Vision I (Fall 2018-2019)
- Advanced Computer Vision (Spring 2019)
- Machine Learning Frontiers (Fall 2024-2025)
- Representation Learning (Fall 2020-2022)
Funding
- National Science Foundation
- Defense Advanced Research Projects Agency
- Toyota Research Institute
- Amazon Research

Arjun Mani
Junbang Liang
Lennart Schulze
Sreehari Rammohan
Sruthi Sudhakar






































