2026

2025

2024

2023

2022

  • UPGRADVISOR: Early Adopting Dependency Updates Using Hybrid Program Analysis and Hardware Tracing
    Yaniv David, Xudong Sun, Raphael J. Sofaer, Aditya Senthilnathan, Junfeng Yang, Zhiqiang Zuo, Guoqing Harry Xu, Jason Nieh, Ronghui Gu
    OSDI, 2022 PDFBibTeX
    Highlight

    Applications often have fast-paced release schedules, but adoption of software dependency updates can lag by years, leaving applications susceptible to security risks and unexpected breakage.Upgradvisor is a system that largely automates software dependency updates a novel co-designed static analysis and dynamic tracing mechanism to gauge the scope and effect of dependency updates on an application.

  • XRP: In-Kernel Storage Functions with eBPF
    Yuhong Zhong, Haoyu Li, Yu Jian Wu, Ioannis Zarkadas, Jeffrey Tao, Evan Mesterhazy, Michael Makris, Junfeng Yang, Amy Tai, Ryan Stutsman, Asaf Cidon
    OSDI, 2022 PDFBibTeX

    OSDI Best Paper Award

    Highlight

    With the emergence of microsecond-scale NVMe storage devices, the Linux kernel storage stack overhead has become significant, almost doubling access times. XRP is a framework that allows applications to execute user-defined storage functions, such as index lookups or aggregations, from an eBPF hook in the NVMe driver, safely bypassing most of the kernel’s storage stack.

  • NeuDep: Neural Binary Memory Dependence Analysis
    Kexin Pei, Dongdong She, Michael Wang, Scott Geng, Zhou Xuan, Yaniv David, Junfeng Yang, Suman Jana, Baishakhi Ray
    ESEC/FSE, 2022 PDFBibTeX
  • A Tale of Two Models: Constructing Evasive Attacks on Edge Models
    Wei Hao, Aahil Awatramani, Jiayang Hu, Chengzhi Mao, Pin-Chun Chen, Eyal Cidon, Asaf Cidon, Junfeng Yang
    MLSys, 2022 PDFBibTeX
    Highlight

    Full-precision deep learning models are typically compressed, pruned, or quantized to run on edge devices. DIVA is a new evasive attack that exploits these differences in edge adaptation to trick the adapted model running on the edge, but will be virtually undetectable by the original model, which typically serves as the authoritative model version, used for validation, debugging and retraining.

  • Neuroshard: Towards Automatic Multi-objective Sharding with Deep Reinforcement Learning
    Tamer Eldeeb, Zhengneng Chen, Asaf Cidon, Junfeng Yang
    5th International Workshop on Exploiting Artificial Intelligence Techniques for Data Management (aiDM), 2022 PDFBibTeX
    Highlight

    Horizontal sharding is a decades-old technique to scale production databases. Neuroshard is the first system that learns shard assignments directly from the workload using reinforcement learning, and optimizes for multiple sharding objectives simultaneously.

  • Causal Transportability for Visual Recognition
    Chengzhi Mao, Kevin Xia, James Wang, Hao Wang, Junfeng Yang, Elias Bareinboim, Carl Vondrick
    CVPR, 2022 PDFBibTeX
    Highlight

    Visual representations underlie object recognition tasks, but they often contain both robust and non-robust features. We develop an algorithm to estimate the causal effect for image classification, which is transportable (i.e., invariant) across source and target environments.

  • Using Multiple Self-Supervised Tasks Improves Model Robustness
    Matthew Lawhon, Chengzhi Mao, Junfeng Yang
    Workshop on Privacy, Accountability, Interpretability, Robustness, Reasoning on Structured Data (PAIR2Struct), held with ICLR, 2022 PDFBibTeX
    Highlight

    Extends our test-time defense to use multiple self-supervised learning tasks for attack reversal.

2021

  • Adversarial Attacks are Reversible with Natural Supervision
    Chengzhi Mao, Mia Chiquier, Hao Wang, Junfeng Yang, Carl Vondrick
    ICCV, 2021 PDFBibTeX
    Highlight

    We propose a test-time defense that reverses adversarial attacks by restoring the intrinsic structures collatorally damaged by attack vectors. Our defense is still effective even if the attacker is aware of the defense mechanism. Since our defense is deployed during inference instead of training, it is compatible with pre-trained networks as well as most other defenses. Our results suggest deep networks are vulnerable to adversarial examples partly because their representations do not enforce the natural structure of images.

  • StateFormer: Fine-Grained Type Recovery from Binaries Using Generative State Modeling
    Kexin Pei, Jonas Guan, Matthew Broughton, Zhongtian Chen, Songchen Yao, David Williams-King, Vikas Ummadisetty, Junfeng Yang, Baishakhi Ray, Suman Jana
    ESEC/FSE, 2021 PDFBibTeX
    Highlight

    StateFormer recovers source-level types from stripped binaries leveraging transfer learning. Inspired by how human analysts reverse-engineer binaries, we propose a pretraining task called Generative State Modeling (GSM) to teach an ML model assembly code operational semantics, and then transfer the learned knowledge for type inference.

  • Argus: Debugging Performance Issues in Modern Desktop Applications with Annotated Causal Tracing
    Lingmei Weng, Peng Huang, Jason Nieh, Junfeng Yang
    USENIX ATC, 2021 PDFBibTeX

    USENIX ATC Best Paper Award

    Highlight

    We identify inherent imprecision in prior causal tracing work, and build Argus, a novel system that differentiates weak and strong causal edges in the traced graphs for combating this imprecision. Evaluation shows that Argus effectively helps diagnose open spinning-cursor issues in modern MacOS applications.

  • BPF for Storage: An Exokernel-Inspired Approach
    Yuhong Zhong, Hongyi Wang, Yu Jian Wu, Asaf Cidon, Ryan Stutsman, Amy Tai, Junfeng Yang
    HotOS, 2021 PDFBibTeX
    Highlight

    This paper explores an approach inspired by the exokernel file systems to leverage BPF for new, super fast storage devices.

  • Generative Interventions for Causal Learning
    Chengzhi Mao, Augustine Cha, Amogh Gupta, Hao Wang, Junfeng Yang, Carl Vondrick
    CVPR, 2021 PDFBibTeX
    Highlight

    We introduce a framework for learning robust visual representations that generalize to new viewpoints, backgrounds, and scene contexts. Key insight is to steer generative models to manufacture interventions on features caused by confounding factors. Experiments, visualizations, and theoretical results show this method learns robust representations more consistent with the underlying causal relationships.

  • XDA: Accurate, Robust Disassembly with Transfer Learning
    Kexin Pei, Jonas Guan, David Williams-King, Junfeng Yang, Suman Jana
    NDSS, 2021 PDFBibTeX
    Highlight

    XDA is a transfer-learning-based binary disassembler. Key in XDA is to model the problem of going from bits to assembly code as a language translation problem. Evaluation on real-world binaries compiled by four compilers with different optimization levels shows that XDA achieves extremely high accuracy, substantially surpassing the prior art, and is 38 times faster than hand-written disassemblers such as IDA Pro.

2020

  • Multitask Learning Strengthens Adversarial Robustness
    Chengzhi Mao, Amogh Gupta, Vikram Nitin, Baishakhi Ray, Shuran Song, Junfeng Yang, Carl Vondrick
    ECCV, 2020 PDFBibTeX

    ECCV Oral Presentation (top 2%)

    Highlight

    We present both theoretical and empirical analyses that connect the adversarial robustness of a model to the number of tasks that it is trained on. Experiments on two datasets show that attack difficulty increases as the number of target tasks increase. Moreover, our results suggest that when models are trained on multiple tasks at once, they become more robust to adversarial attacks on individual tasks. While adversarial defense remains an open challenge, our results suggest that deep networks are vulnerable partly because they are trained on too few tasks.

  • Lambdata: Optimizing Serverless Computing by Making Data Intents Explicit
    Yang Tang, Junfeng Yang
    13th IEEE International Conference on Cloud Computing (CLOUD), 2020 PDFBibTeX

    IEEE YESC Best Student Paper Award

    Highlight

    Lambadata is a serverless computing system that enables developers to declare a cloud function’s data intents, including both data read and data written. Once data intents are made explicit, Lambadata performs a variety of optimizations to improve speed, including caching data locally and scheduling functions based on code and data locality.

  • Fooling Semantic Segmentation in One Step via Manipulating Nuisance Factors
    Guangyu Shen, Chengzhi Mao, Junfeng Yang, Baishakhi Ray
    Workshop on Adversarial Robustness in the Real World (AROW), held with ECCV, 2020 PDFBibTeX
    Highlight

    We demonstrate a simple and effective method that changes the nuisance factors for a GAN to fool semantic segmentation models in a single step, even if the models have been adversarially trained for defense. It is the first method able to attack semantic segmentation models online, 100x faster than prior attacks. Earlier version of the paper is on arxiv.

  • What does CNN Shift Invariance Look Like? A Visualization Study
    Jake Lee, Junfeng Yang, Zhangyang Wang
    ECCV Workshops, 2020 PDFBibTeX
    Highlight

    We study the shift invariance of representations learned by CNNs and find that (1) state-of-the-art antialiasing improves local shift invariance but not global shift invariance; (2) horizontally shifted images have more similar representations than vertical translation; and (3) learned representations exhibit what we call 'feature arithemetic' properties that addition or substraction in the feature-space visualize to pixel-space addition or subtraction. [Results]

  • Live Trojan Attacks on Deep Neural Networks
    Robby Costales, Chengzhi Mao, Raphael Norwitz, Bryan Kim, Junfeng Yang
    CVPR Workshops, 2020 PDFBibTeX
    Highlight

    Presents a new attack that leverages classic buffer overruns to corrupt the weight matrices of deep neural nets and insert trojans on the fly, causing the nets to mispredict on certain trigger inputs. Key is a regularization method to minimize the amount of buffer overruns needed for the attack.

  • Egalito: Layout-Agnostic Binary Recompilation
    David Williams-King, Hidenori Kobayashi, Kent Williams-King, Graham Patterson, Frank Spano, Yu Jian Wu, Junfeng Yang, Vasileios P. Kemerlis
    ASPLOS, 2020 PDFBibTeX
    Highlight

    Egalito is a binary recompiler that leverages metadata widely present in modern binaries for full disassembly and transformations. We demonstrate nine binary tools including a novel continuous code randomization technique where Egalito transforms itself, and software emulation of the control-flow integrity in upcoming hardware

  • Effective Concurrency Testing for Distributed Systems
    Xinhao Yuan, Junfeng Yang
    ASPLOS, 2020 PDFBibTeX
    Highlight

    Morpheus is the first concurrency testing tool leveraging partial order sampling, a randomized testing method formally analyzed and empirically validated to provide strong probabilistic guarantees of error-detection, for real-world distributed systems. It found previously unknown errors in four Erlang systems including RabbitMQ and Mnesia, 11 total, all of which are flaws in their core protocols that may cause deadlocks, unexpected crashes, or inconsistenlivet states

2019

2018

2017

2016

2015

2014

2013

2012

2011

2010

2009

2008

2006

2005

2004

  • Using Model Checking to Find Serious File System Errors
    Junfeng Yang, Paul Twohey, Dawson Engler, Madanlal Musuvathi
    OSDI, 2004 PDFTalkBibTeX

    OSDI Best Paper Award

    Highlight

    FiSC leveraged model checking to systematically explore crash states beyond conventional testing. It found serious bugs in every file system checked—32 across ext3, JFS, and ReiserFS—including failures that could irrecoverably destroy entire directories, even the file-system root; most were patched within a day.

  • Correlation exploitation in error ranking
    Ted Kremenek, Ken Ashcraft, Junfeng Yang, Dawson Engler
    FSE, 2004 PDFBibTeX
    Highlight

    Describes how we can exploit the correlations of error messages emitted by static analysis tools to cluster false positives together, thus improve the effectiveness of the static analysis tools.

2003

2001