Abstract
Deep learning (DL) systems are increasingly deployed in safety- and security-critical domains including self-driving cars and malware detection, where the correctness and predictability of a system’s behavior for corner case inputs are of great importance. Existing DL testing depends heavily on manually labeled data and therefore often fails to expose erroneous behaviors for rare inputs. We design, implement, and evaluate DeepXplore, the first whitebox framework for systematically testing real-world DL systems. First, we introduce neuron coverage for systematically measuring the parts of a DL system exercised by test inputs. Next, we leverage multiple DL systems with similar functionality as cross-referencing oracles to avoid manual checking. Finally, we demonstrate how finding inputs for DL systems that both trigger many differential behaviors and achieve high neuron coverage can be represented as a joint optimization problem and solved efficiently using gradient-based search techniques. DeepXplore efficiently finds thousands of incorrect corner case behaviors (e.g., self-driving cars crashing into guard rails and malware masquerading as benign software) in state-of-the-art DL models with thousands of neurons trained on five popular datasets including ImageNet and Udacity self-driving challenge data. For all tested DL models, on average, DeepXplore generated one test input demonstrating incorrect behavior within one second while running only on a commodity laptop. We further show that the test inputs generated by DeepXplore can also be used to retrain the corresponding DL model to improve the model’s accuracy by up to 3%.
Resources
Other Versions
These publications report or revisit the same underlying work.
- DeepXplore: Automated Whitebox Testing of Deep Learning Systems Communications of the ACM, Volume 62, Number 11, November, 2019 PDFPublisher BibTeX
- DeepXplore: Automated Whitebox Testing of Deep Learning Systems GetMobile: Mobile Computing and Communications, Volume 22, Number 3, January, 2019 PDFPublisher BibTeX
Recognition
- 2018CSAW 2018 Applied Research Second Place
- 2019CACM Research Highlight
- 2017SOSP Best Paper Award
Coverage
- Scientific American When AI Steers Us Astray
- IEEE Spectrum A New Way to Find Bugs in Self-Driving AI Could Save Lives
- Newsweek Scientists May Have Found a Way to Stop Artificial Intelligence from Becoming Racist and Sexist
- Communications of the ACM CACM Nov. 2019 - DeepXplore: Automated Whitebox Testing of Deep Learning Systems
- TechRadar Self-Driving Cars: Your Complete Guide to Autonomous Vehicles
- The Next Web Science May Have Cured Biased AI
- Columbia News Researchers Unveil Tool to Debug ‘Black Box’ Deep Learning Algorithms
- The Morning Paper DeepXplore: Automated Whitebox Testing of Deep Learning Systems
- The Foretellix CTO Blog DeepXplore and New Ideas for Verifying ML Systems
- EurekAlert First White-Box Testing Model Finds Thousands of Errors in Self-Driving Cars
- EurekAlert Researchers Unveil Tool to Debug ‘Black Box’ Deep Learning Algorithms
- Columbia Engineering Magazine Machine Learning 2.0
- LeiPhone 帮助自动驾驶找出算法中的 Bug,这种方法可堪大任吗?
- ChainNews 学界 | 新研究提出DeepXplore:首个系统性测试现实深度学习系统的白箱框架
- Sohu DeepXplore最新动态:科学有望打开AI“黑盒子”,消除AI偏见
- Sina 深度 | 详解首个系统性测试现实深度学习系统的白箱框架DeepXplore
- CCTV Hello AI Documentary — DeepXplore Segment
- e15.cz DeepXplore hledá chyby a nebezpečí v neznámých útrobách hlubokých neuronových sítí
- N+1 Тестирование нейросетей указало на дефекты в «мышлении» беспилотных автомобилей
- Columbia Engineering Luca Carloni and Junfeng Yang Elected 2025 ACM Fellows