Explainable deep learning improves human mental models of self-driving cars
Abstract
Self-driving cars increasingly rely on deep neural networks to achieve human-like driving. The opacity of these black-box planners makes it challenging to accurately anticipate when they will fail, with potentially catastrophic consequences. Although research into interpreting these systems has surged, most of it is confined to simulations or toy setups because of the difficulty of real-world deployment, leaving the practical utility of these techniques unknown. Here, we introduce the Concept-Wrapper Network (CW-Net), a method for faithfully explaining the behaviour of machine-learning-based planners that causally grounds their reasoning in human-interpretable concepts without sacrificing performance. We deploy CW-Net on a real self-driving car and show that the resulting explanations improve the human driver’s mental model of the vehicle, allowing them to better predict its behaviour, particularly in surprising situations. This demonstrates that explainable deep learning integrated into self-driving cars can be both understandable and useful in a realistic deployment setting. We anticipate our method could be applied to other safety-critical systems, such as autonomous drones and robotic surgeons, as well as to other architectures, such as end-to-end learning systems and vision–language–action models. Overall, our study establishes a deployment-validated pathway to interpretability for autonomous agents, which could help make them more transparent and safe.
For More information, please visit:https://www.nature.com/articles/s41586-026-10950-5
Published: 02 September 2026
DOI:10.1038/s41586-026-10950-5
Disclaimer:
Partial content of this page is transferred from the network, only for the use of scientific communication, if there is infringement, please contact us to delete. See the Privacy Policy for more information.