Explainable Reaction Atlas
Jump to Discussion
Abstract
Exploring chemical reaction latent spaces is essential for understanding the organization of complex reaction landscapes. Previous work has shown that transformer-based models can generate high-quality latent representations, known as reaction fingerprints, that enable the visualization of reaction spaces as tree-structured atlases. However, interpreting these reaction atlases remains largely manual and time-consuming, limiting their scalability in modern cheminformatics workflows. Although users can readily identify branches, neighborhoods, and transitions, understanding their underlying chemical significance still requires examining numerous individual reactions. We present a framework that combines interactive reaction-atlas visualization with large language model (LLM)-based explanations. Building on reaction atlases, our approach introduces two complementary LLM-driven components: a Neighborhood Explainer, which summarizes and compares local regions of the atlas, and a Path Explainer, which analyzes transitions of reactions with similar latent space representations along paths defined by user-selected endpoints. The framework supports both label-free and label-aware settings, operating either directly on reaction SMILES (Simplified Molecular Input Line Entry System) representations or by incorporating reaction class labels. We demonstrate the framework on the Schneider 50K dataset, a curated benchmark comprising approximately 50,000 atom-mapped chemical reactions, and show that explanation-guided exploration enables users to interpret local neighborhoods, compare related branches, analyze reaction-space transitions, and identify locally inconsistent regions within chemical reaction latent spaces.
