The system, dubbed Tautomer‑Predictor, was developed by a team led by chemistry professor Yingkai Zhang and postdoctoral researcher Xiaolin Pan. By mining more than 1.1 million experimentally resolved tautomeric states from the Cambridge Structural Database, the team trained a graph‑neural‑network model to infer the most stable tautomer directly from a 2‑D molecular representation, eliminating the need for costly quantum‑mechanical calculations or 3‑D structural input.

In benchmark tests covering crystal‑structure, aqueous‑solution and drug‑like datasets, the model demonstrated strong predictive performance, indicating that it captures chemically meaningful rules governing tautomer stability. When applied to 5,075 ligands from the PDBbind collection—protein‑ligand complexes stored in the Protein Data Bank—the AI identified 126 cases (about 2.5 %) where the originally assigned tautomer was likely incorrect, often resulting in more plausible hydrogen‑bonding patterns after reassignment.

Correct tautomer assignment is critical for molecular modeling, docking, free‑energy calculations and virtual screening. Even a single hydrogen shift can alter how a molecule interacts with a protein target, potentially skewing simulation results and downstream drug‑design decisions. Zhang emphasized that while the protein structures themselves are not necessarily wrong, the chemical representation of the bound ligand may need revision.

The open‑source tool also proved its scalability: processing the 4.6 million‑compound Enamine library in just 3.2 hours on a single GPU‑enabled node. This speed opens the door for large‑scale tautomer correction in commercial and academic drug‑discovery pipelines, where rapid, accurate chemical representation is a prerequisite for reliable computational studies.

The researchers published their findings in the journal Chemical Science under the title “Deep learning of tautomer stability from crystallographic proton positions.” They argue that crystallographic proton placements, long underutilized, constitute a rich source of experimental data that can be leveraged to overcome the persistent challenge of rapid tautomer identification in molecular design.