BioChemGraph: Unifying structural and bioactivity data to accelerate drug discovery
In the era of data-driven biology, integrating information from different resources is essential yet often challenging. The BioChemGraph project addresses this challenge by creating infrastructure that consolidates structural, functional, and biochemical annotations for small molecules and their targets from key resources: the Protein Data Bank (PDB), the ChEMBL database, and the Cambridge Structural Database (CSD).
This collaborative effort builds on the existing community-driven PDBeKnowledge Base resource (PDBe-KB), established in 2018 and maintained by the PDBe team at EMBL-EBI. PDBe-KB focuses on integrating and enriching 3D-structure data with annotations to establish their biological context.
BioChemGraph promotes interoperability and unlocks new research opportunities by adopting common data standards. The project delivers improved findability and accessibility of small molecule annotations through uniform data access mechanisms (such as RESTful APIs, PDBe-KB JSON files and FTP).
The project team has also developed intuitive and user-friendly web interfaces to facilitate quick access to relevant information from these resources, advancing research in drug target validation, development and repurposing.
ChEMBL brings drug bioactivity data to PDBe
ChEMBL is an open data resource of manually curated, high-quality, large-scale bioactivity data from literature or individual depositions for bioactive molecules with drug-like properties. ChEMBL is based at EMBL-EBI and is a Global Core Biodata Resource.
By establishing direct mappings via UniProt Accession and compound InChIKey, we’ve linked more than 17,000 experimentally determined protein-ligand complexes from the PDB to about 39,000 ChEMBL bioactivity records. The ChEMBL records encompass a range of experimental data, including binding affinity, functional assays, inhibition, antagonist, and displacement assays.

The integrated data is presented as a simple report with mean pChEMBL value providing a standardised measure of compound potency derived from the relevant activity data in the ChEMBL database. Higher pChEMBL values indicate greater potency (lower concentration required to achieve the biological effect). The pChEMBL value helps put the compound activities on a consistent scale, making it easier to compare, assess and rank the potency of different compounds across various bioassays (more details about pChEMBL can be found here).
This powerful synergy enables researchers to delve deeper into drug mechanisms, identify promising drug candidates, and explore potential drug repurposing opportunities.
The integrated dataset is now available in FTP, with all bioactivities reported in full report- ‘bioactivity_report_full.tsv’ and a simplified version- ‘bioactivity_report_simple.tsv’ with aggregated values for each target-ligand complex. Besides, direct links to individual data items in the ChEMBL data resource are available to facilitate further exploration.
These bioactivity report files are available to download from our FTP services. ChEMBL and PDBe have collaborated to set up an automatic pipeline for generating these data. As a result, the data will be updated weekly, in sync with the PDBe release every Wednesday at 00.00 UTC.
CCDC data is now linked to PDBe, ChEMBL and other sources via UniChem
The Cambridge Crystallographic Data Centre (CCDC) is a non-profit organisation dedicated to advancing structural science by curating and maintaining the Cambridge Structural Database (CSD), a globally recognised resource containing over 1.3 million experimentally determined 3D crystal structures of small-molecule organic and metal-organic compounds. Additionally, the CCDC develops software and services that enable structural chemistry knowledge derived from the CSD to be applied to research challenges across domains.
To further enhance the applicability of CCDC data, the BioChemGraph initiative has now linked over 235,000 CSD identifiers to their corresponding entries in UniChem, a “universal translator” for chemistry using InChIs to connect chemical structures and their identifiers across various databases. UniChem enables researchers to seamlessly access information about a specific molecule, regardless of their initial search platform.
To enable this integration, the CCDC team developed a protocol that generates reliable InChIs for CSD entries via the CSD Python API, utilising 2D diagrams for connectivity and 3D models for stereochemistry. An automated pipeline ensures that CSD InChIs are shared with the ChEMBL team to support regular updates of the UniChem service.
By integrating CSD identifiers in UniChem, it is now easier to identify the availability of crystal structures that contain a specific molecule of interest. As a result, links to crystal structures in the CSD will be integrated into the PDBe-KB for around 32,000 small molecules present in either ChEMBL or PDBe.

By bridging the gap between small molecule data resources and macromolecular information (PDB), BioChemGraph empowers researchers to explore the intricate relationships between these structural domains. This will ultimately accelerate research in diverse fields, including drug discovery materials science, and advance our understanding of fundamental biological processes.
To understand the data available, visit the UniChem website.
BioChemGraph training material is available to users
To help users navigate and utilise the newly integrated data within BioChemGraph, comprehensive training materials are now available.
These resources are meant to guide researchers through analysing integrated small molecule data from ChEMBL, PDBe, and the CCDC, focusing on addressing complex questions in drug discovery and molecular interactions. The training demonstrates how to leverage the combined data effectively to gain deeper insights into drug mechanisms, identify potential drug targets, and accelerate research in related fields.
You can find the BioChemGraph training material available here.
More data coming
Each of the data points highlighted in this article is powered by the new infrastructure established across the three sites to facilitate seamless data exchange. This setup enables us to provide our users with up-to-date data and regular automatic updates. Additionally, more data will be added as the system expands.
The project team will soon launch PDBe-KB ligand pages, a novel set of web pages that will aggregate small-molecule data and provide biological context for the small molecules in the PDB. These pages will also serve as access points for integrated data from the CCDC, ChEMBL, and the PDB. The team is also integrating residue-level annotation for fragment hotspot maps, generated using CCDC software, for the entire PDB archive into PDBe-KB, which will be released later this year.
Edit