Machine learning method identifies evidence for developmental disorders in the G2P database

An extensive collection of peer-reviewed publications describing developmental disorders has been identified and integrated into G2P to help clinicians and researchers better understand the genetic basis of these conditions

Compiling evidence of gene-disease associations from the scientific literature is essential for rare disease diagnosis and research but is time consuming to do at scale.

A new machine learning approach has been developed and employed to search the peer-reviewed literature and identify publications describing patients with specific rare developmental disorders. 

These publications have been integrated into gene-disease models in the G2P database, where they can be easily searched and browsed to provide greater understanding of these complex, genetically heterogeneous conditions.  

Gene disease curation

Multiple reports linking a gene with a disease are required to be confident of the association and enable its use in diagnostic pipelines or as a target for novel therapies. For over a decade clinical and scientific curators have been manually searching the scientific literature to find reports on rare diseases, then extracting and evaluating relevant information to use as the basis of G2P gene-disease association records. 

The identification of relevant literature needs to be performed regularly to seek further evidence which may increase the confidence of existing associations as well as to find new associations. This is a time consuming process which is challenging to do at scale. 

New machine learning method

This new machine learning method, has been developed to scan the scientific literature and identify reports describing the molecular basis of developmental disorders. The pipeline employs a BERT classifier and crossencoder, both fine-tuned to identify descriptions of these rare disorders, to find likely disease candidates and a constrained large language model to map  these candidates to G2P records. Around 38 million abstracts were filtered to identify nearly 69,000 relevant reports. This collection has been integrated into gene disease models in the G2P database, greatly increasing the manually identified 9,000 publications. It can be searched using gene or disease names, providing convenient access to this rich dataset. 

Disease types described in the ~69,000 mined publications now added to G2P, clustered using HDBSCAN; visualisation created with UMAP. Image from Yates et al., 2025

Impact

This work provides an extensive set of highly relevant publications which can improve understanding of more recently described gene disease associations with little curated evidence and highlight novel treatment options for more established conditions. Integration into G2P enables easy discovery and access.

This new method will also allow routine surveillance of the scientific literature to simply identify further evidence for gene disease associations, enabling their confident use in diagnostic screening panels. Work is also ongoing to extract key attributes from each paper to further accelerate the curation process and accelerate rare disease diagnosis, research and therapy development.

Example page in G2P showing mined publications. Image from www.ebi.ac.uk/gene2phenotype/lgd/G2P01943
Edit

Source article(s)

Tags: bioinformatics, data resources, data service, embl-ebi, G2P, genomics,