New MGnify Proteins web resources launched

EMBL-EBI’s microbiome data resource MGnify produces a valuable trove of protein sequence data through its ongoing analysis of microbiome derived data. Two new MGnify Proteins web resources make this dataset easily accessible and searchable, both for large queries across the entire database…
MGnify makes its protein sequence data more easily accessible.

EMBL-EBI’s microbiome data resource MGnify produces a valuable trove of protein sequence data through its ongoing analysis of microbiome derived data. Two new MGnify Proteins web resources make this dataset easily accessible and searchable, both for large queries across the entire database and for inspecting the details of a proteins’ structure and microbial context in a single place.

The MGnify Proteins database contains 2.5 billion non-redundant protein sequences, and its latest release from mid 2024 clusters these into 718 million clusters.

This database was a key source of information for the revolution in protein-folding prediction which saw Google Deepmind’s AlphaFold and Meta AI’s ESMFold bring high quality structure predictions to a large scale. 

“MGnify is an incredible resource for the scientific community, cataloguing billions of unknown protein sequences. This allowed us to predict structures for hundreds of millions of proteins with AI, and can help us to see deep into the immense diversity of natural proteins at a scale that has not been possible before.”
– Alex Rives, Chief Scientist, EvolutionaryScale (at the time: Research Scientist at Meta)

Since then, the MGnify Proteins database has continued to grow, as have the number of groups making use of it for their own structural research including through the community-wide CASP (Critical Assessment of Structure Prediction) programme.

Beyond its value to developments in structure prediction, the database also serves as a rich resource for proteins’ environmental and genomic contexts. The MGnify Proteins database includes information on which microbiomes (e.g. environments and host species) each protein has been found in, and annotations of Pfam domains. Genomic location information links each protein to the assemblies and genomic coordinates where it was predicted.

The availability and linking of these metadata enable original environment-specific protein research, like hunting for novel plastic-degrading enzymes (plastizymes) to tackle the proliferation of plastic pollution in waste water and marine environments. The MGnify Proteins database is already being used at scale to develop machine learning approaches to find these plastizymes.

“We already know some enzymes that can degrade certain types of plastics, but we want to find better versions, combining data, protein structures and machine learning techniques to determine what makes a good plastic degrading enzyme,” said Rob Finn, Team Leader at EMBL-EBI. “We want to generalise this approach to then tackle different types of plastic that are harder to break down.”

The scale of genomic context information available in MGnify Proteins makes it suitable for building generalised approaches to uncover gene function.

“In our work, we show that genomic context information can be leveraged to learn contextualised functional representations of proteins. MGnify’s database allowed for the large-scale language modelling of metagenomic sequences, uncovering the complex relationship between genomic loci and gene function.”
– Yunha Hwang, co-founder, Tatta Bio

Until now this information has only been distributed in bulk, as a set of flat files suitable for large machine learning projects like these. 

This month, we’ve released two new modes to access the data. In partnership with the Google Cloud Public Datasets Programme, the latest MGnify Proteins database release can now be accessed via BigQuery. This allows users to efficiently search across the proteins and their metadata with a SQL-like query syntax, without having to self-host the entire database. 

Our release of the MGnify Proteins website enables access to the 718 million clustered proteins for the first time – again without downloading the entire database. Proteins can be looked up via their MGYP-prefixed accessions, and are linked from the MGnify Sequence Search results such that users can quickly search a protein query sequence against the database and immediately browse the metadata and structures (where available) of each result. Protein structures (where available) are currently retrieved from the ESM Atlas, following a collaborative effort to release the structures of MGnify Proteins’ 2023_02 data release. ESM Atlas’s open dataset and licensing, along with extensive coverage of the MGnify Proteins sequence clusters, have enabled this integration. The MGnify Proteins web resource is designed such that structure predictions from other resources that meet these criteria could be integrated in future.

The MGnify Proteins database – protein sequence predictions and their metadata – is available for bulk download, can be searched using HMMER, queried via BigQuery, and the new web resource includes structure predictions from ESM Atlas.

These new access modes open up the database to more users, including making it accessible to those without the compute resources needed to work directly with the database in its entirety. They also improve the findability of the proteins’ metadata by providing a URL for the details of each protein cluster. Users can more easily search for, explore, save and share their discoveries from the database, accelerating their research cycles.

“One of the most frequent questions we’re asked by users is how to navigate between sequence data, contextual metadata, and predicted data like protein annotations and structures. The MGnify Proteins web resource we’re launching today addresses this directly by presenting information from these three data types in one place.”

MGnify Proteins, now in public beta, can be accessed at www.ebi.ac.uk/metagenomics/proteins 

The dataset can now also be queried using BigQuery at: https://console.cloud.google.com/marketplace/product/bigquery-public-data/ebi-mgnify 

Edit

Related links

Tags: ai, artificial intelligence, big data, bioinformatics, data resources, embl-ebi, FAIR data, metagenomics, mgnify, open data, protein sequence, services,