Improving metadata for genomic sequences
Summary
- The European Nucleotide Archive is set to introduce mandatory spatio-temporal information for all new samples
- The change aims to enrich the scientific value of the data, especially for scientists working in the areas of infectious disease, biodiversity and ecology
- The team is seeking input from the community
1 December, Cambridge – The European Nucleotide Archive (ENA) is an immense collection of genomic sequence data from all species, used by life scientists all over the world. Because the ENA is open access, any researcher can submit the data they have generated, and share it with the global scientific community. They can also access vast amounts of data generated by others.
The ENA also underpins a number of other data resources, including:
- European COVID-19 Data Platform – one of the world’s largest collections of molecular data about the SARs-CoV-2 virus which causes COVID-19,
- Darwin Tree of Life Data Portal, which provides access to data emerging from comprehensive sequencing of all species in Great Britain and Ireland
- MGnify – the data resource for the analysis and exploration of microbiome data.
Rich and relevant data
To enrich the scientific value of the data it archives, the ENA will introduce mandatory spatio-temporal metadata for new data submissions. The drive for this change has come from the user community, specifically from researchers working in the fields of infectious disease, ecology and biodiversity, for whom this information is particularly valuable because it allows them to correlate genotypes to specific environmental conditions.
This change will serve over time to increase the level to which datasets in ENA are Findable, Accessible, Interoperable and Reusable (FAIR).
Over the coming year, all new submissions to ENA will require the following information for each sample:
- The country or region where the sample was collected, using standardised country names from a controlled list.
- The collection date of the sample, recording at least the year of collection.
This change will only apply to new submissions, and will not be done retrospectively. The goal is to ensure this exists for all new sequences by the end of 2022.
“The pandemic has really galvanised the sequencing community and has shown the value of metadata for fields like virology and public health,” explained Guy Cochrane, Head of ENA at EMBL-EBI. "Knowing when and where a sample was taken is essential for understanding the emergence of a pathogen, the spread of disease and guides vaccines, treatments and public health interventions. In parallel, researchers working on biodiversity and ecology are also requiring this information more often to intersect genomic features with data beyond the molecular, such as biodiversity observation and climate data.
“The ENA is now adapting its data submission process to the needs of these communities. Open data underpin a huge proportion of life science research, but if we want data to be findable, reusable and truly valuable, we need rich and relevant metadata. We want to work with our submitters to make any changes as seamless as possible.”
Good news for the scientific community
This change is aligned with the International Nucleotide Sequence Database Consortium’s objective to significantly increase the number of sequences for which the origin of the sample can be precisely located in time and space.
The INSDC is a long-standing foundational initiative that operates between the DNA Data Bank of Japan, EMBL-EBI in Europe and the National Center for Biotechnology Information (NCBI) in the United States. INSDC databases contain raw genomic data, alignments and assemblies, as well as functional annotation, enriched with contextual information relating to samples and experimental configurations.
Although the spatio-temporal information will become mandatory in most cases, some exceptions will be allowed when it is deemed necessary and the exception indicated to users.
“Genomics was an incredibly powerful tool in the pandemic response,” explained Kostas Glinos, Head of Open Science Unit at the European Commission. “To fully maximise its potential in tackling infectious diseases and future pandemics, we need to ensure researchers can easily find and analyse relevant datasets. The INSDC’s commitment to introduce spatio-temporal information for new data submissions is a step in the right direction.”
“Biodiversity data from the non-model organisms is improved if spatiotemporal parameters accompany the DNA data,” said Joe Miller, Director at the Global Biodiversity Information Facility (GBIF). “GBIF welcomes the new INSDC policy and encourages all current and future contributors of DNA sequences to INSDC to provide full metadata including coordinates and date of collection. This will support discoverability, reuse and citation of DNA derived data thanks to recently improved ENA to GBIF linkages.”
"This move by the INSDC represents a win-win for science and policymakers,” said Amber Hartman Scholz, Deputy Director of the Leibniz-Institut DSMZ German Collection of Microorganisms and Cell Cultures. “Increasing the transparency of the country of origin for our samples builds trust with partner countries and simultaneously improves the quality of the data. The geographical and temporal information also help users to ensure compliance with existing national laws and requirements, especially around access and benefit-sharing. I applaud INSDC for taking this concrete step to engage the scientific community to improve the scientific record on digital sequence information."
“The devil is in the detail and the new system has to work for our data submitters,” continued Cochrane. “We encourage users to let us know whether these measures will be feasible for them and how they would like to see this happen.”
ENA users are invited to submit their feedback to ena-collaborations@ebi.ac.uk.
Further details regarding the change will be released on or before 1 April 2022.
Edit