RNAcentral is EMBL-EBI’s open resource for non-coding RNA (ncRNA) sequences. It brings together data from 57 expert databases of the RNAcentral Consortium, presenting them under stable identifiers with a common set of annotations, cross-references and analysis tools. This gives researchers a single point of access to the ncRNA sequence space, whether they are looking up an individual transcript or building a dataset for large-scale analysis.
RNAcentral 27 released
What’s new in RNAcentral?
New data and literature updates
Release 27 brings RNAcentral to more than 55 million sequences, an increase of roughly 10 million since the previous release. For the first time the resource includes substantial circular RNA sequences, through the expert databases CIRCpedia and circAtlas. It also welcomes mirtronDB, which provides mirtrons, a subclass of miRNAs derived from small introns, and JaponicusDB, which contributes sequences from the fission yeast Schizosaccharomyces japonicus. Alongside these additions, more than 25 existing member databases have been updated, including ENA, Rfam, Ensembl and RefSeq.
Gene-level annotation has expanded considerably. The previous release described around 107,000 non-coding RNA genes across 203 species; release 27 covers more than one million genes across over 800 species, with LitSumm AI-generated transcript summaries now shown on gene pages as well.
RNAcentral’s literature component, LitScan, has also grown. It now scans just over 23 million identifiers across more than 1.5 million papers, with tailored queries added for TarBase, LncBook and DictyBase so that gene synonyms and database accessions specific to those resources are matched correctly.
Assessing coding potential
Distinguishing genuine non-coding RNAs from unannotated protein-coding sequences remains a persistent problem for any ncRNA resource. Release 27 strengthens RNAcentral’s quality control with two additional coding-potential tools alongside the existing check.
To make these results easy to interpret, a quality-control summary table now appears in the QC Status section of every sequence page, reporting coding potential, contamination and Rfam checks at a glance.
AI-ready data access
Several changes in this release make RNAcentral data more AI-ready. Search results can now be exported directly as Hugging Face datasets in Parquet format, with an accompanying dataset README, so that a set of sequences can be taken straight into a machine-learning workflow without manual conversion. We also provide all active sequences in a single HuggingFace dataset for this release.
This release additionally introduces the RNAcentral MCP Server, which exposes the database as a set of tools that large language models can call directly. Built on the Model Context Protocol, it allows AI assistants to query sequences, map identifiers, find ncRNAs overlapping a genomic region, retrieve secondary structures and access literature summaries. It is available on PyPI and is listed in BioContextAI, a community effort connecting agentic AI to biomedical resources.
Website and sequence search improvements
Sequence search, one of RNAcentral’s most heavily used features, now runs on EMBL-EBI’s Job Dispatcher service. Moving the search onto shared institute infrastructure has made it both quicker and more robust under load, and RNAcentral’s nhmmer-based search now appears in the Job Dispatcher tools listing alongside EMBL-EBI’s other search and analysis services.
The embedded genome browser has gained a new layer of data. RNA modification sites from Sci-ModoM are now available as a track hub in IGV, offering a holistic view of the epitranscriptomic landscape. This new track hub covers nearly eight million reported sites for human and mouse, including data from the Human RNome Project. Navigating the browser is easier too: the transcript you are viewing is highlighted in yellow, so it stands out from the other transcripts sharing the same genomic region rather than blending into them. Finally, a new target filter on the Relationships tab lets you jump straight to specific interactions instead of paging through hundreds of entries.
The complete set of new features can be found in our blog.
Edit