
<!DOCTYPE article
  PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with MathML3 v1.4 20241031//EN" "JATS-archivearticle1-4-mathml3.dtd">
<article xml:lang="EN" article-type="research-article" dtd-version="1.4"><processing-meta base-tagset="archiving" mathml-version="3.0" table-model="xhtml" tagset-family="jats"><restricted-by>pmc</restricted-by></processing-meta><front><journal-meta><journal-id journal-id-type="nlm-ta">J Chem Inf Model</journal-id><journal-id journal-id-type="iso-abbrev">J Chem Inf Model</journal-id><journal-id journal-id-type="pmc-domain-id">822</journal-id><journal-id journal-id-type="pmc-domain">acssd</journal-id><journal-id journal-id-type="nlm-id">101230060</journal-id><journal-id journal-id-type="publisher-id">ci</journal-id><journal-title-group><journal-title>Journal of Chemical Information and Modeling</journal-title></journal-title-group><issn pub-type="ppub">1549-9596</issn><issn pub-type="epub">1549-960X</issn><?publisher_abbrev acs?><custom-meta-group><custom-meta><meta-name>pmc-is-collection-domain</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-collection-title</meta-name><meta-value>ACS AuthorChoice</meta-value></custom-meta></custom-meta-group></journal-meta><article-meta><article-id pub-id-type="pmcid">PMC11094721</article-id><article-id pub-id-type="pmcid-ver">PMC11094721.1</article-id><article-id pub-id-type="pmcaid">11094721</article-id><article-id pub-id-type="pmcaiid">11094721</article-id><article-id pub-id-type="pmid">38648189</article-id><article-id pub-id-type="doi">10.1021/acs.jcim.4c00458</article-id><article-version article-version-type="pmc-version">1</article-version><article-categories><subj-group><subject>Article</subject></subj-group></article-categories><title-group><article-title>DEBFold: Computational
Identification of RNA Secondary
Structures for Sequences across Structural Families Using Deep Learning</article-title></title-group><contrib-group><contrib contrib-type="author" corresp="yes" id="ath1"><contrib-id contrib-id-type="orcid" authenticated="true">https://orcid.org/0000-0001-9420-196X</contrib-id><name name-style="western"><surname>Yang</surname><given-names initials="TH">Tzu-Hsien</given-names></name><xref rid="cor1" ref-type="other">*</xref><xref rid="aff1" ref-type="aff">†</xref><xref rid="aff2" ref-type="aff">‡</xref></contrib><aff id="aff1"><label>†</label>Department
of Biomedical Engineering, <institution>National Cheng
Kung University</institution>, No.1, University Road, Tainan 701, <country>Taiwan</country></aff><aff id="aff2"><label>‡</label>Medical
Device Innovation Center, <institution>National Cheng
Kung University</institution>, No.1,
University Road, Tainan 701, <country>Taiwan</country></aff></contrib-group><author-notes><corresp id="cor1"><label>*</label>Email: <email>thyangza@gs.ncku.edu.tw</email>.</corresp></author-notes><pub-date pub-type="epub"><day>22</day><month>04</month><year>2024</year></pub-date><pub-date pub-type="collection"><day>13</day><month>05</month><year>2024</year></pub-date><volume>64</volume><issue>9</issue><issue-id pub-id-type="pmc-issue-id">462391</issue-id><fpage>3756</fpage><lpage>3766</lpage><history><date date-type="received"><day>15</day><month>03</month><year>2024</year></date><date date-type="accepted"><day>09</day><month>04</month><year>2024</year></date><date date-type="rev-recd"><day>09</day><month>04</month><year>2024</year></date></history><pub-history><event event-type="pmc-release"><date><day>15</day><month>05</month><year>2024</year></date></event><event event-type="pmc-live"><date><day>16</day><month>05</month><year>2024</year></date></event><event event-type="pmc-last-change"><date iso-8601-date="2024-05-16 10:25:13.820"><day>16</day><month>05</month><year>2024</year></date></event></pub-history><permissions><copyright-statement>© 2024 The Author. Published by American Chemical Society</copyright-statement><copyright-year>2024</copyright-year><copyright-holder>The Author</copyright-holder><license><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/" specific-use="textmining" content-type="ccbylicense">https://creativecommons.org/licenses/by/4.0/</ali:license_ref><license-p>Permits the broadest form of re-use including for commercial purposes, provided that author attribution and integrity are maintained (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link>).</license-p></license></permissions><self-uri xmlns:xlink="http://www.w3.org/1999/xlink" content-type="pmc-pdf" xlink:href="ci4c00458.pdf"><?pdf-name ci4c00458.pdf?><?pdf-size 2717783?><?pdf-md5 ef8beaab96f9d65f6cdbcc22c7614ba6?><?pdf-image-server-status NEVER_LOAD?><?pdf-cloudpmc-urn urn:app:08a5/11094721/ef8beaab96f9/ci4c00458.pdf?></self-uri><abstract><p content-type="toc-graphic"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" id="ab-tgr1" position="float" orientation="portrait" xlink:href="ci4c00458_0002.jpg"><?image-name ci4c00458_0002.jpg?><?image-size 119280?><?image-md5 9055edc9fb54af6f810b02c24ae1e5cd?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 539?><?image-original-width 997?><?image-scaled-height 359?><?image-scaled-width 664?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/9055edc9fb54/ci4c00458_0002.jpg?><?thumb-name ci4c00458_0002.gif?><?thumb-size 17526?><?thumb-md5 9d9f7eda88529628de6899103eb34c04?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 79?><?thumb-scaled-width 147?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/9d9f7eda8852/ci4c00458_0002.gif?></graphic></p><p>It is now known that RNAs play more active roles in cellular
pathways
beyond simply serving as transcription templates. These biological
mechanisms might be mediated by higher RNA stereo conformations, triggering
the need to understand RNA secondary structures first. However, experimental
protocols for solving RNA structures are unavailable for large-scale
investigation due to their high costs and time-consuming nature. Various
computational tools were thus developed to predict the RNA secondary
structures from sequences. Recently, deep networks have been investigated
to help predict RNA structures directly from their sequences. However,
existing deep-learning-based tools are more or less suffering from
model overfitting due to their complicated problem formulation and
defective model training processes, limiting their applications across
sequences from different structural families. In this research, we
designed a two-stage RNA structure prediction strategy called DEBFold
(deep ensemble boosting and folding) based on convolution encoding/decoding
and self-attention mechanisms to enhance the existing thermodynamic
structure models. Moreover, the model training process followed rigorous
steps to achieve an acceptable prediction generalization. On the family-wise
reserved test sets and the PDB-derived test set, DEBFold achieves
better structure prediction performance over traditional tools and
existing deep-learning methods. In summary, we obtained a cutting-edge
deep-learning-based structure prediction tool with supreme across-family
generalization performance. The DEBFold tool can be accessed at <uri xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://cobis.bme.ncku.edu.tw/DEBFold/">https://cobis.bme.ncku.edu.tw/DEBFold/</uri>.</p></abstract><funding-group><award-group><funding-source><institution-wrap><institution>Ministry of Education</institution><institution-id institution-id-type="doi">10.13039/501100002701</institution-id></institution-wrap></funding-source><award-id>NA</award-id></award-group></funding-group><funding-group><award-group><funding-source><institution-wrap><institution>National Science and Technology Council</institution><institution-id institution-id-type="doi">10.13039/501100020950</institution-id></institution-wrap></funding-source><award-id>NSTC 112-2221-E-006-129-MY2</award-id></award-group></funding-group><funding-group><award-group><funding-source><institution-wrap><institution>National Science and Technology Council</institution><institution-id institution-id-type="doi">10.13039/501100020950</institution-id></institution-wrap></funding-source><award-id>MOST 111-2221-E-006-231</award-id></award-group></funding-group><funding-group><award-group><funding-source><institution-wrap><institution>National Science and Technology Council</institution><institution-id institution-id-type="doi">10.13039/501100020950</institution-id></institution-wrap></funding-source><award-id>MOST 110-2222-E-006-017</award-id></award-group></funding-group><funding-group><award-group><funding-source><institution-wrap><institution>National Cheng Kung University</institution><institution-id institution-id-type="doi">10.13039/501100007750</institution-id></institution-wrap></funding-source><award-id>NA</award-id></award-group></funding-group><custom-meta-group><custom-meta><meta-name>pmc-status-qastatus</meta-name><meta-value>0</meta-value></custom-meta><custom-meta><meta-name>pmc-status-live</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-status-embargo</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-status-released</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-open-access</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-olf</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-manuscript</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-legally-suppressed</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-has-pdf</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-has-supplement</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-pdf-only</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-suppress-copyright</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-is-real-version</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-is-scanned-article</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-preprint</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-in-epmc</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-license-ref</meta-name><meta-value>CC BY</meta-value></custom-meta><custom-meta><meta-name>document-id-old-9</meta-name><meta-value>ci4c00458</meta-value></custom-meta><custom-meta><meta-name>document-id-new-14</meta-name><meta-value>ci4c00458</meta-value></custom-meta><custom-meta><meta-name>ccc-price</meta-name><meta-value/></custom-meta></custom-meta-group></article-meta></front><body><sec id="sec1"><title>Introduction</title><p>RNA (ribonucleic acid) molecules in cells
can serve as not only
transcription templates but also noncoding RNAs.<sup><xref ref-type="bibr" rid="ref1">1</xref></sup> Many of them carry out specific cellular functions by using
their sequence patterns or folding structures.<sup><xref ref-type="bibr" rid="ref2">2</xref>,<xref ref-type="bibr" rid="ref3">3</xref></sup> For
example, internal ribosome entry site elements can regulate the translation
initiation of transcripts under stress conditions via their nonconservative
sequence patterns and partial secondary structures.<sup><xref ref-type="bibr" rid="ref4">4</xref>,<xref ref-type="bibr" rid="ref5">5</xref></sup> Since
RNA secondary structures can form the subdomains in tertiary structures<sup><xref ref-type="bibr" rid="ref6">6</xref></sup> and can carry out critical cellular functions
by themselves,<sup><xref ref-type="bibr" rid="ref7">7</xref></sup> figuring out the potential
secondary structures of newly found RNA sequences becomes routine
in RNA molecule studies.<sup><xref ref-type="bibr" rid="ref8">8</xref></sup></p><p>Experimental
approaches for solving RNA structures include X-ray
crystallography, nuclear magnetic resonance (NMR), and cryo-electron
microscopy.<sup><xref ref-type="bibr" rid="ref9">9</xref></sup> These experiments can recognize
the intramolecular RNA base–pair interactions within several
angstroms.<sup><xref ref-type="bibr" rid="ref10">10</xref></sup> However, a long period of
technician training time and high costs are required to carry out
these experimental approaches. Recently, several chemical probing
methods were developed, including the dimethyl sulfate (DMS) and the
selective 2′-hydroxyl analyzed by primer extension (SHAPE)
reagents treatment, to help identify the existence of structure pairings
on sequence locations by detecting reverse transcriptase-stopping
sites.<sup><xref ref-type="bibr" rid="ref1">1</xref></sup> When coupled with next-generation
sequencing techniques, genome-wide structural information can be obtained.<sup><xref ref-type="bibr" rid="ref11">11</xref></sup> Nevertheless, these chemical probing approaches
indicate only the existence of some structural forms. They also require
additional folding prediction to provide RNA structure–function
insights. Because of these reasons, in silico algorithms for identifying
RNA structures are essential in the large-scale study of RNA sequences.</p><p>Traditional knowledge-based RNA secondary structure prediction
algorithms utilize two approaches. The first type of algorithm assumes
that an RNA sequence folds into the structure with the minimum free
energy. The structures occupying the least free energy states are
found through dynamic programming<sup><xref ref-type="bibr" rid="ref12">12</xref></sup> or
Boltzmann partition functions.<sup><xref ref-type="bibr" rid="ref13">13</xref></sup> Researchers
also developed a second type of algorithm that further incorporates
multiple sequence alignment results when finding the minimum free
energy structures.<sup><xref ref-type="bibr" rid="ref14">14</xref></sup> This type of tool
sifts out the consensus structure either by simultaneous folding and
aligning or aligning before folding.<sup><xref ref-type="bibr" rid="ref15">15</xref></sup> Recently,
with the development of chemical structure probing reagents, many
tools can also consider probing read scores as soft constraints to
enhance structure prediction performance.<sup><xref ref-type="bibr" rid="ref16">16</xref></sup> Although a bundle of tools based on these two approaches has been
proposed, these tools bear drawbacks that deteriorate prediction accuracy.
First, free energy minimization algorithms may suffer from incomplete
parametrization of the thermal models, triggering lower sensitivity.<sup><xref ref-type="bibr" rid="ref7">7</xref>,<xref ref-type="bibr" rid="ref17">17</xref>,<xref ref-type="bibr" rid="ref18">18</xref></sup> Second, a well-chosen homologous
sequence family is indispensable for homology alignment-based tools.<sup><xref ref-type="bibr" rid="ref19">19</xref></sup> This requirement is usually not feasible for
large-scale RNA profiling. Because of the aforementioned issues, researchers
have tried to adopt data-oriented methods based on machine learning
and deep-learning techniques as another way to solve the RNA structure
prediction problem.</p><p>Early machine learning-based methods for
RNA secondary structure
prediction assisted the prediction process by refitting some predefined
parameters in the thermal models.<sup><xref ref-type="bibr" rid="ref20">20</xref></sup> Later,
some of these data-driven methods further tried to deal with the prediction
problem directly from the sequence level.<sup><xref ref-type="bibr" rid="ref21">21</xref>,<xref ref-type="bibr" rid="ref22">22</xref></sup> As a further step from feature-based algorithms, deep-learning tools
have become promising for solving RNA structures in an end-to-end
manner. However, existing deep-learning-based tools formulated the
structure representation using 2D matrix encoding. Since there are
far more unpaired position tuples, the 2D matrix structure encoding
can easily result in unbalanced positive and negative targets, leading
to the need for a large diversity in the training data sets. In a
recent study, Szikszai et al.<sup><xref ref-type="bibr" rid="ref23">23</xref></sup> found that
structure diversities within the data set are fundamental for training
structure prediction models. They pinpointed that existing tools were
trained only on data with sequence base variations, which suggested
the occurrence of model overfitting. Because of the high parametrization
and defect training processes, no robust result for diverse RNA structures
was ready in deep-learning-based tools.<sup><xref ref-type="bibr" rid="ref23">23</xref></sup></p><p>In this research, we devised a two-stage strategy called DEBFold
(deep ensemble boosting and folding) to overcome the prevalent overfitting
issue in deep-learning-based structure prediction tools. Instead of
assuming a sparse 2D structure matrix formulation, DEBFold is based
on a more compact 1D linear representation and was trained with a
family-wise prepared ground-truth data set. We arranged deep convolution
encoding/decoding and self-attention mechanisms to integrate and boost
existing thermal structure models in DEBFold for better RNA structure
prediction. Compared with traditional knowledge-based or deep-learning-based
structure prediction methods, DEBFold achieved better F1 score performance
on the reserved structure-independent test sets and the PDB-derived
test set, indicating good generalization across RNA structural families.
Moreover, we demonstrated that DEBFold is robust in its pipeline design.
These results suggest that DEBFold is the state-of-the-art deep-learning-based
tool that handles cross-family RNA structure prediction. Finally,
we implemented a web interface (<uri xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://cobis.bme.ncku.edu.tw/DEBFold/">https://cobis.bme.ncku.edu.tw/DEBFold/</uri>) that is freely available for researchers to facilitate the usage
of DEBFold.</p></sec><sec id="sec2"><title>Methods and Data Sets</title><sec id="sec2.1"><title>DEBFold Workflow</title><p>In this research, we designed a deep-learning
workflow that integrates knowledge of different thermodynamic structure
models with constrained optimization to tackle the RNA secondary structure
folding problem. The designed integration algorithm is called DEBFold
(deep ensemble boosting and folding). The DEBFold workflow comprises
two stages to convert the given RNA sequences into their potential
secondary structures (<xref rid="fig1" ref-type="fig">Figure <xref rid="fig1" ref-type="fig">1</xref></xref>): (1) structure location folding probability estimation.
DEBFold was designed to take an RNA sequence into a deep-learning
network that helps to integrate knowledge from various thermal models.
Then, it generates the folding constraints similar to the SHAPE experiment
results. (2) SHAPE-like constrained optimization. In the second stage,
free energy minimization constrained by the predicted structure location
folding probability will help provide the final structure prediction.
These two stages are elaborated on in the following subsections.</p><fig id="fig1" position="float" orientation="portrait"><label>Figure 1</label><caption><p>Overview
of DEBFold. The example RNA structure is visualized by <monospace>forna</monospace>.<sup><xref ref-type="bibr" rid="ref24">24</xref></sup> Notations: Conv—convolution
layers; FC—fully connected layers; TransConv—transpose
convolution layers; MFE: minimum free energy.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" id="gr1" position="float" orientation="portrait" xlink:href="ci4c00458_0001.jpg"><?image-name ci4c00458_0001.jpg?><?image-size 288562?><?image-md5 f040c486522f7d070190d68f680c85ef?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 2290?><?image-original-width 996?><?image-scaled-height 1527?><?image-scaled-width 664?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/f040c486522f/ci4c00458_0001.jpg?><?thumb-name ci4c00458_0001.gif?><?thumb-size 18785?><?thumb-md5 17fa8f330923d125c074c578f862f58a?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 230?><?thumb-scaled-width 100?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/17fa8f330923/ci4c00458_0001.gif?></graphic></fig><sec id="sec2.1.1"><title>Stage 1: Structure Location Folding Probability Estimation</title><p>In Stage 1, DEBFold tries to identify the potential locations that
are involved in the final structure pairings based on a structure
location identification deep network. DEBFold takes an RNA sequence
as the single input. Before the deep network, the pipeline prepares
the network input tensor based on the given sequence. The sequence
of the given RNA is one-hot encoded into an <italic toggle="yes">l</italic> x 4
tensor, where <italic toggle="yes">l</italic> is the length of the sequence. Then,
to integrate existing thermodynamics structure models, DEBFold collects
the prediction results suggested by 6 different single-sequence prediction
tools. The prediction results from the following tools were selected
to be part of the DEBFold-prepared input tensors: RNAfold,<sup><xref ref-type="bibr" rid="ref25">25</xref></sup> IPknot,<sup><xref ref-type="bibr" rid="ref6">6</xref></sup> MaxExpect,<sup><xref ref-type="bibr" rid="ref26">26</xref></sup> ProbKnot,<sup><xref ref-type="bibr" rid="ref27">27</xref></sup> RNAprob,<sup><xref ref-type="bibr" rid="ref28">28</xref></sup> and Fold.<sup><xref ref-type="bibr" rid="ref29">29</xref></sup> These
single-sequence input structure prediction tools were selected based
on their prediction accuracy, calculation efficiency, and ease of
incorporation. In our previous research,<sup><xref ref-type="bibr" rid="ref3">3</xref></sup> it was evaluated that RNAfold,<sup><xref ref-type="bibr" rid="ref25">25</xref></sup> IPknot,<sup><xref ref-type="bibr" rid="ref6">6</xref></sup> MaxExpect,<sup><xref ref-type="bibr" rid="ref26">26</xref></sup> and ProbKnot<sup><xref ref-type="bibr" rid="ref27">27</xref></sup> achieved acceptable structure prediction performance
within acceptable time requirements. We further included RNAprob<sup><xref ref-type="bibr" rid="ref28">28</xref></sup> and Fold<sup><xref ref-type="bibr" rid="ref29">29</xref></sup> to diversify
the thermal models while still limiting the time required to generate
these inputs. Although ShapeKnots and HotKnots can provide additional
thermal model diversity, these two tools have a much longer processing
time for long sequences and thus are not integrated. The prediction
results of the 6 tools under consideration are separately encoded
into individual tensor <italic toggle="yes">P</italic><sub><italic toggle="yes">i</italic></sub>, where <italic toggle="yes">P</italic><sub><italic toggle="yes">ij</italic></sub> = 1 if the <italic toggle="yes">j</italic>th location is said to be paired in the predicted structure
of tool <italic toggle="yes">i</italic>. At the end of this step, the sequence
one-hot encoding and the <italic toggle="yes">P</italic><sub><italic toggle="yes">i</italic></sub>’s are stacked into one <italic toggle="yes">l</italic> × 10
tensor. Since GPU parallel computation requires the lengths of the
training batch tensors to be the same, we zero-padded the length of
the network input tensor to 512 during the cross-validation training
process. For prediction, this model does not need this padding. Therefore,
variable-length RNAs can be fed into the structure location identification
network when used in prediction. However, because the ground-truth
data set collected in this research mainly consists of sequences shorter
than 512 bps, the most suitable value of <italic toggle="yes">l</italic> is suggested
to be less than 512 nt long.</p><p>The structure location identification
deep network can be divided into 3 parts: feature extraction, self-attention,
and structure location folding probability formation. The first part
of the deep network mixes the input information to extract valuable
features for the structure location identification. In previous research,
convolution operations have been successfully used to extract sequence
pattern features.<sup><xref ref-type="bibr" rid="ref30">30</xref>,<xref ref-type="bibr" rid="ref31">31</xref></sup> Therefore, consecutive convolution
operations (<xref rid="eq1" ref-type="disp-formula">eqs <xref rid="eq1" ref-type="disp-formula">1</xref></xref> and <xref rid="eq2" ref-type="disp-formula">2</xref>) were adopted in the deep network of DEBFold for
extracting sequence and structure features<disp-formula id="eq1"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m001.jpg"><?image-name ci4c00458_m001.jpg?><?image-size 1914?><?image-md5 c6dd4127e1eabf814ccdd4ca885e7634?><?image-image-server-status NEVER_LOAD?><?image-original-height 15?><?image-original-width 158?><?image-scaled-height 15?><?image-scaled-width 158?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/c6dd4127e1ea/ci4c00458_m001.jpg?><?thumb-name ci4c00458_m001.gif?><?thumb-size 767?><?thumb-md5 5827252a8ac859e232aecc958983cb1d?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 15?><?thumb-scaled-width 158?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/5827252a8ac8/ci4c00458_m001.gif?></graphic><label>1</label></disp-formula><disp-formula id="eq2"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m002.jpg"><?image-name ci4c00458_m002.jpg?><?image-size 2088?><?image-md5 ba2121a77ea5bda0791d91ba717d3f51?><?image-image-server-status NEVER_LOAD?><?image-original-height 15?><?image-original-width 171?><?image-scaled-height 15?><?image-scaled-width 171?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/ba2121a77ea5/ci4c00458_m002.jpg?><?thumb-name ci4c00458_m002.gif?><?thumb-size 811?><?thumb-md5 827685fd1e26ce15ed9d2ab33d59fe22?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 15?><?thumb-scaled-width 171?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/827685fd1e26/ci4c00458_m002.gif?></graphic><label>2</label></disp-formula>where <italic toggle="yes">I</italic> is the <italic toggle="yes">l</italic> × 10 input tensor and BN_ReLU refers to the operation of ReLU
activation, followed by batch normalization. In the above equation,
the convolution operation conv(<italic toggle="yes">n</italic>,<italic toggle="yes">T</italic>) with <italic toggle="yes">n</italic> 1D kernels of size 3 on the <italic toggle="yes">l</italic> × <italic toggle="yes">c</italic> tensor <italic toggle="yes">T</italic> (using stride
2) is defined as previously suggested<sup><xref ref-type="bibr" rid="ref32">32</xref>,<xref ref-type="bibr" rid="ref33">33</xref></sup> (<xref rid="eq3" ref-type="disp-formula">eq <xref rid="eq3" ref-type="disp-formula">3</xref></xref>)<disp-formula id="eq3"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m003.jpg"><?image-name ci4c00458_m003.jpg?><?image-size 3477?><?image-md5 a7bba6b403e45b3f4494a4496a35e45d?><?image-image-server-status NEVER_LOAD?><?image-original-height 41?><?image-original-width 264?><?image-scaled-height 41?><?image-scaled-width 264?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/a7bba6b403e4/ci4c00458_m003.jpg?><?thumb-name ci4c00458_m003.gif?><?thumb-size 2263?><?thumb-md5 0e4a21918609908640a5345b477308fd?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 31?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/0e4a21918609/ci4c00458_m003.gif?></graphic><label>3</label></disp-formula>where <italic toggle="yes">T</italic>(−1) is the
zero-padded element, <inline-formula id="d34e344"><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="ci4c00458_m004.gif"><?image-name ci4c00458_m004.gif?><?image-size 329?><?image-md5 c59e4d86ff263f160b0b5af9dae0e0c9?><?image-image-server-status NEVER_LOAD?><?image-scaled-height 20?><?image-scaled-width 58?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/c59e4d86ff26/ci4c00458_m004.gif?><?thumb-name ci4c00458_m004.gif?><?thumb-size 329?><?thumb-md5 c59e4d86ff263f160b0b5af9dae0e0c9?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 20?><?thumb-scaled-width 58?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/c59e4d86ff26/ci4c00458_m004.gif?></inline-graphic></inline-formula>, <italic toggle="yes">j</italic> is in the range of
0 and (<italic toggle="yes">n</italic> – 1), and <italic toggle="yes">K</italic><sub><italic toggle="yes">j</italic></sub> is the <italic toggle="yes">j</italic>th kernel.</p><p>After the features centered at each location are extracted (as <italic toggle="yes">X</italic><sub>2</sub> in <xref rid="eq2" ref-type="disp-formula">eq <xref rid="eq2" ref-type="disp-formula">2</xref></xref>), the mutual interactions between each location are considered
by the self-attention operation (<xref rid="eq4" ref-type="disp-formula">eqs <xref rid="eq4" ref-type="disp-formula">4</xref></xref> and <xref rid="eq6" ref-type="disp-formula">6</xref>). A self-attention block
helps compute the mutual influence between two nucleotide locations
within the given sequence. In order to avoid the gradient vanishing
problem in deep neural networks,<sup><xref ref-type="bibr" rid="ref34">34</xref></sup> a residual
architecture was also implemented in the self-attention block (<xref rid="eq5" ref-type="disp-formula">eqs <xref rid="eq5" ref-type="disp-formula">5</xref></xref> and <xref rid="eq7" ref-type="disp-formula">7</xref>)<disp-formula id="eq4"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m005.jpg"><?image-name ci4c00458_m005.jpg?><?image-size 2283?><?image-md5 f5294072fb8dfe576435da11e31b109f?><?image-image-server-status NEVER_LOAD?><?image-original-height 16?><?image-original-width 202?><?image-scaled-height 16?><?image-scaled-width 202?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/f5294072fb8d/ci4c00458_m005.jpg?><?thumb-name ci4c00458_m005.gif?><?thumb-size 2038?><?thumb-md5 43bb84865ed77fe7e2c7f38366b233eb?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 16?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/43bb84865ed7/ci4c00458_m005.gif?></graphic><label>4</label></disp-formula><disp-formula id="eq5"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m006.jpg"><?image-name ci4c00458_m006.jpg?><?image-size 1798?><?image-md5 d7c17f0cc83b8693f38ab8894ac6ebf9?><?image-image-server-status NEVER_LOAD?><?image-original-height 17?><?image-original-width 148?><?image-scaled-height 17?><?image-scaled-width 148?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/d7c17f0cc83b/ci4c00458_m006.jpg?><?thumb-name ci4c00458_m006.gif?><?thumb-size 738?><?thumb-md5 dfa63a209204e16c13b44845f3025a19?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 17?><?thumb-scaled-width 148?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/dfa63a209204/ci4c00458_m006.gif?></graphic><label>5</label></disp-formula><disp-formula id="eq6"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m007.jpg"><?image-name ci4c00458_m007.jpg?><?image-size 2283?><?image-md5 220aae03828d8872f1816ecc3f0e3678?><?image-image-server-status NEVER_LOAD?><?image-original-height 16?><?image-original-width 201?><?image-scaled-height 16?><?image-scaled-width 201?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/220aae03828d/ci4c00458_m007.jpg?><?thumb-name ci4c00458_m007.gif?><?thumb-size 2032?><?thumb-md5 d2758b8d3b4f28bf7885675fcb7179dd?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 16?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/d2758b8d3b4f/ci4c00458_m007.gif?></graphic><label>6</label></disp-formula><disp-formula id="eq7"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m008.jpg"><?image-name ci4c00458_m008.jpg?><?image-size 1764?><?image-md5 8818d6cf9000a967d857593c7640339b?><?image-image-server-status NEVER_LOAD?><?image-original-height 17?><?image-original-width 143?><?image-scaled-height 17?><?image-scaled-width 143?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/8818d6cf9000/ci4c00458_m008.jpg?><?thumb-name ci4c00458_m008.gif?><?thumb-size 732?><?thumb-md5 7857f5fc3036c52552db2dbe7b770c68?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 17?><?thumb-scaled-width 143?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/7857f5fc3036/ci4c00458_m008.gif?></graphic><label>7</label></disp-formula>where (<italic toggle="yes">W</italic><sub>1</sub>, <italic toggle="yes">W</italic><sub>2</sub>) and (<italic toggle="yes">W</italic><sub>3</sub>, <italic toggle="yes">W</italic><sub>4</sub>) are the trainable weights with internal positional
feed-forward dimension 1024, and multi_atten(.) is the multihead attention
operation. In DEBFold, the self-attention operation was repeated twice
in the network architecture. The multihead attention module, denoted
as multi_atten(<italic toggle="yes">Q</italic>, <italic toggle="yes">K</italic>, <italic toggle="yes">V</italic>, <italic toggle="yes">n</italic>), is defined as the following<sup><xref ref-type="bibr" rid="ref35">35</xref></sup><disp-formula id="eq8"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m009.jpg"><?image-name ci4c00458_m009.jpg?><?image-size 3332?><?image-md5 9e45d80c67fae80a133b3ca2c381135d?><?image-image-server-status NEVER_LOAD?><?image-original-height 42?><?image-original-width 205?><?image-scaled-height 42?><?image-scaled-width 205?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/9e45d80c67fa/ci4c00458_m009.jpg?><?thumb-name ci4c00458_m009.gif?><?thumb-size 2871?><?thumb-md5 e811f2f3a8683cd4e6e7818b16ea6dcc?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 41?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/e811f2f3a868/ci4c00458_m009.gif?></graphic><label>8</label></disp-formula><disp-formula id="eq9"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m010.jpg"><?image-name ci4c00458_m010.jpg?><?image-size 2448?><?image-md5 571bbed1f9ef7c0fafec814a293c6671?><?image-image-server-status NEVER_LOAD?><?image-original-height 15?><?image-original-width 210?><?image-scaled-height 15?><?image-scaled-width 210?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/571bbed1f9ef/ci4c00458_m010.jpg?><?thumb-name ci4c00458_m010.gif?><?thumb-size 2135?><?thumb-md5 05bbe682a4cb9b4bcbd18f5abf21a40f?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 14?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/05bbe682a4cb/ci4c00458_m010.gif?></graphic><label>9</label></disp-formula>where <italic toggle="yes">W</italic><sub>Q,<italic toggle="yes">i</italic></sub>/<italic toggle="yes">W</italic><sub>K,<italic toggle="yes">i</italic></sub>/<italic toggle="yes">W</italic><sub>V,<italic toggle="yes">i</italic></sub> are the query/key/value
weight matrices that map Q, K, and V into tensors of dimension 512, <italic toggle="yes">l</italic> is the length of the Q/K/V tensors, <italic toggle="yes">n</italic> is the number of heads used in this operation, ⊕ represents
the tensor concatenation operation, and <italic toggle="yes">H</italic> is the
computed multihead attention tensor. The self-attention operation
is obtained by setting the <italic toggle="yes">K</italic>, <italic toggle="yes">Q</italic>,
and <italic toggle="yes">V</italic> tensors to be identical. The number of heads
in this network was picked to be 16. After the self-attention block,
the sequence features for each location are weighted by the potential
influences among different locations.</p><p>In the last step of this
stage, the locations involved in the final
structure pairings are identified in the form of potential pairing
probabilities based on the final attended hidden feature tensor <italic toggle="yes">A</italic> (from <xref rid="eq7" ref-type="disp-formula">eq <xref rid="eq7" ref-type="disp-formula">7</xref></xref>). Since the final structure locations are marked segments along
the sequence position, DEBFold resorts to consecutive 1D transpose
convolutions (<xref rid="eq10" ref-type="disp-formula">eqs <xref rid="eq10" ref-type="disp-formula">10</xref></xref> and <xref rid="eq11" ref-type="disp-formula">11</xref>. See <xref rid="eq12" ref-type="disp-formula">eqs <xref rid="eq12" ref-type="disp-formula">12</xref></xref>–<xref rid="eq14" ref-type="disp-formula">14</xref> for its
definition) to help decode <italic toggle="yes">A</italic> (from <xref rid="eq7" ref-type="disp-formula">eq <xref rid="eq7" ref-type="disp-formula">7</xref></xref>) to produce segment marking along
the positions<disp-formula id="eq10"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m011.jpg"><?image-name ci4c00458_m011.jpg?><?image-size 2250?><?image-md5 b555777ba801141a436a2d2678eed4ea?><?image-image-server-status NEVER_LOAD?><?image-original-height 15?><?image-original-width 186?><?image-scaled-height 15?><?image-scaled-width 186?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/b555777ba801/ci4c00458_m011.jpg?><?thumb-name ci4c00458_m011.gif?><?thumb-size 882?><?thumb-md5 5f3af71773835cb5b5979283272c8a52?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 15?><?thumb-scaled-width 186?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/5f3af7177383/ci4c00458_m011.gif?></graphic><label>10</label></disp-formula><disp-formula id="eq11"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m012.jpg"><?image-name ci4c00458_m012.jpg?><?image-size 2269?><?image-md5 deba01c60ff25c4234cab2657f3570be?><?image-image-server-status NEVER_LOAD?><?image-original-height 15?><?image-original-width 186?><?image-scaled-height 15?><?image-scaled-width 186?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/deba01c60ff2/ci4c00458_m012.jpg?><?thumb-name ci4c00458_m012.gif?><?thumb-size 887?><?thumb-md5 1acdbfaa5576437a85d8a0d842d9099e?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 15?><?thumb-scaled-width 186?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/1acdbfaa5576/ci4c00458_m012.gif?></graphic><label>11</label></disp-formula></p><p>In the above equations, the 1D transpose
convolution transConv(<italic toggle="yes">n</italic>, <italic toggle="yes">F</italic>) using <italic toggle="yes">n</italic> kernels
of size 3 on the <italic toggle="yes">l</italic> × <italic toggle="yes">c</italic> tensor <italic toggle="yes">F</italic> (with stride 2) is calculated based on the following formula<disp-formula id="eq12"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m013.jpg"><?image-name ci4c00458_m013.jpg?><?image-size 2747?><?image-md5 57e97c60b938b091f7eaaf1f8721636c?><?image-image-server-status NEVER_LOAD?><?image-original-height 13?><?image-original-width 341?><?image-scaled-height 13?><?image-scaled-width 341?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/57e97c60b938/ci4c00458_m013.jpg?><?thumb-name ci4c00458_m013.gif?><?thumb-size 1487?><?thumb-md5 6e51eb2ac10a8f159473139aa1486156?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 8?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/6e51eb2ac10a/ci4c00458_m013.gif?></graphic><label>12</label></disp-formula><disp-formula id="eq13"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m014.jpg"><?image-name ci4c00458_m014.jpg?><?image-size 1598?><?image-md5 8967e88399fb8aa3a14191f6bb384a7d?><?image-image-server-status NEVER_LOAD?><?image-original-height 13?><?image-original-width 181?><?image-scaled-height 13?><?image-scaled-width 181?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/8967e88399fb/ci4c00458_m014.jpg?><?thumb-name ci4c00458_m014.gif?><?thumb-size 559?><?thumb-md5 f6ea11f9c22b2689cdcecd7bd945d170?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 13?><?thumb-scaled-width 181?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/f6ea11f9c22b/ci4c00458_m014.gif?></graphic><label>13</label></disp-formula><disp-formula id="eq14"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m015.jpg"><?image-name ci4c00458_m015.jpg?><?image-size 3669?><?image-md5 313cc7ce2d568335682c72e03314437f?><?image-image-server-status NEVER_LOAD?><?image-original-height 41?><?image-original-width 281?><?image-scaled-height 41?><?image-scaled-width 281?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/313cc7ce2d56/ci4c00458_m015.jpg?><?thumb-name ci4c00458_m015.gif?><?thumb-size 2204?><?thumb-md5 7c7f479d6902048c24209adcfb66a82e?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 29?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/7c7f479d6902/ci4c00458_m015.gif?></graphic><label>14</label></disp-formula>where <italic toggle="yes">E</italic> is the extension
tensor for <italic toggle="yes">F</italic>, 0 ≤ <italic toggle="yes">i</italic> ≤
2<italic toggle="yes">l</italic> – 1, <italic toggle="yes">j</italic> is in the range
of 0 and (<italic toggle="yes">n</italic> – 1), and <italic toggle="yes">P</italic><sub><italic toggle="yes">j</italic></sub> is the <italic toggle="yes">j</italic> kernel of the
transpose convolution operation. Finally, the structure location folding
probability of each location <italic toggle="yes">i</italic> in the output tensor <italic toggle="yes">Y</italic> is individually computed based on <italic toggle="yes">O</italic> (from <xref rid="eq11" ref-type="disp-formula">eq <xref rid="eq11" ref-type="disp-formula">11</xref></xref>) as<disp-formula id="uneq1"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m016.jpg"><?image-name ci4c00458_m016.jpg?><?image-size 932?><?image-md5 d007dd1c2a0d2f2395582d3317c0c111?><?image-image-server-status NEVER_LOAD?><?image-original-height 17?><?image-original-width 70?><?image-scaled-height 17?><?image-scaled-width 70?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/d007dd1c2a0d/ci4c00458_m016.jpg?><?thumb-name ci4c00458_m016.gif?><?thumb-size 3348?><?thumb-md5 a93e0b88c7bea5b7cd29503673a18022?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 49?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/a93e0b88c7be/ci4c00458_m016.gif?></graphic></disp-formula>where <italic toggle="yes">W</italic><sub>5</sub> is the
256 × 1 trainable weight matrix shared among all base locations
and σ is the sigmoid function. The choice of a weight matrix
of the same size as the internal hidden dimension fulfills the need
to accept RNAs with variant lengths, making DEBFold unlimited by RNA
sequence lengths.</p></sec><sec id="sec2.1.2"><title>Stage 2: Score-Constrained Optimization Folding</title><p>It
has been shown that structure probing values as soft folding constraints
can benefit the overall RNA secondary structure prediction.<sup><xref ref-type="bibr" rid="ref36">36</xref></sup> DEBFold utilizes this concept and incorporates
a two-stage pipeline to predict the RNA secondary structures. In Stage
1 of DEBFold, we obtain a structure location folding probability tensor,
indicating the possibility of each location being involved in the
final structure. The probabilities are inversely related to the read
scores measured in probing experiments such as SHAPE-seq. Hence, we
defined the SHAPE-like score tensor <italic toggle="yes">S</italic> by <italic toggle="yes">S</italic> = 1 – <italic toggle="yes">Y</italic>. The computed score tensor <italic toggle="yes">S</italic> is then provided as the SHAPE score file option for the
minimum free energy structure prediction optimization tools to be
used in soft constraint calculation. In DEBFold, we adopted the linear
equation proposed by Deigan et al.<sup><xref ref-type="bibr" rid="ref37">37</xref></sup> and
used the default slope and intercept values in each tool for this
equation to convert the SHAPE-like scores into pseudoenergy terms
in the optimization process. With the guidance of score tensor <italic toggle="yes">S</italic>, more accurate RNA structure predictions can be obtained
in the constrained optimization. In the final design of DEBFold, the
Fold algorithm<sup><xref ref-type="bibr" rid="ref29">29</xref></sup> is chosen as the default
constrained optimization algorithm to be coupled with the score tensor <italic toggle="yes">S</italic> for generating the final structure prediction.</p></sec><sec id="sec2.1.3"><title>Model Training Hyperparameters</title><p>DEBFold Stage 1 is formed
by a deep neural network, and we utilized cross-validation (CV) and
learning-curve techniques to select the proper hyperparameters. The
sum of the binary cross-entropy for each base location was selected
as the overall loss function to optimize the deep network. Since the
available structural families are not abundant, 25-CV was adopted
to fully utilize the data set. In order to avoid overfitting, the
average training and validation learning curves were monitored based
on the partial loss and partial F1 scores, which exclude the loss
and F1 calculation of the padded locations after the original sequence
length in each RNA. The final hyperparameters chosen in DEBFold are
summarized as follows: (1) training epochs: 250; (2) batch training
size: 512; (3) the optimization update algorithm: Adam; (4) learning
rate and decay schedule: cosine decay (minimum learning rate = 1 ×
10<sup>–13</sup>) in the first 227 epochs and then exponential
decay (multiplicative factor 1 × 10<sup>–10</sup>) for
the rest of the 23 epochs; (5) dropout layers with dropout rate =
0.1 and 0.4 were added to the residual layers and convolution/transpose
convolution layers, respectively, to regularize the deep network;
and (6) positive weights = 1.8 in the binary cross-entropy loss. The
training process was carried out using a NVIDIA RTX 4080 GPU.</p></sec></sec><sec id="sec2.2"><title>Family-Wise Processed RNA Structure Ground-Truth Data Set</title><p>Model training in data-based learning approaches requires the data
diversity in the training set to mimic the degree of variations in
real problems.<sup><xref ref-type="bibr" rid="ref38">38</xref></sup> Researchers have experimented
that a large number of sequences from a small number of structural
families can easily lead to model overfitting.<sup><xref ref-type="bibr" rid="ref23">23</xref></sup> The possible reason for the situation is that structurally
similar sequences provide limited variations for deep models to fetch
the inherent patterns that map sequences to the corresponding structures.
In order to prepare a ground-truth data set that incorporates a high
structure diversity, we downloaded the data deposited in bpRNA-1m<sup><xref ref-type="bibr" rid="ref39">39</xref></sup> and retrieved the sequences with known Rfam
families. In bpRNA-1m, RNAs curated from Rfam 12.2<sup><xref ref-type="bibr" rid="ref40">40</xref></sup> were selected and processed. In total, 43,269 valid RNA
sequences from 2125 families were found.</p><p>In order to tune and
evaluate DEBFold in a structural family independent manner, the ground-truth
data set is divided and sampled into two parts: the training-validation
and test sets. Since RNAs in the same family are structurally similar,<sup><xref ref-type="bibr" rid="ref23">23</xref></sup> the data processing was carried out in a family-wise
manner. We resorted to 25-fold cross-validation in DEBFold Stage 1
model training due to the scarcity of known distinct RNA structural
families in the collected ground-truth data set. However, the average
of the cross-validation results provides only an optimistic estimation
of the model performance. An extra test set should be set aside as
totally clean data for a fair performance evaluation. Therefore, among
the gathered 2125 families, 44 (2%, or 1/50 due to 25-fold cross-validation,
of the total 2125 families) of them were reserved in the test set.
These 44 families were selected by stratified sampling that preserves
the length distribution of all 2125 RNA families. In the calculation
of the length distribution of RNA families, the longest sequence within
a structural family was chosen as the representative. These 44 families
were also ensured to be in distinct Rfam clans from the rest of the
remaining families. The remaining 2081 families were used as the training-validation
folds in cross-validation. We sampled at most 5 sequences from the
2081 families to be included in the training-validation set (8782
sequences in total). Finally, we took one sequence from each of the
44 test families for testing, resulting in a test set consisting of
44 RNAs. This test set is termed TestSetα in this research.
We further confirmed the sequence similarity between the training-validation
and test RNA sequences using CD-HIT.<sup><xref ref-type="bibr" rid="ref41">41</xref></sup> Under
the threshold of 0.85, TestSetα contains no RNA with sequence
similarity to any member of the training-validation set. In order
to have more extensive testing, we further prepared two additional
independent test sets (TestSetβ and TestSetγ) as described
in the following two subsections.</p><sec id="sec2.2.1"><title>Contamination-Free Family-Wise Independent Test Set for Evaluating
Existing Deep-Learning-Based Structure Prediction Tools</title><p>In
this research, we also tried to compare the performance of DEBFold
with existing deep-learning-based structure prediction tools. While
SPOT-RNA and SPOT-RNA2 were originally trained on the bpRNA-1m data
set and there is no retraining code for these tools, it is necessary
to gather some other independent test set than TestSetα, which
was derived from bpRNA-1m, for ensuring a fair generalization evaluation.<sup><xref ref-type="bibr" rid="ref38">38</xref></sup> For this purpose, we downloaded the bpRNA-new
data set from the MXfold2 work<sup><xref ref-type="bibr" rid="ref22">22</xref></sup> and prepared
an independent test set for fairly evaluating the performance of DEBFold
with existing deep-learning-based prediction tools. The so-called
bpRNA-new data set from the MXfold2 work includes RNA sequences from
Rfam 14.4 but excluded those already in Rfam 12.2. We further reorganized
the bpRNA-new data set to better help evaluate the deep-learning model
performance. In the downloaded bpRNA-new data set, there are a total
of 489 families for the collected sequences. Among them, 8 families
also appear in the bpRNA-1m data set. Hence, only 481 families with
5265 sequences were suitable and thus retained. These 481 families
were also ensured to be in different Rfam clans from sequences in
bpRNA-1m. Since most deep-learning-based prediction tools can only
handle sequences composed of the four basic nucleotide alphabets,
sequences with alphabets other than AU(T)GC were abandoned. For each
family, one sequence from each family was selected as the representative.
As previously cautioned, benchmark sets with too many short sequences
tend to provide biased performance evaluation results.<sup><xref ref-type="bibr" rid="ref42">42</xref></sup> In order to eliminate the length biases for
the remaining bpRNA-new sequences in a family-wise manner, the remaining
481 families are binned with a length of 50 nts, resulting in 10 bins
in total (max-length = 489 nts). In these 10 bins, the minimum number
of families within a bin is 4. Therefore, to balance the number of
sequences in each length bin, we sampled 4 families from each of these
10 bins to prepare the final test set, resulting in 40 test families
(i.e., 40 sequences). The test set consisting of these 40 sequences
from different families is called TestSetβ in this research.
We also confirmed the sequence similarity between the training-validation
and test RNA sequences using CD-HIT.<sup><xref ref-type="bibr" rid="ref41">41</xref></sup> Under
the threshold of 0.85, TestSetβ contains no RNA sequence similar
to any member of the training-validation set.</p></sec><sec id="sec2.2.2"><title>PDB-Derived Source-Independent Test Set</title><p>Since RNA structures
deposited in Rfam are actually only a mixture of experimental results,
predictions, and homologue modeling inference,<sup><xref ref-type="bibr" rid="ref43">43</xref></sup> we sought another test set that gathers RNA structures
identified with pure experimental evidence. In the bpRNA database,<sup><xref ref-type="bibr" rid="ref39">39</xref></sup> Danaee et al. also collected some PDB-derived
RNA secondary structures that were parsed only from the 3D structures
obtained using X-ray crystallography and NMR techniques. We downloaded
these PDB-derived results and prepared a source-independent test set
named TestSetγ. As in TestSetβ, only sequences represented
using AU(T)GC were retained. In addition, to have a fair test set
that does not include sequences already used in the training process
of any tool, we excluded sequences that have more than 85% CD-HIT-calculated
similarity to any sequence from the training-validation set used in
this research and the TR1 data set used in the training process of
SPOT-RNA. In total, 147 PDB-derived RNA structures were available.
We also sought to avoid length biases in TestSetγ by equally
length-partitioning these sequences into 10 bins (max-length = 338
nts). We take only 3 sequences within each length bin in order to
balance the number of sequences in each bin, resulting in 15 independent
sequences (several bins have zero sequence in them) in TestSetγ.
Overall, TestSetγ serves as a source-independent test set that
helps evaluate the tools in predicting structures directly inferred
from experiments.</p></sec></sec><sec id="sec2.3"><title>Structure Prediction Evaluation Metrics</title><p>To evaluate
the RNA secondary structure prediction performance, we resorted to
the structure F1 score calculation. The precision, recall, and F1
score for a predicted structure are computed as follows<sup><xref ref-type="bibr" rid="ref3">3</xref>,<xref ref-type="bibr" rid="ref8">8</xref></sup><disp-formula id="eq15"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m017.jpg"><?image-name ci4c00458_m017.jpg?><?image-size 3150?><?image-md5 ba6b96949cd8a3daa474186b5e384b8c?><?image-image-server-status NEVER_LOAD?><?image-original-height 29?><?image-original-width 261?><?image-scaled-height 29?><?image-scaled-width 261?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/ba6b96949cd8/ci4c00458_m017.jpg?><?thumb-name ci4c00458_m017.gif?><?thumb-size 1922?><?thumb-md5 d219d7af9f15b6996e01201fff809001?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 22?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/d219d7af9f15/ci4c00458_m017.gif?></graphic><label>15</label></disp-formula><disp-formula id="eq16"><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="anchor" orientation="portrait" xlink:href="ci4c00458_m018.jpg"><?image-name ci4c00458_m018.jpg?><?image-size 3120?><?image-md5 09176ddaa0eda168dcd4787fa669c1c1?><?image-image-server-status NEVER_LOAD?><?image-original-height 33?><?image-original-width 177?><?image-scaled-height 33?><?image-scaled-width 177?><?image-cloudpmc-urn urn:cdn:blobs/08a5/11094721/09176ddaa0ed/ci4c00458_m018.jpg?><?thumb-name ci4c00458_m018.gif?><?thumb-size 1142?><?thumb-md5 049d3ff860184fa48b504b9d3edc81de?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 33?><?thumb-scaled-width 177?><?thumb-cloudpmc-urn urn:cdn:blobs/08a5/11094721/049d3ff86018/ci4c00458_m018.gif?></graphic><label>16</label></disp-formula>where TP, FP, and FN stand for the numbers
of true-positive, false-positive, and false–negative pairs,
respectively. For a predicted structure and its corresponding real
structure, TP counts the number of correctly predicted pairings, FP
is the number of predicted pairings that do not appear in the real
structure, and FN considers the number of real base pairings missed
by the prediction. The final F1 score can balance the trade-off between
the FP and FN values and is thus used as the overall evaluation metric.</p></sec></sec><sec id="sec3"><title>Results and Discussion</title><sec id="sec3.1"><title>DEBFold Outperforms Previous Thermodynamics-Based RNA Structure
Prediction Tools</title><p>We first evaluated the generalization prediction
performance of DEBFold and compared its performance with classical
thermodynamics-based prediction tools on prepared TestSetα,
TestSetβ, and TestSetγ. Nineteen prediction tools that
were still publicly available were compared in this section: (1) tools
from single-sequence thermodynamics models: RNAfold,<sup><xref ref-type="bibr" rid="ref25">25</xref></sup> RNALfold,<sup><xref ref-type="bibr" rid="ref44">44</xref></sup> Fold,<sup><xref ref-type="bibr" rid="ref29">29</xref></sup> MaxExpect,<sup><xref ref-type="bibr" rid="ref26">26</xref></sup> RNAprob,<sup><xref ref-type="bibr" rid="ref28">28</xref></sup> RME (both the PARS-model and the DMS-model),<sup><xref ref-type="bibr" rid="ref45">45</xref></sup> PKNOTS,<sup><xref ref-type="bibr" rid="ref46">46</xref></sup> IPknot,<sup><xref ref-type="bibr" rid="ref6">6</xref></sup> HotKnots,<sup><xref ref-type="bibr" rid="ref47">47</xref></sup> IterativeHFold,<sup><xref ref-type="bibr" rid="ref48">48</xref></sup> ShapeKnots,<sup><xref ref-type="bibr" rid="ref49">49</xref></sup> and
ProbKnot;<sup><xref ref-type="bibr" rid="ref27">27</xref></sup> (2) tools based on homologous
sequence alignment and folding: RNAalifold,<sup><xref ref-type="bibr" rid="ref14">14</xref></sup> TurboFold,<sup><xref ref-type="bibr" rid="ref29">29</xref></sup> SPARSE,<sup><xref ref-type="bibr" rid="ref50">50</xref></sup> aliFreeFold,<sup><xref ref-type="bibr" rid="ref15">15</xref></sup> LocARNA,<sup><xref ref-type="bibr" rid="ref51">51</xref></sup> comRNA,<sup><xref ref-type="bibr" rid="ref52">52</xref></sup> and MXSCARNA.<sup><xref ref-type="bibr" rid="ref53">53</xref></sup> Algorithms that fail to provide executable codes
were not included in this comparison. Default parameters suggested
by the authors were adopted in these tools for a fair comparison.
For IterativeHFold, we set RNAfold as the restriction structure generator
for it. For tools that require multiple sequence alignment results,
sequences in the same Rfam family were taken. In some Rfam families,
the sequences are nearly identical, even though they were collected
from different species. These nearly identical sequences within the
same family were used in tools that required homologous alignments
if no other homologous sequence was available for that family. If
no sequence is available in the same RNA structural family of the
given sequence, the required homologues were searched against the
RNAcentral database<sup><xref ref-type="bibr" rid="ref54">54</xref></sup> using BLAST as suggested
in our previous work.<sup><xref ref-type="bibr" rid="ref3">3</xref></sup></p><p>The results
of different tools on TestSetα, TestSetβ, and TestSetγ
are summarized in <xref rid="tbl1" ref-type="other">Table <xref rid="tbl1" ref-type="other">1</xref></xref>. In <xref rid="tbl1" ref-type="other">Table <xref rid="tbl1" ref-type="other">1</xref></xref>, the
performance of each tool was summarized by its median rank and overall
median F1 score on the three test sets. Among single-sequence input
tools [RNAfold, RNALfold, Fold, MaxExpect, RNAprob, RME (both the
PARS-model and the DMS-model), PKNOTS, IPknot, HotKnots, IterativeHFold,
ShapeKnots, and ProbKnot], DEBFold achieves the overall top-one median
F1 score rank on the three test sets. The improvement is supposed
to be from the integration of various existing thermal knowledge followed
by deep learning boosting. Compared with tools that utilize homologous
sequences (RNAalifold, TurboFold, SPARSE, aliFreeFold, LocARNA, comRNA,
and MXSCARNA), DEBFold still has a better overall median F1 score
rank. Combing these comparisons, DEBFold provides cutting-edge performance
over existing tools on the test sets. Notice that in TestSetα,
tools that rely on sequence alignment results generally perform better
than single-sequence input tools since the RNA sequences from Rfam
12.2 mostly come with good family-wise homologous sequences. This
advantage may diminish in many real-world use cases, such as in TestSetγ.
DEBFold overcomes this restriction by using a deep learning model
that integrates the results of many single-sequence structure prediction
tools. It is worthwhile noting that DEBFold also boosts the prediction
F1 score over the 6 predictions (RNAfold, IPknot, MaxExpect, ProbKnot,
RNAprob, and Fold) integrated in the input encoding tensor of the
deep network. In summary, DEBFold is concluded to outperform existing
classical structure prediction tools with less input requirements.</p><table-wrap id="tbl1" position="float" orientation="portrait"><label>Table 1</label><caption><title>Test Set Median F1 Score Performance
Comparison between DEBFold and Other Available Thermodynamics-Based
RNA Secondary Structure Prediction Tools on the Three Prepared Test
Sets<xref rid="t1fn1" ref-type="table-fn">a</xref></title></caption><table frame="hsides" rules="groups" border="0"><colgroup span="1"><col align="left" span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="left" span="1"/><col align="left" span="1"/><col align="left" span="1"/><col align="left" span="1"/><col align="left" span="1"/><col align="left" span="1"/><col align="left" span="1"/><col align="left" span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/></colgroup><thead><tr><th style="border:none;" align="center" colspan="1" rowspan="1">structure
prediction tool</th><th colspan="4" align="center" char="." rowspan="1">TestSetα<hr/></th><th colspan="4" align="center" rowspan="1">TestSetβ<hr/></th><th colspan="4" align="center" rowspan="1">TestSetγ<hr/></th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">overall median F1 (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">median rank</th></tr><tr><th style="border:none;" align="center" colspan="1" rowspan="1"> </th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">R (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">rank</th><th style="border:none;" align="center" colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" colspan="1" rowspan="1">R (%)</th><th style="border:none;" align="center" colspan="1" rowspan="1">rank</th><th style="border:none;" align="center" colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" colspan="1" rowspan="1">R (%)</th><th style="border:none;" align="center" colspan="1" rowspan="1">rank</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1"> </th><th style="border:none;" align="center" char="." colspan="1" rowspan="1"> </th></tr></thead><tbody><tr><td style="border:none;" align="left" colspan="1" rowspan="1">DEBFold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">62.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">67.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">4</td><td style="border:none;" align="left" colspan="1" rowspan="1">55.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">56.4</td><td style="border:none;" align="left" colspan="1" rowspan="1">56.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">1</td><td style="border:none;" align="left" colspan="1" rowspan="1">77.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">83.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">73.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">1</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">RNAalifold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">62.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">67.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">5</td><td style="border:none;" align="left" colspan="1" rowspan="1">55.6</td><td style="border:none;" align="left" colspan="1" rowspan="1">50.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">62.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">2</td><td style="border:none;" align="left" colspan="1" rowspan="1">59.6</td><td style="border:none;" align="left" colspan="1" rowspan="1">56.4</td><td style="border:none;" align="left" colspan="1" rowspan="1">63.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">14</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">59.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">5</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">ProbKnot</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">60.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">60.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">60.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">13</td><td style="border:none;" align="left" colspan="1" rowspan="1">53.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">50.4</td><td style="border:none;" align="left" colspan="1" rowspan="1">58.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">6</td><td style="border:none;" align="left" colspan="1" rowspan="1">72.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">77.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">68.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">60.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">6</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">MaxExpect</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">54.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">82.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">7</td><td style="border:none;" align="left" colspan="1" rowspan="1">55.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">43.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">81.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">4</td><td style="border:none;" align="left" colspan="1" rowspan="1">69.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">73.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">65.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">7</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">ShapeKnots</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">61.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">54.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">71.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">12</td><td style="border:none;" align="left" colspan="1" rowspan="1">55.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">45.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">69.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">3</td><td style="border:none;" align="left" colspan="1" rowspan="1">70.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">73.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">68.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">61.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">7</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">IterativeHFold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">53.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">82.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">8</td><td style="border:none;" align="left" colspan="1" rowspan="1">52.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">39.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">81.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">8</td><td style="border:none;" align="left" colspan="1" rowspan="1">77.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">83.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">73.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">8</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">RNAfold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">53.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">82.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">8</td><td style="border:none;" align="left" colspan="1" rowspan="1">52.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">39.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">81.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">8</td><td style="border:none;" align="left" colspan="1" rowspan="1">77.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">83.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">73.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">8</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">RME (PARS model)</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">56.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">75.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">10</td><td style="border:none;" align="left" colspan="1" rowspan="1">54.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">43.1</td><td style="border:none;" align="left" colspan="1" rowspan="1">81.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">5</td><td style="border:none;" align="left" colspan="1" rowspan="1">69.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">73.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">65.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">8</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">RME (DMS model)</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">53.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">82.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">11</td><td style="border:none;" align="left" colspan="1" rowspan="1">53.1</td><td style="border:none;" align="left" colspan="1" rowspan="1">38.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">87.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">7</td><td style="border:none;" align="left" colspan="1" rowspan="1">69.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">73.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">65.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">8</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">SPARSE</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">65.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">66.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">2</td><td style="border:none;" align="left" colspan="1" rowspan="1">51.4</td><td style="border:none;" align="left" colspan="1" rowspan="1">38.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">80.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">10</td><td style="border:none;" align="left" colspan="1" rowspan="1">60.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">72.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">52.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">12</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">60.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">10</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">TurboFold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">54.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">83.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">6</td><td style="border:none;" align="left" colspan="1" rowspan="1">47.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">57.1</td><td style="border:none;" align="left" colspan="1" rowspan="1">45.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">13</td><td style="border:none;" align="left" colspan="1" rowspan="1">63.6</td><td style="border:none;" align="left" colspan="1" rowspan="1">85.4</td><td style="border:none;" align="left" colspan="1" rowspan="1">50.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">11</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">11</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">LocARNA</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">65.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">66.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">2</td><td style="border:none;" align="left" colspan="1" rowspan="1">49.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">47.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">54.1</td><td style="border:none;" align="left" colspan="1" rowspan="1">11</td><td style="border:none;" align="left" colspan="1" rowspan="1">48.1</td><td style="border:none;" align="left" colspan="1" rowspan="1">53.4</td><td style="border:none;" align="left" colspan="1" rowspan="1">43.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">19</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">49.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">11</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">HotKnots</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">59.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">54.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">75.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">14</td><td style="border:none;" align="left" colspan="1" rowspan="1">48.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">40.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">62.1</td><td style="border:none;" align="left" colspan="1" rowspan="1">12</td><td style="border:none;" align="left" colspan="1" rowspan="1">76.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">82.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">70.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">59.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">12</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">IPknot</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">57.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">47.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">81.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">15</td><td style="border:none;" align="left" colspan="1" rowspan="1">47.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">44.1</td><td style="border:none;" align="left" colspan="1" rowspan="1">52.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">14</td><td style="border:none;" align="left" colspan="1" rowspan="1">76.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">83.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">66.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">4</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">57.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">14</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">aliFreeFold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">65.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">59.4</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">67.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">1</td><td style="border:none;" align="left" colspan="1" rowspan="1">44.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">41.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">49.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">16</td><td style="border:none;" align="left" colspan="1" rowspan="1">48.1</td><td style="border:none;" align="left" colspan="1" rowspan="1">53.4</td><td style="border:none;" align="left" colspan="1" rowspan="1">43.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">19</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">48.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">16</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">MXSCARNA</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">47.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">37.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">16</td><td style="border:none;" align="left" colspan="1" rowspan="1">34.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">29.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">43.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">19</td><td style="border:none;" align="left" colspan="1" rowspan="1">54.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">100.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">37.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">15</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">47.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">16</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">Fold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">45.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">36.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">73.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">17</td><td style="border:none;" align="left" colspan="1" rowspan="1">43.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">43.6</td><td style="border:none;" align="left" colspan="1" rowspan="1">42.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">17</td><td style="border:none;" align="left" colspan="1" rowspan="1">54.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">72.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">42.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">15</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">45.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">17</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">RNAprob</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">45.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">36.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">73.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">17</td><td style="border:none;" align="left" colspan="1" rowspan="1">43.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">43.6</td><td style="border:none;" align="left" colspan="1" rowspan="1">42.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">17</td><td style="border:none;" align="left" colspan="1" rowspan="1">54.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">72.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">42.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">15</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">45.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">17</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">PKNOTs</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">42.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">37.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">51.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">20</td><td style="border:none;" align="left" colspan="1" rowspan="1">47.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">42.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">55.6</td><td style="border:none;" align="left" colspan="1" rowspan="1">15</td><td style="border:none;" align="left" colspan="1" rowspan="1">54.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">65.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">46.4</td><td style="border:none;" align="left" colspan="1" rowspan="1">18</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">47.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">18</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">RNALfold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">45.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">51.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">66.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">19</td><td style="border:none;" align="left" colspan="1" rowspan="1">25.4</td><td style="border:none;" align="left" colspan="1" rowspan="1">31.8</td><td style="border:none;" align="left" colspan="1" rowspan="1">21.1</td><td style="border:none;" align="left" colspan="1" rowspan="1">20</td><td style="border:none;" align="left" colspan="1" rowspan="1">60.5</td><td style="border:none;" align="left" colspan="1" rowspan="1">81.2</td><td style="border:none;" align="left" colspan="1" rowspan="1">48.1</td><td style="border:none;" align="left" colspan="1" rowspan="1">13</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">45.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">19</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">comRNA</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">29.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">24.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">37.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">21</td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">29.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">21</td></tr></tbody></table><table-wrap-foot><fn id="t1fn1"><label>a</label><p>The median rank column summarizes
the rank median of each tool in the three test sets, and the overall
median F1 column computes the median F1 score of the three test-set
median F1 scores. For comRNA, no reasonable result was provided in
TestSetβ and TestSet γ. Notations: P—precision,
R—recall.</p></fn></table-wrap-foot></table-wrap></sec><sec id="sec3.2"><title>DEBFold Has Better Generalization Performance than Existing
Deep-Learning-Based Attempts</title><p>Recently, more and more deep-learning-based
structure prediction tools were developed to help predict RNA secondary
structures. We also sought to compare DEBFold to these tools. To the
best of our knowledge, existing deep-learning-based RNA secondary
structure prediction tools include the following: SPOT-RNA,<sup><xref ref-type="bibr" rid="ref21">21</xref></sup> SPOT-RNA2,<sup><xref ref-type="bibr" rid="ref55">55</xref></sup> MXfold2,<sup><xref ref-type="bibr" rid="ref22">22</xref></sup> GCNFold,<sup><xref ref-type="bibr" rid="ref56">56</xref></sup> UFold,<sup><xref ref-type="bibr" rid="ref57">57</xref></sup> REDFold,<sup><xref ref-type="bibr" rid="ref58">58</xref></sup> and e2eFold.<sup><xref ref-type="bibr" rid="ref59">59</xref></sup> Since SPOT-RNA and SPOT-RNA2 released by the
authors do not encompass any retraining codes and the tools themselves
incorporate the whole bpRNA-1m data set in its model training process,
independent test sets other than TestSetα, which was prepared
from the bpRNA-1m Rfam 12.2 part, should be used to ensure fair comparison
between DEBFold and SPOT-RNA/SPOT-RNA2. For this purpose, we downloaded
the bpRNA-new data set from the MXfold2 work and prepared TestSetβ
based on Rfam 14.4 (excluding those already in Rfam 12.2) for fairly
evaluating the performance of DEBFold with existing deep-learning-based
prediction tools. In TestSetβ, both the length bias and data
contamination issues were eliminated. Details of the TestSetβ
preparation steps can be found in the “<xref rid="sec2.2.1" ref-type="other">Contamination-Free Family-Wise Independent Test Set for Evaluating
Existing Deep-Learning-Based Structure Prediction Tools</xref>”
section. Another PDB-derived independent test set, TestSetγ,
was also prepared for this purpose. The preparation steps of TestSetγ
can be found in the “<xref rid="sec2.2.2" ref-type="other">PDB-Derived Source-Independent
Test Set</xref>” section.</p><p>The final test results of DEBFold
and the existing deep-learning-based prediction tools on TestSetα,
TestSetβ, and TestSetγ are listed in <xref rid="tbl2" ref-type="other">Table <xref rid="tbl2" ref-type="other">2</xref></xref>. For SPOT-RNA, SPOT-RNA2, and
UFold, TestSetα results were not reported due to complete data
contamination caused by using the whole bpRNA-1m data set in the training
process. REDFold has potential slightly optimistic evaluation results
on TestSetα and TestSetβ since a small portion (122 families,
whose details were not clearly reported by the original paper) of
the Rfam 14.4 families were included in the training process. MXfold2
also included a small portion of the Rfam data (22 Rfam families from
Rfam 10.0 and 151 sequences from the S-151Rfam data set, whose details
were not clearly reported by the original paper), which might result
in potential slightly optimistic results for TestSetα and TestSetβ.
We calculated the median performance ranking and median F1 score of
each tool on suitable test sets in the comparison. For tools that
are suitable to be evaluated on all three test sets (DEBFold, MXfold2,
REDFold, GCNfold, and e2efold), DEBFold has superior median F1 score
results. Compared with SPOT-RNA, UFold, and SPOT-RNA2, DEBFold shows
the top overall median F1 score rank. Among these tools, SPOT-RNA
was first pretrained on the bpRNA-1m data set. Then, these pretrained
models were fine-tuned and aggregated on the PDB data set. In other
words, both the RNA structural families from bpRNA-1m and the PDB-identified
RNA sequences were considered in the tool. Because of the final fine-tuning
process, the final published SPOT-RNA tool somewhat favors 3D-structure-derived
RNA base pairings. However, in PDB-derived TestSetγ, DEBFold
still outperforms SPOT-RNA. Therefore, we conclude that DEBFold can
provide better results than SPOT-RNA even on the 3D-structure-derived
base pairs. It is worth noting that SPOT-RNA2 can achieve good results
by considering the homologous sequence information. Nonetheless, SPOT-RNA2
only gets the same median performance rank as DEBFold while demanding
more than 7500 times longer execution time than DEBFold (see <xref rid="tbl5" ref-type="other">Table <xref rid="tbl5" ref-type="other">5</xref></xref>). In view of the
overall consideration of prediction efficiency and accuracy, DEBFold
still has better results than SPOT-RNA2. Overall, the performance
improvements in DEBFold are supposed to be attributed to a novel two-stage
pipeline that utilizes a 1D structure representation and deep network
integration of various thermal models. These analyses indicate that
DEBFold can achieve state-of-the-art generalization structure prediction
performance over that of currently available deep-learning-based prediction
tools.</p><table-wrap id="tbl2" position="float" orientation="portrait"><label>Table 2</label><caption><title>Test Set Median F1 Score Performance
Comparison between DEBFold and Other Deep-Learning-Based RNA Structure
Prediction Tools on the Three Prepared Test Sets<xref rid="t2fn1" ref-type="table-fn">a</xref></title></caption><table frame="hsides" rules="groups" border="0"><colgroup span="1"><col align="left" span="1"/><col align="left" span="1"/><col align="left" span="1"/><col align="left" span="1"/><col align="left" span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/></colgroup><thead><tr><th style="border:none;" align="center" colspan="1" rowspan="1">structure prediction tool</th><th colspan="4" align="center" rowspan="1">TestSetα<hr/></th><th colspan="4" align="center" char="." rowspan="1">TestSetβ<hr/></th><th colspan="4" align="center" char="." rowspan="1">TestSetγ<hr/></th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">overall median F1 (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">median rank</th></tr><tr><th style="border:none;" align="center" colspan="1" rowspan="1"> </th><th style="border:none;" align="center" colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" colspan="1" rowspan="1">R (%)</th><th style="border:none;" align="center" colspan="1" rowspan="1">rank</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">R (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">rank</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">R (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">rank</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1"> </th><th style="border:none;" align="center" char="." colspan="1" rowspan="1"> </th></tr></thead><tbody><tr><td style="border:none;" align="left" colspan="1" rowspan="1"><bold>DEBFold</bold></td><td style="border:none;" align="left" colspan="1" rowspan="1">64.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">62.1</td><td style="border:none;" align="left" colspan="1" rowspan="1">67.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">55.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">56.4</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">56.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">77.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">83.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">73.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">1</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">MXfold2</td><td style="border:none;" align="left" colspan="1" rowspan="1">58.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">62.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">55.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">54.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">43.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">73.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">78.4</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">80.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">76.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">58.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">2</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">REDFold</td><td style="border:none;" align="left" colspan="1" rowspan="1">40.4</td><td style="border:none;" align="left" colspan="1" rowspan="1">46.7</td><td style="border:none;" align="left" colspan="1" rowspan="1">37.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">41.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">50.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">35.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">53.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">56.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">51.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">41.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">3</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">GCNfold</td><td style="border:none;" align="left" colspan="1" rowspan="1">31.9</td><td style="border:none;" align="left" colspan="1" rowspan="1">83.3</td><td style="border:none;" align="left" colspan="1" rowspan="1">20.0</td><td style="border:none;" align="left" colspan="1" rowspan="1">4</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">27.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">61.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">18.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">4</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">40.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">81.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">25.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">4</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">31.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">4</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">e2eFold</td><td style="border:none;" align="left" colspan="1" rowspan="1">0</td><td style="border:none;" align="left" colspan="1" rowspan="1">0</td><td style="border:none;" align="left" colspan="1" rowspan="1">0</td><td style="border:none;" align="left" colspan="1" rowspan="1">5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">1.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">2.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">0.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">9.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">26.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">5.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">1.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">5</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1"> </td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1"><bold>DEBFold</bold></td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">55.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">56.4</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">56.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">77.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">83.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">73.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">66.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">1.5</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">SPOT-RNA2</td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">51.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">44.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">84.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">84.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">84.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">68.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">1.5</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">SPOT-RNA</td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">50.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">57.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">52.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">73.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">93.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">60.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">62.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">3</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">UFold</td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="left" colspan="1" rowspan="1"> </td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">45.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">33.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">77.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">4</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">48.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">59.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">40.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">4</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">47.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">4</td></tr></tbody></table><table-wrap-foot><fn id="t2fn1"><label>a</label><p>The median rank column summarizes
the rank median of each tool in the three test sets, and the overall
median F1 column computes the median F1 score of the three test-set
median F1 scores. For e2eFold, no reasonable result was provided in
TestSetα. Notations: P—precision, R—recall.</p></fn></table-wrap-foot></table-wrap></sec><sec id="sec3.3"><title>DEBFold Is Robust against Different Thermodynamics-Constrained
Optimization Algorithms</title><p>The second step of DEBFold feeds
the computed structure location folding probabilities into thermodynamically
constrained optimization algorithms to obtain the final RNA secondary
structure prediction. Previous sections show that DEBFold can have
a better structure prediction over the integrated constitutional prediction
results by using more accurate folding scores as the optimization
soft constraints. We next evaluated the robustness of DEBFold against
different constrained optimization algorithms. In DEBFold, Fold was
the adopted algorithm for the thermodynamically constrained optimization
of the final predicted structures. Besides Fold, RNAfold and ShapeKnots
can also help perform thermodynamics-constrained optimization based
on soft constraints. These three constrained optimization algorithms
were separately combined with DEBFold Stage I (DEBFold-Fold, DEBFold-RNAfold,
and DEBFold-ShapeKnots) and then evaluated on TestSetα, TestSetβ,
and TestSetγ. The comparison results are summarized in <xref rid="tbl3" ref-type="other">Table <xref rid="tbl3" ref-type="other">3</xref></xref>. As shown in <xref rid="tbl3" ref-type="other">Table <xref rid="tbl3" ref-type="other">3</xref></xref>, DEBFold achieves
nearly identical median F1 score results (within one percent) on the
three test sets when different thermodynamics-constrained optimization
tools are used. From this comparison, it is suggested that DEBFold
can provide accurate folding scores that help boost the final structure
prediction and is robust against different thermodynamics-constrained
optimization tools.</p><table-wrap id="tbl3" position="float" orientation="portrait"><label>Table 3</label><caption><title>Performance Evaluation for DEBFold
Algorithm Robustness Using the Median F1 Scores on TestSetα,
TestSetβ, and TestSetγ<xref rid="t3fn1" ref-type="table-fn">a</xref></title></caption><table frame="hsides" rules="groups" border="0"><colgroup span="1"><col align="left" span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/></colgroup><thead><tr><th style="border:none;" align="center" colspan="1" rowspan="1">folding
algorithm used</th><th colspan="3" align="center" char="." rowspan="1">TestSetα<hr/></th><th colspan="3" align="center" char="." rowspan="1">TestSetβ<hr/></th><th colspan="3" align="center" char="." rowspan="1">TestSetγ<hr/></th></tr><tr><th style="border:none;" align="center" colspan="1" rowspan="1"> </th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">R (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">R (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">R (%)</th></tr></thead><tbody><tr><td style="border:none;" align="left" colspan="1" rowspan="1">DEBFold-Fold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">62.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">67.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">55.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">56.4</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">56.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">77.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">83.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">73.2</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">DEBFold-RNAfold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">62.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">67.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">56.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">54.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">77.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">83.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">73.2</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">DEBFold-ShapeKnots</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">64.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">62.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">67.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">55.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">55.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">56.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">77.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">83.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">73.2</td></tr></tbody></table><table-wrap-foot><fn id="t3fn1"><label>a</label><p>DEBFold Stage I deep network combined
with different thermodynamics-constrained optimization algorithms
were compared. Notations: P—precision, R—recall.</p></fn></table-wrap-foot></table-wrap></sec><sec id="sec3.4"><title>DEBFold Provides Superior Structure Location Labeling Results
to the Baseline and the Existing Model</title><p>In DEBFold Stage 1,
the base locations where the structure pairings occur are labeled.
The higher performance of the labeling process usually leads to better
constrained optimization results. We compared the structure location
labeling performance of DEBFold Stage 1 with the simple consensus
model and another existing technique called GRASP.<sup><xref ref-type="bibr" rid="ref60">60</xref></sup> In DEBFold Stage 1, the probability of a base location
being involved in the final structure pairings is predicted by a deep
convolutional network. In previous research, Ke et al. proposed using
the XGBoost model on the flattened one-hot encoding tensor of a sequence
window to get the pairing probability of an RNA base location and
implemented this method as a tool called GRASP.<sup><xref ref-type="bibr" rid="ref60">60</xref></sup> In both DEBFold Stage 1 and GRASP, a fair threshold of
0.5 was adopted when the existence of base pairings was labeled for
different RNA base locations. On the other hand, the simple consensus
model labels the existence of the base pairing for an RNA base location
if all 6 tools (RNAfold,<sup><xref ref-type="bibr" rid="ref25">25</xref></sup> IPknot,<sup><xref ref-type="bibr" rid="ref6">6</xref></sup> MaxExpect,<sup><xref ref-type="bibr" rid="ref26">26</xref></sup> ProbKnot,<sup><xref ref-type="bibr" rid="ref27">27</xref></sup> RNAProb,<sup><xref ref-type="bibr" rid="ref28">28</xref></sup> and Fold<sup><xref ref-type="bibr" rid="ref29">29</xref></sup>) agree on the pairing existence. In order to
evaluate the location labeling results, we resort to the confusion
matrix generated from the labeling results of a specific tool for
each RNA. For structure location labeling in each RNA, TP is the number
of correctly labeled pairing locations, FN represents the number of
missed pairing locations, FP shows the number of nonpairing locations
mistakenly marked to be paired, and TN counts the number of correctly
labeled nonpairing locations. Based on the defined confusion matrix,
the recall, precision, and F1 score values for the structure location
labeling problem are defined similarly to those in <xref rid="eq15" ref-type="disp-formula">eqs <xref rid="eq15" ref-type="disp-formula">15</xref></xref> and <xref rid="eq16" ref-type="disp-formula">16</xref>.
We calculated the evaluation metrics for each RNA on the three prepared
test sets (TestSetα, TestSetβ, and TestSetγ) and
obtained the median F1 score of all of the RNAs in each test set.
As summarized in <xref rid="tbl4" ref-type="other">Table <xref rid="tbl4" ref-type="other">4</xref></xref>, DEBFold Stage 1 achieves F1 scores superior to those of both the
consensus model and GRASP on all three test sets. By this comparison,
we conclude that DEBFold Stage 1 is better than the simple consensus
model and the existing tool GRASP in labeling the existence of the
base pairings for RNA base locations.</p><table-wrap id="tbl4" position="float" orientation="portrait"><label>Table 4</label><caption><title>Labeling Performance Comparison among
DEBFold, GRASP, and the Simple Consensus Model on the Three Test Sets
(TestSetα, TestSetβ, and TestSetγ)<xref rid="t4fn1" ref-type="table-fn">a</xref></title></caption><table frame="hsides" rules="groups" border="0"><colgroup span="1"><col align="left" span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/><col align="char" char="." span="1"/></colgroup><thead><tr><th style="border:none;" align="center" colspan="1" rowspan="1">folding
algorithm used</th><th colspan="3" align="center" char="." rowspan="1">TestSetα<hr/></th><th colspan="3" align="center" char="." rowspan="1">TestSetβ<hr/></th><th colspan="3" align="center" char="." rowspan="1">TestSetγ<hr/></th></tr><tr><th style="border:none;" align="center" colspan="1" rowspan="1"> </th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">R (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">R (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">F1 (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">P (%)</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">R (%)</th></tr></thead><tbody><tr><td style="border:none;" align="left" colspan="1" rowspan="1"><bold>DEBFold Stage 1</bold></td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">77.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">72.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">82.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">71.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">61.8</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">85.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">79.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">87.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">73.2</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">the simple consensus model</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">70.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">76.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">65.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">55.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">50.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">77.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">88.9</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">68.3</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">GRASP</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">52.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">45.7</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">61.5</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">51.0</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">48.1</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">54.2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.3</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">77.6</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">53.5</td></tr></tbody></table><table-wrap-foot><fn id="t4fn1"><label>a</label><p>The median value is recorded for
each test set. Notations: P – precision, R – recall.</p></fn></table-wrap-foot></table-wrap></sec><sec id="sec3.5"><title>Speed Comparison of DEBFold with the Existing Tools</title><p>Due to the advancement of sequencing technology, lots of RNA sequences
are found, leading to the need for large-scale structure investigation
directly on the sequences. To fulfill this need, the execution time
required for each prediction tool should also be inspected in addition
to structure prediction accuracy. We evaluated and compared the execution
time of DEBFold and different structure prediction tools on the collected
147 PDB-derived sequences using a workstation with Intel i7 CPU cores
and 128GB RAM. The results are summarized in <xref rid="tbl5" ref-type="other">Table <xref rid="tbl5" ref-type="other">5</xref></xref>. On these 147 RNA sequences, most single-input folding algorithms
used less than a minute to process the predictions. Since DEBFold
integrates the outcomes of six different single-input folding prediction
results, it took around 1 min to finish all predictions. For algorithms
that consider homologous sequences (RNAalifold,<sup><xref ref-type="bibr" rid="ref14">14</xref></sup> TurboFold,<sup><xref ref-type="bibr" rid="ref29">29</xref></sup> SPARSE,<sup><xref ref-type="bibr" rid="ref50">50</xref></sup> aliFreeFold,<sup><xref ref-type="bibr" rid="ref15">15</xref></sup> LocARNA,<sup><xref ref-type="bibr" rid="ref51">51</xref></sup> comRNA,<sup><xref ref-type="bibr" rid="ref52">52</xref></sup> and MXSCARNA<sup><xref ref-type="bibr" rid="ref53">53</xref></sup>), around an additional 10 min was demanded to
get the homologous sequences using BLAST. Therefore, DEBFold has better
time efficiency than these homology-based tools and achieves better
performance. In deep-learning-based predictions, SPOT-RNA required
about 3 times more execution time than did DEBFold. Among all tools,
SPOT-RNA2 had very low execution efficiency and needed more than 142
h, which is more than 7500 times as long as the execution time of
DEBFold, to predict merely the structures of the 15 TestSetγ
sequences selected by stratified sampling from all 147 RNAs. More
execution time is required to finish all 147 RNAs. The high execution
time of SPOT-RNA2 suggests that SPOT-RNA2 is probably not ready for
a large-scale RNA structure investigation. Moreover, SPOT-RNA2 requires
more than 2TB of disk space to store the information needed for its
algorithm. In summary, DEBFold not only achieves state-of-the-art
prediction accuracy but also retains an acceptable execution time
for large-scale RNA sequence investigation.</p><table-wrap id="tbl5" position="float" orientation="portrait"><label>Table 5</label><caption><title>Execution Time Comparison between
DEBFold and Other Available RNA Secondary Structure Prediction Tools
on the Collected 147 PDB-Derived Sequences</title></caption><table frame="hsides" rules="groups" border="0"><colgroup span="1"><col align="left" span="1"/><col align="char" char="." span="1"/></colgroup><thead><tr><th style="border:none;" align="center" colspan="1" rowspan="1">structure prediction tool</th><th style="border:none;" align="center" char="." colspan="1" rowspan="1">execution time (s)</th></tr></thead><tbody><tr><td style="border:none;" align="left" colspan="1" rowspan="1"><bold>DEBFold</bold></td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">67.3</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">RNAfold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">7.7</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">RNALfold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">9.2</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">PKNOTS</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">12.2</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">Fold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">13.4</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">IterativeHFold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">15.5</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">IPknot</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">17.3</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">ProbKnot</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">20.5</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">MaxExpect</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">20.6</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">RNAProb</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">22.1</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">ShapeKnots</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">382.8</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">HotKnots</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">2824.0</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">RME (PARS model + DMS model)</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">74906.6</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">(Homology finding
using BLAST)</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">591.1</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">RNAalifold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">6.1</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">SPARSE</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">10.1</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">MXSCARNA</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">14.9</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">LocARNA</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">22.9</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">aliFreeFold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">31.6</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">TurboFold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">190.0</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">comRNA</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">904.3</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">MXfold2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">33.6</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">GCNfold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">63.8</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">REDFold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">72.9</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">UFold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">102.8</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">e2eFold</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">123.5</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">SPOT-RNA</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">192.8</td></tr><tr><td style="border:none;" align="left" colspan="1" rowspan="1">SPOT-RNA2</td><td style="border:none;" align="char" char="." colspan="1" rowspan="1">511970.3<xref rid="t5fn1" ref-type="table-fn">a</xref></td></tr></tbody></table><table-wrap-foot><fn id="t5fn1"><label>a</label><p>Note that the execution time accumulated
for SPOT-RNA2 is only from the prediction of the 15 TestSetγ
sequences selected by stratified sampling from the total available
147 PDB-derived sequences.</p></fn></table-wrap-foot></table-wrap></sec><sec id="sec3.6"><title>Limitations of DEBFold</title><p>In this work, we developed a
deep-learning model that can integrate the prediction results from
different single-sequence folding tools to more accurately predict
RNA secondary structures. DEBFold is a two-stage pipeline for RNA
secondary structure prediction. In the first stage, DEBFold identifies
the locations where structural pairings occur. Then, the pairing probabilities
for the base locations of a given sequence are transformed into the
SHAPE-like score tensor. Based on the SHAPE-like score tensor, the
structure prediction bearing the minimum free energy under these soft
constraints is given by Fold. In RNA structure prediction, pseudoknots
(intertwined helices on a plane) and noncanonical pairings (the base
pairs formed by hydrogen bonding differing from the patterns of standard
Watson–Crick rules<sup><xref ref-type="bibr" rid="ref61">61</xref></sup>) present significant
challenges for prediction algorithms. Since DEBFold Stage 1 does not
distinguish pseudoknot, canonical, and noncanonical RNA pairings,
these pairings are all potentially considered in this stage. It is
possible to adopt ShapeKnots or other minimum free energy structure
prediction optimizers in DEBFold Stage 2 if these challenging types
of base pairings are to be considered in the final optimization process.
However, since there are insufficient noncanonical or pseudoknot pairings
in the ground-truth data set, DEBFold currently mainly focuses on
canonical base pairings.</p><p>Previously, Szikszai et al.<sup><xref ref-type="bibr" rid="ref23">23</xref></sup> pointed out that deep-learning models for RNA
secondary structure prediction can easily overfit a data set if a
sequence-based cross-validation and a test process are adopted. This
type of overfitting occurs even when a massive number of sequences
from only a few families are provided. Szikszai et al. argued that
the chief cause for the phenomenon lies within the structural similarities
among RNA sequences in the same family. Based on their research, we
adopted a family-wise approach in developing DEBFold and further prepared
test sets that are more suitable for evaluating this problem in existing
deep-learning-based prediction tools. Although DEBFold overcomes the
overfitting problem and provides state-of-the-art performance, inherent
limitations remain. Currently, the known RNA families are far smaller
than the known sequences in the community. Despite the large number
of RNAs with known sequences available, most of them belong to the
same structural families. Because of the limited RNA structural families
available (only 2606 different families in Rfam 14.4), the performance
of single-sequence structure prediction using deep learning is largely
restricted. The limited RNA structural families also constrain the
prediction accuracy of long noncoding RNAs (lncRNAs) since the overfitting
problem boils down to the lack of structural family diversity. Currently,
popular data sets contain scarce numbers of structural families for
long RNAs, inhibiting the structural prediction of lncRNAs. The current
version of DEBFold is suitable only for sequences up to 512 bps. Although
some of the existing deep-learning-based tools claim to be able to
deal with lncRNAs, the moderate performance observed in <xref rid="tbl2" ref-type="other">Table <xref rid="tbl2" ref-type="other">2</xref></xref> indicates that the originally
claimed prediction accuracy may still need more attention. Notice
that DEBFold performs better than classical folding algorithms that
incorporate multiple-sequence alignments. Therefore, whether multiple-sequence
alignments can better help suggest structure modules in the DEBFold
pipeline requires further study.</p></sec></sec><sec id="sec4"><title>Conclusions</title><p>In this research, we designed and implemented
a deep learning tool
called DEBFold to provide accurate RNA secondary structure prediction
for sequences across RNA structural families. DEBFold is verified
to be free of the overfitting pitfall occurring in many deep-learning-based
structure prediction tools and outperforms the currently available
tools. We believe that the development of DEBFold can significantly
accelerate the use of artificial intelligence to help us understand
structure–function relations for RNAs.</p></sec></body><back><notes notes-type="data-availability" id="notes2"><title>Data Availability Statement</title><p>The implemented
DEBFold pipeline and the processed RNA structure ground-truth data
sets (including the training-validation set, TestSetα, TestSetβ,
and TestSetγ) are available at <uri xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://cobis.bme.ncku.edu.tw/DEBFold">https://cobis.bme.ncku.edu.tw/DEBFold</uri> and <uri xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/cobisLab/DEBFold">https://github.com/cobisLab/DEBFold</uri>.</p></notes><notes notes-type="COI-statement" id="notes1"><p>The author
declares no competing financial interest.</p></notes><ack><title>Acknowledgments</title><p>The author would like to thank Li-Chang
Teng, Zhan-Yi
Liao, and Min Hsia for their help in the early stage exploration of
the works related to this research. This study was supported by National
Cheng Kung University and the National Science and Technology Council
of Taiwan (MOST 110-2222-E-006-017, MOST 111-2221-E-006-231, and NSTC
112-2221-E-006-129-MY2). The study was also supported by the Headquarters
of University Advancement at National Cheng Kung University and the
Ministry of Education, Taiwan.</p></ack><ref-list><title>References</title><ref id="ref1"><mixed-citation publication-type="journal" id="cit1"><name name-style="western"><surname>Kwok</surname><given-names>C. K.</given-names></name>; <name name-style="western"><surname>Tang</surname><given-names>Y.</given-names></name>; <name name-style="western"><surname>Assmann</surname><given-names>S. M.</given-names></name>; <name name-style="western"><surname>Bevilacqua</surname><given-names>P. C.</given-names></name>
<article-title>The RNA
structurome: transcriptome-wide structure probing with next-generation
sequencing</article-title>. <source>Trends Biochem. Sci.</source>
<year>2015</year>, <volume>40</volume>, <fpage>221</fpage>–<lpage>232</lpage>. <pub-id pub-id-type="doi">10.1016/j.tibs.2015.02.005</pub-id>.<pub-id pub-id-type="pmid">25797096</pub-id>
</mixed-citation></ref><ref id="ref2"><mixed-citation publication-type="journal" id="cit2"><name name-style="western"><surname>Wan</surname><given-names>Y.</given-names></name>; <name name-style="western"><surname>Kertesz</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Spitale</surname><given-names>R. C.</given-names></name>; <name name-style="western"><surname>Segal</surname><given-names>E.</given-names></name>; <name name-style="western"><surname>Chang</surname><given-names>H. Y.</given-names></name>
<article-title>Understanding
the transcriptome through RNA structure</article-title>. <source>Nat.
Rev. Genet.</source>
<year>2011</year>, <volume>12</volume>, <fpage>641</fpage>–<lpage>655</lpage>. <pub-id pub-id-type="doi">10.1038/nrg3049</pub-id>.<pub-id pub-id-type="pmid">21850044</pub-id>
<pub-id pub-id-type="pmcid">PMC3858389</pub-id></mixed-citation></ref><ref id="ref3"><mixed-citation publication-type="journal" id="cit3"><name name-style="western"><surname>Yang</surname><given-names>T.-H.</given-names></name>
<article-title>An aggregation
method to identify the RNA meta-stable secondary structure and its
functionally interpretable structure ensemble</article-title>. <source>IEEE/ACM Trans. Comput. Biol. Bioinf.</source>
<year>2022</year>, <volume>19</volume>, <fpage>75</fpage>–<lpage>86</lpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2021.3082396</pub-id>.<pub-id pub-id-type="pmid">34014829</pub-id></mixed-citation></ref><ref id="ref4"><mixed-citation publication-type="journal" id="cit4"><name name-style="western"><surname>Baird</surname><given-names>S. D.</given-names></name>; <name name-style="western"><surname>Lewis</surname><given-names>S. M.</given-names></name>; <name name-style="western"><surname>Turcotte</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Holcik</surname><given-names>M.</given-names></name>
<article-title>A search for
structurally
similar cellular internal ribosome entry sites</article-title>. <source>Nucleic Acids Res.</source>
<year>2007</year>, <volume>35</volume>, <fpage>4664</fpage>–<lpage>4677</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkm483</pub-id>.<pub-id pub-id-type="pmid">17591613</pub-id>
<pub-id pub-id-type="pmcid">PMC1950536</pub-id></mixed-citation></ref><ref id="ref5"><mixed-citation publication-type="journal" id="cit5"><name name-style="western"><surname>Yang</surname><given-names>T.-H.</given-names></name>; <name name-style="western"><surname>Wang</surname><given-names>C.-Y.</given-names></name>; <name name-style="western"><surname>Tsai</surname><given-names>H.-C.</given-names></name>; <name name-style="western"><surname>Liu</surname><given-names>C.-T.</given-names></name>
<article-title>Human IRES Atlas:
an integrative platform for studying IRES-driven translational regulation
in humans</article-title>. <source>Database</source>
<year>2021</year>, <volume>2021</volume>, <fpage>baab025</fpage><pub-id pub-id-type="doi">10.1093/database/baab025</pub-id>.<pub-id pub-id-type="pmid">33942874</pub-id>
<pub-id pub-id-type="pmcid">PMC8094437</pub-id></mixed-citation></ref><ref id="ref6"><mixed-citation publication-type="journal" id="cit6"><name name-style="western"><surname>Sato</surname><given-names>K.</given-names></name>; <name name-style="western"><surname>Kato</surname><given-names>Y.</given-names></name>; <name name-style="western"><surname>Hamada</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Akutsu</surname><given-names>T.</given-names></name>; <name name-style="western"><surname>Asai</surname><given-names>K.</given-names></name>
<article-title>IPknot: fast
and accurate prediction of RNA secondary structures with pseudoknots
using integer programming</article-title>. <source>Bioinformatics</source>
<year>2011</year>, <volume>27</volume>, <fpage>i85</fpage>–<lpage>i93</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btr215</pub-id>.<pub-id pub-id-type="pmid">21685106</pub-id>
<pub-id pub-id-type="pmcid">PMC3117384</pub-id></mixed-citation></ref><ref id="ref7"><mixed-citation publication-type="journal" id="cit7"><name name-style="western"><surname>Vandivier</surname><given-names>L. E.</given-names></name>; <name name-style="western"><surname>Anderson</surname><given-names>S. J.</given-names></name>; <name name-style="western"><surname>Foley</surname><given-names>S. W.</given-names></name>; <name name-style="western"><surname>Gregory</surname><given-names>B. D.</given-names></name>
<article-title>The conservation
and function of RNA secondary structure in plants</article-title>. <source>Annu. Rev. Plant Biol.</source>
<year>2016</year>, <volume>67</volume>, <fpage>463</fpage>–<lpage>488</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-arplant-043015-111754</pub-id>.<pub-id pub-id-type="pmid">26865341</pub-id>
<pub-id pub-id-type="pmcid">PMC5125251</pub-id></mixed-citation></ref><ref id="ref8"><mixed-citation publication-type="journal" id="cit8"><name name-style="western"><surname>Yang</surname><given-names>T.-H.</given-names></name>; <name name-style="western"><surname>Lin</surname><given-names>Y.-C.</given-names></name>; <name name-style="western"><surname>Hsia</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Liao</surname><given-names>Z.-Y.</given-names></name>
<article-title>SSRTool: a web tool
for evaluating RNA secondary structure predictions based on species-specific
functional interpretability</article-title>. <source>Comput. Struct.
Biotechnol. J.</source>
<year>2022</year>, <volume>20</volume>, <fpage>2473</fpage>–<lpage>2483</lpage>. <pub-id pub-id-type="doi">10.1016/j.csbj.2022.05.028</pub-id>.<pub-id pub-id-type="pmid">35664227</pub-id>
<pub-id pub-id-type="pmcid">PMC9136272</pub-id></mixed-citation></ref><ref id="ref9"><mixed-citation publication-type="journal" id="cit9"><name name-style="western"><surname>Dagenais</surname><given-names>P.</given-names></name>; <name name-style="western"><surname>Girard</surname><given-names>N.</given-names></name>; <name name-style="western"><surname>Bonneau</surname><given-names>E.</given-names></name>; <name name-style="western"><surname>Legault</surname><given-names>P.</given-names></name>
<article-title>Insights into RNA structure
and dynamics from recent NMR and X-ray studies of the Neurospora Varkud
satellite ribozyme</article-title>. <source>Wiley Interdiscip. Rev.:
RNA</source>
<year>2017</year>, <volume>8</volume>, <elocation-id>e1421</elocation-id><pub-id pub-id-type="doi">10.1002/wrna.1421</pub-id>.<pub-id pub-id-type="pmid">28382748</pub-id>
<pub-id pub-id-type="pmcid">PMC5573960</pub-id></mixed-citation></ref><ref id="ref10"><mixed-citation publication-type="journal" id="cit10"><name name-style="western"><surname>Ma</surname><given-names>H.</given-names></name>; <name name-style="western"><surname>Jia</surname><given-names>X.</given-names></name>; <name name-style="western"><surname>Zhang</surname><given-names>K.</given-names></name>; <name name-style="western"><surname>Su</surname><given-names>Z.</given-names></name>
<article-title>Cryo-EM advances in RNA structure
determination</article-title>. <source>Signal Transduction Targeted
Ther.</source>
<year>2022</year>, <volume>7</volume>, <fpage>58</fpage><pub-id pub-id-type="doi">10.1038/s41392-022-00916-0</pub-id>.<pub-id pub-id-type="pmcid">PMC8864457</pub-id><pub-id pub-id-type="pmid">35197441</pub-id></mixed-citation></ref><ref id="ref11"><mixed-citation publication-type="journal" id="cit11"><name name-style="western"><surname>Rouskin</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Zubradt</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Washietl</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Kellis</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Weissman</surname><given-names>J. S.</given-names></name>
<article-title>Genome-wide
probing of RNA structure reveals active unfolding of mRNA structures
in vivo</article-title>. <source>Nature</source>
<year>2014</year>, <volume>505</volume>, <fpage>701</fpage>–<lpage>705</lpage>. <pub-id pub-id-type="doi">10.1038/nature12894</pub-id>.<pub-id pub-id-type="pmid">24336214</pub-id>
<pub-id pub-id-type="pmcid">PMC3966492</pub-id></mixed-citation></ref><ref id="ref12"><mixed-citation publication-type="journal" id="cit12"><name name-style="western"><surname>Zuker</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Stiegler</surname><given-names>P.</given-names></name>
<article-title>Optimal computer folding of large RNA sequences using
thermodynamics and auxiliary information</article-title>. <source>Nucleic
Acids Res.</source>
<year>1981</year>, <volume>9</volume>, <fpage>133</fpage>–<lpage>148</lpage>. <pub-id pub-id-type="doi">10.1093/nar/9.1.133</pub-id>.<pub-id pub-id-type="pmid">6163133</pub-id>
<pub-id pub-id-type="pmcid">PMC326673</pub-id></mixed-citation></ref><ref id="ref13"><mixed-citation publication-type="journal" id="cit13"><name name-style="western"><surname>McCaskill</surname><given-names>J. S.</given-names></name>
<article-title>The equilibrium
partition function and base pair binding probabilities for RNA secondary
structure</article-title>. <source>Biopolymers</source>
<year>1990</year>, <volume>29</volume>, <fpage>1105</fpage>–<lpage>1119</lpage>. <pub-id pub-id-type="doi">10.1002/bip.360290621</pub-id>.<pub-id pub-id-type="pmid">1695107</pub-id>
</mixed-citation></ref><ref id="ref14"><mixed-citation publication-type="journal" id="cit14"><name name-style="western"><surname>Bernhart</surname><given-names>S. H.</given-names></name>; <name name-style="western"><surname>Hofacker</surname><given-names>I. L.</given-names></name>; <name name-style="western"><surname>Will</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Gruber</surname><given-names>A. R.</given-names></name>; <name name-style="western"><surname>Stadler</surname><given-names>P. F.</given-names></name>
<article-title>RNAalifold:
improved consensus structure prediction for RNA alignments</article-title>. <source>BMC Bioinf.</source>
<year>2008</year>, <volume>9</volume> (<issue>1</issue>), <fpage>474</fpage><pub-id pub-id-type="doi">10.1186/1471-2105-9-474</pub-id>.<pub-id pub-id-type="pmcid">PMC2621365</pub-id><pub-id pub-id-type="pmid">19014431</pub-id></mixed-citation></ref><ref id="ref15"><mixed-citation publication-type="journal" id="cit15"><name name-style="western"><surname>Glouzon</surname><given-names>J.-P. S.</given-names></name>; <name name-style="western"><surname>Ouangraoua</surname><given-names>A.</given-names></name>
<article-title>aliFreeFold:
an alignment-free approach to predict
secondary structure from homologous RNA sequences</article-title>. <source>Bioinformatics</source>
<year>2018</year>, <volume>34</volume>, <fpage>i70</fpage>–<lpage>i78</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty234</pub-id>.<pub-id pub-id-type="pmid">29949960</pub-id>
<pub-id pub-id-type="pmcid">PMC6022685</pub-id></mixed-citation></ref><ref id="ref16"><mixed-citation publication-type="journal" id="cit16"><name name-style="western"><surname>Lorenz</surname><given-names>R.</given-names></name>; <name name-style="western"><surname>Luntzer</surname><given-names>D.</given-names></name>; <name name-style="western"><surname>Hofacker</surname><given-names>I. L.</given-names></name>; <name name-style="western"><surname>Stadler</surname><given-names>P. F.</given-names></name>; <name name-style="western"><surname>Wolfinger</surname><given-names>M. T.</given-names></name>
<article-title>SHAPE directed
RNA folding</article-title>. <source>Bioinformatics</source>
<year>2016</year>, <volume>32</volume>, <fpage>145</fpage>–<lpage>147</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btv523</pub-id>.<pub-id pub-id-type="pmid">26353838</pub-id>
<pub-id pub-id-type="pmcid">PMC4681990</pub-id></mixed-citation></ref><ref id="ref17"><mixed-citation publication-type="journal" id="cit17"><name name-style="western"><surname>Hollams</surname><given-names>E. M.</given-names></name>; <name name-style="western"><surname>Giles</surname><given-names>K. M.</given-names></name>; <name name-style="western"><surname>Thomson</surname><given-names>A. M.</given-names></name>; <name name-style="western"><surname>Leedman</surname><given-names>P. J.</given-names></name>
<article-title>MRNA stability
and
the control of gene expression: implications for human disease</article-title>. <source>Neurochem. Res.</source>
<year>2002</year>, <volume>27</volume>, <fpage>957</fpage>–<lpage>980</lpage>. <pub-id pub-id-type="doi">10.1023/A:1020992418511</pub-id>.<pub-id pub-id-type="pmid">12462398</pub-id>
</mixed-citation></ref><ref id="ref18"><mixed-citation publication-type="journal" id="cit18"><name name-style="western"><surname>Draper</surname><given-names>D. E.</given-names></name>
<article-title>A guide
to ions and RNA structure</article-title>. <source>RNA</source>
<year>2004</year>, <volume>10</volume>, <fpage>335</fpage>–<lpage>343</lpage>. <pub-id pub-id-type="doi">10.1261/rna.5205404</pub-id>.<pub-id pub-id-type="pmid">14970378</pub-id>
<pub-id pub-id-type="pmcid">PMC1370927</pub-id></mixed-citation></ref><ref id="ref19"><mixed-citation publication-type="journal" id="cit19"><name name-style="western"><surname>Jabbari</surname><given-names>H.</given-names></name>; <name name-style="western"><surname>Wark</surname><given-names>I.</given-names></name>; <name name-style="western"><surname>Montemagno</surname><given-names>C.</given-names></name>
<article-title>RNA secondary structure prediction
with pseudoknots: contribution of algorithm versus energy model</article-title>. <source>PLoS One</source>
<year>2018</year>, <volume>13</volume>, <elocation-id>e0194583</elocation-id><pub-id pub-id-type="doi">10.1371/journal.pone.0194583</pub-id>.<pub-id pub-id-type="pmid">29621250</pub-id>
<pub-id pub-id-type="pmcid">PMC5886407</pub-id></mixed-citation></ref><ref id="ref20"><mixed-citation publication-type="journal" id="cit20"><name name-style="western"><surname>Zhao</surname><given-names>Q.</given-names></name>; <name name-style="western"><surname>Zhao</surname><given-names>Z.</given-names></name>; <name name-style="western"><surname>Fan</surname><given-names>X.</given-names></name>; <name name-style="western"><surname>Yuan</surname><given-names>Z.</given-names></name>; <name name-style="western"><surname>Mao</surname><given-names>Q.</given-names></name>; <name name-style="western"><surname>Yao</surname><given-names>Y.</given-names></name>
<article-title>Review of machine learning methods for RNA secondary structure prediction</article-title>. <source>PLoS Comput. Biol.</source>
<year>2021</year>, <volume>17</volume>, <elocation-id>e1009291</elocation-id><pub-id pub-id-type="doi">10.1371/journal.pcbi.1009291</pub-id>.<pub-id pub-id-type="pmid">34437528</pub-id>
<pub-id pub-id-type="pmcid">PMC8389396</pub-id></mixed-citation></ref><ref id="ref21"><mixed-citation publication-type="journal" id="cit21"><name name-style="western"><surname>Singh</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Hanson</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Paliwal</surname><given-names>K.</given-names></name>; <name name-style="western"><surname>Zhou</surname><given-names>Y.</given-names></name>
<article-title>RNA secondary structure
prediction using an ensemble of two-dimensional deep neural networks
and transfer learning</article-title>. <source>Nat. Commun.</source>
<year>2019</year>, <volume>10</volume>, <fpage>5407</fpage><pub-id pub-id-type="doi">10.1038/s41467-019-13395-9</pub-id>.<pub-id pub-id-type="pmid">31776342</pub-id>
<pub-id pub-id-type="pmcid">PMC6881452</pub-id></mixed-citation></ref><ref id="ref22"><mixed-citation publication-type="journal" id="cit22"><name name-style="western"><surname>Sato</surname><given-names>K.</given-names></name>; <name name-style="western"><surname>Akiyama</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Sakakibara</surname><given-names>Y.</given-names></name>
<article-title>RNA secondary
structure prediction
using deep learning with thermodynamic integration</article-title>. <source>Nat. Commun.</source>
<year>2021</year>, <volume>12</volume> (<issue>1</issue>), <fpage>941</fpage><pub-id pub-id-type="doi">10.1038/s41467-021-21194-4</pub-id>.<pub-id pub-id-type="pmid">33574226</pub-id>
<pub-id pub-id-type="pmcid">PMC7878809</pub-id></mixed-citation></ref><ref id="ref23"><mixed-citation publication-type="journal" id="cit23"><name name-style="western"><surname>Szikszai</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Wise</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Datta</surname><given-names>A.</given-names></name>; <name name-style="western"><surname>Ward</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Mathews</surname><given-names>D. H.</given-names></name>
<article-title>Deep learning
models for RNA secondary structure prediction (probably) do not generalize
across families</article-title>. <source>Bioinformatics</source>
<year>2022</year>, <volume>38</volume>, <fpage>3892</fpage>–<lpage>3899</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btac415</pub-id>.<pub-id pub-id-type="pmid">35748706</pub-id>
<pub-id pub-id-type="pmcid">PMC9364374</pub-id></mixed-citation></ref><ref id="ref24"><mixed-citation publication-type="journal" id="cit24"><name name-style="western"><surname>Kerpedjiev</surname><given-names>P.</given-names></name>; <name name-style="western"><surname>Hammer</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Hofacker</surname><given-names>I. L.</given-names></name>
<article-title>Forna (force-directed
RNA): simple
and effective online RNA secondary structure diagrams</article-title>. <source>Bioinformatics</source>
<year>2015</year>, <volume>31</volume>, <fpage>3377</fpage>–<lpage>3379</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btv372</pub-id>.<pub-id pub-id-type="pmid">26099263</pub-id>
<pub-id pub-id-type="pmcid">PMC4595900</pub-id></mixed-citation></ref><ref id="ref25"><mixed-citation publication-type="journal" id="cit25"><name name-style="western"><surname>Lorenz</surname><given-names>R.</given-names></name>; <name name-style="western"><surname>Bernhart</surname><given-names>S. H.</given-names></name>; <name name-style="western"><surname>Höner zu Siederdissen</surname><given-names>C.</given-names></name>; <name name-style="western"><surname>Tafer</surname><given-names>H.</given-names></name>; <name name-style="western"><surname>Flamm</surname><given-names>C.</given-names></name>; <name name-style="western"><surname>Stadler</surname><given-names>P. F.</given-names></name>; <name name-style="western"><surname>Hofacker</surname><given-names>I. L.</given-names></name>
<article-title>ViennaRNA
Package 2.0</article-title>. <source>Algorithms Mol. Biol.</source>
<year>2011</year>, <volume>6</volume>, <fpage>1</fpage>–<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1186/1748-7188-6-26</pub-id>.<pub-id pub-id-type="pmid">22115189</pub-id>
<pub-id pub-id-type="pmcid">PMC3319429</pub-id></mixed-citation></ref><ref id="ref26"><mixed-citation publication-type="journal" id="cit26"><name name-style="western"><surname>Lu</surname><given-names>Z. J.</given-names></name>; <name name-style="western"><surname>Gloor</surname><given-names>J. W.</given-names></name>; <name name-style="western"><surname>Mathews</surname><given-names>D. H.</given-names></name>
<article-title>Improved RNA secondary structure
prediction by maximizing expected pair accuracy</article-title>. <source>RNA</source>
<year>2009</year>, <volume>15</volume>, <fpage>1805</fpage>–<lpage>1813</lpage>. <pub-id pub-id-type="doi">10.1261/rna.1643609</pub-id>.<pub-id pub-id-type="pmid">19703939</pub-id>
<pub-id pub-id-type="pmcid">PMC2743040</pub-id></mixed-citation></ref><ref id="ref27"><mixed-citation publication-type="journal" id="cit27"><name name-style="western"><surname>Bellaousov</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Mathews</surname><given-names>D. H.</given-names></name>
<article-title>ProbKnot: fast prediction
of RNA secondary structure
including pseudoknots</article-title>. <source>RNA</source>
<year>2010</year>, <volume>16</volume>, <fpage>1870</fpage>–<lpage>1880</lpage>. <pub-id pub-id-type="doi">10.1261/rna.2125310</pub-id>.<pub-id pub-id-type="pmid">20699301</pub-id>
<pub-id pub-id-type="pmcid">PMC2941096</pub-id></mixed-citation></ref><ref id="ref28"><mixed-citation publication-type="journal" id="cit28"><name name-style="western"><surname>Deng</surname><given-names>F.</given-names></name>; <name name-style="western"><surname>Ledda</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Vaziri</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Aviran</surname><given-names>S.</given-names></name>
<article-title>Data-directed RNA secondary
structure prediction using probabilistic modeling</article-title>. <source>RNA</source>
<year>2016</year>, <volume>22</volume>, <fpage>1109</fpage>–<lpage>1119</lpage>. <pub-id pub-id-type="doi">10.1261/rna.055756.115</pub-id>.<pub-id pub-id-type="pmid">27251549</pub-id>
<pub-id pub-id-type="pmcid">PMC4931104</pub-id></mixed-citation></ref><ref id="ref29"><mixed-citation publication-type="journal" id="cit29"><name name-style="western"><surname>Reuter</surname><given-names>J. S.</given-names></name>; <name name-style="western"><surname>Mathews</surname><given-names>D. H.</given-names></name>
<article-title>RNAstructure: software
for RNA secondary structure
prediction and analysis</article-title>. <source>BMC Bioinf.</source>
<year>2010</year>, <volume>11</volume> (<issue>1</issue>), <fpage>129</fpage><pub-id pub-id-type="doi">10.1186/1471-2105-11-129</pub-id>.<pub-id pub-id-type="pmcid">PMC2984261</pub-id><pub-id pub-id-type="pmid">20230624</pub-id></mixed-citation></ref><ref id="ref30"><mixed-citation publication-type="journal" id="cit30"><name name-style="western"><surname>Yang</surname><given-names>T.-H.</given-names></name>; <name name-style="western"><surname>Shiue</surname><given-names>S.-C.</given-names></name>; <name name-style="western"><surname>Chen</surname><given-names>K.-Y.</given-names></name>; <name name-style="western"><surname>Tseng</surname><given-names>Y.-Y.</given-names></name>; <name name-style="western"><surname>Wu</surname><given-names>W.-S.</given-names></name>
<article-title>Identifying
piRNA targets on mRNAs in C. elegans using a deep multi-head attention
network</article-title>. <source>BMC Bioinf.</source>
<year>2021</year>, <volume>22</volume>, <fpage>503</fpage><pub-id pub-id-type="doi">10.1186/s12859-021-04428-6</pub-id>.<pub-id pub-id-type="pmcid">PMC8520261</pub-id><pub-id pub-id-type="pmid">34656087</pub-id></mixed-citation></ref><ref id="ref31"><mixed-citation publication-type="journal" id="cit31"><name name-style="western"><surname>Yang</surname><given-names>T.-H.</given-names></name>; <name name-style="western"><surname>Chen</surname><given-names>J.-C.</given-names></name>; <name name-style="western"><surname>Lee</surname><given-names>Y.-H.</given-names></name>; <name name-style="western"><surname>Lu</surname><given-names>S.-Y.</given-names></name>; <name name-style="western"><surname>Wu</surname><given-names>S.-H.</given-names></name>; <name name-style="western"><surname>Chang</surname><given-names>F.-Y.</given-names></name>; <name name-style="western"><surname>Huang</surname><given-names>Y.-C.</given-names></name>; <name name-style="western"><surname>Lee</surname><given-names>M.-H.</given-names></name>; <name name-style="western"><surname>Tseng</surname><given-names>Y.-Y.</given-names></name>; <name name-style="western"><surname>Wu</surname><given-names>W.-S.</given-names></name>
<article-title>Identifying human miRNA target sites via learning the interaction
patterns between miRNA and mRNA segments</article-title>. <source>J.
Chem. Inf. Model.</source>
<year>2024</year>, <volume>64</volume>, <fpage>2445</fpage>–<lpage>2453</lpage>. <pub-id pub-id-type="doi">10.1021/acs.jcim.3c01150</pub-id>.<pub-id pub-id-type="pmid">37903033</pub-id>
</mixed-citation></ref><ref id="ref32"><mixed-citation publication-type="journal" id="cit32"><name name-style="western"><surname>Yang</surname><given-names>T.-H.</given-names></name>; <name name-style="western"><surname>Yang</surname><given-names>Y.-C.</given-names></name>; <name name-style="western"><surname>Tu</surname><given-names>K.-C.</given-names></name>
<article-title>regCNN:
identifying Drosophila genome-wide
cis-regulatory modules via integrating the local patterns in epigenetic
marks and transcription factor binding motifs</article-title>. <source>Comput. Struct. Biotechnol. J.</source>
<year>2022</year>, <volume>20</volume>, <fpage>296</fpage>–<lpage>308</lpage>. <pub-id pub-id-type="doi">10.1016/j.csbj.2021.12.015</pub-id>.<pub-id pub-id-type="pmid">35035784</pub-id>
<pub-id pub-id-type="pmcid">PMC8724954</pub-id></mixed-citation></ref><ref id="ref33"><mixed-citation publication-type="journal" id="cit33"><name name-style="western"><surname>Yang</surname><given-names>T.-H.</given-names></name>; <name name-style="western"><surname>Yu</surname><given-names>Y.-H.</given-names></name>; <name name-style="western"><surname>Wu</surname><given-names>S.-H.</given-names></name>; <name name-style="western"><surname>Zhang</surname><given-names>F.-Y.</given-names></name>
<article-title>CFA: An explainable
deep learning model for annotating the transcriptional roles of cis-regulatory
modules based on epigenetic codes</article-title>. <source>Comput. Biol.
Med.</source>
<year>2023</year>, <volume>152</volume>, <fpage>106375</fpage><pub-id pub-id-type="doi">10.1016/j.compbiomed.2022.106375</pub-id>.<pub-id pub-id-type="pmid">36502693</pub-id>
</mixed-citation></ref><ref id="ref34"><mixed-citation publication-type="conf-proc" id="cit34"><person-group person-group-type="allauthors"><name name-style="western"><surname>He</surname><given-names>K.</given-names></name>; <name name-style="western"><surname>Zhang</surname><given-names>X.</given-names></name>; <name name-style="western"><surname>Ren</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Sun</surname><given-names>J.</given-names></name></person-group><article-title>Deep residual learning for
image recognition</article-title>. <source>2016 IEEE Conference on
Computer Vision and Pattern Recognition
(CVPR)</source>. <year>2016</year>; pp <fpage>770</fpage>–<lpage>778</lpage>, <pub-id pub-id-type="doi">10.1109/cvpr.2016.90</pub-id>.</mixed-citation></ref><ref id="ref35"><mixed-citation publication-type="journal" id="cit35"><name name-style="western"><surname>Vaswani</surname><given-names>A.</given-names></name>; <name name-style="western"><surname>Shazeer</surname><given-names>N.</given-names></name>; <name name-style="western"><surname>Parmar</surname><given-names>N.</given-names></name>; <name name-style="western"><surname>Uszkoreit</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Jones</surname><given-names>L.</given-names></name>; <name name-style="western"><surname>Gomez</surname><given-names>A. N.</given-names></name>; <name name-style="western"><surname>Kaiser</surname><given-names>Ł.</given-names></name>; <name name-style="western"><surname>Polosukhin</surname><given-names>I.</given-names></name>
<article-title>Attention
is all you need</article-title>. <source>Adv. Neur.</source>
<year>2017</year>, <volume>30</volume>, <fpage>30</fpage>.</mixed-citation></ref><ref id="ref36"><mixed-citation publication-type="journal" id="cit36"><name name-style="western"><surname>Lorenz</surname><given-names>R.</given-names></name>; <name name-style="western"><surname>Hofacker</surname><given-names>I. L.</given-names></name>; <name name-style="western"><surname>Stadler</surname><given-names>P. F.</given-names></name>
<article-title>RNA folding with hard and soft constraints</article-title>. <source>Algorithms Mol. Biol.</source>
<year>2016</year>, <volume>11</volume>, <fpage>8</fpage><pub-id pub-id-type="doi">10.1186/s13015-016-0070-z</pub-id>.<pub-id pub-id-type="pmid">27110276</pub-id>
<pub-id pub-id-type="pmcid">PMC4842303</pub-id></mixed-citation></ref><ref id="ref37"><mixed-citation publication-type="journal" id="cit37"><name name-style="western"><surname>Deigan</surname><given-names>K. E.</given-names></name>; <name name-style="western"><surname>Li</surname><given-names>T. W.</given-names></name>; <name name-style="western"><surname>Mathews</surname><given-names>D. H.</given-names></name>; <name name-style="western"><surname>Weeks</surname><given-names>K. M.</given-names></name>
<article-title>Accurate SHAPE-directed
RNA structure determination</article-title>. <source>Proc. Natl. Acad.
Sci. U.S.A.</source>
<year>2009</year>, <volume>106</volume>, <fpage>97</fpage>–<lpage>102</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0806929106</pub-id>.<pub-id pub-id-type="pmid">19109441</pub-id>
<pub-id pub-id-type="pmcid">PMC2629221</pub-id></mixed-citation></ref><ref id="ref38"><mixed-citation publication-type="book" id="cit38"><person-group person-group-type="allauthors"><name name-style="western"><surname>Abu-Mostafa</surname><given-names>Y. S.</given-names></name>; <name name-style="western"><surname>Magdon-Ismail</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Lin</surname><given-names>H.-T.</given-names></name></person-group><source>Learning from Data</source>; <publisher-name>AMLBook</publisher-name>: <publisher-loc>New York</publisher-loc>, <year>2012</year>; Vol. <volume>4</volume>.</mixed-citation></ref><ref id="ref39"><mixed-citation publication-type="journal" id="cit39"><name name-style="western"><surname>Danaee</surname><given-names>P.</given-names></name>; <name name-style="western"><surname>Rouches</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Wiley</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Deng</surname><given-names>D.</given-names></name>; <name name-style="western"><surname>Huang</surname><given-names>L.</given-names></name>; <name name-style="western"><surname>Hendrix</surname><given-names>D.</given-names></name>
<article-title>bpRNA: large-scale
automated annotation and analysis
of RNA secondary structure</article-title>. <source>Nucleic Acids Res.</source>
<year>2018</year>, <volume>46</volume>, <fpage>5381</fpage>–<lpage>5394</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky285</pub-id>.<pub-id pub-id-type="pmid">29746666</pub-id>
<pub-id pub-id-type="pmcid">PMC6009582</pub-id></mixed-citation></ref><ref id="ref40"><mixed-citation publication-type="journal" id="cit40"><name name-style="western"><surname>Nawrocki</surname><given-names>E. P.</given-names></name>; <name name-style="western"><surname>Burge</surname><given-names>S. W.</given-names></name>; <name name-style="western"><surname>Bateman</surname><given-names>A.</given-names></name>; <name name-style="western"><surname>Daub</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Eberhardt</surname><given-names>R. Y.</given-names></name>; <name name-style="western"><surname>Eddy</surname><given-names>S. R.</given-names></name>; <name name-style="western"><surname>Floden</surname><given-names>E. W.</given-names></name>; <name name-style="western"><surname>Gardner</surname><given-names>P. P.</given-names></name>; <name name-style="western"><surname>Jones</surname><given-names>T. A.</given-names></name>; <name name-style="western"><surname>Tate</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Finn</surname><given-names>R. D.</given-names></name>
<article-title>Rfam 12.0: updates to the RNA families
database</article-title>. <source>Nucleic Acids Res.</source>
<year>2015</year>, <volume>43</volume>, <fpage>D130</fpage>–<lpage>D137</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gku1063</pub-id>.<pub-id pub-id-type="pmid">25392425</pub-id>
<pub-id pub-id-type="pmcid">PMC4383904</pub-id></mixed-citation></ref><ref id="ref41"><mixed-citation publication-type="journal" id="cit41"><name name-style="western"><surname>Fu</surname><given-names>L.</given-names></name>; <name name-style="western"><surname>Niu</surname><given-names>B.</given-names></name>; <name name-style="western"><surname>Zhu</surname><given-names>Z.</given-names></name>; <name name-style="western"><surname>Wu</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Li</surname><given-names>W.</given-names></name>
<article-title>CD-HIT: accelerated
for clustering the next-generation sequencing data</article-title>. <source>Bioinformatics</source>
<year>2012</year>, <volume>28</volume>, <fpage>3150</fpage>–<lpage>3152</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bts565</pub-id>.<pub-id pub-id-type="pmid">23060610</pub-id>
<pub-id pub-id-type="pmcid">PMC3516142</pub-id></mixed-citation></ref><ref id="ref42"><mixed-citation publication-type="journal" id="cit42"><name name-style="western"><surname>Löwes</surname><given-names>B.</given-names></name>; <name name-style="western"><surname>Chauve</surname><given-names>C.</given-names></name>; <name name-style="western"><surname>Ponty</surname><given-names>Y.</given-names></name>; <name name-style="western"><surname>Giegerich</surname><given-names>R.</given-names></name>
<article-title>The BRaliBase
dent—a tale of benchmark design and interpretation</article-title>. <source>Brief. Bioinform.</source>
<year>2017</year>, <volume>18</volume> (<issue>2</issue>), <fpage>306</fpage>–<lpage>311</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbw022</pub-id>.<pub-id pub-id-type="pmid">26984616</pub-id>
<pub-id pub-id-type="pmcid">PMC5444242</pub-id></mixed-citation></ref><ref id="ref43"><mixed-citation publication-type="journal" id="cit43"><name name-style="western"><surname>Kalvari</surname><given-names>I.</given-names></name>; <name name-style="western"><surname>Nawrocki</surname><given-names>E. P.</given-names></name>; <name name-style="western"><surname>Ontiveros-Palacios</surname><given-names>N.</given-names></name>; <name name-style="western"><surname>Argasinska</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Lamkiewicz</surname><given-names>K.</given-names></name>; <name name-style="western"><surname>Marz</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Griffiths-Jones</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Toffano-Nioche</surname><given-names>C.</given-names></name>; <name name-style="western"><surname>Gautheret</surname><given-names>D.</given-names></name>; <name name-style="western"><surname>Weinberg</surname><given-names>Z.</given-names></name>; <name name-style="western"><surname>Rivas</surname><given-names>E.</given-names></name>; <name name-style="western"><surname>Eddy</surname><given-names>S. R.</given-names></name>; <name name-style="western"><surname>Finn</surname><given-names>R.</given-names></name>; <name name-style="western"><surname>Bateman</surname><given-names>A.</given-names></name>; <name name-style="western"><surname>Petrov</surname><given-names>A. I.</given-names></name>
<article-title>Rfam 14:
expanded coverage of metagenomic, viral and microRNA families</article-title>. <source>Nucleic Acids Res.</source>
<year>2021</year>, <volume>49</volume>, <fpage>D192</fpage>–<lpage>D200</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkaa1047</pub-id>.<pub-id pub-id-type="pmid">33211869</pub-id>
<pub-id pub-id-type="pmcid">PMC7779021</pub-id></mixed-citation></ref><ref id="ref44"><mixed-citation publication-type="journal" id="cit44"><name name-style="western"><surname>Hofacker</surname><given-names>I. L.</given-names></name>; <name name-style="western"><surname>Priwitzer</surname><given-names>B.</given-names></name>; <name name-style="western"><surname>Stadler</surname><given-names>P. F.</given-names></name>
<article-title>Prediction of locally stable RNA
secondary structures for genome-wide surveys</article-title>. <source>Bioinformatics</source>
<year>2004</year>, <volume>20</volume>, <fpage>186</fpage>–<lpage>190</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btg388</pub-id>.<pub-id pub-id-type="pmid">14734309</pub-id>
</mixed-citation></ref><ref id="ref45"><mixed-citation publication-type="journal" id="cit45"><name name-style="western"><surname>Wu</surname><given-names>Y.</given-names></name>; <name name-style="western"><surname>Shi</surname><given-names>B.</given-names></name>; <name name-style="western"><surname>Ding</surname><given-names>X.</given-names></name>; <name name-style="western"><surname>Liu</surname><given-names>T.</given-names></name>; <name name-style="western"><surname>Hu</surname><given-names>X.</given-names></name>; <name name-style="western"><surname>Yip</surname><given-names>K. Y.</given-names></name>; <name name-style="western"><surname>Yang</surname><given-names>Z. R.</given-names></name>; <name name-style="western"><surname>Mathews</surname><given-names>D. H.</given-names></name>; <name name-style="western"><surname>Lu</surname><given-names>Z. J.</given-names></name>
<article-title>Improved prediction
of RNA secondary structure by integrating the free energy model with
restraints derived from experimental probing data</article-title>. <source>Nucleic Acids Res.</source>
<year>2015</year>, <volume>43</volume>, <fpage>7247</fpage>–<lpage>7259</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkv706</pub-id>.<pub-id pub-id-type="pmid">26170232</pub-id>
<pub-id pub-id-type="pmcid">PMC4551937</pub-id></mixed-citation></ref><ref id="ref46"><mixed-citation publication-type="journal" id="cit46"><name name-style="western"><surname>Rivas</surname><given-names>E.</given-names></name>; <name name-style="western"><surname>Eddy</surname><given-names>S. R.</given-names></name>
<article-title>A dynamic programming
algorithm for RNA structure prediction
including pseudoknots 1 1Edited by I. Tinoco</article-title>. <source>J. Mol. Biol.</source>
<year>1999</year>, <volume>285</volume>, <fpage>2053</fpage>–<lpage>2068</lpage>. <pub-id pub-id-type="doi">10.1006/jmbi.1998.2436</pub-id>.<pub-id pub-id-type="pmid">9925784</pub-id>
</mixed-citation></ref><ref id="ref47"><mixed-citation publication-type="journal" id="cit47"><name name-style="western"><surname>Ren</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Rastegari</surname><given-names>B.</given-names></name>; <name name-style="western"><surname>Condon</surname><given-names>A.</given-names></name>; <name name-style="western"><surname>Hoos</surname><given-names>H. H.</given-names></name>
<article-title>HotKnots: heuristic
prediction of RNA secondary structures including pseudoknots</article-title>. <source>RNA</source>
<year>2005</year>, <volume>11</volume>, <fpage>1494</fpage>–<lpage>1504</lpage>. <pub-id pub-id-type="doi">10.1261/rna.7284905</pub-id>.<pub-id pub-id-type="pmid">16199760</pub-id>
<pub-id pub-id-type="pmcid">PMC1370833</pub-id></mixed-citation></ref><ref id="ref48"><mixed-citation publication-type="journal" id="cit48"><name name-style="western"><surname>Jabbari</surname><given-names>H.</given-names></name>; <name name-style="western"><surname>Condon</surname><given-names>A.</given-names></name>
<article-title>A fast and robust iterative
algorithm for prediction
of RNA pseudoknotted secondary structures</article-title>. <source>BMC
Bioinf.</source>
<year>2014</year>, <volume>15</volume>, <fpage>147</fpage><pub-id pub-id-type="doi">10.1186/1471-2105-15-147</pub-id>.<pub-id pub-id-type="pmcid">PMC4064103</pub-id><pub-id pub-id-type="pmid">24884954</pub-id></mixed-citation></ref><ref id="ref49"><mixed-citation publication-type="journal" id="cit49"><name name-style="western"><surname>Hajdin</surname><given-names>C. E.</given-names></name>; <name name-style="western"><surname>Bellaousov</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Huggins</surname><given-names>W.</given-names></name>; <name name-style="western"><surname>Leonard</surname><given-names>C. W.</given-names></name>; <name name-style="western"><surname>Mathews</surname><given-names>D. H.</given-names></name>; <name name-style="western"><surname>Weeks</surname><given-names>K. M.</given-names></name>
<article-title>Accurate SHAPE-directed RNA secondary
structure modeling,
including pseudoknots</article-title>. <source>Proc. Natl. Acad. Sci.
U.S.A.</source>
<year>2013</year>, <volume>110</volume>, <fpage>5498</fpage>–<lpage>5503</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1219988110</pub-id>.<pub-id pub-id-type="pmid">23503844</pub-id>
<pub-id pub-id-type="pmcid">PMC3619282</pub-id></mixed-citation></ref><ref id="ref50"><mixed-citation publication-type="journal" id="cit50"><name name-style="western"><surname>Will</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Otto</surname><given-names>C.</given-names></name>; <name name-style="western"><surname>Miladi</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Möhl</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Backofen</surname><given-names>R.</given-names></name>
<article-title>SPARSE: quadratic time simultaneous
alignment and folding
of RNAs without sequence-based heuristics</article-title>. <source>Bioinformatics</source>
<year>2015</year>, <volume>31</volume>, <fpage>2489</fpage>–<lpage>2496</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btv185</pub-id>.<pub-id pub-id-type="pmid">25838465</pub-id>
<pub-id pub-id-type="pmcid">PMC4514930</pub-id></mixed-citation></ref><ref id="ref51"><mixed-citation publication-type="journal" id="cit51"><name name-style="western"><surname>Will</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Joshi</surname><given-names>T.</given-names></name>; <name name-style="western"><surname>Hofacker</surname><given-names>I. L.</given-names></name>; <name name-style="western"><surname>Stadler</surname><given-names>P. F.</given-names></name>; <name name-style="western"><surname>Backofen</surname><given-names>R.</given-names></name>
<article-title>LocARNA-P:
accurate boundary prediction and improved detection of structural
RNAs</article-title>. <source>RNA</source>
<year>2012</year>, <volume>18</volume>, <fpage>900</fpage>–<lpage>914</lpage>. <pub-id pub-id-type="doi">10.1261/rna.029041.111</pub-id>.<pub-id pub-id-type="pmid">22450757</pub-id>
<pub-id pub-id-type="pmcid">PMC3334699</pub-id></mixed-citation></ref><ref id="ref52"><mixed-citation publication-type="journal" id="cit52"><name name-style="western"><surname>Ji</surname><given-names>Y.</given-names></name>; <name name-style="western"><surname>Xu</surname><given-names>X.</given-names></name>; <name name-style="western"><surname>Stormo</surname><given-names>G. D.</given-names></name>
<article-title>A graph
theoretical approach for predicting common
RNA secondary structure motifs including pseudoknots in unaligned
sequences</article-title>. <source>Bioinformatics</source>
<year>2004</year>, <volume>20</volume>, <fpage>1591</fpage>–<lpage>1602</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bth131</pub-id>.<pub-id pub-id-type="pmid">14962926</pub-id>
</mixed-citation></ref><ref id="ref53"><mixed-citation publication-type="journal" id="cit53"><name name-style="western"><surname>Tabei</surname><given-names>Y.</given-names></name>; <name name-style="western"><surname>Kiryu</surname><given-names>H.</given-names></name>; <name name-style="western"><surname>Kin</surname><given-names>T.</given-names></name>; <name name-style="western"><surname>Asai</surname><given-names>K.</given-names></name>
<article-title>A fast structural
multiple
alignment method for long RNA sequences</article-title>. <source>BMC
Bioinf.</source>
<year>2008</year>, <volume>9</volume> (<issue>1</issue>), <fpage>33</fpage><pub-id pub-id-type="doi">10.1186/1471-2105-9-33</pub-id>.<pub-id pub-id-type="pmcid">PMC2375124</pub-id><pub-id pub-id-type="pmid">18215258</pub-id></mixed-citation></ref><ref id="ref54"><mixed-citation publication-type="journal" id="cit54"><name name-style="western"><surname>Sweeney</surname><given-names>B. A.</given-names></name>; <name name-style="western"><surname>Petrov</surname><given-names>A. I.</given-names></name>; <name name-style="western"><surname>Burkov</surname><given-names>B.</given-names></name>; <name name-style="western"><surname>Finn</surname><given-names>R. D.</given-names></name>; <name name-style="western"><surname>Bateman</surname><given-names>A.</given-names></name>; <name name-style="western"><surname>Szymanski</surname><given-names>M.</given-names></name>; <name name-style="western"><surname>Karlowski</surname><given-names>W. M.</given-names></name>; <name name-style="western"><surname>Gorodkin</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Seemann</surname><given-names>S. E.</given-names></name>; <name name-style="western"><surname>Cannone</surname><given-names>J. J.</given-names></name>; et al. <article-title>RNAcentral: a hub of information for non-coding
RNA sequences</article-title>. <source>Nucleic Acids Res.</source>
<year>2019</year>, <volume>47</volume>, <fpage>D221</fpage>–<lpage>D229</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky1034</pub-id>.<pub-id pub-id-type="pmid">30395267</pub-id>
<pub-id pub-id-type="pmcid">PMC6324050</pub-id></mixed-citation></ref><ref id="ref55"><mixed-citation publication-type="journal" id="cit55"><name name-style="western"><surname>Singh</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Paliwal</surname><given-names>K.</given-names></name>; <name name-style="western"><surname>Zhang</surname><given-names>T.</given-names></name>; <name name-style="western"><surname>Singh</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Litfin</surname><given-names>T.</given-names></name>; <name name-style="western"><surname>Zhou</surname><given-names>Y.</given-names></name>
<article-title>Improved RNA secondary structure
and tertiary base-pairing prediction
using evolutionary profile, mutational coupling and two-dimensional
transfer learning</article-title>. <source>Bioinformatics</source>
<year>2021</year>, <volume>37</volume>, <fpage>2589</fpage>–<lpage>2600</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btab165</pub-id>.<pub-id pub-id-type="pmid">33704363</pub-id>
</mixed-citation></ref><ref id="ref56"><mixed-citation publication-type="journal" id="cit56"><name name-style="western"><surname>Yang</surname><given-names>E.</given-names></name>; <name name-style="western"><surname>Zhang</surname><given-names>H.</given-names></name>; <name name-style="western"><surname>Zang</surname><given-names>Z.</given-names></name>; <name name-style="western"><surname>Zhou</surname><given-names>Z.</given-names></name>; <name name-style="western"><surname>Wang</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Liu</surname><given-names>Z.</given-names></name>; <name name-style="western"><surname>Liu</surname><given-names>Y.</given-names></name>
<article-title>GCNfold: A
novel lightweight model with valid extractors
for RNA secondary structure prediction</article-title>. <source>Comput.
Biol. Med.</source>
<year>2023</year>, <volume>164</volume>, <fpage>107246</fpage><pub-id pub-id-type="doi">10.1016/j.compbiomed.2023.107246</pub-id>.<pub-id pub-id-type="pmid">37487383</pub-id>
</mixed-citation></ref><ref id="ref57"><mixed-citation publication-type="journal" id="cit57"><name name-style="western"><surname>Fu</surname><given-names>L.</given-names></name>; <name name-style="western"><surname>Cao</surname><given-names>Y.</given-names></name>; <name name-style="western"><surname>Wu</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Peng</surname><given-names>Q.</given-names></name>; <name name-style="western"><surname>Nie</surname><given-names>Q.</given-names></name>; <name name-style="western"><surname>Xie</surname><given-names>X.</given-names></name>
<article-title>UFold: fast
and accurate RNA secondary structure prediction with deep learning</article-title>. <source>Nucleic Acids Res.</source>
<year>2022</year>, <volume>50</volume>, <elocation-id>e14</elocation-id><pub-id pub-id-type="doi">10.1093/nar/gkab1074</pub-id>.<pub-id pub-id-type="pmid">34792173</pub-id>
<pub-id pub-id-type="pmcid">PMC8860580</pub-id></mixed-citation></ref><ref id="ref58"><mixed-citation publication-type="journal" id="cit58"><name name-style="western"><surname>Chen</surname><given-names>C.-C.</given-names></name>; <name name-style="western"><surname>Chan</surname><given-names>Y.-M.</given-names></name>
<article-title>REDfold: accurate
RNA secondary structure prediction
using residual encoder-decoder network</article-title>. <source>BMC
Bioinf.</source>
<year>2023</year>, <volume>24</volume> (<issue>1</issue>), <fpage>122</fpage><pub-id pub-id-type="doi">10.1186/s12859-023-05238-8</pub-id>.<pub-id pub-id-type="pmcid">PMC10044938</pub-id><pub-id pub-id-type="pmid">36977986</pub-id></mixed-citation></ref><ref id="ref59"><mixed-citation publication-type="journal" id="cit59"><name name-style="western"><surname>Chen</surname><given-names>X.</given-names></name>; <name name-style="western"><surname>Li</surname><given-names>Y.</given-names></name>; <name name-style="western"><surname>Umarov</surname><given-names>R.</given-names></name>; <name name-style="western"><surname>Gao</surname><given-names>X.</given-names></name>; <name name-style="western"><surname>Song</surname><given-names>L.</given-names></name>
<article-title>RNA secondary structure
prediction by learning unrolled algorithms</article-title>. <source>arXiv</source>
<year>2020</year>, <fpage>2002.05810</fpage><pub-id pub-id-type="doi">10.48550/arXiv.2002.05810</pub-id>.</mixed-citation></ref><ref id="ref60"><mixed-citation publication-type="journal" id="cit60"><name name-style="western"><surname>Ke</surname><given-names>Y.</given-names></name>; <name name-style="western"><surname>Rao</surname><given-names>J.</given-names></name>; <name name-style="western"><surname>Zhao</surname><given-names>H.</given-names></name>; <name name-style="western"><surname>Lu</surname><given-names>Y.</given-names></name>; <name name-style="western"><surname>Xiao</surname><given-names>N.</given-names></name>; <name name-style="western"><surname>Yang</surname><given-names>Y.</given-names></name>
<article-title>Accurate prediction
of genome-wide RNA secondary structure profile based on extreme gradient
boosting</article-title>. <source>Bioinformatics</source>
<year>2020</year>, <volume>36</volume>, <fpage>4576</fpage>–<lpage>4582</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa534</pub-id>.<pub-id pub-id-type="pmid">32467966</pub-id>
</mixed-citation></ref><ref id="ref61"><mixed-citation publication-type="journal" id="cit61"><name name-style="western"><surname>Lemieux</surname><given-names>S.</given-names></name>; <name name-style="western"><surname>Major</surname><given-names>F.</given-names></name>
<article-title>RNA canonical and non-canonical
base pairing types: a recognition
method and complete repertoire</article-title>. <source>Nucleic Acids
Res.</source>
<year>2002</year>, <volume>30</volume>, <fpage>4250</fpage>–<lpage>4263</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkf540</pub-id>.<pub-id pub-id-type="pmid">12364604</pub-id>
<pub-id pub-id-type="pmcid">PMC140540</pub-id></mixed-citation></ref></ref-list></back></article>