<?xml version="1.0" encoding="UTF-8"?><article xml:lang="en" article-type="research-article"><front><journal-meta><journal-id journal-id-type="pmc-domain-id">32</journal-id><journal-id journal-id-type="pmc-domain">bmcgeno</journal-id><journal-title-group><journal-title>BMC Genomics</journal-title><abbrev-journal-title>BMC Genomics</abbrev-journal-title></journal-title-group><publisher><publisher-name>BMC</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="pmcid">PMC13425921</article-id><article-id pub-id-type="pmcaid">13425921</article-id><article-id pub-id-type="pmcaiid">13425921</article-id><article-id pub-id-type="pmid">42215857</article-id><article-id pub-id-type="doi">10.1186/s12864-026-12976-5</article-id><title-group><article-title>PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space</article-title></title-group><contrib-group content-type="author"><contrib><name name-style="western"><surname>Gibson</surname><given-names initials="K">Kaetlyn</given-names></name><xref ref-type="aff" rid="Aff1">1</xref></contrib><contrib><name name-style="western"><surname>Li</surname><given-names initials="PE">Po-E</given-names></name><xref ref-type="aff" rid="Aff1">1</xref></contrib><contrib><name name-style="western"><surname>Li</surname><given-names initials="V">Valerie</given-names></name><xref ref-type="aff" rid="Aff1">1</xref></contrib><contrib><name name-style="western"><surname>Dix</surname><given-names initials="M">Martha</given-names></name><xref ref-type="aff" rid="Aff1">1</xref></contrib><contrib><name name-style="western"><surname>Hung</surname><given-names initials="LW">Li-Wei</given-names></name><xref ref-type="aff" rid="Aff2">2</xref></contrib><contrib><name name-style="western"><surname>Stelle</surname><given-names initials="GW">George Widgery</given-names></name><xref ref-type="aff" rid="Aff3">3</xref></contrib><contrib><name name-style="western"><surname>Babinski</surname><given-names initials="M">Michal</given-names></name><xref ref-type="aff" rid="Aff1">1</xref></contrib><contrib><name name-style="western"><surname>Chain</surname><given-names initials="P">Patrick</given-names></name><xref ref-type="aff" rid="Aff1">1</xref></contrib><contrib><name name-style="western"><surname>Hu</surname><given-names initials="B">Bin</given-names></name><xref ref-type="aff" rid="Aff1">1</xref><xref ref-type="author-notes" rid="_fncrsp93pmc__">✉</xref></contrib></contrib-group><aff id="Aff1"><label>1</label>Genomics and Bioanalytics Group, Bioscience Division, Los Alamos National Laboratory, Los Alamos, NM 87545 USA </aff><aff id="Aff2"><label>2</label>Biochemistry and Biotechnology Group, Bioscience Division, Los Alamos National Laboratory, Los Alamos, NM 87545 USA </aff><aff id="Aff3"><label>3</label>Applied Computer Science Group, Computer, Computational and Statistical Sciences Division, Los Alamos National Laboratory, Los Alamos, NM 87545 USA </aff><author-notes><fn id="_fncrsp93pmc__"><label>✉</label><p>Corresponding author.</p></fn></author-notes><pub-date><day>30</day><month>5</month><year>2026</year></pub-date><volume>27</volume><fpage>645</fpage><page-range>645</page-range><pub-history><event event-type="pmc-release"><date><day>1</day><month>8</month><year>2026</year></date></event></pub-history><permissions><copyright-statement>© This is a U.S. Government work and not under copyright protection in the US; foreign copyright protection may apply 2026</copyright-statement><license><license-p><bold>Open Access</bold> This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://creativecommons.org/licenses/by/4.0/" ext-link-type="uri">http://creativecommons.org/licenses/by/4.0/</ext-link>.</license-p></license></permissions><self-uri xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="12864_2026_Article_12976.pdf" content-type="pmc-pdf"><?cloudpmc-path 886e/13425921/0fe0ce12ebed/12864_2026_Article_12976.pdf?><?cloudpmc-bucket app?><?size 2687319?></self-uri><abstract id="Abs1"><title>Abstract</title><sec id="sec1" disp-level="2"><title>Background</title><p id="Par1">Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility.</p></sec><sec id="sec2" disp-level="2"><title>Results</title><p id="Par2">We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R<sup>2</sup> values (&gt; 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R<sup>2</sup> ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R<sup>2</sup> &lt;0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods.</p></sec><sec id="sec3" disp-level="2"><title>Conclusion</title><p id="Par3">PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.</p></sec><sec id="sec4" disp-level="2"><title>Supplementary Information</title><p>The online version contains supplementary material available at 10.1186/s12864-026-12976-5.</p></sec><sec id="kwd-group1" xml:lang="en" sec-type="kwd-group" disp-level="2"><p><bold>Keywords:</bold> Protein language models, Biosurveillance, Phenotype prediction, Gomology leakage, Position-stratified validation, Host tropism prediction, Model fine-tuning, Machine learning efficiency</p></sec></abstract><custom-meta-group><custom-meta><meta-name>status</meta-name><meta-value>released</meta-value></custom-meta><custom-meta><meta-name>display-pdf</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>is-olf</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-manuscript</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-preprint</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-journal-matter</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-scanned</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-retracted</meta-name><meta-value>no</meta-value></custom-meta></custom-meta-group></article-meta><notes notes-type="article-notes"><sec id="historyarticle-meta1" sec-type="history" disp-level="2"><p>Received 2026 Mar 12; Accepted 2026 May 18; Collection date 2026.</p></sec></notes></front><body><sec id="Sec1" disp-level="1"><title>Background</title><p id="Par4">Pathogen evolution continuously generates variants with altered phenotypic properties that fundamentally shift transmission dynamics [<xref rid="CR1" ref-type="bibr">1</xref>–<xref rid="CR5" ref-type="bibr">5</xref>], immune interactions [<xref rid="CR6" ref-type="bibr">6</xref>–<xref rid="CR9" ref-type="bibr">9</xref>], and host range [<xref rid="CR10" ref-type="bibr">10</xref>–<xref rid="CR12" ref-type="bibr">12</xref>]. Current computational biosurveillance strategies, while effective for phylogenetic tracking, face significant limitations in predicting how specific genetic changes translate into the phenotypic consequences that drive outbreak dynamics. Traditional multiple sequence alignment (MSA) and phylogenetic methods [<xref rid="CR13" ref-type="bibr">13</xref>–<xref rid="CR16" ref-type="bibr">16</xref>] are inherently reactive; they often fail to capture genotype-to-phenotype relationships between pathogens that share functional properties but are phylogenetically distant [<xref rid="CR4" ref-type="bibr">4</xref>, <xref rid="CR5" ref-type="bibr">5</xref>, <xref rid="CR17" ref-type="bibr">17</xref>–<xref rid="CR19" ref-type="bibr">19</xref>]. Furthermore, these approaches scale poorly [<xref rid="CR20" ref-type="bibr">20</xref>] when applied to the massive, diverse sequence datasets generated during real-time pandemic monitoring and may introduce artifacts when analyzing hypervariable viral genomes under selective pressure [<xref rid="CR21" ref-type="bibr">21</xref>]. Recent computational efforts have begun to address this gap through proactive approaches, including machine learning models for forecasting future dominant SARS-CoV-2 lineages [<xref rid="CR22" ref-type="bibr">22</xref>], anomaly-detection frameworks for flagging emerging variants before they reach epidemiological significance, and generative models that anticipate plausible future viral sequences to expand surveillance coverage beyond observed diversity [<xref rid="CR23" ref-type="bibr">23</xref>]. While these approaches represent important advances, they have not systematically addressed the evaluation bias introduced by homology leakage, nor have they benchmarked the conditions under which PLM-derived representations genuinely generalize to novel mutational sites.</p><p id="Par5">The receptor-binding domains (RBDs) [<xref rid="CR24" ref-type="bibr">24</xref>–<xref rid="CR27" ref-type="bibr">27</xref>] of betacoronavirus spike proteins illustrate these challenges. The SARS-CoV-2 RBD binds the human ACE2 receptor [<xref rid="CR28" ref-type="bibr">28</xref>] while the related MERS-CoV binds to dipeptidyl peptidase (DPP4) [<xref rid="CR29" ref-type="bibr">29</xref>, <xref rid="CR30" ref-type="bibr">30</xref>], illustrating how discrete sequence variations in homologous domains dictate host specificity. The evolutionary adaptability of the RBD is further underscored by the emergence of the Omicron (B.1.1.529) lineage, where pivotal mutations within the RBD enabled widespread escape by significantly attenuating antibody neutralization efficacy and compromising vaccine-induced immunity [<xref rid="CR31" ref-type="bibr">31</xref>–<xref rid="CR34" ref-type="bibr">34</xref>].</p><p id="Par6">Deep mutational scanning (DMS) has emerged as a powerful experimental tool to systematically survey how random mutations across a sequence impacts protein function, providing rich datasets that capture complex, non-linear molecular interactions [<xref rid="CR35" ref-type="bibr">35</xref>–<xref rid="CR38" ref-type="bibr">38</xref>]. However, extracting actionable patterns from these high-dimensional, combinatorial datasets requires computational frameworks capable of modeling intricate biophysical constraints.</p><p id="Par7">Protein language models (PLMs) such as the Evolutionary Scale Modeling (ESM) family [<xref rid="CR39" ref-type="bibr">39</xref>–<xref rid="CR44" ref-type="bibr">44</xref>], have shown immense promise in capturing sequence-function relationships by learning from millions of protein sequences across the tree of life [<xref rid="CR45" ref-type="bibr">45</xref>]. Despite their success in general protein modeling and structure prediction, direct application to pathogen-specific sequences often yields limited predictive power. The subtle sequence variations that distinguish viral lineages or confer immune escape represent only a small fraction of the sequence space covered by general foundation models. Moreover, many pathogen sequences are intentionally excluded from the large-scale training sets due to biosecurity concerns, creating a critical resolution gap in legitimate biosurveillance applications.</p><p id="Par8">A primary, yet often overlooked, obstacle in applying PLMs to viral data is the phenomenon of “homology leakage”. Due to the extreme sequence similarity within viral families, standard random training-test splits often allow models to achieve deceptively high accuracy by memorizing site-specific fitness effects rather than learning generalizable biophysical rules. This results in performance metrics that fail to translate into predictive utility for novel, emerging variants. These concerns are corroborated by recent evidence demonstrating that pretrained PLMs can inflate performance scores in paired biological tasks through the memorization of training sequences [<xref rid="CR46" ref-type="bibr">46</xref>]. Critically, such models have been shown to fail in generalizing to the effect of point mutations on binding affinities, highlighting a fundamental gap in current evaluation protocols. To address this directly, we introduce a position-stratified splitting strategy that partitions mutation sites — rather than sequences — into disjoint train and test sets, ensuring the model is never evaluated on positions encountered during training (Fig. <xref rid="Fig1" ref-type="fig">1</xref>). Unlike random splits, where the same mutated positions routinely appear in both partitions and artificially inflate performance estimates, position-stratified splits provide a more stringent test of generalization to unseen mutational sites.</p><fig id="Fig1" position="float"><?disp-level 2?><label>Fig. 1</label><caption><p>Overview of random and position-stratified splitting strategies for SARS-CoV-2 DMS data. <bold>A</bold> Example sequence variants derived from a shortened wild-type (WT) SARS-CoV-2 RBD sequence. Each variant carries one or more amino acid substitutions (circled) relative to the WT; blue circles indicate positions assigned to the train-position set and orange circles indicate positions assigned to the test-position set. <bold>B</bold> In a random split, variants are assigned to train or test partitions at the sequence level without consideration of mutation positions. Consequently, the same mutated positions appear in both partitions — for example, positions 2, 5, and 7 are shared across train and test — creating positional information leakage that inflates apparent model performance. <bold>C</bold> In a position-stratified split, all mutation positions across the dataset are first identified, then partitioned into disjoint train-position and test-position sets using KFold indexing (illustrated here with one fold for clarity; five folds are used in Table 1, Supplementary Table 6, and Supplementary Figure 10). Variants are then assigned based on the positions of all mutations they carry: variants mutated exclusively at train positions are assigned to train, variants mutated exclusively at test positions are assigned to test, and variants carrying mutations from both position sets are excluded to prevent positional overlap. This design ensures that the model is evaluated only on mutation sites never seen during training, providing a more realistic estimate of generalization to novel mutational positions</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d33e379" xlink:href="12864_2026_12976_Fig1_HTML.jpg"><?cloudpmc-path blobs/886e/13425921/0bc37cdceb6f/12864_2026_12976_Fig1_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1251?><?original-width 1985?><?scaled-height 1251?><?scaled-width 1985?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="12864_2026_12976_Fig1_HTML.gif"><?cloudpmc-path blobs/886e/13425921/dfaed9bda1a7/12864_2026_12976_Fig1_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par10">Here, we present <bold>PRIME</bold> (Protein Representation Inference for Mutation Evaluation), a framework that attempts to address these limitations by combining domain-specific fine-tuning of state-of-the-art PLMs with a rigorous position-stratified validation protocol. Our approach demonstrates three key advances: (1) we show that fine-tuning is essential for models to predict phenotypes at previously “unseen” mutational sites; (2) we demonstrate that neural networks built upon these specialized embeddings achieve meaningful predictive power on stratified benchmarks where base foundation models fail; and (3) we show that embedding-based functional clustering can identify host tropism patterns and provide a complementary perspective on functional similarity among bat coronavirus sequences relative to traditional phylogenetic approaches. PRIME complements emerging proactive surveillance methods by establishing a rigorous evaluation standard — position-stratified validation — that can serve as a benchmark for any PLM-based framework operating in high-homology viral sequence space.</p></sec><sec id="Sec2" disp-level="1"><title>Results</title><sec id="Sec3" disp-level="2"><title>Benchmark data collection for viral phenotype prediction</title><p id="Par11">We collected and curated three complementary sequencing datasets, “outbreak”, “DMS” and “betaCov”, to comprehensively analyze RBD sequences. The <bold>outbreak dataset</bold> comprises 347,432 unique RBD sequences curated from GISAID [<xref rid="CR47" ref-type="bibr">47</xref>] and the Sequence Read Archive (SRA) [<xref rid="CR48" ref-type="bibr">48</xref>] as of February 3rd, 2023, with all duplicates and sequences containing ambiguous amino acids removed. This dataset represents viral sequences under real-time selection pressure during the SARS-CoV-2 pandemic, spanning all major SARS-CoV-2 lineages, with Omicron (52.70%), Delta (35.71%), and Alpha (7.20%) variants accounting for 95.61% of the sequences (Fig. <xref rid="Fig2" ref-type="fig">2</xref>A). Amino acid distribution analysis revealed representation of all 20 canonical amino acids, though with notable variability—Leucine (9.48%), Valine (9.04%), and Serine (7.12%) were most abundant, while Methionine appeared in only 0.01% of positions and fewer than 1% of all sequences (Fig. <xref rid="Fig2" ref-type="fig">2</xref>B). Sequence length distribution was highly uniform, with 99.75% of the sequences consisting of 223 amino acids.</p><fig id="Fig2" position="float"><?disp-level 3?><label>Fig. 2</label><caption><p>SARS-CoV-2 RBD sequence analysis and language model performance comparison. <bold>A</bold> Distribution of viral lineages in the outbreak dataset (n = 347,432 sequences), showing predominance of Omicron (52.70%) and Delta (35.71%) variants. <bold>B</bold> Amino acid frequency distribution within RBD sequences, highlighting variability in representation from most common (Asparagine, 9.48%) to rare occurrences (Methionine, 0.01%). <bold>C</bold> Confusion matrix of the non-fine-tuned ESM-2 8M model amino acid prediction accuracy. Diagonal values represent correct prediction rates for each amino acid, with overall accuracy of 5.06%. <bold>D</bold> Confusion matrix for the PRIME fine-tuned ESM-RBD (ESM-2 8M) model after 100 epochs, demonstrating a 19-fold improvement in sequence understanding with an overall masked language modeling accuracy of 96.33%</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d33e428" xlink:href="12864_2026_12976_Fig2_HTML.jpg"><?cloudpmc-path blobs/886e/13425921/ba84f1ecc8f9/12864_2026_12976_Fig2_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1328?><?original-width 1985?><?scaled-height 1328?><?scaled-width 1985?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="12864_2026_12976_Fig2_HTML.gif"><?cloudpmc-path blobs/886e/13425921/3e6962e64965/12864_2026_12976_Fig2_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par13">For evaluating model performance in predicting functional properties, we utilized the <bold>DMS dataset</bold> containing 116,257 unique RBD sequences with measured ACE2 binding affinities and 105,525 unique RBD sequences with quantified expression levels from yeast in vitro experiments [<xref rid="CR36" ref-type="bibr">36</xref>].</p><p id="Par14">To assess cross-species coronavirus predictions, we assembled the <bold>betaCov dataset</bold> containing 75 RBD sequences from diverse Sarbecovirus members, including SARS-CoV-2 (Wuhan-HU-1 strain [<xref rid="CR49" ref-type="bibr">49</xref>], RaTG13v (a bat coronavirus reportedly closely related to SARS-CoV-2 [<xref rid="CR50" ref-type="bibr">50</xref>]), SARS-CoV-1 [<xref rid="CR51" ref-type="bibr">51</xref>], MERS-CoV [<xref rid="CR52" ref-type="bibr">52</xref>], Pangolin-CoV [<xref rid="CR53" ref-type="bibr">53</xref>], and various bat coronaviruses used in a recent survey of sarbecovirus ACE2 binding [<xref rid="CR36" ref-type="bibr">36</xref>]. This dataset represents RBD sequences targeting different mammalian hosts, including bat, pangolin, civet, camel, and human (<bold>Supplementary Table 1</bold>), with some SARS-CoV-1 and MERS-CoV sequences demonstrating cross-species receptor binding.</p></sec><sec id="Sec4" disp-level="2"><title>Scaling and fine-tuning overcome the generalization gap</title><p id="Par15">To evaluate the predictive limits of Protein Language Models (PLMs) in the context of high sequence homology, we benchmarked five PLM architectures spanning two architectures: ESM-2 [<xref rid="CR40" ref-type="bibr">40</xref>] (8 M, 150 M, 650 M) and ESM-C [<xref rid="CR52" ref-type="bibr">52</xref>] (300 M, 600 M), under both random and position-stratified regimes. To address “homology leakage”, we implemented <bold>position-stratified cross-validation</bold>, where all mutations at specific residue positions are held out from the training set to measure genuine biological generalization.</p><p id="Par16">We first assessed whether domain-specific fine-tuning could improve the model’s representation of the RBD sequence space. The original ESM-2 8 M model, without fine-tuning, achieved only 5.06% accuracy in masked language modeling (MLM) tasks. After 100 epochs of fine-tuning, the resulting <bold>ESM-RBD</bold> model achieved 96.33% MLM accuracy—a 19-fold improvement (Fig. <xref rid="Fig2" ref-type="fig">2</xref>C and D).</p><p id="Par17">We compared this to an in-house BERT [<xref rid="CR54" ref-type="bibr">54</xref>]-based model (<bold>BERT-RBD</bold>) trained from scratch on only RBD sequences. While BERT-RBD reached 94.09% MLM accuracy (<bold>Supplementary Fig. 1A</bold>,<bold> 1B</bold>), it remained inferior to the fine-tuned ESM models in capturing rare biochemical signatures. Most notably, ESM-RBD predicted Methionine with 33.04% accuracy, whereas BERT-RBD failed to predict it entirely. This highlights a key advantage of the PRIME framework: fine-tuning a pre-trained foundation model leverages broad, evolutionarily informed representations of protein sequence patterns that are difficult to recover when training from scratch on narrow, pathogen-specific datasets.</p><p id="Par18">The necessity of both scale and fine-tuning becomes most apparent when confronting the generalization gap in phenotype prediction (Table <xref rid="Tab1" ref-type="table">1</xref>). Our results demonstrate that performance metrics from random splits are highly inflated due to data leakage. For the <bold>ESM-C 600M</bold> model, using mean embedding across all amino acid residues in the RBD, a random split yielded a deceptively high R<sup>2</sup> = 0.7711 for binding affinity. However, under the rigorous position-stratified split, the base foundation model (without domain-specific fine-tuning) failed to provide any predictive signal, resulting in a negative correlation coefficient (R<sup>2</sup>=-0.1050) (Table <xref rid="Tab1" ref-type="table">1</xref>).</p><table-wrap id="Tab1" position="float"><?disp-level 3?><label>Table 1</label><caption><p>Benchmarking PRIME across different model scales and validation regimes for mutated RBD binding and expression</p></caption><table frame="hsides" rules="groups"><thead><tr><th align="left" rowspan="3" colspan="1">Model</th><th align="left" rowspan="3" colspan="1">EmbedMethod</th><th align="left" rowspan="3" colspan="1">Fine-tuned</th><th align="left" colspan="4" rowspan="1">Random Split</th><th align="left" colspan="4" rowspan="1">Position-Stratified Split</th></tr><tr><th align="left" colspan="2" rowspan="1">Binding</th><th align="left" colspan="2" rowspan="1">Expression</th><th align="left" colspan="2" rowspan="1">Binding</th><th align="left" colspan="2" rowspan="1">Expression</th></tr><tr><th align="left" colspan="1" rowspan="1">R²</th><th align="left" colspan="1" rowspan="1">RMSE</th><th align="left" colspan="1" rowspan="1">R²</th><th align="left" colspan="1" rowspan="1">RMSE</th><th align="left" colspan="1" rowspan="1">R²</th><th align="left" colspan="1" rowspan="1">RMSE</th><th align="left" colspan="1" rowspan="1">R²</th><th align="left" colspan="1" rowspan="1">RMSE</th></tr></thead><tbody><tr><td align="left" rowspan="4" colspan="1">ESM-2 8M</td><td align="left" rowspan="2" colspan="1">CLS</td><td align="left" colspan="1" rowspan="1">×</td><td align="left" colspan="1" rowspan="1">0.5876 ± 0.07</td><td align="left" colspan="1" rowspan="1">1.2191 ± 0.10</td><td align="left" colspan="1" rowspan="1">0.5984 ± 0.07</td><td align="left" colspan="1" rowspan="1">0.6272 ± 0.05</td><td align="left" colspan="1" rowspan="1">-0.2018 ± 0.04</td><td align="left" colspan="1" rowspan="1">1.9452 ± 0.04</td><td align="left" colspan="1" rowspan="1">-0.0324 ± 0.05</td><td align="left" colspan="1" rowspan="1">1.0916 ± 0.02</td></tr><tr><td align="left" colspan="1" rowspan="1">✓</td><td align="left" colspan="1" rowspan="1">0.5897 ± 0.03</td><td align="left" colspan="1" rowspan="1">1.2194 ± 0.04</td><td align="left" colspan="1" rowspan="1">0.5302 ± 0.03</td><td align="left" colspan="1" rowspan="1">0.6808 ± 0.02</td><td align="left" colspan="1" rowspan="1">-0.2294 ± 0.03</td><td align="left" colspan="1" rowspan="1">1.9653 ± 0.03</td><td align="left" colspan="1" rowspan="1">-0.1928 ± 0.02</td><td align="left" colspan="1" rowspan="1">1.1756 ± 0.01</td></tr><tr><td align="left" rowspan="2" colspan="1">Mean</td><td align="left" colspan="1" rowspan="1">×</td><td align="left" colspan="1" rowspan="1">0.6794 ± 0.02</td><td align="left" colspan="1" rowspan="1">1.0777 ± 0.04</td><td align="left" colspan="1" rowspan="1">0.6576 ± 0.04</td><td align="left" colspan="1" rowspan="1">0.5807 ± 0.03</td><td align="left" colspan="1" rowspan="1">0.0248 ± 0.01</td><td align="left" colspan="1" rowspan="1">1.7519 ± 0.01</td><td align="left" colspan="1" rowspan="1">0.0967 ± 0.02</td><td align="left" colspan="1" rowspan="1">1.0221 ± 0.01</td></tr><tr><td align="left" colspan="1" rowspan="1">✓</td><td align="left" colspan="1" rowspan="1">0.6225 ± 0.07</td><td align="left" colspan="1" rowspan="1">1.1659 ± 0.10</td><td align="left" colspan="1" rowspan="1">0.5545 ± 0.07</td><td align="left" colspan="1" rowspan="1">0.6613 ± 0.05</td><td align="left" colspan="1" rowspan="1">0.0071 ± 0.05</td><td align="left" colspan="1" rowspan="1">1.7688 ± 0.06</td><td align="left" colspan="1" rowspan="1">0.1013 ± 0.06</td><td align="left" colspan="1" rowspan="1">1.0213 ± 0.03</td></tr><tr><td align="left" rowspan="4" colspan="1">ESM-2 150M</td><td align="left" rowspan="2" colspan="1">CLS</td><td align="left" colspan="1" rowspan="1">×</td><td align="left" colspan="1" rowspan="1">0.5788 ± 0.14</td><td align="left" colspan="1" rowspan="1">1.2167 ± 0.22</td><td align="left" colspan="1" rowspan="1">0.5409 ± 0.08</td><td align="left" colspan="1" rowspan="1">0.6709 ± 0.06</td><td align="left" colspan="1" rowspan="1">-0.2449 ± 0.06</td><td align="left" colspan="1" rowspan="1">1.9777 ± 0.03</td><td align="left" colspan="1" rowspan="1">-0.0754 ± 0.04</td><td align="left" colspan="1" rowspan="1">1.1169 ± 0.02</td></tr><tr><td align="left" colspan="1" rowspan="1">✓</td><td align="left" colspan="1" rowspan="1">0.5836 ± 0.28</td><td align="left" colspan="1" rowspan="1">1.1639 ± 0.40</td><td align="left" colspan="1" rowspan="1">0.3354 ± 0.45</td><td align="left" colspan="1" rowspan="1">0.7664 ± 0.26</td><td align="left" colspan="1" rowspan="1">-0.2675 ± 0.07</td><td align="left" colspan="1" rowspan="1">2.0022 ± 0.04</td><td align="left" colspan="1" rowspan="1">-0.3805 ± 0.04</td><td align="left" colspan="1" rowspan="1">1.2480 ± 0.02</td></tr><tr><td align="left" rowspan="2" colspan="1">Mean</td><td align="left" colspan="1" rowspan="1">×</td><td align="left" colspan="1" rowspan="1">0.6713 ± 0.13</td><td align="left" colspan="1" rowspan="1">1.0693 ± 0.22</td><td align="left" colspan="1" rowspan="1">0.6131 ± 0.07</td><td align="left" colspan="1" rowspan="1">0.6150 ± 0.06</td><td align="left" colspan="1" rowspan="1">-0.0695 ± 0.03</td><td align="left" colspan="1" rowspan="1">1.8359 ± 0.02</td><td align="left" colspan="1" rowspan="1">0.0195 ± 0.03</td><td align="left" colspan="1" rowspan="1">1.0649 ± 0.02</td></tr><tr><td align="left" colspan="1" rowspan="1">✓</td><td align="left" colspan="1" rowspan="1">0.7461 ± 0.08</td><td align="left" colspan="1" rowspan="1">0.9480 ± 0.15</td><td align="left" colspan="1" rowspan="1">0.6693 ± 0.04</td><td align="left" colspan="1" rowspan="1">0.5704 ± 0.03</td><td align="left" colspan="1" rowspan="1">-0.0080 ± 0.00</td><td align="left" colspan="1" rowspan="1">1.7863 ± 0.01</td><td align="left" colspan="1" rowspan="1">0.1039 ± 0.04</td><td align="left" colspan="1" rowspan="1">1.0165 ± 0.02</td></tr><tr><td align="left" rowspan="4" colspan="1">ESM-2 650M</td><td align="left" rowspan="2" colspan="1">CLS</td><td align="left" colspan="1" rowspan="1">×</td><td align="left" colspan="1" rowspan="1">0.7694 ± 0.13</td><td align="left" colspan="1" rowspan="1">0.8757 ± 0.27</td><td align="left" colspan="1" rowspan="1">0.7139 ± 0.13</td><td align="left" colspan="1" rowspan="1">0.5171 ± 0.12</td><td align="left" colspan="1" rowspan="1">-0.1591 ± 0.04</td><td align="left" colspan="1" rowspan="1">1.9063 ± 0.02</td><td align="left" colspan="1" rowspan="1">-0.0004 ± 0.03</td><td align="left" colspan="1" rowspan="1">1.0781 ± 0.01</td></tr><tr><td align="left" colspan="1" rowspan="1">✓</td><td align="left" colspan="1" rowspan="1">0.7584 ± 0.11</td><td align="left" colspan="1" rowspan="1">0.9104 ± 0.22</td><td align="left" colspan="1" rowspan="1">0.6758 ± 0.10</td><td align="left" colspan="1" rowspan="1">0.5582 ± 0.09</td><td align="left" colspan="1" rowspan="1">-0.0399 ± 0.06</td><td align="left" colspan="1" rowspan="1">1.8084 ± 0.06</td><td align="left" colspan="1" rowspan="1">-0.0175 ± 0.06</td><td align="left" colspan="1" rowspan="1">1.0869 ± 0.03</td></tr><tr><td align="left" rowspan="2" colspan="1">Mean</td><td align="left" colspan="1" rowspan="1">×</td><td align="left" colspan="1" rowspan="1">0.7981 ± 0.13</td><td align="left" colspan="1" rowspan="1">0.8091 ± 0.28</td><td align="left" colspan="1" rowspan="1">0.7337 ± 0.15</td><td align="left" colspan="1" rowspan="1">0.4942 ± 0.14</td><td align="left" colspan="1" rowspan="1">-0.0865 ± 0.04</td><td align="left" colspan="1" rowspan="1">1.8488 ± 0.03</td><td align="left" colspan="1" rowspan="1">-0.0103 ± 0.03</td><td align="left" colspan="1" rowspan="1">1.0807 ± 0.01</td></tr><tr><td align="left" colspan="1" rowspan="1">✓</td><td align="left" colspan="1" rowspan="1">0.8189 ± 0.10</td><td align="left" colspan="1" rowspan="1">0.7798 ± 0.22</td><td align="left" colspan="1" rowspan="1">0.7354 ± 0.11</td><td align="left" colspan="1" rowspan="1">0.4999 ± 0.10</td><td align="left" colspan="1" rowspan="1">0.0588 ± 0.04</td><td align="left" colspan="1" rowspan="1">1.7200 ± 0.03</td><td align="left" colspan="1" rowspan="1">0.0779 ± 0.04</td><td align="left" colspan="1" rowspan="1">1.0336 ± 0.02</td></tr><tr><td align="left" rowspan="4" colspan="1">ESM-C 300M</td><td align="left" rowspan="2" colspan="1">CLS</td><td align="left" colspan="1" rowspan="1">×</td><td align="left" colspan="1" rowspan="1">0.5869 ± 0.15</td><td align="left" colspan="1" rowspan="1">1.2032 ± 0.23</td><td align="left" colspan="1" rowspan="1">0.6026 ± 0.08</td><td align="left" colspan="1" rowspan="1">0.6230 ± 0.06</td><td align="left" colspan="1" rowspan="1">-0.1954 ± 0.06</td><td align="left" colspan="1" rowspan="1">1.9366 ± 0.04</td><td align="left" colspan="1" rowspan="1">-0.0313 ± 0.04</td><td align="left" colspan="1" rowspan="1">1.0924 ± 0.02</td></tr><tr><td align="left" colspan="1" rowspan="1">✓</td><td align="left" colspan="1" rowspan="1">0.7647 ± 0.08</td><td align="left" colspan="1" rowspan="1">0.9104 ± 0.16</td><td align="left" colspan="1" rowspan="1">0.6678 ± 0.07</td><td align="left" colspan="1" rowspan="1">0.5691 ± 0.06</td><td align="left" colspan="1" rowspan="1">0.0019 ± 0.05</td><td align="left" colspan="1" rowspan="1">1.7766 ± 0.04</td><td align="left" colspan="1" rowspan="1">0.0916 ± 0.02</td><td align="left" colspan="1" rowspan="1">1.0265 ± 0.02</td></tr><tr><td align="left" rowspan="2" colspan="1">Mean</td><td align="left" colspan="1" rowspan="1">×</td><td align="left" colspan="1" rowspan="1">0.7155 ± 0.07</td><td align="left" colspan="1" rowspan="1">1.0075 ± 0.13</td><td align="left" colspan="1" rowspan="1">0.6546 ± 0.04</td><td align="left" colspan="1" rowspan="1">0.5830 ± 0.03</td><td align="left" colspan="1" rowspan="1">-0.0162 ± 0.01</td><td align="left" colspan="1" rowspan="1">1.7874 ± 0.02</td><td align="left" colspan="1" rowspan="1">0.0647 ± 0.02</td><td align="left" colspan="1" rowspan="1">1.0410 ± 0.01</td></tr><tr><td align="left" colspan="1" rowspan="1">✓</td><td align="left" colspan="1" rowspan="1">0.8057 ± 0.06</td><td align="left" colspan="1" rowspan="1">0.8278 ± 0.14</td><td align="left" colspan="1" rowspan="1">0.7165 ± 0.06</td><td align="left" colspan="1" rowspan="1">0.5261 ± 0.05</td><td align="left" colspan="1" rowspan="1">0.1949 ± 0.03</td><td align="left" colspan="1" rowspan="1">1.5918 ± 0.03</td><td align="left" colspan="1" rowspan="1">0.2313 ± 0.01</td><td align="left" colspan="1" rowspan="1">0.9449 ± 0.01</td></tr><tr><td align="left" rowspan="4" colspan="1">ESM-C 600M</td><td align="left" rowspan="2" colspan="1">CLS</td><td align="left" colspan="1" rowspan="1">×</td><td align="left" colspan="1" rowspan="1">0.6616 ± 0.12</td><td align="left" colspan="1" rowspan="1">1.0863 ± 0.22</td><td align="left" colspan="1" rowspan="1">0.6030 ± 0.09</td><td align="left" colspan="1" rowspan="1">0.6219 ± 0.07</td><td align="left" colspan="1" rowspan="1">-0.1106 ± 0.02</td><td align="left" colspan="1" rowspan="1">1.8726 ± 0.00</td><td align="left" colspan="1" rowspan="1">-0.0167 ± 0.03</td><td align="left" colspan="1" rowspan="1">1.0823 ± 0.01</td></tr><tr><td align="left" colspan="1" rowspan="1">✓</td><td align="left" colspan="1" rowspan="1">0.7552 ± 0.07</td><td align="left" colspan="1" rowspan="1">0.9308 ± 0.15</td><td align="left" colspan="1" rowspan="1">0.5388 ± 0.21</td><td align="left" colspan="1" rowspan="1">0.6577 ± 0.15</td><td align="left" colspan="1" rowspan="1">0.1165 ± 0.06</td><td align="left" colspan="1" rowspan="1">1.6616 ± 0.05</td><td align="left" colspan="1" rowspan="1">0.1478 ± 0.01</td><td align="left" colspan="1" rowspan="1">0.9922 ± 0.01</td></tr><tr><td align="left" rowspan="2" colspan="1">Mean</td><td align="left" colspan="1" rowspan="1">×</td><td align="left" colspan="1" rowspan="1">0.7355 ± 0.12</td><td align="left" colspan="1" rowspan="1">0.9525 ± 0.23</td><td align="left" colspan="1" rowspan="1">0.6610 ± 0.08</td><td align="left" colspan="1" rowspan="1">0.5738 ± 0.07</td><td align="left" colspan="1" rowspan="1">-0.0973 ± 0.03</td><td align="left" colspan="1" rowspan="1">1.8558 ± 0.02</td><td align="left" colspan="1" rowspan="1">0.0666 ± 0.02</td><td align="left" colspan="1" rowspan="1">1.0398 ± 0.02</td></tr><tr><td align="left" colspan="1" rowspan="1">✓</td><td align="left" colspan="1" rowspan="1">0.8346 ± 0.07</td><td align="left" colspan="1" rowspan="1">0.7560 ± 0.17</td><td align="left" colspan="1" rowspan="1">0.7342 ± 0.06</td><td align="left" colspan="1" rowspan="1">0.5087 ± 0.06</td><td align="left" colspan="1" rowspan="1">0.2120 ± 0.03</td><td align="left" colspan="1" rowspan="1">1.5745 ± 0.02</td><td align="left" colspan="1" rowspan="1">0.2479 ± 0.02</td><td align="left" colspan="1" rowspan="1">0.9334 ± 0.02</td></tr><tr><td align="left" colspan="3" rowspan="1">One-Hot Mean</td><td align="left" colspan="1" rowspan="1">0.1136 ± 0.03</td><td align="left" colspan="1" rowspan="1">1.7930 ± 0.03</td><td align="left" colspan="1" rowspan="1">0.0617 ± 0.02</td><td align="left" colspan="1" rowspan="1">0.9625 ± 0.01</td><td align="left" colspan="1" rowspan="1">-0.2397 ± 0.01</td><td align="left" colspan="1" rowspan="1">1.9754 ± 0.00</td><td align="left" colspan="1" rowspan="1">-0.1774 ± 0.01</td><td align="left" colspan="1" rowspan="1">1.1675 ± 0.00</td></tr></tbody></table><table-wrap-foot><fn id="_fn_p24"><p>Performance is reported as the coefficient of determination (R2) and Root Mean Square Error (RMSE), with ± values representing the standard deviation across random seeds seen in Supplementary Table 6. Models used either mean embeddings, where sequence-level feature was derived by calculating the mean amino acid representation (average of all site-specific embeddings) across the entire sequence domain, or the standard CLS token embedding as the sequence-level representation. If the model is not designated as fine-tuned, the base ESM model is used as the embedder</p></fn></table-wrap-foot></table-wrap><p id="Par20">Domain-specific fine-tuning on the outbreak dataset significantly restored predictive power in the stratified context. For the ESM-C 600 M model, fine-tuning improved stratified binding affinity R<sup>2</sup> from no measurable signal (-0.1050) to a modest predictive value of 0.1823 ± 0.2149 using mean amino acid representation and 0.2087 ± 0.1451 using sequence level CLS embeddings. Similar improvement was found in fine-tuned ESM-C models under a stratified expression split (Table <xref rid="Tab1" ref-type="table">1</xref>). These findings demonstrate that while foundation models are essential starting points, pathogen-specific adaptation is needed to capture the nuanced biophysical constraints necessary for prospective biosurveillance.</p></sec><sec id="Sec5" disp-level="2"><title>PLM scale and fine-tuning resolve viral lineages and diversity</title><p id="Par21">We evaluated whether the learned representations of the fine-tuned ESM-RBD model could enable rapid identification of evolutionary relationships and distinguish major SARS-CoV-2 lineages based solely on RBD sequence embeddings. To address the imbalance in lineage abundance, we randomly downsampled the Omicron (<italic>n</italic> = 160,016) and Delta (<italic>n</italic> = 109,448) sequences to match the number of Alpha lineage sequences (<italic>n</italic> = 22,075 per lineage). Visualization of these sequence-level latent representations using t-SNE produced three distinct clusters corresponding to the Alpha, Delta, and Omicron lineages (Fig. <xref rid="Fig3" ref-type="fig">3</xref>A). HDBSCAN clustering of these representations showed high agreement with established lineage assignments, with Cluster 0 corresponding predominantly to Delta, Cluster 1 to Omicron, and Cluster 2 to Alpha (Fig. <xref rid="Fig3" ref-type="fig">3</xref>B). We validated the robustness of this organization through nine iterations of random downsampling (<bold>Supplementary Fig. 2</bold>) and across a range of t-SNE perplexities, where optimal separation was achieved at a perplexity of 750 (Adjusted Rand Index [ARI] = 0.98) (<bold>Supplementary Fig. 3–5</bold>).</p><fig id="Fig3" position="float"><?disp-level 3?><label>Fig. 3</label><caption><p>Fine-tuned ESM-RBD (ESM-2 8M) embeddings resolve viral lineages and capture temporal sub-population dynamics. <bold>A</bold>, <bold>B</bold>: t-SNE visualization and HDBSCAN clustering of sequence-level latent representations for Alpha (blue), Delta (orange), and Omicron (green) lineages. Fine-tuning enables the model to effectively partition major lineages with an Adjusted Rand Index of 0.98. <bold>C</bold> HDBSCAN clustering applied to t-SNE representations of all available Omicron sequences (n = 160,016) reveals 12 clusters (0-11) with some outliers (-1), demonstrating significant diversity within the Omicron lineage. <bold>D</bold> Temporal dynamics of Omicron clusters from November 2021 to May 2023, showing the emergence and succession of specific sub-lineages (e.g., Cluster 7) during the pandemic</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d33e1061" xlink:href="12864_2026_12976_Fig3_HTML.jpg"><?cloudpmc-path blobs/886e/13425921/04b078881fe1/12864_2026_12976_Fig3_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1324?><?original-width 1985?><?scaled-height 1324?><?scaled-width 1985?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="12864_2026_12976_Fig3_HTML.gif"><?cloudpmc-path blobs/886e/13425921/2913117326b1/12864_2026_12976_Fig3_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par23">To further explore diversity within the rapidly evolving Omicron lineage, we applied our clustering approach to the complete set of Omicron sequences (<italic>n</italic> = 160,016), revealing 12 distinct sub-clusters (Fig. <xref rid="Fig3" ref-type="fig">3</xref>C). Tracking the temporal distribution of these clusters from November 2021 to May 2023 demonstrated clear patterns of variant succession, such as the early dominance of Clusters 2 and 5 followed by the rise of Cluster 7 in May 2022 (Fig. <xref rid="Fig3" ref-type="fig">3</xref>D). These sub-clusters showed significantly higher correspondence with established Pango lineage classifications—achieving &gt; 95% purity for lineages such as BA.1.1* and XBB.1.5—compared to MSA-based approaches (<bold>Supplementary Fig. 6</bold>,<bold> Supplementary Table 2</bold>).</p><p id="Par24">Systematic comparison demonstrated that domain-specific fine-tuning is valuable for high-resolution lineage discrimination. While the non-fine-tuned ESM baseline could separate major lineages, it exhibited lower resolution within the Omicron lineage, identifying only 10 clusters compared to the 12 identified by ESM-RBD (<bold>Supplementary Fig. 7</bold>). Furthermore, ESM-RBD outperformed MSA-based clustering (<bold>Supplementary Fig. 8</bold>), which failed to capture fine-grained evolutionary relationships and yielded lower overall clustering metrics (ARI = 0.90 for MSA vs. ARI = 0.98 for ESM-RBD) (<bold>Supplementary Fig. 9</bold>). These results indicate that fine-tuning enables the model to identify functionally and evolutionarily relevant sequence patterns that are obscured by general-purpose models without fine-tuning and traditional MSA-based methods do not provide this type of resolution.</p></sec><sec id="Sec6" disp-level="2"><title>Optimization of neural architectures for real-time deployment</title><p id="Par25">To identify which components drive predictive performance, we conducted a systematic ablation across four neural network architectures: fully connected networks (FCN), graph convolutional networks (GCN) [<xref rid="CR55" ref-type="bibr">55</xref>], bidirectional long short-term memory networks (BLSTM) [<xref rid="CR56" ref-type="bibr">56</xref>], BERT-RBD), two learning rates, and both frozen and fine-tuned encoder conditions, using base ESM and ESM-RBD initializations (<bold>Supplementary Tables 3 and 4</bold>). This design allows us to isolate the contribution of embedding quality from that of the downstream regressor architecture. The FCN architecture with sequence-level (CLS) representations achieved superior predictive performance across both functional properties while maintaining a significantly smaller parameter footprint (<bold>Supplementary Tables 3 and 4</bold>).</p><p id="Par26">We further evaluated the impact of multi-task learning by training a single FCN to simultaneously predict ACE2 binding affinity and protein expression levels using the DMS dataset [<xref rid="CR36" ref-type="bibr">36</xref>]. This multi-task approach reduced total computational requirements by approximately 50% without a significant loss in accuracy compared to single-task models (Fig. <xref rid="Fig4" ref-type="fig">4</xref>A and B). Training dynamics showed stable convergence, with optimal performance for binding and expression prediction achieved at epoch 402 and epoch 155, respectively (Fig. <xref rid="Fig4" ref-type="fig">4</xref>A and B, <bold>Supplementary Tables 3</bold>, and <bold>Supplementary Table 4</bold>).</p><fig id="Fig4" position="float"><?disp-level 3?><label>Fig. 4</label><caption><p>Multi-task architecture optimization and fitting capacity for RBD phenotypes. <bold>A</bold> Training progression for the multi-task ESM-RBD-FCN (ESM-2 8M) model showing RMSE values for binding (blue), expression (orange), and combined loss (green). The model achieved lowest validation RMSE at epoch 402 for binding (RMSE = 0.5059) and epoch 930 for expression (RMSE = 0.3180), with optimal combined performance at epoch 632 (RMSE = 0.6015). <bold>B</bold> Performance comparison of single-task models optimized for binding (blue, optimal at epoch 288, RMSE = 0.5081) and expression (orange, optimal at epoch 155, RMSE = 0.3163), demonstrating that the multi-task approach maintains performance while halving computational overhead. <bold>C</bold>, <bold>D</bold> Correlation between experimentally measured and predicted scores for binding affinity (R2=0.9288) and protein expression (R2=0.8954) under a random-split regime. Insets: Error distributions show that over 95% of predictions fall within 5% of experimental values. Outliers (&gt;10% error, gray) predominantly correspond to residues at functionally critical sites such as N501</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d33e1145" xlink:href="12864_2026_12976_Fig4_HTML.jpg"><?cloudpmc-path blobs/886e/13425921/c44d270dfff7/12864_2026_12976_Fig4_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1329?><?original-width 1985?><?scaled-height 1329?><?scaled-width 1985?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="12864_2026_12976_Fig4_HTML.gif"><?cloudpmc-path blobs/886e/13425921/c6f82094076a/12864_2026_12976_Fig4_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par28">While the correlation coefficients observed in random-split evaluations (R<sup>2</sup> = 0.9288 for binding; R<sup>2</sup> = 0.8954 for expression) are inflated by sequence homology, the error analysis provides critical insight into the model’s high-fidelity representation of known mutational space (Fig. <xref rid="Fig4" ref-type="fig">4</xref>C and D). Over 95% of predictions fell within a 5% error threshold of experimental measurements, with error distributions tightly centered around zero. Systematic analysis of outliers revealed that residues with errors exceeding 10% were predominantly located at functionally critical sites, such as N501, which are known to exhibit complex, non-linear effects on ACE2 binding (<bold>Supplementary Table 5</bold>) [<xref rid="CR57" ref-type="bibr">57</xref>]. This indicates that the FCN architecture effectively captures the majority of the biochemical landscape, but systematic failures pinpoint specific residues where more advanced biophysical modeling is required.</p></sec><sec id="Sec7" disp-level="2"><title>ESM-RBD embeddings identify candidate bat coronavirus sequences for experimental follow-up</title><p id="Par29">To evaluate whether fine-tuned PLMs can detect functional signatures associated with diverse host spillover risk, we analyzed the latent representations of 75 betacoronavirus RBD sequences spanning diverse host species (<bold>Supplementary Table 1</bold>). We first established a baseline using the non-fine-tuned ESM model, where dimensionality reduction and HDBSCAN [<xref rid="CR58" ref-type="bibr">58</xref>] clustering of embeddings exhibited limited resolution and failed to clearly separate groups (Fig. <xref rid="Fig5" ref-type="fig">5</xref>A and B). In contrast, clustering of the fine-tuned ESM-RBD embeddings identified three distinct clusters that more effectively recapitulated established taxonomic and functional relationships (Fig. <xref rid="Fig5" ref-type="fig">5</xref>C and D). Cluster 0 comprised MERS-CoV and hibecovirus sequences, while Cluster 1 contained SARS-CoV-1, SARS-CoV-2, and pangolin coronaviruses. Cluster 2 consisted primarily of bat coronaviruses.</p><fig id="Fig5" position="float"><?disp-level 3?><label>Fig. 5</label><caption><p>Domain-specific fine-tuning resolves host tropism patterns and zoonotic signatures among betacoronaviruses. <bold>A</bold>, <bold>B</bold> Baseline representation: t-SNE visualization and HDBSCAN clustering of 75 betacoronavirus RBD sequences using the non-fine-tuned ESM-2 8M model. Baseline embeddings exhibit limited resolution, failing to clearly delineate functional groups or identify outliers with zoonotic potential. <bold>C</bold>, <bold>D</bold> PRIME representation: t-SNE visualization and HDBSCAN clustering using the fine-tuned ESM-RBD (ESM-2 8M) model. <bold>E</bold> Cross-tabulation of host annotations versus predicted clusters as seen in 5D. PRIME identifies 3.03% of bat coronaviruses that cluster with known human-infective lineages (Cluster 1), suggesting functional similarity to human-infective strains in embedding space; external experimental validation would be required to assess zoonotic risk. <bold>F</bold> Maximum-likelihood phylogeny of the betacoronavirus dataset</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d33e1206" xlink:href="12864_2026_12976_Fig5_HTML.jpg"><?cloudpmc-path blobs/886e/13425921/60cc10aebb8c/12864_2026_12976_Fig5_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1999?><?original-width 1985?><?scaled-height 1999?><?scaled-width 1985?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="12864_2026_12976_Fig5_HTML.gif"><?cloudpmc-path blobs/886e/13425921/934b67590390/12864_2026_12976_Fig5_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par31">Qualitative comparison of the t-SNE projections between the baseline and fine-tuned models suggests improved cluster separation following fine-tuning; we note that this comparison is visualization-supported and exploratory rather than a formal quantitative metric, as t-SNE distances are not directly interpretable (Fig. <xref rid="Fig5" ref-type="fig">5</xref>A and C). This value was calculated by comparing the maximum range of the t-SNE axis components, where the fine-tuned model projected sequence into a broader, more discriminative embedding space. Quantitative analysis of this refined space revealed high clustering consistency with host annotations, achieving 100% accuracy for MERS-CoV, hibecovirus, and pangolin coronavirus sequences (Fig. <xref rid="Fig5" ref-type="fig">5</xref>E). Significantly, 3.03% of sequences annotated as bat coronaviruses clustered with the zoonotic SARS-CoV-1/SARS-CoV-2 group (Cluster 1) rather than the primary bat virus cluster (Fig. <xref rid="Fig5" ref-type="fig">5</xref>E). These sequences cluster in closer proximity to known human-infective strains in embedding space (Fig. <xref rid="Fig5" ref-type="fig">5</xref>D), suggesting functional similarity that warrants further experimental investigation, e.g. ACE2 binding compatibility assays. We note that clustering proximity alone does not establish spillover potential, and external validation would be required to assess true zoonotic risk.</p><p id="Par32">We compared this approach to a traditional maximum likelihood phylogenetic tree (Fig. <xref rid="Fig5" ref-type="fig">5</xref>F). The phylogeny and clustering method corroborated the broad taxonomic groupings— notably placing MERS-CoV in a distinct, distantly related clade. However, PRIME’s embedding space uniquely focuses on the biophysical similarity between human pathogens and select bat coronaviruses. The two methods offer complementary perspectives: PRIME’s embedding space uniquely emphasizes the functional similarity or representation geometry whereas the phylogeny emphasizes evolutionary relationships. These methods can be used in combination for improved biosurveillance.</p><p id="Par33">PRIME may offer a complementary perspective for evaluating functional similarity to human-infective strains: the embedding-based clustering places these specific bat viruses closer to the SARS-CoV-1/SARS-CoV-2 group (Fig. <xref rid="Fig5" ref-type="fig">5</xref>D) than to the primary bat-associated cluster, whereas phylogeny reflects evolutionary proximity without directly capturing this functional distinction. Furthermore, the embedding analysis showed remarkable robustness to sequence divergence. The hibecovirus sequence, which failed standard composition tests during phylogenetic reconstruction, was seamlessly incorporated and accurately clustered by our model (Fig. <xref rid="Fig5" ref-type="fig">5</xref>C). Although these observations are based on a limited dataset (<italic>n</italic> = 75), they suggest that fine-tuned PLM embeddings may capture biologically relevant structural and functional constraints and could provide a complementary perspective for assessing spillover potential. Further validation on larger datasets will be necessary to confirm these trends.</p></sec></sec><sec id="Sec8" disp-level="1"><title>Discussion</title><p id="Par34">The central challenge of applying protein language models to viral pathogens like SARS-CoV-2 is the deceptive nature of sequence homology. As highlighted by our benchmarking of random vs. position-stratified splits, standard validation protocols often lead to “homology leakage,” where models achieve high accuracy (R2 &gt; 0.90) by memorizing site-specific fitness rather than learning the underlying biophysical rules. PRIME addresses this by establishing the importance of <bold>position-stratified validation</bold>.</p><p id="Par35">While previous studies have noted that PLM-based models struggle to generalize to point mutations in protein-protein interactions [<xref rid="CR46" ref-type="bibr">46</xref>], our results suggest that this is not an inherent limitation of the architecture, but rather a consequence of insufficient domain adaptation and improper validation splitting. By implementing position-stratified splits, we demonstrate that domain-specific fine-tuning can indeed recover a meaningful predictive signal (R<sup>2</sup> ≈ 0.23 for <bold>ESM-C 600M</bold>) for binding affinity on entirely unseen sites. Our findings suggest that base models fail to generalize to novel mutational sites (R<sup>2</sup> &lt; 0), whereas domain-specific fine-tuning on large-scale “outbreak” datasets lead to improved predictive performance.</p><p id="Par36">While domain-specific fine-tuning consistently improves stratified performance relative to base models — which show no predictive signal (R² &lt; 0) — the resulting R² values of approximately 0.18–0.23 for binding affinity on unseen mutational sites reflect a meaningful but modest level of position-stratified generalization — which, while not strictly zero-shot given the domain-specific fine-tuning applied, represents a meaningful test of the model’s ability to extrapolate to entirely unseen mutational positions. These results should not be interpreted as a resolution of the generalization challenge, but rather as evidence that fine-tuning on large-scale outbreak data provides a foundation for genuine prospective prediction. True generalization to entirely novel mutational contexts remains an open problem, likely requiring richer structural priors, larger domain-specific corpora, or multi-modal training signals beyond sequence alone. We therefore frame PRIME’s stratified benchmark as a more honest baseline from which future work can measure real progress.</p><p id="Par37">Our systematic evaluation reveals that neither model scale nor pre-training alone is sufficient for pathogen-specific accuracy. General-purpose PLMs, trained on “tree of life” scale data with limited coverage of viral sequence variations, often lack the resolution to distinguish the subtle variations that determine viral host specificity or immune escape. We found that even state-of-the-art architectures like ESM-C 600 M require fine-tuning to move from negative to positive correlation on stratified tasks. This suggests that fine-tuning allows the model to capture nuanced biophysical constraints—such as those at functionally critical sites like N501—that are otherwise invisible to base foundation models. Furthermore, our analysis confirms that mean pooling consistently outperforms CLS-token representations across all five ESM model scales under five-fold cross-validation. This advantage likely reflects the nature of the RBD itself: phenotypic properties such as ACE2 binding affinity are determined by the collective contribution of residues distributed across the domain rather than any single site. Mean pooling aggregates these distributed site-specific signals into a single representation, whereas the CLS token — originally designed to capture global sequence context during pre-training — may not optimally encode this kind of position-averaged biochemical information when repurposed as a downstream regression feature. The robustness of these findings across three random seeds and multiple fold assignments (Supplementary Table 6, Supplementary Fig. 10) further supports the stability of this conclusion.</p><p id="Par38">PRIME’s embedding-based clustering and traditional phylogenetic methods address related but distinct questions: phylogeny aims to reconstruct evolutionary history, whereas PRIME’s embedding space captures functional and biochemical similarity that may not strictly follow neutral genetic drift. These approaches are therefore best viewed as complementary tools for biosurveillance. Both methods successfully separated MERS-CoV and hibecoviruses from SARS-associated viruses, consistent with their divergent receptor usage (DPP4 vs. ACE2). Where phylogeny provides evolutionary context and ancestral relationships, PRIME may offer additional resolution for identifying sequences with convergent functional properties — such as shared receptor binding profiles — regardless of phylogenetic distance. By focusing on features related to host receptor interactions rather than neutral sequence similarity, PRIME can help highlight mutations that may be functionally relevant even when phylogenetic signal is limited or misleading.</p><p id="Par39">The betacoronavirus analysis should be interpreted with appropriate caution. The dataset is small (<italic>n</italic> = 75 RBD sequences) and host labels are coarse, and the boundary separating bat-associated sequences from human-infective clusters in the embedding space is conceptual rather than quantitatively derived. While PRIME identified 3.03% of bat coronavirus sequences clustering near known human-infective strains, this observation reflects functional similarity in embedding space and does not constitute evidence of elevated zoonotic potential. These sequences are best understood as candidates for experimental prioritization — for example, for ACE2 binding assays — rather than confirmed high-risk variants. Future work with larger, more diverse datasets and paired experimental validation would be needed to determine whether embedding-space proximity is genuinely predictive of host-jump capacity.</p><p id="Par40">Notably, PRIME may provide a more informative and functional assessment of zoonotic risk than baseline foundation models. In the non-fine-tuned baseline (Fig. <xref rid="Fig4" ref-type="fig">4</xref>B), several bat coronavirus sequences appeared proximal to the SARS-CoV-1 group. However, following domain-specific fine-tuning, PRIME re-assigned these sequences to the primary bat-associated Cluster 2 (Fig. <xref rid="Fig4" ref-type="fig">4</xref>D). This suggests that PRIME may better differentiate between neutral genetic homology and the specific functional requirements needed for efficient human ACE2 binding.</p><p id="Par41">While no bat viruses were directly assigned to the human-pathogen Cluster 1, PRIME identified 3.03% of bat coronavirus sequences that exhibited high functional similarity to human pathogens as outliers (Cluster − 1), pinpointing specific candidates within the “human-infective” embedding space (Fig. <xref rid="Fig5" ref-type="fig">5</xref>D and E). This group includes RaTG-13, which occupies a unique position in the embedding space near the hypothetical zoonotic boundary, reflecting its high similarity to human pathogens while accurately signaling its functional distinctness. Similarly, PRIME classified 50% of pangolin-derived sequences within Cluster 1 (Fig. <xref rid="Fig5" ref-type="fig">5</xref>D and E), aligning with experimental reports that specific pangolin-CoV strains exhibit high affinity for human ACE2 [<xref rid="CR59" ref-type="bibr">59</xref>]. By focusing on “human vs. non-human” functional profiles rather than over-interpreting evolutionary distance, PRIME identifies candidates for experimental prioritization based on their embedding-space proximity to human-infective strains, offering a potentially complementary signal for early-stage surveillance.</p><p id="Par42">The computational efficiency of PRIME suggests potential for near real-time deployment. Embedding-based clustering scales significantly better than maximum likelihood phylogenetics, allowing for the analysis of thousands of sequences in minutes. Additionally, the framework’s robustness to sequence divergence—demonstrated by the seamless integration of highly divergent hibecoviruses—makes it ideal for monitoring rapidly evolving pathogens that challenge traditional alignment tools.</p><p id="Par43">PRIME is designed as an analytical tool to assess phenotypic risk from sequence data rather than a generative model for sequence design. All sequence data used in this study are drawn from publicly accessible databases (GISAID, GenBank) under their respective data access agreements. Model weights for the ESM-RBD model (8 M only) and code are released under the MIT license with the intent of supporting legitimate biosurveillance research. Limitations on GitHub repo size and file size prevent the addition of other weights. As PLMs become more capable, integrating built-in guardrail, such as restricting applications to surveillance rather than design contexts, and ensuring compliance with institutional biosafety review, will be essential for their responsible use and deployment.</p></sec><sec id="Sec9" disp-level="1"><title>Conclusions</title><p id="Par44">PRIME establishes a scalable framework for pathogen biosurveillance that prioritizes true biological generalization over deceptive pattern memorization. By combining state-of-the-art model scale with domain-specific fine-tuning and stratified validation, it provides a framework that may assist in prioritizing sequences for experimental follow-up in the context of viral surveillance.</p></sec><sec id="Sec10" disp-level="1"><title>Methods</title><sec id="Sec11" disp-level="2"><title>Data collection and curation</title><p id="Par45">Three distinct datasets were utilized to develop and validate the PRIME framework:</p><sec id="Sec12" disp-level="3"><title>Outbreak dataset</title><p id="Par46">This dataset comprises 347,432 unique SARS-CoV-2 RBD sequences (amino acids 319–541of the spike protein) curated from GISAID and SRA as of February 3rd, 2023. All sequences containing ambiguous amino acids or duplicates were removed to ensure data integrity.</p></sec><sec id="Sec13" disp-level="3"><title>DMS dataset</title><p id="Par47">For phenotype prediction, we utilized experimental Deep Mutational Scanning (DMS) data containing 116,257 unique RBD sequences with measured ACE2 binding affinities and <bold>105</bold>,<bold>525 sequences</bold> with quantified protein expression levels from in vitro yeast display experiments.</p></sec><sec id="Sec14" disp-level="3"><title>BetaCov dataset</title><p id="Par48">To assess cross-species generalization, we assembled <bold>75 RBD sequences</bold> from diverse Sarbecovirus members, including SARS-CoV-1, SARS-CoV-2, Pangolin virus, Hibecovirus, and various bat viruses were sourced from a recent survey of sarbecovirus ACE2 binding [<xref rid="CR36" ref-type="bibr">36</xref>]. In total, there are 63 GenBank/NCBI sequences and 3 GISAID sequences, accessible through <italic>doi</italic>: 10.55876/gis8.250709on. MERS sequences were manually retrieved by querying the accession number <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.ncbi.nlm.nih.gov/protein/ALA49374.1" ext-link-type="uri">ALA49374.1</ext-link> using NIH’s BLASTp. From the top few hundred results, we carefully filtered out duplicates and identified relevant host species. The complete list of this dataset is presented in <bold>Supplementary Table 1</bold>.</p></sec></sec><sec id="Sec15" disp-level="2"><title>Model architecture and scale</title><p id="Par49">We systematically benchmarked five protein language model (PLM) architectures to evaluate the impact of parameter scale:</p><sec id="Sec16" disp-level="3"><title>ESM-2 family</title><p id="Par50">We utilized the 8 M (6 layers), 150 M (30 layers), and 650 M (33 layers) parameter models as baseline foundations.</p></sec><sec id="Sec17" disp-level="3"><title>ESM-C family</title><p id="Par51">We incorporated the state-of-the-art ESM-C 300 M and ESM-C 600 M models to assess the benefits of modern architecture design on viral representation learning.</p></sec><sec id="Sec18" disp-level="3"><title>BERT-RBD</title><p id="Par52">A BERT-based model was trained from scratch with a 320-dimensional embedding space as a domain-specific control.</p></sec></sec><sec id="Sec19" disp-level="2"><title>Domain-specific fine-tuning</title><p id="Par53">Fine-tuning was performed using a Masked Language Modeling (MLM) objective. For the ESM-RBD 8 M model, 15% of amino acids were randomly masked, and the model was tasked with predicting these residues based on contextual information. Fine-tuning was conducted over 100 epochs at a learning rate of 1 × 10<sup>− 5</sup>. For scaling tests, fine-tuning was performed over 15 epochs at the same learning rate. Fine-tuning used the AdamW optimizer with β1 = 0.9, β2 = 0.999 and no warm-up schedule. The effective batch size was 512, or 64 sequences per GPU across eight A100 GPUs using Distributed Data Parallel. Model selection was based on highest validation MLM accuracy across training epochs.</p></sec><sec id="Sec20" disp-level="2"><title>Position-stratified validation protocol</title><p id="Par54">To mitigate “homology leakage,” we implemented a rigorous position-stratified split for all phenotype precision tasks.</p><sec id="Sec21" disp-level="3"><title>Leakage control</title><p id="Par55">Unlike random splits, where mutations at the same site can appear in both training and test sets, position-stratified splitting ensures that all mutations at a specific residue position are held out together.</p></sec><sec id="Sec22" disp-level="3"><title>Cross-validation</title><p id="Par56">We utilized <bold>5-fold cross-validation</bold>. Residue positions were randomly partitioned into five subsets, and sequences were assigned to training or testing sets based on whether their mutated positions belonged exclusively to those subsets. To assess sensitivity to fold definition, we repeated the stratified evaluation across three independent random seeds (seeds 0, 1, 2) for fold assignment. Supplementary Fig. 10 confirms that mutation count distributions across partitions remain consistent across seeds and folds, indicating that reported performance metrics are not sensitive to the specific fold construction.</p></sec><sec id="Sec23" disp-level="3"><title>Input features</title><p id="Par57">Frozen embeddings from the last hidden layer (or mean embeddings where indicated) served as input to a multi-task FCN to jointly predict binding affinity and expression. The FCN consisted of 5 fully connected layers with hidden dimensions determined by the ESM embedding size (320 for ESM2 8 M, 640 for ESM2 150 M, 1280 for ESM2 650 M, 960 for ESMC 300 M, and 1152 for ESMC 600 M), ReLU activations, and no dropout, trained using mean squared error loss.Within each cross-validation split, the model was reinitialized, trained for 1000 epochs with a learning rate of 1e<sup>-5</sup>, and evaluated on the corresponding test set. Performance metrics, including R<sup>2</sup> and RMSE, were computed per fold and then averaged across folds.</p></sec></sec><sec id="Sec24" disp-level="2"><title>Clustering and phylogenetic analysis</title><sec id="Sec25" disp-level="3"><title>Latent space visualization</title><p id="Par58">High-dimensional embeddings were projected into 2D space using t-SNE [<xref rid="CR60" ref-type="bibr">60</xref>] with perplexity optimized between 30 and 750. Clustering accuracies were calculated by identifying the percentage of sequences correctly assigned to their majority clusters (100% for MERS, SARS-CoV-1, SARS-CoV-2, Pangolin Virus, and Hibecovirus; 69.7% for Bat Virus). This yielded an overall clustering accuracy of 73.7% when weighted by the number of sequences in each virus category. We note that t-SNE projections are used here for visualization and exploratory analysis. Clustering was performed in the original high-dimensional embedding space via HDBSCAN, with t-SNE used solely for 2D visualization of the resulting structure.</p></sec><sec id="Sec26" disp-level="3"><title>Clustering</title><p id="Par59">HDBSCAN [<xref rid="CR58" ref-type="bibr">58</xref>] was applied to t-SNE projections to identify de novo clusters. Performance was quantified using the Adjusted Rand Index (ARI) and Silhouette Coefficient (SC). To assess robustness to hyperparameter choices, we performed a systematic grid search over HDBSCAN min_cluster_size and min_sample parameters (Supplementary Figs. 9, 7E–F, 8E–F) and swept t-SNE perplexity from 30 to 750 (Supplementary Figs. 3–5). Lineage separation results were additionally validated across nine independent random downsampling seeds at fixed perplexity (Supplementary Fig. 2). All parameter sweep code is available in the GitHub repository.</p></sec><sec id="Sec27" disp-level="3"><title>Phylogenetics</title><p id="Par60">Sequences were aligned using <bold>MUSCLE (v5.3)</bold> with default settings [<xref rid="CR61" ref-type="bibr">61</xref>]. A maximum likelihood tree was constructed using <bold>IQ-TREE (v2.4.0)</bold> [<xref rid="CR62" ref-type="bibr">62</xref>] with the <bold>WAG+R2</bold> substitution model recommended by ModelFinder [<xref rid="CR63" ref-type="bibr">63</xref>] and 1000 ultrafast bootstrap [<xref rid="CR64" ref-type="bibr">64</xref>].</p><p id="Par61">The tree was midpoint rooted using DendroPy (v5.0.6) [<xref rid="CR65" ref-type="bibr">65</xref>] and visualized with ggtree (v3.10.1) [<xref rid="CR66" ref-type="bibr">66</xref>].</p></sec></sec><sec id="Sec28" disp-level="2"><title>Computational resources and reproducibility</title><p id="Par62">PRIME is implemented in PyTorch Lightning [2.5.1] (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/Lightning-AI/pytorch-lightning" ext-link-type="uri">https://github.com/Lightning-AI/pytorch-lightning</ext-link><italic>).</italic> All experiments (excluding UMAP generation and split comparisons post initial embedding of sequences) used Python 3.11.13, PyTorch [2.6.0 + cu124], PyTorch Lightning [2.5.1], and ESM library version [3.2.1]. For UMAP and split comparisons, we used Python 3.13.12, PyTorch [2.11.0 + cu130], and scikit-learn [1.8.0]. UMAP and HDBSCAN GPU versions provided by RAPIDS cuML [26.04.000]. Random seeds for fold assignment were set to 0, 1, and 2 for reproducibility. For model training, random seed was set to 0. Model weights, source code, and curated datasets are available under the MIT license in the GitHub repository (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/lanl/prime" ext-link-type="uri">https://github.com/lanl/prime</ext-link><bold><underline>)</underline></bold> to ensure compliance with FAIR data standards. We used eight NVIDIA A100 GPUs via a Distributed Data Parallel (DDP) strategy implemented in Pytorch Lightning. NVIDIA RTX Pro 6000 GPUs (not using DDP) were used for UMAP generation (available in the repository at <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/lanl/prime/tree/main/notebooks/clustering/blackwell" ext-link-type="uri">https://github.com/lanl/prime/tree/main/notebooks/clustering/blackwell</ext-link><bold><underline>)</underline></bold> as well as the comparison between random split and position-stratified split (Table <xref rid="Tab1" ref-type="table">1</xref> and <bold>Supplementary Table 6</bold>). The saved checkpoint with the highest validation accuracy from the ESM-RBD model was used for downstream tasks.</p><p id="Par63">Calculation of binding affinity and expression level errors (Fig. 4C and D) were determined using the formula: (predicted-measured)/measured * 100%.</p></sec><sec id="Sec29" disp-level="2"><title>LLM statement</title><p id="Par64">LLMs were used to aid code development.</p></sec></sec><sec id="Sec33" disp-level="1"><title>Supplementary Information</title>
<supplementary-material id="MOESM1" position="float"><media xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="12864_2026_12976_MOESM1_ESM.docx" mimetype="application" mime-subtype="vnd.openxmlformats-officedocument.wordprocessingml.document"><?cloudpmc-path 886e/13425921/809bbac21418/12864_2026_12976_MOESM1_ESM.docx?><?cloudpmc-bucket app?><?size 7546833?><caption><p>Supplementary Material 1.</p></caption></media></supplementary-material>

<supplementary-material id="MOESM2" position="float"><media xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="12864_2026_12976_MOESM2_ESM.xlsx" mimetype="application" mime-subtype="vnd.openxmlformats-officedocument.spreadsheetml.sheet"><?cloudpmc-path 886e/13425921/b4f3a5fdfe34/12864_2026_12976_MOESM2_ESM.xlsx?><?cloudpmc-bucket app?><?size 366970?><caption><p>Supplementary Material 2.</p></caption></media></supplementary-material>
</sec><sec id="ack1" sec-type="ack" disp-level="1"><title>Acknowledgements</title><p>We gratefully acknowledge all data contributors, i.e., the Authors and their Originating laboratories responsible for obtaining the specimens, and their Submitting laboratories for generating the genetic sequence and metadata and sharing via the GISAID Initiative, on which this research is based.</p></sec><sec id="notes1" disp-level="1"><title>Authors' contributions</title><p>B.H. developed the concept of this study. K.G., M.B., and B.H. implemented all the models. K.G. tuned and benchmarked all model performances. G.W.S. and K.G. developed model parallel execution scripts. P. L. and V.L. collected and curated data. M.D. performed phylogeny analysis. K.G., L.H., P.C. and B.H. interpreted the model results. All authors contributed to the writing of the paper.</p></sec><sec id="notes2" disp-level="1"><title>Funding</title><p>The authors appreciate funding from Los Alamos National Laboratory (LDRD Director Initiated Research, 20250637DI, 20250638DI, 20250639DI, 20240734DI, 20210767DI, 20200732ER). Part of this work was funded by the DOE Office of Science through the National Virtual Biotechnology Laboratory, a consortium of DOE national laboratories focused on response to COVID-19, with funding provided by the Coronavirus CARES Act. This project has been funded in part with Federal funds under Interagency Agreement No. 22FED2200087IPD (RRJJ) between the Centers for Disease Control and Prevention, Influenza Division and the US Department of Energy, Los Alamos National Laboratory. The GPU cluster was purchased by bioassurance funding through DOD. This work is approved for public release under LA-UR-26-21490.</p></sec><sec id="notes3" disp-level="1"><title>Data availability</title><p>All genome sequences and associated metadata in the outbreak dataset are published in GISAID’s EpiCoV database with the GISAID identifier: EPI_SET_240219yp. It is composed of 347,409 individual genome sequences with collection dates ranging from 2019-06-25 to 2023-05-17; data were collected in 202 countries and territories. To view the contributors of each individual sequence with details such as accession number, Virus name, Collection date, Originating Lab and Submitting Lab and the list of Authors, visit doi:10.55876/gis8.240219ypThe reference sequence for SARS-CoV-2 used in this study is hCoV-19/Wuhan/WIV04/2019 (WIV04), the official reference sequence employed by GISAID (EPI_ISL_402124, https://gisaid.org/WIV04).All GISAID sequences and associated metadata used in the BetaCov dataset are published in GISAID’s EpiCoV database with the identifier: EPI_SET_250709on. It is composed of 3 individual genome sequences with collection dates ranging from 2017 to 2019-06-25; the data was collected in 1 countries and territories. To view the contributors of each individual sequence with details such as accession number, Virus name, Collection date, Originating Lab and Submitting Lab and the list of Authors, visit doi: 10.55876/gis8.250709on. PRIME source code is available under the MIT license in the GitHub repository (https://github.com/lanl/prime).</p></sec><sec id="notes4" disp-level="1"><title>Declarations</title><sec id="FPar1" disp-level="2"><title>Ethics approval and consent to participate</title><p id="Par65">Not applicable.</p></sec><sec id="FPar2" disp-level="2"><title>Consent for publication</title><p id="Par66">Not applicable.</p></sec><sec id="FPar3" disp-level="2"><title>Competing interests</title><p id="Par67">The authors declare no competing interests.</p></sec></sec><sec id="fn-group1" sec-type="fn-group" disp-level="1"><title>Footnotes</title><fn-group><fn id="fn1"><p><bold>Publisher’s note</bold></p><p>Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p></fn></fn-group></sec><sec id="Bib1" sec-type="ref-list" disp-level="1"><title>References</title><sec id="Bib1_sec2" disp-level="2"><ref-list><ref id="CR1"><label>1.</label><mixed-citation id="mc-CR1"><named-content content-type="citation-string">Pybus OG, Rambaut A. Evolutionary analysis of the dynamics of viral infectious disease. Nat Rev Genet. 2009;10:540–50. 10.1038/nrg2583.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/nrg2583"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7097015"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="19564871"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Pybus OG, Rambaut A. Evolutionary analysis of the dynamics of viral infectious disease. Nat Rev Genet. 2009;10:540–50. 10.1038/nrg2583."/></mixed-citation></ref><ref id="CR2"><label>2.</label><mixed-citation id="mc-CR2"><named-content content-type="citation-string">Tsetsarkin KA, Vanlandingham DL, McGee CE, Higgs S. A Single Mutation in Chikungunya Virus Affects Vector Specificity and Epidemic Potential. PLOS Pathog. 2007;3:e201. 10.1371/journal.ppat.0030201.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1371/journal.ppat.0030201"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC2134949"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="18069894"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Tsetsarkin KA, Vanlandingham DL, McGee CE, Higgs S. A Single Mutation in Chikungunya Virus Affects Vector Specificity and Epidemic Potential. PLOS Pathog. 2007;3:e201. 10.1371/journal.ppat.0030201."/></mixed-citation></ref><ref id="CR3"><label>3.</label><mixed-citation id="mc-CR3"><named-content content-type="citation-string">Zhao S, Lou J, Cao L, Zheng H, Chong MKC, Chen Z, et al. Real-time quantification of the transmission advantage associated with a single mutation in pathogen genomes: a case study on the D614G substitution of SARS-CoV-2. BMC Infect Dis. 2021;21:1039. 10.1186/s12879-021-06729-w.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1186/s12879-021-06729-w"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8495436"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="34620109"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Zhao S, Lou J, Cao L, Zheng H, Chong MKC, Chen Z, et al. Real-time quantification of the transmission advantage associated with a single mutation in pathogen genomes: a case study on the D614G substitution of SARS-CoV-2. BMC Infect Dis. 2021;21:1039. 10.1186/s12879-021-06729-w."/></mixed-citation></ref><ref id="CR4"><label>4.</label><mixed-citation id="mc-CR4"><named-content content-type="citation-string">Korber B, Fischer WM, Gnanakaran S, Yoon H, Theiler J, Abfalterer W, et al. Tracking Changes in SARS-CoV-2 Spike: Evidence that D614G Increases Infectivity of the COVID-19 Virus. Cell. 2020;182:812–e82719. 10.1016/j.cell.2020.06.043.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.cell.2020.06.043"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7332439"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32697968"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Korber B, Fischer WM, Gnanakaran S, Yoon H, Theiler J, Abfalterer W, et al. Tracking Changes in SARS-CoV-2 Spike: Evidence that D614G Increases Infectivity of the COVID-19 Virus. Cell. 2020;182:812–e82719. 10.1016/j.cell.2020.06.043."/></mixed-citation></ref><ref id="CR5"><label>5.</label><mixed-citation id="mc-CR5"><named-content content-type="citation-string">Volz E, Hill V, McCrone JT, Price A, Jorgensen D, O’Toole Á, et al. Evaluating the Effects of SARS-CoV-2 Spike Mutation D614G on Transmissibility and Pathogenicity. Cell. 2021;184:64–e7511. 10.1016/j.cell.2020.11.020.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.cell.2020.11.020"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7674007"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33275900"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Volz E, Hill V, McCrone JT, Price A, Jorgensen D, O’Toole Á, et al. Evaluating the Effects of SARS-CoV-2 Spike Mutation D614G on Transmissibility and Pathogenicity. Cell. 2021;184:64–e7511. 10.1016/j.cell.2020.11.020."/></mixed-citation></ref><ref id="CR6"><label>6.</label><mixed-citation id="mc-CR6"><named-content content-type="citation-string">Kleanthous H, Silverman JM, Makar KW, Yoon I-K, Jackson N, Vaughn DW. Scientific rationale for developing potent RBD-based vaccines targeting COVID-19. Npj Vaccines. 2021;6:1–10. 10.1038/s41541-021-00393-6.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41541-021-00393-6"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8553742"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="34711846"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Kleanthous H, Silverman JM, Makar KW, Yoon I-K, Jackson N, Vaughn DW. Scientific rationale for developing potent RBD-based vaccines targeting COVID-19. Npj Vaccines. 2021;6:1–10. 10.1038/s41541-021-00393-6."/></mixed-citation></ref><ref id="CR7"><label>7.</label><mixed-citation id="mc-CR7"><named-content content-type="citation-string">Yang J, Wang W, Chen Z, Lu S, Yang F, Bi Z, et al. A vaccine targeting the RBD of the S protein of SARS-CoV-2 induces protective immunity. Nature. 2020;586:572–7. 10.1038/s41586-020-2599-8.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-020-2599-8"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32726802"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Yang J, Wang W, Chen Z, Lu S, Yang F, Bi Z, et al. A vaccine targeting the RBD of the S protein of SARS-CoV-2 induces protective immunity. Nature. 2020;586:572–7. 10.1038/s41586-020-2599-8."/></mixed-citation></ref><ref id="CR8"><label>8.</label><mixed-citation id="mc-CR8"><named-content content-type="citation-string">Carabelli AM, Peacock TP, Thorne LG, Harvey WT, Hughes J, de Silva TI, et al. SARS-CoV-2 variant biology: immune escape, transmission and fitness. Nat Rev Microbiol. 2023;21:162–77. 10.1038/s41579-022-00841-7.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41579-022-00841-7"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC9847462"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="36653446"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Carabelli AM, Peacock TP, Thorne LG, Harvey WT, Hughes J, de Silva TI, et al. SARS-CoV-2 variant biology: immune escape, transmission and fitness. Nat Rev Microbiol. 2023;21:162–77. 10.1038/s41579-022-00841-7."/></mixed-citation></ref><ref id="CR9"><label>9.</label><mixed-citation id="mc-CR9"><named-content content-type="citation-string">McCallum M, Walls AC, Sprouse KR, Bowen JE, Rosen LE, Dang HV, et al. Molecular basis of immune evasion by the Delta and Kappa SARS-CoV-2 variants. Science. 2021. 10.1126/science.abl8506.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1126/science.abl8506"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC12240541"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="34751595"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="McCallum M, Walls AC, Sprouse KR, Bowen JE, Rosen LE, Dang HV, et al. Molecular basis of immune evasion by the Delta and Kappa SARS-CoV-2 variants. Science. 2021. 10.1126/science.abl8506."/></mixed-citation></ref><ref id="CR10"><label>10.</label><mixed-citation id="mc-CR10"><named-content content-type="citation-string">Longdon B, Brockhurst MA, Russell CA, Welch JJ, Jiggins FM. The Evolution and Genetics of Virus Host Shifts. PLOS Pathog. 2014;10:e1004395. 10.1371/journal.ppat.1004395.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1371/journal.ppat.1004395"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC4223060"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="25375777"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Longdon B, Brockhurst MA, Russell CA, Welch JJ, Jiggins FM. The Evolution and Genetics of Virus Host Shifts. PLOS Pathog. 2014;10:e1004395. 10.1371/journal.ppat.1004395."/></mixed-citation></ref><ref id="CR11"><label>11.</label><mixed-citation id="mc-CR11"><named-content content-type="citation-string">Long JS, Mistry B, Haslam SM, Barclay WS. Host and viral determinants of influenza A virus species specificity. Nat Rev Microbiol. 2019;17:67–81. 10.1038/s41579-018-0115-z.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41579-018-0115-z"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="30487536"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Long JS, Mistry B, Haslam SM, Barclay WS. Host and viral determinants of influenza A virus species specificity. Nat Rev Microbiol. 2019;17:67–81. 10.1038/s41579-018-0115-z."/></mixed-citation></ref><ref id="CR12"><label>12.</label><mixed-citation id="mc-CR12"><named-content content-type="citation-string">Peacock TP, Moncla L, Dudas G, VanInsberghe D, Sukhova K, Lloyd-Smith JO, et al. The global H5N1 influenza panzootic in mammals. Nature. 2025;637:304–13. 10.1038/s41586-024-08054-z.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-024-08054-z"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="39317240"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Peacock TP, Moncla L, Dudas G, VanInsberghe D, Sukhova K, Lloyd-Smith JO, et al. The global H5N1 influenza panzootic in mammals. Nature. 2025;637:304–13. 10.1038/s41586-024-08054-z."/></mixed-citation></ref><ref id="CR13"><label>13.</label><mixed-citation id="mc-CR13"><named-content content-type="citation-string">Feng DF, Doolittle RF. Progressive sequence alignment as a prerequisite to correct phylogenetic trees. J Mol Evol. 1987;25:351–60. 10.1007/BF02603120.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1007/BF02603120"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="3118049"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Feng DF, Doolittle RF. Progressive sequence alignment as a prerequisite to correct phylogenetic trees. J Mol Evol. 1987;25:351–60. 10.1007/BF02603120."/></mixed-citation></ref><ref id="CR14"><label>14.</label><mixed-citation id="mc-CR14"><named-content content-type="citation-string">Hall BG, Barlow M. Phylogenetic Analysis as a Tool in Molecular Epidemiology of Infectious Diseases. Ann Epidemiol. 2006;16:157–69. 10.1016/j.annepidem.2005.04.010.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.annepidem.2005.04.010"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="16099674"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Hall BG, Barlow M. Phylogenetic Analysis as a Tool in Molecular Epidemiology of Infectious Diseases. Ann Epidemiol. 2006;16:157–69. 10.1016/j.annepidem.2005.04.010."/></mixed-citation></ref><ref id="CR15"><label>15.</label><mixed-citation id="mc-CR15"><named-content content-type="citation-string">Phillips A, Janies D, Wheeler W. Multiple Sequence Alignment in Phylogenetic Analysis. Mol Phylogenet Evol. 2000;16:317–30. 10.1006/mpev.2000.0785.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1006/mpev.2000.0785"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="10991785"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Phillips A, Janies D, Wheeler W. Multiple Sequence Alignment in Phylogenetic Analysis. Mol Phylogenet Evol. 2000;16:317–30. 10.1006/mpev.2000.0785."/></mixed-citation></ref><ref id="CR16"><label>16.</label><mixed-citation id="mc-CR16"><named-content content-type="citation-string">Ashkenazy H, Sela I, Levy Karin E, Landan G, Pupko T. Multiple Sequence Alignment Averaging Improves Phylogeny Reconstruction. Syst Biol. 2019;68:117–30. 10.1093/sysbio/syy036.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/sysbio/syy036"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC6657586"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="29771363"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Ashkenazy H, Sela I, Levy Karin E, Landan G, Pupko T. Multiple Sequence Alignment Averaging Improves Phylogeny Reconstruction. Syst Biol. 2019;68:117–30. 10.1093/sysbio/syy036."/></mixed-citation></ref><ref id="CR17"><label>17.</label><mixed-citation id="mc-CR17"><named-content content-type="citation-string">Thomson EC, Rosen LE, Shepherd JG, Spreafico R, da Silva Filipe A, Wojcechowskyj JA, et al. Circulating SARS-CoV-2 spike N439K variants maintain fitness while evading antibody-mediated immunity. Cell. 2021;184:1171–e118720. 10.1016/j.cell.2021.01.037.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.cell.2021.01.037"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7843029"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33621484"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Thomson EC, Rosen LE, Shepherd JG, Spreafico R, da Silva Filipe A, Wojcechowskyj JA, et al. Circulating SARS-CoV-2 spike N439K variants maintain fitness while evading antibody-mediated immunity. Cell. 2021;184:1171–e118720. 10.1016/j.cell.2021.01.037."/></mixed-citation></ref><ref id="CR18"><label>18.</label><mixed-citation id="mc-CR18"><named-content content-type="citation-string">Hadfield J, Megill C, Bell SM, Huddleston J, Potter B, Callender C, et al. Nextstrain: real-time tracking of pathogen evolution. Bioinforma Oxf Engl. 2018;34:4121–3. 10.1093/bioinformatics/bty407.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/bioinformatics/bty407"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC6247931"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="29790939"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Hadfield J, Megill C, Bell SM, Huddleston J, Potter B, Callender C, et al. Nextstrain: real-time tracking of pathogen evolution. Bioinforma Oxf Engl. 2018;34:4121–3. 10.1093/bioinformatics/bty407."/></mixed-citation></ref><ref id="CR19"><label>19.</label><mixed-citation id="mc-CR19"><named-content content-type="citation-string">Steenwyk JL, Li Y, Zhou X, Shen X-X, Rokas A. Incongruence in the phylogenomics era. Nat Rev Genet. 2023;24:834–50. 10.1038/s41576-023-00620-x.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41576-023-00620-x"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC11499941"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="37369847"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Steenwyk JL, Li Y, Zhou X, Shen X-X, Rokas A. Incongruence in the phylogenomics era. Nat Rev Genet. 2023;24:834–50. 10.1038/s41576-023-00620-x."/></mixed-citation></ref><ref id="CR20"><label>20.</label><mixed-citation id="mc-CR20"><named-content content-type="citation-string">De Maio N, Kalaghatgi P, Turakhia Y, Corbett-Detig R, Minh BQ, Goldman N. Maximum likelihood pandemic-scale phylogenetics. Nat Genet. 2023;55:746–52. 10.1038/s41588-023-01368-0.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41588-023-01368-0"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC10181937"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="37038003"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="De Maio N, Kalaghatgi P, Turakhia Y, Corbett-Detig R, Minh BQ, Goldman N. Maximum likelihood pandemic-scale phylogenetics. Nat Genet. 2023;55:746–52. 10.1038/s41588-023-01368-0."/></mixed-citation></ref><ref id="CR21"><label>21.</label><mixed-citation id="mc-CR21"><named-content content-type="citation-string">Posada D, Crandall KA. The Effect of Recombination on the Accuracy of Phylogeny Estimation. J Mol Evol. 2002;54:396–402. 10.1007/s00239-001-0034-9.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1007/s00239-001-0034-9"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="11847565"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Posada D, Crandall KA. The Effect of Recombination on the Accuracy of Phylogeny Estimation. J Mol Evol. 2002;54:396–402. 10.1007/s00239-001-0034-9."/></mixed-citation></ref><ref id="CR22"><label>22.</label><mixed-citation id="mc-CR22"><named-content content-type="citation-string">Rancati S, Nicora G, Prosperi M, Bellazzi R, Salemi M, Marini S. Forecasting dominance of SARS-CoV-2 lineages by anomaly detection using deep AutoEncoders. Brief Bioinform. 2024;25:bbae535. 10.1093/bib/bbae535.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/bib/bbae535"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC11500442"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="39446192"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Rancati S, Nicora G, Prosperi M, Bellazzi R, Salemi M, Marini S. Forecasting dominance of SARS-CoV-2 lineages by anomaly detection using deep AutoEncoders. Brief Bioinform. 2024;25:bbae535. 10.1093/bib/bbae535."/></mixed-citation></ref><ref id="CR23"><label>23.</label><mixed-citation id="mc-CR23"><named-content content-type="citation-string">Rancati S, Nicora G, Bergomi L, Buonocore TM, Czyz DM, Parimbelli E, et al. SARITA: a large language model for generating the S1 subunit of the SARS-CoV-2 spike protein. Brief Bioinform. 2025;26:bbaf384. 10.1093/bib/bbaf384.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/bib/bbaf384"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC12319310"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="40755284"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Rancati S, Nicora G, Bergomi L, Buonocore TM, Czyz DM, Parimbelli E, et al. SARITA: a large language model for generating the S1 subunit of the SARS-CoV-2 spike protein. Brief Bioinform. 2025;26:bbaf384. 10.1093/bib/bbaf384."/></mixed-citation></ref><ref id="CR24"><label>24.</label><mixed-citation id="mc-CR24"><named-content content-type="citation-string">Fung TS, Liu DX. Human Coronavirus: Host-Pathogen Interaction. Annu Rev Microbiol. 2019;73:529–57. 10.1146/annurev-micro-020518-115759.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1146/annurev-micro-020518-115759"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="31226023"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Fung TS, Liu DX. Human Coronavirus: Host-Pathogen Interaction. Annu Rev Microbiol. 2019;73:529–57. 10.1146/annurev-micro-020518-115759."/></mixed-citation></ref><ref id="CR25"><label>25.</label><mixed-citation id="mc-CR25"><named-content content-type="citation-string">Cui J, Li F, Shi Z-L. Origin and evolution of pathogenic coronaviruses. Nat Rev Microbiol. 2019;17:181–92. 10.1038/s41579-018-0118-9.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41579-018-0118-9"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7097006"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="30531947"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Cui J, Li F, Shi Z-L. Origin and evolution of pathogenic coronaviruses. Nat Rev Microbiol. 2019;17:181–92. 10.1038/s41579-018-0118-9."/></mixed-citation></ref><ref id="CR26"><label>26.</label><mixed-citation id="mc-CR26"><named-content content-type="citation-string">Shang J, Wan Y, Luo C, Ye G, Geng Q, Auerbach A, et al. Cell entry mechanisms of SARS-CoV-2. Proc Natl Acad Sci U S A. 2020;117:11727–34. 10.1073/pnas.2003138117.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1073/pnas.2003138117"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7260975"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32376634"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Shang J, Wan Y, Luo C, Ye G, Geng Q, Auerbach A, et al. Cell entry mechanisms of SARS-CoV-2. Proc Natl Acad Sci U S A. 2020;117:11727–34. 10.1073/pnas.2003138117."/></mixed-citation></ref><ref id="CR27"><label>27.</label><mixed-citation id="mc-CR27"><named-content content-type="citation-string">Starr TN, Zepeda SK, Walls AC, Greaney AJ, Alkhovsky S, Veesler D, et al. ACE2 binding is an ancestral and evolvable trait of sarbecoviruses. Nature. 2022;603:913–8. 10.1038/s41586-022-04464-z.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-022-04464-z"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8967715"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="35114688"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Starr TN, Zepeda SK, Walls AC, Greaney AJ, Alkhovsky S, Veesler D, et al. ACE2 binding is an ancestral and evolvable trait of sarbecoviruses. Nature. 2022;603:913–8. 10.1038/s41586-022-04464-z."/></mixed-citation></ref><ref id="CR28"><label>28.</label><mixed-citation id="mc-CR28"><named-content content-type="citation-string">Lan J, Ge J, Yu J, Shan S, Zhou H, Fan S, et al. Structure of the SARS-CoV-2 spike receptor-binding domain bound to the ACE2 receptor. Nature. 2020;581:215–20. 10.1038/s41586-020-2180-5.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-020-2180-5"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32225176"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Lan J, Ge J, Yu J, Shan S, Zhou H, Fan S, et al. Structure of the SARS-CoV-2 spike receptor-binding domain bound to the ACE2 receptor. Nature. 2020;581:215–20. 10.1038/s41586-020-2180-5."/></mixed-citation></ref><ref id="CR29"><label>29.</label><mixed-citation id="mc-CR29"><named-content content-type="citation-string">Mou H, Raj VS, van Kuppeveld FJM, Rottier PJM, Haagmans BL, Bosch BJ. The receptor binding domain of the new Middle East respiratory syndrome coronavirus maps to a 231-residue region in the spike protein that efficiently elicits neutralizing antibodies. J Virol. 2013;87:9379–83. 10.1128/JVI.01277-13.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1128/JVI.01277-13"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC3754068"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="23785207"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Mou H, Raj VS, van Kuppeveld FJM, Rottier PJM, Haagmans BL, Bosch BJ. The receptor binding domain of the new Middle East respiratory syndrome coronavirus maps to a 231-residue region in the spike protein that efficiently elicits neutralizing antibodies. J Virol. 2013;87:9379–83. 10.1128/JVI.01277-13."/></mixed-citation></ref><ref id="CR30"><label>30.</label><mixed-citation id="mc-CR30"><named-content content-type="citation-string">Wang N, Shi X, Jiang L, Zhang S, Wang D, Tong P, et al. Structure of MERS-CoV spike receptor-binding domain complexed with human receptor DPP4. Cell Res. 2013;23:986–93. 10.1038/cr.2013.92.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/cr.2013.92"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC3731569"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="23835475"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Wang N, Shi X, Jiang L, Zhang S, Wang D, Tong P, et al. Structure of MERS-CoV spike receptor-binding domain complexed with human receptor DPP4. Cell Res. 2013;23:986–93. 10.1038/cr.2013.92."/></mixed-citation></ref><ref id="CR31"><label>31.</label><mixed-citation id="mc-CR31"><named-content content-type="citation-string">Cao Y, Wang J, Jian F, Xiao T, Song W, Yisimayi A, et al. Omicron escapes the majority of existing SARS-CoV-2 neutralizing antibodies. Nature. 2022;602:657–63. 10.1038/s41586-021-04385-3.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-021-04385-3"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8866119"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="35016194"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Cao Y, Wang J, Jian F, Xiao T, Song W, Yisimayi A, et al. Omicron escapes the majority of existing SARS-CoV-2 neutralizing antibodies. Nature. 2022;602:657–63. 10.1038/s41586-021-04385-3."/></mixed-citation></ref><ref id="CR32"><label>32.</label><mixed-citation><named-content content-type="citation-string">Dejnirattisai W, Huo J, Zhou D, Zahradník J, Supasa P, Liu C, et al. SARS-CoV-2 Omicron-B.1.1.529 leads to widespread escape from neutralizing antibody responses. Cell. 2022;0. 10.1016/j.cell.2021.12.046.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.cell.2021.12.046"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8723827"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="35081335"/></mixed-citation></ref><ref id="CR33"><label>33.</label><mixed-citation><named-content content-type="citation-string">McCallumM, Czudnochowski N, Rosen LE, Zepeda SK, Bowen JE, Walls AC et al. Structural basis of SARS-CoV-2 Omicron immune evasion and receptor engagement. Science. 2022; 375(6583):864-8 . 10.1126/science.abn8652.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1126/science.abn8652"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC9427005"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="35076256"/></mixed-citation></ref><ref id="CR34"><label>34.</label><mixed-citation id="mc-CR34"><named-content content-type="citation-string">Willett BJ, Grove J, MacLean OA, Wilkie C, De Lorenzo G, Furnon W, et al. SARS-CoV-2 Omicron is an immune escape variant with an altered cell entry pathway. Nat Microbiol. 2022;7:1161–79. 10.1038/s41564-022-01143-7.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41564-022-01143-7"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC9352574"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="35798890"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Willett BJ, Grove J, MacLean OA, Wilkie C, De Lorenzo G, Furnon W, et al. SARS-CoV-2 Omicron is an immune escape variant with an altered cell entry pathway. Nat Microbiol. 2022;7:1161–79. 10.1038/s41564-022-01143-7."/></mixed-citation></ref><ref id="CR35"><label>35.</label><mixed-citation id="mc-CR35"><named-content content-type="citation-string">Fowler DM, Fields S. Deep mutational scanning: a new style of protein science. Nat Methods. 2014;11:801–7. 10.1038/nmeth.3027.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/nmeth.3027"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC4410700"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="25075907"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Fowler DM, Fields S. Deep mutational scanning: a new style of protein science. Nat Methods. 2014;11:801–7. 10.1038/nmeth.3027."/></mixed-citation></ref><ref id="CR36"><label>36.</label><mixed-citation id="mc-CR36"><named-content content-type="citation-string">Starr TN, Greaney AJ, Hilton SK, Ellis D, Crawford KHD, Dingens AS, et al. Deep Mutational Scanning of SARS-CoV-2 Receptor Binding Domain Reveals Constraints on Folding and ACE2 Binding. Cell. 2020;182:1295–e131020. 10.1016/j.cell.2020.08.012.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.cell.2020.08.012"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7418704"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32841599"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Starr TN, Greaney AJ, Hilton SK, Ellis D, Crawford KHD, Dingens AS, et al. Deep Mutational Scanning of SARS-CoV-2 Receptor Binding Domain Reveals Constraints on Folding and ACE2 Binding. Cell. 2020;182:1295–e131020. 10.1016/j.cell.2020.08.012."/></mixed-citation></ref><ref id="CR37"><label>37.</label><mixed-citation id="mc-CR37"><named-content content-type="citation-string">Frank F, Keen MM, Rao A, Bassit L, Liu X, Bowers HB, et al. Deep mutational scanning identifies SARS-CoV-2 Nucleocapsid escape mutations of currently available rapid antigen tests. Cell. 2022;185:3603–e361613. 10.1016/j.cell.2022.08.010.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.cell.2022.08.010"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC9420710"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="36084631"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Frank F, Keen MM, Rao A, Bassit L, Liu X, Bowers HB, et al. Deep mutational scanning identifies SARS-CoV-2 Nucleocapsid escape mutations of currently available rapid antigen tests. Cell. 2022;185:3603–e361613. 10.1016/j.cell.2022.08.010."/></mixed-citation></ref><ref id="CR38"><label>38.</label><mixed-citation id="mc-CR38"><named-content content-type="citation-string">Dadonaite B, Brown J, McMahon TE, Farrell AG, Figgins MD, Asarnow D, et al. Spike deep mutational scanning helps predict success of SARS-CoV-2 clades. Nature. 2024;631:617–26. 10.1038/s41586-024-07636-1.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-024-07636-1"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC11254757"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="38961298"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Dadonaite B, Brown J, McMahon TE, Farrell AG, Figgins MD, Asarnow D, et al. Spike deep mutational scanning helps predict success of SARS-CoV-2 clades. Nature. 2024;631:617–26. 10.1038/s41586-024-07636-1."/></mixed-citation></ref><ref id="CR39"><label>39.</label><mixed-citation id="mc-CR39"><named-content content-type="citation-string">Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583–9. 10.1038/s41586-021-03819-2.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-021-03819-2"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8371605"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="34265844"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583–9. 10.1038/s41586-021-03819-2."/></mixed-citation></ref><ref id="CR40"><label>40.</label><mixed-citation id="mc-CR40"><named-content content-type="citation-string">Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630:493–500. 10.1038/s41586-024-07487-w.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-024-07487-w"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC11168924"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="38718835"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630:493–500. 10.1038/s41586-024-07487-w."/></mixed-citation></ref><ref id="CR41"><label>41.</label><mixed-citation><named-content content-type="citation-string">Rives A, Meier J, Sercu T, Goyal S, Lin Z, Liu J, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc Natl Acad Sci. 2021;118(15):e2016239118 . 10.1073/pnas.2016239118.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1073/pnas.2016239118"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8053943"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33876751"/></mixed-citation></ref><ref id="CR42"><label>42.</label><mixed-citation id="mc-CR42"><named-content content-type="citation-string">Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379:1123–30. 10.1126/science.ade2574.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1126/science.ade2574"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="36927031"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379:1123–30. 10.1126/science.ade2574."/></mixed-citation></ref><ref id="CR43"><label>43.</label><mixed-citation id="mc-CR43"><named-content content-type="citation-string">Brandes N, Ofer D, Peleg Y, Rappoport N, Linial M. ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics. 2022;38:2102–10. 10.1093/bioinformatics/btac020.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/bioinformatics/btac020"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC9386727"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="35020807"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Brandes N, Ofer D, Peleg Y, Rappoport N, Linial M. ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics. 2022;38:2102–10. 10.1093/bioinformatics/btac020."/></mixed-citation></ref><ref id="CR44"><label>44.</label><mixed-citation><named-content content-type="citation-string">Rao R, Bhattacharya N, Thomas N, Duan Y, Chen X, Canny J et al. Evaluating Protein Transfer Learning with TAPE. 2019. 10.48550/arXiv.1906.08230</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7774645"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33390682"/></mixed-citation></ref><ref id="CR45"><label>45.</label><mixed-citation id="mc-CR45"><named-content content-type="citation-string">Shaw R, Love SD, McWhite CD. Evaluating Pretrained Protein Language Model Embeddings as Proxies for Functional Similarity. J Mol Evol. 2025;93:765–76. 10.1007/s00239-025-10282-4.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1007/s00239-025-10282-4"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC12756192"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="41273410"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Shaw R, Love SD, McWhite CD. Evaluating Pretrained Protein Language Model Embeddings as Proxies for Functional Similarity. J Mol Evol. 2025;93:765–76. 10.1007/s00239-025-10282-4."/></mixed-citation></ref><ref id="CR46"><label>46.</label><mixed-citation id="mc-CR46"><named-content content-type="citation-string">Szymborski J, Emad A. A flaw in using pretrained protein language models in protein–protein interaction inference models. Nat Mach Intell. 2026;8:197–208. 10.1038/s42256-025-01176-7.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Szymborski J, Emad A. A flaw in using pretrained protein language models in protein–protein interaction inference models. Nat Mach Intell. 2026;8:197–208. 10.1038/s42256-025-01176-7."/></mixed-citation></ref><ref id="CR47"><label>47.</label><mixed-citation id="mc-CR47"><named-content content-type="citation-string">Elbe S, Buckland-Merrett G. Data, disease and diplomacy: GISAID’s innovative contribution to global health. Glob Chall. 2017;1:33–46. 10.1002/gch2.1018.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1002/gch2.1018"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC6607375"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="31565258"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Elbe S, Buckland-Merrett G. Data, disease and diplomacy: GISAID’s innovative contribution to global health. Glob Chall. 2017;1:33–46. 10.1002/gch2.1018."/></mixed-citation></ref><ref id="CR48"><label>48.</label><mixed-citation><named-content content-type="citation-string">Leinonen R, Sugawara H, Shumway M. The Sequence Read Archive. Nucleic Acids Res. 2011;39. 10.1093/nar/gkq1019. Database issue:D19–21.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/nar/gkq1019"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC3013647"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="21062823"/></mixed-citation></ref><ref id="CR49"><label>49.</label><mixed-citation id="mc-CR49"><named-content content-type="citation-string">Wu F, Zhao S, Yu B, Chen Y-M, Wang W, Song Z-G, et al. A new coronavirus associated with human respiratory disease in China. Nature. 2020;579:265–9. 10.1038/s41586-020-2008-3.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-020-2008-3"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7094943"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32015508"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Wu F, Zhao S, Yu B, Chen Y-M, Wang W, Song Z-G, et al. A new coronavirus associated with human respiratory disease in China. Nature. 2020;579:265–9. 10.1038/s41586-020-2008-3."/></mixed-citation></ref><ref id="CR50"><label>50.</label><mixed-citation id="mc-CR50"><named-content content-type="citation-string">Zhou P, Yang X-L, Wang X-G, Hu B, Zhang L, Zhang W, et al. A pneumonia outbreak associated with a new coronavirus of probable bat origin. Nature. 2020;579:270–3. 10.1038/s41586-020-2012-7.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-020-2012-7"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7095418"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32015507"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Zhou P, Yang X-L, Wang X-G, Hu B, Zhang L, Zhang W, et al. A pneumonia outbreak associated with a new coronavirus of probable bat origin. Nature. 2020;579:270–3. 10.1038/s41586-020-2012-7."/></mixed-citation></ref><ref id="CR51"><label>51.</label><mixed-citation id="mc-CR51"><named-content content-type="citation-string">Skowronski DM, Astell C, Brunham RC, Low DE, Petric M, Roper RL, et al. Severe Acute Respiratory Syndrome (SARS): A Year in Review. Annu Rev Med. 2005;56:357–81. 10.1146/annurev.med.56.091103.134135.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1146/annurev.med.56.091103.134135"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="15660517"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Skowronski DM, Astell C, Brunham RC, Low DE, Petric M, Roper RL, et al. Severe Acute Respiratory Syndrome (SARS): A Year in Review. Annu Rev Med. 2005;56:357–81. 10.1146/annurev.med.56.091103.134135."/></mixed-citation></ref><ref id="CR52"><label>52.</label><mixed-citation id="mc-CR52"><named-content content-type="citation-string">Zaki AM, van Boheemen S, Bestebroer TM, Osterhaus ADME, Fouchier RAM. Isolation of a novel coronavirus from a man with pneumonia in Saudi Arabia. N Engl J Med. 2012;367:1814–20. 10.1056/NEJMoa1211721.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1056/NEJMoa1211721"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="23075143"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Zaki AM, van Boheemen S, Bestebroer TM, Osterhaus ADME, Fouchier RAM. Isolation of a novel coronavirus from a man with pneumonia in Saudi Arabia. N Engl J Med. 2012;367:1814–20. 10.1056/NEJMoa1211721."/></mixed-citation></ref><ref id="CR53"><label>53.</label><mixed-citation id="mc-CR53"><named-content content-type="citation-string">Xiao K, Zhai J, Feng Y, Zhou N, Zhang X, Zou J-J, et al. Isolation of SARS-CoV-2-related coronavirus from Malayan pangolins. Nature. 2020;583:286–9. 10.1038/s41586-020-2313-x.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-020-2313-x"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32380510"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Xiao K, Zhai J, Feng Y, Zhou N, Zhang X, Zou J-J, et al. Isolation of SARS-CoV-2-related coronavirus from Malayan pangolins. Nature. 2020;583:286–9. 10.1038/s41586-020-2313-x."/></mixed-citation></ref><ref id="CR54"><label>54.</label><mixed-citation><named-content content-type="citation-string">DevlinJ, Chang M-W, Lee K, Toutanova K, BERT. Pre-training of Deep Bidirectional Transformers for Language Understanding. ArXiv181004805 Cs. 2019. 10.48550/arXiv.1810.04805.</named-content></mixed-citation></ref><ref id="CR55"><label>55.</label><mixed-citation id="mc-CR55"><named-content content-type="citation-string">Zhang S, Tong H, Xu J, Maciejewski R. Graph convolutional networks: a comprehensive review. Comput Soc Netw. 2019;6:11. 10.1186/s40649-019-0069-y.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1186/s40649-019-0069-y"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC10615927"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="37915858"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Zhang S, Tong H, Xu J, Maciejewski R. Graph convolutional networks: a comprehensive review. Comput Soc Netw. 2019;6:11. 10.1186/s40649-019-0069-y."/></mixed-citation></ref><ref id="CR56"><label>56.</label><mixed-citation id="mc-CR56"><named-content content-type="citation-string">Hochreiter S, Schmidhuber J. Long Short-Term Memory. Neural Comput. 1997;9:1735–80. 10.1162/neco.1997.9.8.1735.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1162/neco.1997.9.8.1735"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="9377276"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Hochreiter S, Schmidhuber J. Long Short-Term Memory. Neural Comput. 1997;9:1735–80. 10.1162/neco.1997.9.8.1735."/></mixed-citation></ref><ref id="CR57"><label>57.</label><mixed-citation id="mc-CR57"><named-content content-type="citation-string">Gu H, Chen Q, Yang G, He L, Fan H, Deng Y-Q, et al. Adaptation of SARS-CoV-2 in BALB/c mice for testing vaccine efficacy. Science. 2020;369:1603–7. 10.1126/science.abc4730.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1126/science.abc4730"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7574913"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32732280"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Gu H, Chen Q, Yang G, He L, Fan H, Deng Y-Q, et al. Adaptation of SARS-CoV-2 in BALB/c mice for testing vaccine efficacy. Science. 2020;369:1603–7. 10.1126/science.abc4730."/></mixed-citation></ref><ref id="CR58"><label>58.</label><mixed-citation id="mc-CR58"><named-content content-type="citation-string">McInnes L, Healy J, Astels S. hdbscan: Hierarchical density based clustering. J Open Source Softw. 2017;2:205. 10.21105/joss.00205.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="McInnes L, Healy J, Astels S. hdbscan: Hierarchical density based clustering. J Open Source Softw. 2017;2:205. 10.21105/joss.00205."/></mixed-citation></ref><ref id="CR59"><label>59.</label><mixed-citation id="mc-CR59"><named-content content-type="citation-string">Zhang T, Wu Q, Zhang Z. Probable Pangolin Origin of SARS-CoV-2 Associated with the COVID-19 Outbreak. Curr Biol. 2020;30:1346–e13512. 10.1016/j.cub.2020.03.022.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.cub.2020.03.022"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7156161"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32197085"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Zhang T, Wu Q, Zhang Z. Probable Pangolin Origin of SARS-CoV-2 Associated with the COVID-19 Outbreak. Curr Biol. 2020;30:1346–e13512. 10.1016/j.cub.2020.03.022."/></mixed-citation></ref><ref id="CR60"><label>60.</label><mixed-citation id="mc-CR60"><named-content content-type="citation-string">van der Maaten L. Visualizing Data using t-SNE. J Mach Learn Res. 2008;9:2579–605.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="van der Maaten L. Visualizing Data using t-SNE. J Mach Learn Res. 2008;9:2579–605."/></mixed-citation></ref><ref id="CR61"><label>61.</label><mixed-citation id="mc-CR61"><named-content content-type="citation-string">Edgar RC. MUSCLE: a multiple sequence alignment method with reduced time and space complexity. BMC Bioinformatics. 2004;5:113. 10.1186/1471-2105-5-113.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1186/1471-2105-5-113"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC517706"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="15318951"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Edgar RC. MUSCLE: a multiple sequence alignment method with reduced time and space complexity. BMC Bioinformatics. 2004;5:113. 10.1186/1471-2105-5-113."/></mixed-citation></ref><ref id="CR62"><label>62.</label><mixed-citation id="mc-CR62"><named-content content-type="citation-string">Minh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, von Haeseler A, et al. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Mol Biol Evol. 2020;37:1530–4. 10.1093/molbev/msaa015.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/molbev/msaa015"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7182206"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32011700"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Minh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, von Haeseler A, et al. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Mol Biol Evol. 2020;37:1530–4. 10.1093/molbev/msaa015."/></mixed-citation></ref><ref id="CR63"><label>63.</label><mixed-citation id="mc-CR63"><named-content content-type="citation-string">Kalyaanamoorthy S, Minh BQ, Wong TKF, von Haeseler A, Jermiin LS. ModelFinder: fast model selection for accurate phylogenetic estimates. Nat Methods. 2017;14:587–9. 10.1038/nmeth.4285.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/nmeth.4285"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC5453245"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="28481363"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Kalyaanamoorthy S, Minh BQ, Wong TKF, von Haeseler A, Jermiin LS. ModelFinder: fast model selection for accurate phylogenetic estimates. Nat Methods. 2017;14:587–9. 10.1038/nmeth.4285."/></mixed-citation></ref><ref id="CR64"><label>64.</label><mixed-citation id="mc-CR64"><named-content content-type="citation-string">Hoang DT, Chernomor O, von Haeseler A, Minh BQ, Vinh LS. UFBoot2: Improving the Ultrafast Bootstrap Approximation. Mol Biol Evol. 2018;35:518–22. 10.1093/molbev/msx281.
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/molbev/msx281"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC5850222"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="29077904"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Hoang DT, Chernomor O, von Haeseler A, Minh BQ, Vinh LS. UFBoot2: Improving the Ultrafast Bootstrap Approximation. Mol Biol Evol. 2018;35:518–22. 10.1093/molbev/msx281."/></mixed-citation></ref><ref id="CR65"><label>65.</label><mixed-citation id="mc-CR65"><named-content content-type="citation-string">Moreno MA, Holder MT, Sukumaran J. DendroPy 5: a mature Python library for phylogenetic computing. J Open Source Softw. 2024;9:6943. 10.21105/joss.06943.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Moreno MA, Holder MT, Sukumaran J. DendroPy 5: a mature Python library for phylogenetic computing. J Open Source Softw. 2024;9:6943. 10.21105/joss.06943."/></mixed-citation></ref><ref id="CR66"><label>66.</label><mixed-citation id="mc-CR66"><named-content content-type="citation-string">Yu G. Data Integration, Manipulation and Visualization of Phylogenetic Trees. New York: Chapman and Hall/CRC; 2022. 10.1201/9781003279242.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Yu G. Data Integration, Manipulation and Visualization of Phylogenetic Trees. New York: Chapman and Hall/CRC; 2022. 10.1201/9781003279242."/></mixed-citation></ref></ref-list></sec></sec><sec id="_ad93_" xml:lang="en" sec-type="associated-data" disp-level="1"><title>Associated Data</title><sec id="_adsm93_" xml:lang="en" sec-type="supplementary-materials" disp-level="2"><title>Supplementary Materials</title><supplementary-material id="db_ds_supplementary-material1_reqid_" position="float"><media xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="12864_2026_12976_MOESM1_ESM.docx" mimetype="application" mime-subtype="vnd.openxmlformats-officedocument.wordprocessingml.document"><?cloudpmc-path 886e/13425921/809bbac21418/12864_2026_12976_MOESM1_ESM.docx?><?cloudpmc-bucket app?><?size 7546833?><caption><p>Supplementary Material 1.</p></caption></media></supplementary-material><supplementary-material id="db_ds_supplementary-material2_reqid_" position="float"><media xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="12864_2026_12976_MOESM2_ESM.xlsx" mimetype="application" mime-subtype="vnd.openxmlformats-officedocument.spreadsheetml.sheet"><?cloudpmc-path 886e/13425921/b4f3a5fdfe34/12864_2026_12976_MOESM2_ESM.xlsx?><?cloudpmc-bucket app?><?size 366970?><caption><p>Supplementary Material 2.</p></caption></media></supplementary-material></sec><sec id="_adda93_" xml:lang="en" sec-type="data-availability-statement" disp-level="2"><title>Data Availability Statement</title><p>All genome sequences and associated metadata in the outbreak dataset are published in GISAID’s EpiCoV database with the GISAID identifier: EPI_SET_240219yp. It is composed of 347,409 individual genome sequences with collection dates ranging from 2019-06-25 to 2023-05-17; data were collected in 202 countries and territories. To view the contributors of each individual sequence with details such as accession number, Virus name, Collection date, Originating Lab and Submitting Lab and the list of Authors, visit doi:10.55876/gis8.240219ypThe reference sequence for SARS-CoV-2 used in this study is hCoV-19/Wuhan/WIV04/2019 (WIV04), the official reference sequence employed by GISAID (EPI_ISL_402124, https://gisaid.org/WIV04).All GISAID sequences and associated metadata used in the BetaCov dataset are published in GISAID’s EpiCoV database with the identifier: EPI_SET_250709on. It is composed of 3 individual genome sequences with collection dates ranging from 2017 to 2019-06-25; the data was collected in 1 countries and territories. To view the contributors of each individual sequence with details such as accession number, Virus name, Collection date, Originating Lab and Submitting Lab and the list of Authors, visit doi: 10.55876/gis8.250709on. PRIME source code is available under the MIT license in the GitHub repository (https://github.com/lanl/prime).</p></sec></sec></body></article>