<!DOCTYPE article
 PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.2 20190208//EN" "JATS-archivearticle1.dtd">
<article xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="preprint"><?all-math-mml yes?><?use-mml?><?origin ukpmcpa?><front><journal-meta><journal-id journal-id-type="nlm-ta">bioRxiv</journal-id><journal-title-group><journal-title>bioRxiv : the preprint server for biology</journal-title></journal-title-group><issn pub-type="epub">2692-8205</issn></journal-meta><article-meta><article-id pub-id-type="manuscript">EMS215481</article-id><article-id pub-id-type="doi">10.64898/2026.03.02.709004</article-id><article-id pub-id-type="archive">PPR1221387</article-id><article-version-alternatives><article-version article-version-type="status">preprint</article-version><article-version article-version-type="number">1</article-version></article-version-alternatives><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title-group><article-title>Structural Plausibility Without Binding Specificity: Limits of AI-Based Antibody-Antigen Structure Prediction Confidence Scores</article-title></title-group><contrib-group><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-5457-5163</contrib-id><name><surname>Smorodina</surname><given-names>Eva</given-names></name><xref ref-type="aff" rid="A1">1</xref><xref ref-type="corresp" rid="CR1">#</xref><xref ref-type="fn" rid="FN1">☯</xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0004-9022-3896</contrib-id><name><surname>Ali</surname><given-names>Montader</given-names></name><xref ref-type="aff" rid="A2">2</xref><xref ref-type="fn" rid="FN1">☯</xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-1866-4094</contrib-id><name><surname>Brumat</surname><given-names>Klara Kropivšek</given-names></name><xref ref-type="aff" rid="A3">3</xref><xref ref-type="fn" rid="FN1">☯</xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-4444-3885</contrib-id><name><surname>Salicari</surname><given-names>Leonardo</given-names></name><xref ref-type="aff" rid="A4">4</xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0002-6381-8443</contrib-id><name><surname>Miklavc</surname><given-names>Samo</given-names></name><xref ref-type="aff" rid="A5">5</xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0004-1415-0082</contrib-id><name><surname>Kappassov</surname><given-names>Aibek</given-names></name><xref ref-type="aff" rid="A1">1</xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0005-5551-6957</contrib-id><name><surname>Fu</surname><given-names>Chengcheng</given-names></name><xref ref-type="aff" rid="A6">6</xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-6228-2221</contrib-id><name><surname>Sormanni</surname><given-names>Pietro</given-names></name><xref ref-type="aff" rid="A2">2</xref><xref ref-type="aff" rid="A7">7</xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-7729-819X</contrib-id><name><surname>de Marco</surname><given-names>Ario</given-names></name><xref ref-type="aff" rid="A3">3</xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-2622-5032</contrib-id><name><surname>Greiff</surname><given-names>Victor</given-names></name><xref ref-type="aff" rid="A1">1</xref><xref ref-type="aff" rid="A8">8</xref><xref ref-type="corresp" rid="CR1">#</xref></contrib></contrib-group><aff id="A1"><label>1</label>Department of Immunology, <institution-wrap><institution-id institution-id-type="ror">https://ror.org/01xtthb56</institution-id><institution>University of Oslo</institution></institution-wrap> and <institution-wrap><institution-id institution-id-type="ror">https://ror.org/00j9c2840</institution-id><institution>Oslo University Hospital</institution></institution-wrap>, <city>Oslo</city>, <country country="NO">Norway</country></aff><aff id="A2"><label>2</label>Yusuf Hamied Department of Chemistry, <institution-wrap><institution-id institution-id-type="ror">https://ror.org/013meh722</institution-id><institution>University of Cambridge</institution></institution-wrap>, <city>Cambridge</city><postal-code>CB2 1EW</postal-code>, <country country="GB">UK</country></aff><aff id="A3"><label>3</label>Laboratory for Environmental and Life Sciences, <institution-wrap><institution-id institution-id-type="ror">https://ror.org/00mw0tw28</institution-id><institution>University of Nova Gorica</institution></institution-wrap>, <city>Nova Gorica</city>, <country country="SI">Slovenia</country></aff><aff id="A4"><label>4</label><institution-wrap><institution-id institution-id-type="ror">https://ror.org/02f013h18</institution-id><institution>CINECA</institution></institution-wrap>, <city>Bologna</city>, <country country="IT">Italy</country></aff><aff id="A5"><label>5</label>Supercomputing department, IZUM, Maribor, Slovenia</aff><aff id="A6"><label>6</label>College of Engineering, <institution-wrap><institution-id institution-id-type="ror">https://ror.org/01zkghx44</institution-id><institution>Georgia Institute of Technology</institution></institution-wrap>, <city>Atlanta</city>, <state>GA</state>, <country country="US">USA</country></aff><aff id="A7"><label>7</label>Department of Chemical Engineering, <institution-wrap><institution-id institution-id-type="ror">https://ror.org/041kmwe10</institution-id><institution>Imperial College London</institution></institution-wrap>, <addr-line>South Kensington Campus</addr-line>, <city>London</city>, <country country="GB">UK</country></aff><aff id="A8"><label>8</label>Imprint Labs, LLC, New York, NY, USA</aff><author-notes><corresp id="CR1"><label>#</label>Correspondence: <email>victor.greiff@medisin.uio.no</email>, <email>eva.smorodina@medisin.uio.no</email></corresp><fn fn-type="equal" id="FN1"><label>☯</label><p id="P1">Contributed equally to this work</p></fn></author-notes><pub-date pub-type="nihms-submitted"><day>09</day><month>06</month><year>2026</year></pub-date><pub-date pub-type="preprint"><day>03</day><month>03</month><year>2026</year></pub-date><permissions><ali:free_to_read/><license><ali:license_ref>https://creativecommons.org/licenses/by/4.0/</ali:license_ref><license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0 International license</ext-link>.</license-p></license></permissions><abstract><p id="P2">Antibody-antigen binding prediction remains a central challenge for AI-driven therapeutic discovery, particularly in discriminating cognate interactions from structurally plausible but incorrect pairings. We present a controlled, AI-method- and antibody-format-agnostic evaluation framework that measures binding specificity under realistic conditions. Using 106 experimentally determined single-chain antibody (nanobody)-antigen complexes and 11,342 shuffled non-cognate pairings, we benchmarked publicly-available state-of-the-art structure prediction methods (AlphaFold3, Boltz-2, Chai-1). Although the methods tested often generated geometrically plausible complexes, internal confidence metrics (ipTM) frequently failed to discriminate correct from incorrect pairings. Increased sampling improved structural refinement but not pairing discrimination, indicating that computational resources are better allocated across independent seeds and explicit negative controls. We conclude that internal confidence scores are not inherently calibrated to binding specificity and require validation against realistic decoys. To enable community benchmarking and method development, we release ~1.8 million AI-generated complex structures and guidance for the benchmarks ahead.</p></abstract></article-meta></front><body><sec id="S1" sec-type="intro"><title>Introduction</title><p id="P3">Antibodies are key immunotherapeutic biomolecules characterized by antigen-specific binding. This specificity enables antibodies to identify and bind molecular targets such as tumor- or pathogen-associated antigens <sup><xref ref-type="bibr" rid="R1">1</xref></sup>, underpinning a wide range of therapeutic applications and positioning antibodies as the largest class of biotherapeutics <sup><xref ref-type="bibr" rid="R2">2</xref></sup>. With the continued growth of monoclonal antibody (mAb) therapeutics, there is increasing interest in developing in silico antibody discovery and design methods <sup><xref ref-type="bibr" rid="R3">3</xref>–<xref ref-type="bibr" rid="R7">7</xref></sup>. While experimental discovery pipelines still rely heavily on antibody libraries and screening <sup><xref ref-type="bibr" rid="R8">8</xref></sup>, improved prediction of paratope-epitope interactions may enable in silico discovery strategies <sup><xref ref-type="bibr" rid="R3">3</xref></sup>.</p><p id="P4">Antigen recognition is primarily mediated by the antibody’s six hypervariable complementarity-determining region (CDR) loops spread over heavy and light chains that together form the paratope <sup><xref ref-type="bibr" rid="R9">9</xref></sup>. Other classes of antibodies (such as nanobodies also known as VHHs<sup><xref ref-type="bibr" rid="R10">10</xref></sup>) have only 3 CDRs (the heavy chain loops only). These loops engage antigen surface regions (the epitope) through a combination of shape complementarity, physicochemical interactions, and conformational adaptability <sup><xref ref-type="bibr" rid="R11">11</xref>–<xref ref-type="bibr" rid="R14">14</xref></sup>. Although recent advances in deep learning have revolutionized protein structure prediction for many molecular systems <sup><xref ref-type="bibr" rid="R15">15</xref>–<xref ref-type="bibr" rid="R17">17</xref></sup>, accurately predicting antibody/nanobody-antigen complexes in general and identifying the correct paratope-epitope interface in particular remains challenging <sup><xref ref-type="bibr" rid="R18">18</xref>,<xref ref-type="bibr" rid="R19">19</xref></sup>. Previous benchmarking studies report success rates of approximately 20% for antibody-antigen docking using AlphaFold-Multimer (v2.3.0) and Rosetta-based protocols <sup><xref ref-type="bibr" rid="R20">20</xref>–<xref ref-type="bibr" rid="R23">23</xref></sup>. More recent evaluations indicate improved performance <sup><xref ref-type="bibr" rid="R24">24</xref>,<xref ref-type="bibr" rid="R25">25</xref></sup>, with ~35% success using a single stochastic seed and up to ~60% success when extensive sampling (up to 1000 seeds) is combined with confidence-based ranking <sup><xref ref-type="bibr" rid="R15">15</xref>,<xref ref-type="bibr" rid="R26">26</xref></sup>, albeit at substantial computational cost.</p><p id="P5">Recently, several state-of-the-art molecular structure prediction models have emerged, including AlphaFold2/3 <sup><xref ref-type="bibr" rid="R15">15</xref>,<xref ref-type="bibr" rid="R27">27</xref></sup>, Boltz-1/2 <sup><xref ref-type="bibr" rid="R16">16</xref>,<xref ref-type="bibr" rid="R28">28</xref></sup>, and Chai-1/2 <sup><xref ref-type="bibr" rid="R29">29</xref>,<xref ref-type="bibr" rid="R30">30</xref></sup>. Several performance evaluations on diverse benchmarks spanning protein monomers, multimers, and small-molecule interactions have been made <sup><xref ref-type="bibr" rid="R31">31</xref>–<xref ref-type="bibr" rid="R34">34</xref></sup> and suggested that some of these models outperform earlier generations such as, for example, AlphaFold-Multimer <sup><xref ref-type="bibr" rid="R35">35</xref></sup>. However, systematic evaluation of their ability to identify correct antibody-antigen interfaces and their inability to identify incorrect interactomes remains absent. Evaluating these computational complex-prediction models can highlight key metrics that can be used to discriminate cognate binding partners (“real” complexes) from non-cognate ones (“shuffled” complexes, negative controls) <sup><xref ref-type="bibr" rid="R36">36</xref>,<xref ref-type="bibr" rid="R37">37</xref></sup>. This distinction is particularly important in discovery and screening contexts, where large numbers of candidate antibody-antigen pairings must be evaluated to separate binders from non-binders and wrong-epitope binders, with false positives proving to be a persistent challenge <sup><xref ref-type="bibr" rid="R26">26</xref>,<xref ref-type="bibr" rid="R38">38</xref>,<xref ref-type="bibr" rid="R39">39</xref></sup>.</p><p id="P6">To explicitly address the problem of binding partner discrimination, we introduce a benchmarking framework that enables ground truth distinction between real and shuffled nanobody (here: VHH) and antigen complexes (<xref ref-type="fig" rid="F1">Figure 1</xref>). Real complexes are defined as cognate VHH-antigen pairs extracted from experimentally solved structures, where the VHH and antigen are known biological binding partners within the same Protein Data Bank (PDB) entry. In contrast, shuffled non-cognate complexes are artificially generated, non-cognate VHH-antigen pairs created by pairing a VHH sequence from one experimental complex with an antigen sequence from a different experimental complex. These non-cognate shuffled pairs do not correspond to known biological interactions and serve as ground-truth decoys. This benchmark design allows direct comparison of computational predictions across cognate and non-cognate pairings under identical modeling conditions.</p><p id="P7">We benchmarked AlphaFold3, Boltz-2, and Chai-1 for predicting VHH-antigen paratope-epitope interactions, quantifying their ability to distinguish real cognate from shuffled non-cognate pairs. This allowed us to assess structural and interface accuracy (DockQ <sup><xref ref-type="bibr" rid="R40">40</xref>,<xref ref-type="bibr" rid="R41">41</xref></sup>), test whether confidence metrics (e.g., ipTM <sup><xref ref-type="bibr" rid="R35">35</xref>,<xref ref-type="bibr" rid="R42">42</xref></sup>) can discriminate real from shuffled complexes, and identify sequence/structure features linked to high or low confidence independent of pairing correctness. Overall, we present a ground-truth based benchmark pipeline that exposes current model limitations and highlights opportunities to improve antibody/nanobody-antigen complex prediction for practical screening workflows.</p></sec><sec id="S2" sec-type="results"><title>Results</title><sec id="S3"><title>Computational predictors largely fail to discriminate real from “shuffled” nanobody-antigen complexes</title><p id="P8">To assess whether publicly available state-of-the-art structure prediction methods can distinguish biologically observed nanobody-antigen interactions from shuffled complexes, we evaluated three AI-based complex structure prediction tools (AF3, Boltz-2, and Chai-1) across a panel of 561,800 (106<sup>2</sup>×50, 561,800×3=1,685,400 for all 3 tools) VHH-antigen pairings (<xref ref-type="fig" rid="F2">Figure 2</xref>): 106 real systems (complexes) from 91 unique PDB ids and 11,130 (1062=11,236, 11,236-106=11,130) shuffled complexes of 50 replicates each, where a replicate is defined as a prediction sample generated by a model. The selection of 106 real VHH-antigen complexes was based on their sequence and structural diversity, size, and uniqueness (<xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 1</xref>, <xref ref-type="sec" rid="S13">Methods</xref>, <xref ref-type="supplementary-material" rid="SD1">Supplementary Table 1</xref>).</p><p id="P9">The larger number of systems (106) compared to PDB IDs (91) is due to 15 PDB entries containing more than one VHH bound to the same antigen at different epitopes. For all complexes, information on whether they were present in the training set of the respective AI tool was also available (see below). In the following, we assume that VHHs and antigens of “shuffled complexes” are non-binders.</p><p id="P10">First, we found that interface score landscapes (ipTM) largely lack specificity for real complexes (<xref ref-type="fig" rid="F2">Figure 2A</xref>). We intentionally used only 91 unique PDB entries (not 106 systems that include the same antigens with different VHHs) to avoid mixing different analytical levels and to maintain dataset simplicity. This ensured a clean comparison of whether current structure prediction tools can distinguish real from shuffled interactions, without overrepresenting antigen contexts containing multiple VHHs. All-VHH-vs-all-Ags heatmaps of predicted interface confidence scores revealed broadly similar interaction landscapes for real and shuffled complexes across all three tools. AF3 and Chai-1 produced sparse, heterogeneous score patterns with isolated high-scoring interactions distributed throughout the matrix, while Boltz-2 assigned uniformly high scores across most VHH-antigen combinations. While AF3 showed some enrichment of scores along the diagonal of the heatmap (real cognate pairs), none of the methods displayed systematic diagonal enrichment compared to the shuffled complexes, indicating, qualitatively, a lack of specificity for biologically correct pairings.</p><p id="P11">To quantify whether cognate (diagonal) pairs are enriched relative to shuffled (off-diagonal) pairs (<xref ref-type="fig" rid="F2">Figure 2A</xref>), we evaluated each tool using precision-recall (PR) curves and their area under the PR curve (PR-AUC; Average Precision, AP), which are less sensitive to class imbalance than ROC-based metrics and avoid reliance on an arbitrarily chosen cutoff (<xref ref-type="fig" rid="F2">Figure 2B</xref>). Across all three predictors, PR performance remained low. AF3 achieved the highest PR-AUC (AP = 0.187), followed by Chai-1 (AP = 0.067) and Boltz-2 (AP = 0.026). Because positives are rare in this all-vs-all setting, the random baseline in PR space equals the positive prevalence and is correspondingly low (baseline precision ≈ 0.011; ~89-91 positives out of ~8041-8281 evaluated pairs, depending on the tool after excluding missing values). Thus, while AF3’s AP is above baseline, all tools remain far from reliably separating real from shuffled complexes, consistent with the weak diagonal enrichment observed in the heatmaps.</p><p id="P12">For interpretability, we additionally overlaid three reference thresholds <sup><xref ref-type="bibr" rid="R42">42</xref>,<xref ref-type="bibr" rid="R43">43</xref></sup> (ipTM≥0.3, 0.6, 0.8, denoted t1-t3, chosen according to the scoring metric literature <sup><xref ref-type="bibr" rid="R42">42</xref>,<xref ref-type="bibr" rid="R43">43</xref></sup>) onto the PR curves (<xref ref-type="fig" rid="F2">Figure 2B</xref>). These markers illustrate the expected threshold trade-off: lower thresholds increase recall at the expense of precision (many false positives), whereas higher thresholds improve precision but sharply reduce recall (many false negatives). This effect is especially pronounced under extreme class imbalance, making single-threshold summaries (e.g., confusion matrices, F1, specificity) highly threshold-dependent and potentially misleading when comparing tools. In Boltz-2, uniformly elevated scores drive near-maximal recall at low-to-intermediate thresholds but with extremely low precision, reflecting widespread false positives rather than meaningful discrimination. Overall, PR curves and PR-AUC confirm that none of the evaluated tools achieve robust separation of real nanobody-antigen complexes from shuffled mismatches.</p><p id="P13">Next, we asked to what extent scores differed between pairs that belong to the train or test sets within each dataset split (<xref ref-type="fig" rid="F2">Figure 2C</xref>). Dataset splits were defined using the original train-test splits of each tool: complexes were labeled as “train” when both the VHH and antigen were in complex (real) and present in the training set of that model respectively (as of its training set date-cutoff); as “test” when both were in complex (real) but not seen by the model, as “both” when one of the two binding pairs was drawn from the training set and the other from the test set (shuffled). For this analysis, we used the following datasplits: Chai-1- 25 train and 81 test systems; Boltz-2-64 train and 42 test systems; AF3 - 30 train and 76 test systems. Violin plots summarizing ipTM score distributions confirm substantial similarity of interface scores between real and shuffled complexes for all tools and across training, test, and mixed (both) splits. AF3 showed bimodal distributions for real complexes, with a subset of high-scoring interactions, but shuffled complexes still occupy overlapping score ranges. Boltz-2 assigns consistently high ipTM values to both real and shuffled interactions (mean ~0.85-0.91), with minimal separation. Chai-1 yielded uniformly low scores for both classes, again with nearly identical distributions. Across all tools, mean scores and variances were comparable between real and shuffled complexes, indicating that score magnitude alone does not encode interaction authenticity.</p><p id="P14">We used TopModel’s clash score <sup><xref ref-type="bibr" rid="R44">44</xref></sup> to evaluate the structural quality across real and shuffled nanobody-antigen complexes (<xref ref-type="fig" rid="F2">Figure 2D</xref>). The clash score accounts for both the number of steric clashes between all atoms and protein length, with lower values indicating better geometric quality. Across all tools and dataset splits, clash score distributions were broad and overlapped between real and shuffled complexes. For AF3, top-ranked (highest ipTM) models show mean clash scores of 25.7±9.1 (train), 26.7±9.8 (test), and 26.8±9.7 (both) for shuffled complexes, compared with 22.7±9.4, 24.2±9.9, and 26.8±9.7 for real complexes across the same splits. Boltz-2 produced systematically higher clash scores than AF3, with shuffled complexes averaging 30.0±9.5 (train), 36.1±12.4 (test), and 33.0±11.3 (both), while real complexes average 27.9±10.7, 34.8±12.2, and 32.9±11.3, respectively. Chai-1 exhibits markedly poorer structural quality by this metric. Mean clash scores for Chai-1 were 4-5 times higher and more variable, reaching ~80-90 on average, with extreme dispersion (e.g., 62.0±65.0 and 90.7±125.2 in the test split). These elevated and highly variable scores were observed for both real and shuffled complexes, indicating frequent steric clashes in top-ranked predictions and a lack of effective internal filtering for geometric plausibility. Importantly, real complexes did not systematically achieve lower clash scores than shuffled complexes in any tool or dataset split. In several cases, shuffled complexes displayed comparable or even slightly lower average clash scores than real complexes, despite being biologically incorrect.</p><p id="P15">To assess whether different prediction tools agree on which nanobody-antigen pairs are confident or uncertain, we directly compared ipTM score matrices generated by AF3, Boltz-2, and Chai-1 (<xref ref-type="fig" rid="F2">Figure 2A</xref>). Global correlations between flattened score matrices were uniformly low (pearson r=0.14 for Chai-1-Boltz-2, r=0.18 for Chai-1-AF3, and r=0.13 for Boltz-2-AF3), indicating low agreement across tools in their assessment of interaction confidence.</p><p id="P16">Per-system cross-tool correlation analysis, which examines whether tools predict the same systems with similar levels of confidence or lack thereof, further revealed substantial heterogeneity across both VHHs and antigens. For each pair of tools, many systems exhibited weak or even negative correlations, highlighting disagreement in how individual VHHs or antigens are scored (<xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 2A</xref>). Importantly, these discrepancies were not confined to a small subset of problematic systems but were broadly distributed across the benchmark, suggesting that tool-specific inductive biases strongly shape confidence assignment. Outlier analysis reinforces this conclusion. High-confidence “shuffled” outliers identified by one tool rarely overlapped with those from another. Specifically, only four shuffled outliers overlapped between Chai-1 and Boltz-2, fourteen between Chai-1 and AF3, and just one between Boltz-2 and AF3, despite all tools being evaluated on the same set of VHH-antigen pairings. Representative examples illustrate the magnitude of these discrepancies. For instance, the system (4NC2, 7R24) was assigned high confidence by Boltz-2 (ipTM≈0.81) but substantially lower confidence by Chai-1 (ipTM≈0.60), while (9EMY, 8UKV) achieved high confidence in Chai-1 (ipTM≈0.73) but was scored markedly lower by Boltz-2 (ipTM≈0.44). Similarly, several systems such as (7TGF, 6OBO) and (8K4Q, 8EW6) were consistently high-confidence outliers in AF3 (ipTM≥0.8) yet did not emerge as outliers in the other tools (<xref ref-type="supplementary-material" rid="SD3">Supplementary Table 2</xref>).</p><p id="P17">Together, these results show that both ipTM-confidence and structural accuracy remain limited across tools, with AF3 producing the lowest average clash scores, Boltz-2 intermediate values, and Chai-1 the poorest overall structural quality. However, even the best-performing method by this metric fails to reliably discriminate real nanobody-antigen complexes from shuffled mismatches. This inability to distinguish real from shuffled complexes persists across training, test, and combined datasets, indicating that the observed overlap is unlikely to result from overfitting or data leakage. Instead, it reflects a limitation of current structure-based confidence scores and simple structural quality measures, which capture generic interface properties but ultimately fail to reflect biological correctness in nanobody-antigen complex prediction.</p></sec><sec id="S4"><title>Cross-tool score agreement is weak and sampling improves structure but not confidence alignment</title><p id="P18">Next we focused on a thorough evaluation of the real complexes (the diagonal systems of the <xref ref-type="fig" rid="F2">Figure 2A</xref> matrices) against their experimental references. We asked whether models “know” when they fail - that is, whether confidence scores (ipTM) reliably indicate structural prediction quality (DockQ) under different sampling regimes (<xref ref-type="fig" rid="F3">Figure 3</xref>). Confidence calibration - the correspondence between predicted confidence and actual quality <sup><xref ref-type="bibr" rid="R45">45</xref>,<xref ref-type="bibr" rid="R46">46</xref></sup> is critical for practical screening workflows where ground-truth access to structural information is absent. Thus, confidence calibration analysis was only performed on the real complexes, as shuffled pairings lack experimental reference structures against which to compute DockQ <sup><xref ref-type="bibr" rid="R41">41</xref></sup>. We found that confidence calibration differs systematically across tools when we display correlation of of DockQ versus ipTM for the best, initial (sample<sub>0</sub>), and worst of 50 predictions (<xref ref-type="fig" rid="F3">Figure 3A</xref>).</p><p id="P19">To interpret calibration patterns, we defined four diagnostic quadrants based on thresholds of DockQ=0.23 (acceptable quality) and ipTM=0.5 (confident prediction): Q1 (confident and correct), Q2 (overconfident failures), Q3 (uncertain and incorrect), and Q4 (underconfident success) (<xref ref-type="supplementary-material" rid="SD2">Supplementary Table 3</xref>). Among the three methods, AF3 exhibited the strongest calibration, with a pearson correlation of r=0.736 for best-DockQ structures and relatively few overconfident failures (Q2 - 2%, Q4 - 28%). AF3 maintained high correlation even at the single-sample level (r=0.888 for sample<sub>0</sub>), indicating comparatively robust confidence assignment without extensive sampling. In contrast, Boltz-2 was systematically overconfident. Although its best-DockQ calibration remained moderate (pearson r=0.665), a substantial fraction of predictions fell into the overconfident failure regime (Q2 - 18%), where ipTM is high despite poor structural quality. This effect was exacerbated for single-sample predictions, where correlation dropped sharply (pearson r=0.493), highlighting the unreliability of Boltz-2 confidence scores in low-sampling screening workflows. Conversely, Chai-1 exhibited systematic underconfidence. Although a subset of predictions achieved high structural quality (DockQ≈0.4-0.9), these models were frequently assigned low ipTM scores, resulting in a substantial fraction of underconfident successes (Q4; 24%). Consequently, the relationship between confidence and best-case structural quality was weaker for Chai-1 (best-DockQ vs. ipTM, r=0.612), meaning that confidence-based filtering would exclude many structurally accurate predictions.</p><p id="P20">Extreme quadrant cases further illustrate observed failure modes. Overconfident failures (Q2) represented geometrically poor structures assigned high confidence (e.g., AF3 system_24_24_6V7Y), whereas underconfident successes (Q4) correspond to high-quality complexes that would be missed by thresholding (e.g., AF3 system_27_27_7B5G). These cases provide concrete examples of how confidence scores can diverge from structural reality in both directions (See <xref ref-type="supplementary-material" rid="SD2">Supplementary Table 3</xref> for the classification of all systems in the four quadrants).</p><p id="P21">Next, we found that sampling improves structure but not confidence alignment. Across all three tools, saturation sampling improved best-case DockQ (<xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 3A</xref>), confirming that additional sampling can refine geometry. DockQ improvement distributions are right-skewed for all models, with AF3 and Boltz-2 exhibiting the largest absolute gains, while Chai-1 showed smaller but consistent improvements. The proportion of systems achieving acceptable or higher quality (DockQ&gt;0.23) increased substantially after saturation sampling (<xref ref-type="fig" rid="F3">Figure 3B</xref>, <xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 3B</xref>). For AF3, incorrect predictions decreased by nearly 20%, with Boltz-2 showing a comparable reduction in incorrect predictions and a marked increase in medium- and high-quality outcomes; in contrast, Chai-1 remains dominated by incorrect predictions despite sampling. This reflects systems “rescued” above the quality threshold through sampling alone (the “improvers”, <xref ref-type="fig" rid="F3">Figure 3B</xref>). However, the relationship between structural improvement and confidence during saturation sampling remains weak. Whilst all three models showed significant improvements in DockQ scores (Wilcoxon p&lt;0.001; <xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 3C</xref>), paired analyses showed no significant improvement in ipTM for AF3 or Chai-1 (p = 0.179 and p = 0.222, respectively), and although Boltz-2 showed a statistically significant change (p &lt; 0.001),the ipTM score values slightly decreased (<xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 3D</xref>). In particular, correlations between changes in DockQ and changes in ipTM were near zero across all models (r=-0.027, -0.040, and -0.019 for AF3, Boltz-2, and Chai-1, respectively; <xref ref-type="fig" rid="F3">Figure 3C</xref>), indicating that confidence scores were largely insensitive to structural improvements achieved through sampling; they remain “locked-in” to the initial structural hypothesis.</p><p id="P22">To assess whether confidence scores yielded consistent rankings across tools, we compared ipTM values for each system (<xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 3E</xref>). Cross-model correlations were modest (pearson’s r = 0.29-0.39), and importantly, higher confidence in one model did not reliably correspond to higher structural quality, as measured by DockQ. From a screening perspective, this behavior has important implications. Single-sample predictions from Boltz-2 are frequently overconfident relative to achieved DockQ, whereas AF3 displays comparatively better calibration even prior to saturation sampling. Nonetheless, for all tools, confidence does not reliably track structural refinement, limiting its utility as a post-sampling ranking criterion.</p><p id="P23">Finally, we examined prediction consistency by measuring epitope-paratope variation across 15 antigens, each crystallized with two distinct VHHs binding distinct epitopes (<xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 4A</xref>, <xref ref-type="sec" rid="S13">Methods</xref>). This provides a natural test of whether models that succeed with one VHH can generalize to alternative binders of the same antigen. AF3 produced acceptable predictions (DockQ for both pairs &gt;0.23) in 10 out of 15 cases (67%), compared to 6/15 (40%) for Boltz-2 and 2/15 (13%) for Chai-1 (<xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 4B</xref>). Boltz-2 and, especially, Chai-1 more frequently succeeded with only one VHH or failed on both partners (<xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 4A</xref>), suggesting that AF3 better captures the underlying antigen structure independently of the specific paratope geometry.</p><p id="P24">To summarize, we show that confident failures are largely tool-specific rather than reflecting a shared set of universally difficult or ambiguous systems. The low overlap in outlier systems and the weak agreement in confidence-based rankings across predictors highlight the absence of a common confidence landscape, underscoring the risk of relying on any single model’s confidence scores for candidate selection, particularly after saturation sampling, where structural quality improves, but confidence remains misaligned.</p></sec><sec id="S5"><title>Structural accuracy scores are uniformly poor, while interface contacts do not reflect specificity</title><p id="P25">We next evaluated whether predicted nanobody-antigen complexes recovered the experimentally defined antigen epitope and whether this information can discriminate real from shuffled pairings (<xref ref-type="fig" rid="F4">Figure 4</xref>). DockQ provides a stringent, structure-level measure of docking correctness by comparing a predicted complex to the experimentally solved reference, and therefore is only defined for cognate (“real”) complexes that have a ground-truth structure. In our analysis, DockQ can be performed only on real pairs, whereas shuffled pairs lack a native reference against which to compute DockQ. However, DockQ can remain low even when a model contacts a broadly correct antigen surface region, because it penalizes global pose errors and interface misalignment. To evaluate interface recovery in a way that (i) is directly interpretable in terms of epitope footprint and (ii) can be applied uniformly across both real and shuffled pairings, we therefore quantify “epitope recall” as the fraction of experimentally defined antigen epitope residues recovered by the predictions, counting a residue as recovered only if it forms a VHH contact in at least 50% of stochastic replicate predictions (n=50) (<xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 2B</xref>). This ensemble-consensus definition intentionally emphasizes reproducible contact preferences over single-sample noise, but it also collapses geometry to a residue set and does not penalize additional (non-epitope) contacts. High epitope recall does not necessarily imply a correct docking mode, but provides insight into what epitope regions are relatively more likely to be hit in predicted complexes. Epitope recall was quantified as the fraction of experimental epitope residues recovered by model predictions, where a residue was considered recovered if it contacted the VHH in at least 50% of stochastic replicate predictions (n=50). Across all three tools, epitope recall values were generally low and sparsely distributed, with substantial overlap between real and shuffled complexes (<xref ref-type="fig" rid="F4">Figure 4A</xref>, <xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 2B</xref>). Although real complexes lie along the diagonal of the epitope recall matrices, off-diagonal (shuffled) complexes frequently exhibited comparable levels of epitope recovery. This indicates that models often identify plausible antigen-contact regions even in non-cognate pairings, limiting the specificity of epitope-based discrimination. Consistent with this observation, direct comparison of epitope recall distributions showed that shuffled complexes frequently overlapped or exceeded the recall observed for real complexes across all models and data splits (<xref ref-type="fig" rid="F4">Figure 4B</xref>). Focusing on the cognate pairings and whether the respective models recovered at least one residue from the true epitope (epitope recall &gt;0), AF3 hit the most epitopes (at least one residue in the ground truth experimental epitope) in the test set of complexes (n=29 out of 76), followed by Chai-1 (n=26 out of 81) and Boltz-2 (n=17 out of 42).</p><p id="P26">Epitope recall scores were then correlated with ipTM scores of each model respectively (on the cognate pairings alone) to assess whether each model’s internal confidence tracks interface recovery. Across cognate pairs, ipTM was only weakly associated with epitope recall for AF3 (pearson and spearman of 0.18 and 0.15 respectively) and Boltz-2 (pearson and spearman of 0.08 and 0.14 respectively), indicating that high ipTM scores frequently occur even when the predicted interface fails to recover the experimental epitope. In contrast, Chai-1 exhibited a moderate positive relationship between ipTM and epitope recall (pearson and spearman 0.46 and 0.48 respectively), suggesting the model’s confidence signal better reflects antigen side-contact recovery than the other two models. Overall, the dispersion of the correlations implies that ipTM is at best a coarse proxy for specificity-relevant interface recovery across different models rather than a reliable discriminator on its own.</p><p id="P27">To assess the impact of data leakage, we examined epitope recall distributions for real complexes split into train-leakage and true test sets (<xref ref-type="fig" rid="F4">Figure 4B</xref>). Across models, train-leakage complexes showed slightly higher median epitope recall than true test complexes, consistent with partial memorization or bias toward known interfaces. However, the overall distributions remained broad, and substantial overlap persisted between train-leakage and true test predictions. Notably, shuffled complexes spanned similar recall ranges, indicating that leakage alone does not explain the limited discriminatory power of epitope recall. Thus, high epitope recall does not reliably indicate cognate binding even in the absence of explicit train/test overlap.</p><p id="P28">We next tested whether increasing stochastic sampling improves epitope discrimination by assessing epitope enrichment significance as a function of replicate count (<xref ref-type="fig" rid="F4">Figure 4C</xref>). Enrichment significance was computed using a one-sided binomial test for each VHH-antigen pair. Let <italic>N</italic> denote the number of stochastic structure predictions and <italic>K</italic> the number of predictions containing at least one epitope contact. Under a null model in which contacts occur uniformly across the antigen surface with probability <italic>p<sub>epitope</sub></italic> (estimated from the fraction of solvent-accessible antigen residues (rSASA) belonging to the epitope), the probability of observing <italic>K</italic> or more epitope hits is given by the binomial survival function <italic>P</italic>(<italic>X</italic> ≥ K | X ~ Binomial(N,p<sub>epitope</sub>)). Note that this is a deliberately coarse null. In practice, VHH–antigen contacts are spatially heterogeneous due to local geometry and physicochemical ‘hotspots’ (e.g., concavities, charge patches, domain accessibility), so a uniform-over-SASA model can miscalibrate absolute enrichment p-values. We therefore interpret these enrichment statistics primarily as an operational, antigen-normalized measure of epitope-seeking behavior, rather than a physically faithful generative null. Because the same null is applied to real and shuffled pairs for a given antigen, this limitation is less likely to explain the similarity of the saturation curves between real and shuffled cohorts, but it does limit interpretation of FDR as a calibrated significance level.</p><p id="P29">Resulting p-values were corrected across all VHH-antigen pairs using the Benjamini-Hochberg procedure, and pairs with FDR &lt; 0.05 were considered significantly enriched. To assess the effect of ensemble size, we recomputed binomial p-values for hypothetical replicate counts <italic>m</italic> by rescaling the observed hit fraction (K/N) and repeating the same multiple-testing correction. This analysis also treats stochastic replicates as independent Bernoulli trials with a common hit probability. Diffusion replicates may exhibit non-trivial correlation because they share the same conditioning signal (sequence/MSA/embeddings) and differ only by sampling noise, reducing the effective sample size relative to N. As a result, p-values (and the extrapolation to larger m) may be optimistic, and therefore emphasize relative, within-model comparisons and the qualitative saturation behavior rather than interpreting FDR thresholds as strictly calibrated. For all three models, the fraction of VHH-antigen pairs declared significantly enriched increased rapidly with the number of stochastic predictions and was saturated by approximately 10-15 replicates. Importantly, real and shuffled complexes exhibited highly similar saturation behavior, regardless of whether the shuffled pairs involved train leakage of the VHH, the antigen, both, or neither. This indicates that increasing ensemble size primarily reinforces existing contact preferences rather than sharpening discrimination between cognate and non-cognate interactions. AF3 and Boltz-2 retained a modest residual separation between real and shuffled complexes across replicate counts, with real pairs showing slightly higher significant-enrichment rates than any shuffled cohort. In contrast, Chai-1 showed near-complete convergence: shuffled complexes rapidly achieved significance frequencies comparable to (or exceeding) those of real complexes, consistent with an epitope-seeking prior under this enrichment definition that limits the interpretability of enrichment-based metrics for this model. Notably, the leakage-stratified shuffled cohorts largely overlapped within each model, further indicating that these effects are not primarily driven by explicit train/test leakage but instead reflect intrinsic model priors that are amplified by ensembling.</p><p id="P30">Epitope recall of each model was then correlated against the ipTM values of the predicted, real, complexes. <xref ref-type="supplementary-material" rid="SD4">Figure S2B</xref> compares each model’s ipTM (diagonal) to system-level epitope-recall (diagonal). AF3 and Boltz-2 show weak association (pearson r≈0.18 and 0.08; spearman ρ≈0.15 and 0.14 respectively), whereas Chai-1 shows a moderate correlation (r≈0.46; ρ≈0.48). However, ipTM is designed as a global complex-confidence measure rather than an epitope-localization score, so high ipTM can occur even when the interface is confidently placed on the wrong antigen region. Moreover, ipTM is range-restricted for many systems (notably Boltz-2), and epitope-recall is bounded with a substantial mass near zero, making linear trendlines and correlation coefficients descriptive rather than calibrated. Additionally, since multiple pairs can share the same antigen (and/or related antigens), observations are not strictly independent and correlation magnitudes should be interpreted qualitatively.</p><p id="P31">Finally, we examined the relationship between structural consistency across predictions and epitope recovery within each model (<xref ref-type="fig" rid="F4">Figure 4D</xref>). Structural consistency was quantified as the mean pairwise RMSD across n=50 stochastic replicates for each complex. Across all models and data splits, epitope recall was negatively correlated with replicate RMSD, indicating that complexes predicted with consistent orientations tend to recover a larger fraction of epitope residues. However, this relationship holds for both train-leakage and true test complexes, and does not differentiate real from shuffled pairings. pearson correlation coefficients ranged from r=-0.22 to -0.30 for AF3, r=-0.14 to -0.27 for Boltz-2, and r=-0.39 to -0.59 for Chai-1, demonstrating that while structural consistency improves epitope recovery, it does not guarantee biological correctness. It is worth noting that RMSD distributions for 50 replicas of real and shuffled complexes show slightly higher values for the shuffled complexes than for the real complexes (<xref ref-type="fig" rid="F4">Figure 4E</xref>), suggesting a potential metric for distinguishing cognate and non-cognate interfaces.</p><p id="P32">Together, these analyses show that epitope recovery and enrichment metrics are driven primarily by generic contact biases and structural consistency rather than by interaction specificity. Increasing stochastic sampling improves apparent epitope enrichment but does so similarly for real and shuffled complexes, limiting its utility as a discriminative signal. These findings reinforce the conclusion that current structure prediction tools preferentially identify plausible antigen contact regions, but struggle to distinguish true cognate paratope-epitope interactions from non-cognate alternatives.</p></sec><sec id="S6"><title>Computational cost and sampling efficiency differ substantially across structure prediction tools</title><p id="P33">Using ML structure prediction tools for high-throughput antibody/nanobody screening requires understanding not only prediction accuracy but also computational cost. With GPU resources unevenly distributed across institutions <sup><xref ref-type="bibr" rid="R47">47</xref></sup> and growing concern over the energy footprint of large-scale computation <sup><xref ref-type="bibr" rid="R48">48</xref>,<xref ref-type="bibr" rid="R49">49</xref></sup>, identifying cost-effective sampling strategies has practical importance. Saturation sampling and seed selection can both improve prediction quality <sup><xref ref-type="bibr" rid="R15">15</xref>,<xref ref-type="bibr" rid="R26">26</xref>,<xref ref-type="bibr" rid="R31">31</xref>,<xref ref-type="bibr" rid="R50">50</xref></sup>, yet their relative contributions and associated costs remain uncharacterized. We therefore evaluated each system (n=106, 91 unique PDBs; no deposition-year bias observed; <xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 6A</xref>) across five random independent seeds with increasing numbers of diffusion samples (N = 1, 10, 25, 50, and 100, around 100,000 additional structures in total), monitoring GPU energy consumption throughout (<xref ref-type="fig" rid="F5">Figure 5</xref>). For this analysis, we additionally included Boltz-1, the predecessor to Boltz-2, to check the consistency of saturation effects across model versions. Because each saturation level uses an independent seed, this design enables analysis of combined seed and saturation effects; extracting only the first sample from each seed isolates seed-dependent variation.</p><p id="P34">To characterize how computational cost scales with sampling depth, we computed the median energy consumption across all systems at each saturation level (N = 1, 10, 25, 50, and 100). Energy consumption varied markedly between tools, revealing distinct cost profiles. At the highest saturation level (N=100 samples), median energy usage per system ranged from 23.0 Wh for AF3 to 82.9 Wh for Chai-1, corresponding to a 3.6-fold difference (<xref ref-type="fig" rid="F5">Figure 5A</xref>, <xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 5A-C</xref>). Boltz-1 and Boltz-2 exhibited intermediate energy-usage profiles, with substantially lower cost than Chai-1 at high saturation but higher baseline costs than AF3. When aggregated across all 106 real systems and saturation levels, total energy consumption ranged from 13.6 kWh for AF3 to 22.7 kWh for Chai-1.</p><p id="P35">These differences reflect different architectural choices. AF3 incurred a high baseline energy cost for MSA computation (82.4 Wh per system, computed once regardless of sampling depth), but showed relatively low marginal cost for additional diffusion samples. In contrast, Chai-1 exhibited a low baseline cost (6.6 Wh) but a steep increase in energy consumption with increasing saturation, consistent with its embedding-based workflow that omits MSA computation (<xref ref-type="sec" rid="S13">Methods</xref>). Boltz-based models showed intermediate behavior. Across all tools, both energy consumption and time scaled approximately linearly with system length, with pearson correlation coefficients ranging from r=0.60 for AF3 to r=0.98 for Boltz-1 (<xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 5D,E</xref>).</p><p id="P36">Increasing saturation sampling improved prediction quality for all models with diminishing returns, as measured by DockQ (<xref ref-type="fig" rid="F5">Figure 5B</xref>). We quantified this by identifying, for each system, the maximum DockQ achieved at each saturation level, then computing the median across systems. Median maximum DockQ increased from baseline (N=1) to N=100 by Δ=+0.43 for AF3 (0.24 to 0.68), Δ=+0.23 for Boltz-2 (0.57 to 0.80), Δ=+0.20 for Boltz-1 (0.06 to 0.26), and Δ=+0.15 for Chai-1 (0.04 to 0.19). While all models benefited from increased sampling, the magnitude of improvement differed substantially, with AF3 showing the largest gains.</p><p id="P37">Marginal quality gain was computed as the per-system difference in maximum DockQ between consecutive saturation levels, then summarized as the median across systems. Across models, the largest marginal improvements occurred at low saturation levels. The transition from baseline to N=10 samples captured the majority of achievable quality improvement, while additional sampling beyond N=25 yielded progressively smaller gains (<xref ref-type="fig" rid="F5">Figure 5C</xref>). This pattern is reflected in marginal ΔDockQ per sampling transition, which peaks at the base to 10 transitions across all tools and declines sharply thereafter. Consistent with this pattern, the proportion of incorrect predictions (DockQ&lt;0.23) decreased sharply from baseline to N=10 but showed more modest (~5-15%) reductions at higher saturation levels (N=25-100, <xref ref-type="fig" rid="F5">Figure 5D</xref>).</p><p id="P38">Analyzing cumulative DockQ improvement as a function of cumulative energy expenditure revealed a consistent “efficiency frontier” across tools (<xref ref-type="fig" rid="F5">Figure 5C</xref>). All models exhibited steep initial slopes, indicating high quality gains per unit energy at low sampling depths, followed by pronounced flattening as energy expenditure increased. Despite differing absolute costs, the efficiency frontiers converge across tools, indicating that early sampling dominates in terms of quality gains regardless of architecture. These trends indicate that N~10-25 samples capture most of the attainable quality improvement at a fraction of the computational cost required for N=100, regardless of model architecture.</p><p id="P39">Beyond sampling depth, seed choice had a substantial independant effect on prediction outcomes (<xref ref-type="fig" rid="F5">Figure 5E</xref>). To isolate seed effects from saturation, we extracted only the first sample (sample<sub>0</sub>) from each seed and computed the per-system DockQ range (maximum minus minimum across seeds). The median range was 0.04-0.05, though extreme cases spanned nearly the full quality spectrum (e.g. system_70_70_8EN2: DockQ 0.01-0.93 across five different AF3 seeds). Across models, approximately 10-15% of systems were strongly seed-sensitive and benefited disproportionately from deeper saturation sampling (<xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 6D, Supplementary Figure 7A</xref>).</p><p id="P40">Cross-seed trajectory analysis further showed that deeper sampling can partially rescue poor initial seed choices. For each seed, we tracked cumulative maximum DockQ (the best quality achieved up to each sample number) and computed the median across systems (<xref ref-type="fig" rid="F5">Figure 5F</xref>). While seeds producing high-quality structures early tended to maintain their advantage throughout saturation, seeds that began in low-quality regions could still achieve substantial improvements with deeper sampling, though typically plateauing at lower absolute quality than favorable seeds (<xref ref-type="fig" rid="F5">Figure 5F</xref>, <xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 7B</xref>). Together, these results indicate that seed selection and saturation sampling act as orthogonal optimization mechanisms: seed choice determines which region of the solution landscape is explored, while saturation sampling refines predictions within that region and can partially mitigate but rarely fully overcome suboptimal initial trajectories (~85-90% of systems remained below their best-seed ceiling).</p><p id="P41">Finally, we assessed whether confidence metrics reflect structural improvements achieved through saturation sampling. Across all models, changes in ipTM from baseline to N = 100 showed near-zero correlation with corresponding changes in DockQ (pearson r=0.00-0.16; <xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 6B,C</xref>). This indicates that confidence scores are largely determined by properties of the initial prediction and remain insensitive to substantial gains in quality from additional sampling. Reliable confidence tracking would enable early stopping without ground-truth structures, but the observed ipTM-DockQ decoupling precludes such use. This behavior mirrors the DockQ-ipTM misalignment observed above, reinforcing that current confidence metrics do not reflect convergence toward higher-quality docking solutions.</p><p id="P42">Together, these results demonstrate that moderate sampling (N~10-25) captures most achievable quality gains at a fraction of full saturation cost, seed selection and saturation act as orthogonal optimization mechanisms, and current model confidence metrics fail to track these improvements for VHH-antigen predictions.</p></sec></sec><sec id="S7" sec-type="discussion"><title>Discussion</title><p id="P43">This work benchmarks whether modern AI complex-prediction tools can discriminate cognate nanobody-antigen binding and recover correct paratope-epitope interfaces. Our central findings are: across tools, high-level confidence scores frequently fail to separate real from mismatched (“shuffled”) complexes, and sampling can improve best-case geometry without resolving the upstream problem of selecting the correct binding mode. Importantly, even when structural quality improves substantially under aggressive sampling (as, for example, measured by DockQ for real complexes) model confidence scores often remain unchanged, revealing a disconnect between structural refinement and confidence calibration. Below, we place these results in the broader “AI biologics” landscape and outline practical implications for drug discovery and model development.</p><sec id="S8"><title>Where the field stands: disconnect between de novo generation and de novo prediction</title><p id="P44">The AI biologics ecosystem is now shaped by strong competition among major players (e.g., Latent, Chai, Nabla, Isomorphic, Boltz), with rapid iteration cycles. Although de novo generation success rates are increasing into the double digit realm (~10-15% on hard interface tasks <sup><xref ref-type="bibr" rid="R30">30</xref>,<xref ref-type="bibr" rid="R51">51</xref></sup>), our results indicate that there is a disconnect between de novo design (task: “generate an antibody for a given target without exploiting other information”) and de novo prediction (task: “indicate all the antigens that can be bound by a given antibody, and vice-versa”). Our “real vs shuffled” discrimination benchmark shows that models can generate interfaces that appear plausible, albeit incorrect, across many pairings, resulting in an abundance of false positives that are challenging to triage.</p><p id="P45">An emerging industry screening workflow is: generate thousands to millions of candidates, then filter using internal confidence scores (e.g., pLDDT, ipTM, PAE). Our results suggest that this strategy is fragile for antibody-antigen binding, because confidence does not necessarily equate to accurate biology. In our many-VHH-versus-many-antigen screening setting, shuffled complexes often reach confidence levels comparable to cognate pairs, while genuinely high-quality structures can remain underconfident depending on the prediction tool. This breaks the assumption that “high-confidence docking implies correct binding” and it helps explain why confidence-guided selection can produce many experimental failures even when predicted structures look polished. Consistent with recent proposals for alternative evaluation metrics <sup><xref ref-type="bibr" rid="R52">52</xref>,<xref ref-type="bibr" rid="R53">53</xref></sup> our results underscore the need for “better scores” that more directly reflect biological meaning rather than internal structural self-consistency.</p><p id="P46">Crucially, drug viability depends on properties that these scores do not directly encode: induced fit compatibility, entropic penalties, off-target propensity, developability constraints, and the ability to tolerate antigen dynamics or conformational selection <sup><xref ref-type="bibr" rid="R54">54</xref>,<xref ref-type="bibr" rid="R55">55</xref></sup>. Confidence metrics were not designed as surrogates for molecular functionality, and our data reinforce (with regard to off-target binding potential, for example) that treating them as such inflates false positives, especially in interface-dependent problems.</p></sec><sec id="S9"><title>Why evolutionary signals and sampling help less than hoped for antibodies</title><p id="P47">A common intuition is that better MSAs and more evolutionary information should improve AI docking and design. However, both antibody and nanobody paratopes (especially CDR loops) represent the most variable regions in biology, and much of the binding specificity arises from flexibility, conformational diversity, and context-dependent loop rearrangements rather than sequence conservation <sup><xref ref-type="bibr" rid="R56">56</xref>–<xref ref-type="bibr" rid="R58">58</xref></sup>. This is consistent with the idea that the utility of MSAs in antibody-focused tasks is more nuanced than in general protein structure prediction. While MSAs are clearly critical for accurate monomer modeling (removing them in AF3 leads to large degradations in structural accuracy <sup><xref ref-type="bibr" rid="R59">59</xref></sup>), their contribution to antibody-antigen docking is constrained by the high sequence and conformational diversity of CDRs, which limits the availability of informative evolutionary signals. This suggests that MSA-derived confidence estimates may be inherently less reliable for antibody-antigen interfaces than for other protein complexes <sup><xref ref-type="bibr" rid="R60">60</xref></sup>. As a result, alternative MSA-strategies could provide better guidance; especially in cases of evolutionary complex/highly mutating proteins such as antibodies.</p><p id="P48">On the other side, sampling helps, but exposes a deeper bottleneck: path selection vs refinement. Saturation/diffusion sampling improved best-case DockQ in many systems, confirming that sampling is valuable as a refinement mechanism. However, we observe essentially zero correlation between changes in DockQ (ΔDockQ) and changes in ipTM (ΔipTM) across all three models (pearson’s r=-0.03, -0.04, and -0.02 for AF3, Boltz-2 and Chai-1), indicating that confidence scores do not track actual structural improvement. The persistence of “non-improver” systems supports a two-stage view: i) path/seed selection determines the qualitative docking mode (often wrong) and ii) saturation refines within that chosen mode.</p><p id="P49">In this sense, current models appear to have a “fixed mindset”: once a confidence level is assigned to a given seed or trajectory, it remains largely invariant, even when the resulting structure improves substantially. This likely reflects an architectural constraint: confidence scores (like ipTM) are derived from MSA-based pairwise representations computed in the trunk network, which remains fixed during diffusion sampling <sup><xref ref-type="bibr" rid="R61">61</xref></sup>. Indeed, MSA representations can be directly optimized via gradient descent to manipulate confidence outputs <sup><xref ref-type="bibr" rid="R61">61</xref>,<xref ref-type="bibr" rid="R62">62</xref></sup>, confirming this upstream dependency. If the model commits early to an incorrect binding mode, additional diffusion samples often cannot rescue it. This has immediate practical implications: “just sample more” may be an expensive way to get diminishing returns, and large seed sweeps, while sometimes recommended, can be prohibitive at screening scale-without guaranteeing specificity or providing confidence-based validation that sampling actually helped. Development of confidence measures sensitive to binding mode quality-or orthogonal strategies such as consensus across independent seeds-remains an important challenge for the field.</p></sec><sec id="S10"><title>Structural biology perspective: flexibility is not optional for paratope-epitope recovery</title><p id="P50">Antibody recognition is not merely geometric matching; it is often a dynamic process where CDR loops adapt to the antigen surface and exploit transient conformations <sup><xref ref-type="bibr" rid="R63">63</xref></sup>. This is particularly relevant for targets with disordered regions or multiple accessible states. Therefore, a major unresolved challenge is distinguishing hallucinated (geometrically plausible but physically or functionally implausible) complexes from viable ones that can maintain binding under realistic dynamics.</p><p id="P51">This points to a missing ingredient in many AI pipelines: a standardized representation of flexibility and experimentally grounded dynamic behavior. We argue that molecular dynamics (MD)-informed filtering is a promising bridge between AI-generated hypotheses and functional plausibility <sup><xref ref-type="bibr" rid="R4">4</xref>,<xref ref-type="bibr" rid="R64">64</xref>,<xref ref-type="bibr" rid="R65">65</xref></sup>. However, today’s AI-generated “MD-like” outputs are often not physically faithful enough to serve as ground truth. These insights suggest that future antibody-antigen discovery pipelines should move beyond static confidence metrics toward standardized, dynamics-aware frameworks, such as MD-informed benchmarks and datasets like DINO <sup><xref ref-type="bibr" rid="R65">65</xref></sup>, that prioritize biophysical plausibility, interface stability, and functional robustness to enable efficient selection of truly viable therapeutic candidates.</p></sec><sec id="S11"><title>Limitations and outlook</title><p id="P52">Our study focuses on VHH-antigen complexes under a controlled real-versus-shuffled pairing design, which is intentionally challenging: it tests specificity (off-diagonal), not just docking plausibility (on-diagonal). In other words, our benchmark quantifies the extent to which generalized (or universal) binding prediction is feasible. That said, our study may overestimate performance limitations in narrower settings where the antigen epitope is known, constraints are available, or the target is rigid and structured. Furthermore, some of the assumed non-binding shuffled complexes may indeed be binding (off-target effect is typically less than 1% <sup><xref ref-type="bibr" rid="R66">66</xref></sup>). Additionally, similar benchmarking should be extended to more complex antibody formats to assess whether these findings generalize beyond VHHs.</p><p id="P53">Recent large-scale resources such as AbSet <sup><xref ref-type="bibr" rid="R53">53</xref></sup> (a dataset of over 800,000 antibody structures including both experimental and computational models) highlight an additional challenge: the sheer volume of predicted antibody structures does not imply reliability. Systematic assessment of how many computational antibody structures in such datasets are structurally sound, interface-correct, or functionally meaningful remains limited. Our results are consistent with the view that antibody structure prediction, particularly in complex with antigens, is not yet a solved problem. Quantifying the extent to which prediction performance is a function of structural data quality in train and test sets remains to be investigated.</p><p id="P54">Another important limitation concerns the structural evaluation metrics employed in this study. While DockQ provides an established composite measure of interface quality, it remains a geometry-centric metric and does not directly assess specificity, energetic plausibility, or functional robustness. Future work should therefore systematically investigate alternative and complementary interface-focused measures <sup><xref ref-type="bibr" rid="R36">36</xref></sup>, such as ipSAE <sup><xref ref-type="bibr" rid="R61">61</xref></sup>, FNAT (fraction of native contacts) <sup><xref ref-type="bibr" rid="R67">67</xref></sup>, CAPRI-style classifications <sup><xref ref-type="bibr" rid="R68">68</xref></sup>, interface RMSD variants <sup><xref ref-type="bibr" rid="R69">69</xref></sup>, and other contact-overlap <sup><xref ref-type="bibr" rid="R56">56</xref></sup> or energy-informed metrics <sup><xref ref-type="bibr" rid="R70">70</xref>–<xref ref-type="bibr" rid="R74">74</xref></sup>, to determine whether they better correlate with cognate discrimination and biological plausibility than trunk-derived confidence scores (e.g., ipTM). A broader metric landscape may reveal evaluation signals that are more sensitive to binding mode correctness or to subtle interface rearrangements that are not captured by current confidence outputs.</p><p id="P55">Looking forward, progress will likely require: (1) interface-specific training objectives that incorporate hard mutation-diverse negatives (shuffled non-cognate pairings) across several affinity ranges <sup><xref ref-type="bibr" rid="R75">75</xref>–<xref ref-type="bibr" rid="R79">79</xref></sup>, (2) confidence estimates calibrated to specificity and functional plausibility rather than structural self-consistency, and (3) standardized dynamic datasets <sup><xref ref-type="bibr" rid="R65">65</xref></sup> enabling models and filters to account for flexibility and induced fit. The most impactful near-term improvement may not be a new confidence score, but a robust, scalable post-prediction filtering layer grounded in biophysics and dynamics.</p></sec><sec id="S12"><title>Guidance on sampling, compute-accuracy tradeoffs, and confidence calibration</title><p id="P56">Our results have immediate consequences for how AI-based structure prediction is operationalized in antibody and antibody discovery pipelines. In current industry practice, large libraries of candidates are frequently filtered using internal confidence scores (e.g., ipTM, PAE, or pLDDT), often combined with extensive stochastic sampling to improve predicted geometry. However, the analyses presented here indicate that such workflows risk conflating structural plausibility with binding specificity. Because shuffled non-cognate complexes frequently achieve confidence scores comparable to real complexes, and because confidence remains largely insensitive to structural improvements achieved through sampling, we argue that future screening strategies must explicitly decouple geometric refinement from interaction prioritization.</p><p id="P57">A key practical implication concerns the role of stochastic sampling. Across all models, moderate sampling improved best-case structural quality, confirming that diffusion sampling is valuable as a refinement mechanism. Yet additional samples primarily refine an already selected structural hypothesis rather than enabling discrimination between cognate and non-cognate partners. Note that the reported gains reflect best-of-ensemble outcomes (i.e., selecting the highest-quality structure among samples using ground-truth evaluation), which is not directly available in prospective pipelines. Deeper sampling only improves decision-making to the extent that downstream filters can reliably identify the better mode among the sampled structures; otherwise, additional samples mainly increase the number of plausible-looking false positives. In practice, this suggests that sampling should be used as a mode-exploration step rather than as a simple confidence amplifier. For high-throughput workflows, shallow ensembles (≈10-25 stochastic predictions) capture most achievable DockQ improvement at a fraction of the computational cost of deep saturation runs, while larger ensembles mainly yield diminishing returns. Importantly, seed choice and sampling depth act as orthogonal optimization mechanisms: independent seeds explore different regions of the solution landscape, whereas deeper sampling refines within those regions. Consequently, distributing computational budget across multiple seeds is often more informative than increasing sampling depth within a single trajectory. Treating sampling as a tool to detect structural convergence rather than as a guarantee of specificity may reduce unnecessary GPU expenditure without compromising structural insight.</p><p id="P58">These observations also inform compute-accuracy tradeoffs in large-scale biologics programs. The efficiency frontiers observed across tools show steep early gains in quality per unit energy, followed by rapid flattening at higher sampling depths. From a discovery perspective, this implies that the most cost-effective strategy is a staged workflow: early, low-cost ensembles to identify stable interface modes, followed by selective deeper sampling only for candidates that satisfy additional biological or experimental constraints. Because confidence scores do not track structural refinement, escalating compute solely to improve ipTM values is unlikely to yield more reliable candidate prioritization. Instead, compute allocation should be conditioned on whether additional sampling changes the structural hypothesis (for example, by resolving competing interface clusters or rescuing seed-sensitive systems) rather than on absolute confidence thresholds.</p><p id="P59">Perhaps the most consequential implication relates to confidence calibration. Our benchmark demonstrates that ipTM scores cannot be interpreted as probabilities of correct binding or as reliable indicators of specificity in realistic discovery settings. More precisely, our analyses show that these scores are not calibrated to specificity (cognate vs non-cognate binding) in this discovery-like setting. Converting ipTM (or any internal score) into a probability of correct binding would require target- and program-specific calibration data with explicit negatives and, ideally, affinity/function labels; such a mapping is not identified by structure-only benchmarking. Any learned calibration is also likely to drift across model releases and prompting/inference settings, motivating routine re-calibration within each program. Overconfident failures, underconfident successes, and weak cross-tool agreement all highlight that current confidence metrics capture internal structural consistency rather than biological correctness <sup><xref ref-type="bibr" rid="R80">80</xref></sup>. For pharma-facing workflows, this necessitates a shift from absolute confidence thresholds toward relative or context-aware calibration. One practical approach is to compare each candidate against a panel of negative controls (such as shuffled pairings, homologous off-targets, or unrelated antigens) and to evaluate whether predicted interfaces are exceptional relative to these decoys. Under such a framework, confidence becomes a comparative signal embedded within a program-specific landscape rather than a universal scalar metric. This paradigm aligns with our real-versus-shuffled benchmarking strategy and reflects the reality that specificity can only be assessed in the presence of realistic alternatives.</p><p id="P60">Beyond score calibration, ensemble-derived signals may provide a more robust basis for decision-making. Structural consistency across stochastic predictions, dominance of a single interface cluster, and reproducibility of epitope contacts emerge as practical indicators of hypothesis stability, even though they do not by themselves guarantee biological correctness. Incorporating these ensemble-level features into downstream filtering layers may help distinguish fragile, hallucinated complexes from hypotheses worthy of experimental follow-up. Notably, cross-tool agreement on interface mode rather than agreement on raw confidence magnitude may represent a stronger indicator of robustness, given the low correlation in ipTM landscapes observed here.</p><p id="P61">Taken together, these considerations suggest a reframing of AI-assisted antibody discovery workflows. Rather than relying on confidence scores as standalone decision metrics, future pipelines may benefit from explicitly separating three conceptual layers: (i) geometry confidence, describing structural self-consistency, (ii) mode confidence, reflecting ensemble convergence toward a stable interface hypothesis, and (iii) specificity confidence, derived from comparative evaluation against realistic decoys or alternative partners. Such a multi-layered framework acknowledges the limitations identified in this study while preserving the strengths of modern structure prediction tools as generators of plausible structural hypotheses. In the near term, the most impactful improvements may therefore arise not from deeper sampling or higher raw confidence scores, but from better-calibrated filtering strategies that integrate ensemble behavior, negative controls, and biological context into a unified decision process for AI-guided biologics development.</p></sec></sec><sec id="S13" sec-type="methods"><title>Methods</title><sec id="S14"><title>VHH-antigen dataset curation</title><p id="P62">For benchmarking nanobody-antigen complex prediction, VHH-antigen structures were curated from two complementary sources: SAbDab-nano <sup><xref ref-type="bibr" rid="R81">81</xref></sup>, downloaded March 2025, 90% sequence redundancy cutoff, bound complexes only) and the Antigen-nanobody Complex Database (AACDB) <sup><xref ref-type="bibr" rid="R82">82</xref></sup>. To minimize overlap with the training data of AlphaFold 3 (AF3), Boltz-2, and Chai-1, only post-October 2021depositions were retained from SAbDab-nano, corresponding to the earliest training cutoff among the evaluated tools (Boltz-2). To explicitly probe potential memorization effects, pre-cutoff structures from AACDB were retained and later used to define train-leakage subsets.</p><p id="P63">All structures were filtered to include protein antigens only, with crystallographic resolution ≤ 3.0 Å, VHH length between 110-150 amino acids, and antigen length between 100-400 amino acids. The two datasets were merged and exact PDB-chain duplicates were removed. For epitope-paratope variation analysis, we defined a set of distinct VHHs binding the same antigen within a single PDB entry (<xref ref-type="supplementary-material" rid="SD1">Supplementary Table 1</xref>). This yielded 15 PDBs containing such replicates (30 systems).</p><p id="P64">To maintain computational feasibility, the total number of unique systems was capped below 125. After manual inspection, 17 systems were excluded due to structural artifacts, VHHs contacting multiple antigen chains, or ambiguous chain annotation. The final benchmark comprised 106 VHH-antigen systems from 91 unique PDB entries.</p><p id="P65">Antigen secondary structure was assigned using DSSP, amino-acid composition was computed directly from sequences, and PDB novelty was inferred from the first character of the PDB accession code.</p></sec><sec id="S15"><title>Real vs shuffled complex generation</title><p id="P66">To assess interaction specificity, we constructed a full combinatorial pairing matrix between the curated VHHs and antigens. For each VHH, the experimentally observed pairing with its cognate antigen was designated as the real complex, while all other VHH-antigen pairings were designated as shuffled complexes. This design preserves realistic molecular interfaces while systematically breaking biological specificity.</p><p id="P67">Unless otherwise stated, all analyses were performed on the complete 106 × 106 VHH-antigen matrix, enabling direct comparison between real and shuffled predictions under identical modeling conditions.</p></sec><sec id="S16"><title>Structure prediction</title><p id="P68">Structure predictions for the full combinatorial dataset (106 VHHs × 106 antigens, 50 samples per complex) were performed using three independent structure prediction frameworks: AlphaFold3 (AF3), Boltz-2, and Chai-1. Predictions were executed across multiple high-performance computing environments: <list list-type="bullet" id="L1"><list-item><p id="P69">AF3: Leonardo Booster partition (CINECA, Italy)</p></list-item><list-item><p id="P70">Boltz-2: VEGA supercomputer (IZUM, Slovenia/University of Oslo)</p></list-item><list-item><p id="P71">Chai-1 : Immunohub cluster (University of Oslo)</p></list-item></list></p><p id="P72">Exact hardware specifications, runtime stacks, and software versions were logged at job start for reproducibility.</p><sec id="S17"><title>Chai-1</title><p id="P73">Chai-1 predictions were performed using version 0.6.1. All complexes were run on the University of Oslo Immunohub cluster using 1 GPU per job. Each system was predicted with 50 models, using the parameters: num_trunk_samples = 5, num_diffn_samples = 10, num_trunk_recycles = 3, num_diffn_timesteps = 200, seed = 42. No MSAs were computed; instead, ESM-based embeddings were used under Chai-1 ‘s default inference mode</p></sec><sec id="S18"><title>Boltz-2</title><p id="P74">Predictions were performed using the Boltz CLI v2.2.0. Boltz-1 and Boltz-2 were selected via --model boltz1 or --model boltz2, respectively. Unless otherwise noted, all runs used: --use_msa_server, recycling_steps = 3, sampling_steps = 200, diffusion_samples = 50, fixed random seed (42). Outputs were written in PDB or mmCIF format, with full per-residue and interaction-level scores exported using --write_full_pde. Boltz-2 jobs ran on the Immunohub and Boltz-1 on the VEGA.</p></sec><sec id="S19"><title>AlphaFold3</title><p id="P75">AlphaFold 3 predictions were performed using the alphafold3-3.0.1 via Singularity container on the Leonardo Booster partition of CINECA. Each system was predicted with 50 diffusion samples using parameters: <monospace>num_diffusion_samples = 50, run_data_pipeline = true, run_inference = true</monospace>. Each job was run with a fixed random seed (seed = 1). MSAs were computed using the built-in data pipeline (<monospace>--run_data_pipeline=true</monospace>).</p></sec><sec id="S20"><title>Train, test, and mixed (“both”) labeling</title><p id="P76">Each VHH-antigen system was assigned to train, test, or both categories based on the training cutoffs of the corresponding model: i) AF3: ~30 September 2021, ii) Chai-1: ~12 January 2021, iii) Boltz-2: ~1 June 2023. A system was labeled train if both the VHH and antigen were present in the model’s training set, test if neither was present, and both if one component was present while the other was not.</p></sec><sec id="S21"><title>Structural similarity between replicas</title><p id="P77">To quantify structural reproducibility, we computed geometric and contact-based similarity metrics across independently generated predictions. All analyses were performed on Cα atomic coordinates extracted from PDB files using custom Python scripts (NumPy <sup><xref ref-type="bibr" rid="R83">83</xref></sup>, SciPy <sup><xref ref-type="bibr" rid="R84">84</xref></sup>, pandas <sup><xref ref-type="bibr" rid="R85">85</xref>,<xref ref-type="bibr" rid="R86">86</xref></sup>). For each complex, all pairwise comparisons among replicate structures were computed and aggregated to yield per-system metrics.</p></sec><sec id="S22"><title>Root Mean Square Deviation (RMSD)</title><p id="P78">Global structural similarity was quantified using RMSD after superposition on the antigen via the Kabsch algorithm: <disp-formula id="FD1"><mml:math id="M1"><mml:mrow><mml:mi>R</mml:mi><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>q</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>×</mml:mo><mml:mi>Σ</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mtext>i</mml:mtext></mml:msub><mml:mo>−</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>R</mml:mi><mml:mo>·</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mtext>i</mml:mtext></mml:msub><mml:mo>+</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula> where <italic>X</italic> and <italic>Y</italic> are corresponding Cα coordinates, and <italic>R, t</italic> are the optimal rotation and translation. Lower RMSD values indicate higher structural consistency. For each system, RMSD values were averaged across all replicate pairs to obtain mean ± SD estimates.</p></sec><sec id="S23"><title>Epitope enrichment</title><p id="P79">We computed epitope enrichment for each VHH-antigen pair using a one-tailed binomial test. For each complex (real or shuffled), a replicate prediction was scored as a hit if any contacted antigen residue overlapped the experimental epitope for that antigen. For each pair, we aggregated K (number of hit replicates) out of N (50 replicates). We assessed statistical enrichment using a one-sided binomial test under a null model in which contacts are uniformly distributed over the antigen surface. The null epitope-contact probability p<sub>epitope</sub> was estimated per antigen structure as the fraction of antigen solvent-accessible surface area (SASA) attributable to epitope residues: p<sub>epitope</sub> = SASA<sub>epitope</sub>/SASA<sub>antigen</sub>. Associated summary statistics including expected hits (N<sub>pepitope</sub>), enrichment lift ((K/N)/p<sub>epitope</sub>), and −log<sub>10</sub>(p). To control for multiple testing, p-values were corrected across all evaluated VHH-antigen pairs using the Benjamini-Hochberg procedure to obtain FDR-adjusted q-values.</p><p id="P80">To separate effects driven by cognate pairing versus dataset split leakage, each VHH-antigen pair was assigned a cohort label based on whether the pairing was real (cognate) or shuffled (non-cognate) and on the train/test split assignments of the individual binding partners.</p><p id="P81">Replicate-sensitivity (“significance vs. replicates”) analysis (<xref ref-type="fig" rid="F4">Figure 4C</xref>) set out to evaluate how apparent enrichment depends on ensemble size; we performed a replicate-scaling analysis for hypothetical replicate counts <italic>m</italic>=1,…,N. For each pair, we approximated the number of hits at depth <italic>m</italic> by rescaling the observed hit fraction: Km=round((K/N) <italic>m</italic>). We recomputed one-sided binomial p-values using <italic>m</italic> and the same pepitope, applied Benjamini-Hochberg FDR correction across all pairs separately at each <italic>m</italic>, and recorded the fraction of pairs deemed significant (q&lt;0.05) within each shuffled cohort. These cohort-stratified curves were used to summarize how quickly significance accumulates with increasing stochastic sampling for each model.</p></sec></sec><sec id="S24"><title>Saturation analysis and computational cost measurement</title><p id="P82">To assess the relationship between computational cost and structural prediction quality, we performed saturation runs on VEGA (IZUM, Slovenia), with each job allocated a single NVIDIA A100 GPU (40 GB).</p><sec id="S25"><title>Sampling configurations</title><p id="P83">We tested diffusion sample counts of {n= 1 (baseline), 10, 25, 50, 100} for AlphaFold3 and Boltz-1/2. For Chai-1, we additionally varied trunk samples, testing configurations of {5x1, 5x2, 5x5, 5x10, 5x20} (trunk x diffusion samples), to match the structure prediction of the whole 106x106 matrix (as described above).</p></sec><sec id="S26"><title>MSA strategy</title><p id="P84">To isolate inference costs from alignment overhead: i) AlphaFold3: MSAs were generated once per system during baseline run (diffusion_samples =1) and reused for subsequent runs via --run_data_pipeline=false, ii) Boltz-1/2: a shared MSA-cache (--cache) was populated during baseline runs and reused across all saturation runs, iii) Chai-1: ESM embeddings were used without MSA search (--use-ESM-embeddings)</p></sec><sec id="S27"><title>Seed strategy</title><p id="P85">Each prediction job used independent random seed initialization to ensure independent sampling across runs. AlphaFold3 and Chai-1 used explicitly generated seeds at submission time; for Boltz-1/2 we used default random initialization. Within each run, multiple diffusion samples explore different conformations along independent stochastic trajectories.</p></sec><sec id="S28"><title>Energy monitoring</title><p id="P86">GPU telemetry (power draw, utilization, memory usage) was sampled every 5 seconds using nvidia-smi and written to per-run CSV files. Wall-clock timing and system metrics were captured via /usr/bin/time -v. Total energy consumption (Wh) was calculated by trapezoidal integration of instantaneous power readings over the run duration: <disp-formula id="FD2"><mml:math id="M2"><mml:mrow><mml:msub><mml:mtext>E</mml:mtext><mml:mo>_</mml:mo></mml:msub><mml:mtext>run</mml:mtext><mml:mspace width="0.2em"/><mml:mo>=</mml:mo><mml:mspace width="0.2em"/><mml:mo>Σ</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>P</mml:mtext><mml:mo>_</mml:mo><mml:mtext>i</mml:mtext><mml:mo>+</mml:mo><mml:mtext>P</mml:mtext><mml:mo>_</mml:mo><mml:mo>{</mml:mo><mml:mtext>i</mml:mtext><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>}</mml:mo></mml:mrow><mml:mo>)/</mml:mo></mml:mrow><mml:mn>2</mml:mn><mml:mo>·</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>t</mml:mtext><mml:mo>_</mml:mo><mml:mo>{</mml:mo><mml:mtext>i</mml:mtext><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>}-</mml:mo><mml:mtext>t</mml:mtext><mml:mo>_</mml:mo><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]/</mml:mo></mml:mrow><mml:mn>3600</mml:mn></mml:mrow></mml:math></disp-formula></p><p id="P87">Runs were aggregated by (model × diffusion sample count) and reported as mean ± SD. Due to scheduling, initialization, and I/O overhead, reported values reflect relative efficiency under identical cluster conditions, not theoretical lower bounds.</p></sec></sec><sec id="S29"><title>Data analysis and statistical aggregation</title><p id="P88">All analyses were implemented in Python using the NumPy, SciPy, matplotlib and pandas libraries. Pairwise structure comparisons were parallelized using the multiprocessing module. The workflow automatically parsed PDB files, performed optimal structural alignments, computed the defined similarity metrics, and aggregated results per system. For each complex system, all pairwise metric values were averaged to produce mean and standard deviation estimates.</p><sec id="S30"><title>Computational cost and sampling efficiency analysis</title><p id="P89">As mentioned above, energy consumption was measured using NVIDIA’s nvidia-smi tool, sampling GPU power draws at 5-second intervals throughout each prediction run.</p><p id="P90">Structural quality was assessed using DockQ v[2.1.3]<sup><xref ref-type="bibr" rid="R41">41</xref></sup>, which integrates interface RMSD, ligand RMSD, and fraction of native contacts into a single score ranging from 0 (incorrect) to 1 (perfect). Quality categories follow CAPRI conventions: incorrect (&lt;0.23), acceptable (0.23-0.49), medium (0.49-0.80), and high (≥0.80).</p><p id="P91">For saturation analysis, each system was evaluated with five independent random seeds, each assigned to a single saturation level (N = 1, 10, 25, 50, and 100 diffusion samples for seeds 1-5, respectively). To isolate seed-dependent variation from saturation effects, we extracted only the first sample (sample<sub>0</sub>) from each saturation run and computed per-system DockQ and ipTM ranges (maximum minus minimum across five seeds). Chai-1 uses a fixed five-trunk architecture. Requiring a modified sampling scheme to achieve equivalent saturation levels. We varied the number of diffusion samples per trunk (1, 2, 5, 10 and 20 samples x 5 trunks = N of 5, 10, 25, 50, and 100 total samples). For seed-isolation analysis, we extracted sample<sub>0</sub> from trunk<sub>0</sub> as the baseline prediction for each seed.</p><p id="P92">Marginal quality gains were computed as the per-system difference in maximum DockQ between consecutive saturation levels, then summarized as the median across all systems. Efficiency frontiers were done by plotting cumulative median ΔDockQ (relative to baseline) against cumulative median energy expenditure at each saturation level. Cross-seed trajectory analysis tracked the cumulative maximum DockQ (i.e. the best DockQ achieved up to and including each sample number) within each seed’s saturation run, then computed the median across systems. Correlations between confidence score changes (ΔipTM) and quality changes (ΔDockQ) were assessed using pearson correlation coefficients.</p></sec></sec></sec><sec sec-type="supplementary-material" id="SM"><title>Supplementary Material</title><supplementary-material content-type="local-data" id="SD1"><label>Supplementary Table 1</label><media xlink:href="EMS215481-supplement-Supplementary_Table_1.csv" mimetype="text" mime-subtype="csv; charset=utf-8" id="d89aAcEbB" position="anchor"/></supplementary-material><supplementary-material content-type="local-data" id="SD2"><label>Supplementary Table 3</label><media xlink:href="EMS215481-supplement-Supplementary_Table_3.csv" mimetype="text" mime-subtype="csv; charset=utf-8" id="d89aAcEcB" position="anchor"/></supplementary-material><supplementary-material content-type="local-data" id="SD3"><label>Supplementary Table 4</label><media xlink:href="EMS215481-supplement-Supplementary_Table_4.csv" mimetype="text" mime-subtype="csv; charset=utf-8" id="d89aAcEdB" position="anchor"/></supplementary-material><supplementary-material content-type="local-data" id="SD4"><label>Supplementary Materials</label><media xlink:href="EMS215481-supplement-Supplementary_Materials.pdf" mimetype="application" mime-subtype="pdf" id="d89aAcEeB" position="anchor"/></supplementary-material></sec></body><back><ack id="S31"><title>Acknowledgements</title><p>We thank Prof. Charlotte Deane and Henriette Capel (OPIG, University of Oxford) for valuable discussions. We are grateful to Žiga Zebec (IZUM, Slovenia) for guidance on assessing VEGA and EuroHPC resources.</p><sec id="S32"><title>Funding</title><p>This work was supported by grants from the Norwegian Cancer Society Grant (#215817, to VG), Research Council of Norway projects (#300740, #331890 to VG). This project has received funding (to VG) from the Innovative Medicines Initiative 2 Joint Undertaking under grant agreement No 101007799 (Inno4Vac). This Joint Undertaking receives support from the European Union’s Horizon 2020 research and innovation programme and EFPIA. This communication reflects the author’s view and neither IMI nor the European Union, EFPIA, or any Associated Partners are responsible for any use that may be made of the information contained therein. Funded by the European Union (ERC, AB-AG-INTERACT, 101125630, to VG). KKB has received funding from the European Union’s Horizon Europe research and innovation programme under the Marie Skłodowska-Curie COFUND Postdoctoral Programme grant agreement No. 101081355-SMASH and from the Republic of Slovenia and the European Union from the European Regional Development Fund. KKB acknowledges the EuroHPC Joint Undertaking for awarding project EHPC-DEV-2025D08-047 access to the EuroHPC supercomputers LEONARDO, hosted by CINECA (Italy), and Vega, hosted by IZUM (Slovenia), through a EuroHPC Development Access call. Computational support was provided by the EPICURE project, which received funding from the EuroHPC Joint Undertaking under grant agreement No. 101139786. KKB also acknowledges the HPC RIVR consortium for providing computing resources through the SLING project S25R01-01. AdM was supported by the grants P3-0428, J4-50144, N4-0282 and N4-0325 provided by the Slovenian Research and Innovation Agency (ARIS). PS is a Royal Society University Research Fellow (grant no. URF\R\251013) and acknowledges funding from UK Research and Innovation (UKRI) Engineering and Physical Sciences Research Council (EPSRC grant no. EP/X024733/1).</p></sec></ack><sec id="S33" sec-type="data-availability"><title>Data availability</title><p id="P93">Supplementary tables, experimental and computational data, along with related scripts and pipelines, can be found on Zenodo (<ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.5281/zenodo.18390239">10.5281/zenodo.18390239</ext-link>) and GitHub (<ext-link ext-link-type="uri" xlink:href="https://github.com/csi-greifflab/ab_ag_champloo">https://github.com/csi-greifflab/ab_ag_champloo</ext-link>).</p></sec><fn-group><fn id="FN2"><p id="P94">Disclaimer: Co-funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Research Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.</p></fn><fn id="FN3" fn-type="con"><p id="P95"><bold>Author contributions</bold></p><p id="P96">E.S. and V.G. conceived the study. E.S. led the study, performed structure prediction using Chai-1 and Boltz-2, conducted data analysis, prepared figures, and wrote the manuscript. M.A. performed data analysis, prepared figures, and contributed to manuscript writing. K.K.B. performed structure prediction using AF3, conducted saturation and calibration studies, performed data analysis, prepared figures, and contributed to manuscript writing. L.S. and S.M. supported AF3 and Boltz-1 computations, including optimization and execution of prediction runs. A.K. and C.F. performed sequence analyses and generated preliminary results. P.S. and A.d.M. supervised the study and provided scientific guidance. V.G. supervised the project, contributed to conceptualization, and participated in manuscript writing. All authors reviewed and approved the final version of the manuscript.</p></fn><fn id="FN4" fn-type="conflict"><p id="P97"><bold>Disclosure statement</bold></p><p id="P98">V.G. declares advisory board positions in aiNET GmbH, Enpicom B.V, Absci, Omniscope, and Diagonal Therapeutics. V.G. is a consultant for Adaptyv Biosystems, Specifica Inc, Roche/Genentech, immunai, Proteinea, LabGenius, and FairJourney Biologics. V.G. is an employee of Imprint LLC.</p></fn></fn-group><ref-list><ref id="R1"><label>1</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Meng</surname><given-names>F</given-names></name><etal/></person-group><article-title>A comprehensive overview of recent advances in generative models for antibodies</article-title><source>Comput Struct Biotechnol. J</source><year>2024</year><volume>23</volume><fpage>2648</fpage><lpage>2660</lpage><pub-id pub-id-type="doi">10.1016/j.csbj.2024.06.016</pub-id><pub-id pub-id-type="pmcid">PMC11254834</pub-id><pub-id pub-id-type="pmid">39027650</pub-id></element-citation></ref><ref id="R2"><label>2</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bielska</surname><given-names>W</given-names></name><etal/></person-group><article-title>Applying computational protein design to therapeutic antibody discovery -current state and perspectives</article-title><source>arXiv [q-bio.BM]</source><year>2025</year><pub-id pub-id-type="doi">10.3389/fimmu.2025.1571371</pub-id><pub-id pub-id-type="pmcid">PMC12137305</pub-id><pub-id pub-id-type="pmid">40475769</pub-id></element-citation></ref><ref id="R3"><label>3</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Akbar</surname><given-names>R</given-names></name><etal/></person-group><article-title>Progress and challenges for the machine learning-based design of fit-for-purpose monoclonal antibodies</article-title><source>MAbs</source><year>2022</year><volume>14</volume><elocation-id>2008790</elocation-id><pub-id pub-id-type="doi">10.1080/19420862.2021.2008790</pub-id><pub-id pub-id-type="pmcid">PMC8928824</pub-id><pub-id pub-id-type="pmid">35293269</pub-id></element-citation></ref><ref id="R4"><label>4</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bashour</surname><given-names>H</given-names></name><etal/></person-group><article-title>Biophysical cartography of the native and human-engineered antibody landscapes quantifies the plasticity of antibody developability</article-title><source>Commun. Biol</source><year>2024</year><volume>7</volume><fpage>922</fpage><pub-id pub-id-type="doi">10.1038/s42003-024-06561-3</pub-id><pub-id pub-id-type="pmcid">PMC11291509</pub-id><pub-id pub-id-type="pmid">39085379</pub-id></element-citation></ref><ref id="R5"><label>5</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Greiff</surname><given-names>V</given-names></name><name><surname>Yaari</surname><given-names>G</given-names></name><name><surname>Cowell</surname><given-names>LG</given-names></name></person-group><article-title>Mining adaptive immune receptor repertoires for biological and clinical information using machine learning</article-title><source>Curr Opin Syst Biol</source><year>2020</year><volume>24</volume><fpage>109</fpage><lpage>119</lpage></element-citation></ref><ref id="R6"><label>6</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Levine</surname><given-names>S</given-names></name><etal/></person-group><article-title>Origin-1 : a generative AI platform for <italic>de novo</italic> antibody design against novel epitopes</article-title><source>bioRxiv</source><year>2026</year><pub-id pub-id-type="doi">10.64898/2026.01.14.699389</pub-id></element-citation></ref><ref id="R7"><label>7</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Overath</surname><given-names>MD</given-names></name><etal/></person-group><article-title>Accelerating multi-objective V<sub>H</sub> H discovery via integrated high-throughput selection and AlphaFold3-guided structure prediction</article-title><source>bioRxiv</source><year>2026</year><pub-id pub-id-type="doi">10.64898/2026.01.19.700436</pub-id></element-citation></ref><ref id="R8"><label>8</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Laustsen</surname><given-names>AH</given-names></name><name><surname>Greiff</surname><given-names>V</given-names></name><name><surname>Karatt-Vellatt</surname><given-names>A</given-names></name><name><surname>Muyldermans</surname><given-names>S</given-names></name><name><surname>Jenkins</surname><given-names>TP</given-names></name></person-group><article-title>Animal immunization, in vitro display technologies, and machine learning for antibody discovery</article-title><source>Trends Biotechnol</source><year>2021</year><volume>39</volume><fpage>1263</fpage><lpage>1273</lpage><pub-id pub-id-type="pmid">33775449</pub-id></element-citation></ref><ref id="R9"><label>9</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ruffolo</surname><given-names>JA</given-names></name><name><surname>Chu</surname><given-names>L-S</given-names></name><name><surname>Mahajan</surname><given-names>SP</given-names></name><name><surname>Gray</surname><given-names>JJ</given-names></name></person-group><article-title>Fast, accurate antibody structure prediction from deep learning on massive set of natural antibodies</article-title><source>Nat Commun</source><year>2023</year><volume>14</volume><elocation-id>2389</elocation-id><pub-id pub-id-type="doi">10.1038/s41467-023-38063-x</pub-id><pub-id pub-id-type="pmcid">PMC10129313</pub-id><pub-id pub-id-type="pmid">37185622</pub-id></element-citation></ref><ref id="R10"><label>10</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hamers-Casterman</surname><given-names>C</given-names></name><etal/></person-group><article-title>Naturally occurring antibodies devoid of light chains</article-title><source>Nature</source><year>1993</year><volume>363</volume><fpage>446</fpage><lpage>448</lpage><pub-id pub-id-type="pmid">8502296</pub-id></element-citation></ref><ref id="R11"><label>11</label><element-citation publication-type="confproc"><person-group person-group-type="author"><name><surname>Smorodina</surname><given-names>E</given-names></name><etal/></person-group><source>Structural modeling of antibody variant epitope specificity with complementary experimental and computational techniques</source><conf-name>ICLR 2025 Workshop on Generative and Experimental Perspectives for Biomolecular Design</conf-name><year>2025</year></element-citation></ref><ref id="R12"><label>12</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kuroda</surname><given-names>D</given-names></name><name><surname>Gray</surname><given-names>JJ</given-names></name></person-group><article-title>Shape complementarity and hydrogen bond preferences in protein-protein interfaces: implications for antibody modeling and protein-protein docking</article-title><source>Bioinformatics</source><year>2016</year><volume>32</volume><fpage>2451</fpage><lpage>2456</lpage><pub-id pub-id-type="doi">10.1093/bioinformatics/btw197</pub-id><pub-id pub-id-type="pmcid">PMC4978935</pub-id><pub-id pub-id-type="pmid">27153634</pub-id></element-citation></ref><ref id="R13"><label>13</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Zavrtanik</surname><given-names>U</given-names></name><name><surname>Lukan</surname><given-names>J</given-names></name><name><surname>Loris</surname><given-names>R</given-names></name><name><surname>Lah</surname><given-names>J</given-names></name><name><surname>Hadži</surname><given-names>S</given-names></name></person-group><article-title>Structural Basis of Epitope Recognition by Heavy-Chain Camelid Antibodies</article-title><source>J Mol Biol</source><year>2018</year><volume>430</volume><fpage>4369</fpage><lpage>4386</lpage><pub-id pub-id-type="pmid">30205092</pub-id></element-citation></ref><ref id="R14"><label>14</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Mitchell</surname><given-names>LS</given-names></name><name><surname>Colwell</surname><given-names>LJ</given-names></name></person-group><article-title>Analysis of nanobody paratopes reveals greater diversity than classical antibodies</article-title><source>Protein Eng Des Sel</source><year>2018</year><volume>31</volume><fpage>267</fpage><lpage>275</lpage><pub-id pub-id-type="doi">10.1093/protein/gzy017</pub-id><pub-id pub-id-type="pmcid">PMC6277174</pub-id><pub-id pub-id-type="pmid">30053276</pub-id></element-citation></ref><ref id="R15"><label>15</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Abramson</surname><given-names>J</given-names></name><etal/></person-group><article-title>Accurate structure prediction of biomolecular interactions with AlphaFold 3</article-title><source>Nature</source><year>2024</year><volume>630</volume><fpage>493</fpage><lpage>500</lpage><pub-id pub-id-type="doi">10.1038/s41586-024-07487-w</pub-id><pub-id pub-id-type="pmcid">PMC11168924</pub-id><pub-id pub-id-type="pmid">38718835</pub-id></element-citation></ref><ref id="R16"><label>16</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Passaro</surname><given-names>S</given-names></name><etal/></person-group><article-title>Boltz-2: Towards accurate and efficient binding affinity prediction</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.1101/2025.06.14.659707</pub-id></element-citation></ref><ref id="R17"><label>17</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Discovery</surname><given-names>C</given-names></name><etal/></person-group><article-title>Chai-1: Decoding the molecular interactions of life</article-title><source>bioRxiv</source><year>2024</year><pub-id pub-id-type="doi">10.1101/2024.10.10.615955</pub-id></element-citation></ref><ref id="R18"><label>18</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Smorodina</surname><given-names>E</given-names></name><etal/></person-group><article-title>Structural informatic study of determined and AlphaFold2 predicted molecular structures of 13 human solute carrier transporters and their water-soluble QTY variants</article-title><source>Sci Rep</source><year>2022</year><volume>12</volume><elocation-id>20103</elocation-id><pub-id pub-id-type="doi">10.1038/s41598-022-23764-y</pub-id><pub-id pub-id-type="pmcid">PMC9684436</pub-id><pub-id pub-id-type="pmid">36418372</pub-id></element-citation></ref><ref id="R19"><label>19</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Yin</surname><given-names>R</given-names></name><name><surname>Pierce</surname><given-names>BG</given-names></name></person-group><article-title>Evaluation of AlphaFold antibody-antigen modeling with implications for improving predictive accuracy</article-title><source>Protein Sci</source><year>2024</year><volume>33</volume><elocation-id>e4865</elocation-id><pub-id pub-id-type="doi">10.1002/pro.4865</pub-id><pub-id pub-id-type="pmcid">PMC10751731</pub-id><pub-id pub-id-type="pmid">38073135</pub-id></element-citation></ref><ref id="R20"><label>20</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hitawala</surname><given-names>FN</given-names></name><name><surname>Gray</surname><given-names>JJ</given-names></name></person-group><article-title>What does AlphaFold3 learn about antigen and nanobody docking, and what remains unsolved?</article-title><source>Bioengineering</source><year>2024</year><pub-id pub-id-type="doi">10.1080/19420862.2025.2545601</pub-id><pub-id pub-id-type="pmcid">PMC12360200</pub-id><pub-id pub-id-type="pmid">40814020</pub-id></element-citation></ref><ref id="R21"><label>21</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Harmalkar</surname><given-names>A</given-names></name><name><surname>Lyskov</surname><given-names>S</given-names></name><name><surname>Gray</surname><given-names>JJ</given-names></name></person-group><article-title>Reliable protein-protein docking with AlphaFold, Rosetta, and replica-exchange</article-title><source>bioRxivorg</source><year>2023</year><pub-id pub-id-type="doi">10.7554/eLife.94029</pub-id><pub-id pub-id-type="pmcid">PMC12113263</pub-id><pub-id pub-id-type="pmid">40424178</pub-id></element-citation></ref><ref id="R22"><label>22</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Weitzner</surname><given-names>BD</given-names></name><etal/></person-group><article-title>Modeling and docking of antibody structures with Rosetta</article-title><source>Nat Protoc</source><year>2017</year><volume>12</volume><fpage>401</fpage><lpage>416</lpage><pub-id pub-id-type="doi">10.1038/nprot.2016.180</pub-id><pub-id pub-id-type="pmcid">PMC5739521</pub-id><pub-id pub-id-type="pmid">28125104</pub-id></element-citation></ref><ref id="R23"><label>23</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ambrosetti</surname><given-names>F</given-names></name><name><surname>Jiménez-García</surname><given-names>B</given-names></name><name><surname>Roel-Touris</surname><given-names>J</given-names></name><name><surname>Bonvin</surname><given-names>AMJJ</given-names></name></person-group><article-title>Modeling antibody-antigen complexes by information-driven docking</article-title><source>Structure</source><year>2020</year><volume>28</volume><fpage>119</fpage><lpage>129</lpage><elocation-id>e2</elocation-id><pub-id pub-id-type="pmid">31727476</pub-id></element-citation></ref><ref id="R24"><label>24</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname><given-names>X</given-names></name><etal/></person-group><article-title>Protenix -advancing structure prediction through a comprehensive AlphaFold3 reproduction</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.1101/2025.01.08.631967</pub-id></element-citation></ref><ref id="R25"><label>25</label><element-citation publication-type="web"><source>The Isomorphic Labs Drug Design Engine unlocks a new frontier beyond AlphaFold</source><comment><ext-link ext-link-type="uri" xlink:href="https://www.isomorphiclabs.com/articles/the-isomorphic-labs-drug-design-engine-unlocks-a-new-frontier">https://www.isomorphiclabs.com/articles/the-isomorphic-labs-drug-design-engine-unlocks-a-new-frontier</ext-link></comment></element-citation></ref><ref id="R26"><label>26</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Hitawala</surname><given-names>FN</given-names></name><name><surname>Gray</surname><given-names>JJ</given-names></name></person-group><article-title>What does AlphaFold3 learn about antibody and nanobody docking, and what remains unsolved?</article-title><source>MAbs</source><year>2025</year><volume>17</volume><elocation-id>2545601</elocation-id><pub-id pub-id-type="doi">10.1080/19420862.2025.2545601</pub-id><pub-id pub-id-type="pmcid">PMC12360200</pub-id><pub-id pub-id-type="pmid">40814020</pub-id></element-citation></ref><ref id="R27"><label>27</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Jumper</surname><given-names>J</given-names></name><etal/></person-group><article-title>Highly accurate protein structure prediction with AlphaFold</article-title><source>Nature</source><year>2021</year><volume>596</volume><fpage>583</fpage><lpage>589</lpage><pub-id pub-id-type="doi">10.1038/s41586-021-03819-2</pub-id><pub-id pub-id-type="pmcid">PMC8371605</pub-id><pub-id pub-id-type="pmid">34265844</pub-id></element-citation></ref><ref id="R28"><label>28</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Wohlwend</surname><given-names>J</given-names></name><etal/></person-group><article-title>Boltz-1: Democratizing Biomolecular Interaction Modeling</article-title><source>Biophysics</source><year>2024</year></element-citation></ref><ref id="R29"><label>29</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Boitreaud</surname><given-names>J</given-names></name><etal/></person-group><article-title>Chai-1 : Decoding the molecular interactions of life</article-title><source>Synthetic Biology</source><year>2024</year></element-citation></ref><ref id="R30"><label>30</label><element-citation publication-type="journal"><collab>Chai Discovery Team</collab><article-title>Zero-shot antibody design in a 24-well plate</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.1101/2025.07.05.663018</pub-id></element-citation></ref><ref id="R31"><label>31</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Fromm</surname><given-names>S</given-names></name><name><surname>Ludaic</surname><given-names>M</given-names></name><name><surname>Elofsson</surname><given-names>A</given-names></name></person-group><article-title>Evaluating deep learning based structure prediction methods on antibody-antigen complexes</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.1093/bioinformatics/btag136</pub-id><pub-id pub-id-type="pmcid">PMC13061134</pub-id><pub-id pub-id-type="pmid">41863324</pub-id></element-citation></ref><ref id="R32"><label>32</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Unsal</surname><given-names>S</given-names></name><name><surname>Holland</surname><given-names>B</given-names></name><name><surname>Sardag</surname><given-names>I</given-names></name><name><surname>Timucin</surname><given-names>E</given-names></name></person-group><article-title>Confidence scoring for AI-predicted antibody-antigen complexes: AntiConf as a precision-driven metric</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.1093/bib/bbag137</pub-id><pub-id pub-id-type="pmcid">PMC13032827</pub-id><pub-id pub-id-type="pmid">41903187</pub-id></element-citation></ref><ref id="R33"><label>33</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Grieswelle</surname><given-names>M</given-names></name><etal/></person-group><article-title>A new benchmark for deep learning based affinity prediction: Solving the inter-protein scoring noise problem</article-title><source>ChemRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.26434/chemrxiv-2025-sf3cs</pub-id></element-citation></ref><ref id="R34"><label>34</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bradley</surname><given-names>P</given-names></name></person-group><article-title>Structure-based prediction of T cell receptor:peptide-MHC interactions</article-title><source>Elife</source><year>2023</year><volume>12</volume><pub-id pub-id-type="doi">10.7554/eLife.82813</pub-id><pub-id pub-id-type="pmcid">PMC9859041</pub-id><pub-id pub-id-type="pmid">36661395</pub-id></element-citation></ref><ref id="R35"><label>35</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Evans</surname><given-names>R</given-names></name><etal/></person-group><article-title>Protein complex prediction with AlphaFold-Multimer</article-title><source>bioRxiv</source><year>2021</year><pub-id pub-id-type="doi">10.1101/2021.10.04.463034</pub-id></element-citation></ref><ref id="R36"><label>36</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Overath</surname><given-names>MD</given-names></name><etal/></person-group><article-title>Predicting experimental success in DE Novo binder design: A meta-analysis of 3,766 experimentally characterised binders</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.1101/2025.08.14.670059</pub-id></element-citation></ref><ref id="R37"><label>37</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Jarboe</surname><given-names>B</given-names></name><name><surname>Dunbrack</surname><given-names>RL</given-names><suffix>Jr</suffix></name></person-group><article-title>Structure-guided analysis and prediction of human E2–E3 ligase pairing specificity</article-title><source>bioRxiv</source><year>2026</year><pub-id pub-id-type="doi">10.64898/2026.02.10.700855</pub-id></element-citation></ref><ref id="R38"><label>38</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Holt</surname><given-names>CM</given-names></name><etal/></person-group><article-title>Contrastive Learning Enables Epitope Overlap Predictions for Targeted Antibody Discovery</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.1016/j.patter.2025.101419</pub-id><pub-id pub-id-type="pmcid">PMC12921510</pub-id><pub-id pub-id-type="pmid">41726099</pub-id></element-citation></ref><ref id="R39"><label>39</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Dang</surname><given-names>X</given-names></name><etal/></person-group><article-title>Epitope mapping of monoclonal antibodies: a comprehensive comparison of different technologies</article-title><source>MAbs</source><year>2023</year><volume>15</volume><elocation-id>2285285</elocation-id><pub-id pub-id-type="doi">10.1080/19420862.2023.2285285</pub-id><pub-id pub-id-type="pmcid">PMC10730160</pub-id><pub-id pub-id-type="pmid">38010385</pub-id></element-citation></ref><ref id="R40"><label>40</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Basu</surname><given-names>S</given-names></name><name><surname>Wallner</surname><given-names>B</given-names></name></person-group><article-title>DockQ: A Quality Measure for Protein-Protein Docking Models</article-title><source>PLoS One</source><year>2016</year><volume>11</volume><elocation-id>e0161879</elocation-id><pub-id pub-id-type="doi">10.1371/journal.pone.0161879</pub-id><pub-id pub-id-type="pmcid">PMC4999177</pub-id><pub-id pub-id-type="pmid">27560519</pub-id></element-citation></ref><ref id="R41"><label>41</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Mirabello</surname><given-names>C</given-names></name><name><surname>Wallner</surname><given-names>B</given-names></name></person-group><article-title>DockQ v2: improved automatic quality measure for protein multimers, nucleic acids, and small molecules</article-title><source>Bioinformatics</source><year>2024</year><volume>40</volume><pub-id pub-id-type="doi">10.1093/bioinformatics/btae586</pub-id><pub-id pub-id-type="pmcid">PMC11467047</pub-id><pub-id pub-id-type="pmid">39348158</pub-id></element-citation></ref><ref id="R42"><label>42</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Genz</surname><given-names>LR</given-names></name><name><surname>Nair</surname><given-names>S</given-names></name><name><surname>Nagar</surname><given-names>N</given-names></name><name><surname>Topf</surname><given-names>M</given-names></name></person-group><article-title>Assessing scoring metrics for AlphaFold2 and AlphaFold3 protein complex predictions</article-title><source>Protein Sci</source><year>2025</year><volume>34</volume><elocation-id>e70327</elocation-id><pub-id pub-id-type="doi">10.1002/pro.70327</pub-id><pub-id pub-id-type="pmcid">PMC12516916</pub-id><pub-id pub-id-type="pmid">41081541</pub-id></element-citation></ref><ref id="R43"><label>43</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Weeratunga</surname><given-names>S</given-names></name><etal/></person-group><article-title>Interrogation and validation of the interactome of neuronal Munc18-interacting Mint proteins with AlphaFold2</article-title><source>J Biol Chem</source><year>2024</year><volume>300</volume><elocation-id>105541</elocation-id><pub-id pub-id-type="doi">10.1016/j.jbc.2023.105541</pub-id><pub-id pub-id-type="pmcid">PMC10820826</pub-id><pub-id pub-id-type="pmid">38072052</pub-id></element-citation></ref><ref id="R44"><label>44</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Fernández-Quintero</surname><given-names>ML</given-names></name><etal/></person-group><article-title>Challenges in antibody structure prediction</article-title><source>MAbs</source><year>2023</year><volume>15</volume><elocation-id>2175319</elocation-id><pub-id pub-id-type="doi">10.1080/19420862.2023.2175319</pub-id><pub-id pub-id-type="pmcid">PMC9928471</pub-id><pub-id pub-id-type="pmid">36775843</pub-id></element-citation></ref><ref id="R45"><label>45</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname><given-names>C</given-names></name><name><surname>Pleiss</surname><given-names>G</given-names></name><name><surname>Sun</surname><given-names>Y</given-names></name><name><surname>Weinberger</surname><given-names>KQ</given-names></name></person-group><article-title>On calibration of modern neural networks</article-title><source>arXiv [cs.LG]</source><year>2017</year><pub-id pub-id-type="doi">10.48550/ARXIV.1706.04599</pub-id></element-citation></ref><ref id="R46"><label>46</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Kull</surname><given-names>M</given-names></name><etal/></person-group><article-title>Beyond temperature scaling: Obtaining well-calibrated multiclass probabilities with Dirichlet calibration</article-title><source>arXiv [cs.LG]</source><year>2019</year><pub-id pub-id-type="doi">10.48550/ARXIV.1910.12656</pub-id></element-citation></ref><ref id="R47"><label>47</label><element-citation publication-type="other"><comment>Website. <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/html/2512.11892v1">https://arxiv.org/html/2512.11892v1</ext-link></comment></element-citation></ref><ref id="R48"><label>48</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Morand</surname><given-names>C</given-names></name><name><surname>Ligozat</surname><given-names>A-L</given-names></name><name><surname>Névéol</surname><given-names>A</given-names></name></person-group><article-title>How green can AI be? A study of trends in Machine Learning environmental impacts</article-title><source>arXiv [cs.LG]</source><year>2024</year><pub-id pub-id-type="doi">10.48550/ARXIV.2412.17376</pub-id></element-citation></ref><ref id="R49"><label>49</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Lannelongue</surname><given-names>L</given-names></name><name><surname>Inouye</surname><given-names>M</given-names></name></person-group><article-title>Environmental Impacts of Machine Learning Applications in Protein Science</article-title><source>Cold Spring Harb Perspect Biol</source><year>2023</year><volume>15</volume><pub-id pub-id-type="doi">10.1101/cshperspect.a041473</pub-id><pub-id pub-id-type="pmcid">PMC10691472</pub-id><pub-id pub-id-type="pmid">38040454</pub-id></element-citation></ref><ref id="R50"><label>50</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Johansson-Åkhe</surname><given-names>I</given-names></name><name><surname>Wallner</surname><given-names>B</given-names></name></person-group><article-title>Improving peptide-protein docking with AlphaFold-Multimer using forced sampling</article-title><source>Front Bioinform</source><year>2022</year><volume>2</volume><elocation-id>959160</elocation-id><pub-id pub-id-type="doi">10.3389/fbinf.2022.959160</pub-id><pub-id pub-id-type="pmcid">PMC9580857</pub-id><pub-id pub-id-type="pmid">36304330</pub-id></element-citation></ref><ref id="R51"><label>51</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Stark</surname><given-names>H</given-names></name><etal/></person-group><article-title>BoltzGen: Toward Universal Binder Design</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.1101/2025.11.20.689494</pub-id></element-citation></ref><ref id="R52"><label>52</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname><given-names>X</given-names></name><name><surname>Coratella</surname><given-names>I</given-names></name><name><surname>Reys</surname><given-names>V</given-names></name><name><surname>Bonvin</surname><given-names>AM</given-names></name></person-group><article-title>DeepRank-Ab: a scoring function for antibody-antigen complexes based on geometric deep learning</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="pmid">42230751</pub-id></element-citation></ref><ref id="R53"><label>53</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Almeida</surname><given-names>DS</given-names></name><etal/></person-group><article-title>AbSet: A Standardized Data Set of Antibody Structures for Machine Learning Applications</article-title><source>J Chem Inf Model</source><year>2025</year><volume>65</volume><fpage>4767</fpage><lpage>4774</lpage><pub-id pub-id-type="doi">10.1021/acs.jcim.5c00410</pub-id><pub-id pub-id-type="pmcid">PMC12117563</pub-id><pub-id pub-id-type="pmid">40349368</pub-id></element-citation></ref><ref id="R54"><label>54</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Chungyoun</surname><given-names>M</given-names></name><name><surname>Gray</surname><given-names>J</given-names></name></person-group><article-title>Fitness Landscape for Antibodies 2: Benchmarking Reveals That Protein AI Models Cannot Yet Consistently Predict Developability Properties</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.64898/2025.12.27.696706</pub-id></element-citation></ref><ref id="R55"><label>55</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Tennenhouse</surname><given-names>A</given-names></name><etal/></person-group><article-title>Structure-based design of antibody repertoires with drug-like properties</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.64898/2025.12.10.693474</pub-id></element-citation></ref><ref id="R56"><label>56</label><element-citation publication-type="confproc"><person-group person-group-type="author"><name><surname>Smorodina</surname><given-names>E</given-names></name><etal/></person-group><source>Structural modeling of antibody variant epitope specificity with complementary experimental and computational techniques</source><conf-name>ICLR 2025 Workshop on Generative and Experimental Perspectives for Biomolecular Design</conf-name><year>2025</year></element-citation></ref><ref id="R57"><label>57</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Spoendlin</surname><given-names>FC</given-names></name><etal/></person-group><article-title>Predicting the conformational flexibility of antibody and T cell receptor complementarity-determining regions</article-title><source>Nat Mach Intell</source><year>2025</year><volume>7</volume><fpage>1755</fpage><lpage>1767</lpage><pub-id pub-id-type="doi">10.1038/s42256-025-01131-6</pub-id><pub-id pub-id-type="pmcid">PMC12552124</pub-id><pub-id pub-id-type="pmid">41143207</pub-id></element-citation></ref><ref id="R58"><label>58</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Fernández-Quintero</surname><given-names>ML</given-names></name><name><surname>Georges</surname><given-names>G</given-names></name><name><surname>Varga</surname><given-names>JM</given-names></name><name><surname>Liedl</surname><given-names>KR</given-names></name></person-group><article-title>Ensembles in solution as a new paradigm for antibody structure prediction and design</article-title><source>MAbs</source><year>2021</year><volume>13</volume><elocation-id>1923122</elocation-id><pub-id pub-id-type="doi">10.1080/19420862.2021.1923122</pub-id><pub-id pub-id-type="pmcid">PMC8158028</pub-id><pub-id pub-id-type="pmid">34030577</pub-id></element-citation></ref><ref id="R59"><label>59</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ali</surname><given-names>M</given-names></name><etal/></person-group><article-title>Improving nanobody structure prediction with self-distillation</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.64898/2025.12.01.691162</pub-id></element-citation></ref><ref id="R60"><label>60</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Wayment-Steele</surname><given-names>HK</given-names></name><etal/></person-group><article-title>Predicting multiple conformations via sequence clustering and AlphaFold2</article-title><source>Nature</source><year>2024</year><volume>625</volume><fpage>832</fpage><lpage>839</lpage><pub-id pub-id-type="doi">10.1038/s41586-023-06832-9</pub-id><pub-id pub-id-type="pmcid">PMC10808063</pub-id><pub-id pub-id-type="pmid">37956700</pub-id></element-citation></ref><ref id="R61"><label>61</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Dunbrack</surname><given-names>RL</given-names><suffix>Jr</suffix></name></person-group><article-title>What’s wrong with AlphaFold’s score and how to fix it</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.1101/2025.02.10.637595</pub-id></element-citation></ref><ref id="R62"><label>62</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bryant</surname><given-names>P</given-names></name><name><surname>Noé</surname><given-names>F</given-names></name></person-group><article-title>Improved protein complex prediction with AlphaFold-multimer by denoising the MSA profile</article-title><source>PLoS Comput Biol</source><year>2024</year><volume>20</volume><elocation-id>e1012253</elocation-id><pub-id pub-id-type="doi">10.1371/journal.pcbi.1012253</pub-id><pub-id pub-id-type="pmcid">PMC11302914</pub-id><pub-id pub-id-type="pmid">39052676</pub-id></element-citation></ref><ref id="R63"><label>63</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname><given-names>C‘nan</given-names></name><name><surname>Denzler</surname><given-names>LM</given-names></name><name><surname>Hood</surname><given-names>OEC</given-names></name><name><surname>Martin</surname><given-names>ACR</given-names></name></person-group><article-title>Do antibody CDR loops change conformation upon binding?</article-title><source>MAbs</source><year>2024</year><volume>16</volume><elocation-id>2322533</elocation-id><pub-id pub-id-type="doi">10.1080/19420862.2024.2322533</pub-id><pub-id pub-id-type="pmcid">PMC10939163</pub-id><pub-id pub-id-type="pmid">38477253</pub-id></element-citation></ref><ref id="R64"><label>64</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Park</surname><given-names>E</given-names></name><name><surname>Izadi</surname><given-names>S</given-names></name></person-group><article-title>Molecular surface descriptors to predict antibody developability: sensitivity to parameters, structure models, and conformational sampling</article-title><source>MAbs</source><year>2024</year><volume>16</volume><elocation-id>2362788</elocation-id><pub-id pub-id-type="doi">10.1080/19420862.2024.2362788</pub-id><pub-id pub-id-type="pmcid">PMC11168226</pub-id><pub-id pub-id-type="pmid">38853585</pub-id></element-citation></ref><ref id="R65"><label>65</label><element-citation publication-type="book"><person-group person-group-type="author"><name><surname>Smorodina</surname><given-names>E</given-names></name><name><surname>Greiff</surname><given-names>V</given-names></name><name><surname>Akbar</surname><given-names>R</given-names></name></person-group><source>DINO: dynamics-informed dataset to overcome the limitations of static molecular data in AI-driven drug discovery</source><publisher-name>NeurIPS 2025 AI for Science Workshop</publisher-name><year>2025</year></element-citation></ref><ref id="R66"><label>66</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Della Pia</surname><given-names>EA</given-names></name><name><surname>Martinez</surname><given-names>KL</given-names></name></person-group><article-title>Single domain antibodies as a powerful tool for high quality surface plasmon resonance studies</article-title><source>PLoS One</source><year>2015</year><volume>10</volume><elocation-id>e0124303</elocation-id><pub-id pub-id-type="doi">10.1371/journal.pone.0124303</pub-id><pub-id pub-id-type="pmcid">PMC4378939</pub-id><pub-id pub-id-type="pmid">25822527</pub-id></element-citation></ref><ref id="R67"><label>67</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Jandova</surname><given-names>Z</given-names></name><name><surname>Vargiu</surname><given-names>AV</given-names></name><name><surname>Bonvin</surname><given-names>AMJJ</given-names></name></person-group><article-title>Native or Non-Native Protein-Protein Docking Models? Molecular Dynamics to the Rescue</article-title><source>J Chem Theory Comput 1</source><year>2021</year><volume>7</volume><fpage>5944</fpage><lpage>5954</lpage><pub-id pub-id-type="doi">10.1021/acs.jctc.1c00336</pub-id><pub-id pub-id-type="pmcid">PMC8444332</pub-id><pub-id pub-id-type="pmid">34342983</pub-id></element-citation></ref><ref id="R68"><label>68</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Janin</surname><given-names>J</given-names></name><etal/></person-group><article-title>CAPRI: a Critical Assessment of PRedicted Interactions</article-title><source>Proteins</source><year>2003</year><volume>52</volume><fpage>2</fpage><lpage>9</lpage><pub-id pub-id-type="pmid">12784359</pub-id></element-citation></ref><ref id="R69"><label>69</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Verburgt</surname><given-names>J</given-names></name><name><surname>Kihara</surname><given-names>D</given-names></name></person-group><article-title>Benchmarking of structure refinement methods for protein complex models</article-title><source>Proteins</source><year>2022</year><volume>90</volume><fpage>83</fpage><lpage>95</lpage><pub-id pub-id-type="doi">10.1002/prot.26188</pub-id><pub-id pub-id-type="pmcid">PMC8671191</pub-id><pub-id pub-id-type="pmid">34309909</pub-id></element-citation></ref><ref id="R70"><label>70</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Chaves</surname><given-names>EJF</given-names></name><etal/></person-group><article-title>Estimating Absolute Protein-Protein Binding Free Energies by a Super Learner Model</article-title><source>J Chem Inf Model</source><year>2025</year><volume>65</volume><fpage>2602</fpage><lpage>2609</lpage><pub-id pub-id-type="doi">10.1021/acs.jcim.4c01641</pub-id><pub-id pub-id-type="pmcid">PMC11898044</pub-id><pub-id pub-id-type="pmid">39973292</pub-id></element-citation></ref><ref id="R71"><label>71</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Chaudhury</surname><given-names>S</given-names></name><name><surname>Lyskov</surname><given-names>S</given-names></name><name><surname>Gray</surname><given-names>JJ</given-names></name></person-group><article-title>PyRosetta: a script-based interface for implementing molecular modeling algorithms using Rosetta</article-title><source>Bioinformatics</source><year>2010</year><volume>26</volume><fpage>689</fpage><lpage>691</lpage><pub-id pub-id-type="doi">10.1093/bioinformatics/btq007</pub-id><pub-id pub-id-type="pmcid">PMC2828115</pub-id><pub-id pub-id-type="pmid">20061306</pub-id></element-citation></ref><ref id="R72"><label>72</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sampson</surname><given-names>JM</given-names></name><etal/></person-group><article-title>Robust Prediction of Relative Binding Energies for Protein-Protein Complex Mutations Using Free Energy Perturbation Calculations</article-title><source>J Mol Biol</source><year>2024</year><volume>436</volume><elocation-id>168640</elocation-id><pub-id pub-id-type="doi">10.1016/j.jmb.2024.168640</pub-id><pub-id pub-id-type="pmcid">PMC11339910</pub-id><pub-id pub-id-type="pmid">38844044</pub-id></element-citation></ref><ref id="R73"><label>73</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Tennenhouse</surname><given-names>A</given-names></name><etal/></person-group><article-title>Energy-guided combinatorial co-optimization of antibody affinity and stability</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.1101/2025.11.26.690765</pub-id></element-citation></ref><ref id="R74"><label>74</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Sarma</surname><given-names>S</given-names></name><etal/></person-group><article-title>Can We Extract Physics-like Energies from Generative Protein Diffusion Models?</article-title><source>bioRxiv</source><year>2025</year><pub-id pub-id-type="doi">10.1101/2025.11.28.690021</pub-id></element-citation></ref><ref id="R75"><label>75</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Ursu</surname><given-names>E</given-names></name><etal/></person-group><article-title>Training data composition determines machine learning generalization and biological rule discovery</article-title><source>bioRxiv</source><year>2024</year><elocation-id>2024.06.17.599333</elocation-id><pub-id pub-id-type="doi">10.1101/2024.06.17.599333</pub-id></element-citation></ref><ref id="R76"><label>76</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>O’Donnell</surname><given-names>TJ</given-names></name><etal/></person-group><article-title>Reading the repertoire: Progress in adaptive immune receptor analysis using machine learning</article-title><source>Cell Syst</source><year>2024</year><volume>15</volume><fpage>1168</fpage><lpage>1189</lpage><pub-id pub-id-type="pmid">39701034</pub-id></element-citation></ref><ref id="R77"><label>77</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Mason</surname><given-names>DM</given-names></name><name><surname>Reddy</surname><given-names>ST</given-names></name></person-group><article-title>Predicting adaptive immune receptor specificities by machine learning is a data generation problem</article-title><source>Cell Syst</source><year>2024</year><volume>15</volume><fpage>1190</fpage><lpage>1197</lpage><pub-id pub-id-type="pmid">39701035</pub-id></element-citation></ref><ref id="R78"><label>78</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Janusz</surname><given-names>B</given-names></name><etal/></person-group><article-title>AbDesign: database of point mutants of antibodies with associated structures reveals poor generalization of binding predictions from machine learning models</article-title><source>MAbs</source><year>2025</year><volume>17</volume><elocation-id>2567319</elocation-id><pub-id pub-id-type="doi">10.1080/19420862.2025.2567319</pub-id><pub-id pub-id-type="pmcid">PMC12520099</pub-id><pub-id pub-id-type="pmid">41058476</pub-id></element-citation></ref><ref id="R79"><label>79</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Czerwiński</surname><given-names>A</given-names></name><etal/></person-group><article-title>ASD: antigen-specific antibody database</article-title><source>MAbs</source><year>2026</year><volume>18</volume><elocation-id>2623330</elocation-id><pub-id pub-id-type="doi">10.1080/19420862.2026.2623330</pub-id><pub-id pub-id-type="pmcid">PMC12915772</pub-id><pub-id pub-id-type="pmid">41689452</pub-id></element-citation></ref><ref id="R80"><label>80</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Bio</surname><given-names>A-A</given-names></name></person-group><article-title>Overconfident: What Structure Prediction Confidence Scores Tell Us About Binding</article-title><source>To Affinity And Beyond</source><year>2026</year><comment><ext-link ext-link-type="uri" xlink:href="https://aalphabio.substack.com/p/overconfident-what-structure-prediction">https://aalphabio.substack.com/p/overconfident-what-structure-prediction</ext-link></comment></element-citation></ref><ref id="R81"><label>81</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Schneider</surname><given-names>C</given-names></name><name><surname>Raybould</surname><given-names>MIJ</given-names></name><name><surname>Deane</surname><given-names>CM</given-names></name></person-group><article-title>SAbDab in the age of biotherapeutics: updates including SAbDab-nano, the nanobody structure tracker</article-title><source>Nucleic Acids Res</source><year>2022</year><volume>50</volume><fpage>D1368</fpage><lpage>D1372</lpage><pub-id pub-id-type="doi">10.1093/nar/gkab1050</pub-id><pub-id pub-id-type="pmcid">PMC8728266</pub-id><pub-id pub-id-type="pmid">34986602</pub-id></element-citation></ref><ref id="R82"><label>82</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname><given-names>Y</given-names></name><etal/></person-group><article-title>A comprehensive antigen-antibody complex database unlocking insights into interaction interface</article-title><source>Elife</source><year>2025</year><volume>14</volume><pub-id pub-id-type="doi">10.7554/eLife.104934</pub-id><pub-id pub-id-type="pmcid">PMC12097784</pub-id><pub-id pub-id-type="pmid">40402563</pub-id></element-citation></ref><ref id="R83"><label>83</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Harris</surname><given-names>CR</given-names></name><etal/></person-group><article-title>Array programming with NumPy</article-title><source>Nature</source><year>2020</year><volume>585</volume><fpage>357</fpage><lpage>362</lpage><pub-id pub-id-type="doi">10.1038/s41586-020-2649-2</pub-id><pub-id pub-id-type="pmcid">PMC7759461</pub-id><pub-id pub-id-type="pmid">32939066</pub-id></element-citation></ref><ref id="R84"><label>84</label><element-citation publication-type="journal"><person-group person-group-type="author"><name><surname>Virtanen</surname><given-names>P</given-names></name><etal/></person-group><article-title>SciPy 1.0: fundamental algorithms for scientific computing in Python</article-title><source>Nat Methods</source><year>2020</year><volume>17</volume><fpage>261</fpage><lpage>272</lpage><pub-id pub-id-type="doi">10.1038/s41592-019-0686-2</pub-id><pub-id pub-id-type="pmcid">PMC7056644</pub-id><pub-id pub-id-type="pmid">32015543</pub-id></element-citation></ref><ref id="R85"><label>85</label><element-citation publication-type="book"><collab>The pandas development team</collab><source>Pandas-Dev/pandas: Pandas</source><publisher-name>Zenodo</publisher-name><year>2026</year><pub-id pub-id-type="doi">10.5281/ZENODO.3509134</pub-id></element-citation></ref><ref id="R86"><label>86</label><element-citation publication-type="confproc"><person-group person-group-type="author"><name><surname>McKinney</surname><given-names>W</given-names></name></person-group><source>Data Structures for Statistical Computing in Python</source><conf-name>Proceedings of the Python in Science Conference 56-61</conf-name><conf-sponsor>SciPy</conf-sponsor><year>2010</year></element-citation></ref></ref-list></back><floats-group><fig id="F1" position="float"><label>Figure 1</label><caption><title>Paratope-epitope identification challenge.</title><p>Schematic overview of the real-versus-shuffled nanobody-antigen pairing strategy used to assess whether computational methods can correctly prioritize cognate paratope-epitope interactions. Starting from 106 experimentally resolved nanobody-antigen complexes deposited in the PDB (left), nanobodies (VHH) and antigens are separated and then recombined pairwise. Original nanobody-antigen pairings correspond to real complexes (<xref ref-type="supplementary-material" rid="SD1">Supplementary Table 1</xref>), while shuffled nanobody-antigen combinations form shuffled (non-cognate) complexes. The central grid illustrates representative examples of real (✔) and shuffled (✘) pairings, highlighting cases where a nanobody paratope (“para”) is presented with the correct or incorrect antigen epitope (“epi”). Nanobodies are shown in purple/magenta, antigens in blue/cyan, with paratopes and epitopes highlighted. The benchmark tests whether computational metrics can (i) distinguish real from shuffled complexes, (ii) identify the correct epitope for a given paratope, and (iii) penalize incorrect binding modes that are geometrically plausible but biologically incorrect.</p></caption><graphic xlink:href="EMS215481-f001"/></fig><fig id="F2" position="float"><label>Figure 2</label><caption><title>Current structure prediction tools and scores do not capture nanobody-antigen specificity.</title><p>A. <italic>Global interaction score landscapes reveal limited specificity across prediction tools.</italic> Heatmaps show the best predicted interface confidence scores (ipTM) out of 50 replicas (structure prediction samples) per each system for all pairwise VHH-antigen combinations generated by AF3, Boltz-2, and Chai-1 (fixed seed, MSAs were generated automatically for everything except Chai-1, instead embeddings were used with Chai-1). Rows correspond to VHHs and columns to antigens. For all three tools, high-scoring interactions are broadly distributed throughout the matrices, with no clear enrichment along the diagonal corresponding to cognate (real) complexes. AF3 and Chai-1 display sparse, heterogeneous high-score patterns, whereas Boltz-2 assigns uniformly high scores across most combinations, irrespective of biological relevance. Color scales indicate ipTM values for each tool (darker color -higher values). B. <italic>Precision-recall analysis shows limited discrimination between real and shuffled complexes.</italic> Precision-recall (PR) curves summarize performance across all possible ipTM thresholds for AF3, Boltz-2, and Chai-1. Because positives are rare in the all-vs-all evaluation, the dashed horizontal line indicates the random baseline corresponding to class prevalence (~0.011 precision). PR-AUC (Average Precision, AP) values quantify overall discrimination ability, with AF3 achieving the highest performance (AP=0.187), followed by Chai-1 (AP=0.067) and Boltz-2 (AP=0.026). Reference thresholds (t1-t3; ipTM ≥0.3, ≥0.6, ≥0.8) are marked on the curves to illustrate the trade-off between precision and recall: lower thresholds increase recall but introduce many false positives, whereas higher thresholds improve precision at the cost of sharply reduced recall. Across all tools, curves remain close to the random baseline, indicating limited ability to distinguish real from shuffled complexes. C. <italic>Score distributions overlap between real and shuffled complexes across dataset splits.</italic> Violin plots summarize ipTM score distributions for real and shuffled complexes across training, test, and mixed (both) splits, defined based on the original train-test split of each tool. Real and shuffled distributions strongly overlap for all predictors, with comparable means and variances across splits. AF3 exhibits bimodal distributions for real complexes but substantial overlap with shuffled scores, Boltz-2 assigns consistently high scores to both classes, and Chai-1 yields uniformly low scores, indicating limited discriminatory power across all methods. D. <italic>Structural accuracy scores are non-discriminative across prediction tools.</italic> Violin plots show the distribution of clash scores for the top-ranked predicted nanobody-antigen complex generated by AF3, Boltz-2, and Chai-1, evaluated across training-leakage, test, and mixed (“both”) dataset splits. “Both” data split refer to either the antigen or the VHH present in the training of the model. Clash scores quantify structural quality by accounting for both the number of steric clashes and protein length, with lower values indicating better geometric plausibility. For all tools, clash score distributions are broad and strongly overlapping between real and shuffled complexes, indicating that basic steric quality does not distinguish cognate from non-cognate pairings.</p></caption><graphic xlink:href="EMS215481-f002"/></fig><fig id="F3" position="float"><label>Figure 3</label><caption><title>Cross-tool disagreement and confidence calibration of nanobody-antigen docking predictions.</title><p>A. <italic>Confidence calibration differs systematically across ML structure prediction tools.</italic> Scatter plots show DockQ versus ipTM for AF3, Boltz-2, and Chai-1 under three sampling regimes: best DockQ, initial sample (sample<sub>0</sub>), and worst DockQ. Points are colored by the official DockQ quality category (incorrect, acceptable, medium, high) and sized according to the saturation range (ΔDockQ from sample<sub>0</sub> to best). Yellow stars mark quadrant extremes, highlighting overconfident failures (high ipTM, low DockQ) and underconfident successes (low ipTM, high DockQ). Dashed lines indicate DockQ=0.23 (acceptable quality threshold) and ipTM=0.5 (confidence threshold). pearson correlation coefficients (r) are shown for each panel. AF3 exhibits the strongest calibration, Boltz-2 shows systematic overconfidence, and Chai-1 displays systematic underconfidence, with calibration degrading substantially in low-sampling regimes. B. <italic>Saturation outcome categories reveal when sampling rescues docking quality.</italic> Line plots show per-system trajectories from the initial prediction (sample<sub>0</sub>) to the best of the sampled predictions, grouped into three outcome classes: non-improvers (red) remain below the acceptable-quality threshold (DockQ&lt;0.23) across all samples; improvers (green) are rescued from incorrect to acceptable quality (cross DockQ=0.23); and neutral systems (grey) are already acceptable at sample<sub>0</sub> (DockQ≥0.23) and may further refine with sampling. Systems with known structural artifacts and runs that failed to complete were excluded (14 systems; see <xref ref-type="sec" rid="S13">Methods</xref>). A comprehensive list of improvement categories is provided in <xref ref-type="supplementary-material" rid="SD3">Supplementary Table 4</xref>. C. <italic>DockQ gains from sampling do not translate into confidence-score gains.</italic> Scatter plots show, for each system and model, the relationship between ΔDockQ (best -sample<sub>0</sub>) and ΔipTM (best - sample<sub>0</sub>), with linear regression fits and 95% confidence intervals. Near-zero correlations indicate that confidence scores largely fail to track structural quality improvements achieved through saturation sampling (AF3: N=106, r=-0.027; Boltz-2: N=106, r=-0.040; Chai-1: N=103, r=-0.019), consistent with confidence being “locked in” to early trajectory choices rather than reflecting refinement.</p></caption><graphic xlink:href="EMS215481-f001"/></fig><fig id="F4" position="float"><label>Figure 4</label><caption><title>Epitope recovery and enrichment fail to distinguish cognate from non-cognate nanobody-antigen interactions.</title><p>A. <italic>Epitope recall across real and shuffled complexes.</italic> Heatmaps show epitope recall for predicted nanobody-antigen complexes, defined as the fraction of experimentally determined epitope residues recovered by the model according to the equation shown. An epitope residue is considered recovered if it is in contact with the VHH in at least 50% of the n=50 stochastic replicate predictions. Diagonal entries correspond to real (cognate) nanobody-antigen pairs, whereas off-diagonal entries represent shuffled (non-cognate) complexes. Across all models, shuffled complexes frequently achieve epitope recall comparable to real complexes, indicating limited specificity of epitope recovery. B. <italic>Epitope recall across real and shuffled complexes.</italic> Distribution of epitope recall across real and shuffled complexes. Boxplots summarize epitope recall values for real and shuffled VHH-antigen pairs, stratified by train-leakage and true test sets for each model. Shuffled complexes frequently overlap or exceed the recall observed for real complexes, demonstrating that high epitope recall alone does not reliably distinguish cognate interactions. The number of points when epitope predicted correctly (points &gt;0): AF3: train - n=12, test - n=29, Boltz-2: train - n=28, test - n=17, Chai-1: train - n=9, test - n=26. C. <italic>Epitope enrichment significance saturates rapidly with stochastic sampling.</italic> Line plots show the fraction of VHH-antigen pairs declared significantly enriched (FDR &lt; 0.05) as a function of the number of stochastic predictions per complex, up to n = 50. Six distributions are evaluated per model: real complexes split into train-leakage and true test sets, and shuffled complexes split according to whether the VHH, the antigen, both, or neither were present in the training data. Across all three models, enrichment significance increases rapidly and saturates by approximately 10-15 replicates, with shuffled complexes exhibiting saturation behavior comparable to real complexes. Notably, Chai-1 shows near-complete saturation for shuffled complexes at low replicate counts, indicating a strong epitope-seeking bias that reduces the discriminative value of enrichment-based metrics. These trends suggest that increasing ensemble size primarily amplifies pre-existing contact biases rather than improving identification of true cognate interactions. D. <italic>Structural consistency correlates with epitope recovery but not interaction correctness.</italic> Scatter plots show the relationship between epitope recall and structural consistency across predictions, quantified as the mean pairwise RMSD of all n=50 VHH replicates after superposition on the antigen (Cα atoms only). A case study of the real complex of PDB 9ETJ is shown-predicted by the different models. The first 10 structures from the 50 replicates are shown, where the predicted complexes are superimposed on the antigen from the first predicted complex, and the predicted nanobodies are shown complexed to that antigen respectively. Points correspond to individual VHH-antigen complexes, with linear fits shown separately for train-leakage and true test sets. Reported values indicate pearson correlation coefficients. Across all models, epitope recall is negatively correlated with RMSD, indicating improved epitope recovery for complexes predicted with consistent orientations; however, this relationship holds for both real and shuffled complexes and does not imply correct paratope-epitope pairing. E. <italic>Structural diversity of predicted complexes.</italic> Structural diversity of predicted complexes. Violin plots show RMSD distributions for real, shuffled, and mixed (“both”) complexes across train-leakage and true test sets for each model. Similar RMSD distributions between real and shuffled conditions indicate that structural consistency alone is insufficient to discriminate cognate from non-cognate nanobody-antigen interactions.</p></caption><graphic xlink:href="EMS215481-f004"/></fig><fig id="F5" position="float"><label>Figure 5</label><caption><title>Computational cost, sampling efficiency, and seed-dependent optimization in nanobody-antigen structure prediction.</title><p><italic>A. Energy consumption differs substantially across tools and sampling depth.</italic> Median energy usage per predicted complex is shown as a function of diffusion sampling depth (N=1, 10, 25, 50, 100), with each saturation level corresponding to an independent random seed. Shaded regions indicate interquartile ranges across 106 systems. At the highest saturation level (N=100), median energy consumption spans from 23.0 Wh for AF3 to 82.9 Wh for Chai-1, reflecting distinct baseline and scaling behaviors across architectures. B. <italic>Saturation sampling improves best-case structural quality with model-dependent magnitude.</italic> Median maximum DockQ is shown as a function of sampling depth, with shaded regions indicating interquartile ranges across systems. Because each saturation level uses an independent seed, curves reflect combined effects of seed selection and increased sampling. Total improvements from baseline to N=100 (ΔDockQ) are indicated for each model, with AF3 showing the largest gain. Horizontal dashed lines mark DockQ quality thresholds separating incorrect (&lt;0.23), acceptable (0.23-0.49), medium (0.49-0.80), and high (≥0.80) quality regimes. C. <italic>Quality gains per unit energy peak at low saturation levels.</italic> Bar plots show the median gain in maximum DockQ between consecutive saturation levels across systems. The largest improvements occur between baseline and low saturation, with progressively smaller gains at higher sampling depths. Each transition reflects both a change in seed and an increase in the number of samples. Line plot shows that efficiency frontiers relate cumulative DockQ improvement relative to baseline (seed 1, N=1) to cumulative energy expenditure. Each point corresponds to a saturation level and its associated seed. Steeper initial slopes indicate high computational efficiency at low sampling depths, followed by pronounced flattening at higher energy costs, consistent with diminishing returns beyond moderate saturation. D. <italic>Saturation sampling reduces incorrect predictions but plateaus at higher depth.</italic> Stacked bar plots show the distribution of DockQ quality categories across saturation levels, with each level corresponding to an independent seed and its associated sampling depth. The fraction of incorrect predictions (DockQ&lt;0.23) decreases substantially at low saturation but shows limited further reduction at higher sampling depths. E. <italic>Seed selection introduces substantial variability in prediction quality.</italic> Boxplots show the range of DockQ values obtained from five independent seeds using only the first sample (sample<sub>0</sub>), isolating seed effects from saturation. For each system, the DockQ range (maximum minus minimum across seeds) highlights that a subset of predictions exhibits large seed-dependent variation without additional computational cost. F. <italic>Different seeds explore distinct solution landscapes.</italic> Cross-seed saturation trajectories for AF3 show median cumulative maximum DockQ across systems as a function of sample number within each seed’s saturation run. Seeds that achieve higher quality early maintain their advantage throughout sampling, whereas poor initial seeds plateau at lower quality ceilings, indicating that seed selection and saturation sampling act as orthogonal optimization mechanisms. Corresponding analyses for other models are shown in <xref ref-type="supplementary-material" rid="SD4">Supplementary Figure 6</xref>.</p></caption><graphic xlink:href="EMS215481-f005"/></fig></floats-group></article>