<?xml version="1.0" encoding="UTF-8"?><article xml:lang="en" article-type="research-article"><front><journal-meta><journal-id journal-id-type="pmc-domain-id">2873</journal-id><journal-id journal-id-type="pmc-domain">ncomms</journal-id><journal-title-group><journal-title>Nature Communications</journal-title><abbrev-journal-title>Nat Commun</abbrev-journal-title></journal-title-group><publisher><publisher-name>Nature Publishing Group</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="pmcid">PMC9750992</article-id><article-id pub-id-type="pmcaid">9750992</article-id><article-id pub-id-type="pmcaiid">9750992</article-id><article-id pub-id-type="pmid">36517480</article-id><article-id pub-id-type="doi">10.1038/s41467-022-35422-y</article-id><title-group><article-title>Merging enzymatic and synthetic chemistry with computational synthesis planning</article-title></title-group><contrib-group content-type="author"><contrib><name name-style="western"><surname>Levin</surname><given-names initials="I">Itai</given-names></name><xref ref-type="aff" rid="Aff1">1</xref><xref ref-type="aff" rid="Aff2">2</xref></contrib><contrib><name name-style="western"><surname>Liu</surname><given-names initials="M">Mengjie</given-names></name><xref ref-type="aff" rid="Aff2">2</xref></contrib><contrib><name name-style="western"><surname>Voigt</surname><given-names initials="CA">Christopher A</given-names></name><xref ref-type="aff" rid="Aff1">1</xref></contrib><contrib><name name-style="western"><surname>Coley</surname><given-names initials="CW">Connor W</given-names></name><xref ref-type="aff" rid="Aff2">2</xref><xref ref-type="aff" rid="Aff3">3</xref><xref ref-type="author-notes" rid="_fncrsp93pmc__">✉</xref></contrib></contrib-group><aff id="Aff1"><label>1</label>Synthetic Biology Center, Department of Biological Engineering, Massachusetts Institute of Technology, Cambridge, MA USA </aff><aff id="Aff2"><label>2</label>Department of Chemical Engineering, Massachusetts Institute of Technology, Cambridge, MA USA </aff><aff id="Aff3"><label>3</label>Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MA USA </aff><author-notes><fn id="_fncrsp93pmc__"><label>✉</label><p>Corresponding author.</p></fn></author-notes><pub-date><day>14</day><month>12</month><year>2022</year></pub-date><volume>13</volume><fpage>7747</fpage><page-range>7747</page-range><pub-history><event event-type="pmc-release"><date><day>16</day><month>12</month><year>2022</year></date></event></pub-history><permissions><copyright-statement>© The Author(s) 2022</copyright-statement><license><license-p><bold>Open Access</bold> This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons license, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons license and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this license, visit <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://creativecommons.org/licenses/by/4.0/" ext-link-type="uri">http://creativecommons.org/licenses/by/4.0/</ext-link>.</license-p></license></permissions><self-uri xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="41467_2022_Article_35422.pdf" content-type="pmc-pdf"><?cloudpmc-path 46c9/9750992/befb946c5f6d/41467_2022_Article_35422.pdf?><?cloudpmc-bucket app?><?size 1466269?></self-uri><abstract id="Abs1"><title>Abstract</title><p id="Par1">Synthesis planning programs trained on chemical reaction data can design efficient routes to new molecules of interest, but are limited in their ability to leverage rare chemical transformations. This challenge is acute for enzymatic reactions, which are valuable due to their selectivity and sustainability but are few in number. We report a retrosynthetic search algorithm using two neural network models for retrosynthesis–one covering 7984 enzymatic transformations and one 163,723 synthetic transformations–that balances the exploration of enzymatic and synthetic reactions to identify hybrid synthesis plans. This approach extends the space of retrosynthetic moves by thousands of uniquely enzymatic one-step transformations, discovers routes to molecules for which synthetic or enzymatic searches find none, and designs shorter routes for others. Application to (-)-Δ<sup>9</sup> tetrahydrocannabinol (THC) (dronabinol) and R,R-formoterol (arformoterol) illustrates how our strategy facilitates the replacement of metal catalysis, high step counts, or costly enantiomeric resolution with more elegant hybrid proposals.</p><sec id="kwd-group1" sec-type="kwd-group" disp-level="2"><p><bold>Subject terms:</bold> Biosynthesis, Cheminformatics, Chemical synthesis</p></sec></abstract><abstract id="Abs2" abstract-type="web-summary"><p id="Par2">The identification of synthetic routes combining enzymatic and non-enzymatic reactions has been challenging and requiring expert knowledge. Here, the authors describe a computational retrosynthetic approach relying on neural network models for planning synthetic routes using both strategies.</p></abstract><custom-meta-group><custom-meta><meta-name>status</meta-name><meta-value>released</meta-value></custom-meta><custom-meta><meta-name>display-pdf</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>is-olf</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-manuscript</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-preprint</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-journal-matter</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-scanned</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-retracted</meta-name><meta-value>no</meta-value></custom-meta></custom-meta-group></article-meta><notes notes-type="article-notes"><sec id="historyarticle-meta1" sec-type="history" disp-level="2"><p>Received 2022 Sep 13; Accepted 2022 Nov 30; Collection date 2022.</p></sec></notes></front><body><sec id="Sec1" disp-level="1"><title>Introduction</title><p id="Par3">Enzymatic and non-enzymatic synthetic organic (“synthetic”) reactions can be synergistically combined to leverage the unique strengths of each. Many efficient hybrid syntheses have been reported from the discovery to the process scale<sup><xref rid="CR1" ref-type="bibr">1</xref>–<xref rid="CR8" ref-type="bibr">8</xref></sup>. Enzymes can be used to introduce stereochemistry at key points of an otherwise synthetic process. An industrially-relevant example is the evolution and use of a transaminase to selectively catalyze the formation of the chiral amine in sitagliptin, an anti-diabetes drug, from the chemically-derived pro-sitagliptin<sup><xref rid="CR9" ref-type="bibr">9</xref></sup>. Enzymes can also catalyze reactions with superior regioselectivity. This was leveraged for the synthesis of the antiviral compound islatravir, where a purine nucleoside phosphorylase and a phosphopentomutase were evolved to catalyze the regio- and stereoselective installation of an unnatural purine moiety on a chemically<sup><xref rid="CR10" ref-type="bibr">10</xref></sup> or enzymatically<sup><xref rid="CR11" ref-type="bibr">11</xref></sup> synthesized unprotected, unnatural deoxyribose analog. Combining enzymatic and synthetic steps can unlock a more efficient process overall than using only one or the other. A striking example of this is the implementation of an ex vivo enzymatic cascade to convert chemically fixed carbon dioxide into starch at higher rates than maize<sup><xref rid="CR12" ref-type="bibr">12</xref></sup>. In addition to enabling unique or uniquely selective transformations, enzymes can improve the environmental sustainability of a chemical process as they are renewable and biodegradable catalysts that can operate under mild conditions<sup><xref rid="CR13" ref-type="bibr">13</xref></sup>.</p><p id="Par4">However, identifying synthetic routes that use both enzymatic and synthetic organic reaction steps remains a largely manual, intuition-driven process despite the emergence of computer-aided synthesis planning (CASP) tools.<sup><xref rid="CR14" ref-type="bibr">14</xref>–<xref rid="CR17" ref-type="bibr">17</xref></sup> Retrosynthesis is a search through the space of possible chemical precursors. The starting position is a target molecule and the goal is to find a path to viable starting materials. The search space grows exponentially with the search depth, so a brute force enumeration of all precursors quickly becomes computationally intractable. Most CASP algorithms emulate chemists’ retrosynthetic analysis process; starting from the target molecule, the algorithms recursively generate the most plausible precursors until pathways to suitably simple starting materials are found. As has been previously reviewed<sup><xref rid="CR18" ref-type="bibr">18</xref>–<xref rid="CR23" ref-type="bibr">23</xref></sup>, CASP tools differ along many axes, including how one-step retrosynthetic moves are proposed, the class of models used to rank single-step retrosynthetic suggestions, what is considered an acceptable starting material, and the algorithms used to efficiently piece together single steps to navigate a retrosynthetic search space.</p><p id="Par5">A primary differentiator is whether CASP methods are template-based or template-free. Template-based models assign scores to a pre-defined set of reaction templates—generalized reaction rules—which can be applied to the input target molecule to produce precursors. When templates are algorithmically extracted, each reaction suggested by a template-based model can be linked to precedent examples to be interpreted and reviewed by a chemist. Examples of template-based synthetic organic tools include ASKCOS,<sup><xref rid="CR24" ref-type="bibr">24</xref></sup> which uses a set of templates automatically extracted from Reaxys, and Synthia<sup><xref rid="CR21" ref-type="bibr">21</xref></sup>, which uses expert-curated templates; analogous data-driven and expert bioretrosynthetic tools include Retropath<sup><xref rid="CR25" ref-type="bibr">25</xref>,<xref rid="CR26" ref-type="bibr">26</xref></sup>, which uses templates extracted from metabolic pathway databases<sup><xref rid="CR27" ref-type="bibr">27</xref></sup> and RetroBioCat<sup><xref rid="CR28" ref-type="bibr">28</xref></sup>, which focuses on industrially-relevant biocatalytic reactions. Template-free models instead learn to generate reactant molecules from an input product molecule end-to-end<sup><xref rid="CR29" ref-type="bibr">29</xref>,<xref rid="CR30" ref-type="bibr">30</xref></sup>. The lack of pre-defined reaction rules theoretically allows template-free models to predict novel reactions, but complicates the task of linking predictions to existing reaction data. Examples of this approach include IBM’s RXN<sup><xref rid="CR31" ref-type="bibr">31</xref>,<xref rid="CR32" ref-type="bibr">32</xref></sup> and BioNavi-NP<sup><xref rid="CR33" ref-type="bibr">33</xref></sup>, both of which employ sequence-to-sequence Transformer models<sup><xref rid="CR32" ref-type="bibr">32</xref></sup> to generate precursor SMILES strings directly from product SMILES strings.</p><p id="Par6">Data-driven CASP tools are limited in finding hybrid synthesis routes because distinct sets of CASP software tools have been designed for fully synthetic organic synthesis planning<sup><xref rid="CR24" ref-type="bibr">24</xref>,<xref rid="CR32" ref-type="bibr">32</xref>,<xref rid="CR34" ref-type="bibr">34</xref>–<xref rid="CR38" ref-type="bibr">38</xref></sup> and for fully enzymatic synthesis planning<sup><xref rid="CR16" ref-type="bibr">16</xref>,<xref rid="CR26" ref-type="bibr">26</xref>,<xref rid="CR33" ref-type="bibr">33</xref>,<xref rid="CR39" ref-type="bibr">39</xref></sup>. CASP tools such as ASKCOS<sup><xref rid="CR24" ref-type="bibr">24</xref></sup> and AiZynthFinder<sup><xref rid="CR37" ref-type="bibr">37</xref></sup> were developed based on sets of reactions such as the USPTO<sup><xref rid="CR40" ref-type="bibr">40</xref></sup> or Reaxys<sup><xref rid="CR41" ref-type="bibr">41</xref></sup> where enzymatic reactions represent a fraction of the total dataset (e.g., Reaxys contains ~5 × 10<sup>4</sup> enzymatic reactions compared to &gt;10<sup>7</sup> total reactions), whereas enzymatic CASP tools such as BNICE.ch<sup><xref rid="CR42" ref-type="bibr">42</xref></sup>, RetroPath<sup><xref rid="CR25" ref-type="bibr">25</xref>,<xref rid="CR26" ref-type="bibr">26</xref></sup>, or the similarity-based retrosynthesis tool from ref. <xref rid="CR43" ref-type="bibr">43</xref> use reaction databases such as the Kyoto Encyclopedia of Genes and Genomes (KEGG)<sup><xref rid="CR44" ref-type="bibr">44</xref></sup>, MetaNetX<sup><xref rid="CR45" ref-type="bibr">45</xref></sup>, or Rhea<sup><xref rid="CR46" ref-type="bibr">46</xref></sup> that contain exclusively enzymatic reactions. Stitching together the results from a synthetic and enzymatic CASP search for the same molecule is insufficient, as part of the challenge of hybrid synthesis planning is identifying routes where one set of reactions leads to an intermediate that can be used by the other set, yielding a route that would have remained undiscovered if only one set of reactions were considered at a time. Probst et al.<sup><xref rid="CR31" ref-type="bibr">31</xref></sup> use transfer learning to pretrain on synthetic reactions and fine-tune enzymatic reactions, such that the model used can suggest non-enzymatic reactions if it predicts low confidence for enzymatic suggestions. RetroBioCat allows users to manually generate hybrid networks. However, no tool is designed to automatically search hybrid retrosynthesis networks. New CASP algorithms are needed to integrate and balance the two complementary synthesis strategies.</p><p id="Par7">Here, we introduce a synthesis planning algorithm to generate multi-step synthesis plans that leverage the breadth of known synthetic and enzymatic chemistry (Fig. <xref rid="Fig1" ref-type="fig">1</xref>). We trained a template-based, enzymatic retrosynthesis neural network<sup><xref rid="CR47" ref-type="bibr">47</xref></sup> using enzymatic reaction data from the BKMS database<sup><xref rid="CR48" ref-type="bibr">48</xref></sup> to rank single enzymatic retrosynthetic steps. We show that the chemistry captured by this model expands upon chemistry captured by the synthetic chemistry retrosynthesis model from ASKCOS<sup><xref rid="CR24" ref-type="bibr">24</xref></sup>, adding 4169 unique templates. We then designed a multi-step search algorithm that uses both the enzymatic retrosynthesis model and a synthetic retrosynthesis model to prioritize the possible retrosynthetic steps in a way that balances the exploration of enzymatic and synthetic steps. We find that this hybrid search identifies routes to molecules for which no routes are found using only enzymatic or synthetic organic chemistry. Further, the hybrid search identifies shorter pathways where enzymatic steps replace multiple synthetic steps. Finally, we demonstrate how our search algorithm can suggest promising hybrid synthesis plans which were not found otherwise, using dronabinol and arformoterol as case studies.</p><fig id="Fig1" position="float"><?disp-level 2?><label>Fig. 1</label><caption><title>Machine learning approach to hybrid synthesis planning.</title><p><bold>a</bold> Development workflow of the hybrid synthesis planner. A database of enzymatic reactions was parsed into machine-readable format. Reaction templates were algorithmically extracted from the reactions in the database. A neural network template prioritizer<sup><xref rid="CR47" ref-type="bibr">47</xref></sup> was trained to predict the reaction template associated with each product molecule in the reaction database. The enzymatic template prioritizer and a previously trained synthetic template prioritizer<sup><xref rid="CR24" ref-type="bibr">24</xref></sup> are used in tandem to predict hybrid synthesis plans. <bold>b</bold> Multi-model-guided tree search strategy used to explore the retrosynthetic search space for multi-step synthesis planning from an input molecule (yellow circle). The possible retrosynthetic reaction templates (squares) are scored using template prioritizer neural networks. Different colors correspond to different template sets (e.g., synthetic and enzymatic). (i) The leaf node from the highest-scoring path is selected. (ii) The selected retrosynthetic template is applied to the product molecule to generate the predicted precursor. The precursor is added to the search tree and the retrosynthetic templates are scored by their corresponding template prioritizer with the precursor as input. (iii) Visit counts are updated for the explored nodes. The visit counts are used in scoring to balance exploration and exploitation. Steps i, ii, and iii are repeated until a stopping criterion for the search is met. All pathways that connect the input molecule to allowed starting materials (gray circles) are returned.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d32e496" xlink:href="41467_2022_35422_Fig1_HTML.jpg"><?cloudpmc-path blobs/46c9/9750992/1048ebe74d12/41467_2022_35422_Fig1_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 877?><?original-width 1748?><?scaled-height 351?><?scaled-width 699?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="41467_2022_35422_Fig1_HTML.gif"><?cloudpmc-path blobs/46c9/9750992/4329388afb92/41467_2022_35422_Fig1_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec><sec id="Sec2" disp-level="1"><title>Results</title><sec id="Sec3" disp-level="2"><title>Single-step enzymatic retrosynthetic expansion</title><p id="Par8">An enzymatic reaction dataset was curated from the BKMS<sup><xref rid="CR48" ref-type="bibr">48</xref></sup> database. BKMS contains approximately 37,000 enzyme-catalyzed reactions aggregated from BRENDA<sup><xref rid="CR49" ref-type="bibr">49</xref></sup>, the Kyoto Encyclopedia of Genes and Genomes (KEGG)<sup><xref rid="CR44" ref-type="bibr">44</xref></sup>, Metacyc<sup><xref rid="CR50" ref-type="bibr">50</xref></sup>, and SABIO-RK<sup><xref rid="CR51" ref-type="bibr">51</xref></sup>. We processed the reaction data by removing biological cofactors, converting the reactions to standardized SMILES<sup><xref rid="CR52" ref-type="bibr">52</xref></sup> strings, and performing atom-atom mapping to track which atoms in the reactants corresponded to which atoms in the products of each reaction (see Methods). Our final dataset contains 15,309 unique, single-product, atom-mapped reaction SMILES strings.</p><p id="Par9">Using RDChiral<sup><xref rid="CR53" ref-type="bibr">53</xref></sup>, reaction templates summarizing the chemistry of each reaction were automatically extracted from these atom-mapped reactions as generalized SMARTS strings (examples shown in Fig. <xref rid="Fig2" ref-type="fig">2</xref>a–d). A total of 7984 unique reaction templates were sufficient to describe the 15,309 enzymatic reactions. This method generates a single template per reaction with chiral information and a heuristically determined, variable amount of context around the reaction center as opposed to RetroRules<sup><xref rid="CR27" ref-type="bibr">27</xref></sup>, which stores multiple fixed-diameter templates per reaction. The generalized reaction templates approximate the range of possible chemical transformations that enzymes can catalyze by representing only the reaction center and its adjacent context from the reactant and product. This template-based approach was chosen to maintain a link between retrosynthetic suggestions and precedent reactions from the database; this link makes the model’s suggestions more interpretable and actionable as starting points for enzyme selection and optimization.</p><fig id="Fig2" position="float"><?disp-level 3?><label>Fig. 2</label><caption><title>Reaction templates automatically extracted from the BKMS biochemical reaction database.</title><p><bold>a</bold>–<bold>d</bold> Examples of reactions and reaction templates extracted from the BKMS database. Cofactor reactant-product pairs (e.g., SAM and SAH) were automatically identified and removed. Chemical names were converted to SMILES strings, shown here by molecular structures. Reaction templates were automatically extracted from the atom-mapped SMILES string as SMARTS strings, shown here as retrosynthetic reaction fragments. Reaction rules are linked to the reaction database, so metadata such as associated enzymes and EC number for the reaction can be easily retrieved. <bold>e</bold> The number of examples for a reaction template in the reaction database as a function of the rank of the template’s popularity. <bold>f</bold> Fraction of the reactions in the BKMS database that can be described by reaction templates that have greater than a threshold number of examples.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d32e564" xlink:href="41467_2022_35422_Fig2_HTML.jpg"><?cloudpmc-path blobs/46c9/9750992/29fbeb5a2472/41467_2022_35422_Fig2_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1776?><?original-width 2000?><?scaled-height 710?><?scaled-width 800?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="41467_2022_35422_Fig2_HTML.gif"><?cloudpmc-path blobs/46c9/9750992/295f44d2874e/41467_2022_35422_Fig2_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par10">When building template-based models, a common post-processing step is to remove templates for which too few precedent examples exist. In our enzymatic reaction dataset, nearly 80% of reaction templates have only one precedent (Fig. <xref rid="Fig2" ref-type="fig">2</xref>e). Even if rare reactions were assigned more generalized templates and stereochemical information was removed as described in the methods section, requiring that extracted templates have <italic>n</italic> &gt; 1 precedent would make the filtered set of templates unable to describe nearly 20% of reactions in the database. (Fig. <xref rid="Fig2" ref-type="fig">2</xref>f). Hence, we did not remove templates based on the number of precedents to maximize the diversity of enzymatic reactions that were captured.</p><p id="Par11">For each enzymatic reaction, the product molecule and the extracted reaction template used in the dataset to synthesize that molecule were used as input-label pairs to train a multi-layer perceptron (MLP) classification model<sup><xref rid="CR47" ref-type="bibr">47</xref></sup>. The MLP template prioritizer model was trained to predict which template was used to synthesize each product in the training set based on the product’s molecular structure. At inference, given a new molecular structure, the template prioritizer outputs a softmax normalized score for each of the 7984 reaction templates that can be interpreted as the probability that a given template is the best retrosynthetic move from the target molecule. After tuning the model’s hyperparameters using distinct training and validation sets, an MLP model was trained on all of the available data with a fixed number of epochs. This model was used for all subsequent analyses using synthetic targets selected from ZINC, MOSES, and FDA-approved drugs. Including all of the available data from the BKMS dataset improved the model’s ability to predict rare reactions during multi-step pathway prediction.</p></sec><sec id="Sec4" disp-level="2"><title>Comparing enzymatic and synthetic transformations</title><p id="Par12">A simple comparison between the size of the chemical reaction template set (163,723) and the enzymatic reaction template set (7984) supports the observation that synthetic organic chemistry enables a much broader set of transformations than known enzymatic chemistry. However, it is not obvious whether enzymes catalyze different reactions or simply catalyze reactions with improved specificity and efficiency. We sought to better understand what fraction of the reactions in the BKMS dataset are captured by the reaction templates in the Reaxys dataset and whether including enzymatic reactions in a retrosynthetic search could potentially expand the accessible chemical space.</p><p id="Par13">To identify which reactions from BKMS comprise “unique” chemistry, we assessed whether any synthetic reaction templates from Reaxys could reproduce the same reactants given the product molecule. To this end, all of the synthetic reaction templates were applied to each of the product molecules from our BKMS dataset. If any of the reaction templates reproduced the original reactant molecule(s), the enzymatic reaction was marked as recovered and, therefore, not unique. Of the 14,601 single-step, non-spontaneous, non-generic reactions in the enzymatic reaction database, 9095 were recovered with synthetic reaction templates (not considering charge or stereochemistry). Enzyme catalysts may offer enhanced selectivity for these processes, but the chemical transformation could be achieved  without enzymes.</p><p id="Par14">Of the remaining 5506 enzymatic reactions (corresponding to 4169 unique reaction templates), it is likely that some could be achieved with synthetic organic chemistry. Certain reagents, cofactors, and leaving groups are omitted from reaction definitions so reactions can be modeled as single-product (Methods). This is required when performing iterative retrosynthesis. Thus, certain atoms appear only in the reactant side or product side of a reaction definition (Fig. <xref rid="Fig3" ref-type="fig">3</xref>a, b). For example, in 228 of the reactions in our BKMS dataset, the molecular structure corresponding to Coenzyme A appears in the reactants side and not the products side of the reaction. Only a single one of these reactions was recovered by the chemical reaction rules (Supplementary Fig. <xref rid="MOESM1" ref-type="supplementary-material">1)</xref>. This is not because the hydrolysis of a thioester bond represents uniquely enzymatic chemistry. Rather, in the Reaxys dataset, in the reaction rules representing the hydrolysis of thioesters, the acyl group is treated as the eliminated group and the product that is kept is the thiol, whereas in this work, because of our automated identification of common biological cofactors, the thiol (specifically CoA) is defined as the eliminated group and the acyl group is kept (Fig. <xref rid="Fig3" ref-type="fig">3</xref>c). This is more consistent with the role of CoA in biochemical reactions. The presence of eliminated and added groups confounds the automated comparison of chemical transformations across datasets, as it makes truly unique chemistry indistinguishable from the consequences of choices made during the dataset preparation (Fig. <xref rid="Fig3" ref-type="fig">3</xref>c–h).</p><fig id="Fig3" position="float"><?disp-level 3?><label>Fig. 3</label><caption><title>Comparison of synthetic and enzymatic reaction sets.</title><p>Most commonly, <bold>a</bold> eliminated and <bold>b</bold> added substructures in reactions from the BKMS dataset that were not captured by the Reaxys dataset reaction templates. <bold>c</bold>–<bold>j</bold> Reactions from the BKMS dataset that were not recovered by applying any of the Reaxys dataset reaction templates to the product molecules. Reactions where the underlying chemistry exists in the Reaxys dataset but the reaction is not captured because of how eliminated/added groups are handled are highlighted in orange and reactions which are truly not captured by the chemistry in the Reaxys dataset are highlighted in blue.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d32e625" xlink:href="41467_2022_35422_Fig3_HTML.jpg"><?cloudpmc-path blobs/46c9/9750992/5c39246ecb07/41467_2022_35422_Fig3_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1182?><?original-width 1353?><?scaled-height 591?><?scaled-width 676?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="41467_2022_35422_Fig3_HTML.gif"><?cloudpmc-path blobs/46c9/9750992/06a3cee22a81/41467_2022_35422_Fig3_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par15">To estimate a lower bound on the number of unique enzymatic transformations, we identified that of the 3466 enzymatic reactions with no addition or loss of heavy atoms between reactants and products (e.g., excluding reactions with leaving groups), 968 reactions could not be described by the Reaxys reaction templates, corresponding to 824 unique enzymatic reaction templates. Partially due to the constraints placed on this set of reactions, these are predominantly unimolecular templates that include oxidation, reduction, isomerization, and intramolecular cyclization reactions. We show two examples, the reaction of the isochorismate synthase and griseophenone synthase in Fig. <xref rid="Fig3" ref-type="fig">3</xref>i, j. This suggests that the enzymes expand the scope of possible organic transformations in the synthetic chemist’s toolbox and do not merely offer an alternate set of conditions for existing reactions.</p></sec><sec id="Sec5" disp-level="2"><title>Hybrid route search guided by enzymatic and synthetic models</title><p id="Par16">To explore the retrosynthetic search space efficiently, we expanded on the tree search algorithm implemented in ASKCOS<sup><xref rid="CR24" ref-type="bibr">24</xref></sup>. Whereas the original algorithm uses one template prioritizer model to score the potential retrosynthetic moves at each step and guide the search, the algorithm developed in this work can use arbitrarily many models to guide a search. A search tree rooted at the input target molecule is constructed by the iterative selection, update, and expansion (Fig. <xref rid="Fig1" ref-type="fig">1</xref>b) (Methods). At the expansion step, our algorithm scores the enzymatic templates and synthetic chemistry templates using their corresponding template prioritizer models. The scores are normalized with a softmax function such that the sum of all the scores for a given template set is 1. The distinct template sets are then combined and sorted based on these scores. This means that at each step of retrosynthetic expansion, moves from either template set can be selected.</p><p id="Par17">The multi-model search algorithm directly compares the scores from the template prioritizer models to decide whether to explore a synthetic or enzymatic step. To identify hybrid pathways, this strategy relies on the scores (probabilities) from the two models being scaled appropriately. We show that this is the case for our models by comparing the scores of models’ top-1 recommendations for two external test sets: 48,869 small organic molecules from the MOSES dataset<sup><xref rid="CR54" ref-type="bibr">54</xref></sup> and 45,035 molecules annotated as biogenic (natural products) from the ZINC15 catalog<sup><xref rid="CR55" ref-type="bibr">55</xref></sup> that were not seen by either model during training. An enzymatic reaction template was in the top 3 suggestions for 88% of the small organic molecules and 96% of the natural product molecules. Conversely, a synthetic reaction template was in the top three suggestions for 99% of small organic molecules and 95% of natural product molecules (Fig. <xref rid="Fig4" ref-type="fig">4</xref>a, b). Hence, for most of the molecules in these sets, both synthetic and enzymatic steps would be considered in a pathway search that considers at least three possible moves from a molecule, allowing the algorithm to find hybrid plans when appropriate search parameters are selected.</p><fig id="Fig4" position="float"><?disp-level 3?><label>Fig. 4</label><caption><title>Distribution of output scores from the synthetic chemistry and enzymatic one-step retrosynthesis models.</title><p>Comparison of predicted output scores (probabilities) from the synthetic chemistry and enzymatic template prioritizers illustrate their balance for different chemical spaces. The likelihood that at least one synthetic chemistry template (black) or enzymatic template (blue) will rank among the top templates when templates from both sets are combined and sorted by their probabilities for <bold>a</bold> small organic molecules from the MOSES dataset and for <bold>b</bold> biogenic molecules from ZINC. Means for the top ten probability scores using the enzymatic (blue square) and organic (black triangle) template prioritizers on <bold>c</bold> the MOSES subset and on <bold>d</bold> the ZINC subset. Distribution of top-1 probability scores using the enzymatic (blue) and organic (gray) template prioritizers when the scores are computed for <bold>e</bold> the MOSES subset and for the <bold>f</bold> ZINC subset.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d32e685" xlink:href="41467_2022_35422_Fig4_HTML.jpg"><?cloudpmc-path blobs/46c9/9750992/5d2f44b779c2/41467_2022_35422_Fig4_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1972?><?original-width 1349?><?scaled-height 985?><?scaled-width 674?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="41467_2022_35422_Fig4_HTML.gif"><?cloudpmc-path blobs/46c9/9750992/f5414dcd699c/41467_2022_35422_Fig4_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par18">An intuitive, slight bias appears when comparing the scores from the two models on the different chemical spaces. For small organic molecules, the top retrosynthetic proposal is more likely to be synthetic than enzymatic (Fig. <xref rid="Fig4" ref-type="fig">4</xref>c, e). Conversely, for the natural products, the top retrosynthetic suggestion is more likely to be enzymatic (Fig. <xref rid="Fig4" ref-type="fig">4</xref>d, f). This trend suggests that the models display a (justified) bias that leads to the prioritization of enzymatic chemistries for molecules that are more structurally similar to natural products.</p><p id="Par19">We demonstrate how the interaction of the two models recovers the hybrid synthesis of fluoropyridinyl tryoptoline as synthesized by ref. <xref rid="CR56" ref-type="bibr">56</xref> The retrosynthetic paths that reached viable starting materials are shown in Fig. <xref rid="Fig5" ref-type="fig">5</xref>, where the allowed starting materials were a set of buyable compounds from eMolecules and Sigma-Aldrich. The experimental pathway was recovered and is highlighted in the figure. None of the enzymatic reactions shown are present in the model’s training data, meaning that the model was able to generalize to unseen products and intermediates. The template for the enzymatic aryl bromination links back to 8 reactions from the reaction database belonging to EC classes 1.14.19.55, 1.14.19.58, and 1.97.1. The most similar reaction in the database is the bromination of tryptophan. The Brenda<sup><xref rid="CR49" ref-type="bibr">49</xref></sup> entry for this reaction points to the Uniprot entry for PyrH, a flavin-dependent tryptophan halogenase that shares a mechanism, reaction, and 38% sequence identity with the flavin-dependent tryptophan halogenase that ref. <xref rid="CR56" ref-type="bibr">56</xref> used, RebH (Supplementary Fig. <xref rid="MOESM1" ref-type="supplementary-material">2)</xref>. This example offers a retrospective look at how our hybrid retrosynthesis algorithm could have been used as a successful starting point for route development.</p><fig id="Fig5" position="float"><?disp-level 3?><label>Fig. 5</label><caption><title>Example hybrid pathway search results for fluoropyridinyl tryptoline.</title><p>Pathways displayed from a pathway search using the enzymatic and synthetic template prioritizers (see Methods and <xref rid="MOESM1" ref-type="supplementary-material">SI</xref> for more details). Some reactions are removed from the graph for visual clarity. The recovered experimentally validated pathway from ref. <xref rid="CR56" ref-type="bibr">56</xref> is highlighted with red dashes.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d32e729" xlink:href="41467_2022_35422_Fig5_HTML.jpg"><?cloudpmc-path blobs/46c9/9750992/47bc205cf63d/41467_2022_35422_Fig5_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1085?><?original-width 1757?><?scaled-height 434?><?scaled-width 702?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="41467_2022_35422_Fig5_HTML.gif"><?cloudpmc-path blobs/46c9/9750992/519a2dcc0625/41467_2022_35422_Fig5_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par20">To better understand how the search space of our multi-model algorithm compared with that of the single-model algorithm, we generated a dataset of predicted synthetic routes to 1000 molecules chosen at random from the “boutique” subset of ZINC15 using a search guided by just the synthetic, just the enzymatic, or both template prioritizer models. We compared the number of molecules for which each search was able to identify synthesis pathways and the number of steps in the shortest synthesis pathway found to each target molecule (Fig. <xref rid="Fig6" ref-type="fig">6</xref>).</p><fig id="Fig6" position="float"><?disp-level 3?><label>Fig. 6</label><caption><title>Comparison of routes found with hybrid and single-model searches.</title><p><bold>a</bold> Number of molecules for which synthesis routes were found out of 1000 molecules randomly sampled from the ZINC “boutique” subset. Each pathway search was performed with the same parameters using only the synthetic template prioritizer, only the enzymatic template prioritizer, or both template prioritizers at the same time (hybrid). Comparison of the number of steps in the shortest pathway found with the hybrid pathway planner (containing ≥1 enzymatic step) compared to the shortest pathway found with the organic pathway planner for molecules for which pathways were found by both (431 total) from the ZINC15 boutique sample when <bold>b</bold> each step in a sequence of enzymatic steps is counted and <bold>c</bold> consecutive enzymatic steps (cascades) are counted as a single synthetic step.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d32e753" xlink:href="41467_2022_35422_Fig6_HTML.jpg"><?cloudpmc-path blobs/46c9/9750992/c0f40854d83f/41467_2022_35422_Fig6_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 772?><?original-width 2001?><?scaled-height 257?><?scaled-width 667?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="41467_2022_35422_Fig6_HTML.gif"><?cloudpmc-path blobs/46c9/9750992/ba0fa4000869/41467_2022_35422_Fig6_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par21">Out of the 1000 molecules from ZINC, given the same search parameters, including wall time (Supplementary Table <xref rid="MOESM1" ref-type="supplementary-material">2)</xref>, the enzymatic synthesis planner found pathways to 317, the synthetic planner found pathways to 535, and the hybrid planner found pathways to 493 (Fig. <xref rid="Fig6" ref-type="fig">6</xref>a). The enzymatic path planner finding paths to the smallest number of molecules is expected given that the set of enzymatic templates represents the most limited transformation space (7984 reaction templates compared to 163,723 synthetic organic reaction templates). That the hybrid planner finds routes to fewer compounds than the synthetic planner is also not wholly surprising; the hybrid path planner, in principle, can find a superset of paths found by the other two planners, but this will not be the case when performing a time-limited search. Including a larger number of templates in the search increases the number of possible moves considered in the search and can lead the search algorithm in unproductive directions. Nevertheless, when comparing the targets reached with just the synthetic search and the hybrid search, the hybrid search found routes to 56 molecules, for which none were found with the synthetic search. Of these, all of the routes found to 11 molecules would not be possible without the reaction templates in BKMS (Supplementary Fig. <xref rid="MOESM1" ref-type="supplementary-material">3</xref>), and six molecules required transformations that were present in the constrained set of 824 uniquely enzymatic reactions. Routes to 15 molecules were found for which no routes were found with either individual strategy (Supplementary Fig. <xref rid="MOESM1" ref-type="supplementary-material">4)</xref>. This demonstrates that the hybrid search navigates a different retrosynthetic search space than the individual model searches in a manner that is beneficial for certain molecular targets.</p><p id="Par22">Another metric by which to compare the outputs from the different search strategies is by the total number of reactions in the synthesis plans. Fewer reactions correlate with fewer reagents and fewer purification steps, which in turn correlates with cheaper, more efficient syntheses<sup><xref rid="CR57" ref-type="bibr">57</xref></sup>. The length of a pathway is an incomplete measure of quality. However, the only difference between the purely synthetic and hybrid search is the inclusion of enzymatic steps, so path lengths are particularly informative when comparing pathways in this constrained context. We compare the number of reactions in the shortest pathways found with both the hybrid and the synthetic strategy (431 molecules total) (Fig. <xref rid="Fig6" ref-type="fig">6</xref>b); routes from the hybrid planner were required to have at least one enzymatic reaction. This analysis shows that the hybrid synthesis planner returned an improved shortest path for 73 molecules (17%) and the shortest path of equal length for 162 molecules (38%) out of the 431 boutique molecules. Because successive enzymatic reactions can often proceed in the same solvent without isolation of reaction intermediates, we perform a second analysis comparing step count when counting linear sequences of enzymatic steps as a single step (Fig. <xref rid="Fig6" ref-type="fig">6</xref>c). Recent experimental examples of this strategy include the ex vivo syntheses of the antiviral islatravir<sup><xref rid="CR11" ref-type="bibr">11</xref></sup> and monoterpene commodity chemicals<sup><xref rid="CR58" ref-type="bibr">58</xref></sup>. Considering this perspective, the hybrid synthesis planner found improved shortest paths for 116 (27%) molecules and shortest paths of equal length to 150 (35%) out of the 431 molecules for which both synthetic and hybrid pathways were found.</p></sec><sec id="Sec6" disp-level="2"><title>Case study: dronabinol</title><p id="Par23">Dronabinol (<bold>(−)−1</bold>) is the generic trade name for (−)−Δ<sup>9</sup> tetrahydrocannabinol (THC), indicated for the treatment of anorexia in patients with AIDS and nausea in cancer patients undergoing chemotherapy. It is the decarboxylated form of tetrahydroxycannabinolic acid (THCA), one of 113 cannabinoids naturally produced by the cannabis plant, and represents the main psychoactive agent derived from cannabis<sup><xref rid="CR59" ref-type="bibr">59</xref></sup>. Synthetic routes to dronabinol have been established, using synthetic chemistry<sup><xref rid="CR60" ref-type="bibr">60</xref>–<xref rid="CR65" ref-type="bibr">65</xref></sup>, enzymatic chemistry<sup><xref rid="CR66" ref-type="bibr">66</xref>,<xref rid="CR67" ref-type="bibr">67</xref></sup>, and a combination of both<sup><xref rid="CR60" ref-type="bibr">60</xref></sup> (summary of synthesis routes in Fig. <xref rid="Fig7" ref-type="fig">7</xref>a–g). Briefly, the synthetic syntheses rely on transition-metal catalysts (Cu, Ir, Cr, Mo, and Ru) to set the stereocenters and to form the cyclohexene ring structure either preceding or following ligation to olivetol (<bold>3</bold>) (or a derivative). Alternatively, the fully enzymatic synthesis of (−)−Δ<sup>9</sup> THCA was carried out in <italic>S. cerevisiae</italic> by ref. <xref rid="CR66" ref-type="bibr">66</xref> and the enzymatic synthesis of cannabigerolic acid (CBGA, (<bold>6</bold>)), the immediate metabolic precursor to THCA, has been performed in vitro by ref. <xref rid="CR67" ref-type="bibr">67</xref>.</p><fig id="Fig7" position="float"><?disp-level 3?><label>Fig. 7</label><caption><title>Previously published and newly proposed syntheses of dronabinol.</title><p><bold>a</bold>–<bold>g</bold> Overview of published enantioselective syntheses of dronabinol. <bold>h</bold> Shortest hybrid synthesis plan returned. Dronabinol does not appear in the training set of the enzymatic retrosynthesis model. THCAS tetrahydrocannabinolic acid synthase.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d32e862" xlink:href="41467_2022_35422_Fig7_HTML.jpg"><?cloudpmc-path blobs/46c9/9750992/229b95715920/41467_2022_35422_Fig7_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1617?><?original-width 1994?><?scaled-height 646?><?scaled-width 797?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="41467_2022_35422_Fig7_HTML.gif"><?cloudpmc-path blobs/46c9/9750992/20ca0035118d/41467_2022_35422_Fig7_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par24">We performed an automated multi-step retrosynthetic analysis of dronabinol using an enzymatic, synthetic, or hybrid search. Pathways connecting dronabinol to buyable building blocks were found with the purely enzymatic and hybrid search but not with the purely synthetic search when limiting the search depth to 6. This is consistent with the fact that the published synthetic organic synthesis routes to dronabinol require more than six steps to reach buyable building blocks in our database.</p><p id="Par25">The shortest and qualitatively most promising proposed synthesis route was identified with the hybrid search (Fig. <xref rid="Fig7" ref-type="fig">7</xref>h). The pathway first constructs geraniol (<bold>10</bold>) starting from cheap starting materials <bold>7</bold> and <bold>8</bold> with a Horner–Wadsworth–Emmons reaction followed by a reduction to form <bold>10</bold>. Next, olivetol (<bold>4</bold>) is alkylated with <bold>10</bold> to form cannabigerol <bold>11</bold>. This condensation can be achieved with alumina<sup><xref rid="CR68" ref-type="bibr">68</xref></sup>. The final step is a direct enzymatic ring closure to complete the synthesis of <bold>(−)−1</bold>. There is a single precedent example for this template; following the recommendation back to literature precedent, the recommended enzyme is THCA synthase (EC number 1.21.3.7).</p><p id="Par26">This route is promising as it suggests synthetic chemical transformations to efficiently construct <bold>11</bold>, which would require many enzymes to access from suitable building blocks, and suggests an enzymatic step to both close the ring and set the stereochemistry. All of the routes that were found in our retrosynthetic search relied on the reaction template extracted from the THCA synthase reaction to form the cyclohexene ring. This reaction is uniquely enzymatic and offers a “shortcut” compared to the synthetic chemistry approach with no need for transition-metal catalysts.</p><p id="Par27">Although the THCA synthase reaction with <bold>11</bold> as the substrate is not in the training set, the model believes it may be possible to achieve this transformation enzymatically and the generalized reaction rule extracted from the THCAS reaction fits this product. It has been previously reported that the wild-type THCA synthase does not show activity on <bold>11</bold> and requires the carboxylated analog <bold>6</bold> under the same reaction conditions as the native reaction<sup><xref rid="CR69" ref-type="bibr">69</xref></sup>. However, it remains plausible that a variant of the enzyme could catalyze the desired reaction. Given that the carboxylic acid of <bold>6</bold> is not believed to be directly involved in the catalytic mechanism and that the THCAS residue suspected to interact most strongly with the carboxylic acid is not strictly necessary for enzyme activity<sup><xref rid="CR70" ref-type="bibr">70</xref></sup>, it remains possible that modifying the enzyme’s binding pocket, for example, could confer activity on the decarboxylated substrate <bold>11</bold>. The suggested pathway motivates an enzyme engineering effort to identify a novel variant of THCA synthase capable of catalyzing the proposed reaction on <bold>11</bold> to efficiently access <bold>(−)−1</bold>.</p></sec><sec id="Sec7" disp-level="2"><title>Case study: arformoterol</title><p id="Par28">Arformoterol (R,R-formoterol) is an enantiopure long-acting <italic>β</italic><sub>2</sub> adrenoreceptor agonist prescribed as a bronchodilator for patients with chronic obstructive pulmonary disease (COPD). Few routes have been reported to achieve the diastereomerically pure (R,R) form of formoterol. The three most recent reports all follow a similar, convergent logic (Fig. <xref rid="Fig8" ref-type="fig">8</xref>a–c)<sup><xref rid="CR71" ref-type="bibr">71</xref>–<xref rid="CR73" ref-type="bibr">73</xref></sup>. In one branch of the synthesis, an enantiomerically pure epoxide (<bold>(R)-14a-b</bold>) is synthesized from acetophenone (<bold>12a-b</bold>). In the other branch, an enantiomerically pure <italic>α</italic>-methylphenethylamine (<bold>(R)-15a-c</bold>)is synthesized from 4-methoxyphenylacetone (<bold>13</bold>). The epoxide and amine are reacted in a ring-opening reaction to form a diastereomerically pure amino alcohol which is then subjected either to deprotection reactions (refs. <xref rid="CR71" ref-type="bibr">71</xref>, <xref rid="CR73" ref-type="bibr">73</xref>) or additional functional group interconversions (ref. <xref rid="CR72" ref-type="bibr">72</xref>) to yield <bold>(R,R)-2</bold>.</p><fig id="Fig8" position="float"><?disp-level 3?><label>Fig. 8</label><caption><title>Previously published and newly proposed syntheses of arformoterol.</title><p><bold>a</bold>–<bold>c</bold> Overview of published syntheses of arformoterol. <bold>d</bold> Proposed hybrid synthesis plan. Enzyme names were assigned to steps based on the most structurally similar precedent reaction for reference. Arformoterol does not appear in the training set of the model. AMO alkene monooxygenase, F5H ferulate 5-hydroxylase, LAM lysine 5,6-aminomutase, CYP2D6 Cytochrome P450 2D6, COMT catechol <italic>O</italic>-methyltransferase.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" id="d32e1006" xlink:href="41467_2022_35422_Fig8_HTML.jpg"><?cloudpmc-path blobs/46c9/9750992/259ec85c4d41/41467_2022_35422_Fig8_HTML.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1197?><?original-width 1751?><?scaled-height 479?><?scaled-width 700?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="41467_2022_35422_Fig8_HTML.gif"><?cloudpmc-path blobs/46c9/9750992/56c9f95d398b/41467_2022_35422_Fig8_HTML.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p id="Par29">As with the previous case study, the automated multi-step retrosynthetic analysis of arformoterol was run using the three different search strategies. Pathways were found by the synthetic and hybrid searches, but not the enzymatic search. A manual inspection of the 18 routes returned from the hybrid search showed that they followed a similar logic to the published routes: syntheses of a chiral amine and a chiral epoxide followed by epoxide ring-opening (or a similar nucleophilic substitution), sometimes followed by deprotection and tailoring until the formamide, hydroxyl, and methoxy groups were installed in the proper positions (Supplementary Fig. <xref rid="MOESM1" ref-type="supplementary-material">6)</xref>. Figure <xref rid="Fig8" ref-type="fig">8</xref>d shows one of the predicted synthesis routes that proposes an enzymatic cascade to form the chiral amine and an enzymatic epoxidation step in the formation of the chiral epoxide. Of the proposed routes, this was the most novel compared to the published routes. The introduction of enzymatic reactions unlocked building blocks that were neither found with the synthetic search nor previously reported. The suggestion of a biocatalytic cascade from <bold>18</bold> to <bold>(R)-15b</bold> is promising as it may allow the formation of the chiral amine and the installation of the methoxy group on the aryl ring of <bold>(R)-15b</bold> to take place in one pot with no purification of intermediates.</p><p id="Par30">In contrast with previously reported methods which form the chiral amine from a ketone precursor, the route in Fig. <xref rid="Fig8" ref-type="fig">8</xref> suggests a reaction rule extracted from the reactions of lysine 5,6-aminomutase (LAM, EC number 5.4.3.3) and ornithine 4,5-aminomutase (EC number 5.4.3.5) to form a chiral amine from phenylpropylamine <bold>18</bold>. To our knowledge, no synthetic chemical reaction has been reported to achieve the same selective cross-migration of an amine group and a hydrogen atom; as corroboration, this reaction is included in the set of unique enzymatic transformations we identified. While the usefulness of aminomutases for biotechnology has been explored in the past<sup><xref rid="CR74" ref-type="bibr">74</xref>,<xref rid="CR75" ref-type="bibr">75</xref></sup>, engineering adenosylcobalamin-dependent aminomutases (e.g., LAM) to act on non-natural substrates has not been extensively investigated. This suggestion highlights how the extracted reaction rules can identify motifs on products that may be accessible enzymatically but would likely require modifying a natural enzyme to accommodate the novel substrate. We acknowledge that engineering a variant of LAM for the desired reaction would not be trivial, not only because of the structural difference between phenylpropylamine (<bold>18</bold>) and the native substrate, lysine, but also because the native reaction produces the opposite enantiomer (S) from the desired enantiomer (R). However, the radical mechanism by which the reaction proceeds does not inherently force the production of one enantiomer over the other<sup><xref rid="CR76" ref-type="bibr">76</xref></sup>. Previously, enzyme engineering efforts have succeeded in inverting the natural enantioselectivity of enzymes<sup><xref rid="CR77" ref-type="bibr">77</xref>–<xref rid="CR80" ref-type="bibr">80</xref></sup>. The addition of substrate-promiscuous, enantiocomplementary aminomutases to the biocatalysis toolset would be particularly advantageous given the popularity of chiral amines in small molecule pharmaceuticals<sup><xref rid="CR81" ref-type="bibr">81</xref></sup>.</p><p id="Par31">For the synthesis of the chiral epoxide precursor (<bold>(R)-14c</bold>), the search suggests leveraging an alkene monooxygenase (AMO; E.C. number 1.14.13.69) such as the enzyme from <italic>Rhodococcus</italic> sp. strain AD45, which has been shown to catalyze the enantioselective epoxidation of styrene and p-chlorostyrene to (R)-styrene oxide and (R)-p-chlorostyrene oxide, respectively<sup><xref rid="CR82" ref-type="bibr">82</xref></sup>. In the full sequence from <bold>16</bold> to <bold>(R)-14c</bold>, it may be possible to perform the enzymatic epoxidation and hydroxylation steps as a cascade either before or after the chemical steps. Note that despite being able to easily identify routes with enzyme cascades after they are found, our algorithm does not explicitly promote consecutive enzymatic steps.</p></sec></sec><sec id="Sec8" disp-level="1"><title>Discussion</title><p id="Par32">We have demonstrated a hybrid approach to retrosynthetic planning that generates promising synthesis plans with both enzymatic and synthetic steps to complex molecular targets. Deploying multiple prioritization models within a single retrosynthetic tree search was shown to balance enzymatic and synthetic reaction suggestions to discover pathways with reactions from both sets. The returned hybrid pathways can explore molecular intermediates that would not be accessible with synthetic chemistry or enzymatic chemistry alone.</p><p id="Par33">By comparing the BKMS reaction dataset and the set of reaction templates previously extracted from Reaxys, we showed that the enzymatic reactions include 4196 unique transformations (of which 824 do not eliminate or add heavy atoms) that the synthetic dataset does not. While defining what it means for a transformation to be unique is not straightforward, it seems that enzyme catalysts enable chemical transformations to occur in one step that non-enzymatic reaction conditions cannot. This suggests a role for enzymes in synthetic organic chemistry not just to perform reactions with superior selectivity, but to expand the space of accessible molecules.</p><p id="Par34">The diversity of enzymatic transformations also presents a challenge, as there are few reaction examples for each template; approximately 80% of templates were linked to a single reaction example in our dataset. The extent to which any retrosynthesis recommendation model is able to learn chemical trends for these rare templates is severely limited. Algorithms to generalize overly specific templates have been demonstrated to improve the performance of single- and multi-step template-based synthesis planners<sup><xref rid="CR83" ref-type="bibr">83</xref>,<xref rid="CR84" ref-type="bibr">84</xref></sup>. However, over-generalizing reaction templates may remove necessary chemical context and lead to fewer experimentally implementable suggestions, even if it improves model accuracy metrics. Given the complex interactions that govern enzyme-substrate compatibility, defining enzyme reaction templates at a physically meaningful level of generality may require additional information, such as reaction mechanism or binding pocket structure.</p><p id="Par35">The case studies of dronabinol and arformoterol illustrate how unique enzyme chemistry can unlock routes from novel building blocks or intermediates to compounds of interest. These case studies also illustrate how the template-based retrosynthesis models may suggest enzymatic transformations that would likely require enzyme engineering or screening to implement in the lab—the sort of innovative applications of enzymes to novel substrates that expand the biocatalysis toolbox. Assessing whether an enzyme could plausibly be evolved to perform the desired reaction remains an important challenge for computational modeling and still requires expert knowledge, intuition, and experimentation. Models designed to learn the complex interactions between enzyme sequence and substrate acceptance generalize poorly even in relatively data-rich regimes<sup><xref rid="CR85" ref-type="bibr">85</xref></sup> and simpler molecular similarity-based methods lack the nuance needed to flag projects like the evolution of a transaminase for the synthesis of sitagliptin as worthy pursuits<sup><xref rid="CR43" ref-type="bibr">43</xref></sup>. The nature of the algorithmically-extracted reaction templates used in this work ensures that every recommendation can be traced back to database precedents and the associated literature so users can assess the value of a suggestion based on information which may have escaped cataloging in standardized databases.</p><p id="Par36">The models in this work are data-driven, so the same workflow used to balance synthetic and enzymatic synthesis steps can be applied directly to new datasets. The approach could be used to find routes with other relatively rare sets of reactions, such as “green chemistry” reactions (e.g., photochemical and electrochemical) or routes using data from multiple databases (e.g., both a public and a proprietary set of reactions). There is no theoretical guarantee that the same desirable balance between models would be observed for new models as it was between the synthetic and enzymatic models in this work, but the application of a softmax transform to the scores from each model constrains the range of the model outputs, and the empirical trend of higher model confidence for inputs that are similar to training examples seems likely to persist. Nevertheless, a balancing parameter could be introduced to manually tune the relative scores of the models if there is a known desired outcome. An additional limitation is that combining multiple models in a search expands the search space such that, given the same time limit, the hybrid search may not find pathways that a search guided by a single model would find. This further motivates exploring new techniques, such as reinforcement learning to balance model suggestions.</p><p id="Par37">We believe that hybrid CASP approaches such as ours will accelerate the identification and development of new efficient synthesis routes. Enzymes can catalyze certain transformations that are not otherwise possible and increase the selectivity and efficiency of others, while synthetic chemistry offers a broader, complementary toolkit. Identifying opportunities to apply enzyme chemistry to access novel targets with our algorithm can work synergistically with experimental efforts in high throughput screening and the evolution of enzymes for the discovery of new biocatalysts.</p></sec><sec id="Sec9" disp-level="1"><title>Methods</title><sec id="Sec10" disp-level="2"><title>Processing the BKMS database</title><p id="Par38">The BKMS database<sup><xref rid="CR48" ref-type="bibr">48</xref></sup> is a composite database containing reactions from BRENDA<sup><xref rid="CR49" ref-type="bibr">49</xref></sup>, the Kyoto Encylopedia of Genes and Genomes (KEGG)<sup><xref rid="CR44" ref-type="bibr">44</xref></sup>, Metacyc<sup><xref rid="CR50" ref-type="bibr">50</xref></sup>, and SABIO-RK<sup><xref rid="CR51" ref-type="bibr">51</xref></sup>. The database was retrieved as a flat file with 37,235 enzymatic reactions. Reactions are represented in the form A + B + ... = C + ..., where A, B, and C are unstandardized chemical names.</p><p id="Par39">Cofactor reactant-product pairs were automatically identified by computing the co-occurrence of molecule names as reactants and products (e.g., NAD+ and NADH). Product-reactant pairs that co-occur in at least ten reactions in at least 90% of the reactions they were in were removed from the reactions in which they co-occur (full list in Supplementary Table <xref rid="MOESM1" ref-type="supplementary-material">3)</xref>. This was particularly important to remove chemicals for which SMILES strings could not be found (e.g., “oxidized ferredoxin.”)</p><p id="Par40">A dictionary mapping chemical names to SMILES strings was constructed using the PubChem Identifier Exchange Service<sup><xref rid="CR86" ref-type="bibr">86</xref></sup> and data downloaded from MetaCyc and BRENDA. This dictionary was used to obtain reaction SMILES for all reactions. Reactions were removed if SMILES were not found for all reactants and products.</p><p id="Par41">Reaction SMILES were automatically atom-atom-mapped using the Reaction Decoder Tool<sup><xref rid="CR87" ref-type="bibr">87</xref></sup>. The reverse reaction was added for all reactions that were indicated to be reversible in BKMS.</p><p id="Par42">Further processing was performed to maximize the number of single-product reactions, because multi-product reactions cannot be handled in the tree search. The number of times that each molecule SMILES appeared as a product in a reaction with &gt;1 product was counted. Iteratively, the most frequently co-appearing molecule was removed from all multi-product reactions. During this process, common side-products or leaving groups such as coenzyme A and <sc>d</sc>-glucose were removed from multi-product reactions. Finally, reactant molecules that did not have at least three mapped atoms and 15% of their mapped atoms represented in the products were removed from reactions with &gt;1 reactant. After deduplication, the cleaned reaction database contained 18,719 cleaned, mapped reaction SMILES.</p></sec><sec id="Sec11" disp-level="2"><title>Automatic extraction of reaction rules</title><p id="Par43">Reaction templates were automatically extracted using RDChiral<sup><xref rid="CR53" ref-type="bibr">53</xref></sup>. The template radius was set to 1 and the default special groups were included in the templates. Templates were validated by applying extracted templates back to the product using RDKit<sup><xref rid="CR88" ref-type="bibr">88</xref></sup> and checking that the generated reactants matched the reactants from the database. During validation, stereochemistry was ignored around atoms and bonds that were not matched by the template to avoid SMILES definition mismatches that were not relevant to the reaction. Reactions that led to invalid templates were removed. A total of 7984 valid templates were extracted from 15,309 reactions.</p><p id="Par44">The template label prediction accuracy increases when there are more examples with that label in the training set (Supplementary Fig. <xref rid="MOESM1" ref-type="supplementary-material">7)</xref>. We studied how filtering out reaction templates with few examples affected the breadth of chemistry that the template set could cover. For a range of threshold values, all templates that had fewer examples than the threshold were removed from the data. Then, all templates with more examples than the threshold were applied to the products from all unlabeled reactions using RDKit with and without considering stereochemistry. If any template recovered the reactants, the reaction precedent was reassigned to that template and retained; otherwise, the reaction was removed. The relative number of reactions that were covered at different template popularity thresholds is shown in Figure <xref rid="Fig2" ref-type="fig">2</xref>f. Because of the substantial decrease in reactions covered when any threshold was set, templates were not filtered out based on popularity in any subsequent evaluations.</p></sec><sec id="Sec12" disp-level="2"><title>Training the template prioritizer</title><p id="Par45">The template prioritizer is an MLP model trained to score all the retrosynthetic moves for an input product molecule represented by reaction templates. The input for the template prioritizer is a 2048-bit Morgan fingerprint representation of a product molecule as implemented in RDKit<sup><xref rid="CR88" ref-type="bibr">88</xref></sup> with chiral features. The output is a vector of length 7984, corresponding to the number of templates. A softmax activation is applied to the final layer, such that the score across all templates sums to 1 and can be interpreted as a probability.</p><p id="Par46">The data were initially split 80% in training, 10% in validation, and 10% in test sets using a previously described stratified split<sup><xref rid="CR89" ref-type="bibr">89</xref></sup> to ensure a more even distribution of class labels across the splits. For templates with ten or more examples, the examples were assigned at random to one of the three sets. For templates with fewer than ten examples, at random, one reaction example was assigned to the test and one to the validation set and the rest were kept in the training set. For templates with exactly two examples, one was assigned to the test set and one was assigned to the train set. For templates with a single example, the example was assigned at random to the train, validation, or test sets with a probability proportional to the size of the set.</p><p id="Par47">The hyperparameters tuned were the number of hidden layers (1, 2, or 3), the size of the hidden layers (1024, 2048, or 4096), and the number of highway layers<sup><xref rid="CR90" ref-type="bibr">90</xref></sup> (0, 1, or 3) using a grid search. Models were trained with and without pre-training on template applicability<sup><xref rid="CR91" ref-type="bibr">91</xref></sup> and accuracies were calculated on the validation set. The best performance was achieved with 1 hidden layer of size 4096 and 0 highway layers. Pre-training increased the accuracy of all models tested, so the final network used in the multi-step search was pre-trained.</p><p id="Par48">The single-step retrosynthesis model trained on BKMS data ranked the correct template as the top template for 19% of the product molecules from the test set and ranked the correct template in the top 10 for 53% of the test set molecules. We trained an additional model on all of the available data to maximize the diversity of template labels seen by the model. We set hyperparameters and the number of training epochs based on the tuning performed with the split dataset. This final model, trained on the entire dataset, was only used with external test sets selected independently.</p></sec><sec id="Sec13" disp-level="2"><title>Multi-prioritizer guided tree search</title><p id="Par49">The multi-prioritizer tree search uses an expanded version of the algorithm used in ASKCOS.<sup><xref rid="CR24" ref-type="bibr">24</xref></sup> A search tree is constructed from an input target molecule by iterative rounds of selection, expansion, and update steps. During selection, the leaf node of the highest-scoring pathway is identified by greedily traversing the tree, picking the highest-scoring children nodes until a leaf is reached. The node scores are a function of the neural network model scores MLP(<bold>P</bold>), the visit counts of the product node <italic>N</italic><sub><italic>P</italic></sub> and the child reaction node <italic>N</italic><sub><italic>R</italic></sub>, an exploration factor <italic>c</italic>, and a value estimate for the reaction node <italic>V</italic><sub><italic>R</italic></sub>. <italic>V</italic><sub><italic>R</italic></sub> is assigned 0 if no paths from the node reach buyable starting materials. Otherwise, it is the mean number of building blocks needed for routes that do reach buyable starting materials from the node.</p><disp-formula id="Equ1"><label>1</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M2" display="block"><mml:mspace width="0.25em"/><mml:mstyle><mml:mtext>Score</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>c</mml:mi><mml:msqrt><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msqrt></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle><mml:mtext>MLP</mml:mtext></mml:mstyle><mml:mspace width="0.25em"/><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="bold">P</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>−</mml:mo><mml:msub><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula><p id="Par50">The selected leaf node represents a yet-to-be-applied reaction template. During the update step, node visit counts are incremented along the path from the root node (target molecule) to the selected leaf. At the expansion step, a new precursor is generated by applying the leaf node template to its parent chemical node. With the new precursor as input, all reaction templates from each template set are scored separately by their respective template prioritizer model, then combined and re-sorted. The key difference from the original ASKCOS algorithm is that an unbounded number of template prioritizers can be used instead of exactly one. The number of suggestions grows linearly with the number of models in use. Once the expansion time is elapsed, the tree construction halts. Paths that connect the target molecule to buyable starting materials are retrieved and returned using a depth-first search.</p></sec><sec id="Sec14" disp-level="2"><title>Buyable database</title><p id="Par51">The default buyable compound database from ASKCOS was used for all pathway searches. This database contains 106,750 compounds available for less than $100/g from the vendors eMolecules and Sigma-Aldrich. An estimated price is associated with each compound.</p></sec><sec id="Sec15" disp-level="2"><title>Benchmarking search strategies</title><p id="Par52">One thousand molecules were randomly selected as targets from the set of named, “boutique” compounds from the ZINC15<sup><xref rid="CR55" ref-type="bibr">55</xref></sup> database. Three tree searches were performed for each molecule: one using only the enzymatic retrosynthesis model, one using the synthetic retrosynthesis model, and one using both retrosynthesis models. The parameters for all three searches were otherwise identical. The maximum search depth was ten and the maximum expansion time was 180 s (additional parameters in Supplementary Table <xref rid="MOESM1" ref-type="supplementary-material">2)</xref>. All synthesis plans were performed using Google Cloud Platform VM instances with four cores and 26 GB of memory.</p></sec><sec id="Sec16" disp-level="2"><title>Synthesis planning</title><p id="Par53">The synthesis plans for fluoropyridinyl tryptoline, dronabinol, and arformoterol were automatically generated, using the same parameters as the hybrid search performed in the benchmarking analysis. The only difference was that the search depth was capped at 4, 6, and 7 reactions, respectively for the three targets based on the expected path lengths of previously published routes. Potential enzymes were assigned to reactions from the most similar reaction precedent as defined by the Tanimoto similarity between the RDKit reaction structural fingerprints of the proposed and precedent reactions.</p></sec></sec><sec id="sec17" disp-level="1"><title>Supplementary information</title><sec id="Sec17" disp-level="2">
<supplementary-material id="MOESM1" position="float"><media xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="41467_2022_35422_MOESM1_ESM.pdf" mimetype="application" mime-subtype="pdf"><?cloudpmc-path 46c9/9750992/689845310b63/41467_2022_35422_MOESM1_ESM.pdf?><?cloudpmc-bucket app?><?size 3352902?><caption><p>Supplementary Information</p></caption></media></supplementary-material>
</sec></sec><sec id="ack1" sec-type="ack" disp-level="1"><title>Acknowledgements</title><p>We thank Michael Fortunato for help with setting up the model training environment and Samuel Goldman for preprocessing the BRENDA database. The initial development of this work was supported by the Google Cloud Research Credits program. This work was funded by the Machine Learning for Pharmaceutical Discovery and Synthesis consortium and the Air Force Research Lab award no. FA8650-15-D-5405 through UES, Inc.</p></sec><sec id="notes1" disp-level="1"><title>Author contributions</title><p>I.L., C.A.V., and C.W.C. conceptualized the project. I.L. developed the methods and performed computational experiments and analyses. M.L. supported the development and deployment of the software. I.L., C.A.V., and C.W.C. wrote the manuscript. All authors revised and approved the manuscript.</p></sec><sec id="notes2" disp-level="1"><title>Peer review</title><sec id="FPar1" disp-level="2"><title>Peer review information</title><p id="Par54"><italic>Nature Communications</italic> thanks the anonymous reviewers for their contribution to the peer review of this work.</p></sec></sec><sec id="notes3" disp-level="1"><title>Data availability</title><p>Reaxys reaction data used to build the previously published synthetic organic model is the intellectual property of Elsevier and cannot be shared. The templates extracted from the Reaxys data and all of the BKMS data used to build the enzymatic model and associated templates is available at <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/itai-levin/bkms-data" ext-link-type="uri">https://github.com/itai-levin/bkms-data</ext-link><sup><xref rid="CR92" ref-type="bibr">92</xref></sup>.</p></sec><sec id="notes4" disp-level="1"><title>Code availability</title><p>The source code for deploying ASKCOS is available at <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/itai-levin/chemoenzymatic-askcos" ext-link-type="uri">https://github.com/itai-levin/chemoenzymatic-askcos</ext-link><sup><xref rid="CR93" ref-type="bibr">93</xref></sup>. The code for performing the analyses in this manuscript is available at <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/itai-levin/hybmind" ext-link-type="uri">https://github.com/itai-levin/hybmind</ext-link><sup><xref rid="CR94" ref-type="bibr">94</xref></sup>. This repository contains a detailed README with instructions on how to generate the results presented in the manuscript.</p></sec><sec id="FPar2" disp-level="1"><title>Competing interests</title><p id="Par55">The authors declare no competing interests.</p></sec><sec id="fn-group1" sec-type="fn-group" disp-level="1"><title>Footnotes</title><fn-group><fn id="fn1"><p><bold>Publisher’s note</bold> Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p></fn></fn-group></sec><sec id="sec19" disp-level="1"><title>Supplementary information</title><p>The online version contains supplementary material available at 10.1038/s41467-022-35422-y.</p></sec><sec id="Bib1" sec-type="ref-list" disp-level="1"><title>References</title><sec id="Bib1_sec2" disp-level="2"><ref-list><ref id="CR1"><label>1.</label><mixed-citation><named-content content-type="citation-string">Chakrabarty, S., Romero, E. O., Pyser, J. B., Yazarians, J. A. &amp; Narayan, A. R. H. Chemoenzymatic total synthesis of natural products. <italic>Acc. Chem. Res.</italic><bold>54</bold>, 1374–1384 (2021).</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.accounts.0c00810"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8210581"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33600149"/></mixed-citation></ref><ref id="CR2"><label>2.</label><mixed-citation><named-content content-type="citation-string">Li J, Amatuni A, Renata H. Recent advances in the chemoenzymatic synthesis of bioactive natural products. Curr. Opin. Chem. Biol. 2020;55:111–118. doi: 10.1016/j.cbpa.2020.01.005.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.cbpa.2020.01.005"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7237303"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32086167"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Curr. Opin. Chem. Biol.&amp;title=Recent advances in the chemoenzymatic synthesis of bioactive natural products&amp;author=J Li&amp;author=A Amatuni&amp;author=H Renata&amp;volume=55&amp;publication_year=2020&amp;pages=111-118&amp;pmid=32086167&amp;doi=10.1016/j.cbpa.2020.01.005&amp;"/></mixed-citation></ref><ref id="CR3"><label>3.</label><mixed-citation><named-content content-type="citation-string">Zhang X, et al.  Divergent synthesis of complex diterpenes through a hybrid oxidative approach. Science. 2020;369:799–806. doi: 10.1126/science.abb8271.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1126/science.abb8271"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7569743"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32792393"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Science&amp;title=Divergent synthesis of complex diterpenes through a hybrid oxidative approach&amp;author=X Zhang&amp;volume=369&amp;publication_year=2020&amp;pages=799-806&amp;pmid=32792393&amp;doi=10.1126/science.abb8271&amp;"/></mixed-citation></ref><ref id="CR4"><label>4.</label><mixed-citation><named-content content-type="citation-string">Patel NR, et al.  Synthesis of islatravir enabled by a catalytic, enantioselective alkynylation of a ketone. Org. Lett. 2020;22:4659–4664. doi: 10.1021/acs.orglett.0c01431.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.orglett.0c01431"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32516536"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Org. Lett.&amp;title=Synthesis of islatravir enabled by a catalytic, enantioselective alkynylation of a ketone&amp;author=NR Patel&amp;volume=22&amp;publication_year=2020&amp;pages=4659-4664&amp;pmid=32516536&amp;doi=10.1021/acs.orglett.0c01431&amp;"/></mixed-citation></ref><ref id="CR5"><label>5.</label><mixed-citation><named-content content-type="citation-string">M. Abdelraheem EM, Busch H, Hanefeld U, Tonin F. Biocatalysis explained: from pharmaceutical to bulk chemical production. React. Chem. Eng. 2019;4:1878–1894. doi: 10.1039/C9RE00301K.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1039/C9RE00301K"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=React. Chem. Eng.&amp;title=Biocatalysis explained: from pharmaceutical to bulk chemical production&amp;author=EM M. Abdelraheem&amp;author=H Busch&amp;author=U Hanefeld&amp;author=F Tonin&amp;volume=4&amp;publication_year=2019&amp;pages=1878-1894&amp;doi=10.1039/C9RE00301K&amp;"/></mixed-citation></ref><ref id="CR6"><label>6.</label><mixed-citation><named-content content-type="citation-string">Wu, S., Snajdrova, R., Moore, J. C., Baldenius, K. &amp; Bornscheuer, U. Biocatalysis: enzymatic synthesis for industrial applications. <italic>Angew. Chem. Int. Ed.</italic><bold>60</bold>, 88–119 (2020).</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1002/anie.202006648"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7818486"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32558088"/></mixed-citation></ref><ref id="CR7"><label>7.</label><mixed-citation><named-content content-type="citation-string">Sheldon RA, Brady D, Bode ML. The Hitchhiker’s guide to biocatalysis: recent advances in the use of enzymes in organic synthesis. Chem. Sci. 2020;11:2587–2605. doi: 10.1039/C9SC05746C.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1039/C9SC05746C"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7069372"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32206264"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Chem. Sci.&amp;title=The Hitchhiker’s guide to biocatalysis: recent advances in the use of enzymes in organic synthesis&amp;author=RA Sheldon&amp;author=D Brady&amp;author=ML Bode&amp;volume=11&amp;publication_year=2020&amp;pages=2587-2605&amp;pmid=32206264&amp;doi=10.1039/C9SC05746C&amp;"/></mixed-citation></ref><ref id="CR8"><label>8.</label><mixed-citation><named-content content-type="citation-string">Fryszkowska A, Devine PN. Biocatalysis in drug discovery and development. Curr. Opin. Chem. Biol. 2020;55:151–160. doi: 10.1016/j.cbpa.2020.01.012.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.cbpa.2020.01.012"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32169795"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Curr. Opin. Chem. Biol.&amp;title=Biocatalysis in drug discovery and development&amp;author=A Fryszkowska&amp;author=PN Devine&amp;volume=55&amp;publication_year=2020&amp;pages=151-160&amp;pmid=32169795&amp;doi=10.1016/j.cbpa.2020.01.012&amp;"/></mixed-citation></ref><ref id="CR9"><label>9.</label><mixed-citation><named-content content-type="citation-string">Savile CK, et al.  Biocatalytic asymmetric synthesis of chiral amines from ketones applied to sitagliptin manufacture. Science. 2010;329:305–309. doi: 10.1126/science.1188934.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1126/science.1188934"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="20558668"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Science&amp;title=Biocatalytic asymmetric synthesis of chiral amines from ketones applied to sitagliptin manufacture&amp;author=CK Savile&amp;volume=329&amp;publication_year=2010&amp;pages=305-309&amp;pmid=20558668&amp;doi=10.1126/science.1188934&amp;"/></mixed-citation></ref><ref id="CR10"><label>10.</label><mixed-citation><named-content content-type="citation-string">Nawrat CC, et al.  Nine-step stereoselective synthesis of islatravir from deoxyribose. Org. Lett. 2020;22:2167–2172. doi: 10.1021/acs.orglett.0c00239.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.orglett.0c00239"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32108487"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Org. Lett.&amp;title=Nine-step stereoselective synthesis of islatravir from deoxyribose&amp;author=CC Nawrat&amp;volume=22&amp;publication_year=2020&amp;pages=2167-2172&amp;pmid=32108487&amp;doi=10.1021/acs.orglett.0c00239&amp;"/></mixed-citation></ref><ref id="CR11"><label>11.</label><mixed-citation><named-content content-type="citation-string">Huffman MA, et al.  Design of an in vitro biocatalytic cascade for the manufacture of islatravir. Science. 2019;366:1255–1259. doi: 10.1126/science.aay8484.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1126/science.aay8484"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="31806816"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Science&amp;title=Design of an in vitro biocatalytic cascade for the manufacture of islatravir&amp;author=MA Huffman&amp;volume=366&amp;publication_year=2019&amp;pages=1255-1259&amp;pmid=31806816&amp;doi=10.1126/science.aay8484&amp;"/></mixed-citation></ref><ref id="CR12"><label>12.</label><mixed-citation><named-content content-type="citation-string">Cai T, et al.  Cell-free chemoenzymatic starch synthesis from carbon dioxide. Science. 2021;373:1523–1527. doi: 10.1126/science.abh4049.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1126/science.abh4049"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="34554807"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Science&amp;title=Cell-free chemoenzymatic starch synthesis from carbon dioxide&amp;author=T Cai&amp;volume=373&amp;publication_year=2021&amp;pages=1523-1527&amp;pmid=34554807&amp;doi=10.1126/science.abh4049&amp;"/></mixed-citation></ref><ref id="CR13"><label>13.</label><mixed-citation><named-content content-type="citation-string">Truppo MD. Biocatalysis in the pharmaceutical industry: the need for speed. ACS Med. Chem. Lett. 2017;8:476–480. doi: 10.1021/acsmedchemlett.7b00114.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acsmedchemlett.7b00114"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC5430392"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="28523096"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=ACS Med. Chem. Lett.&amp;title=Biocatalysis in the pharmaceutical industry: the need for speed&amp;author=MD Truppo&amp;volume=8&amp;publication_year=2017&amp;pages=476-480&amp;pmid=28523096&amp;doi=10.1021/acsmedchemlett.7b00114&amp;"/></mixed-citation></ref><ref id="CR14"><label>14.</label><mixed-citation><named-content content-type="citation-string">Struble TJ, et al.  Current and future roles of artificial intelligence in medicinal chemistry synthesis. J. Med. Chem. 2020;63:8667–8682. doi: 10.1021/acs.jmedchem.9b02120.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.jmedchem.9b02120"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7457232"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32243158"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Med. Chem.&amp;title=Current and future roles of artificial intelligence in medicinal chemistry synthesis&amp;author=TJ Struble&amp;volume=63&amp;publication_year=2020&amp;pages=8667-8682&amp;pmid=32243158&amp;doi=10.1021/acs.jmedchem.9b02120&amp;"/></mixed-citation></ref><ref id="CR15"><label>15.</label><mixed-citation><named-content content-type="citation-string">Baum, Z. J. et al. Artificial intelligence in chemistry: current trends and future directions. <italic>J, Chem. Inf. Model.</italic><bold>61</bold>, 3197–3212 (2021).</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.jcim.1c00619"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="34264069"/></mixed-citation></ref><ref id="CR16"><label>16.</label><mixed-citation><named-content content-type="citation-string">Hadadi N, Hatzimanikatis V. Design of computational retrobiosynthesis tools for the design of de novo synthetic pathways. Curr. Opin. Chem. Biol. 2015;28:99–104. doi: 10.1016/j.cbpa.2015.06.025.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.cbpa.2015.06.025"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="26177079"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Curr. Opin. Chem. Biol.&amp;title=Design of computational retrobiosynthesis tools for the design of de novo synthetic pathways&amp;author=N Hadadi&amp;author=V Hatzimanikatis&amp;volume=28&amp;publication_year=2015&amp;pages=99-104&amp;pmid=26177079&amp;doi=10.1016/j.cbpa.2015.06.025&amp;"/></mixed-citation></ref><ref id="CR17"><label>17.</label><mixed-citation><named-content content-type="citation-string">Lin G-M, Warden-Rothman R, Voigt CA. Retrosynthetic design of metabolic pathways to chemicals not found in nature. Curr. Opin. Syst. Biol. 2019;14:82–107. doi: 10.1016/j.coisb.2019.04.004.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.coisb.2019.04.004"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Curr. Opin. Syst. Biol.&amp;title=Retrosynthetic design of metabolic pathways to chemicals not found in nature&amp;author=G-M Lin&amp;author=R Warden-Rothman&amp;author=CA Voigt&amp;volume=14&amp;publication_year=2019&amp;pages=82-107&amp;doi=10.1016/j.coisb.2019.04.004&amp;"/></mixed-citation></ref><ref id="CR18"><label>18.</label><mixed-citation><named-content content-type="citation-string">Cook A, et al.  Computer-aided synthesis design: 40 years on. WIREs Comput. Mol. Sci. 2012;2:79–107. doi: 10.1002/wcms.61.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1002/wcms.61"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=WIREs Comput. Mol. Sci.&amp;title=Computer-aided synthesis design: 40 years on&amp;author=A Cook&amp;volume=2&amp;publication_year=2012&amp;pages=79-107&amp;doi=10.1002/wcms.61&amp;"/></mixed-citation></ref><ref id="CR19"><label>19.</label><mixed-citation><named-content content-type="citation-string">Ravitz O. Data-driven computer aided synthesis design. Drug Discov. Today. Technol. 2013;10:e443–e449. doi: 10.1016/j.ddtec.2013.01.005.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.ddtec.2013.01.005"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="24050141"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Drug Discov. Today. Technol.&amp;title=Data-driven computer aided synthesis design&amp;author=O Ravitz&amp;volume=10&amp;publication_year=2013&amp;pages=e443-e449&amp;pmid=24050141&amp;doi=10.1016/j.ddtec.2013.01.005&amp;"/></mixed-citation></ref><ref id="CR20"><label>20.</label><mixed-citation><named-content content-type="citation-string">Johansson S, et al.  AI-assisted synthesis prediction. Drug Discov. Today. Technol. 2019;32-33:65–72. doi: 10.1016/j.ddtec.2020.06.002.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.ddtec.2020.06.002"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33386096"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Drug Discov. Today. Technol.&amp;title=AI-assisted synthesis prediction&amp;author=S Johansson&amp;volume=32-33&amp;publication_year=2019&amp;pages=65-72&amp;pmid=33386096&amp;doi=10.1016/j.ddtec.2020.06.002&amp;"/></mixed-citation></ref><ref id="CR21"><label>21.</label><mixed-citation><named-content content-type="citation-string">Szymkuć S, et al.  Computer-assisted synthetic planning: the end of the beginning. Angew. Chem. Int. Ed. 2016;55:5904–5937. doi: 10.1002/anie.201506101.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1002/anie.201506101"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="27062365"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Angew. Chem. Int. Ed.&amp;title=Computer-assisted synthetic planning: the end of the beginning&amp;author=S Szymkuć&amp;volume=55&amp;publication_year=2016&amp;pages=5904-5937&amp;pmid=27062365&amp;doi=10.1002/anie.201506101&amp;"/></mixed-citation></ref><ref id="CR22"><label>22.</label><mixed-citation><named-content content-type="citation-string">Coley CW, Green WH, Jensen KF. Machine learning in computer-aided synthesis planning. Acc. Chem. Res. 2018;51:1281–1289. doi: 10.1021/acs.accounts.8b00087.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.accounts.8b00087"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="29715002"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Acc. Chem. Res.&amp;title=Machine learning in computer-aided synthesis planning&amp;author=CW Coley&amp;author=WH Green&amp;author=KF Jensen&amp;volume=51&amp;publication_year=2018&amp;pages=1281-1289&amp;pmid=29715002&amp;doi=10.1021/acs.accounts.8b00087&amp;"/></mixed-citation></ref><ref id="CR23"><label>23.</label><mixed-citation><named-content content-type="citation-string">Shen Y, et al.  Automation and computer-assisted planning for chemical synthesis. Nat. Rev. Methods Prim. 2021;1:23. doi: 10.1038/s43586-021-00022-5.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s43586-021-00022-5"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nat. Rev. Methods Prim.&amp;title=Automation and computer-assisted planning for chemical synthesis&amp;author=Y Shen&amp;volume=1&amp;publication_year=2021&amp;pages=23&amp;doi=10.1038/s43586-021-00022-5&amp;"/></mixed-citation></ref><ref id="CR24"><label>24.</label><mixed-citation><named-content content-type="citation-string">Coley CW, et al.  A robotic platform for flow synthesis of organic compounds informed by AI planning. Science. 2019;365:eaax1566. doi: 10.1126/science.aax1566.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1126/science.aax1566"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="31395756"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Science&amp;title=A robotic platform for flow synthesis of organic compounds informed by AI planning&amp;author=CW Coley&amp;volume=365&amp;publication_year=2019&amp;pages=eaax1566&amp;pmid=31395756&amp;doi=10.1126/science.aax1566&amp;"/></mixed-citation></ref><ref id="CR25"><label>25.</label><mixed-citation><named-content content-type="citation-string">Delépine B, Duigou T, Carbonell P, Faulon J-L. RetroPath2.0: a retrosynthesis workflow for metabolic engineers. Metab. Eng. 2018;45:158–170. doi: 10.1016/j.ymben.2017.12.002.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.ymben.2017.12.002"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="29233745"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Metab. Eng.&amp;title=RetroPath2.0: a retrosynthesis workflow for metabolic engineers&amp;author=B Delépine&amp;author=T Duigou&amp;author=P Carbonell&amp;author=J-L Faulon&amp;volume=45&amp;publication_year=2018&amp;pages=158-170&amp;pmid=29233745&amp;doi=10.1016/j.ymben.2017.12.002&amp;"/></mixed-citation></ref><ref id="CR26"><label>26.</label><mixed-citation><named-content content-type="citation-string">Koch M, Duigou T, Faulon J-L. Reinforcement learning for bioretrosynthesis. ACS Synth. Biol. 2020;9:157–168. doi: 10.1021/acssynbio.9b00447.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acssynbio.9b00447"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="31841626"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=ACS Synth. Biol.&amp;title=Reinforcement learning for bioretrosynthesis&amp;author=M Koch&amp;author=T Duigou&amp;author=J-L Faulon&amp;volume=9&amp;publication_year=2020&amp;pages=157-168&amp;pmid=31841626&amp;doi=10.1021/acssynbio.9b00447&amp;"/></mixed-citation></ref><ref id="CR27"><label>27.</label><mixed-citation><named-content content-type="citation-string">Duigou T, du Lac M, Carbonell P, Faulon J-L. RetroRules: a database of reaction rules for engineering biology. Nucleic Acids Res. 2019;47:D1229–D1235. doi: 10.1093/nar/gky940.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/nar/gky940"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC6323975"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="30321422"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nucleic Acids Res.&amp;title=RetroRules: a database of reaction rules for engineering biology&amp;author=T Duigou&amp;author=M du Lac&amp;author=P Carbonell&amp;author=J-L Faulon&amp;volume=47&amp;publication_year=2019&amp;pages=D1229-D1235&amp;pmid=30321422&amp;doi=10.1093/nar/gky940&amp;"/></mixed-citation></ref><ref id="CR28"><label>28.</label><mixed-citation><named-content content-type="citation-string">Finnigan W, Hepworth LJ, Flitsch SL, Turner NJ. RetroBioCat as a computer-aided synthesis planning tool for biocatalytic reactions and cascades. Nat. Catal. 2021;4:98–104. doi: 10.1038/s41929-020-00556-z.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41929-020-00556-z"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7116764"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33604511"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nat. Catal.&amp;title=RetroBioCat as a computer-aided synthesis planning tool for biocatalytic reactions and cascades&amp;author=W Finnigan&amp;author=LJ Hepworth&amp;author=SL Flitsch&amp;author=NJ Turner&amp;volume=4&amp;publication_year=2021&amp;pages=98-104&amp;pmid=33604511&amp;doi=10.1038/s41929-020-00556-z&amp;"/></mixed-citation></ref><ref id="CR29"><label>29.</label><mixed-citation><named-content content-type="citation-string">Liu B, et al.  Retrosynthetic reaction prediction using neural sequence-to-sequence models. ACS Cent. Sci. 2017;3:1103–1113. doi: 10.1021/acscentsci.7b00303.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acscentsci.7b00303"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC5658761"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="29104927"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=ACS Cent. Sci.&amp;title=Retrosynthetic reaction prediction using neural sequence-to-sequence models&amp;author=B Liu&amp;volume=3&amp;publication_year=2017&amp;pages=1103-1113&amp;pmid=29104927&amp;doi=10.1021/acscentsci.7b00303&amp;"/></mixed-citation></ref><ref id="CR30"><label>30.</label><mixed-citation><named-content content-type="citation-string">Zheng S, Rao J, Zhang Z, Xu J, Yang Y. Predicting retrosynthetic reactions using self-corrected transformer neural networks. J. Chem. Inf. Model. 2020;60:47–55. doi: 10.1021/acs.jcim.9b00949.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.jcim.9b00949"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="31825611"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Chem. Inf. Model.&amp;title=Predicting retrosynthetic reactions using self-corrected transformer neural networks&amp;author=S Zheng&amp;author=J Rao&amp;author=Z Zhang&amp;author=J Xu&amp;author=Y Yang&amp;volume=60&amp;publication_year=2020&amp;pages=47-55&amp;pmid=31825611&amp;doi=10.1021/acs.jcim.9b00949&amp;"/></mixed-citation></ref><ref id="CR31"><label>31.</label><mixed-citation><named-content content-type="citation-string">Probst, D. et al. Biocatalysed synthesis planning using data-driven learning. <italic>Nat. Commun.</italic><bold>13</bold>, 964 (2022)</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41467-022-28536-w"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8857209"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="35181654"/></mixed-citation></ref><ref id="CR32"><label>.32.</label><mixed-citation><named-content content-type="citation-string">Schwaller P, et al.  Predicting retrosynthetic pathways using transformer-based models and a hyper-graph exploration strategy. Chem. Sci. 2020;11:3316–3325. doi: 10.1039/C9SC05704H.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1039/C9SC05704H"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8152799"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="34122839"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Chem. Sci.&amp;title=Predicting retrosynthetic pathways using transformer-based models and a hyper-graph exploration strategy&amp;author=P Schwaller&amp;volume=11&amp;publication_year=2020&amp;pages=3316-3325&amp;pmid=34122839&amp;doi=10.1039/C9SC05704H&amp;"/></mixed-citation></ref><ref id="CR33"><label>33.</label><mixed-citation><named-content content-type="citation-string">Zheng, S. et al. Deep learning driven biosynthetic pathways navigation for natural products with BioNavi-NP. <italic>Nat Commun.</italic><bold>13</bold>, 3342 (2022).</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41467-022-30970-9"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC9187661"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="35688826"/></mixed-citation></ref><ref id="CR34"><label>34.</label><mixed-citation><named-content content-type="citation-string">Corey E, Long A, Rubenstein S. Computer-assisted analysis in organic synthesis. Science. 1985;228:408–418. doi: 10.1126/science.3838594.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1126/science.3838594"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="3838594"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Science&amp;title=Computer-assisted analysis in organic synthesis&amp;author=E Corey&amp;author=A Long&amp;author=S Rubenstein&amp;volume=228&amp;publication_year=1985&amp;pages=408-418&amp;pmid=3838594&amp;doi=10.1126/science.3838594&amp;"/></mixed-citation></ref><ref id="CR35"><label>35.</label><mixed-citation><named-content content-type="citation-string">Bøgevig A, et al.  Route design in the 21st century: the IC SYNTH software tool as an idea generator for synthesis prediction. Org. Process Res. Dev. 2015;19:357–368. doi: 10.1021/op500373e.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/op500373e"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Org. Process Res. Dev.&amp;title=Route design in the 21st century: the IC SYNTH software tool as an idea generator for synthesis prediction&amp;author=A Bøgevig&amp;volume=19&amp;publication_year=2015&amp;pages=357-368&amp;doi=10.1021/op500373e&amp;"/></mixed-citation></ref><ref id="CR36"><label>36.</label><mixed-citation><named-content content-type="citation-string">Segler MHS, Preuss M, Waller MP. Planning chemical syntheses with deep neural networks and symbolic AI. Nature. 2018;555:604–610. doi: 10.1038/nature25978.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/nature25978"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="29595767"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nature&amp;title=Planning chemical syntheses with deep neural networks and symbolic AI&amp;author=MHS Segler&amp;author=M Preuss&amp;author=MP Waller&amp;volume=555&amp;publication_year=2018&amp;pages=604-610&amp;pmid=29595767&amp;doi=10.1038/nature25978&amp;"/></mixed-citation></ref><ref id="CR37"><label>37.</label><mixed-citation><named-content content-type="citation-string">Genheden S, et al.  AiZynthFinder: a fast, robust and flexible open-source software for retrosynthetic planning. J. Cheminformatics. 2020;12:70. doi: 10.1186/s13321-020-00472-1.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1186/s13321-020-00472-1"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7672904"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33292482"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Cheminformatics&amp;title=AiZynthFinder: a fast, robust and flexible open-source software for retrosynthetic planning&amp;author=S Genheden&amp;volume=12&amp;publication_year=2020&amp;pages=70&amp;pmid=33292482&amp;doi=10.1186/s13321-020-00472-1&amp;"/></mixed-citation></ref><ref id="CR38"><label>38.</label><mixed-citation><named-content content-type="citation-string">Mikulak-Klucznik B, et al.  Computational planning of the synthesis of complex natural products. Nature. 2020;588:83–88. doi: 10.1038/s41586-020-2855-y.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-020-2855-y"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33049755"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nature&amp;title=Computational planning of the synthesis of complex natural products&amp;author=B Mikulak-Klucznik&amp;volume=588&amp;publication_year=2020&amp;pages=83-88&amp;pmid=33049755&amp;doi=10.1038/s41586-020-2855-y&amp;"/></mixed-citation></ref><ref id="CR39"><label>39.</label><mixed-citation><named-content content-type="citation-string">Bachmann BO. Biosynthesis: is it time to go retro? Nat. Chem. Biol. 2010;6:390–393. doi: 10.1038/nchembio.377.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/nchembio.377"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="20479744"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nat. Chem. Biol.&amp;title=Biosynthesis: is it time to go retro?&amp;author=BO Bachmann&amp;volume=6&amp;publication_year=2010&amp;pages=390-393&amp;pmid=20479744&amp;doi=10.1038/nchembio.377&amp;"/></mixed-citation></ref><ref id="CR40"><label>40.</label><mixed-citation><named-content content-type="citation-string">Lowe, D. Chemical reactions from US patents (1976-Sep2016). figshare 10.6084/m9.figshare.5104873.v1. (2017).</named-content></mixed-citation></ref><ref id="CR41"><label>41.</label><mixed-citation><named-content content-type="citation-string">Badowski T, Gajewska EP, Molga K, Grzybowski BA. Synergy between expert and machine-learning approaches allows for improved retrosynthetic planning. Angew. Chem. Int. Ed. 2020;59:725–730. doi: 10.1002/anie.201912083.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1002/anie.201912083"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="31750610"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Angew. Chem. Int. Ed.&amp;title=Synergy between expert and machine-learning approaches allows for improved retrosynthetic planning&amp;author=T Badowski&amp;author=EP Gajewska&amp;author=K Molga&amp;author=BA Grzybowski&amp;volume=59&amp;publication_year=2020&amp;pages=725-730&amp;pmid=31750610&amp;doi=10.1002/anie.201912083&amp;"/></mixed-citation></ref><ref id="CR42"><label>42.</label><mixed-citation><named-content content-type="citation-string">Tokic M, et al.  Discovery and evaluation of biosynthetic pathways for the production of five methyl ethyl ketone precursors. ACS Synth. Biol. 2018;7:1858–1873. doi: 10.1021/acssynbio.8b00049.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acssynbio.8b00049"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="30021444"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=ACS Synth. Biol.&amp;title=Discovery and evaluation of biosynthetic pathways for the production of five methyl ethyl ketone precursors&amp;author=M Tokic&amp;volume=7&amp;publication_year=2018&amp;pages=1858-1873&amp;pmid=30021444&amp;doi=10.1021/acssynbio.8b00049&amp;"/></mixed-citation></ref><ref id="CR43"><label>43.</label><mixed-citation><named-content content-type="citation-string">Sankaranarayanan K, et al.  Similarity based enzymatic retrosynthesis. Chem. Sci. 2022;13:6039–6053. doi: 10.1039/D2SC01588A.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1039/D2SC01588A"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC9132021"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="35685792"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Chem. Sci.&amp;title=Similarity based enzymatic retrosynthesis&amp;author=K Sankaranarayanan&amp;volume=13&amp;publication_year=2022&amp;pages=6039-6053&amp;pmid=35685792&amp;doi=10.1039/D2SC01588A&amp;"/></mixed-citation></ref><ref id="CR44"><label>44.</label><mixed-citation><named-content content-type="citation-string">Kanehisa M, Goto S. KEGG: Kyoto Encyclopedia of genes and genomes. Nucleic Acids Res. 2000;28:27–30. doi: 10.1093/nar/28.1.27.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/nar/28.1.27"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC102409"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="10592173"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nucleic Acids Res.&amp;title=KEGG: Kyoto Encyclopedia of genes and genomes&amp;author=M Kanehisa&amp;author=S Goto&amp;volume=28&amp;publication_year=2000&amp;pages=27-30&amp;pmid=10592173&amp;doi=10.1093/nar/28.1.27&amp;"/></mixed-citation></ref><ref id="CR45"><label>45.</label><mixed-citation><named-content content-type="citation-string">Moretti S, Tran V, Mehl F, Ibberson M, Pagni M. MetaNetX/MNXref: unified namespace for metabolites and biochemical reactions in the context of metabolic models. Nucleic Acids Res. 2021;49:D570–D574. doi: 10.1093/nar/gkaa992.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/nar/gkaa992"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7778905"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33156326"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nucleic Acids Res.&amp;title=MetaNetX/MNXref: unified namespace for metabolites and biochemical reactions in the context of metabolic models&amp;author=S Moretti&amp;author=V Tran&amp;author=F Mehl&amp;author=M Ibberson&amp;author=M Pagni&amp;volume=49&amp;publication_year=2021&amp;pages=D570-D574&amp;pmid=33156326&amp;doi=10.1093/nar/gkaa992&amp;"/></mixed-citation></ref><ref id="CR46"><label>46.</label><mixed-citation><named-content content-type="citation-string">Bansal P, et al.  Rhea, the reaction knowledgebase in 2022. Nucleic Acids Res. 2022;50:D693–D700. doi: 10.1093/nar/gkab1016.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/nar/gkab1016"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8728268"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="34755880"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nucleic Acids Res.&amp;title=Rhea, the reaction knowledgebase in 2022&amp;author=P Bansal&amp;volume=50&amp;publication_year=2022&amp;pages=D693-D700&amp;pmid=34755880&amp;doi=10.1093/nar/gkab1016&amp;"/></mixed-citation></ref><ref id="CR47"><label>47.</label><mixed-citation><named-content content-type="citation-string">Segler MHS, Waller MP. Neural-symbolic machine learning for retrosynthesis and reaction prediction. Chemistry. 2017;23:5966–5971. doi: 10.1002/chem.201605499.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1002/chem.201605499"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="28134452"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Chemistry&amp;title=Neural-symbolic machine learning for retrosynthesis and reaction prediction&amp;author=MHS Segler&amp;author=MP Waller&amp;volume=23&amp;publication_year=2017&amp;pages=5966-5971&amp;pmid=28134452&amp;doi=10.1002/chem.201605499&amp;"/></mixed-citation></ref><ref id="CR48"><label>48.</label><mixed-citation><named-content content-type="citation-string">Lang M, Stelzer M, Schomburg D. BKM-react, an integrated biochemical reaction database. BMC Biochem. 2011;12:42. doi: 10.1186/1471-2091-12-42.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1186/1471-2091-12-42"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC3167764"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="21824409"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=BMC Biochem.&amp;title=BKM-react, an integrated biochemical reaction database&amp;author=M Lang&amp;author=M Stelzer&amp;author=D Schomburg&amp;volume=12&amp;publication_year=2011&amp;pages=42&amp;pmid=21824409&amp;doi=10.1186/1471-2091-12-42&amp;"/></mixed-citation></ref><ref id="CR49"><label>49.</label><mixed-citation><named-content content-type="citation-string">Chang A, et al.  BRENDA, the ELIXIR core data resource in 2021: new developments and updates. Nucleic Acids Res. 2021;49:D498–D508. doi: 10.1093/nar/gkaa1025.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/nar/gkaa1025"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7779020"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33211880"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nucleic Acids Res.&amp;title=BRENDA, the ELIXIR core data resource in 2021: new developments and updates&amp;author=A Chang&amp;volume=49&amp;publication_year=2021&amp;pages=D498-D508&amp;pmid=33211880&amp;doi=10.1093/nar/gkaa1025&amp;"/></mixed-citation></ref><ref id="CR50"><label>50.</label><mixed-citation><named-content content-type="citation-string">Karp PD, et al.  The BioCyc collection of microbial genomes and metabolic pathways. Brief. Bioinforma. 2019;20:1085–1093. doi: 10.1093/bib/bbx085.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/bib/bbx085"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC6781571"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="29447345"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Brief. Bioinforma.&amp;title=The BioCyc collection of microbial genomes and metabolic pathways&amp;author=PD Karp&amp;volume=20&amp;publication_year=2019&amp;pages=1085-1093&amp;pmid=29447345&amp;doi=10.1093/bib/bbx085&amp;"/></mixed-citation></ref><ref id="CR51"><label>51.</label><mixed-citation><named-content content-type="citation-string">Wittig U, et al.  SABIO-RK-database for biochemical reaction kinetics. Nucleic Acids Res. 2012;40:D790–D796. doi: 10.1093/nar/gkr1046.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/nar/gkr1046"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC3245076"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="22102587"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nucleic Acids Res.&amp;title=SABIO-RK-database for biochemical reaction kinetics&amp;author=U Wittig&amp;volume=40&amp;publication_year=2012&amp;pages=D790-D796&amp;pmid=22102587&amp;doi=10.1093/nar/gkr1046&amp;"/></mixed-citation></ref><ref id="CR52"><label>52.</label><mixed-citation><named-content content-type="citation-string">Weininger D. SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Model. 1988;28:31–36. doi: 10.1021/ci00057a005.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/ci00057a005"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Chem. Inf. Model.&amp;title=SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules&amp;author=D Weininger&amp;volume=28&amp;publication_year=1988&amp;pages=31-36&amp;doi=10.1021/ci00057a005&amp;"/></mixed-citation></ref><ref id="CR53"><label>53.</label><mixed-citation><named-content content-type="citation-string">Coley CW, Green WH, Jensen KF. RDChiral: an RDKit wrapper for handling stereochemistry in retrosynthetic template extraction and application. J. Chem. Inf. Model. 2019;59:2529–2537. doi: 10.1021/acs.jcim.9b00286.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.jcim.9b00286"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="31190540"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Chem. Inf. Model.&amp;title=RDChiral: an RDKit wrapper for handling stereochemistry in retrosynthetic template extraction and application&amp;author=CW Coley&amp;author=WH Green&amp;author=KF Jensen&amp;volume=59&amp;publication_year=2019&amp;pages=2529-2537&amp;pmid=31190540&amp;doi=10.1021/acs.jcim.9b00286&amp;"/></mixed-citation></ref><ref id="CR54"><label>54.</label><mixed-citation><named-content content-type="citation-string">Polykovskiy, D. et al. Molecular sets (MOSES): A benchmarking platform for molecular generation models. <italic>Front. Pharmacol.</italic><bold>11</bold>, 565644 (2020).</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3389/fphar.2020.565644"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7775580"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33390943"/></mixed-citation></ref><ref id="CR55"><label>55.</label><mixed-citation><named-content content-type="citation-string">Sterling T, Irwin JJ. ZINC 15 - ligand discovery for everyone. J. Chem. Inf. Model. 2015;55:2324–2337. doi: 10.1021/acs.jcim.5b00559.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.jcim.5b00559"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC4658288"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="26479676"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Chem. Inf. Model.&amp;title=ZINC 15 - ligand discovery for everyone&amp;author=T Sterling&amp;author=JJ Irwin&amp;volume=55&amp;publication_year=2015&amp;pages=2324-2337&amp;pmid=26479676&amp;doi=10.1021/acs.jcim.5b00559&amp;"/></mixed-citation></ref><ref id="CR56"><label>56.</label><mixed-citation><named-content content-type="citation-string">Durak LJ, Payne JT, Lewis JC. Late-stage diversification of biologically active molecules via chemoenzymatic C-H functionalization. ACS Catal. 2016;6:1451–1454. doi: 10.1021/acscatal.5b02558.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acscatal.5b02558"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC4890977"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="27274902"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=ACS Catal.&amp;title=Late-stage diversification of biologically active molecules via chemoenzymatic C-H functionalization&amp;author=LJ Durak&amp;author=JT Payne&amp;author=JC Lewis&amp;volume=6&amp;publication_year=2016&amp;pages=1451-1454&amp;pmid=27274902&amp;doi=10.1021/acscatal.5b02558&amp;"/></mixed-citation></ref><ref id="CR57"><label>57.</label><mixed-citation><named-content content-type="citation-string">Cornwall P, Diorazio LJ, Monks N. Route design, the foundation of successful chemical development. Bioorg. Med. Chem. 2018;26:4336–4347. doi: 10.1016/j.bmc.2018.06.006.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.bmc.2018.06.006"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="29925485"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Bioorg. Med. Chem.&amp;title=Route design, the foundation of successful chemical development&amp;author=P Cornwall&amp;author=LJ Diorazio&amp;author=N Monks&amp;volume=26&amp;publication_year=2018&amp;pages=4336-4347&amp;pmid=29925485&amp;doi=10.1016/j.bmc.2018.06.006&amp;"/></mixed-citation></ref><ref id="CR58"><label>58.</label><mixed-citation><named-content content-type="citation-string">Korman TP, Opgenorth PH, Bowie JU. A synthetic biochemistry platform for cell free production of monoterpenes from glucose. Nat. Commun. 2017;8:15526. doi: 10.1038/ncomms15526.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/ncomms15526"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC5458089"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="28537253"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nat. Commun.&amp;title=A synthetic biochemistry platform for cell free production of monoterpenes from glucose&amp;author=TP Korman&amp;author=PH Opgenorth&amp;author=JU Bowie&amp;volume=8&amp;publication_year=2017&amp;pages=15526&amp;pmid=28537253&amp;doi=10.1038/ncomms15526&amp;"/></mixed-citation></ref><ref id="CR59"><label>59.</label><mixed-citation><named-content content-type="citation-string">Aizpurua-Olaizola O, et al.  Evolution of the cannabinoid and terpene content during the growth of Cannabis sativa plants from different chemotypes. J. Nat. Prod. 2016;79:324–331. doi: 10.1021/acs.jnatprod.5b00949.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.jnatprod.5b00949"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="26836472"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Nat. Prod.&amp;title=Evolution of the cannabinoid and terpene content during the growth of Cannabis sativa plants from different chemotypes&amp;author=O Aizpurua-Olaizola&amp;volume=79&amp;publication_year=2016&amp;pages=324-331&amp;pmid=26836472&amp;doi=10.1021/acs.jnatprod.5b00949&amp;"/></mixed-citation></ref><ref id="CR60"><label>60.</label><mixed-citation><named-content content-type="citation-string">Shultz ZP, Lawrence GA, Jacobson JM, Cruz EJ, Leahy JW. Enantioselective total synthesis of cannabinoids-A route for analogue development. Org. Lett. 2018;20:381–384. doi: 10.1021/acs.orglett.7b03668.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.orglett.7b03668"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="29293352"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Org. Lett.&amp;title=Enantioselective total synthesis of cannabinoids-A route for analogue development&amp;author=ZP Shultz&amp;author=GA Lawrence&amp;author=JM Jacobson&amp;author=EJ Cruz&amp;author=JW Leahy&amp;volume=20&amp;publication_year=2018&amp;pages=381-384&amp;pmid=29293352&amp;doi=10.1021/acs.orglett.7b03668&amp;"/></mixed-citation></ref><ref id="CR61"><label>61.</label><mixed-citation><named-content content-type="citation-string">Cheng L-J, Xie J-H, Chen Y, Wang L-X, Zhou Q-L. Enantioselective total synthesis of (-)-Δ8-THC and (-)-Δ9-THC via catalytic asymmetric hydrogenation and SNAr cyclization. Org. Lett. 2013;15:764–767. doi: 10.1021/ol303351y.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/ol303351y"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="23346909"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Org. Lett.&amp;title=Enantioselective total synthesis of (-)-Δ8-THC and (-)-Δ9-THC via catalytic asymmetric hydrogenation and SNAr cyclization&amp;author=L-J Cheng&amp;author=J-H Xie&amp;author=Y Chen&amp;author=L-X Wang&amp;author=Q-L Zhou&amp;volume=15&amp;publication_year=2013&amp;pages=764-767&amp;pmid=23346909&amp;doi=10.1021/ol303351y&amp;"/></mixed-citation></ref><ref id="CR62"><label>62.</label><mixed-citation><named-content content-type="citation-string">Schafroth MA, Zuccarello G, Krautwald S, Sarlah D, Carreira EM. Stereodivergent total synthesis of Δ9-tetrahydrocannabinols. Angew. Chem. Int. Ed. 2014;126:14118–14121. doi: 10.1002/ange.201408380.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1002/ange.201408380"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="25303495"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Angew. Chem. Int. Ed.&amp;title=Stereodivergent total synthesis of Δ9-tetrahydrocannabinols&amp;author=MA Schafroth&amp;author=G Zuccarello&amp;author=S Krautwald&amp;author=D Sarlah&amp;author=EM Carreira&amp;volume=126&amp;publication_year=2014&amp;pages=14118-14121&amp;pmid=25303495&amp;doi=10.1002/ange.201408380&amp;"/></mixed-citation></ref><ref id="CR63"><label>63.</label><mixed-citation><named-content content-type="citation-string">Ametovski A, Lupton DW. Enantioselective total synthesis of (-)-Δ9-tetrahydrocannabinol via N-heterocyclic carbene catalysis. Org. Lett. 2019;21:1212–1215. doi: 10.1021/acs.orglett.9b00198.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.orglett.9b00198"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="30726088"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Org. Lett.&amp;title=Enantioselective total synthesis of (-)-Δ9-tetrahydrocannabinol via N-heterocyclic carbene catalysis&amp;author=A Ametovski&amp;author=DW Lupton&amp;volume=21&amp;publication_year=2019&amp;pages=1212-1215&amp;pmid=30726088&amp;doi=10.1021/acs.orglett.9b00198&amp;"/></mixed-citation></ref><ref id="CR64"><label>64.</label><mixed-citation><named-content content-type="citation-string">Evans DA, et al.  Bis(oxazoline) and bis(oxazolinyl)pyridine copper complexes as enantioselective Diels-Alder catalysts: reaction scope and synthetic applications. J. Am. Chem. Soc. 1999;121:7582–7594. doi: 10.1021/ja991191c.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/ja991191c"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Am. Chem. Soc.&amp;title=Bis(oxazoline) and bis(oxazolinyl)pyridine copper complexes as enantioselective Diels-Alder catalysts: reaction scope and synthetic applications&amp;author=DA Evans&amp;volume=121&amp;publication_year=1999&amp;pages=7582-7594&amp;doi=10.1021/ja991191c&amp;"/></mixed-citation></ref><ref id="CR65"><label>65.</label><mixed-citation><named-content content-type="citation-string">Trost BM, Dogra K. Synthesis of (-)-Δ9-trans-tetrahydrocannabinol: stereocontrol via Mo-catalyzed asymmetric allylic alkylation reaction. Org. Lett. 2007;9:861–863. doi: 10.1021/ol063022k.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/ol063022k"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC2597621"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="17266321"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Org. Lett.&amp;title=Synthesis of (-)-Δ9-trans-tetrahydrocannabinol: stereocontrol via Mo-catalyzed asymmetric allylic alkylation reaction&amp;author=BM Trost&amp;author=K Dogra&amp;volume=9&amp;publication_year=2007&amp;pages=861-863&amp;pmid=17266321&amp;doi=10.1021/ol063022k&amp;"/></mixed-citation></ref><ref id="CR66"><label>66.</label><mixed-citation><named-content content-type="citation-string">Luo X, et al.  Complete biosynthesis of cannabinoids and their unnatural analogues in yeast. Nature. 2019;567:123–126. doi: 10.1038/s41586-019-0978-9.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41586-019-0978-9"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="30814733"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nature&amp;title=Complete biosynthesis of cannabinoids and their unnatural analogues in yeast&amp;author=X Luo&amp;volume=567&amp;publication_year=2019&amp;pages=123-126&amp;pmid=30814733&amp;doi=10.1038/s41586-019-0978-9&amp;"/></mixed-citation></ref><ref id="CR67"><label>67.</label><mixed-citation><named-content content-type="citation-string">Valliere, M. A., Korman, T. P., Arbing, M. A. &amp; Bowie, J. U. A bio-inspired cell-free system for cannabinoid production from inexpensive inputs. <italic>Nat. Chem. Biol.</italic><bold>16</bold>, 1427–1433 (2020).</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/s41589-020-0631-9"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32839605"/></mixed-citation></ref><ref id="CR68"><label>68.</label><mixed-citation><named-content content-type="citation-string">Jentsch NG, Zhang X, Magolan J. Efficient synthesis of cannabigerol, grifolin, and piperogalin via alumina-promoted allylation. J. Nat. Products. 2020;83:2587–2591. doi: 10.1021/acs.jnatprod.0c00131.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.jnatprod.0c00131"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32972142"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Nat. Products&amp;title=Efficient synthesis of cannabigerol, grifolin, and piperogalin via alumina-promoted allylation&amp;author=NG Jentsch&amp;author=X Zhang&amp;author=J Magolan&amp;volume=83&amp;publication_year=2020&amp;pages=2587-2591&amp;pmid=32972142&amp;doi=10.1021/acs.jnatprod.0c00131&amp;"/></mixed-citation></ref><ref id="CR69"><label>69.</label><mixed-citation><named-content content-type="citation-string">Taura F, Morimoto S, Shoyama Y, Mechoulam R. First direct evidence for the mechanism of Δ1-tetrahydrocannabinolic acid biosynthesis. J. Am. Chem. Soc. 1995;117:9766–9767. doi: 10.1021/ja00143a024.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/ja00143a024"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Am. Chem. Soc.&amp;title=First direct evidence for the mechanism of Δ1-tetrahydrocannabinolic acid biosynthesis&amp;author=F Taura&amp;author=S Morimoto&amp;author=Y Shoyama&amp;author=R Mechoulam&amp;volume=117&amp;publication_year=1995&amp;pages=9766-9767&amp;doi=10.1021/ja00143a024&amp;"/></mixed-citation></ref><ref id="CR70"><label>70.</label><mixed-citation><named-content content-type="citation-string">Shoyama Y, et al.  Structure and function of Δ1-tetrahydrocannabinolic acid (THCA) synthase, the enzyme controlling the psychoactivity of Cannabis sativa. J. Mol. Biol. 2012;423:96–105. doi: 10.1016/j.jmb.2012.06.030.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.jmb.2012.06.030"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="22766313"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Mol. Biol.&amp;title=Structure and function of Δ1-tetrahydrocannabinolic acid (THCA) synthase, the enzyme controlling the psychoactivity of Cannabis sativa&amp;author=Y Shoyama&amp;volume=423&amp;publication_year=2012&amp;pages=96-105&amp;pmid=22766313&amp;doi=10.1016/j.jmb.2012.06.030&amp;"/></mixed-citation></ref><ref id="CR71"><label>71.</label><mixed-citation><named-content content-type="citation-string">Hett R, Fang QK, Gao Y, Wald SA, Senanayake CH. Large-scale synthesis of enantio- and diastereomerically pure (R, R)-formoterol. Org. Process Res. Dev. 1998;2:96–99. doi: 10.1021/op970116o.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/op970116o"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Org. Process Res. Dev.&amp;title=Large-scale synthesis of enantio- and diastereomerically pure (R, R)-formoterol&amp;author=R Hett&amp;author=QK Fang&amp;author=Y Gao&amp;author=SA Wald&amp;author=CH Senanayake&amp;volume=2&amp;publication_year=1998&amp;pages=96-99&amp;doi=10.1021/op970116o&amp;"/></mixed-citation></ref><ref id="CR72"><label>72.</label><mixed-citation><named-content content-type="citation-string">Campos, F., Bosch, M. P. &amp; Guerrero, A. An effcient enantioselective synthesis of (R,R)-formoterol, a potent bronchodilator, using lipases. <italic>Tetrahedron Asymmetry</italic><bold>13</bold>, 2705–2717 (2000).</named-content></mixed-citation></ref><ref id="CR73"><label>73.</label><mixed-citation><named-content content-type="citation-string">Huang L, et al.  The asymmetric synthesis of (R,R)-formoterol via transfer hydrogenation with polyethylene glycol bound Rh catalyst in PEG2000 and water. Chirality. 2010;22:206–211. doi: 10.1002/chir.20728.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1002/chir.20728"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="19408330"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Chirality&amp;title=The asymmetric synthesis of (R,R)-formoterol via transfer hydrogenation with polyethylene glycol bound Rh catalyst in PEG2000 and water&amp;author=L Huang&amp;volume=22&amp;publication_year=2010&amp;pages=206-211&amp;pmid=19408330&amp;doi=10.1002/chir.20728&amp;"/></mixed-citation></ref><ref id="CR74"><label>74.</label><mixed-citation><named-content content-type="citation-string">Wu B, Szymański W, Heberling MM, Feringa BL, Janssen DB. Aminomutases: mechanistic diversity, biotechnological applications and future perspectives. Trends Biotechnol. 2011;29:352–362. doi: 10.1016/j.tibtech.2011.02.005.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.tibtech.2011.02.005"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="21477876"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Trends Biotechnol.&amp;title=Aminomutases: mechanistic diversity, biotechnological applications and future perspectives&amp;author=B Wu&amp;author=W Szymański&amp;author=MM Heberling&amp;author=BL Feringa&amp;author=DB Janssen&amp;volume=29&amp;publication_year=2011&amp;pages=352-362&amp;pmid=21477876&amp;doi=10.1016/j.tibtech.2011.02.005&amp;"/></mixed-citation></ref><ref id="CR75"><label>75.</label><mixed-citation><named-content content-type="citation-string">Parmeggiani F, Weise NJ, Ahmed ST, Turner NJ. Synthetic and therapeutic applications of ammonia-lyases and aminomutases. Chem. Rev. 2018;118:73–118. doi: 10.1021/acs.chemrev.6b00824.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.chemrev.6b00824"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="28497955"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Chem. Rev.&amp;title=Synthetic and therapeutic applications of ammonia-lyases and aminomutases&amp;author=F Parmeggiani&amp;author=NJ Weise&amp;author=ST Ahmed&amp;author=NJ Turner&amp;volume=118&amp;publication_year=2018&amp;pages=73-118&amp;pmid=28497955&amp;doi=10.1021/acs.chemrev.6b00824&amp;"/></mixed-citation></ref><ref id="CR76"><label>76.</label><mixed-citation><named-content content-type="citation-string">Maity AN, Chen Y-H, Ke S-C. Large-scale domain motions and pyridoxal-5’-phosphate assisted radical catalysis in coenzyme B12-dependent aminomutases. Int. J. Mol. Sci. 2014;15:3064–3087. doi: 10.3390/ijms15023064.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/ijms15023064"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC3958899"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="24562332"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Int. J. Mol. Sci.&amp;title=Large-scale domain motions and pyridoxal-5’-phosphate assisted radical catalysis in coenzyme B12-dependent aminomutases&amp;author=AN Maity&amp;author=Y-H Chen&amp;author=S-C Ke&amp;volume=15&amp;publication_year=2014&amp;pages=3064-3087&amp;pmid=24562332&amp;doi=10.3390/ijms15023064&amp;"/></mixed-citation></ref><ref id="CR77"><label>77.</label><mixed-citation><named-content content-type="citation-string">Kille S, Zilly FE, Acevedo JP, Reetz MT. Regio- and stereoselectivity of P450-catalysed hydroxylation of steroids controlled by laboratory evolution. Nat. Chem. 2011;3:738–743. doi: 10.1038/nchem.1113.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/nchem.1113"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="21860465"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nat. Chem.&amp;title=Regio- and stereoselectivity of P450-catalysed hydroxylation of steroids controlled by laboratory evolution&amp;author=S Kille&amp;author=FE Zilly&amp;author=JP Acevedo&amp;author=MT Reetz&amp;volume=3&amp;publication_year=2011&amp;pages=738-743&amp;pmid=21860465&amp;doi=10.1038/nchem.1113&amp;"/></mixed-citation></ref><ref id="CR78"><label>78.</label><mixed-citation><named-content content-type="citation-string">Zhu D, et al.  Inverting the enantioselectivity of a carbonyl reductase via substrate-enzyme docking-guided point mutation. Org. Lett. 2008;10:525–528. doi: 10.1021/ol702638j.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/ol702638j"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="18205368"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Org. Lett.&amp;title=Inverting the enantioselectivity of a carbonyl reductase via substrate-enzyme docking-guided point mutation&amp;author=D Zhu&amp;volume=10&amp;publication_year=2008&amp;pages=525-528&amp;pmid=18205368&amp;doi=10.1021/ol702638j&amp;"/></mixed-citation></ref><ref id="CR79"><label>79.</label><mixed-citation><named-content content-type="citation-string">Pratter SM, et al.  Inversion of enantioselectivity of a mononuclear non-heme iron(II)-dependent hydroxylase by tuning the interplay of metal-center geometry and protein structure. Angew. Chem. Int. Ed. 2013;125:9859–9863. doi: 10.1002/ange.201304633.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1002/ange.201304633"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="23881738"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Angew. Chem. Int. Ed.&amp;title=Inversion of enantioselectivity of a mononuclear non-heme iron(II)-dependent hydroxylase by tuning the interplay of metal-center geometry and protein structure&amp;author=SM Pratter&amp;volume=125&amp;publication_year=2013&amp;pages=9859-9863&amp;pmid=23881738&amp;doi=10.1002/ange.201304633&amp;"/></mixed-citation></ref><ref id="CR80"><label>80.</label><mixed-citation><named-content content-type="citation-string">May O, Nguyen PT, Arnold FH. Inverting enantioselectivity by directed evolution of hydantoinase for improved production of L-methionine. Nat. Biotechnol. 2000;18:317–320. doi: 10.1038/73773.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1038/73773"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="10700149"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Nat. Biotechnol.&amp;title=Inverting enantioselectivity by directed evolution of hydantoinase for improved production of L-methionine&amp;author=O May&amp;author=PT Nguyen&amp;author=FH Arnold&amp;volume=18&amp;publication_year=2000&amp;pages=317-320&amp;pmid=10700149&amp;doi=10.1038/73773&amp;"/></mixed-citation></ref><ref id="CR81"><label>81.</label><mixed-citation><named-content content-type="citation-string">Ghislieri D, Turner NJ. Biocatalytic approaches to the synthesis of enantiomerically pure chiral amines. Top. Catal. 2014;57:284–300. doi: 10.1007/s11244-013-0184-1.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1007/s11244-013-0184-1"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Top. Catal.&amp;title=Biocatalytic approaches to the synthesis of enantiomerically pure chiral amines&amp;author=D Ghislieri&amp;author=NJ Turner&amp;volume=57&amp;publication_year=2014&amp;pages=284-300&amp;doi=10.1007/s11244-013-0184-1&amp;"/></mixed-citation></ref><ref id="CR82"><label>82.</label><mixed-citation><named-content content-type="citation-string">van Hylckama Vlieg JET, Leemhuis H, Spelberg JHL, Janssen DB. Characterization of the gene cluster involved in isoprene metabolism in Rhodococcus sp. strain AD45. J. Bacteriol. 2000;182:1956–1963. doi: 10.1128/JB.182.7.1956-1963.2000.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1128/JB.182.7.1956-1963.2000"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC101893"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="10715003"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Bacteriol.&amp;title=Characterization of the gene cluster involved in isoprene metabolism in Rhodococcus sp. strain AD45&amp;author=JET van Hylckama Vlieg&amp;author=H Leemhuis&amp;author=JHL Spelberg&amp;author=DB Janssen&amp;volume=182&amp;publication_year=2000&amp;pages=1956-1963&amp;pmid=10715003&amp;doi=10.1128/JB.182.7.1956-1963.2000&amp;"/></mixed-citation></ref><ref id="CR83"><label>83.</label><mixed-citation><named-content content-type="citation-string">Law J, et al.  Route designer: a retrosynthetic analysis tool utilizing automated retrosynthetic rule generation. J. Chem. Inf. Model. 2009;49:593–602. doi: 10.1021/ci800228y.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/ci800228y"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="19434897"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Chem. Inf. Model.&amp;title=Route designer: a retrosynthetic analysis tool utilizing automated retrosynthetic rule generation&amp;author=J Law&amp;volume=49&amp;publication_year=2009&amp;pages=593-602&amp;pmid=19434897&amp;doi=10.1021/ci800228y&amp;"/></mixed-citation></ref><ref id="CR84"><label>84.</label><mixed-citation><named-content content-type="citation-string">Heid E, Liu J, Aude A, Green WH. Influence of template size, canonicalization, and exclusivity for retrosynthesis and reaction prediction applications. J. Chem. Inf. Model. 2022;62:16–26. doi: 10.1021/acs.jcim.1c01192.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.jcim.1c01192"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8757433"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="34939786"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Chem. Inf. Model.&amp;title=Influence of template size, canonicalization, and exclusivity for retrosynthesis and reaction prediction applications&amp;author=E Heid&amp;author=J Liu&amp;author=A Aude&amp;author=WH Green&amp;volume=62&amp;publication_year=2022&amp;pages=16-26&amp;pmid=34939786&amp;doi=10.1021/acs.jcim.1c01192&amp;"/></mixed-citation></ref><ref id="CR85"><label>85.</label><mixed-citation><named-content content-type="citation-string">Goldman S, Das R, Yang KK, Coley CW. Machine learning modeling of family wide enzyme-substrate specificity screens. PLoS Comput. Biol. 2022;18:e1009853. doi: 10.1371/journal.pcbi.1009853.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1371/journal.pcbi.1009853"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8865696"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="35143485"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=PLoS Comput. Biol.&amp;title=Machine learning modeling of family wide enzyme-substrate specificity screens&amp;author=S Goldman&amp;author=R Das&amp;author=KK Yang&amp;author=CW Coley&amp;volume=18&amp;publication_year=2022&amp;pages=e1009853&amp;pmid=35143485&amp;doi=10.1371/journal.pcbi.1009853&amp;"/></mixed-citation></ref><ref id="CR86"><label>86.</label><mixed-citation><named-content content-type="citation-string">NCBI. PubChem identifier exchange service. <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://pubchem.ncbi.nlm.nih.gov/idexchange/idexchange.cgi" ext-link-type="uri">https://pubchem.ncbi.nlm.nih.gov/idexchange/idexchange.cgi</ext-link>.</named-content></mixed-citation></ref><ref id="CR87"><label>87.</label><mixed-citation><named-content content-type="citation-string">Rahman SA, et al.  Reaction decoder tool (RDT): extracting features from chemical reactions. Bioinformatics. 2016;32:2065–2066. doi: 10.1093/bioinformatics/btw096.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1093/bioinformatics/btw096"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC4920114"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="27153692"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Bioinformatics&amp;title=Reaction decoder tool (RDT): extracting features from chemical reactions&amp;author=SA Rahman&amp;volume=32&amp;publication_year=2016&amp;pages=2065-2066&amp;pmid=27153692&amp;doi=10.1093/bioinformatics/btw096&amp;"/></mixed-citation></ref><ref id="CR88"><label>88.</label><mixed-citation><named-content content-type="citation-string">RDKit. <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://www.rdkit.org/" ext-link-type="uri">http://www.rdkit.org/</ext-link>.</named-content></mixed-citation></ref><ref id="CR89"><label>89.</label><mixed-citation><named-content content-type="citation-string">Fortunato, M. E., Coley, C. W. &amp; Barnes, B. C. Machine learned prediction of reaction template applicability for data-driven retrosynthetic predictions of energetic materials. <italic>AIP Conf. Proc</italic>. <bold>2272</bold>, 070014 (2020).</named-content></mixed-citation></ref><ref id="CR90"><label>90.</label><mixed-citation><named-content content-type="citation-string">Srivastava, R. K., Greff, K. &amp; Schmidhuber, J. Training very deep networks. <italic>Advances in Neural Information Processing</italic> (2015).</named-content></mixed-citation></ref><ref id="CR91"><label>91.</label><mixed-citation><named-content content-type="citation-string">Fortunato ME, Coley CW, Barnes BC, Jensen KF. Data augmentation and pretraining for template-based retrosynthetic prediction in computer-aided synthesis planning. J. Chem. Inf. Model. 2020;60:3398–3407. doi: 10.1021/acs.jcim.0c00403.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1021/acs.jcim.0c00403"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32568548"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Chem. Inf. Model.&amp;title=Data augmentation and pretraining for template-based retrosynthetic prediction in computer-aided synthesis planning&amp;author=ME Fortunato&amp;author=CW Coley&amp;author=BC Barnes&amp;author=KF Jensen&amp;volume=60&amp;publication_year=2020&amp;pages=3398-3407&amp;pmid=32568548&amp;doi=10.1021/acs.jcim.0c00403&amp;"/></mixed-citation></ref><ref id="CR92"><label>92.</label><mixed-citation><named-content content-type="citation-string">Levin, I. bkms-data. <italic>Zenodo</italic>. 10.5281/zenodo.7334523 (2022).</named-content></mixed-citation></ref><ref id="CR93"><label>93.</label><mixed-citation><named-content content-type="citation-string">Levin, I. chemoenzymatic-askcos. <italic>Zenodo</italic>. 10.5281/zenodo.7334532 (2022).</named-content></mixed-citation></ref><ref id="CR94"><label>94.</label><mixed-citation><named-content content-type="citation-string">Levin, I. hybmind. <italic>Zenodo</italic>. 10.5281/zenodo.7334538 (2022).</named-content></mixed-citation></ref></ref-list></sec></sec><sec id="_ad93_" xml:lang="en" sec-type="associated-data" disp-level="1"><title>Associated Data</title><sec id="_adsm93_" xml:lang="en" sec-type="supplementary-materials" disp-level="2"><title>Supplementary Materials</title><supplementary-material id="db_ds_supplementary-material1_reqid_" position="float"><media xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="41467_2022_35422_MOESM1_ESM.pdf" mimetype="application" mime-subtype="pdf"><?cloudpmc-path 46c9/9750992/689845310b63/41467_2022_35422_MOESM1_ESM.pdf?><?cloudpmc-bucket app?><?size 3352902?><caption><p>Supplementary Information</p></caption></media></supplementary-material></sec><sec id="_adda93_" xml:lang="en" sec-type="data-availability-statement" disp-level="2"><title>Data Availability Statement</title><p>Reaxys reaction data used to build the previously published synthetic organic model is the intellectual property of Elsevier and cannot be shared. The templates extracted from the Reaxys data and all of the BKMS data used to build the enzymatic model and associated templates is available at <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/itai-levin/bkms-data" ext-link-type="uri">https://github.com/itai-levin/bkms-data</ext-link><sup><xref rid="CR92" ref-type="bibr">92</xref></sup>.</p><p>The source code for deploying ASKCOS is available at <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/itai-levin/chemoenzymatic-askcos" ext-link-type="uri">https://github.com/itai-levin/chemoenzymatic-askcos</ext-link><sup><xref rid="CR93" ref-type="bibr">93</xref></sup>. The code for performing the analyses in this manuscript is available at <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/itai-levin/hybmind" ext-link-type="uri">https://github.com/itai-levin/hybmind</ext-link><sup><xref rid="CR94" ref-type="bibr">94</xref></sup>. This repository contains a detailed README with instructions on how to generate the results presented in the manuscript.</p></sec></sec></body></article>