<?xml version="1.0" encoding="UTF-8"?><article xml:lang="en" article-type="research-article"><front><journal-meta><journal-id journal-id-type="pmc-domain-id">1660</journal-id><journal-id journal-id-type="pmc-domain">sensors</journal-id><journal-title-group><journal-title>Sensors (Basel, Switzerland)</journal-title><abbrev-journal-title>Sensors (Basel)</abbrev-journal-title></journal-title-group><publisher><publisher-name>Multidisciplinary Digital Publishing Institute (MDPI)</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="pmcid">PMC11479280</article-id><article-id pub-id-type="pmcaid">11479280</article-id><article-id pub-id-type="pmcaiid">11479280</article-id><article-id pub-id-type="pmid">39409507</article-id><article-id pub-id-type="doi">10.3390/s24196467</article-id><title-group><article-title>Integrating Automated Labeling Framework for Enhancing Deep Learning Models to Count Corn Plants Using UAS Imagery</article-title></title-group><contrib-group content-type="author"><contrib><name name-style="western"><surname>Katari</surname><given-names initials="S">Sushma</given-names></name><role>Methodology, Software, Validation, Formal analysis, Investigation, Writing – original draft, Visualization</role><xref ref-type="aff" rid="af1-sensors-24-06467">1</xref><xref rid="c1-sensors-24-06467" ref-type="author-notes">*</xref></contrib><contrib><name name-style="western"><surname>Venkatesh</surname><given-names initials="S">Sandeep</given-names></name><role>Methodology, Software, Data curation, Writing – review &amp; editing</role><xref ref-type="aff" rid="af2-sensors-24-06467">2</xref></contrib><contrib><name name-style="western"><surname>Stewart</surname><given-names initials="C">Christopher</given-names></name><role>Methodology, Writing – review &amp; editing, Supervision</role><xref ref-type="aff" rid="af3-sensors-24-06467">3</xref></contrib><contrib><name name-style="western"><surname>Khanal</surname><given-names initials="S">Sami</given-names></name><role>Conceptualization, Methodology, Resources, Writing – review &amp; editing, Supervision, Project administration, Funding acquisition</role><xref ref-type="aff" rid="af1-sensors-24-06467">1</xref><xref rid="c1-sensors-24-06467" ref-type="author-notes">*</xref></contrib></contrib-group><contrib-group content-type="editor"><contrib><name name-style="western"><surname>Chung</surname><given-names initials="Y">Yongwha</given-names></name><role>Academic Editor</role></contrib></contrib-group><aff id="af1-sensors-24-06467"><label>1</label>Department of Food, Agricultural, and Biological Engineering, Ohio State University, 590 Woody Hayes Dr, Columbus, OH 43210, USA</aff><aff id="af2-sensors-24-06467"><label>2</label>Google, Kirkland, WA 98033, USA; kvsandy22@gmail.com</aff><aff id="af3-sensors-24-06467"><label>3</label>Department of Computer Science and Engineering, Ohio State University, 590 Woody Hayes Dr, Columbus, OH 43210, USA; stewart.962@osu.edu</aff><author-notes><fn id="c1-sensors-24-06467"><label>*</label><p>Correspondence: <email>katari.5@osu.edu</email> (S.K.); <email>khanal.3@osu.edu</email> (S.K.)</p></fn></author-notes><pub-date><day>7</day><month>10</month><year>2024</year></pub-date><volume>24</volume><issue>19</issue><fpage>6467</fpage><page-range>6467</page-range><pub-history><event event-type="pmc-release"><date><day>16</day><month>10</month><year>2024</year></date></event></pub-history><permissions><copyright-statement>© 2024 by the authors.</copyright-statement><license><license-p>Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://creativecommons.org/licenses/by/4.0/" ext-link-type="uri">https://creativecommons.org/licenses/by/4.0/</ext-link>).</license-p></license></permissions><self-uri xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="sensors-24-06467.pdf" content-type="pmc-pdf"><?cloudpmc-path beb8/11479280/4abe668e7aa1/sensors-24-06467.pdf?><?cloudpmc-bucket app?><?size 7809462?></self-uri><abstract id="abstract1"><title>Abstract</title><p>Plant counting is a critical aspect of crop management, providing farmers with valuable insights into seed germination success and within-field variation in crop population density, both of which are key indicators of crop yield and quality. Recent advancements in Unmanned Aerial System (UAS) technology, coupled with deep learning techniques, have facilitated the development of automated plant counting methods. Various computer vision models based on UAS images are available for detecting and classifying crop plants. However, their accuracy relies largely on the availability of substantial manually labeled training datasets. The objective of this study was to develop a robust corn counting model by developing and integrating an automatic image annotation framework. This study used high-spatial-resolution images collected with a DJI Mavic Pro 2 at the V2–V4 growth stage of corn plants from a field in Wooster, Ohio. The automated image annotation process involved extracting corn rows and applying image enhancement techniques to automatically annotate images as either corn or non-corn, resulting in 80% accuracy in identifying corn plants. The accuracy of corn stand identification was further improved by training four deep learning (DL) models, including InceptionV3, VGG16, VGG19, and Vision Transformer (ViT), with annotated images across various datasets. Notably, VGG16 outperformed the other three models, achieving an F1 score of 0.955. When the corn counts were compared to ground truth data across five test regions, VGG achieved an R<sup>2</sup> of 0.94 and an RMSE of 9.95. The integration of an automated image annotation process into the training of the DL models provided notable benefits in terms of model scaling and consistency. The developed framework can efficiently manage large-scale data generation, streamlining the process for the rapid development and deployment of corn counting DL models.</p><sec id="kwd-group1" sec-type="kwd-group" disp-level="2"><p><bold>Keywords:</bold> UAS, plant stand count, crop rows, automatic labeling</p></sec></abstract><custom-meta-group><custom-meta><meta-name>status</meta-name><meta-value>released</meta-value></custom-meta><custom-meta><meta-name>display-pdf</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>is-olf</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-manuscript</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-preprint</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-journal-matter</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-scanned</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-retracted</meta-name><meta-value>no</meta-value></custom-meta></custom-meta-group></article-meta><notes notes-type="article-notes"><sec id="historyarticle-meta1" sec-type="history" disp-level="2"><p>Received 2024 Aug 1; Revised 2024 Sep 20; Accepted 2024 Oct 1; Collection date 2024 Oct.</p></sec></notes></front><body><sec id="sec1-sensors-24-06467" disp-level="1"><title>1. Introduction</title><p>Plant counting is a critical task of crop scouting [<xref rid="B1-sensors-24-06467" ref-type="bibr">1</xref>], offering farmers pivotal insights into seed germination success rates, within-field variation in crop population density [<xref rid="B2-sensors-24-06467" ref-type="bibr">2</xref>,<xref rid="B3-sensors-24-06467" ref-type="bibr">3</xref>], and traits related to crop quality and crop yield [<xref rid="B4-sensors-24-06467" ref-type="bibr">4</xref>]. It can also be used as a parameter for understanding the effects of different crop management practices on crop yield and nutrient quality. A study found that varying plant densities influence forage corn’s nutritive quality, with higher densities leading to reduced forage quality [<xref rid="B5-sensors-24-06467" ref-type="bibr">5</xref>]. Conversely, different tillage practices were found to have no notable effect on forage quality.</p><p>The conventional approach of plant counting relies on manual scouting, which is time and labor intensive, and incurs high costs. Recent advancements in remote sensing and machine learning (ML) techniques have offered alternatives by enabling the implementation of ML algorithms for plant counting without the need for on-site scouting. Images captured via Unmanned Aerial Systems (UASs) provide detailed information about crop fields [<xref rid="B6-sensors-24-06467" ref-type="bibr">6</xref>,<xref rid="B7-sensors-24-06467" ref-type="bibr">7</xref>], which is instrumental in estimating the pant stand counts through various techniques such as image enhancement, ML, and deep learning (DL) algorithms [<xref rid="B8-sensors-24-06467" ref-type="bibr">8</xref>].</p><p>Image enhancement techniques employ computer vision approaches such as thresholding, edge detection, and contrast enhancement to map plant densities in crop fields. Donmez et al. (2021) [<xref rid="B9-sensors-24-06467" ref-type="bibr">9</xref>] applied morphological operations such as erosion and dilation on high-resolution (3.7 cm) five-band UAS images to count citrus trees. They used the Connected Components Labeling (CCL) algorithm on morphologically processed images to automatically detect and count citrus trees in dense patches. While image enhancement techniques alone provided accurate plant count estimates, the use of ML or DL models can enhance accuracy further. Xia et al. (2019) [<xref rid="B10-sensors-24-06467" ref-type="bibr">10</xref>] demonstrated this improvement using a Support Vector Machine (SVM) and Maximum Likelihood Classifier (MLC) for counting cotton plants, where the SVM outperformed the MLC. Similarly, Banerjee et al. (2021) [<xref rid="B11-sensors-24-06467" ref-type="bibr">11</xref>] achieved a higher accuracy (R<sup>2</sup> = 0.86) in estimating wheat seedling count using the Gaussian process regression (GPR) model on UAS-based multispectral images when compared to support vector regression and regression trees. Tavus et al. (2015) [<xref rid="B12-sensors-24-06467" ref-type="bibr">12</xref>] combined image morphological and ML techniques on UAS images in estimating plant stand counts and achieved an accuracy of 87.7% using the K-Nearest Neighbor (K-NN) classification algorithm, indicating the potential of merging multiple techniques for precise plant stand count estimation.</p><p>Recent advancements in ML and deep-learning (DL) algorithms have significantly improved plant stand count estimation. Osco et al. (2021) [<xref rid="B13-sensors-24-06467" ref-type="bibr">13</xref>] developed a Convolutional Neural Network (CNN) with a two-branch structure for locating crop plantation rows and counting plant stands, achieving around 90% accuracy across various crops. Zhang et al. (2020) [<xref rid="B14-sensors-24-06467" ref-type="bibr">14</xref>] addressed challenges related to canopy overlap by integrating leaf count information into their CNN model, also achieving an accuracy of over 90% in estimating rapeseed plant count. Another study by Zhang et al. (2020) [<xref rid="B15-sensors-24-06467" ref-type="bibr">15</xref>] evaluated the Scale Sequence Residual U-Net (SS Res U-Net) algorithm, finding it more accurate and computationally efficient than other U-Net variants. </p><p>In contrast, Vong et al. (2021) [<xref rid="B16-sensors-24-06467" ref-type="bibr">16</xref>] observed that U-Net’s performance in segmenting and counting plants varied with residue cover in UAS-based images, noting reduced accuracy with increased residue. Wang et al. (2021) [<xref rid="B17-sensors-24-06467" ref-type="bibr">17</xref>] used YOLOv3 and a Kalman filter for counting seedlings amidst intricate soil backgrounds using images captured by a moving ground platform, achieving 98% accuracy during growth stages V2–V3. However, the use of ground-based machines can be time consuming and challenging to maneuver without disturbing crops.</p><p>The R-CNN family, including Mask R-CNN, has also been used for crop counting. Machefer et al. (2020) [<xref rid="B18-sensors-24-06467" ref-type="bibr">18</xref>] refined the Mask R-CNN model for the precise counting of potato and lettuce crops, achieving an accuracy of 78% and 92%, respectively. Their research emphasized the superiority of fine-tuning the pre-trained Mask R-CNN compared to manually tuned computer vision algorithms. Given the computational expense of training R-CNN models, Lu and Cao (2020) [<xref rid="B19-sensors-24-06467" ref-type="bibr">19</xref>] implemented faster TasselNetV2+ for counting wheat, maize, and sorghum plants and found that it outperformed Faster R-CNN in terms of computation time and performance. Additionally, a study [<xref rid="B20-sensors-24-06467" ref-type="bibr">20</xref>] integrated YOLOv5 with a Convolutional Block Attention Module (CBAM) for rapeseed inflorescence counting, demonstrating superior performance with an R² of 0.96 compared to other DL models, including YOLOv4, TasselNetV2+, CenterNet, and Faster R-CNN. Despite the longer computation time required by these DL models in contrast to ML and image enhancement methods, they have been proven to be instrumental in achieving superior model performance.</p><p>While prior studies have developed various DL models for plant counting, the robustness of these models depends heavily on the quantity and quality of training data. Currently, the training data have been largely based on the manual annotation of images. Challenges such as limited access to ground truth data, high annotation costs, and uncertainties about data quality used for annotation can impede the implementation of state-of-the-art DL models [<xref rid="B21-sensors-24-06467" ref-type="bibr">21</xref>,<xref rid="B22-sensors-24-06467" ref-type="bibr">22</xref>]. </p><p>Automating the image annotation process to create extensive high-quality training datasets ready for DL models holds tremendous opportunities for advancing predictive and prescriptive analytics in precision agriculture. However, only a few studies have explored semisupervised or semi-automated annotation for agricultural use cases including counting [<xref rid="B23-sensors-24-06467" ref-type="bibr">23</xref>]. Some studies [<xref rid="B24-sensors-24-06467" ref-type="bibr">24</xref>,<xref rid="B25-sensors-24-06467" ref-type="bibr">25</xref>] have employed unsupervised domain adaptation (UDA) to enhance model performance utilizing unlabeled datasets from testing environments. Additionally, Shi et al. (2022) [<xref rid="B26-sensors-24-06467" ref-type="bibr">26</xref>] proposed integrating a background-aware domain adaptation (BADA) module into existing DL networks to reduce background plant counting errors and improve model performance. </p><p>Parallelly, studies have made progress in various computer vision techniques to improve object detection, including crop counting. Recent advancements include those of Bai et al. (2022) [<xref rid="B27-sensors-24-06467" ref-type="bibr">27</xref>], who developed the peak detection algorithm for locating crop rows and plant seedlings of sunflower and maize plants using UAS-based RGB images. This is a swift and robust method, achieving an R<sup>2</sup> of approximately 0.8 for both maize and sunflower. Similarly, Wu et al. (2019) [<xref rid="B28-sensors-24-06467" ref-type="bibr">28</xref>] used thresholding, edge detection, and circular hough transformation algorithms to automatically delineate citrus tree boundaries from UAS multispectral images, demonstrating that around 80% of tree boundaries were accurately delineated and hence demonstrating the effectiveness of digital image processing techniques in plant counting. While these techniques have relatively lower accuracy compared to DL models, they require significantly less computation power. This presents an opportunity to combine computer vision methods, particularly image enhancement techniques, with DL models to improve crop counting accuracy without the need for manually labeled data as well as excessive computing resources. </p><p>The objectives of this study are to (1) automate image annotation using image enhancement techniques and (2) evaluate the performance of state-of-the-art DL models, including CNN- and transformer-based architectures, trained on automatically annotated data for crop counting. This study aimed to streamline DL-based crop counting, improving accuracy while reducing the time needed to generate training data.</p></sec><sec id="sec2-sensors-24-06467" disp-level="1"><title>2. Materials and Methods</title><sec id="sec2dot1-sensors-24-06467" disp-level="2"><title>2.1. Study Area</title><p>This study was conducted based on images collected from a corn field located in the Wooster region, Ohio, USA (<xref rid="sensors-24-06467-f001" ref-type="fig">Figure 1</xref>). Corn seeds were sown in May 2019 with a plant population of 81,510 seeds/ha and were harvested around late October 2019.</p><fig id="sensors-24-06467-f001" position="float"><?disp-level 3?><label>Figure 1</label><caption><p>Study area—corn field located in Wooster, Ohio, USA.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g001.jpg"><?cloudpmc-path blobs/beb8/11479280/98762b874330/sensors-24-06467-g001.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1676?><?original-width 2209?><?scaled-height 558?><?scaled-width 736?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g001.gif"><?cloudpmc-path blobs/beb8/11479280/b2475129ba08/sensors-24-06467-g001.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec><sec id="sec2dot2-sensors-24-06467" disp-level="2"><title>2.2. Data Acquisition and Preprocessing</title><p>Images of the corn field were collected on 11 June 2019, using a DJI Mavic Pro 2 (Shenzhen, China) equipped with an RGB camera during the V2 to V4 growth stages. In general, corn growth stages are divided into Vegetative (V) and Reproductive (R) stages [<xref rid="B29-sensors-24-06467" ref-type="bibr">29</xref>], with V(n) stages, where n represents the number of visible leaf collars. The V2–V4 stage, where corn plants develop 2–4 leaves leaf collars, was chosen because plant canopies typically do not overlap, making it easier to accurately identify corn plants in aerial images [<xref rid="B16-sensors-24-06467" ref-type="bibr">16</xref>]. To avoid any potential variation in color across the collected images, UAS flights were conducted around noon with a clear sky at a height of 26 m. Additionally, in the RGB camera settings, the white balance was set to fixed values.</p><p>Pix4Dcapture (4.4) software was used for flight planning with 85% frontal and 75% side overlap, and ground control points were established before UAS surveys to facilitate image geo-rectification. Individual images collected via UAS were stitched using the Pix4Dmapper (4.3) software and a georeferenced RGB image with a pixel resolution of 1.5 cm × 1.5 cm was created to provide a seamless representation of the entire corn field used in this study (<xref rid="sensors-24-06467-f002" ref-type="fig">Figure 2</xref>). </p><fig id="sensors-24-06467-f002" position="float"><?disp-level 3?><label>Figure 2</label><caption><p>Data acquisition and image preprocessing steps: (<bold>a</bold>) UAS used for data collection, (<bold>b</bold>) flight path planning using the Pix4Dcapture software, and (<bold>c</bold>) steps involved in generating an ortho mosaic RGB map of an entire field.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g002.jpg"><?cloudpmc-path blobs/beb8/11479280/39f205315228/sensors-24-06467-g002.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1061?><?original-width 3530?><?scaled-height 236?><?scaled-width 784?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g002.gif"><?cloudpmc-path blobs/beb8/11479280/e87d0ad25b52/sensors-24-06467-g002.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec><sec id="sec2dot3-sensors-24-06467" disp-level="2"><title>2.3. Plant Counting Framework</title><p>The developed framework for counting plants (<xref rid="sensors-24-06467-f003" ref-type="fig">Figure 3</xref>) is divided into two main steps: (i) the automatic annotation of an RGB image into smaller blocks of corn and non-corn using image enhancement techniques (discussed in <xref rid="sec2dot3dot1-sensors-24-06467" ref-type="sec">Section 2.3.1</xref>) and (ii) the training and testing of four DL models including VGG16 [<xref rid="B30-sensors-24-06467" ref-type="bibr">30</xref>], VGG19 [<xref rid="B31-sensors-24-06467" ref-type="bibr">31</xref>], InceptionV3 [<xref rid="B32-sensors-24-06467" ref-type="bibr">32</xref>], and Vision Transformer (ViT) [<xref rid="B33-sensors-24-06467" ref-type="bibr">33</xref>] based on annotated image blocks to identify corn plants and then estimate the corn stand counts (discussed in <xref rid="sec2dot3dot2-sensors-24-06467" ref-type="sec">Section 2.3.2</xref>). </p><fig id="sensors-24-06467-f003" position="float"><?disp-level 3?><label>Figure 3</label><caption><p>Automatic plant counting framework in two steps. (<bold>i</bold>) Step 1: Tasks for automatically annotating the UAS images for identifying corn plant and soil background images. (<bold>ii</bold>) Step 2: Deep learning models used for identifying corn and non-corn blocks within an image.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g003.jpg"><?cloudpmc-path blobs/beb8/11479280/e9684df75c98/sensors-24-06467-g003.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 3609?><?original-width 2863?><?scaled-height 901?><?scaled-width 715?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g003.gif"><?cloudpmc-path blobs/beb8/11479280/94fc58792a3e/sensors-24-06467-g003.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><sec id="sec2dot3dot1-sensors-24-06467" disp-level="3"><title>2.3.1. Automatic Annotation of Images</title><p>To automate image annotation for developing DL models for corn counting, we first extracted corn row location information that was later used to create smaller annotated image blocks of corn and non-corn. To extract the crop rows’ location, an RGB orthomosaic image was rotated until the rows aligned with a perfectly straight horizontal orientation. The extent to which an image should be rotated was derived by estimating the crop row orientation in the center of the orthomosaic image. After rotation, horizontal strips of 1 m in width, matching the spacing between corn rows at the study site, were created. This ensured that each strip contained only one crop row, enabling the accurate extraction of row locations. A green area mask was then generated using upper and lower thresholds (36, 25, 25) and (70, 255, 255) on the RGB bands to identify corn plants. </p><p>The green mask highlighted the green pixel positions that represented corn plant locations. Then, starting and ending corn pixel locations in each horizontal strip were utilized to fit a linear line and extract the crop row. Along the crop row, blocks of size 0.2 m × 0.25 m were created and checked with a 7% green area criterion to label them as corn plant-representing blocks. The image blocks of 0.2 × 0.25 m were represented by approximately 11 × 11 × 3 pixels, which were resized to match the required input size for each of the selected DL models using the cv2.INTER_AREA function (discussed in detail in <xref rid="sec2dot3dot2-sensors-24-06467" ref-type="sec">Section 2.3.2</xref>). </p><p>This selected block size was found optimal for capturing corn plants at the V2–V4 growth stages (<xref rid="sensors-24-06467-f004" ref-type="fig">Figure 4</xref>). All the other blocks with less than 7% green area along the crop row and outside crop row were labeled as non-corn blocks. This labeling method ensured that areas outside the crop rows, where inter-row weeds might be present, were not mistakenly labeled as corn but as non-corn. The flowchart depicting this process is illustrated in <xref rid="sensors-24-06467-f003" ref-type="fig">Figure 3</xref>. These steps were implemented using functions from the OpenCV and Geopandas library in the Python platform.</p><fig id="sensors-24-06467-f004" position="float"><?disp-level 4?><label>Figure 4</label><caption><p>UAS image of sample area showing annotation block sizes of 0.2 m × 0.25 m.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g004.jpg"><?cloudpmc-path blobs/beb8/11479280/8842f558368a/sensors-24-06467-g004.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1070?><?original-width 2806?><?scaled-height 267?><?scaled-width 701?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g004.gif"><?cloudpmc-path blobs/beb8/11479280/7189889447e4/sensors-24-06467-g004.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec><sec id="sec2dot3dot2-sensors-24-06467" disp-level="3"><title>2.3.2. Deep Learning Models for Plant Counting and Their Architectures</title><p><italic>Training and testing datasets:</italic> Given that our automatic annotation framework generates fixed-size image blocks annotated as corn or non-corn, we primarily focused on image classifier models such as VGG, Inceptionv3, and ViT rather than models that combine classification and detection algorithms like YOLO. YOLO requires bounding boxes around individual corn plants for training, whereas image classifier models can be trained on images without the need for such bounding boxes.</p><p>The annotated datasets were selected from various sections of the field, representing diverse field conditions, such as soil types, elevations, and soil backgrounds. This diverse selection ensures that each annotated image captures distinct color gradients, helping to build a robust generalized model capable of performing well across various test datasets. The selected training and validation datasets covered approximately 70% of the field, with the remaining 30% reserved for testing the DL models. Within this 70% dataset used for model training, a ratio of 70:30 was allocated for model training and validation, respectively. The dataset was further augmented by considering rotation, zoom, and horizontal flips using the ImageDataGenerator function in the Keras Python library for capturing variations in the annotated images. There were a total of 18,004 corn and 46,542 non-corn labeled images, which is a ratio greater than 2-to-1. </p><p><italic>Approach to counter data imbalance for the training of DL models:</italic> Studies have shown that highly imbalanced datasets tend to bias models toward the majority class—in our case, non-corn, resulting in a poor performance for the minority class, i.e., corn [<xref rid="B34-sensors-24-06467" ref-type="bibr">34</xref>,<xref rid="B35-sensors-24-06467" ref-type="bibr">35</xref>]. This occurs because the model has fewer features in the minority class to learn from compared to the majority. To improve classification results for the minority class, we addressed this imbalance by applying random undersampling, which reduced the number of images in the majority class in the training dataset. By selecting a random subset of images from the majority class, we ensured that both classes had an equal number of images for training the models. </p><p><italic>DL model architecture:</italic> Three DL architectures were considered for corn stand counting as discussed below:</p><p><bold>VGG16</bold>: It is an object recognition algorithm renowned for its simplicity and effectiveness. It employs multiple convolution layers with a 3 × 3 filter followed by a max-pooling layer with a 2 × 2 filter. At the end of the architecture, it has three fully connected layers followed by a softmax output. VGG19 has an architecture similar to that of VGG16, except for the addition of 3 convolutional layers added at stages 3, 4, and 5 (<xref rid="sensors-24-06467-f005" ref-type="fig">Figure 5</xref>). The input size for VGG16 and VGG19 is 32 × 32 × 3, representing images with a dimension of 32 × 32 and 3 color RGB bands.</p><fig id="sensors-24-06467-f005" position="float"><?disp-level 4?><label>Figure 5</label><caption><p>Architecture of VGG16 and VGG19 (with two extra convolutions and a ReLU layer marked in green) showing the connection of the convolutional, pooling, dense, and softmax layers.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g005.jpg"><?cloudpmc-path blobs/beb8/11479280/c41f577f4ef0/sensors-24-06467-g005.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 841?><?original-width 4196?><?scaled-height 153?><?scaled-width 762?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g005.gif"><?cloudpmc-path blobs/beb8/11479280/a5ce6b387be3/sensors-24-06467-g005.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p><bold>InceptionV3</bold>: Unlike VGG16, InceptionV3 uses convolutions of varying sizes (1 × 1, 3 × 3, and 5 × 5) within the same module, serving as a multi-level feature extractor from input images, which are then aggregated along the channel dimensions (<xref rid="sensors-24-06467-f006" ref-type="fig">Figure 6</xref>). Additionally, InceptionV3 uses comparatively fewer parameters (weights and biases) than VGG models, which can make it more efficient in terms of compute time and resources. It requires a minimum input size of 75 × 75, so 75 × 75 × 3 was used for its training in this study. </p><fig id="sensors-24-06467-f006" position="float"><?disp-level 4?><label>Figure 6</label><caption><p>Architecture of InceptionV3 showing the connection of the convolutional, pooling, concat, and softmax layers.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g006.jpg"><?cloudpmc-path blobs/beb8/11479280/7785c337e983/sensors-24-06467-g006.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1739?><?original-width 4137?><?scaled-height 316?><?scaled-width 752?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g006.gif"><?cloudpmc-path blobs/beb8/11479280/c4f03d76a3e2/sensors-24-06467-g006.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p><xref rid="sensors-24-06467-t001" ref-type="table">Table 1</xref> outlines the input size, trainable parameters, and computational time for each model. VGG16 and VGG19 were configured with a few non-trainable layers, resulting in 7,116,546 training parameters, which were found sufficient for learning corn and non-corn images. In contrast, Inceptionv3 had all layers set as trainable, updating 21,772,450 parameters. Consequently, InceptionV3 exhibited a longer training time compared to VGG16 and VGG19 due to its higher parameter count. Inceptionv3 was trained with more parameters since fewer parameters resulted in poorer model performance. These models were trained with a batch size of 32 and achieved stable training accuracy around 30 epochs.</p><table-wrap id="sensors-24-06467-t001" position="float"><?disp-level 4?><label>Table 1</label><caption><p>Model specifications and computational time.</p></caption><table frame="hsides" rules="groups"><thead><tr><th align="center" valign="middle" style="border-top:solid thin;border-bottom:solid thin" rowspan="1" colspan="1">Model</th><th align="center" valign="middle" style="border-top:solid thin;border-bottom:solid thin" rowspan="1" colspan="1">Image Input Size</th><th align="center" valign="middle" style="border-top:solid thin;border-bottom:solid thin" rowspan="1" colspan="1">Trained <break/> Parameters</th><th align="center" valign="middle" style="border-top:solid thin;border-bottom:solid thin" rowspan="1" colspan="1">Training Time (Mins)</th></tr></thead><tbody><tr><td align="center" valign="middle" rowspan="1" colspan="1">VGG16</td><td align="center" valign="middle" rowspan="1" colspan="1">32 × 32 × 3</td><td align="center" valign="middle" rowspan="1" colspan="1">7,116,546</td><td align="center" valign="middle" rowspan="1" colspan="1">300</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">VGG19</td><td align="center" valign="middle" rowspan="1" colspan="1">32 × 32 × 3</td><td align="center" valign="middle" rowspan="1" colspan="1">7,116,546</td><td align="center" valign="middle" rowspan="1" colspan="1">320</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">InceptionV3</td><td align="center" valign="middle" rowspan="1" colspan="1">75 × 75 × 3</td><td align="center" valign="middle" rowspan="1" colspan="1">21,772,450</td><td align="center" valign="middle" rowspan="1" colspan="1">750</td></tr><tr><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">ViT</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">32 × 32 × 3</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">275,394</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">200</td></tr></tbody></table></table-wrap><p><bold>ViT model:</bold> In addition to the CNN-based models (Inceptionv3, VGG16, and VGG19), we developed a transformer-based ViT model, which uses an attention mechanism in its encoder unit of the transformer architecture. The model processes annotated images by dividing them into user-defined image patches, which are linearly embedded into one-dimensional vectors. These vectors, combined with image positional information, are passed into the transformer encoder. The encoder’s self-attention mechanism estimates the importance of each embedded image patch, enabling the model to learn long-range dependencies between patches (<xref rid="sensors-24-06467-f007" ref-type="fig">Figure 7</xref>). Despite having fewer training parameters than other models, the ViT model did not achieve significant performance gains with additional parameters. An 8 × 8 image patch size was chosen, as it outperformed 4 × 4 patches. The ViT model achieved stable training accuracy after 30 epochs with a batch size of 8. </p><fig id="sensors-24-06467-f007" position="float"><?disp-level 4?><label>Figure 7</label><caption><p>Architecture of ViT showing the position embedding on image patches before passing to the transformer encoder.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g007.jpg"><?cloudpmc-path blobs/beb8/11479280/bd03fe9b82bf/sensors-24-06467-g007.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1920?><?original-width 3291?><?scaled-height 426?><?scaled-width 731?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g007.gif"><?cloudpmc-path blobs/beb8/11479280/bea4580298e9/sensors-24-06467-g007.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>These models were built in a machine with a configuration of 64 GB of RAM and an Intel(R) Xeon(R) Silver 4114 CPU @ 2.20 GHz processor. The keras and scikit-learn Python libraries were used to implement these DL models.</p><p><italic>Evaluating DL model generalizability for corn stand counting using random seeds:</italic> In any DL and ML model, a random seed is used to initialize the model, leading to different starting states and the random selection of training and testing images. This can lead to varying performances of the same model with different datasets. To assess the generalizability of each DL model for corn stand counting, we trained five versions of each DL model, using different training and testing datasets, each selected based on a unique random seed. The models were then compared in terms of loss, accuracy, R<sup>2</sup>, and RMSE (see <xref rid="sec2dot4-sensors-24-06467" ref-type="sec">Section 2.4</xref>). In total, 20 models (five versions for each of the four DL models) were evaluated using different testing datasets to ensure an unbiased assessment of the models. </p><p>After training the models, their performance was evaluated using test data from five randomly selected regions within the field (<xref rid="sensors-24-06467-f008" ref-type="fig">Figure 8</xref>), ensuring that these regions were distinct from those used in training for an unbiased assessment. Following the model’s prediction, a post-processing step was used, which involved filtering out corn plant locations with model probabilities below 90%, effectively removing low-confidence results. </p><fig id="sensors-24-06467-f008" position="float"><?disp-level 4?><label>Figure 8</label><caption><p>Five test regions were selected to compare the manually counted plant stand counts with the model results.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g008.jpg"><?cloudpmc-path blobs/beb8/11479280/f0f592f48130/sensors-24-06467-g008.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1028?><?original-width 3950?><?scaled-height 206?><?scaled-width 790?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g008.gif"><?cloudpmc-path blobs/beb8/11479280/fd74c7da4d3b/sensors-24-06467-g008.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p><italic>Post-processing of DL-based outputs to prevent double counting</italic>: When the orthomosaic image was automatically divided into smaller image blocks for the training and testing of the models, some blocks contained different portions of the same plant. To prevent the double counting of the same corn plant across multiple images, adjacent image blocks representing the same corn plant were removed using the Intersection over Union (IoU) method. This method eliminates overlapping bounding boxes based on their intersection area. After this post-processing step, corn plants were counted in the five test regions and compared with the manual observations.</p></sec></sec><sec id="sec2dot4-sensors-24-06467" disp-level="2"><title>2.4. Model Performance Evaluation</title><p>The performance of DL models was assessed using precision, recall, and F1 score metrics, which were calculated by comparing the model’s corn stand estimations with manually labeled corn plant stands (Equations (1)–(3)). Higher precision, recall, and F1 score correspond to higher model accuracy. In these equations, True Positive (TP) represents the model’s ability to correctly identify corn plants within the ground reference corn blocks, while False Positive (FP) refers to instances where the model incorrectly predicts a corn plant when it is not actually present. False Negative (FN) indicates the model’s failure to detect an actual corn plant. A confusion matrix showing the TP, FP, TN, and FN are explained in <xref rid="sensors-24-06467-f009" ref-type="fig">Figure 9</xref>.
</p><disp-formula id="FD1-sensors-24-06467"><label>(1)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm1" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mrow><mml:mi>Precision</mml:mi><mml:mo>=</mml:mo></mml:mrow><mml:mstyle scriptlevel="0" displaystyle="true"><mml:mfrac><mml:mrow><mml:mrow><mml:mi>True</mml:mi><mml:mo> </mml:mo><mml:mi>Positive</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi>True</mml:mi><mml:mo> </mml:mo><mml:mi>Positive</mml:mi><mml:mo>+</mml:mo><mml:mi>False</mml:mi><mml:mo> </mml:mo><mml:mi>Positive</mml:mi></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mrow></mml:math></disp-formula><disp-formula id="FD2-sensors-24-06467"><label>(2)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm2" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mi>Recall</mml:mi><mml:mo>=</mml:mo><mml:mstyle scriptlevel="0" displaystyle="true"><mml:mfrac><mml:mrow><mml:mrow><mml:mi>True</mml:mi><mml:mo> </mml:mo><mml:mi>Positive</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi>True</mml:mi><mml:mo> </mml:mo><mml:mi>Positive</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>False</mml:mi><mml:mo> </mml:mo><mml:mi>Negative</mml:mi></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mrow></mml:math></disp-formula><disp-formula id="FD3-sensors-24-06467"><label>(3)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm3" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">F</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>=</mml:mo></mml:mrow><mml:mn>2</mml:mn><mml:mo>×</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mstyle scriptlevel="0" displaystyle="true"><mml:mfrac><mml:mrow><mml:mi>Precision</mml:mi><mml:mo>×</mml:mo><mml:mi>Recall</mml:mi></mml:mrow><mml:mrow><mml:mi>Precision</mml:mi><mml:mo>+</mml:mo><mml:mi>Recall</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula><fig id="sensors-24-06467-f009" position="float"><?disp-level 3?><label>Figure 9</label><caption><p>A confusion matrix for corn plant stands. Positive refers to corn plant stand and negative refers to non-corn.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g009.jpg"><?cloudpmc-path blobs/beb8/11479280/3d602a3d9868/sensors-24-06467-g009.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 873?><?original-width 3119?><?scaled-height 218?><?scaled-width 779?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g009.gif"><?cloudpmc-path blobs/beb8/11479280/435171538bc2/sensors-24-06467-g009.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>After predicting corn stand counts using the model and performing post-processing, the results were counted in the five test regions and compared with manually observed corn stands (referred to as the ground truth) to estimate the coefficient of determination (R<sup>2</sup>) and Root Mean Square Error (RMSE) (Equations (4) and (5)):</p><disp-formula id="FD4-sensors-24-06467"><label>(4)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm4" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">R</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>−</mml:mo><mml:mstyle scriptlevel="0" displaystyle="true"><mml:mfrac><mml:mrow><mml:mrow><mml:munderover><mml:mo stretchy="false">∑</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">n</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">x</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi></mml:mrow></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">p</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:munderover><mml:mo stretchy="false">∑</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">n</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">x</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi></mml:mrow></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">x</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mrow></mml:math></disp-formula><disp-formula id="FD5-sensors-24-06467"><label>(5)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm5" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mi>RMSE</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mstyle scriptlevel="0" displaystyle="true"><mml:mfrac><mml:mrow><mml:mrow><mml:munderover><mml:mo stretchy="false">∑</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">n</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">x</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>-</mml:mo><mml:mo> </mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="normal">p</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:mi mathvariant="normal">n</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:msqrt></mml:mrow></mml:mrow></mml:math></disp-formula><p>
where <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm6" overflow="scroll"><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="normal">x</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:math></inline-formula> represents the manual observation of the corn stand count in the ith sample test region, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm7" overflow="scroll"><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="normal">p</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:math></inline-formula> represents the corn stand count estimated from the DL model at the ith sample test region, <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm8" overflow="scroll"><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="normal">x</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:math></inline-formula> denotes the mean stand count of manual observations, and n represents the number of sample test regions. Higher R<sup>2</sup> and lower RMSE values correspond to a better performance of the DL model. By comparing these values across the four different DL models, we can assess the models’ effectiveness in identifying corn stands in the UAS images of the field.</p></sec></sec><sec id="sec3-sensors-24-06467" disp-level="1"><title>3. Results</title><sec id="sec3dot1-sensors-24-06467" disp-level="2"><title>3.1. Comparative Analysis of Automated and Manual Approaches in Curating Data for the Training and Test Set</title><p>The creation of annotated corn and non-corn image blocks for training DL models relied on the extraction of corn rows from the orthomosaic RGB image. The results of this process are shown in <xref rid="sensors-24-06467-f010" ref-type="fig">Figure 10</xref>. These extracted crop rows were examined and overlaid onto the corresponding UAS image using the ArcGIS software (10.7). This evaluation was conducted to assess the effectiveness of the crop row extraction method, which demonstrated an accuracy of 85–90% in representing crop rows compared to UAS images. The method successfully captured most corn rows, particularly when weeds were absent between corn rows. The accuracy of the method, however, declined in the presence of inter-row weeds due to the similarity in color gradients between weeds and corn plants.</p><fig id="sensors-24-06467-f010" position="float"><?disp-level 3?><label>Figure 10</label><caption><p>Sequence of steps considered in automatically generating corn rows from a field image: (<bold>a</bold>) RGB orthomosaic map representing a small section of the corn field, (<bold>b</bold>) rotated RGB image with rows laid horizontal after finding a rotating angle based on the image orientation, (<bold>c</bold>) image strips considering a 1 m inter-crop row spacing on the rotated image, (<bold>d</bold>) mask representing green pixels in each image strip based on upper and lower thresholds of green pixel values, (<bold>e</bold>) fitted linear lines representing corn rows based on the locations of the green masks, and (<bold>f</bold>) identified corn rows overlaid on RGB image.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g010.jpg"><?cloudpmc-path blobs/beb8/11479280/8f3ca9a73ad3/sensors-24-06467-g010.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 2444?><?original-width 3213?><?scaled-height 543?><?scaled-width 714?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g010.gif"><?cloudpmc-path blobs/beb8/11479280/a7424e9372ee/sensors-24-06467-g010.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>Following the extraction of corn rows, image blocks of 0.2 m × 0.25 m size were created and aligned along the corn rows. The blocks containing a minimum of 7% green pixels of the total image area were then annotated as corn plant blocks (<xref rid="sensors-24-06467-f011" ref-type="fig">Figure 11</xref>). Blocks within the row but with less than 7% of green pixels, along with those outside the crop rows, were annotated as non-corn blocks. This process generated annotated training data of two categories, corn and non-corn, across various soil backgrounds within the field. The accuracy of annotated image blocks was assessed by aligning them with the RGB orthomosaic, mirroring a manual approach, at approximately 20 randomly selected locations. This analysis demonstrated that the proposed method accurately captured corn and non-corn image blocks, identifying 80% of total corn plants. While this approach effectively represented corn plants with a decent accuracy, the annotated images were further used to train four DL models to further enhance the classification accuracy of corn stands.</p><fig id="sensors-24-06467-f011" position="float"><?disp-level 3?><label>Figure 11</label><caption><p>Annotated image blocks of corn and non-corn stands in the field. (<bold>a</bold>) RGB image of a small section of the study site, (<bold>b</bold>) corn rows detected using the crop row detection method, (<bold>c</bold>) extracted corn plant blocks intersecting with corn rows, (<bold>d</bold>) extracted non-corn blocks, and (<bold>e</bold>) filtered corn blocks using the percentage of green pixels.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g011.jpg"><?cloudpmc-path blobs/beb8/11479280/84ef8f626630/sensors-24-06467-g011.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1740?><?original-width 3398?><?scaled-height 387?><?scaled-width 755?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g011.gif"><?cloudpmc-path blobs/beb8/11479280/332b184f4b9e/sensors-24-06467-g011.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec><sec id="sec3dot2-sensors-24-06467" disp-level="2"><title>3.2. Performance of Deep Learning Models for the Classification of Corn Plants</title><p>In building on the evaluation of the generalizability of the DL models for corn stand counting, the training and validation accuracy and loss for the four models across the five randomly selected datasets are shown in <xref rid="sensors-24-06467-f012" ref-type="fig">Figure 12</xref>. The figure highlights the central trends and the variability in performance across different datasets. At epoch 30, the final training accuracy was 0.88, 0.92, 0.91, and 0.93 for Inceptionv3, VGG16, VGG19, and ViT, respectively. Throughout the training phase, all four models exhibited minimal variation in loss and accuracy across all datasets.</p><fig id="sensors-24-06467-f012" position="float"><?disp-level 3?><label>Figure 12</label><caption><p>Distribution of loss and accuracy for four DL models over 0 to 30 epochs during training and validation phases. In each graph, the solid lines represent the average loss and accuracy across five versions of the same DL model, with each version trained and validated on different datasets. The gray-shaded areas around the lines indicate the range, showing the maximum and minimum loss and accuracy values across the five versions at each epoch. # refers to the epoch number.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g012.jpg"><?cloudpmc-path blobs/beb8/11479280/d0e36ecff164/sensors-24-06467-g012.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 5249?><?original-width 3389?><?scaled-height 1166?><?scaled-width 753?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g012.gif"><?cloudpmc-path blobs/beb8/11479280/b0d0e9e27f57/sensors-24-06467-g012.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>During the validation phase, VGG16, VGG19, and ViT showed stable accuracy and loss values across the datasets, with the exception of InceptionV3, which displayed more variability. This suggests that the random sampling of training and validation data can significantly influence model learning, particularly for Inceptionv3. Some of this could be attributed to its model architecture involving varying filter sizes and convolution layers [<xref rid="B35-sensors-24-06467" ref-type="bibr">35</xref>]. In contrast, VGG16, VGG19, and ViT consistently exhibited stability across four of the five datasets, showcasing their robust performance. </p><p>Although Inceptionv3 demonstrated higher accuracy when trained with more epochs (50, 100, and 150), its validation accuracy was still unstable, with accuracy ranging between 0.5 and 0.94. In contrast, ViT, VGG16, and VGG19 demonstrated robust performance across different random datasets, with VGG16 consistently outperforming VGG19 and ViT in both training and validation. This stability, combined with lower computational requirements, led to halting the training of Inceptionv3 with additional epochs. Furthermore, increasing the number of epochs and hyperparameter tuning had minimal impact on model performance. </p><p>To further assess the performance of DL models, the trained models were tested on randomly selected regions within the field. Each test image block, which represents a portion of the field, was processed by the trained models to identify blocks containing corn plants (<xref rid="sensors-24-06467-f013" ref-type="fig">Figure 13</xref>). Although these blocks were non-overlapping and might cover only a part of the corn stands, the models were still able to predict the presence of corn plants. Post-processing was applied to remove duplicate predictions based on adjacent overlap information, and the final corn plant blocks were compared with manual visual observations to evaluate the accuracy of the models’ predictions. </p><fig id="sensors-24-06467-f013" position="float"><?disp-level 3?><label>Figure 13</label><caption><p>The models’ performance in detecting corn plant blocks within a sample test region. The colored blocks overlaid on the UAS images represent the corn plant blocks identified by the DL models. Each colored block represents the predicted corn plant blocks identified by a different model.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g013.jpg"><?cloudpmc-path blobs/beb8/11479280/f387ab0d7dd0/sensors-24-06467-g013.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 2849?><?original-width 2071?><?scaled-height 949?><?scaled-width 690?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g013.gif"><?cloudpmc-path blobs/beb8/11479280/4a984f03a726/sensors-24-06467-g013.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>The model performance was further evaluated using metrics, such as accuracy, precision, recall, and F1 score, as depicted in <xref rid="sensors-24-06467-f014" ref-type="fig">Figure 14</xref>. Comparative analyses of these metrics revealed that VGG16 and VGG19 outperformed InceptionV3 and ViT. VGG16, in particular, achieved the highest F1 score of 0.955 for the dataset generated with random seed five. Although all models demonstrated good precision, indicating high accuracy in positive predictions, InceptionV3 had lower recall values, suggesting a higher number of false negatives. In contrast, ViT consistently had high recall. The F1 scores further highlighted the stable performance of VGG16 and VGG19, with the highest values being 0.955 and 0.939, respectively. ViT also performed well, achieving an F1 score of 0.935, and showing stability across various dataset except the fourth. Among the four models, VGG16 exhibited consistently stable and high scores across all performance metrics. </p><fig id="sensors-24-06467-f014" position="float"><?disp-level 3?><label>Figure 14</label><caption><p>Comparison of accuracy, precision, recall, and F1 score of Inceptionv3, VGG16, VGG19, and ViT models on the five random test regions in the corn field. The small black circles represent the outliers, indicating the extremely low and high values observed across five versions of each model.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g014.jpg"><?cloudpmc-path blobs/beb8/11479280/898e11977b86/sensors-24-06467-g014.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 2462?><?original-width 3291?><?scaled-height 547?><?scaled-width 731?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g014.gif"><?cloudpmc-path blobs/beb8/11479280/0a5ba151dc82/sensors-24-06467-g014.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec><sec id="sec3dot3-sensors-24-06467" disp-level="2"><title>3.3. Model Predicted vs. Manually Counted Corn Stands</title><p>The models’ estimated corn plant stands and the manually counted corn plant stands in the five test regions are shown in <xref rid="sensors-24-06467-t002" ref-type="table">Table 2</xref>. The model performance was evaluated using the R<sup>2</sup> and RMSE, with variations across different datasets illustrated in <xref rid="sensors-24-06467-f015" ref-type="fig">Figure 15</xref>. Detailed R<sup>2</sup> and RMSE values for all versions of DL models can be found in <xref rid="app1-sensors-24-06467" ref-type="sec">Table S1</xref>. Notably, Inceptionv3 exhibited a superior performance with the second dataset (i.e., random seed two), but its R<sup>2</sup> and RMSE values varied widely with other datasets, reflecting poor consistency in detecting corn blocks. In contrast, VGG16, VGG19, and ViT (except for only one out of five datasets) outperformed Inceptionv3, with higher R<sup>2</sup> and lower RMSE values. VGG16 and VGG19 consistently exhibited a narrow range of R<sup>2</sup> and RMSE values across five datasets. In particular, VGG16, trained with the second dataset achieved a high R<sup>2</sup> of 0.94 and a lower RMSE of 9.95. On average, Inceptionv3, VGG16, VGG19, and ViT had mean R<sup>2</sup> values of 0.55, 0.93, 0.84, and 0.73, respectively, with corresponding mean RMSE values of 29, 10.86, 16.45, and 20.48. Overall, VGG16 demonstrated high accuracy with minimal variation in model performance across different random datasets, while ViT showed stable performance except for one dataset. For detailed scatterplots comparing model-predicted corn stand counts with the manual counts across all model versions, please refer to <xref rid="app1-sensors-24-06467" ref-type="sec">Figure S1 in the Supplementary Materials</xref>.</p><table-wrap id="sensors-24-06467-t002" position="float"><?disp-level 3?><label>Table 2</label><caption><p>Total corn stand counts at five test regions (TRs) after the training of four deep learning models.</p></caption><table frame="hsides" rules="groups"><thead><tr><th align="center" valign="middle" style="border-top:solid thin;border-bottom:solid thin" rowspan="1" colspan="1">
</th><th align="center" valign="middle" style="border-top:solid thin;border-bottom:solid thin" rowspan="1" colspan="1">
</th><th align="center" valign="middle" style="border-top:solid thin;border-bottom:solid thin" rowspan="1" colspan="1">TR1</th><th align="center" valign="middle" style="border-top:solid thin;border-bottom:solid thin" rowspan="1" colspan="1">TR2</th><th align="center" valign="middle" style="border-top:solid thin;border-bottom:solid thin" rowspan="1" colspan="1">TR3</th><th align="center" valign="middle" style="border-top:solid thin;border-bottom:solid thin" rowspan="1" colspan="1">TR4</th><th align="center" valign="middle" style="border-top:solid thin;border-bottom:solid thin" rowspan="1" colspan="1">TR5</th></tr><tr><th align="left" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">Ground Truth</th><th align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">
</th><th align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">229</th><th align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">274</th><th align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">184</th><th align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">116</th><th align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">140</th></tr><tr><th align="left" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">Models</th><th colspan="2" align="center" valign="middle" style="border-bottom:solid thin" rowspan="1">Version</th><th colspan="4" align="center" valign="middle" style="border-bottom:solid thin" rowspan="1">Estimated Corn Stand Counts </th></tr></thead><tbody><tr><td rowspan="5" align="left" valign="middle" style="border-bottom:solid thin" colspan="1">
<bold>Inceptionv3</bold>
</td><td align="center" valign="middle" rowspan="1" colspan="1">1</td><td align="center" valign="middle" rowspan="1" colspan="1">205</td><td align="center" valign="middle" rowspan="1" colspan="1">246</td><td align="center" valign="middle" rowspan="1" colspan="1">162</td><td align="center" valign="middle" rowspan="1" colspan="1">99</td><td align="center" valign="middle" rowspan="1" colspan="1">121</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">2</td><td align="center" valign="middle" rowspan="1" colspan="1">198</td><td align="center" valign="middle" rowspan="1" colspan="1">236</td><td align="center" valign="middle" rowspan="1" colspan="1">176</td><td align="center" valign="middle" rowspan="1" colspan="1">100</td><td align="center" valign="middle" rowspan="1" colspan="1">116</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">3</td><td align="center" valign="middle" rowspan="1" colspan="1">203</td><td align="center" valign="middle" rowspan="1" colspan="1">255</td><td align="center" valign="middle" rowspan="1" colspan="1">170</td><td align="center" valign="middle" rowspan="1" colspan="1">103</td><td align="center" valign="middle" rowspan="1" colspan="1">133</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">4</td><td align="center" valign="middle" rowspan="1" colspan="1">194</td><td align="center" valign="middle" rowspan="1" colspan="1">203</td><td align="center" valign="middle" rowspan="1" colspan="1">155</td><td align="center" valign="middle" rowspan="1" colspan="1">92</td><td align="center" valign="middle" rowspan="1" colspan="1">94</td></tr><tr><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">5</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">192</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">233</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">150</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">96</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">97</td></tr><tr><td rowspan="5" align="left" valign="middle" colspan="1">
<bold>VGG16</bold>
</td><td align="center" valign="middle" rowspan="1" colspan="1">1</td><td align="center" valign="middle" rowspan="1" colspan="1">209</td><td align="center" valign="middle" rowspan="1" colspan="1">263</td><td align="center" valign="middle" rowspan="1" colspan="1">177</td><td align="center" valign="middle" rowspan="1" colspan="1">111</td><td align="center" valign="middle" rowspan="1" colspan="1">135</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">2</td><td align="center" valign="middle" rowspan="1" colspan="1">211</td><td align="center" valign="middle" rowspan="1" colspan="1">263</td><td align="center" valign="middle" rowspan="1" colspan="1">177</td><td align="center" valign="middle" rowspan="1" colspan="1">111</td><td align="center" valign="middle" rowspan="1" colspan="1">137</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">3</td><td align="center" valign="middle" rowspan="1" colspan="1">210</td><td align="center" valign="middle" rowspan="1" colspan="1">259</td><td align="center" valign="middle" rowspan="1" colspan="1">181</td><td align="center" valign="middle" rowspan="1" colspan="1">111</td><td align="center" valign="middle" rowspan="1" colspan="1">136</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">4</td><td align="center" valign="middle" rowspan="1" colspan="1">209</td><td align="center" valign="middle" rowspan="1" colspan="1">261</td><td align="center" valign="middle" rowspan="1" colspan="1">177</td><td align="center" valign="middle" rowspan="1" colspan="1">110</td><td align="center" valign="middle" rowspan="1" colspan="1">135</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">5</td><td align="center" valign="middle" rowspan="1" colspan="1">210</td><td align="center" valign="middle" rowspan="1" colspan="1">266</td><td align="center" valign="middle" rowspan="1" colspan="1">178</td><td align="center" valign="middle" rowspan="1" colspan="1">111</td><td align="center" valign="middle" rowspan="1" colspan="1">137</td></tr><tr><td rowspan="5" align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin" colspan="1">
<bold>VGG19</bold>
</td><td align="center" valign="middle" style="border-top:solid thin" rowspan="1" colspan="1">1</td><td align="center" valign="middle" style="border-top:solid thin" rowspan="1" colspan="1">202</td><td align="center" valign="middle" style="border-top:solid thin" rowspan="1" colspan="1">255</td><td align="center" valign="middle" style="border-top:solid thin" rowspan="1" colspan="1">177</td><td align="center" valign="middle" style="border-top:solid thin" rowspan="1" colspan="1">111</td><td align="center" valign="middle" style="border-top:solid thin" rowspan="1" colspan="1">132</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">2</td><td align="center" valign="middle" rowspan="1" colspan="1">200</td><td align="center" valign="middle" rowspan="1" colspan="1">252</td><td align="center" valign="middle" rowspan="1" colspan="1">178</td><td align="center" valign="middle" rowspan="1" colspan="1">109</td><td align="center" valign="middle" rowspan="1" colspan="1">135</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">3</td><td align="center" valign="middle" rowspan="1" colspan="1">203</td><td align="center" valign="middle" rowspan="1" colspan="1">251</td><td align="center" valign="middle" rowspan="1" colspan="1">177</td><td align="center" valign="middle" rowspan="1" colspan="1">107</td><td align="center" valign="middle" rowspan="1" colspan="1">135</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">4</td><td align="center" valign="middle" rowspan="1" colspan="1">207</td><td align="center" valign="middle" rowspan="1" colspan="1">249</td><td align="center" valign="middle" rowspan="1" colspan="1">176</td><td align="center" valign="middle" rowspan="1" colspan="1">109</td><td align="center" valign="middle" rowspan="1" colspan="1">135</td></tr><tr><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">5</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">202</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">251</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">173</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">108</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">132</td></tr><tr><td rowspan="5" align="left" valign="middle" style="border-bottom:solid thin" colspan="1">
<bold>ViT</bold>
</td><td align="center" valign="middle" rowspan="1" colspan="1">1</td><td align="center" valign="middle" rowspan="1" colspan="1">206</td><td align="center" valign="middle" rowspan="1" colspan="1">259</td><td align="center" valign="middle" rowspan="1" colspan="1">180</td><td align="center" valign="middle" rowspan="1" colspan="1">108</td><td align="center" valign="middle" rowspan="1" colspan="1">133</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">2</td><td align="center" valign="middle" rowspan="1" colspan="1">210</td><td align="center" valign="middle" rowspan="1" colspan="1">255</td><td align="center" valign="middle" rowspan="1" colspan="1">179</td><td align="center" valign="middle" rowspan="1" colspan="1">110</td><td align="center" valign="middle" rowspan="1" colspan="1">134</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">3</td><td align="center" valign="middle" rowspan="1" colspan="1">204</td><td align="center" valign="middle" rowspan="1" colspan="1">260</td><td align="center" valign="middle" rowspan="1" colspan="1">179</td><td align="center" valign="middle" rowspan="1" colspan="1">108</td><td align="center" valign="middle" rowspan="1" colspan="1">135</td></tr><tr><td align="center" valign="middle" rowspan="1" colspan="1">4</td><td align="center" valign="middle" rowspan="1" colspan="1">146</td><td align="center" valign="middle" rowspan="1" colspan="1">235</td><td align="center" valign="middle" rowspan="1" colspan="1">133</td><td align="center" valign="middle" rowspan="1" colspan="1">87</td><td align="center" valign="middle" rowspan="1" colspan="1">120</td></tr><tr><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">5</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">206</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">260</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">180</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">107</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">133</td></tr></tbody></table><table-wrap-foot><fn id="fn2"><p>Note: There were five versions for each DL model, each trained and validated with different training and validation datasets, selected based on five unique seeds.</p></fn></table-wrap-foot></table-wrap><fig id="sensors-24-06467-f015" position="float"><?disp-level 3?><label>Figure 15</label><caption><p>Performance of Inceptionv3, VGG16, VGG19, and ViT models based on R<sup>2</sup> (<bold>a</bold>) and RMSE (<bold>b</bold>) on five random test regions in the corn field. The orange line in the boxplot represents the median model performance value.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="sensors-24-06467-g015.jpg"><?cloudpmc-path blobs/beb8/11479280/55b7c1f44da0/sensors-24-06467-g015.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1466?><?original-width 3259?><?scaled-height 326?><?scaled-width 724?></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="sensors-24-06467-g015.gif"><?cloudpmc-path blobs/beb8/11479280/429ef63a257c/sensors-24-06467-g015.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec></sec><sec id="sec4-sensors-24-06467" disp-level="1"><title>4. Discussion</title><sec id="sec4dot1-sensors-24-06467" disp-level="2"><title>4.1. Automating High-Quality Annotated Training Data for Corn Counting</title><p>The effectiveness of DL models in crop identification hinges largely on the quality and quantity of the training data. Previous studies have often developed DL models using manually annotated images, which can be labor intensive. For instance, Wu et al. (2019) [<xref rid="B28-sensors-24-06467" ref-type="bibr">28</xref>] estimated a resource requirement of approximately 123 person-hours for annotating images to develop a rice seedling count model. In their study, rice seedlings were manually labeled as points (instead of blocks) in 40 high-resolution images, with seedlings counts varying from 3732 to 16,173. In a separate study [<xref rid="B17-sensors-24-06467" ref-type="bibr">17</xref>], researchers manually labeled only 864 images using Labelme software for training and testing the YOLOv3 model in detecting corn seedlings. This limited dataset poses a potential challenge, as the implemented model may underperform in scenarios where the data lack diversity. </p><p>This study explored automating the generation of high-quality annotated training datasets by leveraging crop row information extracted through image morphological approaches [<xref rid="B27-sensors-24-06467" ref-type="bibr">27</xref>,<xref rid="B36-sensors-24-06467" ref-type="bibr">36</xref>,<xref rid="B37-sensors-24-06467" ref-type="bibr">37</xref>], with the goal of developing DL models trained on these annotated images for accurate corn counting. Unlike previous studies that relied on limited training datasets, our study introduced a workflow that automatically generates 18,004 corn images and 46,542 non-corn images for a 1.5-acre corn field, which were then used to train the corn counting model. Once the annotation method was implemented, extracting and annotating these image blocks took 4–5 h, including the time to generate orthomosaic image, identify crop rows, and refine the labeled images. The primary benefit of this process is its future scalability; the framework can be applied to new fields to quickly generate annotated datasets. With minor adjustments, we believe that this approach can help significantly streamline the process of generating annotated images and improve crop count estimates, thereby reducing significant time and labor costs. </p></sec><sec id="sec4dot2-sensors-24-06467" disp-level="2"><title>4.2. Deep Learning Models for Counting Plant Stands</title><p>In this study, DL models trained on automatically annotated images achieved an F1 score of up to 0.955, an R<sup>2</sup> of up to 0.94, and an RMSE as low as 9.95 in detecting corn plant stands. The performance of these DL models was comparable to the results of previous studies based on manually annotated images [<xref rid="B17-sensors-24-06467" ref-type="bibr">17</xref>,<xref rid="B38-sensors-24-06467" ref-type="bibr">38</xref>,<xref rid="B39-sensors-24-06467" ref-type="bibr">39</xref>]. For example, a recent study [<xref rid="B39-sensors-24-06467" ref-type="bibr">39</xref>] using YOLOv5, YOLOv7, and CenterNet models trained on manually annotated images (using LabelImg tool) achieved an F1 score between 0.90 and 0.95 for detecting cotton seedlings. This highlights the efficacy of our approach in generating annotated images for training DL models. </p></sec><sec id="sec4dot3-sensors-24-06467" disp-level="2"><title>4.3. Differences in Performances with Other Models </title><p>In this study, VGG16 demonstrated the best performance in corn counting, followed by ViT, VGG19, and InceptionV3. However, it is essential to note that ViT was trained with the fewest number of parameters compared to the rest of the models. Hence, if time is a big constraint in building a DL model for corn counting, ViT can be a viable alternative without sacrificing much in performance. It is also worth pointing out that there are other DL architectures, such as YOLO, which could provide a very precise estimate of plant bounding boxes compared to our model-predicted corn plant blocks. Future studies could explore these architectures for further improvement.</p><p>Apart from DL analyses, object-based image analysis (OBIA) has also been used to count crop plants. Koh et al. (2019) [<xref rid="B40-sensors-24-06467" ref-type="bibr">40</xref>] employed template matching for detecting safflower seedlings at early growth stages, achieving an R<sup>2</sup> of ~0.87 and an RSME of ~10. Template matching is an image processing technique wherein a template image is utilized to match smaller segments within a larger image. This study acquired data at a resolution of 0.19 cm/pixel and automated the OBIA algorithm using the proprietary eCognition (9.3) software. The results obtained in our study surpassed these metrics using an open-source Python platform, offering ease of development and deployment to diverse row crop fields.</p></sec></sec><sec id="sec5-sensors-24-06467" disp-level="1"><title>5. Limitations and Future Works</title><p>In this study, the automated image annotation framework was developed and tested for corn stand counting. However, the developed framework can easily be adapted to other crops with minimal parameter adjustments, such as crop row orientation, typical crop coverage in a given growth stage to determine the size of image blocks, and the percent threshold of green pixels within a crop block, to improve labeling. It is also important to note that the image annotation framework assumes that the crop rows are straight rather than curved, as corn fields in the U.S. are typically planted using GPS-based guidance systems [<xref rid="B41-sensors-24-06467" ref-type="bibr">41</xref>], which helps maintain straight lines, improve planting accuracy, minimize crop overlap, and eventually ensure consistent application of inputs and fertilizers throughout the growing season. The field used in this study is representative of many U.S. corn fields and features predominantly straight rows, with the exception of the edges, which are primarily used for turning agricultural machinery. Hence, we used a linear method for row detection. However, this approach may struggle to identify curved rows, potentially reducing the accuracy of corn counting models trained on data annotated using curved rows’ information. Future work could focus on developing methods capable of detecting curved rows, thus enhancing the robustness of crop counting models across various planting scenarios.</p><p>Similarly, the effectiveness of our approach in corn row identification and stand counts may be affected by factors such as the presence of doubles and seed spacing accuracy. Our row crop identification method relies on consistent inter-row spacing and performs best when doubles are absent, and when spacing between corn stands is uniform. While we acknowledge that these issues can influence the accuracy of our method, most row crop fields in the U.S. are planted using modern precision technology designed to minimize doubles through more accurate seed placement. The corn field used for our study was planted using an RTK GPS-based precision planter, and there was no obvious presence of doubles. Nevertheless, future research could focus on developing corn counting models that maintain high accuracy even in the presence of doubles and inconsistent seed spacing. </p><p>Here, our DL models for corn stand counting were trained exclusively on images from a single corn field in V2–V4 growth stages. The robustness of these models can be improved by incorporating images from a range of field sites, environmental conditions, crop growth stages, and management practices. This helps to build a larger, diverse database, which will ultimately enhance model performance and the transferability of a model to new fields. </p></sec><sec id="sec6-sensors-24-06467" disp-level="1"><title>6. Conclusions</title><p>In this study, we developed an automated image annotation framework that utilizes image enhancement techniques to annotate the images, which were used for the training of four DL models, including InceptionV3, VGG16, VGG19, and ViT for the accurate detection and counting of corn stands. While Inceptionv3 exhibited relatively unstable performance and its performance was sensitive to the random selection of training and validation datasets, VGG16, VGG19, and ViT demonstrated more stable performance, indicating their better adaptability to varying training datasets. Notably, VGG16 outperformed the other three models, achieving an F1 score of 0.955, an R<sup>2</sup> of 0.94, and an RMSE of 9.95 when compared to corn stand counts in the test regions. ViT provided the second-best performance, with an R<sup>2</sup> of 0.90 across four out of five training datasets. Overall, the developed automated labeling framework for training the DL model, especially the VGG16 model, demonstrated promising potential for accurate and efficient corn stand detection. This approach paves the way for automating the generation of training data, contributing to the development of a more robust and effective corn plant identification model.</p></sec><sec id="ack1" sec-type="ack" disp-level="1"><title>Acknowledgments</title><p>We would like to thank Kushal KC for collecting the UAS data used in the study and Nuo Xu for preliminary analyses of the images.</p></sec><sec id="app1-sensors-24-06467" sec-type="app" disp-level="1"><title>Supplementary Materials</title><p>The following Supporting Information can be downloaded at <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.mdpi.com/article/10.3390/s24196467/s1" ext-link-type="uri">https://www.mdpi.com/article/10.3390/s24196467/s1</ext-link>: Table S1: R2 and RMSE values observed for all DL model versions; Figure S1: Comparison of corn stand count with the manual count for five test regions observed for each DL model.</p><supplementary-material id="sensors-24-06467-s001" position="float"><media xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="sensors-24-06467-s001.zip" mimetype="application" mime-subtype="zip"><?cloudpmc-path beb8/11479280/55e88c94cddb/sensors-24-06467-s001.zip?><?cloudpmc-bucket app?><?size 528846?></media></supplementary-material></sec><sec id="notes1" disp-level="1"><title>Author Contributions</title><p>Conceptualization, S.K. (Sami Khanal); methodology, S.K. (Sushma Katari), S.V., C.S. and S.K. (Sami Khanal); software, S.K. (Sushma Katari) and S.V.; validation, S.K. (Sushma Katari); formal analysis, S.K. (Sushma Katari); investigation, S.K. (Sushma Katari), S.V. and S.K. (Sami Khanal); resources, S.K. (Sami Khanal); data curation, S.V.; writing—original draft preparation, S.K. (Sushma Katari) and S.K. (Sami Khanal); writing—review and editing, S.V., C.S. and S.K. (Sami Khanal); visualization, S.K. (Sushma Katari); supervision, S.K. (Sami Khanal) and C.S.; project administration, S.K. (Sami Khanal); funding acquisition, S.K. (Sami Khanal). All authors have read and agreed to the published version of the manuscript.</p></sec><sec id="notes2" disp-level="1"><title>Institutional Review Board Statement</title><p>Not applicable. </p></sec><sec id="notes3" disp-level="1"><title>Informed Consent Statement</title><p>Not applicable.</p></sec><sec id="notes4" disp-level="1"><title>Data Availability Statement</title><p>Data used in this study can be made available upon request.</p></sec><sec id="notes5" disp-level="1"><title>Conflicts of Interest</title><p>Author Sandeep Venkatesh was employed by Google. Other authors declare no conflicts of interest.</p></sec><sec id="funding-statement1" xml:lang="en" disp-level="1"><title>Funding Statement</title><p>This research was funded in part by the NRT EmPowerment Fellowship and Ohio Agricultural Research and Development Center (OARDC) Internal Grant Program # 2022-017.</p></sec><sec id="fn-group1" sec-type="fn-group" disp-level="1"><title>Footnotes</title><fn-group><fn id="fn1"><p><bold>Disclaimer/Publisher’s Note:</bold> The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.</p></fn></fn-group></sec><sec id="ref-list1" sec-type="ref-list" disp-level="1"><title>References</title><sec id="ref-list1_sec2" disp-level="2"><ref-list><ref id="B1-sensors-24-06467"><label>1.</label><mixed-citation><named-content content-type="citation-string">Perfect E., McLaughlin N.B. Soil Management Effects on Planting and Emergence of No-till Corn. Trans. ASAE. 1996;39:1611–1615. doi: 10.13031/2013.27676.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.13031/2013.27676"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Trans. ASAE&amp;title=Soil Management Effects on Planting and Emergence of No-till Corn&amp;author=E. Perfect&amp;author=N.B. McLaughlin&amp;volume=39&amp;publication_year=1996&amp;pages=1611-1615&amp;doi=10.13031/2013.27676&amp;"/></mixed-citation></ref><ref id="B2-sensors-24-06467"><label>2.</label><mixed-citation><named-content content-type="citation-string">Poncet A.M., Fulton J.P., McDonald T.P., Knappenberger T., Shaw J.N. Corn Emergence and Yield Response to Row-Unit Depth and Downforce for Varying Field Conditions. Appl. Eng. Agric. 2019;35:399–408. doi: 10.13031/aea.12408.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.13031/aea.12408"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Appl. Eng. Agric.&amp;title=Corn Emergence and Yield Response to Row-Unit Depth and Downforce for Varying Field Conditions&amp;author=A.M. Poncet&amp;author=J.P. Fulton&amp;author=T.P. McDonald&amp;author=T. Knappenberger&amp;author=J.N. Shaw&amp;volume=35&amp;publication_year=2019&amp;pages=399-408&amp;doi=10.13031/aea.12408&amp;"/></mixed-citation></ref><ref id="B3-sensors-24-06467"><label>3.</label><mixed-citation><named-content content-type="citation-string">Mohammadi G.R., Koohi Y., Ghobadi M., Najaphy A. Effects of Seed Priming, Planting Density and Row Spacing on Seedling Emergence and Some Phenological Indices of Corn (Zea mays L.) Philipp. Agric. Sci. 2014;97:300–306.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Philipp. Agric. Sci.&amp;title=Effects of Seed Priming, Planting Density and Row Spacing on Seedling Emergence and Some Phenological Indices of Corn (Zea mays L.)&amp;author=G.R. Mohammadi&amp;author=Y. Koohi&amp;author=M. Ghobadi&amp;author=A. Najaphy&amp;volume=97&amp;publication_year=2014&amp;pages=300-306&amp;"/></mixed-citation></ref><ref id="B4-sensors-24-06467"><label>4.</label><mixed-citation><named-content content-type="citation-string">Lawles K., Raun W., Desta K., Freeman K. Effect of Delayed Emergence on Corn Grain Yields. J. Plant Nutr. 2012;35:480–496. doi: 10.1080/01904167.2012.639926.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1080/01904167.2012.639926"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Plant Nutr.&amp;title=Effect of Delayed Emergence on Corn Grain Yields&amp;author=K. Lawles&amp;author=W. Raun&amp;author=K. Desta&amp;author=K. Freeman&amp;volume=35&amp;publication_year=2012&amp;pages=480-496&amp;doi=10.1080/01904167.2012.639926&amp;"/></mixed-citation></ref><ref id="B5-sensors-24-06467"><label>5.</label><mixed-citation><named-content content-type="citation-string">Baghdadi A., Halim R.A., Majidian M., Daud W.N.W., Ahmad I. Plant Density and Tillage Effects on Forage Corn Quality. J. Food Agric. Environ. 2012;10:366–370.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Food Agric. Environ.&amp;title=Plant Density and Tillage Effects on Forage Corn Quality&amp;author=A. Baghdadi&amp;author=R.A. Halim&amp;author=M. Majidian&amp;author=W.N.W. Daud&amp;author=I. Ahmad&amp;volume=10&amp;publication_year=2012&amp;pages=366-370&amp;"/></mixed-citation></ref><ref id="B6-sensors-24-06467"><label>6.</label><mixed-citation><named-content content-type="citation-string">Yang T., Zhu S., Zhang W., Zhao Y., Song X., Yang G., Yao Z., Wu W., Liu T., Sun C., et al.  Unmanned Aerial Vehicle-Scale Weed Segmentation Method Based on Image Analysis Technology for Enhanced Accuracy of Maize Seedling Counting. Agriculture. 2024;14:175.  doi: 10.3390/agriculture14020175.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/agriculture14020175"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Agriculture&amp;title=Unmanned Aerial Vehicle-Scale Weed Segmentation Method Based on Image Analysis Technology for Enhanced Accuracy of Maize Seedling Counting&amp;author=T. Yang&amp;author=S. Zhu&amp;author=W. Zhang&amp;author=Y. Zhao&amp;author=X. Song&amp;volume=14&amp;publication_year=2024&amp;pages=175&amp;doi=10.3390/agriculture14020175&amp;"/></mixed-citation></ref><ref id="B7-sensors-24-06467"><label>7.</label><mixed-citation><named-content content-type="citation-string">Duffy J.P., Anderson K., Fawcett D., Curtis R.J., Maclean I.M.D. Drones Provide Spatial and Volumetric Data to Deliver New Insights into Microclimate Modelling. Landsc. Ecol. 2021;36:685–702. doi: 10.1007/s10980-020-01180-9.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1007/s10980-020-01180-9"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Landsc. Ecol.&amp;title=Drones Provide Spatial and Volumetric Data to Deliver New Insights into Microclimate Modelling&amp;author=J.P. Duffy&amp;author=K. Anderson&amp;author=D. Fawcett&amp;author=R.J. Curtis&amp;author=I.M.D. Maclean&amp;volume=36&amp;publication_year=2021&amp;pages=685-702&amp;doi=10.1007/s10980-020-01180-9&amp;"/></mixed-citation></ref><ref id="B8-sensors-24-06467"><label>8.</label><mixed-citation><named-content content-type="citation-string">Pathak H., Igathinathane C., Zhang Z., Archer D., Hendrickson J. A Review of Unmanned Aerial Vehicle-Based Methods for Plant Stand Count Evaluation in Row Crops. Comput. Electron. Agric. 2022;198:107064. doi: 10.1016/j.compag.2022.107064.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.compag.2022.107064"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Comput. Electron. Agric.&amp;title=A Review of Unmanned Aerial Vehicle-Based Methods for Plant Stand Count Evaluation in Row Crops&amp;author=H. Pathak&amp;author=C. Igathinathane&amp;author=Z. Zhang&amp;author=D. Archer&amp;author=J. Hendrickson&amp;volume=198&amp;publication_year=2022&amp;pages=107064&amp;doi=10.1016/j.compag.2022.107064&amp;"/></mixed-citation></ref><ref id="B9-sensors-24-06467"><label>9.</label><mixed-citation><named-content content-type="citation-string">Donmez C., Villi O., Berberoglu S., Cilek A. Computer Vision-Based Citrus Tree Detection in a Cultivated Environment Using UAV Imagery. Comput. Electron. Agric. 2021;187:106273. doi: 10.1016/j.compag.2021.106273.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.compag.2021.106273"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Comput. Electron. Agric.&amp;title=Computer Vision-Based Citrus Tree Detection in a Cultivated Environment Using UAV Imagery&amp;author=C. Donmez&amp;author=O. Villi&amp;author=S. Berberoglu&amp;author=A. Cilek&amp;volume=187&amp;publication_year=2021&amp;pages=106273&amp;doi=10.1016/j.compag.2021.106273&amp;"/></mixed-citation></ref><ref id="B10-sensors-24-06467"><label>10.</label><mixed-citation><named-content content-type="citation-string">Xia L., Zhang R., Chen L., Huang Y., Xu G., Wen Y., Yi T. Monitor Cotton Budding Using SVM and UAV Images. Appl. Sci. 2019;9:4312.  doi: 10.3390/app9204312.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/app9204312"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Appl. Sci.&amp;title=Monitor Cotton Budding Using SVM and UAV Images&amp;author=L. Xia&amp;author=R. Zhang&amp;author=L. Chen&amp;author=Y. Huang&amp;author=G. Xu&amp;volume=9&amp;publication_year=2019&amp;pages=4312&amp;doi=10.3390/app9204312&amp;"/></mixed-citation></ref><ref id="B11-sensors-24-06467"><label>11.</label><mixed-citation><named-content content-type="citation-string">Banerjee B.P., Sharma V., Spangenberg G., Kant S. Machine Learning Regression Analysis for Estimation of Crop Emergence Using Multispectral UAV Imagery. Remote Sens. 2021;13:2918.  doi: 10.3390/rs13152918.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/rs13152918"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Remote Sens.&amp;title=Machine Learning Regression Analysis for Estimation of Crop Emergence Using Multispectral UAV Imagery&amp;author=B.P. Banerjee&amp;author=V. Sharma&amp;author=G. Spangenberg&amp;author=S. Kant&amp;volume=13&amp;publication_year=2021&amp;pages=2918&amp;doi=10.3390/rs13152918&amp;"/></mixed-citation></ref><ref id="B12-sensors-24-06467"><label>12.</label><mixed-citation><named-content content-type="citation-string">Tavus M.R., Eker M.E., Senyer N., Karabulut B.  Proceedings of the 2015 23rd Signal Processing and Communications Applications Conference (SIU) IEEE; New York, NY, USA: 2015. Plant Counting By Using k-NN Classification on UAVs Images; pp. 1058–1061.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="title=Proceedings of the 2015 23rd Signal Processing and Communications Applications Conference (SIU)&amp;author=M.R. Tavus&amp;author=M.E. Eker&amp;author=N. Senyer&amp;author=B. Karabulut&amp;publication_year=2015&amp;"/></mixed-citation></ref><ref id="B13-sensors-24-06467"><label>13.</label><mixed-citation><named-content content-type="citation-string">Osco L.P., de Arruda M.d.S., Goncalves D.N., Dias A., Batistoti J., de Souza M., Georges Gomes F.D., Marques Ramos A.P., de Castro Jorge L.A., Liesenberg V., et al.  A CNN Approach to Simultaneously Count Plants and Detect Plantation-Rows from UAV Imagery. ISPRS J. Photogramm. Remote Sens. 2021;174:1–17. doi: 10.1016/j.isprsjprs.2021.01.024.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.isprsjprs.2021.01.024"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=ISPRS J. Photogramm. Remote Sens.&amp;title=A CNN Approach to Simultaneously Count Plants and Detect Plantation-Rows from UAV Imagery&amp;author=L.P. Osco&amp;author=M.d.S. de Arruda&amp;author=D.N. Goncalves&amp;author=A. Dias&amp;author=J. Batistoti&amp;volume=174&amp;publication_year=2021&amp;pages=1-17&amp;doi=10.1016/j.isprsjprs.2021.01.024&amp;"/></mixed-citation></ref><ref id="B14-sensors-24-06467"><label>14.</label><mixed-citation><named-content content-type="citation-string">Zhang J., Zhao B., Yang C., Shi Y., Liao Q., Zhou G., Wang C., Xie T., Jiang Z., Zhang D., et al.  Rapeseed Stand Count Estimation at Leaf Development Stages With UAV Imagery and Convolutional Neural Networks. Front. Plant Sci. 2020;11:617.  doi: 10.3389/fpls.2020.00617.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3389/fpls.2020.00617"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7298076"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="32587594"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Front. Plant Sci.&amp;title=Rapeseed Stand Count Estimation at Leaf Development Stages With UAV Imagery and Convolutional Neural Networks&amp;author=J. Zhang&amp;author=B. Zhao&amp;author=C. Yang&amp;author=Y. Shi&amp;author=Q. Liao&amp;volume=11&amp;publication_year=2020&amp;pages=617&amp;pmid=32587594&amp;doi=10.3389/fpls.2020.00617&amp;"/></mixed-citation></ref><ref id="B15-sensors-24-06467"><label>15.</label><mixed-citation><named-content content-type="citation-string">Zhang C., Atkinson P.M., George C., Wen Z., Diazgranados M., Gerard F. Identifying and Mapping Individual Plants in a Highly Diverse High-Elevation Ecosystem Using UAV Imagery and Deep Learning. ISPRS J. Photogramm. Remote Sens. 2020;169:280–291. doi: 10.1016/j.isprsjprs.2020.09.025.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.isprsjprs.2020.09.025"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=ISPRS J. Photogramm. Remote Sens.&amp;title=Identifying and Mapping Individual Plants in a Highly Diverse High-Elevation Ecosystem Using UAV Imagery and Deep Learning&amp;author=C. Zhang&amp;author=P.M. Atkinson&amp;author=C. George&amp;author=Z. Wen&amp;author=M. Diazgranados&amp;volume=169&amp;publication_year=2020&amp;pages=280-291&amp;doi=10.1016/j.isprsjprs.2020.09.025&amp;"/></mixed-citation></ref><ref id="B16-sensors-24-06467"><label>16.</label><mixed-citation><named-content content-type="citation-string">Vong C.N., Conway L.S., Zhou J., Kitchen N.R., Sudduth K.A. Early Corn Stand Count of Different Cropping Systems Using UAV-Imagery and Deep Learning. Comput. Electron. Agric. 2021;186:106214. doi: 10.1016/j.compag.2021.106214.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.compag.2021.106214"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Comput. Electron. Agric.&amp;title=Early Corn Stand Count of Different Cropping Systems Using UAV-Imagery and Deep Learning&amp;author=C.N. Vong&amp;author=L.S. Conway&amp;author=J. Zhou&amp;author=N.R. Kitchen&amp;author=K.A. Sudduth&amp;volume=186&amp;publication_year=2021&amp;pages=106214&amp;doi=10.1016/j.compag.2021.106214&amp;"/></mixed-citation></ref><ref id="B17-sensors-24-06467"><label>17.</label><mixed-citation><named-content content-type="citation-string">Wang L., Xiang L., Tang L., Jiang H. A Convolutional Neural Network-Based Method for Corn Stand Counting in the Field. Sensors. 2021;21:507.  doi: 10.3390/s21020507.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/s21020507"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7828297"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33450839"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Sensors&amp;title=A Convolutional Neural Network-Based Method for Corn Stand Counting in the Field&amp;author=L. Wang&amp;author=L. Xiang&amp;author=L. Tang&amp;author=H. Jiang&amp;volume=21&amp;publication_year=2021&amp;pages=507&amp;pmid=33450839&amp;doi=10.3390/s21020507&amp;"/></mixed-citation></ref><ref id="B18-sensors-24-06467"><label>18.</label><mixed-citation><named-content content-type="citation-string">Machefer M., Lemarchand F., Bonnefond V., Hitchins A., Sidiropoulos P. Mask R-CNN Refitting Strategy for Plant Counting and Sizing in UAV Imagery. Remote Sens. 2020;12:3015.  doi: 10.3390/rs12183015.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/rs12183015"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Remote Sens.&amp;title=Mask R-CNN Refitting Strategy for Plant Counting and Sizing in UAV Imagery&amp;author=M. Machefer&amp;author=F. Lemarchand&amp;author=V. Bonnefond&amp;author=A. Hitchins&amp;author=P. Sidiropoulos&amp;volume=12&amp;publication_year=2020&amp;pages=3015&amp;doi=10.3390/rs12183015&amp;"/></mixed-citation></ref><ref id="B19-sensors-24-06467"><label>19.</label><mixed-citation><named-content content-type="citation-string">Lu H., Cao Z. TasselNetV2+: A Fast Implementation for High-Throughput Plant Counting From High-Resolution RGB Imagery. Front. Plant Sci. 2020;11:541960.  doi: 10.3389/fpls.2020.541960.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3389/fpls.2020.541960"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7750361"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33365037"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Front. Plant Sci.&amp;title=TasselNetV2+: A Fast Implementation for High-Throughput Plant Counting From High-Resolution RGB Imagery&amp;author=H. Lu&amp;author=Z. Cao&amp;volume=11&amp;publication_year=2020&amp;pages=541960&amp;pmid=33365037&amp;doi=10.3389/fpls.2020.541960&amp;"/></mixed-citation></ref><ref id="B20-sensors-24-06467"><label>20.</label><mixed-citation><named-content content-type="citation-string">Li H., Wang P., Huang C. Comparison of Deep Learning Methods for Detecting and Counting Sorghum Heads in UAV Imagery. Remote Sens. 2022;14:3143.  doi: 10.3390/rs14133143.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/rs14133143"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Remote Sens.&amp;title=Comparison of Deep Learning Methods for Detecting and Counting Sorghum Heads in UAV Imagery&amp;author=H. Li&amp;author=P. Wang&amp;author=C. Huang&amp;volume=14&amp;publication_year=2022&amp;pages=3143&amp;doi=10.3390/rs14133143&amp;"/></mixed-citation></ref><ref id="B21-sensors-24-06467"><label>21.</label><mixed-citation><named-content content-type="citation-string">Robey A., Hassani H., Pappas G.J. Model-Based Robust Deep Learning: Generalizing to Natural, Out-of-Distribution Data. arXiv. 20202005.10247</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=arXiv&amp;title=Model-Based Robust Deep Learning: Generalizing to Natural, Out-of-Distribution Data&amp;author=A. Robey&amp;author=H. Hassani&amp;author=G.J. Pappas&amp;publication_year=2020&amp;"/></mixed-citation></ref><ref id="B22-sensors-24-06467"><label>22.</label><mixed-citation><named-content content-type="citation-string">Wang Y., Bansal M. Robust Machine Comprehension Models via Adversarial Training. arXiv. 20181804.06473</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=arXiv&amp;title=Robust Machine Comprehension Models via Adversarial Training&amp;author=Y. Wang&amp;author=M. Bansal&amp;publication_year=2018&amp;"/></mixed-citation></ref><ref id="B23-sensors-24-06467"><label>23.</label><mixed-citation><named-content content-type="citation-string">Alnaasan N., Lieber M., Shafi A., Subramoni H., Shearer S., Panda D.K. HARVEST: High-Performance Artificial Vision Framework for Expert Labeling Using Semi-Supervised Training; Proceedings of the 2023 IEEE International Conference on Big Data (BigData); Sorrento, Italy. 15–18 December 2023; pp. 139–148.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Proceedings of the 2023 IEEE International Conference on Big Data (BigData)&amp;title=HARVEST: High-Performance Artificial Vision Framework for Expert Labeling Using Semi-Supervised Training&amp;author=N. Alnaasan&amp;author=M. Lieber&amp;author=A. Shafi&amp;author=H. Subramoni&amp;author=S. Shearer&amp;pages=139-148&amp;"/></mixed-citation></ref><ref id="B24-sensors-24-06467"><label>24.</label><mixed-citation><named-content content-type="citation-string">Mei J., Sun K. Unsupervised Adversarial Domain Adaptation Leaf Counting with Bayesian Loss Density Estimation. Signal Image Video Process. 2023;17:1503–1509. doi: 10.1007/s11760-022-02359-0.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1007/s11760-022-02359-0"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Signal Image Video Process.&amp;title=Unsupervised Adversarial Domain Adaptation Leaf Counting with Bayesian Loss Density Estimation&amp;author=J. Mei&amp;author=K. Sun&amp;volume=17&amp;publication_year=2023&amp;pages=1503-1509&amp;doi=10.1007/s11760-022-02359-0&amp;"/></mixed-citation></ref><ref id="B25-sensors-24-06467"><label>25.</label><mixed-citation><named-content content-type="citation-string">Rodriguez-Vazquez J., Fernandez-Cortizas M., Perez-Saura D., Molina M., Campoy P. Overcoming Domain Shift in Neural Networks for Accurate Plant Counting in Aerial Images. Remote Sens. 2023;15:1700.  doi: 10.3390/rs15061700.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/rs15061700"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Remote Sens.&amp;title=Overcoming Domain Shift in Neural Networks for Accurate Plant Counting in Aerial Images&amp;author=J. Rodriguez-Vazquez&amp;author=M. Fernandez-Cortizas&amp;author=D. Perez-Saura&amp;author=M. Molina&amp;author=P. Campoy&amp;volume=15&amp;publication_year=2023&amp;pages=1700&amp;doi=10.3390/rs15061700&amp;"/></mixed-citation></ref><ref id="B26-sensors-24-06467"><label>26.</label><mixed-citation><named-content content-type="citation-string">Shi M., Li X.-Y., Lu H., Cao Z.-G. Background-Aware Domain Adaptation for Plant Counting. Front. Plant Sci. 2022;13:731816.  doi: 10.3389/fpls.2022.731816.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3389/fpls.2022.731816"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC8850787"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="35185973"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Front. Plant Sci.&amp;title=Background-Aware Domain Adaptation for Plant Counting&amp;author=M. Shi&amp;author=X.-Y. Li&amp;author=H. Lu&amp;author=Z.-G. Cao&amp;volume=13&amp;publication_year=2022&amp;pages=731816&amp;pmid=35185973&amp;doi=10.3389/fpls.2022.731816&amp;"/></mixed-citation></ref><ref id="B27-sensors-24-06467"><label>27.</label><mixed-citation><named-content content-type="citation-string">Bai Y., Nie C., Wang H., Cheng M., Liu S., Yu X., Shao M., Wang Z., Wang S., Tuohuti N., et al.  A Fast and Robust Method for Plant Count in Sunflower and Maize at Different Seedling Stages Using High-Resolution UAV RGB Imagery. Precis. Agric. 2022;23:1720–1742. doi: 10.1007/s11119-022-09907-1.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1007/s11119-022-09907-1"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Precis. Agric.&amp;title=A Fast and Robust Method for Plant Count in Sunflower and Maize at Different Seedling Stages Using High-Resolution UAV RGB Imagery&amp;author=Y. Bai&amp;author=C. Nie&amp;author=H. Wang&amp;author=M. Cheng&amp;author=S. Liu&amp;volume=23&amp;publication_year=2022&amp;pages=1720-1742&amp;doi=10.1007/s11119-022-09907-1&amp;"/></mixed-citation></ref><ref id="B28-sensors-24-06467"><label>28.</label><mixed-citation><named-content content-type="citation-string">Wu J., Yang G., Yang X., Xu B., Han L., Zhu Y. Automatic Counting of in Situ Rice Seedlings from UAV Images Based on a Deep Fully Convolutional Neural Network. Remote Sens. 2019;11:691.  doi: 10.3390/rs11060691.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/rs11060691"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Remote Sens.&amp;title=Automatic Counting of in Situ Rice Seedlings from UAV Images Based on a Deep Fully Convolutional Neural Network&amp;author=J. Wu&amp;author=G. Yang&amp;author=X. Yang&amp;author=B. Xu&amp;author=L. Han&amp;volume=11&amp;publication_year=2019&amp;pages=691&amp;doi=10.3390/rs11060691&amp;"/></mixed-citation></ref><ref id="B29-sensors-24-06467"><label>29.</label><mixed-citation><named-content content-type="citation-string">Corn Growth and Development: Crop Staging|Agronomic Crops Network.  [(accessed on 5 September 2024)].  Available online:  <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://agcrops.osu.edu/newsletter/corn-newsletter/2022-18/corn-growth-and-development-crop-staging" ext-link-type="uri">https://agcrops.osu.edu/newsletter/corn-newsletter/2022-18/corn-growth-and-development-crop-staging</ext-link>.</named-content></mixed-citation></ref><ref id="B30-sensors-24-06467"><label>30.</label><mixed-citation><named-content content-type="citation-string">Simonyan K., Zisserman A. Very Deep Convolutional Networks for Large-Scale Image Recognition; Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015); San Diego, CA, USA. 7–9 May 2015.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015)&amp;title=Very Deep Convolutional Networks for Large-Scale Image Recognition&amp;author=K. Simonyan&amp;author=A. Zisserman&amp;"/></mixed-citation></ref><ref id="B31-sensors-24-06467"><label>31.</label><mixed-citation><named-content content-type="citation-string">He K., Zhang X., Ren S., Sun J. Deep Residual Learning for Image Recognition; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; Boston, MA, USA. 7–12 June 2015.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition&amp;title=Deep Residual Learning for Image Recognition&amp;author=K. He&amp;author=X. Zhang&amp;author=S. Ren&amp;author=J. Sun&amp;"/></mixed-citation></ref><ref id="B32-sensors-24-06467"><label>32.</label><mixed-citation><named-content content-type="citation-string">Szegedy C., Liu W., Jia Y., Sermanet P., Reed S., Anguelov D., Erhan D., Vanhoucke V., Rabinovich A. Going Deeper with Convolutions; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; Boston, MA, USA. 7–12 June 2015.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition&amp;title=Going Deeper with Convolutions&amp;author=C. Szegedy&amp;author=W. Liu&amp;author=Y. Jia&amp;author=P. Sermanet&amp;author=S. Reed&amp;"/></mixed-citation></ref><ref id="B33-sensors-24-06467"><label>33.</label><mixed-citation><named-content content-type="citation-string">Dosovitskiy A., Beyer L., Kolesnikov A., Weissenborn D., Zhai X., Unterthiner T., Dehghani M., Minderer M., Heigold G., Gelly S., et al.  An Image Is Worth 16 × 16 Words: Transformers for Image Recognition at Scale. arXiv. 20212010.11929</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=arXiv&amp;title=An Image Is Worth 16 × 16 Words: Transformers for Image Recognition at Scale&amp;author=A. Dosovitskiy&amp;author=L. Beyer&amp;author=A. Kolesnikov&amp;author=D. Weissenborn&amp;author=X. Zhai&amp;publication_year=2021&amp;"/></mixed-citation></ref><ref id="B34-sensors-24-06467"><label>34.</label><mixed-citation><named-content content-type="citation-string">Johnson J.M., Khoshgoftaar T.M. Survey on Deep Learning with Class Imbalance. J. Big Data. 2019;6:27. doi: 10.1186/s40537-019-0192-5.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1186/s40537-019-0192-5"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Big Data&amp;title=Survey on Deep Learning with Class Imbalance&amp;author=J.M. Johnson&amp;author=T.M. Khoshgoftaar&amp;volume=6&amp;publication_year=2019&amp;pages=27&amp;doi=10.1186/s40537-019-0192-5&amp;"/></mixed-citation></ref><ref id="B35-sensors-24-06467"><label>35.</label><mixed-citation><named-content content-type="citation-string">Ahmed W., Karim A. The Impact of Filter Size and Number of Filters on Classification Accuracy in CNN; Proceedings of the 2020 International Conference on Computer Science and Software Engineering (CSASE); Duhok, Iraq. 16–18 April 2020; pp. 88–93.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Proceedings of the 2020 International Conference on Computer Science and Software Engineering (CSASE)&amp;title=The Impact of Filter Size and Number of Filters on Classification Accuracy in CNN&amp;author=W. Ahmed&amp;author=A. Karim&amp;pages=88-93&amp;"/></mixed-citation></ref><ref id="B36-sensors-24-06467"><label>36.</label><mixed-citation><named-content content-type="citation-string">Chen P., Ma X., Wang F., Li J. A New Method for Crop Row Detection Using Unmanned Aerial Vehicle Images. Remote Sens. 2021;13:3526.  doi: 10.3390/rs13173526.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/rs13173526"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Remote Sens.&amp;title=A New Method for Crop Row Detection Using Unmanned Aerial Vehicle Images&amp;author=P. Chen&amp;author=X. Ma&amp;author=F. Wang&amp;author=J. Li&amp;volume=13&amp;publication_year=2021&amp;pages=3526&amp;doi=10.3390/rs13173526&amp;"/></mixed-citation></ref><ref id="B37-sensors-24-06467"><label>37.</label><mixed-citation><named-content content-type="citation-string">Al Mansoori S., Kunhu A., Al Ahmad H.  Automatic Palm Trees Detection from Multispectral UAV Data Using Normalized Difference Vegetation Index and Circular Hough Transform. In: Huang B., Lopez S., Wu Z., editors. Proceedings of the High-Performance Computing in Geoscience and Remote Sensing Viii. Volume 10792. SPIE—The International Society for Optical Engineering; Bellingham, WA, USA: 2018. p. 1079203.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="title=Proceedings of the High-Performance Computing in Geoscience and Remote Sensing Viii&amp;author=S. Al Mansoori&amp;author=A. Kunhu&amp;author=H. Al Ahmad&amp;publication_year=2018&amp;"/></mixed-citation></ref><ref id="B38-sensors-24-06467"><label>38.</label><mixed-citation><named-content content-type="citation-string">Liu X., Ghazali K.H., Han F., Mohamed I.I. Automatic Detection of Oil Palm Tree from UAV Images Based on the Deep Learning Method. Appl. Artif. Intell. 2021;35:13–24. doi: 10.1080/08839514.2020.1831226.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1080/08839514.2020.1831226"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Appl. Artif. Intell.&amp;title=Automatic Detection of Oil Palm Tree from UAV Images Based on the Deep Learning Method&amp;author=X. Liu&amp;author=K.H. Ghazali&amp;author=F. Han&amp;author=I.I. Mohamed&amp;volume=35&amp;publication_year=2021&amp;pages=13-24&amp;doi=10.1080/08839514.2020.1831226&amp;"/></mixed-citation></ref><ref id="B39-sensors-24-06467"><label>39.</label><mixed-citation><named-content content-type="citation-string">Feng Y., Chen W., Ma Y., Zhang Z., Gao P., Lv X. Cotton Seedling Detection and Counting Based on UAV Multispectral Images and Deep Learning Methods. Remote Sens. 2023;15:2680.  doi: 10.3390/rs15102680.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/rs15102680"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Remote Sens.&amp;title=Cotton Seedling Detection and Counting Based on UAV Multispectral Images and Deep Learning Methods&amp;author=Y. Feng&amp;author=W. Chen&amp;author=Y. Ma&amp;author=Z. Zhang&amp;author=P. Gao&amp;volume=15&amp;publication_year=2023&amp;pages=2680&amp;doi=10.3390/rs15102680&amp;"/></mixed-citation></ref><ref id="B40-sensors-24-06467"><label>40.</label><mixed-citation><named-content content-type="citation-string">Koh J.C.O., Hayden M., Daetwyler H., Kant S. Estimation of Crop Plant Density at Early Mixed Growth Stages Using UAV Imagery. Plant Methods. 2019;15:64. doi: 10.1186/s13007-019-0449-1.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1186/s13007-019-0449-1"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC6584986"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="31249606"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Plant Methods&amp;title=Estimation of Crop Plant Density at Early Mixed Growth Stages Using UAV Imagery&amp;author=J.C.O. Koh&amp;author=M. Hayden&amp;author=H. Daetwyler&amp;author=S. Kant&amp;volume=15&amp;publication_year=2019&amp;pages=64&amp;pmid=31249606&amp;doi=10.1186/s13007-019-0449-1&amp;"/></mixed-citation></ref><ref id="B41-sensors-24-06467"><label>41.</label><mixed-citation><named-content content-type="citation-string">Poncet A., Fulton J., Port K., McDonald T., Pate G. Optimizing Field Traffic Patterns to Improve Machinery Efficiency: Path Planning Using Guidance Lines.  [(accessed on 19 September 2024)].  Available online:  <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://ohioline.osu.edu/factsheet/fabe-5531" ext-link-type="uri">https://ohioline.osu.edu/factsheet/fabe-5531</ext-link>.</named-content></mixed-citation></ref></ref-list></sec></sec><sec id="_ad93_" xml:lang="en" sec-type="associated-data" disp-level="1"><title>Associated Data</title><sec id="_adsm93_" xml:lang="en" sec-type="supplementary-materials" disp-level="2"><title>Supplementary Materials</title><supplementary-material id="db_ds_supplementary-material1_reqid_" position="float"><media xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="sensors-24-06467-s001.zip" mimetype="application" mime-subtype="zip"><?cloudpmc-path beb8/11479280/55e88c94cddb/sensors-24-06467-s001.zip?><?cloudpmc-bucket app?><?size 528846?></media></supplementary-material></sec><sec id="_adda93_" xml:lang="en" sec-type="data-availability-statement" disp-level="2"><title>Data Availability Statement</title><p>Data used in this study can be made available upon request.</p></sec></sec></body></article>