
<!DOCTYPE article
  PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with MathML3 v1.4 20241031//EN" "JATS-archivearticle1-4-mathml3.dtd">
<article xml:lang="en" article-type="research-article" dtd-version="1.4"><processing-meta base-tagset="archiving" mathml-version="3.0" table-model="xhtml" tagset-family="jats"><restricted-by>pmc</restricted-by></processing-meta><front><journal-meta><journal-id journal-id-type="nlm-ta">Sensors (Basel)</journal-id><journal-id journal-id-type="iso-abbrev">Sensors (Basel)</journal-id><journal-id journal-id-type="pmc-domain-id">1660</journal-id><journal-id journal-id-type="pmc-domain">sensors</journal-id><journal-id journal-id-type="nlm-id">101204366</journal-id><journal-id journal-id-type="publisher-id">sensors</journal-id><journal-title-group><journal-title>Sensors (Basel, Switzerland)</journal-title></journal-title-group><issn pub-type="epub">1424-8220</issn><?publisher_abbrev mdpi?><publisher><publisher-name>Multidisciplinary Digital Publishing Institute  (MDPI)</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="pmcid">PMC8586972</article-id><article-id pub-id-type="pmcid-ver">PMC8586972.1</article-id><article-id pub-id-type="pmcaid">8586972</article-id><article-id pub-id-type="pmcaiid">8586972</article-id><article-id pub-id-type="pmid">34770306</article-id><article-id pub-id-type="doi">10.3390/s21216999</article-id><article-id pub-id-type="publisher-id">sensors-21-06999</article-id><article-version article-version-type="pmc-version">1</article-version><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title-group><article-title>Central Object Segmentation by Deep Learning to Continuously Monitor Fruit Growth through RGB Images</article-title></title-group><contrib-group><contrib contrib-type="author"><contrib-id contrib-id-type="orcid" authenticated="true">https://orcid.org/0000-0002-6757-1700</contrib-id><name name-style="western"><surname>Fukuda</surname><given-names initials="M">Motohisa</given-names></name><xref rid="af1-sensors-21-06999" ref-type="aff">1</xref><xref rid="c1-sensors-21-06999" ref-type="corresp">*</xref></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid" authenticated="true">https://orcid.org/0000-0002-2002-4662</contrib-id><name name-style="western"><surname>Okuno</surname><given-names initials="T">Takashi</given-names></name><xref rid="af1-sensors-21-06999" ref-type="aff">1</xref></contrib><contrib contrib-type="author"><name name-style="western"><surname>Yuki</surname><given-names initials="S">Shinya</given-names></name><xref rid="af2-sensors-21-06999" ref-type="aff">2</xref></contrib></contrib-group><contrib-group><contrib contrib-type="editor"><name name-style="western"><surname>Moshou</surname><given-names initials="D">Dimitrios</given-names></name><role>Academic Editor</role></contrib></contrib-group><aff id="af1-sensors-21-06999"><label>1</label>Faculty of Science, Yamagata University, 1-4-12 Kojirakawa, Yamagata 990-8560, Japan; <email>okuno@sci.kj.yamagata-u.ac.jp</email></aff><aff id="af2-sensors-21-06999"><label>2</label>Elix Inc., Daini Togo Park Building 3F, 8-34 Yonbancho, Chiyoda-ku, Tokyo 102-0081, Japan; <email>shinya.yuki@elix-inc.com</email></aff><author-notes><corresp id="c1-sensors-21-06999"><label>*</label>Correspondence: <email>fukuda@sci.kj.yamagata-u.ac.jp</email></corresp></author-notes><pub-date pub-type="epub"><day>21</day><month>10</month><year>2021</year></pub-date><pub-date pub-type="collection"><month>11</month><year>2021</year></pub-date><volume>21</volume><issue>21</issue><issue-id pub-id-type="pmc-issue-id">393688</issue-id><elocation-id>6999</elocation-id><history><date date-type="received"><day>12</day><month>9</month><year>2021</year></date><date date-type="accepted"><day>18</day><month>10</month><year>2021</year></date></history><pub-history><event event-type="pmc-release"><date><day>21</day><month>10</month><year>2021</year></date></event><event event-type="pmc-live"><date><day>13</day><month>11</month><year>2021</year></date></event><event event-type="pmc-last-change"><date iso-8601-date="2022-01-04 16:16:40.007"><day>04</day><month>01</month><year>2022</year></date></event></pub-history><permissions><copyright-statement>© 2021 by the authors.</copyright-statement><copyright-year>2021</copyright-year><license><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/" specific-use="textmining" content-type="ccbylicense">https://creativecommons.org/licenses/by/4.0/</ali:license_ref><license-p>Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (<ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link>).</license-p></license></permissions><self-uri xmlns:xlink="http://www.w3.org/1999/xlink" content-type="pmc-pdf" xlink:href="sensors-21-06999.pdf"><?pdf-name sensors-21-06999.pdf?><?pdf-size 10136770?><?pdf-md5 cc08560b47674b7a5f1fcc908bb9ded2?><?pdf-image-server-status NEVER_LOAD?><?pdf-cloudpmc-urn urn:app:0b7e/8586972/cc08560b4767/sensors-21-06999.pdf?></self-uri><abstract><p>Monitoring fruit growth is useful when estimating final yields in advance and predicting optimum harvest times. However, observing fruit all day at the farm via RGB images is not an easy task because the light conditions are constantly changing. In this paper, we present CROP (Central Roundish Object Painter). The method involves image segmentation by deep learning, and the architecture of the neural network is a deeper version of U-Net. CROP identifies different types of central roundish fruit in an RGB image in varied light conditions, and creates a corresponding mask. Counting the mask pixels gives the relative two-dimensional size of the fruit, and in this way, time-series images may provide a non-contact means of automatically monitoring fruit growth. Although our measurement unit is different from the traditional one (length), we believe that shape identification potentially provides more information. Interestingly, CROP can have a more general use, working even for some other roundish objects. For this reason, we hope that CROP and our methodology yield big data to promote scientific advancements in horticultural science and other fields.</p></abstract><kwd-group><kwd>deep learning</kwd><kwd>U-Net</kwd><kwd>image segmentation</kwd><kwd>central object</kwd><kwd>fruit</kwd><kwd>pear</kwd><kwd>growth monitor</kwd><kwd>RGB images</kwd></kwd-group><custom-meta-group><custom-meta><meta-name>pmc-status-qastatus</meta-name><meta-value>0</meta-value></custom-meta><custom-meta><meta-name>pmc-status-live</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-status-embargo</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-status-released</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-open-access</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-olf</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-manuscript</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-legally-suppressed</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-has-pdf</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-has-supplement</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-pdf-only</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-suppress-copyright</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-is-real-version</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-is-scanned-article</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-preprint</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-in-epmc</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-license-ref</meta-name><meta-value>CC BY</meta-value></custom-meta></custom-meta-group></article-meta></front><body><sec sec-type="intro" id="sec1-sensors-21-06999"><title>1. Introduction</title><p>The following <xref rid="sensors-21-06999-f001" ref-type="fig">Figure 1</xref> gives a quick overview of <monospace>CROP</monospace> the neural network to be introduced in this paper. Counting the pixels belonging to the masks, which are represented by red color, gives the relative sizes of the fruit. Applying this method to time-series images captured by fixed cameras, one can keep track of the fruit growth.</p><p>Please note that USDA ARS stands for United States Department of Agriculture, Agricultural Research Service, and the photos were obtained from their image gallery. Use of photos of USDA ARS is not meant to infer or imply USDA ARS endorsement of any product, company, or position. The original photos were cropped and processed.</p><sec id="sec1dot1-sensors-21-06999"><title>1.1. Monitoring Fruit Growth</title><p>Techniques of computer vision have many applications in fruit production, for example, making yield estimates by detecting fruit in the farms [<xref rid="B1-sensors-21-06999" ref-type="bibr">1</xref>]. Non-contact size measurements and shape description are also useful in making yield estimates and distribution plans of fruit [<xref rid="B2-sensors-21-06999" ref-type="bibr">2</xref>,<xref rid="B3-sensors-21-06999" ref-type="bibr">3</xref>,<xref rid="B4-sensors-21-06999" ref-type="bibr">4</xref>]. Indeed, in [<xref rid="B5-sensors-21-06999" ref-type="bibr">5</xref>] using Bertalanffy growth model, predictability of fruit size at the time of harvest based on the measurements at the earlier development stages was discussed. Furthermore, when change of colors cannot be a measure of ripeness, size and volume growth can be good candidates for indexes regarding such predictions; see <xref rid="sec1dot4-sensors-21-06999" ref-type="sec">Section 1.4</xref> for the pear production in Japan. In fact, devices to continuously measure fruit size have been developed, for example stainless frames with potentiometers to be put on fruit [<xref rid="B6-sensors-21-06999" ref-type="bibr">6</xref>], and flexible tapes around fruit to be read by infrared reflex sensors [<xref rid="B7-sensors-21-06999" ref-type="bibr">7</xref>].</p><p>Recently deep learning techniques are applied to monitor growth of mushrooms [<xref rid="B8-sensors-21-06999" ref-type="bibr">8</xref>] and apples [<xref rid="B9-sensors-21-06999" ref-type="bibr">9</xref>]. These two results share the same spirit with ours in the sense that deep neural networks are used to process images captured by fixed cameras to acquire the time-series size data. The former was conducted in the greenhouse with the light source at the top and the images were captured from fixed cameras above, which probably offered stable light conditions. These images were then processed by the object-detection deep neural network <monospace>YOLOv3</monospace> [<xref rid="B10-sensors-21-06999" ref-type="bibr">10</xref>] to locate mushroom caps, and next by a so-called SP algorithm to calculate the diameters. In the latter, the neural network was built upon <monospace>ResNet-50</monospace> [<xref rid="B11-sensors-21-06999" ref-type="bibr">11</xref>] a popular image-classification deep neural network. Then, the images of apples in the orchard captured on cloudy days or at dusk were processed by a Laplace operator and then by the deep neural network, to detect the edges to predict the diameters.</p></sec><sec id="sec1dot2-sensors-21-06999"><title>1.2. Image Segmentation and Our Model</title><p>Initially, computer vision methods for fruit yield estimate are based on pixel-wise color analysis, for example, in [<xref rid="B12-sensors-21-06999" ref-type="bibr">12</xref>] fruit pixels were counted to obtain the occupancy ratios in the images using the thresholds on red, green and blue. Later, techniques of so-called <italic toggle="yes">image segmentation</italic> started being applied; image segmentation means classifying image pixels into segments. With such techniques, one can isolate fruit pixels in images and count the number of blobs for the yield estimates. Such examples include [<xref rid="B13-sensors-21-06999" ref-type="bibr">13</xref>,<xref rid="B14-sensors-21-06999" ref-type="bibr">14</xref>], and [<xref rid="B15-sensors-21-06999" ref-type="bibr">15</xref>] (with the controlled illumination at night). Additionally, techniques of <italic toggle="yes">machine learning</italic> were applied, for example <italic toggle="yes">X-means clustering</italic> in [<xref rid="B16-sensors-21-06999" ref-type="bibr">16</xref>] and <italic toggle="yes">conditional random field</italic> for multi-spectral data in [<xref rid="B17-sensors-21-06999" ref-type="bibr">17</xref>].</p><p>Moreover, image segmentation techniques have also been applied to non-contact size and volume measurement of fruit, or more broadly shape description. Such examples include: via Hough transform [<xref rid="B18-sensors-21-06999" ref-type="bibr">18</xref>], using two cameras [<xref rid="B19-sensors-21-06999" ref-type="bibr">19</xref>], and on smart phones [<xref rid="B20-sensors-21-06999" ref-type="bibr">20</xref>]. In [<xref rid="B21-sensors-21-06999" ref-type="bibr">21</xref>], a machine learning method called <italic toggle="yes">support vector machine</italic> was applied.</p><p>Many of deep learning methods applied in horticultural science fall in the category of object detection, with which one can obtain bounding boxes (usually rectangles) to locate objects in images. <monospace>Faster R-CNN</monospace> [<xref rid="B22-sensors-21-06999" ref-type="bibr">22</xref>] is one of famous DNNs for object detection, and yielded successful applications for example in detecting apples [<xref rid="B23-sensors-21-06999" ref-type="bibr">23</xref>], mangoes with spatial registration [<xref rid="B24-sensors-21-06999" ref-type="bibr">24</xref>], and sweet papers using multi-modal data (RGB colors and Near-infrared) [<xref rid="B25-sensors-21-06999" ref-type="bibr">25</xref>]. However, such rectangular bounding boxes are not descriptive enough to obtain the sizes or shapes of the target fruits. Indeed, in [<xref rid="B8-sensors-21-06999" ref-type="bibr">8</xref>] applying <monospace>YOLOv3</monospace> [<xref rid="B10-sensors-21-06999" ref-type="bibr">10</xref>] was not enough, as described above. Now, <monospace>Mask-RCNN</monospace> [<xref rid="B26-sensors-21-06999" ref-type="bibr">26</xref>] is an implementation of instance segmentation, which locates individual objects and applies image segmentation accordingly. It was applied to identify individual bunches of grapes for more accurate size (volume) estimates in [<xref rid="B27-sensors-21-06999" ref-type="bibr">27</xref>], and to identify individual blue berries for maturity estimations in [<xref rid="B28-sensors-21-06999" ref-type="bibr">28</xref>]. Moreover, in [<xref rid="B29-sensors-21-06999" ref-type="bibr">29</xref>] DNNs were trained by synthesized data to directly count tomatoes.</p><p>For more accurate size measurement and shape description, one needs to put more weight on image segmentation. Recently, deep learning methods started replacing some of the pixel-wise color analysis methods for image segmentation in locating and counting fruit [<xref rid="B30-sensors-21-06999" ref-type="bibr">30</xref>] (see [<xref rid="B31-sensors-21-06999" ref-type="bibr">31</xref>] as well). As discussed above, in [<xref rid="B9-sensors-21-06999" ref-type="bibr">9</xref>] then deep learning was applied for edge detection, which is closely related to image segmentation, in order to predict the diameters of apples. Although there are preferred light conditions and necessary prepossessing in [<xref rid="B9-sensors-21-06999" ref-type="bibr">9</xref>], our neural network, which we call <monospace>CROP</monospace> (Central Roundish Object Painter), conducts image segmentation and makes masks for various roundish fruit located in the center of un-preprocessed images captured in varied surroundings; even with near-infrared light flash in the dark. Although counting the pixels of these masks gives the relative 2-dimensional size of the fruit, which is different from the conventional measures of fruit size, one could keep track of the pear growth as in <xref rid="sec3dot4-sensors-21-06999" ref-type="sec">Section 3.4</xref> and <xref rid="sec3dot5-sensors-21-06999" ref-type="sec">Section 3.5</xref>. Moreover, we believe that this new measure is more stable against measurement disturbances. Furthermore, the methodology would be easily transferred to other types of roundish fruit, which was qualitatively examined in <xref rid="sec3dot3-sensors-21-06999" ref-type="sec">Section 3.3</xref>.</p></sec><sec id="sec1dot3-sensors-21-06999"><title>1.3. About Deep Learning</title><p>Deep learning (DL) is one of machine learning techniques. There are more and more applications of DL not only in fruit production but also in agriculture [<xref rid="B32-sensors-21-06999" ref-type="bibr">32</xref>]. Now, we briefly look into DL. Human-made algorithms require humans to set features to be extracted from raw data. In reality, fruit images in farms vary extensively. Indeed, the fruit can be of any species and at any development stage. Even if these variables are fixed, the surroundings can change depending on time of day, weather conditions, seasons, and other illumination conditions such as shade and light reflection. Therefore, potentially huge number of features are necessary in acquiring desired information from such images. In this case, then it may be simply too much for humans to find and list all important features to write algorithms by hand. In contrast, data-driven methods, such as DL with DNNs (Deep Neural Networks), could extract important features automatically, so that they process data in a ‘black box’, although humans can guess how they process data using for example <monospace>Grad-CAM</monospace> [<xref rid="B33-sensors-21-06999" ref-type="bibr">33</xref>]. Roughly writing, as data flow in a DNN made of many layers, each layer makes the data a little more abstract, so that the DNN yields abstract understanding deep within itself. This way, humans may not have to pick important features by themselves, which distinguishes DL from conventional methods. Interested readers can consult [<xref rid="B34-sensors-21-06999" ref-type="bibr">34</xref>] for more detailed explanations on DL.</p></sec><sec id="sec1dot4-sensors-21-06999"><title>1.4. Pear Production in Japan</title><p><italic toggle="yes">La France</italic> is one of the most popular cultivars of European pear (Pyrus communis) in Japan, its yield is about 70% of European pears in Japan. La France pomes are usually harvested at the mature-green stage and then chilled for stimulation of ethylene biosynthesis prior to being ripened at the room temperature. If the harvest is delayed or too early, the fruit does not ripe properly, and the texture, taste and flavor will be poor. In commercial pear cultivation, harvest time of La France greatly influences both the amount and quality of the harvest. Therefore, it is important to measure the maturity of the fruit precisely to optimize the time to harvest, but criteria to estimate the fruit maturity are limited such as fruit firmness and blooming date. The fruit growth in terms of fruit size is described by an asymmetric sigmoid curve. The growth rate of the fruit on the tree is significantly affected by environmental factors and the physiologically active state. Precise time-lapse measurement of fruit growth should be useful for estimation of the fruit maturation status. To measure the fruit size change as it grows, the size of the same fruit must be measured repeatedly (daily) with a caliper. However instead, we propose to record the area change of the fruit in the time-series images captured by fixed cameras. We hope that our methodology provides horticultural scientists with big data on fruit development processes with least labor.</p></sec></sec><sec sec-type="methods" id="sec2-sensors-21-06999"><title>2. Methods and Materials</title><sec id="sec2dot1-sensors-21-06999"><title>2.1. Neural Networks for Image Segmentation</title><p>In a sense, <italic toggle="yes">image segmentation</italic> can be seen as classification of image pixels. For example, in this study, we want to cut out the central fruit in an image, but it amounts to deciding which class label each pixel belongs to, the central fruit or the region outside it, i.e., the background. In general, the number of class labels can be more than two and such image processing is called <italic toggle="yes">semantic segmentation</italic>. In the rest of this section, we explain about our neural networks and compare them with others.</p><p>Strictly speaking, <monospace>CROP</monospace> is the name for our trained DNNs with special properties, but we call them so even before the training for convenience. As in <xref rid="sensors-21-06999-f002" ref-type="fig">Figure 2</xref>a, <monospace>CROP</monospace> takes in <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm1" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>512</mml:mn><mml:mo>×</mml:mo><mml:mn>512</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula>-pixel RGB images and outputs <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm2" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>512</mml:mn><mml:mo>×</mml:mo><mml:mn>512</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula>-pixel gray-scale images. These inputs and outputs are normalized to take decimal numbers between 0 and 1, unlike usual images taking integer values between 0 and 255. The outputs will turn into masks with some threshold, for example 0.5 in our case. The downward red arrows represent convolution layers with kernel <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm3" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mo>×</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> and stride 2, which double the number of channels. Here, <italic toggle="yes">channel</italic> refers to another dimension than height and width. For example, the number of channels of RGB images is 3 and that of gray-scale images is 1. Similarly, the upward green arrows represent transposed convolution layers with kernel <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm4" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mo>×</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> and stride 2, which however make the number of channels half. The red rectangles are concatenations of two convolution layers with kernel <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm5" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>×</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> and stride 1, which keep the number of channels unchanged. The green rectangles are again concatenations of convolution layers, where the second convolutions are the same as the ones from the red rectangles, but the first ones are a bit different. Their inputs are direct sums (in the channel space) of the outputs of the layers below and those from the left, the latter of which are indicated by horizontal arrows. With these inputs then those concatenated convolution layers make the number of channels half. Pink and light green boxes represent again concatenations of convolution layers with kernel <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm6" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>×</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> and stride 1. Going through these layers the number of channels changes as <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm7" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>→</mml:mo><mml:mn>16</mml:mn><mml:mo>→</mml:mo><mml:mn>16</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm8" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>32</mml:mn><mml:mo>→</mml:mo><mml:mn>16</mml:mn><mml:mo>→</mml:mo><mml:mn>16</mml:mn><mml:mo>→</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula>, respectively.</p><p>The architecture of <monospace>CROP</monospace> is based on <monospace>U-Net</monospace> [<xref rid="B35-sensors-21-06999" ref-type="bibr">35</xref>]. <monospace>U-Net</monospace> was developed for medical image segmentation, and many related research projects were conducted, including <monospace>V-Net</monospace> for 3-dimensional medical images [<xref rid="B36-sensors-21-06999" ref-type="bibr">36</xref>], from which we adopted the loss function. <monospace>U-Net</monospace> (and <monospace>V-net</monospace> as well) belongs to the family of <italic toggle="yes">Convolutional Neural Networks (CNNs)</italic>, which treat 2-dimensional images as they are. Some argue that CNNs can be traced back to [<xref rid="B37-sensors-21-06999" ref-type="bibr">37</xref>]. Historically, CNNs proved to be useful for example in handwriting recognition [<xref rid="B38-sensors-21-06999" ref-type="bibr">38</xref>], and then <monospace>AlexNet</monospace> [<xref rid="B39-sensors-21-06999" ref-type="bibr">39</xref>] improved the error rate dramatically at ImageNet Large Scale Visual Recognition Challenge in 2012. Naturally, this architecture started being adapted to semantic segmentation for example in [<xref rid="B40-sensors-21-06999" ref-type="bibr">40</xref>], and with <italic toggle="yes">encoder-decoder</italic> structure in [<xref rid="B41-sensors-21-06999" ref-type="bibr">41</xref>,<xref rid="B42-sensors-21-06999" ref-type="bibr">42</xref>]. The encoder-decoder structure consists of two parts; in addition to size-invariant convolutions, encoding with convolutions or max pools and decoding with transposed convolutions or up samplings. The encoder extracts image features and is sometimes called <italic toggle="yes">backbone</italic>. Indeed, CNNs trained for image classification, which we suppose have already learned how to extract features, are often placed as backbones, for example <monospace>ResNet-50</monospace> [<xref rid="B11-sensors-21-06999" ref-type="bibr">11</xref>]. In <xref rid="sensors-21-06999-f002" ref-type="fig">Figure 2</xref>a, the left hand side corresponds to the encoder and the right the decoder. Importantly, there are <italic toggle="yes">skip connections</italic> represented by horizontal arrows. They are supposed to transmit location information from the encoder to the decoder, and characterize <monospace>U-Net</monospace>.</p><p><monospace>CROP</monospace> has a similar structure as <monospace>U-Net</monospace>, but the decoder and encoder are deeper than those of <monospace>U-Net</monospace>. In this study, <monospace>CROP-Shallow</monospace> in <xref rid="sensors-21-06999-f002" ref-type="fig">Figure 2</xref>b can be thought to be analogous of <monospace>U-Net</monospace>. Indeed, <monospace>U-Net</monospace> applies max pools 4 times, and at the bottom the number of channels will become 1024 while the size of feature map will be nearly <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm9" display="block" overflow="scroll"><mml:mrow><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>−</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> as small as the input image. In contrast, <monospace>CROP</monospace> applies 7 times convolution operations with kernel <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm10" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mo>×</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> and stride 2, and at the bottom the size of feature map will be <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm11" display="block" overflow="scroll"><mml:mrow><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>−</mml:mo><mml:mn>7</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> as small as the input image. We believe that this difference in depth enables <monospace>CROP</monospace> to have larger receptive fields; see <xref rid="sec3dot2-sensors-21-06999" ref-type="sec">Section 3.2</xref> for quantitative experiments on this matter.</p><p>Before concluding this section, we write about two more well-known DNNs for image (instance) segmentation. First, <monospace>Mask R-CNN</monospace> [<xref rid="B26-sensors-21-06999" ref-type="bibr">26</xref>] is an implementation of <italic toggle="yes">instance segmentation</italic>. Instance segmentation distinguishes individual instances of the same class label, while semantic segmentation does not. Although it can have wide-ranging applications in horticultural science (see <xref rid="sec1dot3-sensors-21-06999" ref-type="sec">Section 1.3</xref> for some), since we do not have to analyze all fruit in the image, we can rather focus on image segmentation. Next, <monospace>DeepLabv3+</monospace> [<xref rid="B43-sensors-21-06999" ref-type="bibr">43</xref>] is an implementation of semantic segmentation and has <italic toggle="yes">atrous convolutions</italic> to capture contextual information. However, because we do not need contextual information for our purpose, we chose to use <monospace>U-Net</monospace> for preciser image segmentation.</p></sec><sec id="sec2dot2-sensors-21-06999"><title>2.2. Loss Functions and Evaluation Criteria</title><p>Choice of loss functions defines how to update the parameters of DNNs, i.e., how they learn. In this project, we used <italic toggle="yes">soft dice loss</italic> [<xref rid="B36-sensors-21-06999" ref-type="bibr">36</xref>]:<disp-formula id="FD1-sensors-21-06999"><label>(1)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm12" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>−</mml:mo><mml:mn>2</mml:mn><mml:mfenced separators="" open="(" close=")"><mml:munder><mml:mo>∑</mml:mo><mml:mi>k</mml:mi></mml:munder><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mfenced><mml:mo stretchy="true">/</mml:mo><mml:mfenced separators="" open="(" close=")"><mml:munder><mml:mo>∑</mml:mo><mml:mi>k</mml:mi></mml:munder><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:munder><mml:mo>∑</mml:mo><mml:mi>k</mml:mi></mml:munder><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:mfenced></mml:mrow></mml:mrow></mml:math></disp-formula>
which is a variant of <italic toggle="yes">dice loss</italic>, where <italic toggle="yes">k</italic> runs over all the pixels; <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm13" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>512</mml:mn><mml:mo>×</mml:mo><mml:mn>512</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> in our case. Additionally <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm14" display="block" overflow="scroll"><mml:mrow><mml:msub><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are outputs of <monospace>CROP</monospace> (gray-scale images), taking values between 0 and 1 after going through the sigmoid function: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm15" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>−</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula>, while <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm16" display="block" overflow="scroll"><mml:mrow><mml:msub><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are ground truth (masks), taking values of only 0 or 1, where 0 and 1 correspond to the background and the central roundish object, respectively; we swapped the roles and took the average. Please note that the loss vanishes if <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm17" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:math></inline-formula> for all <italic toggle="yes">k</italic>’s. Of course, pixel-wise <italic toggle="yes">cross-entropy</italic>: <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm18" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:msub><mml:mo>∑</mml:mo><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo form="prefix">log</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>−</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo form="prefix">log</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>−</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:math></inline-formula> may seem to be a good choice if we consider image segmentation as classification of each pixel; cross-entropy is commonly used for image-classification tasks. However, with cross-entropy each pixel would carry the same share in the loss. Consequently, each mis-classified pixel would make the same amount of contribution to back-propagation (training process) regardless of the size of objects, and then it could let DNNs learn to rather ignore small objects. In contrast, with soft dice loss such imbalance would be compensated by regularization, which appears as the denominator in (<xref rid="FD1-sensors-21-06999" ref-type="disp-formula">1</xref>). For similar reasons, we did not use <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm19" display="block" overflow="scroll"><mml:mrow><mml:msub><mml:mi>l</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> loss (<inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm20" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>≥</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula>): <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm21" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:msub><mml:mo>∑</mml:mo><mml:mi>k</mml:mi></mml:msub><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi></mml:msup></mml:mrow></mml:mrow></mml:math></inline-formula>.</p><p>Finally, <italic toggle="yes">IoU (Intersection over Union)</italic>, or <italic toggle="yes">Jaccard index</italic>:<disp-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm22" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mi>IoU</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>∩</mml:mo><mml:mi>B</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>∪</mml:mo><mml:mi>B</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mspace width="4pt"/><mml:mo>.</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
measures how two sets are close to each other. It takes values between 0 and 1; <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm23" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mi>A</mml:mi><mml:mo>∩</mml:mo><mml:mi>B</mml:mi><mml:mo>=</mml:mo><mml:mo>∅</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm24" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mi>B</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> give 0 and 1, respectively. This is not a loss function because it is not differentiable, but was used in this paper to compare masks made by <monospace>CROP</monospace> and the ground truth masks to evaluate <monospace>CROP</monospace>.</p></sec><sec id="sec2dot3-sensors-21-06999"><title>2.3. Datasets</title><p>In this project, we used three groups of images downloaded from the Internet and captured at the farms (Kaminoyama, Yamagata, Japan).</p><list list-type="simple"><list-item><label>Data_Fruit</label><p>172 images of various fruit downloaded from <monospace>Pixabay</monospace>:</p><p>
<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="sensors-21-06999-i001.jpg"><?image-name sensors-21-06999-i001.jpg?><?image-size 35677?><?image-md5 2e8e3cb77d5ffcad0b2294615022e0ac?><?image-image-server-status NEVER_LOAD?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/2e8e3cb77d5f/sensors-21-06999-i001.jpg?></inline-graphic>
</p></list-item><list-item><label>Data_Pears1</label><p>26 images of pears at the farm in 2018 with <monospace>Brinno BCC100</monospace> (time-lapse mode):</p><p>
<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="sensors-21-06999-i002.jpg"><?image-name sensors-21-06999-i002.jpg?><?image-size 36228?><?image-md5 5e3da7932d143aae732cbfe36547efd0?><?image-image-server-status NEVER_LOAD?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/5e3da7932d14/sensors-21-06999-i002.jpg?></inline-graphic>
</p></list-item><list-item><label>Data_Pears2</label><p>86 images of pears at the farm in 2019 with various cameras:</p><p>
<inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="sensors-21-06999-i003.jpg"><?image-name sensors-21-06999-i003.jpg?><?image-size 37307?><?image-md5 892dd98d365ab5308ffdda44b046fe40?><?image-image-server-status NEVER_LOAD?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/892dd98d365a/sensors-21-06999-i003.jpg?></inline-graphic>
</p></list-item></list><p>These images were all annotated by <monospace>labelme</monospace> [<xref rid="B44-sensors-21-06999" ref-type="bibr">44</xref>] to train and evaluate <monospace>CROP</monospace>. Additionally, <monospace>Brinno BCC100</monospace> is a time-lapse camera previously used by one of the authors to keep track of the development process of pears.</p><p>For training and evaluation, data augmentation was conducted, unless otherwise stated. It includes geometrical changes: size alternations, horizontal and vertical flips, rotations, and various photo-metric changes.</p></sec><sec id="sec2dot4-sensors-21-06999"><title>2.4. Hardware and Software</title><p>For this project, the following two GPUs were used: <monospace>TITAN Xp</monospace> and <monospace>GeForce RTX 2080 Ti</monospace>. The machine learning library was <monospace>PyTorch</monospace> [<xref rid="B45-sensors-21-06999" ref-type="bibr">45</xref>]. The graphs in this paper were drawn by <monospace>Matplotlib</monospace> [<xref rid="B46-sensors-21-06999" ref-type="bibr">46</xref>] and <monospace>Seaborn</monospace> [<xref rid="B47-sensors-21-06999" ref-type="bibr">47</xref>].</p></sec></sec><sec sec-type="results" id="sec3-sensors-21-06999"><title>3. Results</title><sec id="sec3dot1-sensors-21-06999"><title>3.1. Quantitative Analysis</title><sec id="sec3dot1dot1-sensors-21-06999"><title>3.1.1. Training <monospace>CROP</monospace></title><p>For quantitative analysis we divided <monospace>Data_Fruit</monospace> randomly into the training dataset (80%: 137) and the validation dataset (20%: 35). Initial parameters of <monospace>CROP</monospace> were set randomly, where we did not use a pre-trained model for the backbone. The parameters were then updated by <monospace>Adam</monospace> with learning rate <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm25" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>0.001</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula>. The result can be seen in <xref rid="sensors-21-06999-f003" ref-type="fig">Figure 3</xref>, where the best IoU for the validation data were 0.984, achieved at the epoch 9700, where the threshold for masks was 0.5. In this experiment, The individual evaluations were made every 100 epochs and were the average over the 5 independent evaluations with application of random data augmentations, i.e., the training dataset and the validation dataset were practically of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm26" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>137</mml:mn><mml:mo>×</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm27" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>35</mml:mn><mml:mo>×</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula>, respectively. To obtain an idea about the IoU, suppose that we have two 100-pixel masks of the ground truth and a prediction, and that 99 pixels are correctly predicted. Then, the IoU would be:<disp-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm28" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mrow><mml:mi>IoU</mml:mi><mml:mo>(</mml:mo><mml:mi>ground</mml:mi></mml:mrow><mml:mspace width="4.pt"/><mml:mrow><mml:mi>truth</mml:mi><mml:mo>,</mml:mo></mml:mrow><mml:mspace width="4.pt"/><mml:mrow><mml:mi>prediction</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>99</mml:mn><mml:mn>101</mml:mn></mml:mfrac><mml:mo>=</mml:mo><mml:mn>0.98019</mml:mn><mml:mo>…</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula></p><p>This optimal <monospace>CROP</monospace> corresponds to the network dictionary named net_dic_0601_09700, and was applied to test data (<monospace>Data_Pears1</monospace>) in <xref rid="sec3dot4-sensors-21-06999" ref-type="sec">Section 3.4</xref> to give predictions before the fine-tuning. Please note that the network dictionaries store the learned parameters of the neural networks. On the other hand, for qualitative analysis in <xref rid="sec3dot3-sensors-21-06999" ref-type="sec">Section 3.3</xref>, <monospace>CROP</monospace> was trained by all the images of <monospace>Data_Fruit</monospace> until epoch 5000. The corresponding network dictionary was named net_dic_0314_05000. These two dictionaries are available on <monospace>GitHub</monospace>.</p></sec><sec id="sec3dot1dot2-sensors-21-06999"><title>3.1.2. Fine-Tuning <monospace>CROP</monospace></title><p>For fine-tuning, we divided <monospace>Data_Pears2</monospace> randomly into the training dataset (80%: 68) and the validation dataset (20%: 18), and then we trained further the optimal <monospace>CROP</monospace> from <xref rid="sec3dot1dot1-sensors-21-06999" ref-type="sec">Section 3.1.1</xref> with the dictionary named net_dic_0601_09700. The optimization method was <monospace>Adam</monospace> with learning rate 0.0001. One can see the process in <xref rid="sensors-21-06999-f004" ref-type="fig">Figure 4</xref>; the best IoU for the validation data was 0.983 achieved at epoch 5,100. This <monospace>CROP</monospace> was applied to test data (<monospace>Data_Pears1</monospace>) in <xref rid="sec3dot4-sensors-21-06999" ref-type="sec">Section 3.4</xref> to provide predictions after the fine-tuning. This network dictionary was named net_dic_ft_1015_05100. Furthermore, we fine-tuned the network dictionary named net_dic_0314_05000 for 5000 epochs using the whole <monospace>Data_Pears2</monospace> to obtain the network dictionary named net_dic_ft_0328_1_5000, which was used in <xref rid="sec3dot5-sensors-21-06999" ref-type="sec">Section 3.5</xref>. These two dictionaries are also available on <monospace>GitHub</monospace>. One can see the improvement of the IoU in <xref rid="sensors-21-06999-t001" ref-type="table">Table 1</xref>, and the refinement in some examples of test data (<monospace>Data_Pears1</monospace>) in <xref rid="sec3dot4-sensors-21-06999" ref-type="sec">Section 3.4</xref>. Note that data augmentation was not applied in the evaluation.</p></sec></sec><sec id="sec3dot2-sensors-21-06999"><title>3.2. Depth of Neural Networks</title><p>In this section, we report on experiments on <monospace>CROP-Shallow</monospace> (<xref rid="sensors-21-06999-f002" ref-type="fig">Figure 2</xref>b). We trained <monospace>CROP-Shallow</monospace> and <monospace>CROP</monospace> in the same conditions: loss function, optimization method, learning rate and batch size, i.e., they were the same as in <xref rid="sec3dot1dot1-sensors-21-06999" ref-type="sec">Section 3.1.1</xref> except that the batch size was 6, which is the maximum for <monospace>CROP-Shallow</monospace> in terms of the GPU memory capacity. Random were partitions of training and validation datasets, initialization of these neural networks, batches and application of data augmentation. In the same way as <xref rid="sec3dot1dot1-sensors-21-06999" ref-type="sec">Section 3.1.1</xref>, random data augmentation was applied to have <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm29" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>137</mml:mn><mml:mo>×</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> training dataset and the <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm30" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>35</mml:mn><mml:mo>×</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> validation dataset for the individual evaluations. These random methods were adopted to avoid particular biases, and moreover three independent experiments were performed for each neural network. One can see from <xref rid="sensors-21-06999-t002" ref-type="table">Table 2</xref> and <xref rid="sensors-21-06999-f005" ref-type="fig">Figure 5</xref> that <monospace>CROP</monospace> fits our dataset more and the IoUs of <monospace>CROP-Shallow</monospace> seem to have hit the ceiling towards the epoch 2500. We believe the difference came from the fact that segmenting larger objects in images would require larger receptive fields, which can be obtained from deeper CNNs (remember convolution layers act locally), while <monospace>U-Net</monospace> was originally invented to segment rather small cells in biomedical images.</p></sec><sec id="sec3dot3-sensors-21-06999"><title>3.3. Qualitative Analysis</title><p>Now, we apply <monospace>CROP</monospace> to sample images. To make each prediction stable, we fed <monospace>CROP</monospace> the images made by the eight transformations composed by flips and rotations, which keep a right square unchanged in the two-dimensional space. Then, each mask was made from the average of those eight outputs with threshold 0.5. In the figures, the masks were pasted onto the input images; red for the central object and yellow for the background. The implementation of this image processing is also available on <monospace>GitHub</monospace>. Please note that the sample images in this section do not belong to our datasets, and they can be thought of as test data, i.e., they are new to <monospace>CROP</monospace>.</p><p>First, examples in <xref rid="sensors-21-06999-f006" ref-type="fig">Figure 6</xref>a,b,d indicate how <monospace>CROP</monospace> identified individual central grapes. In contrast, <monospace>CROP</monospace> became confused in <xref rid="sensors-21-06999-f006" ref-type="fig">Figure 6</xref>c because the central grape was behind others. Next, one can see in <xref rid="sensors-21-06999-f007" ref-type="fig">Figure 7</xref> that <monospace>CROP</monospace> handled images of various fruit, dealing with peduncle and calyx. However, unsuccessful examples are presented in <xref rid="sensors-21-06999-f008" ref-type="fig">Figure 8</xref>. Finally, it is interesting to point out that although <monospace>CROP</monospace> has been trained solely by 172 fruit images, without transfer learning, it gained some sort of generality as in <xref rid="sensors-21-06999-f009" ref-type="fig">Figure 9</xref>.</p></sec><sec id="sec3dot4-sensors-21-06999"><title>3.4. Fine-Tuning to Pears</title><p>In this section, we make qualitative analysis of fine-tuning explained in <xref rid="sec3dot1dot2-sensors-21-06999" ref-type="sec">Section 3.1.2</xref>. In <xref rid="sensors-21-06999-f010" ref-type="fig">Figure 10</xref>, each triple consists of the original image, and processed images before and after the fine-tuning, placed from left to right.</p></sec><sec id="sec3dot5-sensors-21-06999"><title>3.5. Applying <monospace>CROP</monospace> to Time-Series Images</title><p>In this section, we demonstrate our methodology by processing 510 images on the target pear, which were captured in Kaminoyama, Yamagata, Japan during 12 August 2020 13:49–15 October 2020 8:00; eight a day at 8:00, 9:49, 11:49, 13:49, 15:49, 17:49, 19:49, 21:49, except for the first five and the last one. The images were given IDs from 2 to 511 chronologically. The camera was SINEI SC-MB68 a trail camera. Please note that these images are new to <monospace>CROP</monospace>; see <xref rid="sec2dot3-sensors-21-06999" ref-type="sec">Section 2.3</xref> for our datasets. At the end of this section, we treat only the daytime images and remove outliers based on the coefficient of variation (standard deviation divided by the mean) to obtain a growth curve of the pear.</p><p>The key ideas in processing time-series images by <monospace>CROP</monospace> are:<list list-type="order"><list-item><p>It can detect the central roundish fruit as in <xref rid="sensors-21-06999-f011" ref-type="fig">Figure 11</xref>b,c. By applying this functionality repeatedly, it can keep track of the target fruit.</p></list-item><list-item><p>It can work in different scales. One can take the median of 11 measurements of different scales to exclude outliers as in <xref rid="sensors-21-06999-f012" ref-type="fig">Figure 12</xref>.</p></list-item><list-item><p>It can keep track of the 2D-wise center of mass of the mask, as in <xref rid="sensors-21-06999-f011" ref-type="fig">Figure 11</xref>d.</p></list-item></list></p><fig position="float" id="sensors-21-06999-f011" orientation="portrait"><label>Figure 11</label><caption><p>(<bold>a</bold>) The original image of <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm33" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mn>5200</mml:mn><mml:mo>×</mml:mo><mml:mn>3900</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> pixels. (<bold>b</bold>) Choosing the target fruit roughly manually. (<bold>c</bold>) <monospace>CROP</monospace> will then recognize the central fruit. (<bold>d</bold>) It will measure the area as the number of pixels of the mask and give the 2D-wise center of mass. The pixel number is re-scaled in the scale of (<bold>a</bold>) and hence is in general decimal. See <xref rid="sensors-21-06999-f012" ref-type="fig">Figure 12</xref> for how to obtain the mask. The image ID is 338.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g011.jpg"><?image-name sensors-21-06999-g011.jpg?><?image-size 51774?><?image-md5 b541df1582cbaa29d6c29fac027465cc?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 870?><?image-original-width 3295?><?image-scaled-height 193?><?image-scaled-width 732?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/b541df1582cb/sensors-21-06999-g011.jpg?><?thumb-name sensors-21-06999-g011.gif?><?thumb-size 11406?><?thumb-md5 ebe9f3d5569fc3d095270009af24e780?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 53?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/ebe9f3d5569f/sensors-21-06999-g011.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f012" orientation="portrait"><label>Figure 12</label><caption><p>(<bold>a</bold>) CROP makes 11 measurements in different scale factors; from ×1.0 to ×0.5. (<bold>b</bold>) The median of re-scaled values will be picked to exclude outliers. The corresponding mask will be picked as in <xref rid="sensors-21-06999-f011" ref-type="fig">Figure 11</xref>d. The image ID is 338.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g012.jpg"><?image-name sensors-21-06999-g012.jpg?><?image-size 108307?><?image-md5 3d1d13e83d0e568f3ab7b0e5153f4cb4?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 2160?><?image-original-width 2230?><?image-scaled-height 720?><?image-scaled-width 743?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/3d1d13e83d0e/sensors-21-06999-g012.jpg?><?thumb-name sensors-21-06999-g012.gif?><?thumb-size 8482?><?thumb-md5 0d4009fc34d588595181e51e300bc94a?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 97?><?thumb-scaled-width 100?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/0d4009fc34d5/sensors-21-06999-g012.gif?></graphic></fig><p>The graph in <xref rid="sensors-21-06999-f013" ref-type="fig">Figure 13</xref> shows the growth curve of the target fruit where the alleged outliers remain (some are out of the range). The growth curve here is estimated by the areas of the masks created by <monospace>CROP</monospace>; taking the square root may give the equivalent measure for the diameter or length, though.</p><p>Now, we shall focus on the five days: 8–12 October 2020 (the image IDs are from 455 to 494), where the measurements were rather stable. <xref rid="sensors-21-06999-f014" ref-type="fig">Figure 14</xref> shows the boxplot of all the measurements. Each box represents the 11 measurements as in <xref rid="sensors-21-06999-f012" ref-type="fig">Figure 12</xref>. One can see that the measurements tend to have higher variance during the evening time (17:49, 19:49, 21:49). In addition, the eight images captured on 12 October 2020 and their predicted masks are in <xref rid="sensors-21-06999-f015" ref-type="fig">Figure 15</xref>.</p><p>Also, <xref rid="sensors-21-06999-f016" ref-type="fig">Figure 16</xref>a shows how the target fruit moved around during the whole season, where some possible outliers stay inside and outside the range. On the other hand, <xref rid="sensors-21-06999-f016" ref-type="fig">Figure 16</xref>b describes the positions of the target fruit during the above five days.</p><p>To obtain a nicer growth curve, we treat only 315 daytime images; 5 images per day for 63 days during 13 August–14 October 2020. We used coefficient of variation (CV), which is the standard deviation divided by the mean, to identify outliers; the statistics are made of the 11 measurements (<xref rid="sensors-21-06999-f012" ref-type="fig">Figure 12</xref>). More concretely, we replaced the medians of the 15 images with largest CVs, which amount to less than the 5% of 315, by the medians of the previous images. The histograms of the top 15 CVs and others are found in <xref rid="sensors-21-06999-f017" ref-type="fig">Figure 17</xref>.</p><p>In <xref rid="sensors-21-06999-f018" ref-type="fig">Figure 18</xref>a,b, one can see the processed images corresponding to the two highest CVs. Additionally, we noticed that there was a confident mistake for the ID 393 (<xref rid="sensors-21-06999-f018" ref-type="fig">Figure 18</xref>c), which stood out in the provisional plotting and was modified manually in the same way. After applying these changes, we obtained the growth curve shown in <xref rid="sensors-21-06999-f019" ref-type="fig">Figure 19</xref>. All the boxplots after the modification are placed in <xref rid="app1-sensors-21-06999" ref-type="app">Appendix A</xref>. The 16th largest CV corresponds to the ID 338, which was used in <xref rid="sensors-21-06999-f011" ref-type="fig">Figure 11</xref> and <xref rid="sensors-21-06999-f012" ref-type="fig">Figure 12</xref>. The outliers appearing at ID 338 of <xref rid="sensors-21-06999-f0A6" ref-type="fig">Figure A6</xref> was properly dealt with as in <xref rid="sensors-21-06999-f012" ref-type="fig">Figure 12</xref>; these two sets of measurements were conducted independently, though.</p><p>Finally, we shall emphasize the fact that it took only 791.1025 s, which is less than 14 min, to process the 510 images with <monospace>NVIDIA TITAN Xp</monospace>. All the process was automatic after specifying the target fruit similarly as in <xref rid="sensors-21-06999-f011" ref-type="fig">Figure 11</xref>b for the ID 2. During this process, for individual images, <monospace>CROP</monospace> made the 11 measurements and chose the median and calculated the center of mass, so that it saved all the numerical data as csv files and all the masks as png files.</p></sec></sec><sec sec-type="discussion" id="sec4-sensors-21-06999"><title>4. Discussions</title><p>In this project, <monospace>CROP</monospace> had gained general ability to segment central roundish objects, although we trained it by fruit images. With more training data, it would increase accuracy and generality. Note however that the image processing of <monospace>CROP</monospace> is different from <italic toggle="yes">salient object detection</italic> [<xref rid="B48-sensors-21-06999" ref-type="bibr">48</xref>], as <monospace>CROP</monospace> identifies a small central object such as the one in <xref rid="sensors-21-06999-f006" ref-type="fig">Figure 6</xref>d. This is why we call our method <italic toggle="yes">central image segmentation</italic>. We hope that our non-contact method of size measurement will provide scientific big data for the advancement of science, especially in the field of fruit production. In fact, as of September 2021, pear cultivators in Kaminoyama Yamagata can know, via the smartphone application, the daily sizes of the sampled pears, which are predicted by <monospace>CROP</monospace>. To conclude this paper, we briefly write how one can adapt <monospace>CROP</monospace> to multi/hyper-spectral camera images. In principle, if you have <italic toggle="yes">n</italic>-band images, the number 3 in the pink box in <xref rid="sensors-21-06999-f002" ref-type="fig">Figure 2</xref>a should be replaced by <italic toggle="yes">n</italic>. For example, thermal cameras result in <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm31" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula> and RGBN cameras <inline-formula><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="mm32" display="block" overflow="scroll"><mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:mrow></mml:math></inline-formula>. How large <italic toggle="yes">n</italic> can be depends on how much memory the GPU has. Our GPU <monospace>TITAN Xp</monospace> has 12GB of memory, and in this case we believe <italic toggle="yes">n</italic> could be 10 or so while the training may become unstable because the batch size would be smaller. This problem can be overcome to some extent conducting parallel computation which can deal with batch normalization problem. Apart from the GPU memory issue, we believe when <italic toggle="yes">n</italic> is large one should think whether or not at least the channel number 16 in the pink box in <xref rid="sensors-21-06999-f002" ref-type="fig">Figure 2</xref>a is appropriate, because <italic toggle="yes">n</italic>-mode images, i.e., <italic toggle="yes">n</italic>-dimensional information, would be mapped into 16-dimensional space very at the beginning.</p></sec></body><back><ack><title>Acknowledgments</title><p>M.F. gratefully acknowledges the support of NVIDIA Corporation with the donation of the TITAN Xp GPU. M.F. was financially supported by Leibniz Universität Hannover to present the result and have fruitful discussions in Hannover. M.F. thanks his colleague Richard Jordan for having a discussion with him to name our trained neural network and for suggesting better expressions in the title and abstract. M.F. and T.O. were financially supported by Yamagata University (YU-COE program). T.O. thanks Kazumi Sato and Yota Ozeki, who let him take images of pears in the farms, and Yota Sato for annotating <monospace>Data_Pears2</monospace>. The authors also thank Kazunari Adachi of the engineering department for giving us valuable legal advice concerning this research project, and Hideki Murayama of the agricultural department for providing us with photos of fruits. M.F. thanks Peggy Greb of Agricultural Research Service, USDA for providing consultation for the use of the images.</p></ack><fn-group><fn><p><bold>Publisher’s Note:</bold> MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p></fn></fn-group><notes><title>Author Contributions</title><p>Conceptualization, M.F., T.O. and S.Y.; Data curation, M.F. and T.O.; Formal analysis, M.F.; Investigation, M.F.; Methodology, M.F., T.O. and S.Y.; Project administration, M.F.; Resources, M.F. and T.O.; Software, M.F. and S.Y.; Validation, M.F.; Visualization, M.F.; Writing—original draft, M.F.; Writing—review and editing, M.F., T.O. and S.Y. All authors have read and agreed to the published version of the manuscript.</p></notes><notes><title>Funding</title><p>This research received no external funding.</p></notes><notes><title>Institutional Review Board Statement</title><p>Not applicable.</p></notes><notes><title>Informed Consent Statement</title><p>Not applicable.</p></notes><notes notes-type="data-availability"><title>Data Availability Statement</title><p>Our trained neural network <monospace>CROP</monospace> and the related programs are available on <monospace>GitHub</monospace> (<uri xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/MotohisaFukuda/CROP">https://github.com/MotohisaFukuda/CROP</uri>, accessed on 20 October 2021). Some of the images used for the qualitative analysis in this paper came from the image gallery organized by United States Department of Agriculture, Agricultural Research Service (<uri xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.ars.usda.gov/oc/images/image-gallery">https://www.ars.usda.gov/oc/images/image-gallery</uri>, accessed on 20 October 2021). Data_Fruit the training dataset in this project consists of images downloaded from <monospace>Pixabay</monospace> (<uri xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://pixabay.com">https://pixabay.com</uri>, accessed on 20 October 2021).</p></notes><notes notes-type="COI-statement"><title>Conflicts of Interest</title><p>The authors declare no conflict of interest.</p></notes><app-group><app id="app1-sensors-21-06999"><title>Appendix A. The Boxplots of the Modified Daytime Data</title><p>The boxplots of the modified data are found below. The medians were used for the plot in <xref rid="sensors-21-06999-f019" ref-type="fig">Figure 19</xref>. Yellow indicates removal for high CV, and red the manual removal.</p><fig position="anchor" id="sensors-21-06999-f0A1" orientation="portrait"><label>Figure A1</label><caption><p>13–19 August.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g0A1.jpg"><?image-name sensors-21-06999-g0A1.jpg?><?image-size 32564?><?image-md5 ed5491f77094a747cc2d27ad498adc02?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 968?><?image-original-width 3310?><?image-scaled-height 215?><?image-scaled-width 735?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/ed5491f77094/sensors-21-06999-g0A1.jpg?><?thumb-name sensors-21-06999-g0A1.gif?><?thumb-size 6632?><?thumb-md5 45584fa1107e3805889d3bdd225015d7?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 58?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/45584fa1107e/sensors-21-06999-g0A1.gif?></graphic></fig><fig position="anchor" id="sensors-21-06999-f0A2" orientation="portrait"><label>Figure A2</label><caption><p>20–26 August.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g0A2.jpg"><?image-name sensors-21-06999-g0A2.jpg?><?image-size 31842?><?image-md5 9e261e4522a593b1bdfd31a2c84fd95a?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 953?><?image-original-width 3306?><?image-scaled-height 212?><?image-scaled-width 734?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/9e261e4522a5/sensors-21-06999-g0A2.jpg?><?thumb-name sensors-21-06999-g0A2.gif?><?thumb-size 6708?><?thumb-md5 750a86802c383407604dd18de3d1bc07?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 58?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/750a86802c38/sensors-21-06999-g0A2.gif?></graphic></fig><fig position="anchor" id="sensors-21-06999-f0A3" orientation="portrait"><label>Figure A3</label><caption><p>27 August–2 September.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g0A3.jpg"><?image-name sensors-21-06999-g0A3.jpg?><?image-size 33263?><?image-md5 68288a6c65799e19b14d3929bda4ec1c?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 953?><?image-original-width 3299?><?image-scaled-height 212?><?image-scaled-width 733?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/68288a6c6579/sensors-21-06999-g0A3.jpg?><?thumb-name sensors-21-06999-g0A3.gif?><?thumb-size 6889?><?thumb-md5 c879a46f2efe6ed274042b430e7c0096?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 58?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/c879a46f2efe/sensors-21-06999-g0A3.gif?></graphic></fig><fig position="anchor" id="sensors-21-06999-f0A4" orientation="portrait"><label>Figure A4</label><caption><p>3–9 September.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g0A4.jpg"><?image-name sensors-21-06999-g0A4.jpg?><?image-size 31969?><?image-md5 abbbe17a07a59a2ed6ad81a59d2367e0?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 982?><?image-original-width 3299?><?image-scaled-height 218?><?image-scaled-width 733?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/abbbe17a07a5/sensors-21-06999-g0A4.jpg?><?thumb-name sensors-21-06999-g0A4.gif?><?thumb-size 7343?><?thumb-md5 424692899828ff94bc908d5a8b208906?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 60?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/424692899828/sensors-21-06999-g0A4.gif?></graphic></fig><fig position="anchor" id="sensors-21-06999-f0A5" orientation="portrait"><label>Figure A5</label><caption><p>10–16 September.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g0A5.jpg"><?image-name sensors-21-06999-g0A5.jpg?><?image-size 33544?><?image-md5 41d85da1165ae63b1cf4d3de2977768a?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 971?><?image-original-width 3310?><?image-scaled-height 216?><?image-scaled-width 735?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/41d85da1165a/sensors-21-06999-g0A5.jpg?><?thumb-name sensors-21-06999-g0A5.gif?><?thumb-size 7258?><?thumb-md5 b6c9c93e8da3e47af369f9c90abc4c65?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 59?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/b6c9c93e8da3/sensors-21-06999-g0A5.gif?></graphic></fig><fig position="anchor" id="sensors-21-06999-f0A6" orientation="portrait"><label>Figure A6</label><caption><p>17–23 September.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g0A6.jpg"><?image-name sensors-21-06999-g0A6.jpg?><?image-size 34559?><?image-md5 1dbe1a6399be16f22ea4469667d0b45a?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 971?><?image-original-width 3291?><?image-scaled-height 216?><?image-scaled-width 731?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/1dbe1a6399be/sensors-21-06999-g0A6.jpg?><?thumb-name sensors-21-06999-g0A6.gif?><?thumb-size 7078?><?thumb-md5 c6c78206c6995041b5a244fb5370843d?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 59?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/c6c78206c699/sensors-21-06999-g0A6.gif?></graphic></fig><fig position="anchor" id="sensors-21-06999-f0A7" orientation="portrait"><label>Figure A7</label><caption><p>24–30 September.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g0A7.jpg"><?image-name sensors-21-06999-g0A7.jpg?><?image-size 36175?><?image-md5 102ec369115530f349cb87498350ab50?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 964?><?image-original-width 3291?><?image-scaled-height 214?><?image-scaled-width 731?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/102ec3691155/sensors-21-06999-g0A7.jpg?><?thumb-name sensors-21-06999-g0A7.gif?><?thumb-size 7642?><?thumb-md5 2347b1213d46d27cd62c627ba478fbf4?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 59?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/2347b1213d46/sensors-21-06999-g0A7.gif?></graphic></fig><fig position="anchor" id="sensors-21-06999-f0A8" orientation="portrait"><label>Figure A8</label><caption><p>1–7 October.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g0A8.jpg"><?image-name sensors-21-06999-g0A8.jpg?><?image-size 34744?><?image-md5 edc0a380868ce66972b2b93cc43d10d3?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 957?><?image-original-width 3306?><?image-scaled-height 212?><?image-scaled-width 734?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/edc0a380868c/sensors-21-06999-g0A8.jpg?><?thumb-name sensors-21-06999-g0A8.gif?><?thumb-size 7225?><?thumb-md5 ea9711297d0359085bc603f9dc6e79f9?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 58?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/ea9711297d03/sensors-21-06999-g0A8.gif?></graphic></fig><fig position="anchor" id="sensors-21-06999-f0A9" orientation="portrait"><label>Figure A9</label><caption><p>8–14 October.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g0A9.jpg"><?image-name sensors-21-06999-g0A9.jpg?><?image-size 32245?><?image-md5 598598df414fea4c31651273e60abc9c?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 946?><?image-original-width 3284?><?image-scaled-height 210?><?image-scaled-width 729?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/598598df414f/sensors-21-06999-g0A9.jpg?><?thumb-name sensors-21-06999-g0A9.gif?><?thumb-size 7692?><?thumb-md5 db6135488a3c4a585fdf9c53eddd8f1a?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 58?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/db6135488a3c/sensors-21-06999-g0A9.gif?></graphic></fig></app></app-group><ref-list><title>References</title><ref id="B1-sensors-21-06999"><label>1.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Gongal</surname><given-names>A.</given-names></name>
<name name-style="western"><surname>Amatya</surname><given-names>S.</given-names></name>
<name name-style="western"><surname>Karkee</surname><given-names>M.</given-names></name>
<name name-style="western"><surname>Zhang</surname><given-names>Q.</given-names></name>
<name name-style="western"><surname>Lewis</surname><given-names>K.</given-names></name>
</person-group><article-title>Sensors and systems for fruit detection and localization: A review</article-title><source>Comput. Electron. Agric.</source><year>2015</year><volume>116</volume><fpage>8</fpage><lpage>19</lpage><pub-id pub-id-type="doi">10.1016/j.compag.2015.05.021</pub-id></element-citation></ref><ref id="B2-sensors-21-06999"><label>2.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Mitchell</surname><given-names>P.</given-names></name>
</person-group><article-title>Pear fruit growth and the use of diameter to estimate fruit volume and weight</article-title><source>HortScience</source><year>1986</year><volume>21</volume><fpage>1003</fpage><lpage>1005</lpage></element-citation></ref><ref id="B3-sensors-21-06999"><label>3.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Moreda</surname><given-names>G.</given-names></name>
<name name-style="western"><surname>Ortiz-Cañavate</surname><given-names>J.</given-names></name>
<name name-style="western"><surname>García-Ramos</surname><given-names>F.J.</given-names></name>
<name name-style="western"><surname>Ruiz-Altisent</surname><given-names>M.</given-names></name>
</person-group><article-title>Non-destructive technologies for fruit and vegetable size determination—A review</article-title><source>J. Food Eng.</source><year>2009</year><volume>92</volume><fpage>119</fpage><lpage>136</lpage><pub-id pub-id-type="doi">10.1016/j.jfoodeng.2008.11.004</pub-id></element-citation></ref><ref id="B4-sensors-21-06999"><label>4.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Moreda</surname><given-names>G.</given-names></name>
<name name-style="western"><surname>Muñoz</surname><given-names>M.</given-names></name>
<name name-style="western"><surname>Ruiz-Altisent</surname><given-names>M.</given-names></name>
<name name-style="western"><surname>Perdigones</surname><given-names>A.</given-names></name>
</person-group><article-title>Shape determination of horticultural produce using two-dimensional computer vision—A review</article-title><source>J. Food Eng.</source><year>2012</year><volume>108</volume><fpage>245</fpage><lpage>261</lpage><pub-id pub-id-type="doi">10.1016/j.jfoodeng.2011.08.011</pub-id></element-citation></ref><ref id="B5-sensors-21-06999"><label>5.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Tijskens</surname><given-names>L.</given-names></name>
<name name-style="western"><surname>Unuk</surname><given-names>T.</given-names></name>
<name name-style="western"><surname>Okello</surname><given-names>R.</given-names></name>
<name name-style="western"><surname>Wubs</surname><given-names>A.</given-names></name>
<name name-style="western"><surname>Šuštar</surname><given-names>V.</given-names></name>
<name name-style="western"><surname>Šumak</surname><given-names>D.</given-names></name>
<name name-style="western"><surname>Schouten</surname><given-names>R.</given-names></name>
</person-group><article-title>From fruitlet to harvest: Modelling and predicting size and its distributions for tomato, apple and pepper fruit</article-title><source>Sci. Hortic.</source><year>2016</year><volume>204</volume><fpage>54</fpage><lpage>64</lpage><pub-id pub-id-type="doi">10.1016/j.scienta.2016.03.036</pub-id></element-citation></ref><ref id="B6-sensors-21-06999"><label>6.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Morandi</surname><given-names>B.</given-names></name>
<name name-style="western"><surname>Manfrini</surname><given-names>L.</given-names></name>
<name name-style="western"><surname>Zibordi</surname><given-names>M.</given-names></name>
<name name-style="western"><surname>Noferini</surname><given-names>M.</given-names></name>
<name name-style="western"><surname>Fiori</surname><given-names>G.</given-names></name>
<name name-style="western"><surname>Grappadelli</surname><given-names>L.C.</given-names></name>
</person-group><article-title>A low-cost device for accurate and continuous measurements of fruit diameter</article-title><source>HortScience</source><year>2007</year><volume>42</volume><fpage>1380</fpage><lpage>1382</lpage><pub-id pub-id-type="doi">10.21273/HORTSCI.42.6.1380</pub-id></element-citation></ref><ref id="B7-sensors-21-06999"><label>7.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Thalheimer</surname><given-names>M.</given-names></name>
</person-group><article-title>A new optoelectronic sensor for monitoring fruit or stem radial growth</article-title><source>Comput. Electron. Agric.</source><year>2016</year><volume>123</volume><fpage>149</fpage><lpage>153</lpage><pub-id pub-id-type="doi">10.1016/j.compag.2016.02.028</pub-id></element-citation></ref><ref id="B8-sensors-21-06999"><label>8.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Lu</surname><given-names>C.P.</given-names></name>
<name name-style="western"><surname>Liaw</surname><given-names>J.J.</given-names></name>
</person-group><article-title>A novel image measurement algorithm for common mushroom caps based on convolutional neural network</article-title><source>Comput. Electron. Agric.</source><year>2020</year><volume>171</volume><fpage>105336</fpage><pub-id pub-id-type="doi">10.1016/j.compag.2020.105336</pub-id></element-citation></ref><ref id="B9-sensors-21-06999"><label>9.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Wang</surname><given-names>D.</given-names></name>
<name name-style="western"><surname>Li</surname><given-names>C.</given-names></name>
<name name-style="western"><surname>Song</surname><given-names>H.</given-names></name>
<name name-style="western"><surname>Xiong</surname><given-names>H.</given-names></name>
<name name-style="western"><surname>Liu</surname><given-names>C.</given-names></name>
<name name-style="western"><surname>He</surname><given-names>D.</given-names></name>
</person-group><article-title>Deep Learning Approach for Apple Edge Detection to Remotely Monitor Apple Growth in Orchards</article-title><source>IEEE Access</source><year>2020</year><volume>8</volume><fpage>26911</fpage><lpage>26925</lpage><pub-id pub-id-type="doi">10.1109/ACCESS.2020.2971524</pub-id></element-citation></ref><ref id="B10-sensors-21-06999"><label>10.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Redmon</surname><given-names>J.</given-names></name>
<name name-style="western"><surname>Farhadi</surname><given-names>A.</given-names></name>
</person-group><article-title>Yolov3: An incremental improvement</article-title><source>arXiv</source><year>2019</year><pub-id pub-id-type="arxiv">1804.02767</pub-id></element-citation></ref><ref id="B11-sensors-21-06999"><label>11.</label><element-citation publication-type="confproc"><person-group person-group-type="author">
<name name-style="western"><surname>He</surname><given-names>K.</given-names></name>
<name name-style="western"><surname>Zhang</surname><given-names>X.</given-names></name>
<name name-style="western"><surname>Ren</surname><given-names>S.</given-names></name>
<name name-style="western"><surname>Sun</surname><given-names>J.</given-names></name>
</person-group><article-title>Deep residual learning for image recognition</article-title><source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source><conf-loc>Las Vegas, NV, USA</conf-loc><conf-date>27–30 June 2016</conf-date><fpage>770</fpage><lpage>778</lpage></element-citation></ref><ref id="B12-sensors-21-06999"><label>12.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Dunn</surname><given-names>G.M.</given-names></name>
<name name-style="western"><surname>Martin</surname><given-names>S.R.</given-names></name>
</person-group><article-title>Yield prediction from digital image analysis: A technique with potential for vineyard assessments prior to harvest</article-title><source>Aust. J. Grape Wine Res.</source><year>2004</year><volume>10</volume><fpage>196</fpage><lpage>198</lpage><pub-id pub-id-type="doi">10.1111/j.1755-0238.2004.tb00022.x</pub-id></element-citation></ref><ref id="B13-sensors-21-06999"><label>13.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Payne</surname><given-names>A.B.</given-names></name>
<name name-style="western"><surname>Walsh</surname><given-names>K.B.</given-names></name>
<name name-style="western"><surname>Subedi</surname><given-names>P.</given-names></name>
<name name-style="western"><surname>Jarvis</surname><given-names>D.</given-names></name>
</person-group><article-title>Estimation of mango crop yield using image analysis–segmentation method</article-title><source>Comput. Electron. Agric.</source><year>2013</year><volume>91</volume><fpage>57</fpage><lpage>64</lpage><pub-id pub-id-type="doi">10.1016/j.compag.2012.11.009</pub-id></element-citation></ref><ref id="B14-sensors-21-06999"><label>14.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Dorj</surname><given-names>U.O.</given-names></name>
<name name-style="western"><surname>Lee</surname><given-names>M.</given-names></name>
<name name-style="western"><surname>Yun</surname><given-names>S.S.</given-names></name>
</person-group><article-title>An yield estimation in citrus orchards via fruit detection and counting using image processing</article-title><source>Comput. Electron. Agric.</source><year>2017</year><volume>140</volume><fpage>103</fpage><lpage>112</lpage><pub-id pub-id-type="doi">10.1016/j.compag.2017.05.019</pub-id></element-citation></ref><ref id="B15-sensors-21-06999"><label>15.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Nuske</surname><given-names>S.</given-names></name>
<name name-style="western"><surname>Wilshusen</surname><given-names>K.</given-names></name>
<name name-style="western"><surname>Achar</surname><given-names>S.</given-names></name>
<name name-style="western"><surname>Yoder</surname><given-names>L.</given-names></name>
<name name-style="western"><surname>Narasimhan</surname><given-names>S.</given-names></name>
<name name-style="western"><surname>Singh</surname><given-names>S.</given-names></name>
</person-group><article-title>Automated visual yield estimation in vineyards</article-title><source>J. Field Robot.</source><year>2014</year><volume>31</volume><fpage>837</fpage><lpage>860</lpage><pub-id pub-id-type="doi">10.1002/rob.21541</pub-id></element-citation></ref><ref id="B16-sensors-21-06999"><label>16.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Yamamoto</surname><given-names>K.</given-names></name>
<name name-style="western"><surname>Guo</surname><given-names>W.</given-names></name>
<name name-style="western"><surname>Yoshioka</surname><given-names>Y.</given-names></name>
<name name-style="western"><surname>Ninomiya</surname><given-names>S.</given-names></name>
</person-group><article-title>On plant detection of intact tomato fruits using image analysis and machine learning methods</article-title><source>Sensors</source><year>2014</year><volume>14</volume><fpage>12191</fpage><lpage>12206</lpage><pub-id pub-id-type="doi">10.3390/s140712191</pub-id><pub-id pub-id-type="pmid">25010694</pub-id><pub-id pub-id-type="pmcid">PMC4168514</pub-id></element-citation></ref><ref id="B17-sensors-21-06999"><label>17.</label><element-citation publication-type="confproc"><person-group person-group-type="author">
<name name-style="western"><surname>Hung</surname><given-names>C.</given-names></name>
<name name-style="western"><surname>Nieto</surname><given-names>J.</given-names></name>
<name name-style="western"><surname>Taylor</surname><given-names>Z.</given-names></name>
<name name-style="western"><surname>Underwood</surname><given-names>J.</given-names></name>
<name name-style="western"><surname>Sukkarieh</surname><given-names>S.</given-names></name>
</person-group><article-title>Orchard fruit segmentation using multi-spectral feature learning</article-title><source>Proceedings of the 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems</source><conf-loc>Tokyo, Japan</conf-loc><conf-date>3–7 November 2013</conf-date><fpage>5314</fpage><lpage>5320</lpage></element-citation></ref><ref id="B18-sensors-21-06999"><label>18.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Murillo-Bracamontes</surname><given-names>E.A.</given-names></name>
<name name-style="western"><surname>Martinez-Rosas</surname><given-names>M.E.</given-names></name>
<name name-style="western"><surname>Miranda-Velasco</surname><given-names>M.M.</given-names></name>
<name name-style="western"><surname>Martinez-Reyes</surname><given-names>H.L.</given-names></name>
<name name-style="western"><surname>Martinez-Sandoval</surname><given-names>J.R.</given-names></name>
<name name-style="western"><surname>Cervantes-de Avila</surname><given-names>H.</given-names></name>
</person-group><article-title>Implementation of Hough transform for fruit image segmentation</article-title><source>Procedia Eng.</source><year>2012</year><volume>35</volume><fpage>230</fpage><lpage>239</lpage><pub-id pub-id-type="doi">10.1016/j.proeng.2012.04.185</pub-id></element-citation></ref><ref id="B19-sensors-21-06999"><label>19.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Omid</surname><given-names>M.</given-names></name>
<name name-style="western"><surname>Khojastehnazhand</surname><given-names>M.</given-names></name>
<name name-style="western"><surname>Tabatabaeefar</surname><given-names>A.</given-names></name>
</person-group><article-title>Estimating volume and mass of citrus fruits by image processing technique</article-title><source>J. Food Eng.</source><year>2010</year><volume>100</volume><fpage>315</fpage><lpage>321</lpage><pub-id pub-id-type="doi">10.1016/j.jfoodeng.2010.04.015</pub-id></element-citation></ref><ref id="B20-sensors-21-06999"><label>20.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Z.</given-names></name>
<name name-style="western"><surname>Koirala</surname><given-names>A.</given-names></name>
<name name-style="western"><surname>Walsh</surname><given-names>K.</given-names></name>
<name name-style="western"><surname>Anderson</surname><given-names>N.</given-names></name>
<name name-style="western"><surname>Verma</surname><given-names>B.</given-names></name>
</person-group><article-title>In field fruit sizing using a smart phone application</article-title><source>Sensors</source><year>2018</year><volume>18</volume><elocation-id>3331</elocation-id><pub-id pub-id-type="doi">10.3390/s18103331</pub-id><pub-id pub-id-type="pmid">30301141</pub-id><pub-id pub-id-type="pmcid">PMC6210956</pub-id></element-citation></ref><ref id="B21-sensors-21-06999"><label>21.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Mizushima</surname><given-names>A.</given-names></name>
<name name-style="western"><surname>Lu</surname><given-names>R.</given-names></name>
</person-group><article-title>An image segmentation method for apple sorting and grading using support vector machine and Otsu’s method</article-title><source>Comput. Electron. Agric.</source><year>2013</year><volume>94</volume><fpage>29</fpage><lpage>37</lpage><pub-id pub-id-type="doi">10.1016/j.compag.2013.02.009</pub-id></element-citation></ref><ref id="B22-sensors-21-06999"><label>22.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Ren</surname><given-names>S.</given-names></name>
<name name-style="western"><surname>He</surname><given-names>K.</given-names></name>
<name name-style="western"><surname>Girshick</surname><given-names>R.</given-names></name>
<name name-style="western"><surname>Sun</surname><given-names>J.</given-names></name>
</person-group><article-title>Faster r-cnn: Towards real-time object detection with region proposal networks</article-title><source>Adv. Neural Inf. Process. Syst.</source><year>2015</year><volume>28</volume><fpage>91</fpage><lpage>99</lpage><pub-id pub-id-type="doi">10.1109/TPAMI.2016.2577031</pub-id><pub-id pub-id-type="pmid">27295650</pub-id></element-citation></ref><ref id="B23-sensors-21-06999"><label>23.</label><element-citation publication-type="confproc"><person-group person-group-type="author">
<name name-style="western"><surname>Bargoti</surname><given-names>S.</given-names></name>
<name name-style="western"><surname>Underwood</surname><given-names>J.</given-names></name>
</person-group><article-title>Deep fruit detection in orchards</article-title><source>Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA)</source><conf-loc>Singapore</conf-loc><conf-date>29 May–3 June 2017</conf-date><fpage>3626</fpage><lpage>3633</lpage></element-citation></ref><ref id="B24-sensors-21-06999"><label>24.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Liu</surname><given-names>X.</given-names></name>
<name name-style="western"><surname>Chen</surname><given-names>S.W.</given-names></name>
<name name-style="western"><surname>Liu</surname><given-names>C.</given-names></name>
<name name-style="western"><surname>Shivakumar</surname><given-names>S.S.</given-names></name>
<name name-style="western"><surname>Das</surname><given-names>J.</given-names></name>
<name name-style="western"><surname>Taylor</surname><given-names>C.J.</given-names></name>
<name name-style="western"><surname>Underwood</surname><given-names>J.</given-names></name>
<name name-style="western"><surname>Kumar</surname><given-names>V.</given-names></name>
</person-group><article-title>Monocular Camera Based Fruit Counting and Mapping with Semantic Data Association</article-title><source>IEEE Robot. Autom. Lett.</source><year>2019</year><volume>4</volume><fpage>2296</fpage><lpage>2303</lpage><pub-id pub-id-type="doi">10.1109/LRA.2019.2901987</pub-id></element-citation></ref><ref id="B25-sensors-21-06999"><label>25.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Sa</surname><given-names>I.</given-names></name>
<name name-style="western"><surname>Ge</surname><given-names>Z.</given-names></name>
<name name-style="western"><surname>Dayoub</surname><given-names>F.</given-names></name>
<name name-style="western"><surname>Upcroft</surname><given-names>B.</given-names></name>
<name name-style="western"><surname>Perez</surname><given-names>T.</given-names></name>
<name name-style="western"><surname>McCool</surname><given-names>C.</given-names></name>
</person-group><article-title>Deepfruits: A fruit detection system using deep neural networks</article-title><source>Sensors</source><year>2016</year><volume>16</volume><elocation-id>1222</elocation-id><pub-id pub-id-type="doi">10.3390/s16081222</pub-id><pub-id pub-id-type="pmcid">PMC5017387</pub-id><pub-id pub-id-type="pmid">27527168</pub-id></element-citation></ref><ref id="B26-sensors-21-06999"><label>26.</label><element-citation publication-type="confproc"><person-group person-group-type="author">
<name name-style="western"><surname>Huang</surname><given-names>Z.</given-names></name>
<name name-style="western"><surname>Huang</surname><given-names>L.</given-names></name>
<name name-style="western"><surname>Gong</surname><given-names>Y.</given-names></name>
<name name-style="western"><surname>Huang</surname><given-names>C.</given-names></name>
<name name-style="western"><surname>Wang</surname><given-names>X.</given-names></name>
</person-group><article-title>Mask scoring r-cnn</article-title><source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source><conf-loc>Long Beach, CA, USA</conf-loc><conf-date>16–20 June 2019</conf-date><fpage>6409</fpage><lpage>6418</lpage></element-citation></ref><ref id="B27-sensors-21-06999"><label>27.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Santos</surname><given-names>T.T.</given-names></name>
<name name-style="western"><surname>de Souza</surname><given-names>L.L.</given-names></name>
<name name-style="western"><surname>dos Santos</surname><given-names>A.A.</given-names></name>
<name name-style="western"><surname>Avila</surname><given-names>S.</given-names></name>
</person-group><article-title>Grape detection, segmentation, and tracking using deep neural networks and three-dimensional association</article-title><source>Comput. Electron. Agric.</source><year>2020</year><volume>170</volume><fpage>105247</fpage><pub-id pub-id-type="doi">10.1016/j.compag.2020.105247</pub-id></element-citation></ref><ref id="B28-sensors-21-06999"><label>28.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Ni</surname><given-names>X.</given-names></name>
<name name-style="western"><surname>Li</surname><given-names>C.</given-names></name>
<name name-style="western"><surname>Jiang</surname><given-names>H.</given-names></name>
<name name-style="western"><surname>Takeda</surname><given-names>F.</given-names></name>
</person-group><article-title>Deep learning image segmentation and extraction of blueberry fruit traits associated with harvestability and yield</article-title><source>Hortic. Res.</source><year>2020</year><volume>7</volume><fpage>1</fpage><lpage>14</lpage><pub-id pub-id-type="doi">10.1038/s41438-020-0323-3</pub-id><pub-id pub-id-type="pmid">32637138</pub-id><pub-id pub-id-type="pmcid">PMC7326978</pub-id></element-citation></ref><ref id="B29-sensors-21-06999"><label>29.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Rahnemoonfar</surname><given-names>M.</given-names></name>
<name name-style="western"><surname>Sheppard</surname><given-names>C.</given-names></name>
</person-group><article-title>Deep count: Fruit counting based on deep simulated learning</article-title><source>Sensors</source><year>2017</year><volume>17</volume><elocation-id>905</elocation-id><pub-id pub-id-type="doi">10.3390/s17040905</pub-id><pub-id pub-id-type="pmid">28425947</pub-id><pub-id pub-id-type="pmcid">PMC5426829</pub-id></element-citation></ref><ref id="B30-sensors-21-06999"><label>30.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Bargoti</surname><given-names>S.</given-names></name>
<name name-style="western"><surname>Underwood</surname><given-names>J.P.</given-names></name>
</person-group><article-title>Image segmentation for fruit detection and yield estimation in apple orchards</article-title><source>J. Field Robot.</source><year>2017</year><volume>34</volume><fpage>1039</fpage><lpage>1060</lpage><pub-id pub-id-type="doi">10.1002/rob.21699</pub-id></element-citation></ref><ref id="B31-sensors-21-06999"><label>31.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Chen</surname><given-names>S.W.</given-names></name>
<name name-style="western"><surname>Shivakumar</surname><given-names>S.S.</given-names></name>
<name name-style="western"><surname>Dcunha</surname><given-names>S.</given-names></name>
<name name-style="western"><surname>Das</surname><given-names>J.</given-names></name>
<name name-style="western"><surname>Okon</surname><given-names>E.</given-names></name>
<name name-style="western"><surname>Qu</surname><given-names>C.</given-names></name>
<name name-style="western"><surname>Taylor</surname><given-names>C.J.</given-names></name>
<name name-style="western"><surname>Kumar</surname><given-names>V.</given-names></name>
</person-group><article-title>Counting apples and oranges with deep learning: A data-driven approach</article-title><source>IEEE Robot. Autom. Lett.</source><year>2017</year><volume>2</volume><fpage>781</fpage><lpage>788</lpage><pub-id pub-id-type="doi">10.1109/LRA.2017.2651944</pub-id></element-citation></ref><ref id="B32-sensors-21-06999"><label>32.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Kamilaris</surname><given-names>A.</given-names></name>
<name name-style="western"><surname>Prenafeta-Boldú</surname><given-names>F.X.</given-names></name>
</person-group><article-title>Deep learning in agriculture: A survey</article-title><source>Comput. Electron. Agric.</source><year>2018</year><volume>147</volume><fpage>70</fpage><lpage>90</lpage><pub-id pub-id-type="doi">10.1016/j.compag.2018.02.016</pub-id></element-citation></ref><ref id="B33-sensors-21-06999"><label>33.</label><element-citation publication-type="confproc"><person-group person-group-type="author">
<name name-style="western"><surname>Selvaraju</surname><given-names>R.R.</given-names></name>
<name name-style="western"><surname>Cogswell</surname><given-names>M.</given-names></name>
<name name-style="western"><surname>Das</surname><given-names>A.</given-names></name>
<name name-style="western"><surname>Vedantam</surname><given-names>R.</given-names></name>
<name name-style="western"><surname>Parikh</surname><given-names>D.</given-names></name>
<name name-style="western"><surname>Batra</surname><given-names>D.</given-names></name>
</person-group><article-title>Grad-cam: Visual explanations from deep networks via gradient-based localization</article-title><source>Proceedings of the IEEE International Conference on Computer Vision</source><conf-loc>Venice, Italy</conf-loc><conf-date>22–29 October 2017</conf-date><fpage>618</fpage><lpage>626</lpage></element-citation></ref><ref id="B34-sensors-21-06999"><label>34.</label><element-citation publication-type="book"><person-group person-group-type="author">
<name name-style="western"><surname>Goodfellow</surname><given-names>I.</given-names></name>
<name name-style="western"><surname>Bengio</surname><given-names>Y.</given-names></name>
<name name-style="western"><surname>Courville</surname><given-names>A.</given-names></name>
</person-group><source>Deep Learning</source><publisher-name>MIT Press</publisher-name><publisher-loc>Cambridge, MA, USA</publisher-loc><year>2016</year></element-citation></ref><ref id="B35-sensors-21-06999"><label>35.</label><element-citation publication-type="book"><person-group person-group-type="author">
<name name-style="western"><surname>Ronneberger</surname><given-names>O.</given-names></name>
<name name-style="western"><surname>Fischer</surname><given-names>P.</given-names></name>
<name name-style="western"><surname>Brox</surname><given-names>T.</given-names></name>
</person-group><article-title>U-net: Convolutional networks for biomedical image segmentation</article-title><source>International Conference on Medical Image Computing and Computer-Assisted Intervention</source><publisher-name>Springer</publisher-name><publisher-loc>Cham, Switzerland</publisher-loc><year>2015</year><fpage>234</fpage><lpage>241</lpage></element-citation></ref><ref id="B36-sensors-21-06999"><label>36.</label><element-citation publication-type="confproc"><person-group person-group-type="author">
<name name-style="western"><surname>Milletari</surname><given-names>F.</given-names></name>
<name name-style="western"><surname>Navab</surname><given-names>N.</given-names></name>
<name name-style="western"><surname>Ahmadi</surname><given-names>S.A.</given-names></name>
</person-group><article-title>V-net: Fully convolutional neural networks for volumetric medical image segmentation</article-title><source>Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV)</source><conf-loc>Stanford, CA, USA</conf-loc><conf-date>25–28 October 2016</conf-date><fpage>565</fpage><lpage>571</lpage></element-citation></ref><ref id="B37-sensors-21-06999"><label>37.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Fukushima</surname><given-names>K.</given-names></name>
</person-group><article-title>Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position</article-title><source>Biol. Cybern.</source><year>1980</year><volume>36</volume><fpage>193</fpage><lpage>202</lpage><pub-id pub-id-type="doi">10.1007/BF00344251</pub-id><pub-id pub-id-type="pmid">7370364</pub-id></element-citation></ref><ref id="B38-sensors-21-06999"><label>38.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>LeCun</surname><given-names>Y.</given-names></name>
<name name-style="western"><surname>Bottou</surname><given-names>L.</given-names></name>
<name name-style="western"><surname>Bengio</surname><given-names>Y.</given-names></name>
<name name-style="western"><surname>Haffner</surname><given-names>P.</given-names></name>
</person-group><article-title>Gradient-based learning applied to document recognition</article-title><source>Proc. IEEE</source><year>1998</year><volume>86</volume><fpage>2278</fpage><lpage>2324</lpage><pub-id pub-id-type="doi">10.1109/5.726791</pub-id></element-citation></ref><ref id="B39-sensors-21-06999"><label>39.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Krizhevsky</surname><given-names>A.</given-names></name>
<name name-style="western"><surname>Sutskever</surname><given-names>I.</given-names></name>
<name name-style="western"><surname>Hinton</surname><given-names>G.E.</given-names></name>
</person-group><article-title>Imagenet classification with deep convolutional neural networks</article-title><source>Adv. Neural Inf. Process. Syst.</source><year>2012</year><volume>25</volume><fpage>1097</fpage><lpage>1105</lpage><pub-id pub-id-type="doi">10.1145/3065386</pub-id></element-citation></ref><ref id="B40-sensors-21-06999"><label>40.</label><element-citation publication-type="confproc"><person-group person-group-type="author">
<name name-style="western"><surname>Long</surname><given-names>J.</given-names></name>
<name name-style="western"><surname>Shelhamer</surname><given-names>E.</given-names></name>
<name name-style="western"><surname>Darrell</surname><given-names>T.</given-names></name>
</person-group><article-title>Fully convolutional networks for semantic segmentation</article-title><source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source><conf-loc>Boston, MA, USA</conf-loc><conf-date>7–12 June 2015</conf-date><fpage>3431</fpage><lpage>3440</lpage></element-citation></ref><ref id="B41-sensors-21-06999"><label>41.</label><element-citation publication-type="confproc"><person-group person-group-type="author">
<name name-style="western"><surname>Noh</surname><given-names>H.</given-names></name>
<name name-style="western"><surname>Hong</surname><given-names>S.</given-names></name>
<name name-style="western"><surname>Han</surname><given-names>B.</given-names></name>
</person-group><article-title>Learning deconvolution network for semantic segmentation</article-title><source>Proceedings of the IEEE International Conference on Computer Vision</source><conf-loc>Santiago, Chile</conf-loc><conf-date>7–13 December 2015</conf-date><fpage>1520</fpage><lpage>1528</lpage></element-citation></ref><ref id="B42-sensors-21-06999"><label>42.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Badrinarayanan</surname><given-names>V.</given-names></name>
<name name-style="western"><surname>Kendall</surname><given-names>A.</given-names></name>
<name name-style="western"><surname>Cipolla</surname><given-names>R.</given-names></name>
</person-group><article-title>Segnet: A deep convolutional encoder-decoder architecture for image segmentation</article-title><source>IEEE Trans. Pattern Anal. Mach. Intell.</source><year>2017</year><volume>39</volume><fpage>2481</fpage><lpage>2495</lpage><pub-id pub-id-type="doi">10.1109/TPAMI.2016.2644615</pub-id><pub-id pub-id-type="pmid">28060704</pub-id></element-citation></ref><ref id="B43-sensors-21-06999"><label>43.</label><element-citation publication-type="confproc"><person-group person-group-type="author">
<name name-style="western"><surname>Chen</surname><given-names>L.C.</given-names></name>
<name name-style="western"><surname>Zhu</surname><given-names>Y.</given-names></name>
<name name-style="western"><surname>Papandreou</surname><given-names>G.</given-names></name>
<name name-style="western"><surname>Schroff</surname><given-names>F.</given-names></name>
<name name-style="western"><surname>Adam</surname><given-names>H.</given-names></name>
</person-group><article-title>Encoder-decoder with atrous separable convolution for semantic image segmentation</article-title><source>Proceedings of the European Conference on Computer Vision (ECCV)</source><conf-loc>Munich, Germany</conf-loc><conf-date>8–14 September 2018</conf-date><fpage>801</fpage><lpage>818</lpage></element-citation></ref><ref id="B44-sensors-21-06999"><label>44.</label><element-citation publication-type="webpage"><person-group person-group-type="author">
<name name-style="western"><surname>Wada</surname><given-names>K.</given-names></name>
</person-group><article-title>labelme: Image Polygonal Annotation with Python</article-title><year>2016</year><comment>Available online: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/wkentaro/labelme" ext-link-type="uri">https://github.com/wkentaro/labelme</ext-link></comment><date-in-citation content-type="access-date" iso-8601-date="2021-10-20">(accessed on 20 October 2021)</date-in-citation></element-citation></ref><ref id="B45-sensors-21-06999"><label>45.</label><element-citation publication-type="book"><person-group person-group-type="author">
<name name-style="western"><surname>Paszke</surname><given-names>A.</given-names></name>
<name name-style="western"><surname>Gross</surname><given-names>S.</given-names></name>
<name name-style="western"><surname>Massa</surname><given-names>F.</given-names></name>
<name name-style="western"><surname>Lerer</surname><given-names>A.</given-names></name>
<name name-style="western"><surname>Bradbury</surname><given-names>J.</given-names></name>
<name name-style="western"><surname>Chanan</surname><given-names>G.</given-names></name>
<name name-style="western"><surname>Killeen</surname><given-names>T.</given-names></name>
<name name-style="western"><surname>Lin</surname><given-names>Z.</given-names></name>
<name name-style="western"><surname>Gimelshein</surname><given-names>N.</given-names></name>
<name name-style="western"><surname>Antiga</surname><given-names>L.</given-names></name>
<etal/>
</person-group><article-title>PyTorch: An Imperative Style, High-Performance Deep Learning Library</article-title><source>Advances in Neural Information Processing Systems 32</source><person-group person-group-type="editor">
<name name-style="western"><surname>Wallach</surname><given-names>H.</given-names></name>
<name name-style="western"><surname>Larochelle</surname><given-names>H.</given-names></name>
<name name-style="western"><surname>Beygelzimer</surname><given-names>A.</given-names></name>
<name name-style="western"><surname>d’Alché-Buc</surname><given-names>F.</given-names></name>
<name name-style="western"><surname>Fox</surname><given-names>E.</given-names></name>
<name name-style="western"><surname>Garnett</surname><given-names>R.</given-names></name>
</person-group><publisher-name>Curran Associates, Inc.</publisher-name><publisher-loc>Red Hook, NY, USA</publisher-loc><year>2019</year><fpage>8024</fpage><lpage>8035</lpage></element-citation></ref><ref id="B46-sensors-21-06999"><label>46.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Hunter</surname><given-names>J.D.</given-names></name>
</person-group><article-title>Matplotlib: A 2D graphics environment</article-title><source>Comput. Sci. Eng.</source><year>2007</year><volume>9</volume><fpage>90</fpage><lpage>95</lpage><pub-id pub-id-type="doi">10.1109/MCSE.2007.55</pub-id></element-citation></ref><ref id="B47-sensors-21-06999"><label>47.</label><element-citation publication-type="journal"><person-group person-group-type="author">
<name name-style="western"><surname>Waskom</surname><given-names>M.L.</given-names></name>
</person-group><article-title>seaborn: Statistical data visualization</article-title><source>J. Open Source Softw.</source><year>2021</year><volume>6</volume><fpage>3021</fpage><pub-id pub-id-type="doi">10.21105/joss.03021</pub-id></element-citation></ref><ref id="B48-sensors-21-06999"><label>48.</label><element-citation publication-type="confproc"><person-group person-group-type="author">
<name name-style="western"><surname>Li</surname><given-names>Y.</given-names></name>
<name name-style="western"><surname>Hou</surname><given-names>X.</given-names></name>
<name name-style="western"><surname>Koch</surname><given-names>C.</given-names></name>
<name name-style="western"><surname>Rehg</surname><given-names>J.M.</given-names></name>
<name name-style="western"><surname>Yuille</surname><given-names>A.L.</given-names></name>
</person-group><article-title>The secrets of salient object segmentation</article-title><source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source><conf-loc>Columbus, OH, USA</conf-loc><conf-date>23–28 June 2014</conf-date><fpage>280</fpage><lpage>287</lpage></element-citation></ref></ref-list></back><floats-group><fig position="float" id="sensors-21-06999-f001" orientation="portrait"><label>Figure 1</label><caption><p><monospace>CROP</monospace> identifies and paints central fruit of various kinds. These images were not used for the training or validation, and are considered to be test data. (<bold>a</bold>) pawpaw. (<bold>b</bold>) longan. (<bold>c</bold>) persimmon. (<bold>d</bold>) pear. The original photos of (<bold>a</bold>,<bold>b</bold>) are credited to USDA ARS (Scott Bauer), and those of (<bold>c</bold>,<bold>d</bold>) Hideki Murayama.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g001.jpg"><?image-name sensors-21-06999-g001.jpg?><?image-size 50946?><?image-md5 1c8436f6619a721753f45bea8815ff58?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 560?><?image-original-width 3131?><?image-scaled-height 140?><?image-scaled-width 782?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/1c8436f6619a/sensors-21-06999-g001.jpg?><?thumb-name sensors-21-06999-g001.gif?><?thumb-size 9038?><?thumb-md5 4c3f3f49511a0c90adb1259d5c5c8268?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 36?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/4c3f3f49511a/sensors-21-06999-g001.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f002" orientation="portrait"><label>Figure 2</label><caption><p>The architectures the deep neural networks. (<bold>a</bold>) <monospace>CROP</monospace>. (<bold>b</bold>) <monospace>CROP-Shallow</monospace>. (<bold>c</bold>) Notations. <monospace>ReLU</monospace> and <italic toggle="yes">batch normalization</italic> are applied adequately, which are not explicit in the figure.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g002.jpg"><?image-name sensors-21-06999-g002.jpg?><?image-size 69834?><?image-md5 dbb18aeedd900f8f01948ee8bb069bfd?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 1528?><?image-original-width 2942?><?image-scaled-height 382?><?image-scaled-width 735?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/dbb18aeedd90/sensors-21-06999-g002.jpg?><?thumb-name sensors-21-06999-g002.gif?><?thumb-size 7259?><?thumb-md5 9c2b3dd88c62df0ed0887b9da0befebe?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 80?><?thumb-scaled-width 154?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/9c2b3dd88c62/sensors-21-06999-g002.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f003" orientation="portrait"><label>Figure 3</label><caption><p>Training CROP with the training dataset (80%: 137) and the validation dataset (20%: 35) of Data_Fruit, so that the best IoU for the validation data was 0.984 at epoch 9700. (<bold>a</bold>) Losses for epoch 100–10,000. (<bold>b</bold>) IoUs for epoch 100–10,000. (<bold>c</bold>) Losses for the last half; epoch 5100–10,000. (<bold>d</bold>) IoUs for the last half; epoch 5100–10,000.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g003.jpg"><?image-name sensors-21-06999-g003.jpg?><?image-size 91740?><?image-md5 ab2b26bbe59a3b112db0543dd6c0245c?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 2444?><?image-original-width 3179?><?image-scaled-height 610?><?image-scaled-width 794?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/ab2b26bbe59a/sensors-21-06999-g003.jpg?><?thumb-name sensors-21-06999-g003.gif?><?thumb-size 5720?><?thumb-md5 c34d4bdfc929fe4db9e0a26e370fed8a?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 80?><?thumb-scaled-width 104?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/c34d4bdfc929/sensors-21-06999-g003.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f004" orientation="portrait"><label>Figure 4</label><caption><p>Fine-tuning CROP with the training dataset (80%: 68) and the validation dataset (20%: 18) of Data_Pears2, so that the best IoU for the validation data was 0.983 achieved at epoch 5100. (<bold>a</bold>) Losses for epoch 100–10,000. (<bold>b</bold>) IoUs for epoch 100–10,000.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g004.jpg"><?image-name sensors-21-06999-g004.jpg?><?image-size 56651?><?image-md5 a6ec109d1ee1497960bdb44b3e3b8b0c?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 1197?><?image-original-width 3168?><?image-scaled-height 299?><?image-scaled-width 792?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/a6ec109d1ee1/sensors-21-06999-g004.jpg?><?thumb-name sensors-21-06999-g004.gif?><?thumb-size 7136?><?thumb-md5 98267b7749f5297adc866cd53307ef16?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 76?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/98267b7749f5/sensors-21-06999-g004.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f005" orientation="portrait"><label>Figure 5</label><caption><p>Comparing the training processes of CROP and CROP-Shallow through three experiments for each, numbered from 0 to 2. They were conducted for 2500 epochs with the training dataset (80%: 137) and the validation dataset (20%: 35) of Data_Fruit. (<bold>a</bold>) Losses for epoch 100–2500. (<bold>b</bold>) IoUs for epoch 100–2500.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g005.jpg"><?image-name sensors-21-06999-g005.jpg?><?image-size 111080?><?image-md5 4ad41d35e674d4d5ef88d47b4cc1e41e?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 1691?><?image-original-width 3175?><?image-scaled-height 422?><?image-scaled-width 793?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/4ad41d35e674/sensors-21-06999-g005.jpg?><?thumb-name sensors-21-06999-g005.gif?><?thumb-size 8499?><?thumb-md5 2e7b6eb97c7cbb23d184378c13c1af48?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 80?><?thumb-scaled-width 150?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/2e7b6eb97c7c/sensors-21-06999-g005.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f006" orientation="portrait"><label>Figure 6</label><caption><p>Identifying the central grapes in the bunch of grapes. (<bold>a</bold>) The central grape is identified. (<bold>b</bold>) The central grape in another frame is identified. (<bold>c</bold>) CROP is confused because there is no central fruit in front. (<bold>d</bold>) The central object can be small, and may not have to be salient. The original photo is credited to USDA ARS (Peggy Greb).</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g006.jpg"><?image-name sensors-21-06999-g006.jpg?><?image-size 42767?><?image-md5 616dd9f4be9a9776103664cbe00fd80c?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 564?><?image-original-width 3135?><?image-scaled-height 141?><?image-scaled-width 783?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/616dd9f4be9a/sensors-21-06999-g006.jpg?><?thumb-name sensors-21-06999-g006.gif?><?thumb-size 8948?><?thumb-md5 1e9c756398814f968fe98ea8c0d7b101?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 36?><?thumb-scaled-width 199?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/1e9c75639881/sensors-21-06999-g006.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f007" orientation="portrait"><label>Figure 7</label><caption><p>Examples (<bold>a</bold>–<bold>h</bold>), where peduncle and calyx are properly ignored. Some of the original photos are credited to USDA ARS; (<bold>a</bold>) to Peggy Greb, (<bold>b</bold>) to Mark Ehlenfeldt, (<bold>c</bold>) to Brian Prechtel. Moreover, the original photos of (<bold>e</bold>–<bold>h</bold>) are credited to Hideki Murayama.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g007.jpg"><?image-name sensors-21-06999-g007.jpg?><?image-size 93568?><?image-md5 527a61c019c988f3249dfa0d43d8008c?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 1182?><?image-original-width 3139?><?image-scaled-height 295?><?image-scaled-width 784?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/527a61c019c9/sensors-21-06999-g007.jpg?><?thumb-name sensors-21-06999-g007.gif?><?thumb-size 13482?><?thumb-md5 c424ae390b031c9581051fa2dc6ec2f3?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 75?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/c424ae390b03/sensors-21-06999-g007.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f008" orientation="portrait"><label>Figure 8</label><caption><p>Unsuccessful examples. (<bold>a</bold>) The central fruit is not round enough. (<bold>b</bold>) Two boundaries were mixed up. (<bold>c</bold>) The central object is disrupted by the leaf. (<bold>d</bold>) The boundary is spiky. The original photo of (<bold>a</bold>) is credited to USDA ARS (Keith Weller). The original photos of (<bold>b</bold>–<bold>d</bold>) are credited to Hideki Murayama.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g008.jpg"><?image-name sensors-21-06999-g008.jpg?><?image-size 51921?><?image-md5 7c8957b3ad4b2b8a16c0df5e1285750a?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 557?><?image-original-width 3139?><?image-scaled-height 139?><?image-scaled-width 784?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/7c8957b3ad4b/sensors-21-06999-g008.jpg?><?thumb-name sensors-21-06999-g008.gif?><?thumb-size 8996?><?thumb-md5 baea566567cf87eec7163e0fc7b3ef54?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 35?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/baea566567cf/sensors-21-06999-g008.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f009" orientation="portrait"><label>Figure 9</label><caption><p>Gained generality and works for various objects. (<bold>a</bold>) A garlic. (<bold>b</bold>) A loaf of bread. (<bold>c</bold>) A roasted coffee bean. (<bold>d</bold>) A stone. The original photo of (<bold>a</bold>) is credited to USDA ARS (Scott Bauer).</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g009.jpg"><?image-name sensors-21-06999-g009.jpg?><?image-size 51598?><?image-md5 3ad8c1f14b60db20f92c9b294eef6527?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 560?><?image-original-width 3131?><?image-scaled-height 140?><?image-scaled-width 782?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/3ad8c1f14b60/sensors-21-06999-g009.jpg?><?thumb-name sensors-21-06999-g009.gif?><?thumb-size 9247?><?thumb-md5 da6e8ab236b389707ed315fb77775c2d?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 36?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/da6e8ab236b3/sensors-21-06999-g009.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f010" orientation="portrait"><label>Figure 10</label><caption><p>Before and after the fine-tuning. (<bold>a</bold>) This image was properly processed even before the fine-tuning. (<bold>b</bold>) The bottom contour line became more precise. (<bold>c</bold>) The right boundary became more accurate. (<bold>d</bold>) A good amount of improvement.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g010.jpg"><?image-name sensors-21-06999-g010.jpg?><?image-size 98040?><?image-md5 9c32de7a36e12bbec5ce4c213f258196?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 1568?><?image-original-width 3179?><?image-scaled-height 392?><?image-scaled-width 794?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/9c32de7a36e1/sensors-21-06999-g010.jpg?><?thumb-name sensors-21-06999-g010.gif?><?thumb-size 13139?><?thumb-md5 a1e55249c0b3b9e1a23f7538dd56d02a?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 80?><?thumb-scaled-width 162?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/a1e55249c0b3/sensors-21-06999-g010.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f013" orientation="portrait"><label>Figure 13</label><caption><p>The number of pixels of the masks predicted by <monospace>CROP</monospace> for the period 12 August–15 October 2020. Outliers are plotted inside and outside the range. The five days highlighted by cyan color is treated separately in <xref rid="sensors-21-06999-f014" ref-type="fig">Figure 14</xref>.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g013.jpg"><?image-name sensors-21-06999-g013.jpg?><?image-size 28195?><?image-md5 31129785206645f970c0b7d0a47281cf?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 920?><?image-original-width 3299?><?image-scaled-height 204?><?image-scaled-width 733?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/311297852066/sensors-21-06999-g013.jpg?><?thumb-name sensors-21-06999-g013.gif?><?thumb-size 5831?><?thumb-md5 318a1903f84d94b897168694d98235f3?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 56?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/318a1903f84d/sensors-21-06999-g013.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f014" orientation="portrait"><label>Figure 14</label><caption><p>The box-plot of the 11 measurements (see <xref rid="sensors-21-06999-f012" ref-type="fig">Figure 12</xref>a) during the period 8–12 October 2020. The variance is higher in the evening (17:49, 19:49, 21:49), indicated by the darker background. The higher the variance, the longer the box and the whisker, and the more individual the points. The horizontal line in each box represents the median. The images captured on the last day (from 487 to 494) are found in <xref rid="sensors-21-06999-f015" ref-type="fig">Figure 15</xref>.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g014.jpg"><?image-name sensors-21-06999-g014.jpg?><?image-size 32805?><?image-md5 799f7d5829c6330f79256175ca541b07?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 906?><?image-original-width 3295?><?image-scaled-height 201?><?image-scaled-width 732?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/799f7d5829c6/sensors-21-06999-g014.jpg?><?thumb-name sensors-21-06999-g014.gif?><?thumb-size 7756?><?thumb-md5 a5f5c7c7e233edd19cb51afe090d1c2f?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 55?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/a5f5c7c7e233/sensors-21-06999-g014.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f015" orientation="portrait"><label>Figure 15</label><caption><p>The eight images on 12 October 2020, the last day in <xref rid="sensors-21-06999-f014" ref-type="fig">Figure 14</xref>, whose image IDs are 487 to 494. These masks gave the medians appearing as part of the graph in <xref rid="sensors-21-06999-f013" ref-type="fig">Figure 13</xref> and in <xref rid="sensors-21-06999-f014" ref-type="fig">Figure 14</xref>. Additionally, the blue dots describe the relative positions, which constitute part of the dots in <xref rid="sensors-21-06999-f016" ref-type="fig">Figure 16</xref>.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g015.jpg"><?image-name sensors-21-06999-g015.jpg?><?image-size 70203?><?image-md5 5fe1b0ebc1b2f74417537ceb3823f80c?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 877?><?image-original-width 3328?><?image-scaled-height 195?><?image-scaled-width 739?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/5fe1b0ebc1b2/sensors-21-06999-g015.jpg?><?thumb-name sensors-21-06999-g015.gif?><?thumb-size 13373?><?thumb-md5 ac5949a241564c32d1a8aa63ed1a243d?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 53?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/ac5949a24156/sensors-21-06999-g015.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f016" orientation="portrait"><label>Figure 16</label><caption><p>The positions of the target fruit, predicted by CROP. The coordinates are pixel numbers in the original image (<xref rid="sensors-21-06999-f011" ref-type="fig">Figure 11</xref>a), counting from the top-left corner. (<bold>a</bold>) All the positions were plotted including possible outliers inside and outside (to the below of) the range. The darker the color is, the later the image was captured. (<bold>b</bold>) Visualization of the movement during the five days: 8–12 October 2020.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g016.jpg"><?image-name sensors-21-06999-g016.jpg?><?image-size 49668?><?image-md5 01a8422b0e4795eafeeb325e66b847ce?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 1204?><?image-original-width 3193?><?image-scaled-height 301?><?image-scaled-width 798?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/01a8422b0e47/sensors-21-06999-g016.jpg?><?thumb-name sensors-21-06999-g016.gif?><?thumb-size 8686?><?thumb-md5 f8ae647922bb73d842ecba3a25e2df78?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 75?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/f8ae647922bb/sensors-21-06999-g016.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f017" orientation="portrait"><label>Figure 17</label><caption><p>The histograms of CVs of the 11 measurements for individual images (<xref rid="sensors-21-06999-f012" ref-type="fig">Figure 12</xref>). (<bold>a</bold>) The 300 smaller CVs, which we consider as normal. (<bold>b</bold>) The 15 largest CVs, which we treat as the evidence for outliers.</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g017.jpg"><?image-name sensors-21-06999-g017.jpg?><?image-size 21804?><?image-md5 cfbc8fd3e9c1ae8bb75e1626e6d0b680?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 1008?><?image-original-width 3215?><?image-scaled-height 224?><?image-scaled-width 714?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/cfbc8fd3e9c1/sensors-21-06999-g017.jpg?><?thumb-name sensors-21-06999-g017.gif?><?thumb-size 6119?><?thumb-md5 3f489b235df265506c30a9c5ddb4f6a6?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 63?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/3f489b235df2/sensors-21-06999-g017.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f018" orientation="portrait"><label>Figure 18</label><caption><p>Examples of images with the highest CVs and a confident mistake. (<bold>a</bold>) This blurred image gave the highest CV (id: 381). (<bold>b</bold>) The misty weather condition resulted in the second highest CV (id: 399). (<bold>c</bold>) The unsuccessful case modified manually (id: 393).</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g018.jpg"><?image-name sensors-21-06999-g018.jpg?><?image-size 38977?><?image-md5 e02e8cc75659a207c74a8ab52e59d29d?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 713?><?image-original-width 3237?><?image-scaled-height 158?><?image-scaled-width 719?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/e02e8cc75659/sensors-21-06999-g018.jpg?><?thumb-name sensors-21-06999-g018.gif?><?thumb-size 10241?><?thumb-md5 d62959ef97ac77aa063fb6a012b62866?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 44?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/d62959ef97ac/sensors-21-06999-g018.gif?></graphic></fig><fig position="float" id="sensors-21-06999-f019" orientation="portrait"><label>Figure 19</label><caption><p>The plot of the modified data and the growth curve made by curve-fitting to 5th degree polynomials, with <monospace>Seaborn</monospace> [<xref rid="B47-sensors-21-06999" ref-type="bibr">47</xref>].</p></caption><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="sensors-21-06999-g019.jpg"><?image-name sensors-21-06999-g019.jpg?><?image-size 22539?><?image-md5 44ca846447dd8aa5fca2b2aaeb86a09a?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 928?><?image-original-width 3299?><?image-scaled-height 206?><?image-scaled-width 733?><?image-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/44ca846447dd/sensors-21-06999-g019.jpg?><?thumb-name sensors-21-06999-g019.gif?><?thumb-size 5483?><?thumb-md5 0c4ad009430261d62e4dc1618aca0284?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 56?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/0b7e/8586972/0c4ad0094302/sensors-21-06999-g019.gif?></graphic></fig><table-wrap position="float" id="sensors-21-06999-t001" orientation="portrait"><object-id pub-id-type="pii">sensors-21-06999-t001_Table 1</object-id><label>Table 1</label><caption><p>IoUs for <monospace>Data_Pears1</monospace> before and after the fine-tuning with <monospace>Data_Pears2</monospace>.</p></caption><table frame="hsides" rules="groups"><thead><tr><th align="center" valign="middle" style="border-bottom:solid thin;border-top:solid thin" rowspan="1" colspan="1">
</th><th align="center" valign="middle" style="border-bottom:solid thin;border-top:solid thin" rowspan="1" colspan="1">Before</th><th align="center" valign="middle" style="border-bottom:solid thin;border-top:solid thin" rowspan="1" colspan="1">After</th></tr></thead><tbody><tr><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">IoU</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">0.882</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">0.917</td></tr></tbody></table></table-wrap><table-wrap position="float" id="sensors-21-06999-t002" orientation="portrait"><object-id pub-id-type="pii">sensors-21-06999-t002_Table 2</object-id><label>Table 2</label><caption><p>The best IoUs for the validation datasets with <monospace>CROP</monospace> and <monospace>CROP-Shallow</monospace>, with epoch not more than 2500, using the training dataset (80%: 137) and the validation dataset (20%: 35) of <monospace>Data_Fruit</monospace>. Three experiments, numbered from 0 to 2, were performed for each.</p></caption><table frame="hsides" rules="groups"><thead><tr><th align="center" valign="middle" style="border-bottom:solid thin;border-top:solid thin" rowspan="1" colspan="1">
</th><th colspan="3" align="center" valign="middle" style="border-bottom:solid thin;border-top:solid thin" rowspan="1">
<monospace>CROP</monospace>
</th><th colspan="3" align="center" valign="middle" style="border-bottom:solid thin;border-top:solid thin" rowspan="1">
<monospace>CROP-Shallow</monospace>
</th></tr></thead><tbody><tr><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">experiment</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">0</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">1</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">2</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">0</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">1</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">2</td></tr><tr><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">best IoU</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">0.965</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">0.964</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">0.975</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">0.876</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">0.884</td><td align="center" valign="middle" style="border-bottom:solid thin" rowspan="1" colspan="1">0.877</td></tr></tbody></table></table-wrap></floats-group></article>