
<!DOCTYPE article
  PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with MathML3 v1.4 20241031//EN" "JATS-archivearticle1-4-mathml3.dtd">
<article article-type="data-paper" xml:lang="en" dtd-version="1.4"><processing-meta base-tagset="archiving" mathml-version="3.0" table-model="xhtml" tagset-family="jats"><restricted-by>pmc</restricted-by></processing-meta><front><journal-meta><journal-id journal-id-type="nlm-ta">Data Brief</journal-id><journal-id journal-id-type="iso-abbrev">Data Brief</journal-id><journal-id journal-id-type="pmc-domain-id">2750</journal-id><journal-id journal-id-type="pmc-domain">dib</journal-id><journal-id journal-id-type="nlm-id">101654995</journal-id><journal-title-group><journal-title>Data in Brief</journal-title></journal-title-group><issn pub-type="epub">2352-3409</issn><?publisher_abbrev elsevier?><publisher><publisher-name>Elsevier</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="pmcid">PMC10618421</article-id><article-id pub-id-type="pmcid-ver">PMC10618421.1</article-id><article-id pub-id-type="pmcaid">10618421</article-id><article-id pub-id-type="pmcaiid">10618421</article-id><article-id pub-id-type="pmid">37920387</article-id><article-id pub-id-type="doi">10.1016/j.dib.2023.109688</article-id><article-id pub-id-type="pii">S2352-3409(23)00767-9</article-id><article-id pub-id-type="publisher-id">109688</article-id><article-version article-version-type="pmc-version">1</article-version><article-categories><subj-group subj-group-type="heading"><subject>Data Article</subject></subj-group></article-categories><title-group><article-title>An extensive real-world in field tomato image dataset involving maturity classification and recognition of fresh and defect tomatoes</article-title></title-group><contrib-group><contrib contrib-type="author" id="au0001"><name name-style="western"><surname>Khatun</surname><given-names initials="T">Tania</given-names></name><email>tania.cse@diu.edu.bd</email><xref rid="aff0001" ref-type="aff">a</xref><xref rid="cor0001" ref-type="corresp">⁎</xref></contrib><contrib contrib-type="author" id="au0002"><name name-style="western"><surname>Razzak</surname><given-names initials="A">Abdur</given-names></name><xref rid="aff0001" ref-type="aff">a</xref></contrib><contrib contrib-type="author" id="au0003"><name name-style="western"><surname>Islam</surname><given-names initials="MS">Md. Shofiul</given-names></name><xref rid="aff0001" ref-type="aff">a</xref></contrib><contrib contrib-type="author" id="au0004"><name name-style="western"><surname>Uddin</surname><given-names initials="MS">Mohammad Shorif</given-names></name><xref rid="aff0002" ref-type="aff">b</xref></contrib><aff id="aff0001"><label>a</label>Department of Computer Science and Engineering, Daffodil International University, Dhaka, Bangladesh</aff><aff id="aff0002"><label>b</label>Department of Computer Science and Engineering, Jahangirnagar University, Dhaka, Bangladesh</aff></contrib-group><author-notes><corresp id="cor0001"><label>⁎</label>Corresponding author. <email>tania.cse@diu.edu.bd</email></corresp></author-notes><pub-date pub-type="collection"><month>12</month><year>2023</year></pub-date><pub-date pub-type="epub"><day>15</day><month>10</month><year>2023</year></pub-date><volume>51</volume><issue-id pub-id-type="pmc-issue-id">447014</issue-id><elocation-id>109688</elocation-id><history><date date-type="received"><day>5</day><month>9</month><year>2023</year></date><date date-type="rev-recd"><day>2</day><month>10</month><year>2023</year></date><date date-type="accepted"><day>11</day><month>10</month><year>2023</year></date></history><pub-history><event event-type="pmc-release"><date><day>15</day><month>10</month><year>2023</year></date></event><event event-type="pmc-live"><date><day>02</day><month>11</month><year>2023</year></date></event><event event-type="pmc-last-change"><date iso-8601-date="2025-08-28 20:25:31.720"><day>28</day><month>08</month><year>2025</year></date></event></pub-history><permissions><copyright-statement>© 2023 The Author(s)</copyright-statement><copyright-year>2023</copyright-year><license><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/" specific-use="textmining" content-type="ccbylicense">https://creativecommons.org/licenses/by/4.0/</ali:license_ref><license-p>This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).</license-p></license></permissions><self-uri xmlns:xlink="http://www.w3.org/1999/xlink" content-type="pmc-pdf" xlink:href="main.pdf"><?pdf-name main.pdf?><?pdf-size 3538181?><?pdf-md5 1f5144f63cfb84241cbac7a27aecfaf5?><?pdf-image-server-status NEVER_LOAD?><?pdf-cloudpmc-urn urn:app:e8f4/10618421/1f5144f63cfb/main.pdf?></self-uri><abstract id="abs0001"><p>Tomato, a fruiting plant species within the Solanaceae family, is a widely used ingredient in culinary dishes due to its sweet and acidic flavor profile, as well as its rich nutritional content. Recognized for its potential health benefits, including reducing the risk of coronary artery disease and specific types of cancer, tomatoes have become a staple in global cuisine. Traditional methods for tomato maturity assessment, harvesting, quality grading, and packaging are often labor-intensive and economically inefficient. This paper introduces an extensive dataset of high-resolution tomato images collected over an eight-month period from the demonstration fields of Sher-E-Bangla Agricultural University in Dhaka, Bangladesh, in collaboration with plant breeding experts of the same university. The dataset was meticulously curated to ensure precision and consistency, encompassing various stages of tomato maturity, including images of both fresh and defective tomatoes. This dataset is a valuable resource for researchers, stakeholders, and individuals interested in tomato production in Bangladesh, providing a robust foundation for leveraging computer vision and deep learning techniques in the agriculture sector. The dataset's potential applications extend to automating tasks such as robotic harvesting, quality assessment, and packaging systems, ultimately enhancing the efficiency of tomato production processes.</p></abstract><kwd-group id="keys0001"><title>Keywords</title><kwd>Tomato dataset</kwd><kwd>Agriculture</kwd><kwd>Image recognition</kwd><kwd>Deep learning</kwd><kwd>and Computer vision</kwd></kwd-group><custom-meta-group><custom-meta><meta-name>pmc-status-qastatus</meta-name><meta-value>0</meta-value></custom-meta><custom-meta><meta-name>pmc-status-live</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-status-embargo</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-status-released</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-open-access</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-olf</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-manuscript</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-legally-suppressed</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-has-pdf</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-has-supplement</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-pdf-only</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-suppress-copyright</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-is-real-version</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-is-scanned-article</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-preprint</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>pmc-prop-in-epmc</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>pmc-license-ref</meta-name><meta-value>CC BY</meta-value></custom-meta></custom-meta-group></article-meta></front><body><p id="para0002">Specifications Table<table-wrap position="float" id="utbl0001" orientation="portrait"><table frame="hsides" rules="groups"><tbody><tr><td valign="top" colspan="1" rowspan="1">Subject</td><td valign="top" colspan="1" rowspan="1">Computer science</td></tr><tr><td valign="top" colspan="1" rowspan="1">Specific subject area</td><td valign="top" colspan="1" rowspan="1">Image detection, Robotic harvesting, Image categorization, Ripeness analysis</td></tr><tr><td valign="top" colspan="1" rowspan="1">Data format</td><td valign="top" colspan="1" rowspan="1">Raw images</td></tr><tr><td valign="top" colspan="1" rowspan="1">Type of data</td><td valign="top" colspan="1" rowspan="1">JPEG</td></tr><tr><td valign="top" colspan="1" rowspan="1">Data collection</td><td valign="top" colspan="1" rowspan="1">In collaboration with an expert in the field from Sher-E-Bangla Agricultural University in Dhaka, Bangladesh, images were captured between September 22 and April 23 from the demonstration grounds of the Horticulture Department at the university. It's worth noting that this dataset is entirely new, and no prior research has been conducted using it.</td></tr><tr><td valign="top" colspan="1" rowspan="1">Data source location</td><td valign="top" colspan="1" rowspan="1"><bold>Location:</bold> Sher-E-Bangla Agricultural University<break/><bold>Zone:</bold> Sher-E-Bangla Nagar, Dhaka-1207<break/><bold>Country:</bold> Bangladesh</td></tr><tr><td valign="top" colspan="1" rowspan="1">Data accessibility</td><td valign="top" colspan="1" rowspan="1"><bold>Repository name:</bold> Mendeley Data<break/><bold>Data identification number:</bold><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://doi.org/10.17632/s42kpg8h37.1" id="interref0001a">10.17632/s42kpg8h37.1</ext-link><break/><bold>Direct URL to data:</bold><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://data.mendeley.com/datasets/s42kpg8h37/1" id="interref0001b">https://data.mendeley.com/datasets/s42kpg8h37/1</ext-link><break/><bold>Instructions for accessing these data:</bold> Adhering to the appropriate citation guidelines is crucial when utilizing these datasets.</td></tr></tbody></table></table-wrap></p><sec id="sec0002"><label>1</label><title>Value of the Data</title><p id="para9003">
<list list-type="simple" id="celist0001"><list-item id="celistitem0001"><label>•</label><p id="para0003">Robotic harvesting represents an advanced agricultural technology that offers the potential for substantial enhancements in both quality and productivity, while concurrently reducing production costs and minimizing delays <xref rid="bib0001" ref-type="bibr">[1]</xref>. Achieving optimal results with robotic harvesting hinges on harvesting fruits and vegetables precisely at their peak ripeness; otherwise, substantial losses can be incurred. Timely identification of the appropriate harvest time is thus imperative. In the case of tomatoes, harvest timing is conventionally determined based on their skin color <xref rid="bib0002" ref-type="bibr">[2]</xref>. This dataset encompasses data for tomato maturity identification by categorizing it into two subsets: mature and immature tomatoes, based on their surface color complexion. By analyzing the peel color of these images, machines can effectively distinguish between mature and immature tomatoes in real-world scenarios.</p></list-item><list-item id="celistitem0002"><label>•</label><p id="para0004">Due to the lengthy and time-consuming process of transportation tomatoes become defective very easily which people are not willing to buy. For agricultural production, fruit processing, and packing businesses, the detection of defective fruits is extremely important because it can bring significant economic ramifications <xref rid="bib0003" ref-type="bibr">[3]</xref>. In Bangladesh, the detection of fresh and defective fruits is often performed manually, a labor-intensive and inefficient process for farmers. Hence, there is a compelling need to develop a novel classification model capable of autonomously identifying fruit defects without human intervention, reducing costs and production time. This dataset comprises tomato quality grading data that visually represents both fresh and defective tomatoes. Training machines with this dataset empowers them to readily identify and categorize fresh and deteriorating tomatoes at an early stage.</p></list-item><list-item id="celistitem0003"><label>•</label><p id="para0005">Beyond its immediate applications, researchers can leverage the assembled dataset for various computer vision, machine learning, and deep learning approaches. These techniques hold promise for addressing multiple facets of tomato production, encompassing harvesting optimization, freshness forecasting, packaging automation, and other automated solutions.</p></list-item></list>
</p></sec><sec id="sec0003"><label>2</label><title>Data Description</title><p id="para0006">Tomatoes are widely cherished and nutritionally rich crops cultivated across Bangladesh. Traditionally, they have been primarily grown as a winter vegetable in our nation. Nevertheless, the Bangladesh Agricultural Research Institute (BARl) has recently introduced some varieties suitable for summer cultivation.</p><p id="para0007">The images used in this study were meticulously collected from the Horticulture Department's exhibition grounds at Sher-E-Bangla Agricultural University in Dhaka, Bangladesh, spanning from September 22 to April 23 in collaboration with an expert from the same university. <xref rid="fig0001" ref-type="fig">Fig. 1</xref> shows an image of the actual field conditions where our data was gathered. The primary challenge encountered during data collection pertained to capturing images amidst noisy backgrounds and uneven lighting conditions.<fig id="fig0001" position="float" orientation="portrait"><label>Fig. 1</label><caption><p>The real tomato field from where data were collected.</p></caption><alt-text id="alt0001">Fig 1</alt-text><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="gr1.jpg"><?image-name gr1.jpg?><?image-size 314923?><?image-md5 d85d9e96c1fbd45a12e1e2f3e819d3b9?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 2878?><?image-original-width 2168?><?image-scaled-height 958?><?image-scaled-width 722?><?image-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/d85d9e96c1fb/gr1.jpg?><?thumb-name gr1.gif?><?thumb-size 16098?><?thumb-md5 304fa66757ab175b4d8d5a9bbb5097e7?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 133?><?thumb-scaled-width 100?><?thumb-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/304fa66757ab/gr1.gif?></graphic></fig></p><p id="para0008">Tomato production in Bangladesh is on the rise, offering farmers an additional source of income. Nevertheless, a significant challenge arises from improper storage practices, resulting in substantial losses for farmers. To mitigate these losses, it is crucial to closely monitor the maturity of tomatoes. The condition of vegetable storage and ripening is intricately linked to its level of maturity. In addressing this concern, modern methods have outperformed manual approaches in terms of accuracy, precision, time efficiency, cost-effectiveness, and non-destructiveness.</p><p id="para0009">To facilitate these advancements, in this paper, we have introduced two datasets. The first dataset is called the Tomato Maturity Detection Dataset, and the second is the Tomato Quality Grading Dataset. The level of ripeness and quality are closely associated with the intensity of redness in color and the prominence of flavor <xref rid="bib0002" ref-type="bibr">[2]</xref>. Each of these dataset folders is further split into two subfolders: the original dataset and the augmented dataset. The original dataset folder contains images directly captured with the camera, while the augmented dataset folder contains images generated from the original dataset through data augmentation processes using the software.</p><p id="para0010">Both the original and augmented datasets within the Tomato Maturity Detection Dataset are categorized into two groups: Mature Tomatoes and Immature Tomatoes. Similarly, both the original and augmented datasets within the Tomato Quality Grading Dataset are categorized into two groups: Fresh Tomatoes and Defect Tomatoes. Each of these folders contains relevant tomato images. The images collected come in sizes of 765 × 1024 pixels and 1280 × 957 pixels. <xref rid="bib0004" ref-type="bibr">[4]</xref>, <xref rid="tbl0001" ref-type="table">Table 1</xref> provides a description of mature and immature tomato categories within the dataset.<table-wrap position="float" id="tbl0001" orientation="portrait"><label>Table 1</label><caption><p>Description of tomato maturity detection dataset.</p></caption><alt-text id="alt0006">Table 1:</alt-text><table frame="hsides" rules="groups"><tbody><tr><td valign="top" colspan="1" rowspan="1"><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fx1.gif"><?image-name fx1.gif?><?image-size 8445?><?image-md5 f5e500bc4a8fe382e542a821060698da?><?image-image-server-status NEVER_LOAD?><?image-scaled-height 79?><?image-scaled-width 100?><?image-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/f5e500bc4a8f/fx1.gif?><?thumb-name fx1.gif?><?thumb-size 8445?><?thumb-md5 f5e500bc4a8fe382e542a821060698da?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 79?><?thumb-scaled-width 100?><?thumb-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/f5e500bc4a8f/fx1.gif?><alt-text id="alt1">Image, table 1</alt-text></inline-graphic></td></tr></tbody></table><table-wrap-foot><fn><p>The Tomato Quality Grading Dataset was generated at home by capturing images of tomatoes at different stages every three days, with the main factors determining tomato quality being color, texture, and flavor. <xref rid="bib0004" ref-type="bibr">[4]</xref>, <xref rid="tbl0002" ref-type="table">Table 2</xref> below provides information about fresh and defective classes within the Tomato Quality Grading Dataset.</p></fn></table-wrap-foot></table-wrap><table-wrap position="float" id="tbl0002" orientation="portrait"><label>Table 2</label><caption><p>Description of tomato quality grading dataset.</p></caption><alt-text id="alt0007">Table 2:</alt-text><table frame="hsides" rules="groups"><tbody><tr><td valign="top" colspan="1" rowspan="1"><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fx2.gif"><?image-name fx2.gif?><?image-size 8149?><?image-md5 9bb945599e1c37371113fddffe9187a7?><?image-image-server-status NEVER_LOAD?><?image-scaled-height 80?><?image-scaled-width 105?><?image-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/9bb945599e1c/fx2.gif?><?thumb-name fx2.gif?><?thumb-size 8149?><?thumb-md5 9bb945599e1c37371113fddffe9187a7?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 80?><?thumb-scaled-width 105?><?thumb-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/9bb945599e1c/fx2.gif?><alt-text id="alt2">Image, table 2</alt-text></inline-graphic></td></tr></tbody></table></table-wrap></p></sec><sec id="sec0004"><label>3</label><title>Experimental Design, Materials and Methods</title><sec id="sec0005"><label>3.1</label><title>Camera specification</title><p id="para0011">The dataset was collected using three different smartphones: Samsung Galaxy, Redmi Note-9, and Redmi Y3. Each of these smartphones has specific camera configurations. The Samsung Galaxy is equipped with a triple-lens reflex digital camera that includes a variety of lenses whereas the Redmi Note-9 has a quad-lens reflex digital camera with different lens types. Additionally, the Redmi Y3 features a dual-lens reflex digital camera. All of these cameras come with HDR functionalities, panorama, and LED flash.</p></sec><sec id="sec0006"><label>3.2</label><title>Data augmentation</title><p id="para0012">In order to satisfy the needs of machine vision-based deep learning models, which require a significant number of images, we applied data augmentation methods. Data augmentation serves the purpose of enlarging the dataset size, mitigating overfitting, and enhancing the overall performance of deep learning models <xref rid="bib0005" ref-type="bibr">[5]</xref>. This technique involves actions like rotations, zooming, and mirroring.</p><p id="para0013">During the augmentation process, we carefully adjusted specific parameters, including a probability of 0.07 with maximum left and right rotation angles of 10°, and a probability of 1 with maximum left and right rotation angles of 5° for the rotation component. For zoom and random zoom, the parameters included a probability of 0.5, a minimum factor of 1.1, a maximum factor of 1.5, and a probability of 0.5 with a percentage area of 0.08.</p><p id="para0014">The augmentation process was executed automatically using software, resulting in the creation of 10,000 augmented images derived from the original dataset. <xref rid="fig0002" ref-type="fig">Fig. 2</xref> provides a visual representation of the tomato dataset generation process. Furthermore, <xref rid="tbl0003" ref-type="table">Tables 3</xref> and <xref rid="tbl0004" ref-type="table">4</xref> display the augmented images alongside their corresponding original sample images in each category.<fig id="fig0002" position="float" orientation="portrait"><label>Fig. 2</label><caption><p>The process of tomato dataset generation.</p></caption><alt-text id="alt0002">Fig 2</alt-text><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="gr2.jpg"><?image-name gr2.jpg?><?image-size 41228?><?image-md5 d68cd46d07e80c3853b87764029d2d68?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 940?><?image-original-width 2458?><?image-scaled-height 268?><?image-scaled-width 702?><?image-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/d68cd46d07e8/gr2.jpg?><?thumb-name gr2.gif?><?thumb-size 8053?><?thumb-md5 97c9822d8b8b40cf1cd9f0d21c27b069?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 76?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/97c9822d8b8b/gr2.gif?></graphic></fig><table-wrap position="float" id="tbl0003" orientation="portrait"><label>Table 3</label><caption><p>Augmented images of tomato maturity detection dataset.</p></caption><alt-text id="alt0008">Table 3:</alt-text><table frame="hsides" rules="groups"><tbody><tr><td valign="top" colspan="1" rowspan="1"><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fx3.gif"><?image-name fx3.gif?><?image-size 10206?><?image-md5 164fced4221dbb1962141ee242a0a58b?><?image-image-server-status NEVER_LOAD?><?image-scaled-height 80?><?image-scaled-width 135?><?image-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/164fced4221d/fx3.gif?><?thumb-name fx3.gif?><?thumb-size 10206?><?thumb-md5 164fced4221dbb1962141ee242a0a58b?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 80?><?thumb-scaled-width 135?><?thumb-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/164fced4221d/fx3.gif?><alt-text id="alt3">Image, table 3</alt-text></inline-graphic></td></tr></tbody></table></table-wrap><table-wrap position="float" id="tbl0004" orientation="portrait"><label>Table 4</label><caption><p>Augmented images of tomato quality grading dataset.</p></caption><alt-text id="alt0009">Table 4:</alt-text><table frame="hsides" rules="groups"><tbody><tr><td valign="top" colspan="1" rowspan="1"><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fx4.gif"><?image-name fx4.gif?><?image-size 9755?><?image-md5 d1fea29b6d05664699194266a9536895?><?image-image-server-status NEVER_LOAD?><?image-scaled-height 80?><?image-scaled-width 138?><?image-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/d1fea29b6d05/fx4.gif?><?thumb-name fx4.gif?><?thumb-size 9755?><?thumb-md5 d1fea29b6d05664699194266a9536895?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 80?><?thumb-scaled-width 138?><?thumb-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/d1fea29b6d05/fx4.gif?><alt-text id="alt4">Image, table 4</alt-text></inline-graphic></td></tr></tbody></table></table-wrap></p><p id="para0015">Detailed statistics regarding the image dataset can be found in <xref rid="tbl0005" ref-type="table">Table 5</xref>.<table-wrap position="float" id="tbl0005" orientation="portrait"><label>Table 5</label><caption><p>The statistics of the tomato dataset.</p></caption><alt-text id="alt0010">Table 5:</alt-text><table frame="hsides" rules="groups"><thead><tr><th align="left" valign="top" colspan="1" rowspan="1">Dataset category</th><th valign="top" colspan="1" rowspan="1">Class</th><th valign="top" colspan="1" rowspan="1">No of image in the original dataset</th><th valign="top" colspan="1" rowspan="1">No of image in the augmented dataset</th></tr></thead><tbody><tr><td rowspan="2" align="left" valign="top" colspan="1">Tomato maturity detection dataset</td><td valign="top" colspan="1" rowspan="1">Immature</td><td valign="top" colspan="1" rowspan="1">500</td><td valign="top" colspan="1" rowspan="1">2000</td></tr><tr><td valign="top" colspan="1" rowspan="1">Mature</td><td valign="top" colspan="1" rowspan="1">500</td><td valign="top" colspan="1" rowspan="1">2000</td></tr><tr><td rowspan="2" align="left" valign="top" colspan="1">Tomato quality grading dataset</td><td valign="top" colspan="1" rowspan="1">Fresh</td><td valign="top" colspan="1" rowspan="1">1350</td><td valign="top" colspan="1" rowspan="1">3000</td></tr><tr><td valign="top" colspan="1" rowspan="1">Defect</td><td valign="top" colspan="1" rowspan="1">636</td><td valign="top" colspan="1" rowspan="1">3000</td></tr></tbody></table></table-wrap></p></sec><sec id="sec0007"><label>3.3</label><title>Deep learning model validation</title><p id="para0016">We proposed a deep learning model aimed at efficiently training the dataset to achieve state-of-the-art results. The validation of a deep learning model necessitates a meticulous examination of its output on a dataset, as discussed in <xref rid="bib0006" ref-type="bibr">[6]</xref>. The deep learning model adheres to a five-stage procedure, which involves data preprocessing, data partitioning, model training, assessing performance using a validation set, and finally, testing the model on an entirely separate test set. This rigorous methodology is essential to confirm the model's trustworthiness in delivering precise outcomes and its capability to generalize to new data.</p><p id="para0017">Effective data preprocessing plays a pivotal role in extracting valuable insights from the dataset. In our study, the preprocessing of images encompasses various data transformations, which include tasks like image resizing, contrast enhancement, noise reduction, augmentation, and segmentation. The specific details are outlined below:</p><p id="para0018"><italic toggle="yes">Noise reduction:</italic> We employed the non-linear median filtering technique to effectively remove noise from the images. This choice was made due to the impulsive nature of the noise observed.</p><p id="para0019"><italic toggle="yes">Contrast enhancement:</italic> To rectify uneven illumination and enhance image contrast, we utilized the histogram equalization technique.</p><p id="para0020"><italic toggle="yes">Image resizing:</italic> Since images in the dataset may exhibit varying sizes, we deemed it necessary to resize them as per our requirements. This step was essential to ensure compatibility when training deep learning models.</p><p id="para0021"><italic toggle="yes">Image segmentation:</italic> When necessary, we performed image cropping to eliminate unwanted background elements, enhancing the quality of the dataset.</p><p id="para0022"><italic toggle="yes">Data augmentation:</italic> To augment our dataset, a crucial requirement for training deep learning models, we implemented data augmentation techniques. Further details of this augmentation process can be found in <xref rid="sec0006" ref-type="sec">Section 3.2</xref>.</p><p id="para0023">We partitioned the gathered data into two distinct sets: the training set and the testing set, in an 80:20 ratio. Specifically, 80% of the images were randomly selected for the training dataset, while the remaining 20% constituted the test dataset. The testing set served the purpose of evaluating the model's performance once it had been trained using the training data.</p><p id="para0024"><xref rid="fig0003" ref-type="fig">Fig. 3</xref> provides an overview of the extensive validation procedures applied to our deep-learning model using the tomato image dataset. This validation encompassed tasks such as distinguishing between mature and immature tomatoes and classifying tomatoes as fresh or defective.<fig id="fig0003" position="float" orientation="portrait"><label>Fig. 3</label><caption><p>Working procedure of maturity detection and the recognition of fresh and defect tomatoes.</p></caption><alt-text id="alt0003">Fig 3</alt-text><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="gr3.jpg"><?image-name gr3.jpg?><?image-size 87244?><?image-md5 cc02b4d20d1cbddcda859895267a4d0e?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 1909?><?image-original-width 2146?><?image-scaled-height 636?><?image-scaled-width 715?><?image-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/cc02b4d20d1c/gr3.jpg?><?thumb-name gr3.gif?><?thumb-size 6614?><?thumb-md5 645fbfdc2e74a63ff3ffb87fc4b0ee64?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 89?><?thumb-scaled-width 100?><?thumb-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/645fbfdc2e74/gr3.gif?></graphic></fig></p><sec id="sec0008"><label>3.3.1</label><title>Model description</title><p id="para0025">In this paper, we have implemented the MobileNetV2 architecture for the purpose of detecting tomato maturity and quality. MobileNetV2 is a convolutional neural network (CNN) architecture explicitly crafted for efficient image classification purposes <xref rid="bib0007" ref-type="bibr">[7]</xref>. It builds upon the original MobileNetV2 architecture by introducing innovative architectural elements, prominently featuring inverted residuals. These residuals are composed of the following key components:<list list-type="simple" id="celist0002"><list-item id="celistitem0004"><label>•</label><p id="para0026">Depth-wise convolution: This component entails the application of a separate 3×3 convolution to each input channel. This operation effectively captures spatial information within the data.</p></list-item><list-item id="celistitem0005"><label>•</label><p id="para0027">Point-wise convolution: Following the depth-wise convolution, a 1×1 convolution is applied to amalgamate the output channels. This step serves a dual purpose by reducing dimensionality and introducing non-linearity into the network.</p></list-item></list></p><p id="para0028">In MobileNetV2, batch normalization is applied before activation functions (e.g., ReLU) within each convolutional or fully connected layer. MobileNetV2 typically employs global average pooling to reduce the spatial dimensions of the feature maps. This step converts the feature maps into a fixed-size vector. <xref rid="fig0004" ref-type="fig">Fig. 4</xref> represents MobileNetV2 architecture.<fig id="fig0004" position="float" orientation="portrait"><label>Fig. 4</label><caption><p>MobileNetV2 architecture.</p></caption><alt-text id="alt0004">Fig 4:</alt-text><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="gr4.jpg"><?image-name gr4.jpg?><?image-size 27348?><?image-md5 f7e603a6208363770112859d570d8240?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 481?><?image-original-width 2458?><?image-scaled-height 137?><?image-scaled-width 702?><?image-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/f7e603a62083/gr4.jpg?><?thumb-name gr4.gif?><?thumb-size 6200?><?thumb-md5 209758f1be28ba078898fbe3680386c5?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 39?><?thumb-scaled-width 200?><?thumb-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/209758f1be28/gr4.gif?></graphic></fig></p></sec><sec id="sec0009"><label>3.3.2</label><title>Measurement metrics</title><p id="para0029"><italic toggle="yes">Confusion matrix:</italic> A confusion matrix stands as a crucial instrument in the realm of machine learning and classification endeavors. It serves the purpose of appraising a predictive model's performance. By employing a confusion matrix, one can compute a range of performance indicators like accuracy, precision, recall, and F1 score, which aid in gauging the model's proficiency in accurately categorizing instances and pinpointing potential origins of errors, such as false positives and false negatives.</p><p id="para0030"><italic toggle="yes">Accuracy:</italic> Accuracy, a crucial performance metric in classification, quantifies the fraction of accurately classified instances within the dataset. This metric is determined by dividing the count of correct predictions (comprising both true positives and true negatives) by the total number of instances. Accuracy offers a general evaluation of a model's correctness. It is imperative to also take into account precision, recall, and the F1 score in conjunction with accuracy to gain a comprehensive insight into a model's performance.<disp-formula id="eqn0001"><label>(1)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M1" altimg="si1.svg"><mml:mrow><mml:mrow><mml:mtext>Accuracy</mml:mtext><mml:mo linebreak="badbreak">=</mml:mo></mml:mrow><mml:mfrac><mml:mrow><mml:mtext>TP</mml:mtext><mml:mo>+</mml:mo><mml:mtext>TN</mml:mtext></mml:mrow><mml:mrow><mml:mtext>TP</mml:mtext><mml:mo>+</mml:mo><mml:mtext>FP</mml:mtext><mml:mo>+</mml:mo><mml:mtext>TN</mml:mtext><mml:mo>+</mml:mo><mml:mtext>FN</mml:mtext></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula></p><p id="para0031">Where</p><p id="para0032">True Positive (TP) - the classifier classifies the right class of tomato as right</p><p id="para0033">True Negative (TN) - the classifier classifies the wrong class of tomato as wrong</p><p id="para0034">False Positive (FP) - the classifier classifies the wrong class of tomato as right</p><p id="para0035">False Negative (FN) - the classifier classifies the right class of tomato as wrong</p><p id="para0036"><italic toggle="yes">Precision:</italic> Precision, within the realm of classification tasks, is a performance measure that gauges the precision of positive predictions made by a model. It is calculated as the proportion of true positive predictions relative to the total number of positive predictions (which includes both true positives and false positives). In essence, precision evaluates the model's capacity to accurately pinpoint relevant instances among its positive forecasts. A higher precision signifies a reduced occurrence of false positives and, consequently, a decreased likelihood of incorrectly classifying negative instances as positive.<disp-formula id="eqn0002"><label>(2)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M2" altimg="si2.svg"><mml:mrow><mml:mrow><mml:mtext>Precision</mml:mtext><mml:mo linebreak="badbreak">=</mml:mo></mml:mrow><mml:mfrac><mml:mrow><mml:mtext>True</mml:mtext><mml:mspace width="0.33em"/><mml:mtext>Positives</mml:mtext></mml:mrow><mml:mrow><mml:mtext>True</mml:mtext><mml:mspace width="0.33em"/><mml:mtext>positives</mml:mtext><mml:mo>+</mml:mo><mml:mtext>False</mml:mtext><mml:mspace width="0.33em"/><mml:mtext>Positives</mml:mtext></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula></p><p id="para0037"><italic toggle="yes">Recall:</italic> Recall, also known as true positive rate, quantifies a model's aptitude in accurately recognizing all pertinent instances within a dataset. It is expressed as the fraction of true positive predictions relative to the total number of actual positive instances (comprising both true positives and false negatives). Recall becomes especially valuable when the expense associated with overlooking positive instances (false negatives) is substantial. A heightened recall signifies an increased capability to capture the majority of positive cases.<disp-formula id="eqn0003"><label>(3)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M3" altimg="si3.svg"><mml:mrow><mml:mrow><mml:mtext>Recall</mml:mtext><mml:mo linebreak="badbreak">=</mml:mo></mml:mrow><mml:mfrac><mml:mrow><mml:mtext>True</mml:mtext><mml:mspace width="0.33em"/><mml:mtext>Positives</mml:mtext></mml:mrow><mml:mrow><mml:mtext>True</mml:mtext><mml:mspace width="0.33em"/><mml:mtext>positives</mml:mtext><mml:mo>+</mml:mo><mml:mtext>False</mml:mtext><mml:mspace width="0.33em"/><mml:mtext>Negatives</mml:mtext></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula></p><p id="para0038"><italic toggle="yes">F1 score:</italic> The F1 score serves as a metric that melds together both precision and recall to deliver a well-rounded evaluation of a model's performance. Its computation involves taking the harmonic mean of precision and recall, affording equal importance to both metrics. The F1 score proves particularly valuable when there exists an uneven distribution between the positive and negative classes in the dataset. It offers a single numerical representation that takes into account both false positives and false negatives, establishing it as a dependable gauge of the overall classification performance.<disp-formula id="eqn0004"><label>(4)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M4" altimg="si4.svg"><mml:mrow><mml:mrow><mml:mi mathvariant="normal">F</mml:mi><mml:mo linebreak="badbreak">−</mml:mo><mml:mn>1</mml:mn><mml:mspace width="0.33em"/><mml:mtext>Score</mml:mtext><mml:mo linebreak="badbreak">=</mml:mo></mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>*</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mtext>Precision</mml:mtext><mml:mo>*</mml:mo><mml:mtext>Recall</mml:mtext></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mtext>Precision</mml:mtext><mml:mo>+</mml:mo><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula></p><p id="para0039">The confusion matrices of the MobileNetV2 model is shown in <xref rid="fig0005" ref-type="fig">Fig. 5</xref>.<fig id="fig0005" position="float" orientation="portrait"><label>Fig. 5</label><caption><p>Confusion matrices of the MobileNetV2 model</p></caption><alt-text id="alt0005">Fig 5:</alt-text><graphic xmlns:xlink="http://www.w3.org/1999/xlink" position="float" orientation="portrait" xlink:href="gr5.jpg"><?image-name gr5.jpg?><?image-size 34299?><?image-md5 1590425c9163a493a256d65a91f2bb9f?><?image-image-server-status LOAD_COMPLETED?><?image-original-height 1082?><?image-original-width 2458?><?image-scaled-height 309?><?image-scaled-width 702?><?image-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/1590425c9163/gr5.jpg?><?thumb-name gr5.gif?><?thumb-size 6557?><?thumb-md5 2923c6a209938b2653b0e6a25652b2d4?><?thumb-image-server-status NEVER_LOAD?><?thumb-scaled-height 80?><?thumb-scaled-width 181?><?thumb-cloudpmc-urn urn:cdn:blobs/e8f4/10618421/2923c6a20993/gr5.gif?></graphic></fig></p><p id="para0040"><xref rid="tbl0006" ref-type="table">Table 6</xref> represents the performance metrics for the MobileNetV2 model.<table-wrap position="float" id="tbl0006" orientation="portrait"><label>Table 6</label><caption><p>Performance metrics for MobileNetV2 model.</p></caption><alt-text id="alt0011">Table 6:</alt-text><table frame="hsides" rules="groups"><thead><tr><th valign="top" colspan="1" rowspan="1">Classes</th><th align="left" valign="top" colspan="1" rowspan="1">Model</th><th valign="top" colspan="1" rowspan="1">Precision</th><th valign="top" colspan="1" rowspan="1">Recall</th><th valign="top" colspan="1" rowspan="1">F-1 score</th><th align="left" valign="top" colspan="1" rowspan="1">Accuracy</th></tr></thead><tbody><tr><td valign="top" colspan="1" rowspan="1">Immature</td><td rowspan="2" align="left" valign="top" colspan="1">MobileNetV2</td><td valign="top" colspan="1" rowspan="1">0.99</td><td valign="top" colspan="1" rowspan="1">0.97</td><td valign="top" colspan="1" rowspan="1">0.98</td><td rowspan="2" align="left" valign="top" colspan="1">98%</td></tr><tr><td valign="top" colspan="1" rowspan="1">Mature</td><td valign="top" colspan="1" rowspan="1">0.97</td><td valign="top" colspan="1" rowspan="1">0.99</td><td valign="top" colspan="1" rowspan="1">0.98</td></tr><tr><td valign="top" colspan="1" rowspan="1">Fresh</td><td rowspan="2" align="left" valign="top" colspan="1">MobileNetV2</td><td valign="top" colspan="1" rowspan="1">0.93</td><td valign="top" colspan="1" rowspan="1">1</td><td valign="top" colspan="1" rowspan="1">0.96</td><td rowspan="2" align="left" valign="top" colspan="1">96%</td></tr><tr><td valign="top" colspan="1" rowspan="1">Defect</td><td valign="top" colspan="1" rowspan="1">1</td><td valign="top" colspan="1" rowspan="1">0.91</td><td valign="top" colspan="1" rowspan="1">0.95</td></tr></tbody></table></table-wrap></p><p id="para0041">In the future, we will extensively examine more state-of-the-art deep learning models using this dataset to determine the best technique for practical applications.</p></sec></sec></sec><sec id="sec0010"><title>Limitations</title><p id="para0042">This system is built to focus only on Tomato data.</p></sec><sec id="sec0011"><title>Ethics Statement</title><p id="para0043">None of the authors of this article have conducted any research using humans or animals as subjects. The datasets consulted for this article are accessible to everyone but following the correct citation guidelines is essential.</p></sec><sec id="sec0011a"><title>CRediT authorship contribution statement</title><p id="para0043a"><bold>Tania Khatun:</bold> Conceptualization, Data curation, Methodology, Visualization, Validation, Writing – original draft, Writing – review &amp; editing. <bold>Abdur Razzak:</bold> Methodology, Data curation. <bold>Md. Shofiul Islam:</bold> Methodology, Data curation. <bold>Mohammad Shorif Uddin:</bold> Supervision, Writing – review &amp; editing.</p></sec></body><back><ref-list id="cebibl1"><title>References</title><ref id="bib0001"><label>1</label><element-citation publication-type="journal" id="sbref0001"><person-group person-group-type="author"><name name-style="western"><surname>Altaheri</surname><given-names>H.</given-names></name><name name-style="western"><surname>Alsulaiman</surname><given-names>M.</given-names></name><name name-style="western"><surname>Muhammad</surname><given-names>G.</given-names></name><name name-style="western"><surname>Umar Amin</surname><given-names>S.</given-names></name><name name-style="western"><surname>Bencherif</surname><given-names>M.</given-names></name><name name-style="western"><surname>Mekhtiche</surname><given-names>M.</given-names></name></person-group><article-title>Date fruit dataset for intelligent harvesting</article-title><source>Data Brief</source><volume>26</volume><year>2019</year><object-id pub-id-type="publisher-id">104514</object-id><pub-id pub-id-type="doi">10.1016/j.dib.2019.104514</pub-id><pub-id pub-id-type="pmcid">PMC6811983</pub-id><pub-id pub-id-type="pmid">31667277</pub-id></element-citation></ref><ref id="bib0002"><label>2</label><element-citation publication-type="journal" id="sbref0002"><person-group person-group-type="author"><name name-style="western"><surname>Moreira</surname><given-names>G.</given-names></name><etal/></person-group><article-title>Benchmark of deep learning and a proposed HSV colour space models for the detection and classification of greenhouse tomato</article-title><source>Agronomy</source><volume>12</volume><issue>2</issue><year>2022</year><fpage>356</fpage><pub-id pub-id-type="doi">10.3390/agronomy12020356</pub-id></element-citation></ref><ref id="bib0003"><label>3</label><element-citation publication-type="journal" id="sbref0003"><person-group person-group-type="author"><name name-style="western"><surname>Sultana</surname><given-names>N.</given-names></name><name name-style="western"><surname>Jahan</surname><given-names>M.</given-names></name><name name-style="western"><surname>Uddin</surname><given-names>M.S.</given-names></name></person-group><article-title>An extensive dataset for successful recognition of fresh and defect fruits</article-title><source>Data Brief</source><volume>44</volume><year>2022</year><object-id pub-id-type="publisher-id">108552</object-id><pub-id pub-id-type="doi">10.1016/j.dib.2022.108552</pub-id><pub-id pub-id-type="pmcid">PMC9469664</pub-id><pub-id pub-id-type="pmid">36111284</pub-id></element-citation></ref><ref id="bib0004"><label>4</label><mixed-citation publication-type="other" id="sbref0004">United States standards for grades of fresh tomatoes, <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://www.ams.usda.gov/sites/default/files/media/Tomato_Standard%5B1%5D.pdf" id="interref0003">https://www.ams.usda.gov/sites/default/files/media/Tomato_Standard%5B1%5D.pdf</ext-link>, Accessed 01 October 2023.</mixed-citation></ref><ref id="bib0005"><label>5</label><element-citation publication-type="journal" id="sbref0005"><person-group person-group-type="author"><name name-style="western"><surname>Shorten</surname><given-names>C.</given-names></name><name name-style="western"><surname>Khoshgoftaar</surname><given-names>TM.</given-names></name></person-group><article-title>A survey on image data augmentation for deep learning</article-title><source>J. Big Data</source><volume>60</volume><year>2019</year><fpage>1</fpage><lpage>48</lpage><pub-id pub-id-type="doi">10.1186/s40537-019-0197-0</pub-id><pub-id pub-id-type="pmcid">PMC8287113</pub-id><pub-id pub-id-type="pmid">34306963</pub-id></element-citation></ref><ref id="bib0006"><label>6</label><element-citation publication-type="journal" id="sbref0006"><person-group person-group-type="author"><name name-style="western"><surname>Magalhães</surname><given-names>S.A.</given-names></name><etal/></person-group><article-title>Evaluating the single-shot multibox detector and YOLO deep learning models for the detection of tomatoes in a greenhouse</article-title><source>Sensors</source><volume>21</volume><issue>10</issue><year>2021</year><fpage>3569</fpage><pub-id pub-id-type="doi">10.3390/s21103569</pub-id><pub-id pub-id-type="pmid">34065568</pub-id><pub-id pub-id-type="pmcid">PMC8160895</pub-id></element-citation></ref><ref id="bib0007"><label>7</label><element-citation publication-type="journal" id="sbref0007"><person-group person-group-type="author"><name name-style="western"><surname>Gulzar</surname><given-names>Y.</given-names></name></person-group><article-title>Fruit image classification model based on MobileNetV2 with deep transfer learning technique</article-title><source>Sustainability</source><volume>15</volume><issue>3</issue><year>2023</year><fpage>1906</fpage><pub-id pub-id-type="doi">10.3390/su15031906</pub-id></element-citation></ref></ref-list><sec sec-type="data-availability" id="refdata001"><title>Data Availability</title><p id="para9001">
<list list-type="simple" id="dacelist0001"><list-item id="rdlistitem0001"><p id="para9002"><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://data.mendeley.com/datasets/s42kpg8h37/1" id="interref0002">Tomato Maturity Detection and Quality Grading Dataset (Original data)</ext-link> (Mendeley Data).</p></list-item></list>
</p></sec><ack id="ack0001"><title>Acknowledgments</title><p id="para0045">We are very grateful to the domain expert Prof. Dr. Md. Saiful Islam from the Genetics and Plant Breeding Department of Sher-E-Bangla Agricultural University (SAU), Dhaka, Bangladesh for the valuable feedback and cooperation to accomplish the task.</p><sec id="sec2012"><title>Declaration of Competing Interest</title><p id="para0046">The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.</p></sec></ack></back></article>