<?xml version="1.0" encoding="UTF-8"?><article xml:lang="en" article-type="research-article"><front><journal-meta><journal-id journal-id-type="pmc-domain-id">1787</journal-id><journal-id journal-id-type="pmc-domain">frontplantsci</journal-id><journal-title-group><journal-title>Frontiers in Plant Science</journal-title><abbrev-journal-title>Front Plant Sci</abbrev-journal-title></journal-title-group><publisher><publisher-name>Frontiers Media SA</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="pmcid">PMC12378099</article-id><article-id pub-id-type="pmcaid">12378099</article-id><article-id pub-id-type="pmcaiid">12378099</article-id><article-id pub-id-type="pmid">40874080</article-id><article-id pub-id-type="doi">10.3389/fpls.2025.1626569</article-id><title-group><article-title>LCAMNet: a lightweight model for apple leaf disease classification in natural environments</article-title></title-group><contrib-group content-type="author"><contrib><name name-style="western"><surname>Jiao</surname><given-names initials="Y">Yuanyuan</given-names></name><xref ref-type="aff" rid="aff1">1</xref></contrib><contrib><name name-style="western"><surname>Li</surname><given-names initials="H">Honghui</given-names></name><xref ref-type="aff" rid="aff1">1</xref><xref rid="fn001" ref-type="author-notes">*</xref></contrib><contrib><name name-style="western"><surname>Fu</surname><given-names initials="X">Xueliang</given-names></name><xref ref-type="aff" rid="aff1">1</xref><xref rid="fn001" ref-type="author-notes">*</xref></contrib><contrib><name name-style="western"><surname>Wang</surname><given-names initials="B">Buyu</given-names></name><xref ref-type="aff" rid="aff1">1</xref><xref ref-type="aff" rid="aff2">2</xref></contrib><contrib><name name-style="western"><surname>Hu</surname><given-names initials="K">Kaiwen</given-names></name><xref ref-type="aff" rid="aff1">1</xref></contrib><contrib><name name-style="western"><surname>Zhou</surname><given-names initials="S">Shuncheng</given-names></name><xref ref-type="aff" rid="aff1">1</xref></contrib><contrib><name name-style="western"><surname>Han</surname><given-names initials="D">Daoqi</given-names></name><xref ref-type="aff" rid="aff1">1</xref></contrib></contrib-group><aff id="aff1">
<label>1</label>
College of Computer and Information Engineering, Inner Mongolia Agricultural University, Hohhot, China
</aff><aff id="aff2">
<label>2</label>
Key Laboratory of Smart Animal Husbandry at Universities of Inner Mongolia Autonomous Region, Hohhot, China
</aff><author-notes><fn id="fn1"><p>Edited by: Nebojsa Bacanin, Singidunum University, Serbia</p></fn><fn id="fn2"><p>Reviewed by: Muzafer Saracevic, University of Novi Pazar, Serbia</p><p>Miodrag Zivkovic, Singidunum University, Serbia</p><p>Milos Antonijevic, Singidunum University, Serbia</p></fn><fn id="fn001"><label>✉</label><p>*Correspondence: Honghui Li, <email>lihh@imau.edu.cn</email>; Xueliang Fu, <email>fuxl@imau.edu.cn</email>
</p></fn></author-notes><pub-date><day>12</day><month>8</month><year>2025</year></pub-date><volume>16</volume><fpage>1626569</fpage><page-range>1626569</page-range><pub-history><event event-type="pmc-release"><date><day>27</day><month>8</month><year>2025</year></date></event></pub-history><permissions><copyright-statement>Copyright © 2025 Jiao, Li, Fu, Wang, Hu, Zhou and Han.</copyright-statement><license><license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</license-p></license></permissions><self-uri xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fpls-16-1626569.pdf" content-type="pmc-pdf"><?cloudpmc-path ce99/12378099/c246a2fb2c04/fpls-16-1626569.pdf?><?cloudpmc-bucket app?><?size 13661959?></self-uri><abstract id="abstract1"><title>Abstract</title><p>Apple leaf diseases severely affect the quality and yield of apples, and accurate classification is crucial for reducing losses. However, in natural environments, the similarity between backgrounds and lesion areas makes it difficult for existing models to balance lightweight design and high accuracy, limiting their practical applications. In order to resolve the aforementioned problem, this paper introduces a lightweight converged attention multi-branch network named LCAMNet. The network integrates depthwise separable convolutions and structural re-parameterization techniques to achieve efficient modeling. To avoid feature loss caused by single downsampling operations, a dual-branch downsampling module is designed. A multi-scale structure is introduced to enhance lesion feature diversity representation. An improved triplet attention mechanism is utilized to better capture deep lesion features. Furthermore, a dataset named SCEBD is constructed, containing multiple common disease types and interference factors under natural environments, realistically reflecting orchard conditions. Experimental results show that LCAMNet achieves 92.60% accuracy on the SCEBD and 95.31% on a public dataset, with only 0.03 GFLOPs and 1.30M parameters. The model maintains high accuracy while remaining lightweight, enabling effective apple leaf disease classification in natural environments on devices with limited resources.</p><sec id="kwd-group1" sec-type="kwd-group" disp-level="2"><p><bold>Keywords:</bold> apple leaf disease, image classification, deep learning, triplet attention mechanism, FGVC8 dataset</p></sec></abstract><custom-meta-group><custom-meta><meta-name>status</meta-name><meta-value>released</meta-value></custom-meta><custom-meta><meta-name>display-pdf</meta-name><meta-value>yes</meta-value></custom-meta><custom-meta><meta-name>is-olf</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-manuscript</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-preprint</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-journal-matter</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-scanned</meta-name><meta-value>no</meta-value></custom-meta><custom-meta><meta-name>is-retracted</meta-name><meta-value>no</meta-value></custom-meta></custom-meta-group></article-meta><notes notes-type="article-notes"><sec id="historyarticle-meta1" sec-type="history" disp-level="2"><p>Received 2025 May 11; Accepted 2025 Jul 11; Collection date 2025.</p></sec></notes></front><body><sec id="s1" disp-level="1"><label>1.</label><title>Introduction</title><p>Apple (Malus domestica), a member of the Rosaceae family, is one of the most widely cultivated and commonly consumed fruits worldwide. China is the largest apple producer globally, accounting for 58.3% of the world’s total apple production in 2022, ranking first in the world (<xref rid="B2" ref-type="bibr">Association, 2023</xref>). However, the growth of apple leaves is frequently threatened by pathogens such as fungi and viruses, which can lead to various diseases and result in significant economic losses (<xref rid="B1" ref-type="bibr">Ali et al., 2024</xref>). Therefore, timely detection and accurate identification of apple leaf diseases are of great importance.</p><p>Traditional apple leaf disease detection methods primarily rely on expert visual inspection and experience (<xref rid="B27" ref-type="bibr">Sheng et al., 2022</xref>), which are time-consuming, labor-intensive, and highly susceptible to subjective factors such as fatigue, expertise level, and environmental variability (<xref rid="B41" ref-type="bibr">Zhang et al., 2024</xref>). To improve efficiency, researchers have proposed machine learning-based approaches (<xref rid="B24" ref-type="bibr">Predić et al., 2022a</xref>), which extract handcrafted global features—such as color, texture, and shape—and use traditional image processing techniques combined with classifiers for disease recognition (<xref rid="B3" ref-type="bibr">Bacanin et al., 2022</xref>; <xref rid="B4" ref-type="bibr">Bukumira et al., 2022</xref>). However, these methods have notable limitations: (1) Handcrafted features often lack the descriptive power to capture local characteristics of complex lesions accurately; (2) They are sensitive to noise, lighting variations, and background clutter, resulting in unstable features and reduced classification accuracy.</p><p>With the rise of deep learning, convolutional neural networks (CNNs) have demonstrated strong performance in crop disease classification tasks (<xref rid="B37" ref-type="bibr">Wang et al., 2024</xref>; <xref rid="B23" ref-type="bibr">Petrovic et al., 2024</xref>; <xref rid="B12" ref-type="bibr">Huang et al., 2025</xref>). CNNs can automatically learn discriminative features from raw images, eliminating the need for manual feature engineering. However, existing CNN-based models still face three major challenges in apple leaf disease recognition: (1) Most models are trained on images captured in controlled laboratory environments, lacking high-quality samples collected under real-world field conditions, which limits generalization and practical deployment; (2) Many high-accuracy models are architecturally complex and have large numbers of parameters, making them difficult to deploy on resource-constrained mobile or edge devices. While model compression techniques such as pruning can partially reduce computational demands, it remains challenging to balance accuracy and efficiency (<xref rid="B25" ref-type="bibr">Predić et al., 2022b</xref>).</p><p>To address these issues, this study constructs a real-field apple leaf disease image dataset. Based on this dataset, we propose an efficient and lightweight deep neural network, named LCAMNet. The network integrates depthwise separable convolution and structural re-parameterization techniques to achieve lightweight yet effective modeling. In addition, it incorporates multi-scale downsampling and multi-scale feature extraction modules to enhance the representation of diverse lesion characteristics. An improved triplet attention mechanism is also introduced to strengthen the modeling of deep lesion features. Experimental results demonstrate that LCAMNet achieves an accuracy of 92.60% on the SCEBD dataset and 95.31% on a public dataset, while requiring only 0.03 GFLOPs and 1.30 million parameters, making it highly suitable for deployment in resource-limited environments for apple leaf disease classification.</p><p>The main contributions of this paper are as follows:</p><list list-type="order"><list-item><p>A dual-branch downsampling module is designed. Applying different downsampling operations to channels and using channel shuffle to improve feature fusion between channels. This avoids information loss caused by single downsampling strategies and improves recognition accuracy.</p></list-item><list-item><p>A multi-scale feature extraction module is proposed. Four feature extractors are designed to capture diverse features of the lesion regions from different receptive fields. In addition, channel separation is used to reduce the convolutional computation cost, and the channel shuffling method solves the information isolation issue caused by grouped convolutions, promoting feature fusion across different groups of channels.</p></list-item><list-item><p>An improved triplet attention mechanism is introduced. The original 7x7 convolution is replaced by two cascaded 3x3 convolutions, which not only enhance the deep lesion feature modeling ability but also effectively reduce the model parameter size.</p></list-item><list-item><p>A novel dataset, SCEBD, is developed by aggregating images from four distinct sources and employing a variety of data augmentation techniques. It serves to enable a comprehensive evaluation of LCAMNet and significantly enhances its generalization capability.</p></list-item></list><p>The structure of this paper is as follows: Section 2 reviews related work. Section 3 presents the dataset and the proposed model. Section 4 outlines the experimental setup and results. Section 5 provides the conclusion.</p></sec><sec id="s2" disp-level="1"><label>2.</label><title>Related work</title><p>Since extracting effective features from crop disease images is a critical and challenging task, and deep learning techniques have the capability to automatically learn features from raw images, research in this field primarily focus on designing high-performance model architectures to improve recognition accuracy (<xref rid="B19" ref-type="bibr">Liu et al., 2022b</xref>; <xref rid="B17" ref-type="bibr">Liang and Jiang, 2023</xref>; <xref rid="B14" ref-type="bibr">Li et al., 2024</xref>).</p><p>(<xref rid="B33" ref-type="bibr">Tang et al., 2024</xref>) improve the Inception module based on ResNet50 and integrate the ResNeXt inverted bottleneck module. Their model is capable of identifying seven categories of apple leaves. (<xref rid="B29" ref-type="bibr">Sun et al., 2025</xref>) develop the EMA-DeiT model based on the DeiT, achieving 99.6% accuracy on the PlantVillage dataset for classifying 10 types of tomato diseases and 98.2% accuracy on a dataset containing 6 disease types. (<xref rid="B42" ref-type="bibr">Zhang et al., 2023</xref>) introduce a Dilated Inception module into AlexNet, replacing the fully connected layer with global pooling, which effectively recognizes apple leaf diseases under small sample conditions. (<xref rid="B13" ref-type="bibr">Jiang et al., 2023</xref>) enhance the feature extraction ability for leaf diseases by integrating channel and spatial attention mechanisms to ResNet18, achieving 98.25% classification accuracy on a 5-class apple leaf disease dataset. Although these studies show good performance in terms of classification accuracy, most models have complex architectures and large numbers of parameters, which limit their deployment in real agricultural scenarios. Consequently, research has shifted toward lightweight designs.</p><p>(<xref rid="B5" ref-type="bibr">Cui et al., 2025</xref>) combine CNN and Transformer architectures, achieving high accuracy with fewer parameters. (<xref rid="B16" ref-type="bibr">Li et al., 2023</xref>) use a multi-branch structure to capture diversified features and apply residual connections between layers to ensure maximum information transfer, maintaining fewer parameters while ensuring good generalization. (<xref rid="B6" ref-type="bibr">Dong et al., 2024</xref>) introduce the ECA module into EfficientNetB0 model and apply knowledge distillation to further optimize the model, increasing accuracy without expanding model size. (<xref rid="B35" ref-type="bibr">Ullah et al., 2024</xref>) combine convolutional and ViT blocks to capture both local and global features, achieving 96.38% classification accuracy on the FGVC8 dataset. While these models have made significant progress in lightweight design, they still face challenges such as simple experimental datasets, which limit their adoption in real-world agricultural environments. Some studies have begun to focus on more challenging datasets.</p><p>(<xref rid="B36" ref-type="bibr">Wang and Cui, 2024</xref>) enhance the representational capacity of the model by modifying the convolutional kernels of ShuffleNetV2 and introducing spatial attention and Ghost modules. They also construct a dataset comprising five categories of apple leaf disease images. Experimental results show that the improved model outperforms the baseline model across multiple metrics, with the model parameters totaling only 9.8 MB. (<xref rid="B18" ref-type="bibr">Liu et al., 2022a</xref>) build the ALS module based on ShuffleNetV2 using depthwise separable convolutions and channel shuffling, which reduces the computational cost and number of parameters. Furthermore, a knowledge distillation strategy is employed to train the model, further improving its accuracy. This approach enables real-time, automated monitoring of apple leaf pests and diseases on mobile devices. (<xref rid="B20" ref-type="bibr">Liu et al., 2023</xref>) also propose a method based on MobileNetV3, in which model parameters are progressively optimized using a univariate approach. A flooding technique is introduced as a novel training strategy to prevent excessive loss minimization. This method achieves superior results on both custom and public datasets. (<xref rid="B15" ref-type="bibr">Li et al., 2025</xref>) develop a corn leaf disease recognition model based on MobileNetV3-Large, incorporating a high-frequency feature extraction (HFFE) module to integrate high-frequency image information at the network’s output stage. Additionally, the ACON-C activation function is introduced to enhance the model’s nonlinear representation capacity. Experimental results indicate a 2.1% improvement in average recognition accuracy compared to the baseline model. (<xref rid="B44" ref-type="bibr">Zheng et al., 2023</xref>) propose a network architecture optimized for both training and inference. By employing depthwise separable convolutions and structural re-parameterization techniques, along with embedding a parallel dilated attention module, the model achieves the fastest inference speed on a CPU.</p><p>Current studies have thoroughly validated the effectiveness of deep learning techniques in plant leaf disease classification, particularly highlighting their potential for application in natural environments. However, existing apple leaf disease datasets still fall short in fully capturing the diversity and complexity of real orchard conditions. Moreover, the increasing complexity of models designed to improve classification accuracy poses challenges for deployment on resource-constrained devices. Therefore, this study focuses on the construction of datasets collected under natural environmental conditions and the design of lightweight network architectures, aiming to achieve efficient model deployment while maintaining high classification accuracy, thus contributing to the development needs of smart agriculture.</p></sec><sec id="s3" disp-level="1"><label>3.</label><title>Materials and methods</title><sec id="s3_1" disp-level="2"><label>3.1.</label><title>Image acquisition and preprocessing</title><p>This study conducts experiments on two datasets: a public dataset and a self-constructed dataset with natural environmental backgrounds. The specifics of these datasets are detailed in Sections 3.1.1 and 3.1.2, respectively, while the data preprocessing process is explained in Section 3.1.3.</p><sec id="s3_1_1" disp-level="3"><label>3.1.1.</label><title>Public dataset FGVC8</title><p>The public dataset used in this study is from the CVPR 2021 FGVC8 plant pathology recognition challenge (<xref rid="B34" ref-type="bibr">Thapa et al., 2021</xref>). It consists of 18,632 field-captured apple leaf images. The images are taken at various apple maturation stages and during different times of the day, with non-uniform backgrounds. Most of the images have a resolution of 2676x4000. The dataset includes apple leaf images with various disease categories, including alternaria leaf spot, healthy, powdery mildew, rust, and scab. The selected sample sizes for each category are 489, 529, 485, 503, and 504 images, respectively. After preprocessing, these images are used to form the FGVC8 dataset, with the category distribution shown in <xref rid="T1" ref-type="table">
<bold>Table 1</bold>
</xref>.</p><table-wrap id="T1" position="float"><?disp-level 4?><label>Table 1</label><caption><p>FGVC8 dataset class distribution.</p></caption><table frame="hsides" rules="groups"><thead><tr><th valign="top" align="left" rowspan="1" colspan="1">Class number</th><th valign="top" align="left" rowspan="1" colspan="1">Image type</th><th valign="top" align="left" rowspan="1" colspan="1">Original image</th></tr></thead><tbody><tr><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">Alternaria leaf spot</td><td valign="top" align="left" rowspan="1" colspan="1">489</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">1</td><td valign="top" align="left" rowspan="1" colspan="1">Healthy</td><td valign="top" align="left" rowspan="1" colspan="1">529</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">2</td><td valign="top" align="left" rowspan="1" colspan="1">Powdery mildew</td><td valign="top" align="left" rowspan="1" colspan="1">485</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">3</td><td valign="top" align="left" rowspan="1" colspan="1">Rust</td><td valign="top" align="left" rowspan="1" colspan="1">503</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">4</td><td valign="top" align="left" rowspan="1" colspan="1">Scab</td><td valign="top" align="left" rowspan="1" colspan="1">504</td></tr></tbody></table></table-wrap></sec><sec id="s3_1_2" disp-level="3"><label>3.1.2.</label><title>Self-constructed natural environmental background dataset</title><p>This study also creates a dataset of apple leaf diseases set against a natural environmental background, called SCEBD. The dataset is compiled from four data sources: the FGVC8 dataset, Appleleaf9 (<xref rid="B40" ref-type="bibr">Yang et al., 2022</xref>), ATLDSD (<xref rid="B7" ref-type="bibr">Feng and Chao, 2022</xref>) and self-collected apple leaf disease images. Some images of alternaria leaf spot, healthy, rust, powdery mildew, and scab are from the FGVC8 dataset, with the following sample sizes: 293, 183, 154, 485, and 484 images, respectively. Images of mosaic are sourced from the Appleleaf9 dataset (a total of 105 images), which combines data from four different apple disease datasets, with varying pixel sizes. Images of gray spot are taken from the ATLDSD dataset (a total of 121 images), captured by a Glory V10 smartphone, with images taken from a real orchard, and the pixel size is 256x256. In addition, images of apple leaf diseases, including alternaria leaf spot, brown spot, gray spot, healthy, mosaic, and rust, are collected from real orchards in Yongning Town, Wafangdian City, Liaoning Province, China, using a smartphone (iQOONeo9Pro). The number of images for each disease is as follows: 192 for alternaria leaf spot, 480 for brown spot, 362 for gray spot, 302 for healthy, 376 for mosaic, and 328 for rust. These images are collected in natural environmental settings that include tree leaves, weeds, and the image resolution is not uniform. The distribution of image types and quantities in the SCEBD is shown in <xref rid="T2" ref-type="table">
<bold>Table 2</bold>
</xref>, and examples of apple leaf disease images are shown in <xref rid="f1" ref-type="fig">
<bold>Figure 1</bold>
</xref>.</p><table-wrap id="T2" position="float"><?disp-level 4?><label>Table 2</label><caption><p>SCEBD class distribution.</p></caption><table frame="hsides" rules="groups"><thead><tr><th valign="top" align="left" rowspan="1" colspan="1">Class number</th><th valign="top" align="left" rowspan="1" colspan="1">Image type</th><th valign="top" align="left" rowspan="1" colspan="1">FGVC8</th><th valign="top" align="left" rowspan="1" colspan="1">Private Data</th><th valign="top" align="left" rowspan="1" colspan="1">Appleleaf9</th><th valign="top" align="left" rowspan="1" colspan="1">ATLDSD</th><th valign="top" align="left" rowspan="1" colspan="1">Total images</th></tr></thead><tbody><tr><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">Alternaria leaf spot</td><td valign="top" align="left" rowspan="1" colspan="1">293</td><td valign="top" align="left" rowspan="1" colspan="1">192</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">485</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">1</td><td valign="top" align="left" rowspan="1" colspan="1">Brown spot</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">480</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">480</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">2</td><td valign="top" align="left" rowspan="1" colspan="1">Gray spot</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">362</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">121</td><td valign="top" align="left" rowspan="1" colspan="1">483</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">3</td><td valign="top" align="left" rowspan="1" colspan="1">Healthy</td><td valign="top" align="left" rowspan="1" colspan="1">183</td><td valign="top" align="left" rowspan="1" colspan="1">302</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">485</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">4</td><td valign="top" align="left" rowspan="1" colspan="1">Powdery mildew</td><td valign="top" align="left" rowspan="1" colspan="1">485</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">485</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">5</td><td valign="top" align="left" rowspan="1" colspan="1">Mosaic</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">376</td><td valign="top" align="left" rowspan="1" colspan="1">105</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">481</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">6</td><td valign="top" align="left" rowspan="1" colspan="1">Rust</td><td valign="top" align="left" rowspan="1" colspan="1">154</td><td valign="top" align="left" rowspan="1" colspan="1">328</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">482</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">7</td><td valign="top" align="left" rowspan="1" colspan="1">Scab</td><td valign="top" align="left" rowspan="1" colspan="1">484</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">0</td><td valign="top" align="left" rowspan="1" colspan="1">484</td></tr></tbody></table></table-wrap><fig id="f1" position="float"><?disp-level 4?><label>Figure 1</label><caption><p>Examples of apple leaf disease images: <bold>(A)</bold> Alternaria leaf spot, <bold>(B)</bold> Brown spot, <bold>(C)</bold> Gray spot, <bold>(D)</bold> Health, <bold>(E)</bold> Powdery mildew, <bold>(F)</bold> Mosaic, <bold>(G)</bold> Rust, <bold>(H)</bold> Scab.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g001.jpg"><?cloudpmc-path blobs/ce99/12378099/a4f5b2173344/fpls-16-1626569-g001.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1302?><?original-width 2076?><?scaled-height 434?><?scaled-width 692?><alt-text>A set of eight images labeled A to H, each showing leaves with different conditions. Image A shows a leaf affected by Alternaria leaf spot. Image B displays a yellow leaf with black spots, indicating Brown spot disease. Image C presents a green leaf with minor damage, associated with Gray spot. Image D shows a healthy green leaf. Image E depicts a leaf with a grayish tint, suggesting Powdery mildew infection. Image F features a leaf with yellowish patches, indicating Mosaic disease. Image G shows a leaf with a red circular spot, suggesting Rust. Image H presents a green leaf with black scab-like lesions, associated with Scab disease.</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g001.gif"><?cloudpmc-path blobs/ce99/12378099/dd87ee7aaadc/fpls-16-1626569-g001.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec><sec id="s3_1_3" disp-level="3"><label>3.1.3.</label><title>Data preprocessing</title><p>The collected image data is first cleaned to satisfy criteria, then resized to 224×224 pixels for easier computation. The dataset is split into training, validation, and test sets at a 7:2:1 ratio. Data augmentation is performed only on the training set, while validation and test sets remain unchanged. Next, normalization is performed using the mean and standard deviation for the RGB channels. The data preprocessing flowchart is shown in <xref rid="f2" ref-type="fig">
<bold>Figure 2</bold>
</xref>.</p><fig id="f2" position="float"><?disp-level 4?><label>Figure 2</label><caption><p>Data preprocessing flowchart.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g002.jpg"><?cloudpmc-path blobs/ce99/12378099/0df693a162f0/fpls-16-1626569-g002.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1428?><?original-width 998?><?scaled-height 952?><?scaled-width 665?><alt-text>Flowchart illustrating a data processing pipeline. It begins with Input, followed by Data Cleaning, and Image Scaling. The process then splits into Dataset Splitting, checking if data enhancement is needed. If yes, it proceeds to Data Enhancement, then Data Normalization, leading to Output. If no, it bypasses directly to Data Normalization before reaching Output.</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g002.gif"><?cloudpmc-path blobs/ce99/12378099/7668a7e9bc46/fpls-16-1626569-g002.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>To improve the model’s generalization and minimize noise interference, nine types of data augmentation are used on the training images. These include rotations (90°, 180°, and 270°), gaussian blur, random flips (50% probability for both horizontal and vertical flips), contrast enhancement and reduction, and brightness enhancement and reduction. These augmentations increase the number of training images to 10 times the original size. No data augmentation is applied to the test and validation sets. Examples of the augmented images are shown in <xref rid="f3" ref-type="fig">
<bold>Figure 3</bold>
</xref>.</p><fig id="f3" position="float"><?disp-level 4?><label>Figure 3</label><caption><p>Examples of data augmentation for apple leaf disease images: <bold>(A)</bold> Original, <bold>(B)</bold> Rotated 90, <bold>(C)</bold> Rotated 180, <bold>(D)</bold> Rotated 270, <bold>(E)</bold> Blurred, <bold>(F)</bold> Horizontal flip, <bold>(G)</bold> Contrast high, <bold>(H)</bold> Contrast low, <bold>(I)</bold> Brightness high, <bold>(J)</bold> Brightness low.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g003.jpg"><?cloudpmc-path blobs/ce99/12378099/b4d73eb80cc6/fpls-16-1626569-g003.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1141?><?original-width 2000?><?scaled-height 456?><?scaled-width 800?><alt-text>A series of ten images labeled A to J, each showing a green leaf with a brown spot on a tree branch. The images depict the following variations: (A) original, (B) rotated 90°, (C) rotated 180°, (D) rotated 270°, (E) blurred, (F) horizontally flipped, (G) high contrast, (H) low contrast, (I) high brightness, and (J) low brightness. The background in all images is consistently blurred greenery.</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g003.gif"><?cloudpmc-path blobs/ce99/12378099/505560eab62b/fpls-16-1626569-g003.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>To prevent instability in model training caused by excessively large or small pixel values and to reduce the risk of overfitting, the images are normalized. Specifically, the mean values for the red, green, and blue channels are set to [0.485, 0.456, 0.406], and the standard deviations are [0.229, 0.224, 0.225]. This normalization method helps accelerate model convergence, improves training stability, and enables more efficient learning of image features.</p></sec></sec><sec id="s3_2" disp-level="2"><label>3.2.</label><title>LCAMNet model</title><sec id="s3_2_1" disp-level="3"><label>3.2.1.</label><title>Model structure</title><p>To address the problem of apple leaf disease classification in natural environments, this paper proposes a lightweight converged attention multi-branch network (LCAMNet). The model includes a 3×3 standard convolutional layer (Conv1), a max pooling layer (MaxPooling), three stage modules (Stage2 to Stage4), a 1×1 convolutional layer (Conv5), a global pooling layer (Global Pooling), and a fully connected layer (FC). The overall structure of the model is shown in <xref rid="f4" ref-type="fig">
<bold>Figure 4</bold>
</xref>.</p><fig id="f4" position="float"><?disp-level 4?><label>Figure 4</label><caption><p>Architecture diagram of LCAMNet.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g004.jpg"><?cloudpmc-path blobs/ce99/12378099/1035d5a2899e/fpls-16-1626569-g004.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1059?><?original-width 2066?><?scaled-height 353?><?scaled-width 688?><alt-text>Flowchart illustrating the LCAMNet architecture. It starts with an input image, followed by layers labeled Conv1 and MaxPooling. The network progresses through stages two to four, each containing an MSDM and an MFEN module; additionally, stage four includes an Attention module. Afterward, the data passes through the Conv5 layer. The process concludes with Global Average Pooling, a Fully Connected layer, and an Output.</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g004.gif"><?cloudpmc-path blobs/ce99/12378099/fe6f630114d1/fpls-16-1626569-g004.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>LCAMNet initially processes the input image through Conv1 to extract basic features. It then uses MaxPooling for downsampling, reducing the dimensionality of the output features. The subsequent stages, Stage2 to Stage4, contain several key modules, each consisting of a dual-branch downsampling module (DBDM) and a multi-scale feature extraction module (MFEM). The DBDM module performs downsampling on the extracted features, improving computational efficiency while retaining key features. On top of this, the MFEM module captures multi-scale features related to the disease. Both the DBDM and MFEM modules are stacked three times, with each stack extracting deeper features from the previous layer. At the end of Stage4, LCAMNet introduces an improved triplet attention mechanism to further extract crucial feature information. The model then connects a convolutional layer for feature fusion, followed by a global pooling layer and a fully connected layer for classification, outputting the final class results.</p></sec><sec id="s3_2_2" disp-level="3"><label>3.2.2.</label><title>Dual-branch downsampling module</title><p>Downsampling is often used in convolutional neural networks (CNNs) to reduce the spatial size of feature maps. Pooling is one of the most commonly used downsampling methods, which aggregates pixel values in a local region to decrease the spatial dimensions of feature map, thereby improving model’s robustness to translation variations. Unlikepooling, convolution operations learn convolutional kernel parameters to extract image features, retaining more useful information during the downsampling process and adapting better to various tasks (<xref rid="B16" ref-type="bibr">Li et al., 2023</xref>).</p><p>However, in natural environmental backgrounds, the classification performance of apple leaf disease images is often weak. A single downsampling strategy may cause the loss of important image features, thus affecting the classification results. To address this challenge, LCAMNet incorporates a dual-branch downsampling module, whose structure is shown in <xref rid="f5" ref-type="fig">
<bold>Figure 5</bold>
</xref>. This module performs downsampling on the input feature map through both a pooling branch and a convolutional branch. It then concatenates the output features and applies a channel shuffling operation to achieve feature fusion, producing the final output features. The specific design of the two branches is described in the following paragraphs.</p><fig id="f5" position="float"><?disp-level 4?><label>Figure 5</label><caption><p>Architecture of DBDM.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g005.jpg"><?cloudpmc-path blobs/ce99/12378099/cc3da64a904c/fpls-16-1626569-g005.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1982?><?original-width 2038?><?scaled-height 660?><?scaled-width 679?><alt-text>Flowchart of DSDM. The input is processed through two separate branches: a pooling branch and a convolution branch. The pooling branch performs average and max pooling, followed by concatenation and a 1×1 convolution unit. The convolution branch executes parallel 3×3 depthwise convolutions, followed by batch normalization, a 1×1 convolution unit, and concatenation. Finally, the outputs from both branches are concatenated, undergo channel shuffling, and produce the final output.</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g005.gif"><?cloudpmc-path blobs/ce99/12378099/6144b710e4b6/fpls-16-1626569-g005.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>The pooling branch applies max pooling and average pooling to extract global and salient features from apple leaf disease images. Specifically, average pooling averages the pixels within pooling region, effectively suppressing local noise and enabling stable feature extraction of diseased areas under complex background interference. In contrast, max pooling selects the maximum value from the pooling window to extract salient features, enhancing the most representative texture and morphological changes in the diseased area. This is crucial for distinguishing apple leaf lesions from the salient regions in the complex background. By concatenating the results of max pooling and average pooling, the module achieves a complementary combination of global and salient features. Subsequently, a 1×1 convolution unit is used to aggregate channels, which includes a 1×1 convolution, batch normalization (BN), and a ReLU activation function. This step enhances model’s ability to capture nonlinear feature representations, as shown in <xref rid="f6" ref-type="fig">
<bold>Figure 6</bold>
</xref>.</p><fig id="f6" position="float"><?disp-level 4?><label>Figure 6</label><caption><p>Architecture diagram of the 1×1 conv unit.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g006.jpg"><?cloudpmc-path blobs/ce99/12378099/4c21d710f2c2/fpls-16-1626569-g006.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1019?><?original-width 886?><?scaled-height 679?><?scaled-width 590?><alt-text>Flowchart illustrating the sequence of 1x1 conv unit, starting with “Input,” followed by “1×1 Convolution,” “Batch Normalization (BN),” “ReLU Activation,” and ending with “Output.” Arrows indicate the flow of the process between each step.</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g006.gif"><?cloudpmc-path blobs/ce99/12378099/d85efdce23c3/fpls-16-1626569-g006.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>The convolutional branch consists of convolutional downsampling modules, which aim to reduce dimensions while effectively extracting feature information. Specifically, a 3×3 depthwise convolution with a stride of 2 is used to extract local features, halving the spatial resolution. The parallel convolutional branches double the number of channels via channel concatenation to maintain the same channel dimension as the pooling branch. By concatenating the features extracted from both branches and applying channel shuffling, the model enhances the interaction between channels, effectively preserving the key features of the apple leaf disease areas. The pseudocode for this module is shown in <xref rid="algo1" ref-type="boxed-text">
<bold>Algorithm 1</bold>
</xref>.</p><boxed-text id="algo1" position="float"><?disp-level 4?><label>Algorithm 1</label><caption><title>Downsampling Process of DBDM.</title></caption><p>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g014.jpg"><?cloudpmc-path blobs/ce99/12378099/df5a6c0b4746/fpls-16-1626569-g014.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 885?><?original-width 1024?><?scaled-height 589?><?scaled-width 682?></graphic>
</p></boxed-text></sec><sec id="s3_2_3" disp-level="3"><label>3.2.3.</label><title>Multi-Scale Feature Extraction Module</title><p>Apple leaf disease classification is often weak in complex environments, a single-size convolutional kernel may not effectively extract image features. To address this issue, this study constructs a Multi-Scale Feature Extraction Module (MFEM). Its architecture is shown in <xref rid="f7" ref-type="fig">
<bold>Figure 7</bold>
</xref>.</p><fig id="f7" position="float"><?disp-level 4?><label>Figure 7</label><caption><p>Architecture of MFEM.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g007.jpg"><?cloudpmc-path blobs/ce99/12378099/63c49267d626/fpls-16-1626569-g007.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 2230?><?original-width 1830?><?scaled-height 892?><?scaled-width 732?><alt-text>Flowchart showing the MFEN architecture. It begins with an “Input” node leading to a “Channel Split,” which creates four paths: (1) a skip connection passing input directly, (2) 1×1 Conv Unit, 3×3 DWConv, BN, 1×1 Conv Unit, (3) similar to (2) but with two consecutive 3×3 DWConv layers separated by BN, and (4) MaxPooling, 1×1 Conv Unit. All paths merge at “Concat,” followed by “Channel Shuffle,” and end with “Output.”</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g007.gif"><?cloudpmc-path blobs/ce99/12378099/e53d150bc9b3/fpls-16-1626569-g007.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>The input features first go through a channel separation operation, evenly dividing the channels into four independent branches: the feature-preserving branch, the local detail branch, the deep feature branch, and the salient feature branch. These branches are arranged from left to right, each extracting features at different scales of apple leaf disease image. The feature maps from each branch are fused through concatenation and channel shuffle to improve channel interaction. The detailed design of each branch is explained below.</p><p>In the feature-preserving branch, input features are directly forwarded via a skip connection to produce output. This helps distinguish low-level information and mitigates the degradation problem of deep networks. This branch is crucial for capturing subtle local variations in the image features. For example, in apple leaf disease classification, it helps distinguish minor differences between background noise (such as lighting changes, weeds, or shadows) and disease areas (such as brown spots or leaf discoloration).</p><p>In the local detail branch, the input feature map first goes through a 1×1 conv unit to adjust the number of channels. Then, the feature map is passed through a 3×3 depthwise convolution and BN. The depthwise convolution performs independently on each input channel, enabling the extraction of fine-grained features within each channel. Finally, a 1×1 conv unit is applied to fuse the features extracted by the depthwise convolution from each channel. This branch mainly extracts features related to the edges of lesions, local texture changes, and small disease regions in apple leaf disease images.</p><p>In the deep feature branch, input feature map first goes through a 1×1 conv unit, followed by two 3×3 depthwise convolutions and BN. Afterward, another 1×1 conv unit is applied. The cascading 3×3 convolution layers form a deeper feature extraction unit, capable of capturing more complex local patterns and multi-layered details, such as lesions of varying sizes and morphological changes in apple leaf disease images.</p><p>In the salient feature branch, the input feature map passes through a max-pooling layer to extract features, followed by a 1×1 conv unit to adjust the number of channels. In natural settings, max pooling helps enhance the network’s ability to detect salient disease regions in apple leaf images, especially when the contrast between the background and disease features is low, making it more effective in highlighting key features.</p><p>At last, the feature maps from all four branches are concatenated to merge multi-scale features, constructing a richer global representation. Additionally, channel shuffling is introduced to promote cross-branch feature exchange and reorganization, thereby enhancing the module’s ability to represent features of the apple leaf disease regions. The pseudocode for this module is shown in <xref rid="algo2" ref-type="boxed-text">
<bold>Algorithm 2</bold>
</xref>.</p><boxed-text id="algo2" position="float"><?disp-level 4?><label>Algorithm 2</label><caption><title>The feature extraction process of MFEN.</title></caption><p>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g015.jpg"><?cloudpmc-path blobs/ce99/12378099/e7f92adc4948/fpls-16-1626569-g015.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1150?><?original-width 1024?><?scaled-height 766?><?scaled-width 682?></graphic>
</p></boxed-text></sec><sec id="s3_2_4" disp-level="3"><label>3.2.4.</label><title>Improved triplet attention mechanism</title><p>In recent years, attention mechanisms demonstrate significant advantages in computer vision by helping the model focus on important regions, thus enhancing classification accuracy. Although this study introduces multi-scale branches in the feature extraction layer to capture rich features, challenges remain in capturing deeper features. To solve this, this paper introduces an attention mechanism that captures information from different dimensions, improving the model’s ability to detect apple leaf disease regions.</p><p>The triplet attention mechanism (<xref rid="B21" ref-type="bibr">Misra et al., 2021</xref>) captures interactions between channels and spatial dimensions through two branches. A third branch is used to create spatial attention, and the outputs from all three branches are combined to form the final attention features. This paper improves upon the original triplet attention mechanism. Specifically, we replace the original 7×7 convolution used for feature extraction with two cascading 3×3 convolutions. This modification allows us to maintain sufficient feature extraction capability without increasing the computational burden. Additionally, we introduce the ReLU activation function (<xref rid="B22" ref-type="bibr">Nair and Hinton, 2010</xref>) between the two 3×3 convolutions, which effectively controls the network’s sparsity and enhances its ability for nonlinear transformations. It is worth noting that we remove the original BN module. Experimental analysis reveals that the BN module does not have a significant effect and, in fact, increases computational load. Through these improvements, we significantly enhance the model’s efficiency and practical application performance while ensuring effective cross-dimensional feature modeling. The structure of the improved mechanism is shown in <xref rid="f8" ref-type="fig">
<bold>Figure 8</bold>
</xref>.</p><fig id="f8" position="float"><?disp-level 4?><label>Figure 8</label><caption><p>Architecture of the improved triplet attention mechanism.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g008.jpg"><?cloudpmc-path blobs/ce99/12378099/08dca0f1e518/fpls-16-1626569-g008.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 2326?><?original-width 2078?><?scaled-height 775?><?scaled-width 692?><alt-text>Flowchart of the improved triplet attention mechanism, consisting of three parallel branches that compute attention along the H, W, and C dimensions. Each branch sequentially includes Z-Pool, a 3×3 convolution, ReLU activation, and another 3×3 convolution, replacing the original 7×7 convolution and removing the BN layer. After processing, the outputs of all branches pass through a Sigmoid layer, where the input is multiplied element-wise with the Sigmoid output. Finally, the outputs of all branches are merged after dimension permutation to complete attention feature fusion.</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g008.gif"><?cloudpmc-path blobs/ce99/12378099/7d9d2460fd8f/fpls-16-1626569-g008.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec></sec></sec><sec id="s4" disp-level="1"><label>4.</label><title>Experimental analysis</title><sec id="s4_1" disp-level="2"><label>4.1.</label><title>Experimental environment</title><p>The experimental hardware in this study uses an Intel(R) Core(TM) i7–6700 CPU @ 3.40GHz processor, and the operating system is Windows 10. Model training and testing are accelerated using a GPU, specifically a Tesla V100S-PCIE-32GB graphics card. The software environment includes Python 3.8.19, the PyTorch 2.4.1 framework, and the CUDA toolkit 12.4.</p><p>The number of iterations is set to Epoch=60, with a batch size of 32. The model training uses the Stochastic Gradient Descent (SGD) algorithm, which is one of the most commonly used optimization methods in machine learning. The initial learning rate for SGD is set to 0.0001. A cosine annealing schedule is applied for the first 50 epochs, gradually decreasing the learning rate from 1e-4 to 1e-6, and the learning rate is kept constant at 1e-6 from epoch 50 to 60.</p></sec><sec id="s4_2" disp-level="2"><label>4.2.</label><title>Evaluation criterion</title><p>In this study, six evaluation metrics are employed to assess the performance of the proposed model: Accuracy, Precision, Recall, F1-Score, Kappa, and Matthews Correlation Coefficient (MCC). Their definitions are as follows: Accuracy refers to the ratio of correctly predicted samples to the total number of samples. Precision is the proportion of correctly predicted positive samples among all samples predicted as positive. Recall represents the proportion of correctly predicted positive samples among all actual positive samples. F1-Score is the harmonic mean of Precision and Recall. Kappa measures the agreement between the model’s predictions and the ground truth, while accounting for agreement occurring by chance, thus providing a more objective evaluation of classification performance. MCC evaluates the overall performance of a classification model, particularly suitable for handling imbalanced class distributions. The mathematical formulations of Accuracy, Precision, Recall, F1-Score, Kappa, and MCC are defined as follows (see <xref rid="eq1" ref-type="disp-formula">Equations 1</xref>–<xref rid="eq6" ref-type="disp-formula">6</xref>).</p><disp-formula id="eq1"><label>(1)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M1" display="block" overflow="scroll"><mml:mrow><mml:mtext>Accuracy</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mo>×</mml:mo><mml:mn>100</mml:mn><mml:mo>%</mml:mo></mml:mrow></mml:math></disp-formula><disp-formula id="eq2"><label>(2)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M2" display="block" overflow="scroll"><mml:mrow><mml:mtext>Precision</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac><mml:mo>×</mml:mo><mml:mn>100</mml:mn><mml:mo>%</mml:mo></mml:mrow></mml:math></disp-formula><disp-formula id="eq3"><label>(3)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M3" display="block" overflow="scroll"><mml:mrow><mml:mtext>Recall</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mo>×</mml:mo><mml:mn>100</mml:mn><mml:mo>%</mml:mo></mml:mrow></mml:math></disp-formula><disp-formula id="eq4"><label>(4)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M4" display="block" overflow="scroll"><mml:mrow><mml:mtext>F</mml:mtext><mml:mn>1</mml:mn><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>×</mml:mo><mml:mtext>Precision</mml:mtext><mml:mo>×</mml:mo><mml:mtext>Recall</mml:mtext></mml:mrow><mml:mrow><mml:mtext>Precision</mml:mtext><mml:mo>+</mml:mo><mml:mtext>Recall</mml:mtext></mml:mrow></mml:mfrac><mml:mo>×</mml:mo><mml:mn>100</mml:mn><mml:mo>%</mml:mo></mml:mrow></mml:math></disp-formula><disp-formula id="eq5"><label>(5)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M5" display="block" overflow="scroll"><mml:mrow><mml:mtext>Kappa</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>−</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:mrow></mml:mfrac><mml:mo>×</mml:mo><mml:mn>100</mml:mn><mml:mo>%</mml:mo></mml:mrow></mml:math></disp-formula><disp-formula id="eq6"><label>(6)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M6" display="block" overflow="scroll"><mml:mrow><mml:mtext>MCC</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>×</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>−</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>×</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac><mml:mo>×</mml:mo><mml:mn>100</mml:mn><mml:mo>%</mml:mo></mml:mrow></mml:math></disp-formula><p>TP means the model correctly predicts a positive sample. TN means the model correctly predicts a negative sample. FP means the model wrongly predicts a negative sample as positive. FN means the model wrongly predicts a positive sample as negative. <italic>p</italic>
<sub>0</sub> represents the observed proportion of agreement, i.e., the percentage of instances where the raters reach consensus across all samples. <italic>p<sub>e</sub>
</italic> represents the expected agreement by chance, assuming that the raters classify independently and randomly. In multi-class classification, macro and micro averages are common for evaluation. This study adopts macro average to compute Recall and Precision, and uses the global average value for Accuracy. For example, the macro-averaged Recall is calculated as shown in <xref rid="eq7" ref-type="disp-formula">Equation 7</xref>.</p><disp-formula id="eq7"><label>(7)</label><mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" id="M7" display="block" overflow="scroll"><mml:mrow><mml:mtext>Macro</mml:mtext><mml:mi> </mml:mi><mml:mtext>Recall</mml:mtext><mml:mi> </mml:mi><mml:mo>=</mml:mo><mml:mi> </mml:mi><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:mtext> </mml:mtext><mml:mstyle displaystyle="true"><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mstyle></mml:mrow></mml:math></disp-formula></sec><sec id="s4_3" disp-level="2"><label>4.3.</label><title>Ablation study on SCEBD</title><p>To validate the effectiveness of the proposed method, ablation experiments are conducted on the SCEBD dataset. LCAMNet is an improved architecture based on the ShuffleNetV2 model (<xref rid="B43" ref-type="bibr">Zhang et al., 2018</xref>). The overall network structure is shown in <xref rid="T3" ref-type="table">
<bold>Table 3</bold>
</xref>.</p><table-wrap id="T3" position="float"><?disp-level 3?><label>Table 3</label><caption><p>LCAMNet architecture diagram.</p></caption><table frame="hsides" rules="groups"><thead><tr><th valign="top" align="left" rowspan="1" colspan="1">Layer</th><th valign="top" align="left" rowspan="1" colspan="1">Output size</th><th valign="top" align="left" rowspan="1" colspan="1">Attention</th><th valign="top" align="left" rowspan="1" colspan="1">Ksize</th><th valign="top" align="left" rowspan="1" colspan="1">Stride</th><th valign="top" align="left" rowspan="1" colspan="1">Repeat</th><th valign="top" align="left" rowspan="1" colspan="1">Output channels</th></tr></thead><tbody><tr><td valign="top" align="left" rowspan="1" colspan="1">Image</td><td valign="top" align="left" rowspan="1" colspan="1">224x224</td><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1">3</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">Conv1</td><td valign="top" align="left" rowspan="1" colspan="1">112x112</td><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1">3x3</td><td valign="top" align="left" rowspan="1" colspan="1">2</td><td valign="top" align="left" rowspan="1" colspan="1">1</td><td valign="top" align="left" rowspan="1" colspan="1">24</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">MaxPool</td><td valign="top" align="left" rowspan="1" colspan="1">56x56</td><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1">3x3</td><td valign="top" align="left" rowspan="1" colspan="1">2</td><td valign="top" align="left" rowspan="1" colspan="1">1</td><td valign="top" align="left" rowspan="1" colspan="1">24</td></tr><tr><td valign="middle" align="left" rowspan="1" colspan="1">stage2</td><td valign="top" align="left" rowspan="1" colspan="1">28x28<break/>28x28</td><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1">2<break/>1</td><td valign="top" align="left" rowspan="1" colspan="1">1<break/>1</td><td valign="top" align="left" rowspan="1" colspan="1">48<break/>48</td></tr><tr><td valign="middle" align="left" rowspan="1" colspan="1">Stage3</td><td valign="top" align="left" rowspan="1" colspan="1">14x14<break/>14x14</td><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1">2<break/>1</td><td valign="top" align="left" rowspan="1" colspan="1">1<break/>1</td><td valign="top" align="left" rowspan="1" colspan="1">96<break/>96</td></tr><tr><td valign="middle" align="left" rowspan="1" colspan="1">Stage4</td><td valign="top" align="left" rowspan="1" colspan="1">7x7<break/>7x7<break/>7x7</td><td valign="bottom" align="left" rowspan="1" colspan="1">1</td><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1">2<break/>1</td><td valign="top" align="left" rowspan="1" colspan="1">1<break/>1<break/>1</td><td valign="top" align="left" rowspan="1" colspan="1">192<break/>192<break/>192</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">Conv5</td><td valign="top" align="left" rowspan="1" colspan="1">7x7</td><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1">1x1</td><td valign="top" align="left" rowspan="1" colspan="1">1</td><td valign="top" align="left" rowspan="1" colspan="1">1</td><td valign="top" align="left" rowspan="1" colspan="1">1024</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">GlobalPool</td><td valign="top" align="left" rowspan="1" colspan="1">1x1</td><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1">7x7</td><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1"/></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">FC</td><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1"/><td valign="top" align="left" rowspan="1" colspan="1">8</td></tr></tbody></table></table-wrap><p>To verify the effectiveness of each proposed module, multiple comparative models are designed for experimentation. The baseline model is ShuffleNetV2. Model1 modifies only the downsampling module (SDSM) of ShuffleNetV2. Model2 modifies only the feature extraction module (SFEM). Model3 modifies both the feature extraction and downsampling modules based on ShuffleNetV2. Model4 further incorporates the improved triplet attention mechanism on top of Model3. The detailed configurations of the networks using different strategies are shown in <xref rid="T4" ref-type="table">
<bold>Table 4</bold>
</xref>.</p><table-wrap id="T4" position="float"><?disp-level 3?><label>Table 4</label><caption><p>Network configurations with different strategies.</p></caption><table frame="hsides" rules="groups"><thead><tr><th valign="top" align="left" rowspan="1" colspan="1">Model</th><th valign="top" align="left" rowspan="1" colspan="1">Feature extraction</th><th valign="top" align="left" rowspan="1" colspan="1">Downsampling</th><th valign="top" align="left" rowspan="1" colspan="1">Attention</th></tr></thead><tbody><tr><td valign="top" align="left" rowspan="1" colspan="1">1</td><td valign="top" align="left" rowspan="1" colspan="1">SFEM</td><td valign="top" align="left" rowspan="1" colspan="1">DBDM</td><td valign="top" align="left" rowspan="1" colspan="1">\</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">2</td><td valign="top" align="left" rowspan="1" colspan="1">MFEM</td><td valign="top" align="left" rowspan="1" colspan="1">SDSM</td><td valign="top" align="left" rowspan="1" colspan="1">\</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">3</td><td valign="top" align="left" rowspan="1" colspan="1">MFEM</td><td valign="top" align="left" rowspan="1" colspan="1">DBDM</td><td valign="top" align="left" rowspan="1" colspan="1">\</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">4</td><td valign="top" align="left" rowspan="1" colspan="1">MFEM</td><td valign="top" align="left" rowspan="1" colspan="1">DBDM</td><td valign="top" align="left" rowspan="1" colspan="1">Improved triplet attention</td></tr></tbody></table></table-wrap><p>The experimental results are shown in <xref rid="T5" ref-type="table">
<bold>Table 5</bold>
</xref>. GFLOPs measures the floating-point operations (in billions) and is used to evaluate the model’s computational complexity. Parameters refer to the number of trainable weights in the model and are commonly used to assess the model’s size. As seen in <xref rid="T5" ref-type="table">
<bold>Table 5</bold>
</xref>, Model1 outperforms the baseline, indicating that applying DBDM in the downsampling process of apple leaf disease images better preserves feature information compared to the original SDSM. Model2 also shows better performance than the baseline on the test set, suggesting that MFEM is more effective in extracting features from apple leaf disease images in natural environmental backgrounds than SFEM. Model3 performs better than both Model1 and Model2 on the test set, demonstrating that the combined application of MFEM and DBDM has a synergistic enhancement effect on apple leaf disease image classification, surpassing the performance of each method individually. Model4 shows further improvement over Model3. The addition of the enhanced triplet attention mechanism slightly increases the number of parameters but leads to a significant boost in performance. Additionally, compared to ShuffleNetV2, Model4 reduces both floating-point operations and parameter count while achieving a larger improvement in classification performance. In conclusion, through the stepwise introduction of DBDM, MFEM, and the improved triplet attention mechanism, the classification performance of the model continues to improve, validating the effectiveness of the proposed method.</p><table-wrap id="T5" position="float"><?disp-level 3?><label>Table 5</label><caption><p>Ablation study on the SCEBD.</p></caption><table frame="hsides" rules="groups"><thead><tr><th valign="top" align="left" rowspan="1" colspan="1">Models</th><th valign="top" align="left" rowspan="1" colspan="1">Accuracy (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Recall (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Precision (%)</th><th valign="top" align="left" rowspan="1" colspan="1">F1 (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Kappa (%)</th><th valign="top" align="left" rowspan="1" colspan="1">MCC (%)</th><th valign="top" align="left" rowspan="1" colspan="1">GFLOPs (G)</th><th valign="top" align="left" rowspan="1" colspan="1">Parameters (M)</th></tr></thead><tbody><tr><td valign="top" align="left" rowspan="1" colspan="1">Baseline</td><td valign="top" align="left" rowspan="1" colspan="1">88.01</td><td valign="top" align="left" rowspan="1" colspan="1">88.00</td><td valign="top" align="left" rowspan="1" colspan="1">88.98</td><td valign="top" align="left" rowspan="1" colspan="1">88.21</td><td valign="top" align="left" rowspan="1" colspan="1">86.30</td><td valign="top" align="left" rowspan="1" colspan="1">86.48</td><td valign="top" align="left" rowspan="1" colspan="1">0.04</td><td valign="top" align="left" rowspan="1" colspan="1">1.31</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">Model1</td><td valign="top" align="left" rowspan="1" colspan="1">88.78</td><td valign="top" align="left" rowspan="1" colspan="1">88.80</td><td valign="top" align="left" rowspan="1" colspan="1">89.46</td><td valign="top" align="left" rowspan="1" colspan="1">88.97</td><td valign="top" align="left" rowspan="1" colspan="1">87.17</td><td valign="top" align="left" rowspan="1" colspan="1">87.24</td><td valign="top" align="left" rowspan="1" colspan="1">0.03</td><td valign="top" align="left" rowspan="1" colspan="1">1.29</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">Model2</td><td valign="top" align="left" rowspan="1" colspan="1">91.33</td><td valign="top" align="left" rowspan="1" colspan="1">91.36</td><td valign="top" align="left" rowspan="1" colspan="1">91.83</td><td valign="top" align="left" rowspan="1" colspan="1">91.43</td><td valign="top" align="left" rowspan="1" colspan="1">90.09</td><td valign="top" align="left" rowspan="1" colspan="1">90.13</td><td valign="top" align="left" rowspan="1" colspan="1">0.03</td><td valign="top" align="left" rowspan="1" colspan="1">1.28</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">Model3</td><td valign="top" align="left" rowspan="1" colspan="1">92.35</td><td valign="top" align="left" rowspan="1" colspan="1">92.34</td><td valign="top" align="left" rowspan="1" colspan="1">92.73</td><td valign="top" align="left" rowspan="1" colspan="1">92.41</td><td valign="top" align="left" rowspan="1" colspan="1">91.25</td><td valign="top" align="left" rowspan="1" colspan="1">91.28</td><td valign="top" align="left" rowspan="1" colspan="1">0.03</td><td valign="top" align="left" rowspan="1" colspan="1">1.29</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">Model4</td><td valign="top" align="left" rowspan="1" colspan="1">92.60</td><td valign="top" align="left" rowspan="1" colspan="1">92.64</td><td valign="top" align="left" rowspan="1" colspan="1">92.85</td><td valign="top" align="left" rowspan="1" colspan="1">92.64</td><td valign="top" align="left" rowspan="1" colspan="1">91.55</td><td valign="top" align="left" rowspan="1" colspan="1">91.58</td><td valign="top" align="left" rowspan="1" colspan="1">0.03</td><td valign="top" align="left" rowspan="1" colspan="1">1.30</td></tr></tbody></table></table-wrap><p>To further validate the effectiveness of the LCAMNet model in disease region identification, this study employs the Grad-CAM (<xref rid="B26" ref-type="bibr">Selvaraju et al., 2017</xref>) technique to visualize and analyze the model’s prediction results, comparing the performance before and after the improvements, as shown in <xref rid="f9" ref-type="fig">
<bold>Figure 9</bold>
</xref>. <xref rid="f9" ref-type="fig">
<bold>Figure 9A</bold>
</xref> presents the original image, <xref rid="f9" ref-type="fig">
<bold>Figure 9B</bold>
</xref> shows the visualization result of the baseline model, and <xref rid="f9" ref-type="fig">
<bold>Figure 9C</bold>
</xref> shows the visualization result of the LCAMNet model. Taking a rust disease image as an example, the baseline model focuses only on partial features of the diseased area, leading to missed and false detections, and fails to fully capture the lesion regions. In contrast, the LCAMNet model accurately localizes the key diseased regions of the apple leaf, with more comprehensive and discriminative feature extraction. These results demonstrate that the integration of the MFEM module, DBDM module, and the improved triplet attention mechanism in LCAMNet significantly enhances the model’s ability to focus on diseased areas, thereby improving its feature learning capacity and classification accuracy.</p><fig id="f9" position="float"><?disp-level 3?><label>Figure 9</label><caption><p>Grad-CAM visualization results of different models: <bold>(A)</bold> Original image; <bold>(B)</bold> Heatmap generated by the Baseline model; <bold>(C)</bold> Heatmap generated by Model4.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g009.jpg"><?cloudpmc-path blobs/ce99/12378099/a09b19080af7/fpls-16-1626569-g009.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 780?><?original-width 2048?><?scaled-height 260?><?scaled-width 682?><alt-text>(A) Close-up of green leaves with some yellow spots and a small green fruit in the background. (B) Heatmap overlay generated using the ShuffleNetV2 model, highlighting red and yellow areas on the leaves. (C) Heatmap overlay generated using the LCAMNet model, with red regions more accurately focused on the leaf lesions compared to (B).</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g009.gif"><?cloudpmc-path blobs/ce99/12378099/d60fa0b41622/fpls-16-1626569-g009.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec><sec id="s4_4" disp-level="2"><label>4.4.</label><title>Performance comparison with classic convolutional neural networks on SCEBD</title><p>In this experiment, we compare LCAMNet with several classic convolutional neural network models, including VGG (<xref rid="B28" ref-type="bibr">Simonyan and Zisserman, 2014</xref>), ResNet (<xref rid="B9" ref-type="bibr">He et al., 2016</xref>), ResNext (<xref rid="B38" ref-type="bibr">Xie et al., 2017</xref>), and DenseNet (<xref rid="B11" ref-type="bibr">Huang et al., 2017</xref>), on SCEBD. The experimental results are shown in <xref rid="T6" ref-type="table">
<bold>Table 6</bold>
</xref>. These models exhibit issues when applied to disease classification tasks in natural environments, which are manifested in the following aspects: First, while VGG network has a simple structure, it suffers from a large number of parameters due to its deep architecture, making it prone to overfitting. As a result, its performance is relatively poor in the scenarios with small datasets or complex backgrounds. Second, ResNet and ResNext use residual connections to address the vanishing gradient problem. However, due to the deep network layers, they still struggle to effectively capture features in the presence of high noise and irregular lesions. DenseNet, despite having advantages in feature propagation, has a dense connectivity structure that leads to a large number of parameters and computational overhead, making it inefficient when processing high-resolution images. Thus, these four networks suffer from low computational efficiency and high model complexity.</p><table-wrap id="T6" position="float"><?disp-level 3?><label>Table 6</label><caption><p>Performance comparison of LCAMNet and several classical CNN models on the SCEBD.</p></caption><table frame="hsides" rules="groups"><thead><tr><th valign="top" align="left" rowspan="1" colspan="1">Models</th><th valign="top" align="left" rowspan="1" colspan="1">Accuracy (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Recall (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Precision (%)</th><th valign="top" align="left" rowspan="1" colspan="1">F1 (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Kappa (%)</th><th valign="top" align="left" rowspan="1" colspan="1">MCC (%)</th><th valign="top" align="left" rowspan="1" colspan="1">GFLOPs (G)</th><th valign="top" align="left" rowspan="1" colspan="1">Parameters (M)</th></tr></thead><tbody><tr><td valign="top" align="left" rowspan="1" colspan="1">VGG13</td><td valign="top" align="left" rowspan="1" colspan="1">89.03</td><td valign="top" align="left" rowspan="1" colspan="1">89.04</td><td valign="top" align="left" rowspan="1" colspan="1">89.15</td><td valign="top" align="left" rowspan="1" colspan="1">89.03</td><td valign="top" align="left" rowspan="1" colspan="1">87.64</td><td valign="top" align="left" rowspan="1" colspan="1">87.53</td><td valign="top" align="left" rowspan="1" colspan="1">11.36</td><td valign="top" align="left" rowspan="1" colspan="1">133.05</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">ResNet34</td><td valign="top" align="left" rowspan="1" colspan="1">88.52</td><td valign="top" align="left" rowspan="1" colspan="1">88.54</td><td valign="top" align="left" rowspan="1" colspan="1">89.02</td><td valign="top" align="left" rowspan="1" colspan="1">88.66</td><td valign="top" align="left" rowspan="1" colspan="1">86.88</td><td valign="top" align="left" rowspan="1" colspan="1">86.94</td><td valign="top" align="left" rowspan="1" colspan="1">3.68</td><td valign="top" align="left" rowspan="1" colspan="1">21.80</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">ResNext50</td><td valign="top" align="left" rowspan="1" colspan="1">85.97</td><td valign="top" align="left" rowspan="1" colspan="1">85.99</td><td valign="top" align="left" rowspan="1" colspan="1">86.48</td><td valign="top" align="left" rowspan="1" colspan="1">86.12</td><td valign="top" align="left" rowspan="1" colspan="1">83.96</td><td valign="top" align="left" rowspan="1" colspan="1">84.00</td><td valign="top" align="left" rowspan="1" colspan="1">4.29</td><td valign="top" align="left" rowspan="1" colspan="1">25.03</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">DenseNet201</td><td valign="top" align="left" rowspan="1" colspan="1">90.56</td><td valign="top" align="left" rowspan="1" colspan="1">90.56</td><td valign="top" align="left" rowspan="1" colspan="1">90.67</td><td valign="top" align="left" rowspan="1" colspan="1">90.60</td><td valign="top" align="left" rowspan="1" colspan="1">89.21</td><td valign="top" align="left" rowspan="1" colspan="1">89.24</td><td valign="top" align="left" rowspan="1" colspan="1">4.39</td><td valign="top" align="left" rowspan="1" colspan="1">20.01</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">GoogLeNet</td><td valign="top" align="left" rowspan="1" colspan="1">92.09</td><td valign="top" align="left" rowspan="1" colspan="1">92.09</td><td valign="top" align="left" rowspan="1" colspan="1">92.20</td><td valign="top" align="left" rowspan="1" colspan="1">92.11</td><td valign="top" align="left" rowspan="1" colspan="1">90.96</td><td valign="top" align="left" rowspan="1" colspan="1">90.98</td><td valign="top" align="left" rowspan="1" colspan="1">1.60</td><td valign="top" align="left" rowspan="1" colspan="1">7.01</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">InceptionRes-NetV2</td><td valign="top" align="left" rowspan="1" colspan="1">91.58</td><td valign="middle" align="left" rowspan="1" colspan="1">91.59</td><td valign="middle" align="left" rowspan="1" colspan="1">92.14</td><td valign="middle" align="left" rowspan="1" colspan="1">91.71</td><td valign="middle" align="left" rowspan="1" colspan="1">90.38</td><td valign="middle" align="left" rowspan="1" colspan="1">90.39</td><td valign="middle" align="left" rowspan="1" colspan="1">6.50</td><td valign="middle" align="left" rowspan="1" colspan="1">55.84</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">EfficientNet-B0</td><td valign="top" align="left" rowspan="1" colspan="1">91.33</td><td valign="middle" align="left" rowspan="1" colspan="1">91.35</td><td valign="middle" align="left" rowspan="1" colspan="1">91.54</td><td valign="middle" align="left" rowspan="1" colspan="1">91.34</td><td valign="middle" align="left" rowspan="1" colspan="1">90.09</td><td valign="middle" align="left" rowspan="1" colspan="1">90.13</td><td valign="middle" align="left" rowspan="1" colspan="1">0.42</td><td valign="middle" align="left" rowspan="1" colspan="1">8.43</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">ShuffleNetV2</td><td valign="top" align="left" rowspan="1" colspan="1">88.01</td><td valign="top" align="left" rowspan="1" colspan="1">88.00</td><td valign="top" align="left" rowspan="1" colspan="1">88.98</td><td valign="top" align="left" rowspan="1" colspan="1">88.21</td><td valign="top" align="left" rowspan="1" colspan="1">86.30</td><td valign="top" align="left" rowspan="1" colspan="1">86.48</td><td valign="top" align="left" rowspan="1" colspan="1">0.04</td><td valign="top" align="left" rowspan="1" colspan="1">1.31</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">MobileNetV2</td><td valign="top" align="left" rowspan="1" colspan="1">89.80</td><td valign="top" align="left" rowspan="1" colspan="1">89.79</td><td valign="top" align="left" rowspan="1" colspan="1">90.52</td><td valign="top" align="left" rowspan="1" colspan="1">89.93</td><td valign="top" align="left" rowspan="1" colspan="1">88.34</td><td valign="top" align="left" rowspan="1" colspan="1">88.41</td><td valign="top" align="left" rowspan="1" colspan="1">0.33</td><td valign="top" align="left" rowspan="1" colspan="1">3.51</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">GhostNet</td><td valign="top" align="left" rowspan="1" colspan="1">90.31</td><td valign="top" align="left" rowspan="1" colspan="1">90.35</td><td valign="top" align="left" rowspan="1" colspan="1">91.09</td><td valign="top" align="left" rowspan="1" colspan="1">90.45</td><td valign="top" align="left" rowspan="1" colspan="1">88.92</td><td valign="top" align="left" rowspan="1" colspan="1">88.98</td><td valign="top" align="left" rowspan="1" colspan="1">0.16</td><td valign="top" align="left" rowspan="1" colspan="1">5.18</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">LCAMNet</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>92.60</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>92.64</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>92.85</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>92.64</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>91.55</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>91.58</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>0.03</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>1.30</bold>
</td></tr></tbody></table><table-wrap-foot><fn id="fn3"><p>Bold values indicate the best results in each column. For Accuracy, Recall, Precision, F1, Kappa, and Matthews Correlation Coefficient (MCC), higher values indicate better performance. Conversely, for GFLOPs and Parameters, lower values indicate a more lightweight and efficient model.</p></fn></table-wrap-foot></table-wrap><p>GoogLeNet (<xref rid="B31" ref-type="bibr">Szegedy et al., 2015</xref>) designs the Inception module, which applies multiple convolutional kernels of different sizes in parallel to capture features at various scales. This improves both feature representation and computational performance. Although InceptionResNetV2 (<xref rid="B30" ref-type="bibr">Szegedy et al., 2016</xref>) incorporates optimization techniques such as multi-branch convolutions, residual connections, and depthwise separable convolutions to improve performance, the network’s complexity and large number of parameters limit its advantages in some application scenarios.</p><p>EfficientNetB0 (<xref rid="B32" ref-type="bibr">Tan and Le, 2020</xref>), ShuffleNetV2 (<xref rid="B43" ref-type="bibr">Zhang et al., 2018</xref>), MobileNetV2 (<xref rid="B10" ref-type="bibr">Howard et al., 2017</xref>), and GhostNet (<xref rid="B8" ref-type="bibr">Han et al., 2020</xref>) are lightweight models that have been widely used in recent years. These models optimize the network structure to reduce computational overhead while maintaining high accuracy. They incorporate various lightweight techniques, such as depthwise separable convolutions, channel reparameterization, and residual connections, achieving a good balance between computational efficiency and model accuracy. However, these models may face performance bottlenecks when handling more complex tasks due to limited model capacity and expressive power.</p><p>LCAMNet, based on ShuffleNetV2 model, optimizes the feature extraction and downsampling modules. By using depthwise separable convolutions and structural reparameterization, it reduces network parameters while still extracting sufficient features. Compared to classic convolutional neural networks, LCAMNet has fewer parameters than ShuffleNetV2 and achieves the highest accuracy. This demonstrates that LCAMNet is a high-performance, lightweight model for apple leaf disease classification. The performance comparison of LCAMNet and several classical CNN models on SCEBD is shown in <xref rid="T6" ref-type="table">
<bold>Table 6</bold>
</xref>. The three-dimensional visual analysis of accuracy, GFLOPs, and parameter count on SCEBD is presented in <xref rid="f10" ref-type="fig">
<bold>Figure 10</bold>
</xref>.</p><fig id="f10" position="float"><?disp-level 3?><label>Figure 10</label><caption><p>3D Visualization of accuracy, GFLOPs, and parameter count on the SCEBD.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g010.jpg"><?cloudpmc-path blobs/ce99/12378099/0e3e40f6e848/fpls-16-1626569-g010.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1726?><?original-width 1918?><?scaled-height 690?><?scaled-width 767?><alt-text>A 3D scatter plot comparing various neural networks, including LCAMNet, GoogLeNet, EfficientNetB0, and others based on GFLOPs, parameters, and accuracy. The vertical color bar represents accuracy ranging from eighty-six to ninety-two percent, with different colors indicating accuracy levels.</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g010.gif"><?cloudpmc-path blobs/ce99/12378099/e97c914cb661/fpls-16-1626569-g010.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec><sec id="s4_5" disp-level="2"><label>4.5.</label><title>Performance comparison with similar crop disease image classification models on SCEBD</title><p>This study compares the performance of LCAMNet with similar crop disease image classification models on the SCEBD. Re-GoogLeNet (<xref rid="B39" ref-type="bibr">Yang et al., 2023</xref>) is a network designed for rice image classification in natural environmental backgrounds, based on a series of improvements to GoogLeNet. First, the 7×7 convolution kernel in the first layer of GoogLeNet is replaced with three consecutive 3×3 convolutions. Then, the Inception module is enhanced by adding the ECA attention mechanism and optimized residual connections to strengthen information flow. Lastly, LeakyReLU replaces the ReLU activation to better capture irregular features in diseased leaves. ALS-Net (<xref rid="B18" ref-type="bibr">Liu et al., 2022a</xref>) adds an Inception structure after 3×3 convolution in the ShuffleNetV2 model for multi-scale feature extraction. Additionally, the 3×3 convolutions in the ShuffleNet block are replaced with 5×5 depthwise convolutions to obtain the ALS module. The ELU activation function also replaces ReLU, addressing gradient vanishing and neuron death issues. To further improve classification performance, ALS-Net uses DenseNet161 as a teacher network to guide training and enhance the model’s classification ability. LBMRNet (<xref rid="B16" ref-type="bibr">Li et al., 2023</xref>) is a lightweight algorithm for recognizing tomato leaf diseases, aimed at tackling the problem of significant variation within the same class and minimal variation between different classes in tomato leaf disease images. LBMRNet consists of alternating complementary group dilation residual (CGDR) modules and visual enhancement modules. The CGDR module uses a multi-branch design to capture various features of tomato leaf diseases from different receptive fields. It incorporates multiple residual connections to facilitate better information flow across the network layers. The visual enhancement module combines average pooling, max pooling, and 1×1 convolutions as downsampling strategies, effectively fusing the visual enhancement effects and preventing information loss during the downsampling process, thus improving classification accuracy.</p><p>From <xref rid="T7" ref-type="table">
<bold>Table 7</bold>
</xref>, it could be seen that the performance of ALS-Net and LBMRNet are inferior to LCAMNet (for LBMRNet, only the network structure is restored). This is likely because, although ALS-Net improves ShuffleNetV2, the multi-scale feature extraction is only performed at the initial stages of the network and fails to effectively integrate these features at each stage, limiting its feature extraction capability. Although LBMRNet improves the feature extraction and downsampling layers, its multi-residual structure leads to the continuous retention of irrelevant information, affecting the recognition performance. In comparison, Re-GoogLeNet optimizes GoogLeNet by introducing the attention mechanism to preserve important features and uses residual connections to retain original features. While its performance is better than that of LCAMNet, its higher FLOPs and Parameters may limit its applicability in resource-constrained scenarios. The confusion matrices of the models are shown in <xref rid="f11" ref-type="fig">
<bold>Figure 11</bold>
</xref>.</p><table-wrap id="T7" position="float"><?disp-level 3?><label>Table 7</label><caption><p>Performance comparison of LCAMNet and similar crop disease image classification models on the SCEBD.</p></caption><table frame="hsides" rules="groups"><thead><tr><th valign="top" align="left" rowspan="1" colspan="1">Models</th><th valign="top" align="left" rowspan="1" colspan="1">Accuracy (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Recall (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Precision (%)</th><th valign="top" align="left" rowspan="1" colspan="1">F1 (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Kappa (%)</th><th valign="top" align="left" rowspan="1" colspan="1">MCC (%)</th><th valign="top" align="left" rowspan="1" colspan="1">GFLOPs (G)</th><th valign="top" align="left" rowspan="1" colspan="1">Parameters (M)</th></tr></thead><tbody><tr><td valign="top" align="left" rowspan="1" colspan="1">Re-GoogLeNet</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>93.37</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>93.36</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>93.79</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>93.46</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>92.42</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>92.64</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">2.80</td><td valign="top" align="left" rowspan="1" colspan="1">9.11</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">ALS-Net</td><td valign="top" align="left" rowspan="1" colspan="1">91.07</td><td valign="top" align="left" rowspan="1" colspan="1">91.09</td><td valign="top" align="left" rowspan="1" colspan="1">91.28</td><td valign="top" align="left" rowspan="1" colspan="1">91.13</td><td valign="top" align="left" rowspan="1" colspan="1">89.80</td><td valign="top" align="left" rowspan="1" colspan="1">89.81</td><td valign="top" align="left" rowspan="1" colspan="1">0.73</td><td valign="top" align="left" rowspan="1" colspan="1">1.18</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">LBMRNet</td><td valign="top" align="left" rowspan="1" colspan="1">90.05</td><td valign="top" align="left" rowspan="1" colspan="1">90.06</td><td valign="top" align="left" rowspan="1" colspan="1">91.54</td><td valign="top" align="left" rowspan="1" colspan="1">90.30</td><td valign="top" align="left" rowspan="1" colspan="1">88.63</td><td valign="top" align="left" rowspan="1" colspan="1">88.81</td><td valign="top" align="left" rowspan="1" colspan="1">0.17</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>0.91</bold>
</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">LCAMNet</td><td valign="top" align="left" rowspan="1" colspan="1">92.60</td><td valign="top" align="left" rowspan="1" colspan="1">92.64</td><td valign="top" align="left" rowspan="1" colspan="1">92.85</td><td valign="top" align="left" rowspan="1" colspan="1">92.64</td><td valign="top" align="left" rowspan="1" colspan="1">91.55</td><td valign="top" align="left" rowspan="1" colspan="1">91.58</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>0.03</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">1.30</td></tr></tbody></table><table-wrap-foot><fn id="fn4"><p>Bold values indicate the best results in each column. For Accuracy, Recall, Precision, F1, Kappa, and Matthews Correlation Coefficient (MCC), higher values indicate better performance. Conversely, for GFLOPs and Parameters, lower values indicate a more lightweight and efficient model.</p></fn></table-wrap-foot></table-wrap><fig id="f11" position="float"><?disp-level 3?><label>Figure 11</label><caption><p>Confusion matrices of LCAMNet and similar crop disease image classification models on SCEBD: <bold>(A)</bold> Re-GoogLeNet; <bold>(B)</bold> ALS-Net; <bold>(C)</bold> LBMRNet; <bold>(D)</bold> LCAMNet.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g011.jpg"><?cloudpmc-path blobs/ce99/12378099/e6d629544ff5/fpls-16-1626569-g011.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1980?><?original-width 2105?><?scaled-height 659?><?scaled-width 701?><alt-text>Four confusion matrices labeled A, B, C, and D, compare true versus predicted labels for plant disease classification: Alternaria leaf spot, Brown spot, Grey spot, Healthy, Powdery mildew, Mosaic, Rust, and Scab. Each matrix highlights the number of correct and incorrect predictions, with darker shades indicating higher values.</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g011.gif"><?cloudpmc-path blobs/ce99/12378099/5c9f807e25d1/fpls-16-1626569-g011.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec><sec id="s4_6" disp-level="2"><label>4.6.</label><title>Performance comparison on the public dataset FGVC8</title><p>This study compares the performance of LCAMNet with several classical CNN models on the FGVC8 dataset to verify its effectiveness and superiority. Detailed experimental settings are provided in Section 4.1. <xref rid="T8" ref-type="table">
<bold>Table 8</bold>
</xref> presents a performance comparison between LCAMNet and several classical CNN models on FGVC8. The results show that LCAMNet achieves a classification accuracy comparable to GoogLeNet and InceptionResNetV2, while significantly reducing FLOPs and Parameters compared to these two models. This indicates that LCAMNet maintains a high recognition accuracy while lowering computational resource consumption. In addition, LCAMNet achieves better classification accuracy than ShuffleNetV2 on the public dataset, while maintaining slightly lower FLOPs and parameters. This demonstrates that LCAMNet effectively integrates techniques like multi-branch feature extraction modules, dual-branch downsampling modules, and attention mechanisms into the ShuffleNetV2 architecture, making it suitable for different datasets. Three-dimensional visual analysis of accuracy, GFLOPs, and parameters on FGVC8 is presented in <xref rid="f12" ref-type="fig">
<bold>Figure 12</bold>
</xref>.</p><table-wrap id="T8" position="float"><?disp-level 3?><label>Table 8</label><caption><p>Performance comparison of LCAMNet and several classical CNN models on the FGVC8 dataset.</p></caption><table frame="hsides" rules="groups"><thead><tr><th valign="top" align="left" rowspan="1" colspan="1">Models</th><th valign="top" align="left" rowspan="1" colspan="1">Accuracy (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Recall (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Precision (%)</th><th valign="top" align="left" rowspan="1" colspan="1">F1 (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Kappa (%)</th><th valign="top" align="left" rowspan="1" colspan="1">MCC (%)</th><th valign="top" align="left" rowspan="1" colspan="1">GFLOPs (G)</th><th valign="top" align="left" rowspan="1" colspan="1">Parameters (M)</th></tr></thead><tbody><tr><td valign="top" align="left" rowspan="1" colspan="1">GoogLeNet</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>95.31</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>95.33</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">95.56</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>95.39</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">94.20</td><td valign="top" align="left" rowspan="1" colspan="1">95.39</td><td valign="top" align="left" rowspan="1" colspan="1">1.60</td><td valign="top" align="left" rowspan="1" colspan="1">7.01</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">InceptionRes-NetV2</td><td valign="top" align="left" rowspan="1" colspan="1">94.92</td><td valign="middle" align="left" rowspan="1" colspan="1">94.93</td><td valign="middle" align="left" rowspan="1" colspan="1">95.40</td><td valign="middle" align="left" rowspan="1" colspan="1">95.06</td><td valign="middle" align="left" rowspan="1" colspan="1">93.71</td><td valign="middle" align="left" rowspan="1" colspan="1">95.06</td><td valign="middle" align="left" rowspan="1" colspan="1">6.50</td><td valign="middle" align="left" rowspan="1" colspan="1">55.84</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">EfficientNet-B0</td><td valign="top" align="left" rowspan="1" colspan="1">91.80</td><td valign="middle" align="left" rowspan="1" colspan="1">91.88</td><td valign="middle" align="left" rowspan="1" colspan="1">92.13</td><td valign="middle" align="left" rowspan="1" colspan="1">91.92</td><td valign="middle" align="left" rowspan="1" colspan="1">89.76</td><td valign="middle" align="left" rowspan="1" colspan="1">91.92</td><td valign="middle" align="left" rowspan="1" colspan="1">0.42</td><td valign="middle" align="left" rowspan="1" colspan="1">8.43</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">ShuffleNetV2</td><td valign="top" align="left" rowspan="1" colspan="1">91.80</td><td valign="top" align="left" rowspan="1" colspan="1">91.96</td><td valign="top" align="left" rowspan="1" colspan="1">91.89</td><td valign="top" align="left" rowspan="1" colspan="1">91.86</td><td valign="top" align="left" rowspan="1" colspan="1">89.91</td><td valign="top" align="left" rowspan="1" colspan="1">91.86</td><td valign="top" align="left" rowspan="1" colspan="1">0.04</td><td valign="top" align="left" rowspan="1" colspan="1">1.31</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">LCAMNet</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>95.31</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">95.29</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>95.81</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>95.39</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>94.22</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>95.39</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>0.03</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>1.30</bold>
</td></tr></tbody></table><table-wrap-foot><fn id="fn5"><p>Bold values indicate the best results in each column. For Accuracy, Recall, Precision, F1, Kappa, and Matthews Correlation Coefficient (MCC), higher values indicate better performance. Conversely, for GFLOPs and Parameters, lower values indicate a more lightweight and efficient model.</p></fn></table-wrap-foot></table-wrap><fig id="f12" position="float"><?disp-level 3?><label>Figure 12</label><caption><p>3D visualization of accuracy, FLOPs, and parameters on the FGVC8 dataset.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g012.jpg"><?cloudpmc-path blobs/ce99/12378099/75dd68740a7f/fpls-16-1626569-g012.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 1726?><?original-width 1933?><?scaled-height 690?><?scaled-width 773?><alt-text>3D scatter plot depicting various neural networks comparing GFLOPS, parameters, and accuracy. Networks include LCAMNet, GoogLeNet, InceptionResNetV2, EfficientNetB0, and ShuffleNetV2. Color gradient indicates accuracy, ranging from purple (low) to yellow (high).</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g012.gif"><?cloudpmc-path blobs/ce99/12378099/eef74fc7c604/fpls-16-1626569-g012.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig><p>Furthermore, LCAMNet is compared with similar crop disease image classification models. Experimental results show that LCAMNet outperforms ALS-Net and LBMRNet in overall performance. Although its performance is slightly lower than Re-GoogLeNet, LCAMNet has significantly fewer FLOPs and Parameters, demonstrating a notable optimization in computational resource consumption. In conclusion, the experimental results fully validate the effectiveness of LCAMNet on the public dataset. LCAMNet achieves competitive recognition accuracy while significantly reducing computational complexity, showcasing strong practical application value and potential for broader deployment. Detailed results can be found in <xref rid="T9" ref-type="table">
<bold>Table 9</bold>
</xref>, and the models’ confusion matrices are shown in <xref rid="f13" ref-type="fig">
<bold>Figure 13</bold>
</xref>.</p><table-wrap id="T9" position="float"><?disp-level 3?><label>Table 9</label><caption><p>Performance comparison of similar crop disease image classification models on the FGVC8 dataset.</p></caption><table frame="hsides" rules="groups"><thead><tr><th valign="top" align="left" rowspan="1" colspan="1">Models</th><th valign="top" align="left" rowspan="1" colspan="1">Accuracy (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Recall (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Precision (%)</th><th valign="top" align="left" rowspan="1" colspan="1">F1 (%)</th><th valign="top" align="left" rowspan="1" colspan="1">Kappa (%)</th><th valign="top" align="left" rowspan="1" colspan="1">MCC (%)</th><th valign="top" align="left" rowspan="1" colspan="1">GFLOPs (G)</th><th valign="top" align="left" rowspan="1" colspan="1">Parameters (M)</th></tr></thead><tbody><tr><td valign="top" align="left" rowspan="1" colspan="1">Re-GoogLeNet</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>96.48</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>96.50</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>96.58</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>96.49</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>95.60</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>96.63</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">2.80</td><td valign="top" align="left" rowspan="1" colspan="1">9.11</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">ALS-Net</td><td valign="top" align="left" rowspan="1" colspan="1">92.97</td><td valign="top" align="left" rowspan="1" colspan="1">93.03</td><td valign="top" align="left" rowspan="1" colspan="1">93.55</td><td valign="top" align="left" rowspan="1" colspan="1">93.08</td><td valign="top" align="left" rowspan="1" colspan="1">91.21</td><td valign="top" align="left" rowspan="1" colspan="1">91.32</td><td valign="top" align="left" rowspan="1" colspan="1">0.73</td><td valign="top" align="left" rowspan="1" colspan="1">1.18</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">LBMRNet</td><td valign="top" align="left" rowspan="1" colspan="1">91.02</td><td valign="top" align="left" rowspan="1" colspan="1">91.13</td><td valign="top" align="left" rowspan="1" colspan="1">91.75</td><td valign="top" align="left" rowspan="1" colspan="1">91.30</td><td valign="top" align="left" rowspan="1" colspan="1">88.77</td><td valign="top" align="left" rowspan="1" colspan="1">88.84</td><td valign="top" align="left" rowspan="1" colspan="1">0.17</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>0.91</bold>
</td></tr><tr><td valign="top" align="left" rowspan="1" colspan="1">LCAMNet</td><td valign="top" align="left" rowspan="1" colspan="1">95.31</td><td valign="top" align="left" rowspan="1" colspan="1">95.29</td><td valign="top" align="left" rowspan="1" colspan="1">95.81</td><td valign="top" align="left" rowspan="1" colspan="1">95.39</td><td valign="top" align="left" rowspan="1" colspan="1">94.14</td><td valign="top" align="left" rowspan="1" colspan="1">94. 22</td><td valign="top" align="left" rowspan="1" colspan="1">
<bold>0.03</bold>
</td><td valign="top" align="left" rowspan="1" colspan="1">1.30</td></tr></tbody></table><table-wrap-foot><fn id="fn6"><p>Bold values indicate the best results in each column. For Accuracy, Recall, Precision, F1, Kappa, and Matthews Correlation Coefficient (MCC), higher values indicate better performance. Conversely, for GFLOPs and Parameters, lower values indicate a more lightweight and efficient model.</p></fn></table-wrap-foot></table-wrap><fig id="f13" position="float"><?disp-level 3?><label>Figure 13</label><caption><p>Confusion matrices of similar crop disease image classification models on the FGVC8 dataset: <bold>(A)</bold> Re-GoogLeNet; <bold>(B)</bold> ALS-Net; <bold>(C)</bold> LBMRNet; <bold>(D)</bold> LCAMNet.</p></caption><alternatives><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="image" xlink:href="fpls-16-1626569-g013.jpg"><?cloudpmc-path blobs/ce99/12378099/cd9aa87ecb2b/fpls-16-1626569-g013.jpg?><?cloudpmc-bucket cdn?><?image-server-status LOAD_COMPLETED?><?original-height 2084?><?original-width 2106?><?scaled-height 695?><?scaled-width 702?><alt-text>Four confusion matrices labeled A, B, C, and D, each comparing true labels against predicted labels for five classifications: Alternaria leaf spot, Healthy, Powdery mildew, Rust, and Scab. Each matrix shows a diagonal pattern indicating accurate predictions, with minor off-diagonal values representing errors. The color scale ranges from light to dark blue, indicating fewer to more instances, respectively.</alt-text></graphic><graphic xmlns:xlink="http://www.w3.org/1999/xlink" content-type="thumb" xlink:href="fpls-16-1626569-g013.gif"><?cloudpmc-path blobs/ce99/12378099/53cb2921169a/fpls-16-1626569-g013.gif?><?cloudpmc-bucket cdn?></graphic></alternatives></fig></sec><sec id="s4_7" disp-level="2"><label>4.7.</label><title>Model limitation analysis</title><p>Although LCAMNet achieves outstanding performance on both the SCEBD and FGVC8 datasets, several limitations remain based on the experimental results:</p><p>As shown in <xref rid="T6" ref-type="table">
<bold>Table 6</bold>
</xref>, LCAMNet demonstrates a favorable balance between accuracy and model complexity compared to various CNN models on the SCEBD dataset. However, in the comparative experiments with other crop disease identification models <xref rid="T7" ref-type="table">
<bold>Table 7</bold>
</xref>, although LCAMNet outperforms ALS-Net and LBMRNet, its accuracy is slightly lower than that of Re-GoogLeNet. This indicates that while LCAMNet holds significant advantages in terms of parameter count and computational cost, there is still room for improvement in feature fusion and deep representation capabilities. Moreover, the staged multi-scale attention mechanism present in Re-GoogLeNet significantly enhances performance, suggesting that LCAMNet could benefit from incorporating richer stage-wise fusion strategies for further optimization.</p><p>Furthermore, LCAMNet maintains competitive performance on the FGVC8 dataset (as shown in <xref rid="T8" ref-type="table">
<bold>Tables 8</bold>
</xref>, <xref rid="T9" ref-type="table">
<bold>9</bold>
</xref>), indicating its generalization capability. However, some classification confusion still occurs among certain categories in FGVC8, as evidenced by the confusion matrix in <xref rid="f13" ref-type="fig">
<bold>Figure 13</bold>
</xref>. This reveals a risk of misclassification, especially when dealing with disease symptoms that have similar morphological patterns and subtle color differences. These observations suggest that LCAMNet could further enhance its fine-grained modeling capacity for complex lesion patterns.</p></sec></sec><sec id="s5" disp-level="1"><label>5.</label><title>Conclusion</title><p>The lightweight converged attention multi-branch network, LCAMNet, proposed in this study achieves high accuracy in apple leaf disease classification while significantly reducing model complexity and computational cost. This demonstrates its strong potential for deployment in resource-constrained and complex natural environments. By integrating structural re-parameterization, dual-branch downsampling, and multi-scale attention mechanisms, LCAMNet effectively enhances lesion feature modeling and the expression of feature diversity.</p><p>Despite its outstanding performance in apple leaf disease recognition, there remains room for improvement in LCAMNet’s adaptability and generalization capability. Future research will focus on two main directions: (1) systematically evaluating the model’s robustness and stability under complex natural conditions, such as varying illumination, camera angles, and degrees of leaf occlusion; and (2) extending the application of LCAMNet to the classification of diseases in other crops such as wheat, maize, and rice, in order to explore its cross-crop generalization performance and domain adaptability. These efforts will further promote the practical deployment of LCAMNet in multi-scenario and multi-crop disease recognition, laying a solid foundation for building an intelligent plant disease identification system in precision agriculture.</p></sec><sec id="funding-statement1" xml:lang="en" disp-level="1"><title>Funding Statement</title><p>The author(s) declare that financial support was received for the research and/or publication of this article. This work is supported by the National Natural Science Foundation Project (No. 62041211), the Inner Mongolia Major science and technology project (No. 2021SZD0004), the Inner Mongolia Autonomous Region science and technology plan project (No. 2022YFHH0070), the Basic research expenses of universities directly under the Inner Mongolia Autonomous Region (No. BR22-14-05), the Inner Mongolia Natural Science Foundation Project (No. 2024MS06002), the Inner Mongolia Autonomous Region universities innovative research team project (No. NMGIRT2313) and the Inner Mongolia Natural Science Foundation Project (No. 2025ZD012).</p></sec><sec id="s6" disp-level="1"><title>Data availability statement</title><p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8" ext-link-type="uri">https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8</ext-link>; <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/JasonYangCode/AppleLeaf9" ext-link-type="uri">https://github.com/JasonYangCode/AppleLeaf9</ext-link>; <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.scidb.cn/en/detail?dataSetId=0e1f57004db842f99668d82183afd578" ext-link-type="uri">https://www.scidb.cn/en/detail?dataSetId=0e1f57004db842f99668d82183afd578</ext-link>.</p></sec><sec id="s7" disp-level="1"><title>Author contributions</title><p>YJ: Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft. HL: Funding acquisition, Resources, Writing – review &amp; editing. XF: Project administration, Supervision, Writing – review &amp; editing. BW: Project administration, Writing – review &amp; editing. KH: Conceptualization, Writing – review &amp; editing. SZ: Conceptualization, Writing – review &amp; editing. DH: Conceptualization, Writing – review &amp; editing.</p></sec><sec id="s9" disp-level="1"><title>Conflict of interest</title><p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec><sec id="s10" disp-level="1"><title>Generative AI statement</title><p>The author(s) declare that no Generative AI was used in the creation of this manuscript.</p></sec><sec id="s11" disp-level="1"><title>Publisher’s note</title><p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p></sec><sec id="ref-list1" sec-type="ref-list" disp-level="1"><title>References</title><sec id="ref-list1_sec2" disp-level="2"><ref-list><ref id="B1"><mixed-citation><named-content content-type="citation-string">
Ali M. U., Khalid M., Farrash M., Lahza H. F. M., Zafar A., Kim S.-H. (2024). Appleleafnet: a lightweight and efficient deep learning framework for diagnosing apple leaf diseases. Front. Plant Sci.
15. doi:  10.3389/fpls.2024.1502314, PMID: 
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3389/fpls.2024.1502314"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC11631600"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="39665107"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Front. Plant Sci.&amp;title=Appleleafnet: a lightweight and efficient deep learning framework for diagnosing apple leaf diseases&amp;author=M. U. Ali&amp;author=M. Khalid&amp;author=M. Farrash&amp;author=H. F. M. Lahza&amp;author=A. Zafar&amp;volume=15&amp;publication_year=2024&amp;pmid=39665107&amp;doi=10.3389/fpls.2024.1502314&amp;"/></mixed-citation></ref><ref id="B2"><mixed-citation><named-content content-type="citation-string">
Association, C. A. I  (2023).China’s total apple production reached 47.57 million tons, including 13.03 million tons in Shaanxi. Available online at: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://www.guoye.sn.cn/hydt/40639.jhtml" ext-link-type="uri">http://www.guoye.sn.cn/hydt/40639.jhtml</ext-link> (Accessed March 26 2025).</named-content></mixed-citation></ref><ref id="B3"><mixed-citation><named-content content-type="citation-string">
Bacanin N., Zivkovic M., Sarac M., Petrovic A., Strumberger I., Antonijevic M., et al. (2022). “A novel multiswarm firefly algorithm: An application for plant classification,” in In International Conference on Intelligent and Fuzzy Systems (Springer; ), 1007–1016.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="title=In International Conference on Intelligent and Fuzzy Systems&amp;author=N. Bacanin&amp;author=M. Zivkovic&amp;author=M. Sarac&amp;author=A. Petrovic&amp;author=I. Strumberger&amp;publication_year=2022&amp;"/></mixed-citation></ref><ref id="B4"><mixed-citation><named-content content-type="citation-string">
Bukumira M., Antonijevic M., Jovanovic D., Zivkovic M., Mladenovic D., Kunjadic G. (2022). Carrot grading system using computer vision feature parameters and a cascaded graph convolutional neural network. J. Electronic Imaging
31, 061815–061815. doi:  10.1117/1.JEI.31.6.061815
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1117/1.JEI.31.6.061815"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Electronic Imaging&amp;title=Carrot grading system using computer vision feature parameters and a cascaded graph convolutional neural network&amp;author=M. Bukumira&amp;author=M. Antonijevic&amp;author=D. Jovanovic&amp;author=M. Zivkovic&amp;author=D. Mladenovic&amp;volume=31&amp;publication_year=2022&amp;pages=061815-061815&amp;doi=10.1117/1.JEI.31.6.061815&amp;"/></mixed-citation></ref><ref id="B5"><mixed-citation><named-content content-type="citation-string">
Cui J., Zhang Y., Chen H., Zhang Y., Cai H., Jiang Y., et al. (2025). Cswin-mbconv: A dual-network fusing cnn and transformer for weed recognition. Eur. J. Agron.
164, 127528. doi:  10.1016/j.eja.2025.127528
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.eja.2025.127528"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Eur. J. Agron.&amp;title=Cswin-mbconv: A dual-network fusing cnn and transformer for weed recognition&amp;author=J. Cui&amp;author=Y. Zhang&amp;author=H. Chen&amp;author=Y. Zhang&amp;author=H. Cai&amp;volume=164&amp;publication_year=2025&amp;pages=127528&amp;doi=10.1016/j.eja.2025.127528&amp;"/></mixed-citation></ref><ref id="B6"><mixed-citation><named-content content-type="citation-string">
Dong Q., Gu R., Chen S., Zhu J. (2024). Apple leaf disease diagnosis based on knowledge distillation and attention mechanism. IEEE Access.
12, 65154–65165. doi:  10.1109/ACCESS.2024.3397329
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1109/ACCESS.2024.3397329"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=IEEE Access.&amp;title=Apple leaf disease diagnosis based on knowledge distillation and attention mechanism&amp;author=Q. Dong&amp;author=R. Gu&amp;author=S. Chen&amp;author=J. Zhu&amp;volume=12&amp;publication_year=2024&amp;pages=65154-65165&amp;doi=10.1109/ACCESS.2024.3397329&amp;"/></mixed-citation></ref><ref id="B7"><mixed-citation><named-content content-type="citation-string">
Feng J., Chao X. (2022). Apple tree leaf disease segmentation dataset. doi:  10.11922/sciencedb.01627
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.11922/sciencedb.01627"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Feng J., Chao X. (2022). Apple tree leaf disease segmentation dataset. doi:  10.11922/sciencedb.01627"/></mixed-citation></ref><ref id="B8"><mixed-citation><named-content content-type="citation-string">
Han K., Wang Y., Tian Q., Guo J., Xu C., Xu C. (2020). “Ghostnet: More features from cheap operations,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. 1580–1589, IEEE.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Han K., Wang Y., Tian Q., Guo J., Xu C., Xu C. (2020). “Ghostnet: More features from cheap operations,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. 1580–1589, IEEE."/></mixed-citation></ref><ref id="B9"><mixed-citation><named-content content-type="citation-string">
He K., Zhang X., Ren S., Sun J. (2016). “Deep residual learning for image recognition,” in 2016 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR). 770–778, IEEE.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="He K., Zhang X., Ren S., Sun J. (2016). “Deep residual learning for image recognition,” in 2016 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR). 770–778, IEEE."/></mixed-citation></ref><ref id="B10"><mixed-citation><named-content content-type="citation-string">
Howard A. G., Zhu M., Chen B., Kalenichenko D., Wang W., Weyand T., et al. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=arXiv preprint arXiv:1704.04861&amp;title=Mobilenets: Efficient convolutional neural networks for mobile vision applications&amp;author=A. G. Howard&amp;author=M. Zhu&amp;author=B. Chen&amp;author=D. Kalenichenko&amp;author=W. Wang&amp;publication_year=2017&amp;"/></mixed-citation></ref><ref id="B11"><mixed-citation><named-content content-type="citation-string">
Huang G., Liu Z., van der Maaten L., Weinberger K. Q. (2017). “Densely connected convolutional networks,” in 2017 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR). 2261–2269, IEEE.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Huang G., Liu Z., van der Maaten L., Weinberger K. Q. (2017). “Densely connected convolutional networks,” in 2017 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR). 2261–2269, IEEE."/></mixed-citation></ref><ref id="B12"><mixed-citation><named-content content-type="citation-string">
Huang X., Xu D., Chen Y., Zhang Q., Feng P., Ma Y., et al. (2025). Econv-vit: A strongly generalized apple leaf disease classification model based on the fusion of convnext and transformer. Inf. Process. Agric. doi:  10.1016/j.inpa.2025.03.001
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.inpa.2025.03.001"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Inf. Process. Agric&amp;title=Econv-vit: A strongly generalized apple leaf disease classification model based on the fusion of convnext and transformer&amp;author=X. Huang&amp;author=D. Xu&amp;author=Y. Chen&amp;author=Q. Zhang&amp;author=P. Feng&amp;publication_year=2025&amp;doi=10.1016/j.inpa.2025.03.001&amp;"/></mixed-citation></ref><ref id="B13"><mixed-citation><named-content content-type="citation-string">
Jiang H., Yang X., Ding R., Wang D., Mao W., Qiao Y. (2023). Identification of apple leaf diseases based on improved resnet18. Trans. Chin. Soc Agric. Eng.
54, 295—303. doi:  10.6041/j.issn.1000-1298.2023.04.030
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.6041/j.issn.1000-1298.2023.04.030"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Trans. Chin. Soc Agric. Eng.&amp;title=Identification of apple leaf diseases based on improved resnet18&amp;author=H. Jiang&amp;author=X. Yang&amp;author=R. Ding&amp;author=D. Wang&amp;author=W. Mao&amp;volume=54&amp;publication_year=2023&amp;pages=295—303&amp;doi=10.6041/j.issn.1000-1298.2023.04.030&amp;"/></mixed-citation></ref><ref id="B14"><mixed-citation><named-content content-type="citation-string">
Li D., Zhang C., Li J., Li M., Huang M., Tang Y. (2024). Mccm: multi-scale feature extraction network for disease classification and recognition of chili leaves. Front. Plant Sci.
15. doi:  10.3389/fpls.2024.1367738, PMID: 
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3389/fpls.2024.1367738"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC11165206"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="38863551"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Front. Plant Sci.&amp;title=Mccm: multi-scale feature extraction network for disease classification and recognition of chili leaves&amp;author=D. Li&amp;author=C. Zhang&amp;author=J. Li&amp;author=M. Li&amp;author=M. Huang&amp;volume=15&amp;publication_year=2024&amp;pmid=38863551&amp;doi=10.3389/fpls.2024.1367738&amp;"/></mixed-citation></ref><ref id="B15"><mixed-citation><named-content content-type="citation-string">
Li H., Ruan C., Zhao J., Huang L., Dong Y., Huang W., et al. (2025). Integrating high-frequency detail information for enhanced corn leaf disease recognition: A model utilizing fusion imagery. Eur. J. Agron.
164, 127489. doi:  10.1016/j.eja.2024.127489
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.eja.2024.127489"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Eur. J. Agron.&amp;title=Integrating high-frequency detail information for enhanced corn leaf disease recognition: A model utilizing fusion imagery&amp;author=H. Li&amp;author=C. Ruan&amp;author=J. Zhao&amp;author=L. Huang&amp;author=Y. Dong&amp;volume=164&amp;publication_year=2025&amp;pages=127489&amp;doi=10.1016/j.eja.2024.127489&amp;"/></mixed-citation></ref><ref id="B16"><mixed-citation><named-content content-type="citation-string">
Li M., Zhou G., Chen A., Li L., Hu Y. (2023). Identification of tomato leaf diseases based on lmbrnet. Eng. Appl. Artif. Intell.
123, 106195. doi:  10.1016/j.engappai.2023.106195
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.engappai.2023.106195"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Eng. Appl. Artif. Intell.&amp;title=Identification of tomato leaf diseases based on lmbrnet&amp;author=M. Li&amp;author=G. Zhou&amp;author=A. Chen&amp;author=L. Li&amp;author=Y. Hu&amp;volume=123&amp;publication_year=2023&amp;pages=106195&amp;doi=10.1016/j.engappai.2023.106195&amp;"/></mixed-citation></ref><ref id="B17"><mixed-citation><named-content content-type="citation-string">
Liang J., Jiang W. (2023). A resnet50-dpa model for tomato leaf disease identification. Front. Plant Sci.
14. doi:  10.3389/fpls.2023.1258658, PMID: 
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3389/fpls.2023.1258658"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC10614023"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="37908831"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Front. Plant Sci.&amp;title=A resnet50-dpa model for tomato leaf disease identification&amp;author=J. Liang&amp;author=W. Jiang&amp;volume=14&amp;publication_year=2023&amp;pmid=37908831&amp;doi=10.3389/fpls.2023.1258658&amp;"/></mixed-citation></ref><ref id="B18"><mixed-citation><named-content content-type="citation-string">
Liu B., Jia R., Zhu X., Yu C., Yao Z., Zhang H., et al. (2022. a). Lightweight identification model for apple leaf diseases and pests based on mobile terminals. Trans. Chin. Soc Agric. Eng.
38, 130–139. doi:  10.11975/j.issn.1002-6819.2022.06.015
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.11975/j.issn.1002-6819.2022.06.015"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Trans. Chin. Soc Agric. Eng.&amp;title=Lightweight identification model for apple leaf diseases and pests based on mobile terminals&amp;author=B. Liu&amp;author=R. Jia&amp;author=X. Zhu&amp;author=C. Yu&amp;author=Z. Yao&amp;volume=38&amp;publication_year=2022&amp;pages=130-139&amp;doi=10.11975/j.issn.1002-6819.2022.06.015&amp;"/></mixed-citation></ref><ref id="B19"><mixed-citation><named-content content-type="citation-string">
Liu M., Liang H., Hou M. (2022. b). Research on cassava disease classification using the multiscale fusion model based on efficientnet and attention mechanism. Front. Plant Sci.
13. doi:  10.3389/fpls.2022.1088531, PMID: 
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3389/fpls.2022.1088531"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC9815107"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="36618625"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Front. Plant Sci.&amp;title=Research on cassava disease classification using the multiscale fusion model based on efficientnet and attention mechanism&amp;author=M. Liu&amp;author=H. Liang&amp;author=M. Hou&amp;volume=13&amp;publication_year=2022&amp;pmid=36618625&amp;doi=10.3389/fpls.2022.1088531&amp;"/></mixed-citation></ref><ref id="B20"><mixed-citation><named-content content-type="citation-string">
Liu Y., Wang Z., Wang R., Chen J., Gao H. (2023). Flooding-based mobilenet to identify cucumber diseases from leaf images in natural scenes. Comput. Electron. Agric.
213, 108166. doi:  10.1016/j.compag.2023.108166
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.compag.2023.108166"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Comput. Electron. Agric.&amp;title=Flooding-based mobilenet to identify cucumber diseases from leaf images in natural scenes&amp;author=Y. Liu&amp;author=Z. Wang&amp;author=R. Wang&amp;author=J. Chen&amp;author=H. Gao&amp;volume=213&amp;publication_year=2023&amp;pages=108166&amp;doi=10.1016/j.compag.2023.108166&amp;"/></mixed-citation></ref><ref id="B21"><mixed-citation><named-content content-type="citation-string">
Misra D., Nalamada T., Arasanipalai A. U., Hou Q. (2021). “Rotate to attend: Convolutional triplet attention module,” in 2021 IEEE Winter Conf. Appl. Comput. Vis. (WACV). 3138–3147, IEEE.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Misra D., Nalamada T., Arasanipalai A. U., Hou Q. (2021). “Rotate to attend: Convolutional triplet attention module,” in 2021 IEEE Winter Conf. Appl. Comput. Vis. (WACV). 3138–3147, IEEE."/></mixed-citation></ref><ref id="B22"><mixed-citation><named-content content-type="citation-string">
Nair V., Hinton G. E. (2010). “Rectified linear units improve restricted boltzmann machines,” in Proc. 27th Int. Conf. Mach. Learn. (Madison, WI, USA: ) 807–814.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Nair V., Hinton G. E. (2010). “Rectified linear units improve restricted boltzmann machines,” in Proc. 27th Int. Conf. Mach. Learn. (Madison, WI, USA: ) 807–814."/></mixed-citation></ref><ref id="B23"><mixed-citation><named-content content-type="citation-string">
Petrovic A., Jovanovic L., Bacanin N., Zivkovic M., Malisic S. (2024). “Optimizing convolutional networks using a modified metaheuristic for apple tree leaf disease detection,” in 2024 32nd Telecommunications Forum (TELFOR). 1–4, IEEE.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Petrovic A., Jovanovic L., Bacanin N., Zivkovic M., Malisic S. (2024). “Optimizing convolutional networks using a modified metaheuristic for apple tree leaf disease detection,” in 2024 32nd Telecommunications Forum (TELFOR). 1–4, IEEE."/></mixed-citation></ref><ref id="B24"><mixed-citation><named-content content-type="citation-string">
Predić B., Manić D., Saračević M., Karabašević D., Stanujkić D. (2022. a). Automatic image caption generation based on some machine learning algorithms. Math. Problems Eng.
2022, 4001460. doi:  10.1155/2022/4001460
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1155/2022/4001460"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Math. Problems Eng.&amp;title=Automatic image caption generation based on some machine learning algorithms&amp;author=B. Predić&amp;author=D. Manić&amp;author=M. Saračević&amp;author=D. Karabašević&amp;author=D. Stanujkić&amp;volume=2022&amp;publication_year=2022&amp;pages=4001460&amp;doi=10.1155/2022/4001460&amp;"/></mixed-citation></ref><ref id="B25"><mixed-citation><named-content content-type="citation-string">
Predić B., Vukić U., Saračević M., Karabašević D., Stanujkić D. (2022. b). The possibility of combining and implementing deep neural network compression methods. Axioms
11, 229. doi:  10.3390/axioms11050229
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/axioms11050229"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Axioms&amp;title=The possibility of combining and implementing deep neural network compression methods&amp;author=B. Predić&amp;author=U. Vukić&amp;author=M. Saračević&amp;author=D. Karabašević&amp;author=D. Stanujkić&amp;volume=11&amp;publication_year=2022&amp;pages=229&amp;doi=10.3390/axioms11050229&amp;"/></mixed-citation></ref><ref id="B26"><mixed-citation><named-content content-type="citation-string">
Selvaraju R. R., Cogswell M., Das A., Vedantam R., Parikh D., Batra D. (2017). “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision. 618–626, IEEE.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Selvaraju R. R., Cogswell M., Das A., Vedantam R., Parikh D., Batra D. (2017). “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision. 618–626, IEEE."/></mixed-citation></ref><ref id="B27"><mixed-citation><named-content content-type="citation-string">
Sheng X., Wang F., Ruan H., Fan Y., Zheng J., Zhang Y., et al. (2022). Disease diagnostic method based on cascade backbone network for apple leaf disease classification. Front. Plant Sci.
13. doi:  10.3389/fpls.2022.994227, PMID: 
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3389/fpls.2022.994227"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC9539913"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="36212336"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Front. Plant Sci.&amp;title=Disease diagnostic method based on cascade backbone network for apple leaf disease classification&amp;author=X. Sheng&amp;author=F. Wang&amp;author=H. Ruan&amp;author=Y. Fan&amp;author=J. Zheng&amp;volume=13&amp;publication_year=2022&amp;pmid=36212336&amp;doi=10.3389/fpls.2022.994227&amp;"/></mixed-citation></ref><ref id="B28"><mixed-citation><named-content content-type="citation-string">
Simonyan K., Zisserman A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=arXiv preprint arXiv:1409.1556&amp;title=Very deep convolutional networks for large-scale image recognition&amp;author=K. Simonyan&amp;author=A. Zisserman&amp;publication_year=2014&amp;"/></mixed-citation></ref><ref id="B29"><mixed-citation><named-content content-type="citation-string">
Sun C., Li Y., Song Z., Liu Q., Si H., Yang Y., et al. (2025). Research on tomato disease image recognition method based on deit. Eur. J. Agron.
162, 127400. doi:  10.1016/j.eja.2024.127400
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.eja.2024.127400"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Eur. J. Agron.&amp;title=Research on tomato disease image recognition method based on deit&amp;author=C. Sun&amp;author=Y. Li&amp;author=Z. Song&amp;author=Q. Liu&amp;author=H. Si&amp;volume=162&amp;publication_year=2025&amp;pages=127400&amp;doi=10.1016/j.eja.2024.127400&amp;"/></mixed-citation></ref><ref id="B30"><mixed-citation><named-content content-type="citation-string">
Szegedy C., Ioffe S., Vanhoucke V., Alemi A. (2016). Inception-v4, inception-resnet and the impact of residual connections on learning. arXiv preprint arXiv:1602.07261.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=arXiv preprint arXiv:1602.07261&amp;title=Inception-v4, inception-resnet and the impact of residual connections on learning&amp;author=C. Szegedy&amp;author=S. Ioffe&amp;author=V. Vanhoucke&amp;author=A. Alemi&amp;publication_year=2016&amp;"/></mixed-citation></ref><ref id="B31"><mixed-citation><named-content content-type="citation-string">
Szegedy C., Liu W., Jia Y., Sermanet P., Reed S., Anguelov D., et al. (2015). “Going deeper with convolutions,” in 2015 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR). 1–9, IEEE.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Szegedy C., Liu W., Jia Y., Sermanet P., Reed S., Anguelov D., et al. (2015). “Going deeper with convolutions,” in 2015 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR). 1–9, IEEE."/></mixed-citation></ref><ref id="B32"><mixed-citation><named-content content-type="citation-string">
Tan M., Le Q. V. (2020). Efficientnet: Rethinking model scaling for convolutional neural networks. arXiv preprint arXiv:1905.11946.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=arXiv preprint arXiv:1905.11946&amp;title=Efficientnet: Rethinking model scaling for convolutional neural networks&amp;author=M. Tan&amp;author=Q. V. Le&amp;publication_year=2020&amp;"/></mixed-citation></ref><ref id="B33"><mixed-citation><named-content content-type="citation-string">
Tang L., Yi J., Li X. (2024). Improved multi-scale inverse bottleneck residual network based on triplet parallel attention for apple leaf disease identification. J. Integr. Agric.
23, 901—922. doi:  10.1016/j.jia.2023.06.023
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.jia.2023.06.023"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=J. Integr. Agric.&amp;title=Improved multi-scale inverse bottleneck residual network based on triplet parallel attention for apple leaf disease identification&amp;author=L. Tang&amp;author=J. Yi&amp;author=X. Li&amp;volume=23&amp;publication_year=2024&amp;pages=901—922&amp;doi=10.1016/j.jia.2023.06.023&amp;"/></mixed-citation></ref><ref id="B34"><mixed-citation><named-content content-type="citation-string">
Thapa R., Zhang K., Snavely N., Belongie S., Khan A. (2021). Plant pathology
2021, fgvc8. doi:  10.1002/aps3.11390, PMID: </named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1002/aps3.11390"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC7526434"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="33014634"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Thapa R., Zhang K., Snavely N., Belongie S., Khan A. (2021). Plant pathology 2021, fgvc8. doi:  10.1002/aps3.11390, PMID:"/></mixed-citation></ref><ref id="B35"><mixed-citation><named-content content-type="citation-string">
Ullah W., Javed K., Khan M. A., Alghayadh F. Y., Bhatt M. W., Al Naimi I. S., et al. (2024). Efficient identification and classification of apple leaf diseases using lightweight vision transformer (vit). Discov. Sustain.
5, 116. doi:  10.1007/s43621-024-00307-1
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1007/s43621-024-00307-1"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Discov. Sustain.&amp;title=Efficient identification and classification of apple leaf diseases using lightweight vision transformer (vit)&amp;author=W. Ullah&amp;author=K. Javed&amp;author=M. A. Khan&amp;author=F. Y. Alghayadh&amp;author=M. W. Bhatt&amp;volume=5&amp;publication_year=2024&amp;pages=116&amp;doi=10.1007/s43621-024-00307-1&amp;"/></mixed-citation></ref><ref id="B36"><mixed-citation><named-content content-type="citation-string">
Wang H., Cui Y. (2024). Apple leaf disease identification method based on improved shufflenet v2 model. Jiangsu Agric. Sci.
52, 214–222. doi:  10.15889/j.issn.1002-1302.2024.13.028
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.15889/j.issn.1002-1302.2024.13.028"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Jiangsu Agric. Sci.&amp;title=Apple leaf disease identification method based on improved shufflenet v2 model&amp;author=H. Wang&amp;author=Y. Cui&amp;volume=52&amp;publication_year=2024&amp;pages=214-222&amp;doi=10.15889/j.issn.1002-1302.2024.13.028&amp;"/></mixed-citation></ref><ref id="B37"><mixed-citation><named-content content-type="citation-string">
Wang H., Zhang J., Yin Z., Huang L., Wang J., Ma X. (2024). A deep evidence fusion framework for apple leaf disease classification. Eng. Appl. Artif. Intell.
136, 109011. doi:  10.1016/j.engappai.2024.109011
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.engappai.2024.109011"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Eng. Appl. Artif. Intell.&amp;title=A deep evidence fusion framework for apple leaf disease classification&amp;author=H. Wang&amp;author=J. Zhang&amp;author=Z. Yin&amp;author=L. Huang&amp;author=J. Wang&amp;volume=136&amp;publication_year=2024&amp;pages=109011&amp;doi=10.1016/j.engappai.2024.109011&amp;"/></mixed-citation></ref><ref id="B38"><mixed-citation><named-content content-type="citation-string">
Xie S., Girshick R., Dollár P., Tu Z., He K. (2017). “Aggregated residual transformations for deep neural networks,” in 2017 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR). 5987–5995, IEEE.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Xie S., Girshick R., Dollár P., Tu Z., He K. (2017). “Aggregated residual transformations for deep neural networks,” in 2017 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR). 5987–5995, IEEE."/></mixed-citation></ref><ref id="B39"><mixed-citation><named-content content-type="citation-string">
Yang L., Yu X., Zhang S., Long H., Zhang H., Xu S., et al. (2023). Googlenet based on residual network and attention mechanism identification of rice leaf diseases. Comput. Electron. Agric.
204, 107543. doi:  10.1016/j.compag.2022.107543
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.compag.2022.107543"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Comput. Electron. Agric.&amp;title=Googlenet based on residual network and attention mechanism identification of rice leaf diseases&amp;author=L. Yang&amp;author=X. Yu&amp;author=S. Zhang&amp;author=H. Long&amp;author=H. Zhang&amp;volume=204&amp;publication_year=2023&amp;pages=107543&amp;doi=10.1016/j.compag.2022.107543&amp;"/></mixed-citation></ref><ref id="B40"><mixed-citation><named-content content-type="citation-string">
Yang Q., Duan S., Wang L. (2022). Efficient identification of apple leaf diseases in the wild using convolutional neural networks. Agronomy
12, 2784. doi:  10.3390/agronomy12112784
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3390/agronomy12112784"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Agronomy&amp;title=Efficient identification of apple leaf diseases in the wild using convolutional neural networks&amp;author=Q. Yang&amp;author=S. Duan&amp;author=L. Wang&amp;volume=12&amp;publication_year=2022&amp;pages=2784&amp;doi=10.3390/agronomy12112784&amp;"/></mixed-citation></ref><ref id="B41"><mixed-citation><named-content content-type="citation-string">
Zhang Y., Li X., Wang M., Xu T., Huang K., Sun Y., et al. (2024). Early detection and lesion visualization of pear leaf anthracnose based on multi-source feature fusion of hyperspectral imaging. Front. Plant Sci.
15. doi:  10.3389/fpls.2024.1461855, PMID: 
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.3389/fpls.2024.1461855"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmcid" xlink:href="PMC11493603"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="pmid" xlink:href="39439506"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Front. Plant Sci.&amp;title=Early detection and lesion visualization of pear leaf anthracnose based on multi-source feature fusion of hyperspectral imaging&amp;author=Y. Zhang&amp;author=X. Li&amp;author=M. Wang&amp;author=T. Xu&amp;author=K. Huang&amp;volume=15&amp;publication_year=2024&amp;pmid=39439506&amp;doi=10.3389/fpls.2024.1461855&amp;"/></mixed-citation></ref><ref id="B42"><mixed-citation><named-content content-type="citation-string">
Zhang S., Wang D., Yu C. (2023). Apple leaf disease recognition method based on siamese dilated inception network with less training samples. Comput. Electron. Agric.
213, 108188. doi:  10.1016/j.compag.2023.108188
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.compag.2023.108188"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Comput. Electron. Agric.&amp;title=Apple leaf disease recognition method based on siamese dilated inception network with less training samples&amp;author=S. Zhang&amp;author=D. Wang&amp;author=C. Yu&amp;volume=213&amp;publication_year=2023&amp;pages=108188&amp;doi=10.1016/j.compag.2023.108188&amp;"/></mixed-citation></ref><ref id="B43"><mixed-citation><named-content content-type="citation-string">
Zhang X., Zhou X., Lin M., Sun J. (2018). “Shufflenet: An extremely efficient convolutional neural network for mobile devices,” in 2018 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. 6848–6856, IEEE.</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="Zhang X., Zhou X., Lin M., Sun J. (2018). “Shufflenet: An extremely efficient convolutional neural network for mobile devices,” in 2018 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. 6848–6856, IEEE."/></mixed-citation></ref><ref id="B44"><mixed-citation><named-content content-type="citation-string">
Zheng J., Li K., Wu W., Ruan H. (2023). Repdi: A light-weight cpu network for apple leaf disease identification. Comput. Electron. Agric.
212, 108122. doi:  10.1016/j.compag.2023.108122
</named-content><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="doi" xlink:href="10.1016/j.compag.2023.108122"/><ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="google-scholar" xlink:href="journal=Comput. Electron. Agric.&amp;title=Repdi: A light-weight cpu network for apple leaf disease identification&amp;author=J. Zheng&amp;author=K. Li&amp;author=W. Wu&amp;author=H. Ruan&amp;volume=212&amp;publication_year=2023&amp;pages=108122&amp;doi=10.1016/j.compag.2023.108122&amp;"/></mixed-citation></ref></ref-list></sec></sec><sec id="_ad93_" xml:lang="en" sec-type="associated-data" disp-level="1"><title>Associated Data</title><sec id="_adda93_" xml:lang="en" sec-type="data-availability-statement" disp-level="2"><title>Data Availability Statement</title><p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8" ext-link-type="uri">https://www.kaggle.com/competitions/plant-pathology-2021-fgvc8</ext-link>; <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/JasonYangCode/AppleLeaf9" ext-link-type="uri">https://github.com/JasonYangCode/AppleLeaf9</ext-link>; <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.scidb.cn/en/detail?dataSetId=0e1f57004db842f99668d82183afd578" ext-link-type="uri">https://www.scidb.cn/en/detail?dataSetId=0e1f57004db842f99668d82183afd578</ext-link>.</p></sec></sec></body></article>