1.An efficient and lightweight skin pathology detection method based on multi-scale feature fusion using an improved RT-DETR model.
Yuying REN ; Lingxiao HUANG ; Fang DU ; Xinbo YAO
Journal of Southern Medical University 2025;45(2):409-421
OBJECTIVES:
The presence of multi-scale skin lesion regions and image noise interference and limited resources of auxiliary diagnostic equipment affect the accuracy of skin disease detection in skin disease detection tasks. To solve these problems, we propose a highly efficient and lightweight skin disease detection model using an improved RT-DETR model.
METHODS:
A lightweight FasterNet was introduced as the backbone network and the FasterNetBlock module was parametrically refined. A Convolutional and Attention Fusion Module (CAFM) was used to replace the multi-head self-attention mechanism in the neck network to enhance the ability of the AIFI-CAFM module for capturing global dependencies and local detail information. The DRB-HSFPN feature pyramid network was designed to replace the Cross-Scale Feature Fusion Module (CCFM) to allow the integration of contextual information across different scales to improve the semantic feature expression capacity of the neck network. Finally, combining the advantages of Inner-IoU and EIoU, the Inner-EIoU was used to replace the original loss function GIOU to further enhance the model's inference accuracy and convergence speed.
RESULTS:
The experimental results on the HAM10000 dataset showed that the improved RT-DETR model, as compared with the original model, had increased mAP@50 and mAP@50:95 by 4.5% and 2.8%, respectively, with a detection speed of 59.1 frames per second (FPS). The improved model had a parameter count of 10.9 M and a computational load of 19.3 GFLOPs, which were reduced by 46.0% and 67.2% compared to those of the original model, validating the effectiveness of the improved model.
CONCLUSIONS
The proposed SD-DETR model significantly improves the performance of skin disease detection tasks by effectively extracting and integrating multi-scale features while reducing both parameter count and computational load.
Humans
;
Skin Diseases/diagnosis*
;
Skin/pathology*
;
Neural Networks, Computer
;
Algorithms
2.A multi-scale supervision and residual feedback optimization algorithm for improving optic chiasm and optic nerve segmentation accuracy in nasopharyngeal carcinoma CT images.
Jinyu LIU ; Shujun LIANG ; Yu ZHANG
Journal of Southern Medical University 2025;45(3):632-642
OBJECTIVES:
We propose a novel deep learning segmentation algorithm (DSRF) based on multi-scale supervision and residual feedback strategy for precise segmentation of the optic chiasm and optic nerves in CT images of nasopharyngeal carcinoma (NPC) patients.
METHODS:
We collected 212 NPC CT images and their ground truth labels from SegRap2023, StructSeg2019 and HaN-Seg2023 datasets. Based on a hybrid pooling strategy, we designed a decoder (HPS) to reduce small organ feature loss during pooling in convolutional neural networks. This decoder uses adaptive and average pooling to refine high-level semantic features, which are integrated with primary semantic features to enable network learning of finer feature details. We employed multi-scale deep supervision layers to learn rich multi-scale and multi-level semantic features under deep supervision, thereby enhancing boundary identification of the optic chiasm and optic nerves. A residual feedback module that enables multiple iterations of the network was designed for contrast enhancement of the optic chiasm and optic nerves in CT images by utilizing information from fuzzy boundaries and easily confused regions to iteratively refine segmentation results under supervision. The entire segmentation framework was optimized with the loss from each iteration to enhance segmentation accuracy and boundary clarity. Ablation experiments and comparative experiments were conducted to evaluate the effectiveness of each component and the performance of the proposed model.
RESULTS:
The DSRF algorithm could effectively enhance feature representation of small organs to achieve accurate segmentation of the optic chiasm and optic nerves with an average DSC of 0.837 and an ASSD of 0.351. Ablation experiments further verified the contributions of each component in the DSRF method.
CONCLUSIONS
The proposed deep learning segmentation algorithm can effectively enhance feature representation to achieve accurate segmentation of the optic chiasm and optic nerves in CT images of NPC.
Humans
;
Tomography, X-Ray Computed/methods*
;
Optic Chiasm/diagnostic imaging*
;
Optic Nerve/diagnostic imaging*
;
Algorithms
;
Nasopharyngeal Carcinoma
;
Deep Learning
;
Nasopharyngeal Neoplasms/diagnostic imaging*
;
Neural Networks, Computer
;
Image Processing, Computer-Assisted/methods*
3.A lightweight classification network for single-lead atrial fibrillation based on depthwise separable convolution and attention mechanism.
Yong HONG ; Xin ZHANG ; Mingjun LIN ; Qiucen WU ; Chaomin CHEN
Journal of Southern Medical University 2025;45(3):650-660
OBJECTIVES:
To design a deep learning model that balances model complexity and performance to enable its integration into wearable ECG monitoring devices for automated diagnosis of atrial fibrillation.
METHODS:
This study was performed based on data from 84 patients with atrial fibrillation, 25 patients with atrial fibrillation, and 18 subjects without obvious arrhythmia collected from the publicly available datasets LTAFDB, AFDB, and NSRDB, respectively. A lightweight attention network based on depthwise separable convolution and fusion of channel-spatial information, namely DSC-AttNet, was proposed. Depthwise separable convolution was introduced to replace standard convolution and reduce model parameters and computational complexity to realize high efficiency and light weight of the model. The multilayer hybrid attention mechanism was embedded to compute the attentional weights of the channels and spatial information at different scales to improve the feature expression ability of the model. Ten-fold cross-validation was performed on LTAFDB, and external independent testing was conducted on AFDB and NSRDB datasets.
RESULTS:
DSC-AttNet achieved a ten-fold average accuracy of 97.33% and a precision of 97.30% on the test set, both of which outperformed the other 4 comparison models as well as the 3 classical models. The accuracy of the model on the external test set reached 92.78%, better than those of the 3 classical models. The number of parameters of DSC-AttNet was 1.01M, and the computational volume was 27.19G, both smaller than the 3 classical models.
CONCLUSIONS
This proposed method has a smaller complexity, achieves better classification performance, and has a better generalization ability for atrial fibrillation classification.
Atrial Fibrillation/diagnosis*
;
Humans
;
Electrocardiography
;
Deep Learning
;
Wearable Electronic Devices
;
Neural Networks, Computer
4.AConvLSTM U-Net: a multi-scale jaw cyst segmentation model based on bidirectional dense connection and attention mechanism.
Suqiang LI ; Zhouyang WANG ; Sixian CHAN ; Xiaolong ZHOU
Journal of Southern Medical University 2025;45(5):1082-1092
OBJECTIVES:
We propose a multi-scale jaw cyst segmentation model, AConvLSTM U-Net, which is based on bidirectional dense connections and attention mechanisms to achieve accurate automatic segmentation of mandibular cyst images.
METHODS:
A dataset consisting of 2592 jaw cyst images was used. AConvLSTM U-Net designs a MBC on the encoding path to enhance feature extraction capabilities. A DPD was used to connect the encoder and decoder, and a bidirectional ConvLSTM was introduced in the jump connection to obtain rich semantic information. A decoding block based on scSE was then used on the decoding path to enhance the focus on important information. Finally, a DS was designed, and the model was optimized by integrating a joint loss function to further improve the segmentation accuracy.
RESULTS:
The experiment with AConvLSTM U-Net for jaw cyst lesion segmentation showed a MCC of 93.8443%, a DSC of 93.9067%, and a JSC of 88.5133%, outperforming all the other comparison segmentation models.
CONCLUSIONS
The proposed algorithm shows a high accuracy and robustness on the jaw cyst dataset, demonstrating its superior performance over many existing methods for automatic segmentation of jaw cyst images and its potential to assist clinical diagnosis.
Humans
;
Jaw Cysts/diagnostic imaging*
;
Algorithms
;
Image Processing, Computer-Assisted/methods*
;
Neural Networks, Computer
5.SG-UNet: a melanoma segmentation model enhanced with global attention and self-calibrated convolution.
Huanyu JI ; Rui WANG ; Shengxiang GAO ; Wengang CHE
Journal of Southern Medical University 2025;45(6):1317-1326
OBJECTIVES:
We propose a new melanoma segmentation model, SG-UNet, to enhance the precision of melanoma segmentation in dermascopy images to facilitate early melanoma detection.
METHODS:
We utilized a U-shaped convolutional neural network, UNet, and made improvements to its backbone, skip connections, and downsampling pooling sections. In the backbone, with reference to the structure of VGG, we increased the number of convolutions from 10 to 13 in the downsampling part of UNet to achieve a deepened network hierarchy that allowed capture of more refined feature representations. To further enhance feature extraction and detail recognition, we replaced the traditional convolution the backbone section with self-calibrated convolution to enhance the model's ability to capture both spatial and channel dimensional features. In the pooling part, the original pooling layer was replaced by Haar wavelet downsampling to achieve more effective multi-scale feature fusion and reduce the spatial resolution of the feature map. The global attention mechanism was then incorporated into the skip connections at each layer to enhance the understanding of contextual information of the image.
RESULTS:
The experimental results showed that the SG-UNet model achieved significantly improved segmentation accuracy on ISIC 2017 and ISIC 2018 datasets as compared with other current state-of-the-art segmentation models, with Dice reached 92.41% and 86.62% and IoU reaching 92.31% and 86.48% on the two datasets, respectively.
CONCLUSIONS
The proposed model is capable of effective and accurate segmentation of melanoma from dermoscopy images.
Melanoma/diagnosis*
;
Humans
;
Neural Networks, Computer
;
Dermoscopy/methods*
;
Skin Neoplasms
;
Image Processing, Computer-Assisted/methods*
;
Calibration
;
Algorithms
6.A multi-feature fusion-based model for fetal orientation classification from intrapartum ultrasound videos.
Ziyu ZHENG ; Xiaying YANG ; Shengjie WU ; Shijie ZHANG ; Guorong LYU ; Peizhong LIU ; Jun WANG ; Shaozheng HE
Journal of Southern Medical University 2025;45(7):1563-1570
OBJECTIVES:
To construct an intelligent analysis model for classifying fetal orientation during intrapartum ultrasound videos based on multi-feature fusion.
METHODS:
The proposed model consists of the Input, Backbone Network and Classification Head modules. The Input module carries out data augmentation to improve the sample quality and generalization ability of the model. The Backbone Network was responsible for feature extraction based on Yolov8 combined with CBAM, ECA, PSA attention mechanism and AIFI feature interaction module. The Classification Head consists of a convolutional layer and a softmax function to output the final probability value of each class. The images of the key structures (the eyes, face, head, thalamus, and spine) were annotated with frames by physicians for model training to improve the classification accuracy of the anterior occipital, posterior occipital, and transverse occipital orientations.
RESULTS:
The experimental results showed that the proposed model had excellent performance in the tire orientation classification task with the classification accuracy reaching 0.984, an area under the PR curve (average accuracy) of 0.993, and area under the ROC curve of 0.984, and a kappa consistency test score of 0.974. The prediction results by the deep learning model were highly consistent with the actual classification results.
CONCLUSIONS
The multi-feature fusion model proposed in this study can efficiently and accurately classify fetal orientation in intrapartum ultrasound videos.
Humans
;
Female
;
Ultrasonography, Prenatal/methods*
;
Pregnancy
;
Fetus/diagnostic imaging*
;
Neural Networks, Computer
;
Video Recording
7.A GA-BP neural network model based on spectrum-effect relationship for assessing spectrum-effect score and quality evaluation of Cassia seeds extract.
Haiyan YAN ; Heng WANG ; Chuncai ZOU
Journal of Southern Medical University 2025;45(10):2092-2103
OBJECTIVES:
To construct a GA-BP neural network model based on the spectrum-effect relationship of Cassia seeds extract and test its performance for quality control of Cassia seeds using spectrum-effect score.
METHODS:
The HPLC fingerprints of Cassia seeds extract (0.1, 0.2, and 0.4 g/mL) were established. In a mouse model of 5-Fu-induced liver injury treated with 0.4, 0.8, and 1.6 g/kg of Cassia seeds extract, the pharmacodynamics parameters were measured to calculate the comprehensive efficacy using AHP-EWM. A GA-BP neural network model between the fingerprints and comprehensive efficacy was constructed, and the corresponding predicted comprehensive efficacy was obtained. The spectrum-effect relationship between the fingerprints and the measured and predicted comprehensive efficacy was established using grey correlation method followed by Gaussian fitting analysis. The spectral efficiency score was calculated using the relative peak area of the fingerprints and the correlation degree of the spectral efficiency. The reliability of the data was tested using the Z-ratio score method. The limit range of the spectral efficiency score was determined and the quality of the verification samples was evaluated.
RESULTS:
The error between the predicted value using the GA-BP neural network model and the measured value of the comprehensive efficacy was less than 0.2. Gaussian fitting analysis showed good fitting between the spectrum-effect relationship data of the measured and predicted comprehensive efficacy. The limit of the spectral efficiency score was 6.16-7.30. The prediction results for each verification group were consistent with the experimental results and within the limit of spectral efficiency score, and the results of Z-ratio score analysis demonstrated good data reliability.
CONCLUSIONS
The GA-BP neural network model can effectively predict the comprehensive efficacy of Cassia seeds extract, and the established spectrum-effect scoring method can be used for quality evaluation of samples.
Neural Networks, Computer
;
Animals
;
Seeds/chemistry*
;
Mice
;
Cassia/chemistry*
;
Quality Control
;
Drugs, Chinese Herbal/pharmacology*
;
Plant Extracts/pharmacology*
;
Male
8.A heterogeneous graph method integrating multi-layer semantics and topological information for improving drug-target interaction prediction.
Zihao CHEN ; Yanbu GUO ; Shengli SONG ; Quanming GUO ; Dongming ZHOU
Journal of Southern Medical University 2025;45(11):2394-2404
OBJECTIVES:
To develop a heterogeneous graph prediction method based on the fusion of multi-layer semantics and topological information for addressing the challenges in drug-target interaction prediction, including insufficient modeling of high-order semantic dependencies, lack of adaptive fusion of semantic paths, and over-smoothing of node features.
METHODS:
A heterogeneous graph network with multiple types of entities such as drugs, proteins, side effects, and diseases was constructed, and graph embedding techniques were used to obtain low-dimensional feature representations. An adaptive metapath search module was introduced to automatically discover semantic path combinations for guiding the propagation of high-order semantic information. A semantic aggregation mechanism integrating multi-head attention was designed to automatically learn the importance of each semantic path based on contextual information and achieve differentiated aggregation and dynamic fusion among paths. A structure-aware gated graph convolutional module was then incorporated to regulate the feature propagation intensity for suppressing redundant information and redcuing over-smoothing. Finally, the potential interactions between drugs and targets were predicted through an inner product operation.
RESULTS:
Compared with existing drug-target interaction prediction methods, the proposed method achieved an average improvement of 3.4% and 2.4%, 3.0% and 3.8% in terms of the area under the receiver operating characteristic curve (AUC) and the area under the precision-recall curve (AUPRC) on public datasets, respectively.
CONCLUSIONS
The drug-target interaction prediction method developed in this study can effectively extract complex high-order semantic and topological information from heterogeneous biological networks, thereby improving the accuracy and stability of drug-target interaction prediction. This method provides technical support and theoretical foundation for precise drug target discovery and targeted treatment of complex diseases.
Semantics
;
Humans
;
Drug Interactions
;
Neural Networks, Computer
;
Algorithms
9.A Novel Real-time Phase Prediction Network in EEG Rhythm.
Hao LIU ; Zihui QI ; Yihang WANG ; Zhengyi YANG ; Lingzhong FAN ; Nianming ZUO ; Tianzi JIANG
Neuroscience Bulletin 2025;41(3):391-405
Closed-loop neuromodulation, especially using the phase of the electroencephalography (EEG) rhythm to assess the real-time brain state and optimize the brain stimulation process, is becoming a hot research topic. Because the EEG signal is non-stationary, the commonly used EEG phase-based prediction methods have large variances, which may reduce the accuracy of the phase prediction. In this study, we proposed a machine learning-based EEG phase prediction network, which we call EEG phase prediction network (EPN), to capture the overall rhythm distribution pattern of subjects and map the instantaneous phase directly from the narrow-band EEG data. We verified the performance of EPN on pre-recorded data, simulated EEG data, and a real-time experiment. Compared with widely used state-of-the-art models (optimized multi-layer filter architecture, auto-regress, and educated temporal prediction), EPN achieved the lowest variance and the greatest accuracy. Thus, the EPN model will provide broader applications for EEG phase-based closed-loop neuromodulation.
Humans
;
Electroencephalography/methods*
;
Brain/physiology*
;
Machine Learning
;
Signal Processing, Computer-Assisted
;
Male
;
Adult
;
Neural Networks, Computer
;
Brain Waves/physiology*
10.Prediction of Pharmacoresistance in Drug-Naïve Temporal Lobe Epilepsy Using Ictal EEGs Based on Convolutional Neural Network.
Yiwei GONG ; Zheng ZHANG ; Yuanzhi YANG ; Shuo ZHANG ; Ruifeng ZHENG ; Xin LI ; Xiaoyun QIU ; Yang ZHENG ; Shuang WANG ; Wenyu LIU ; Fan FEI ; Heming CHENG ; Yi WANG ; Dong ZHOU ; Kejie HUANG ; Zhong CHEN ; Cenglin XU
Neuroscience Bulletin 2025;41(5):790-804
Approximately 30%-40% of epilepsy patients do not respond well to adequate anti-seizure medications (ASMs), a condition known as pharmacoresistant epilepsy. The management of pharmacoresistant epilepsy remains an intractable issue in the clinic. Its early prediction is important for prevention and diagnosis. However, it still lacks effective predictors and approaches. Here, a classical model of pharmacoresistant temporal lobe epilepsy (TLE) was established to screen pharmacoresistant and pharmaco-responsive individuals by applying phenytoin to amygdaloid-kindled rats. Ictal electroencephalograms (EEGs) recorded before phenytoin treatment were analyzed. Based on ictal EEGs from pharmacoresistant and pharmaco-responsive rats, a convolutional neural network predictive model was constructed to predict pharmacoresistance, and achieved 78% prediction accuracy. We further found the ictal EEGs from pharmacoresistant rats have a lower gamma-band power, which was verified in seizure EEGs from pharmacoresistant TLE patients. Prospectively, therapies targeting the subiculum in those predicted as "pharmacoresistant" individual rats significantly reduced the subsequent occurrence of pharmacoresistance. These results demonstrate a new methodology to predict whether TLE individuals become resistant to ASMs in a classic pharmacoresistant TLE model. This may be of translational importance for the precise management of pharmacoresistant TLE.
Epilepsy, Temporal Lobe/diagnosis*
;
Animals
;
Drug Resistant Epilepsy/drug therapy*
;
Electroencephalography/methods*
;
Rats
;
Anticonvulsants/pharmacology*
;
Neural Networks, Computer
;
Male
;
Humans
;
Phenytoin/pharmacology*
;
Adult
;
Disease Models, Animal
;
Female
;
Rats, Sprague-Dawley
;
Young Adult
;
Convolutional Neural Networks

Result Analysis
Print
Save
E-mail