Keji Mao

dblp:31/5600 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0002-5021-378XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HSDNet: Hierarchical Semantic Decoupling Network for Camouflaged Object Detection
Zhihu Zhou, Yongbiao Zhao, Junyong Sun, Ruiji Xu, Keji Mao
ICIC (20)6
2026 Automated Detection of Abnormal Wrist Bone Morphology for Bone Age Assessment Using a Dual-Branch Feature Fusion Network
Ruiji Xu, Yan Ling, Keji Mao
ICIC (29)6
2026 Precise bone age assessment via multi-task feature disentanglement and quadruplet ranking
Guoyong Dai, Yantao Shao, Ruiji Xu, Keji Mao, Xiaofeng Qian
Neurocomputing7
2026 MSE-LAM: Multi-scale emotion recognition network based on local attention mask
Lingkang Ying, Zhuchenghao Wang, Yan Ling, Ruiji Xu, Keji Mao, Weiyuan Zhou, Zhitian Zhang
Neurocomputing6
2026 DSTA: A dual-channel deep learning framework with spatiotemporal attention for schizophrenia screening
Runzhe Zhang, Zhou Yuan, Bowei Pan, Yan Ling, Ruiji Xu, Keji Mao
Neurocomputing7
2026 MambaFlow: Cross-frame state space modeling for efficient optical flow estimation
Zhihu Zhou, Zhuchenghao Wang, Juntian Du, Pinyi Chen, Weiping Ye, Ruiji Xu, Yongbiao Zhao, Keji Mao
Neurocomputing8
2026 FaceDepth: A Robust Unimodal Depression Detection Framework Using Invariant Facial Landmark Features
abstract
Although significant progress has been made in automatic diagnosis systems for depression, most of the work focuses on combining features from multiple modalities to improve classification accuracy, which generates a lot of space-time overhead and feature synchronization problems. This research work proposes a unimodal depression detection framework based on facial expression and facial motion features. Firstly, we propose a robust feature extraction method based on the ratio of facial landmark and theoretically prove that this feature has up-down, left-right translation, depth translation, rotation, and flip invariance. The features extracted based on this method maintain the topological structure relationship of facial landmarks in space and maintain the temporal correlation of frames before and after facial landmarks. Then, we provide a novel idea to solve the classification task of large-unit depression videos. The final depression classification result is obtained by decomposing the depression classification task of large-unit videos into the scoring task of multiple short-sequence units and then through the defined score aggregation function. Our key innovations include: (1) theoretically proven invariant facial landmark ratio features, (2) novel video decomposition into short-sequence units with pseudo-labeling, and (3) efficient SRTSNet architecture. On DAIC-WOZ dataset, our framework achieves F1 = 0.85, outperforming all unimodal methods and matching state-of-the-art multimodal approaches while using only facial features.
Ruiji Xu, Runzhe Zhang, Guanglin Dai, Keji Mao
ACM Trans. Multim. Comput. Commun. Appl.5
2025 A Prediction Method for Adult Height of Children Based on ACPSO-SVR
Tianxiang He, Ziqi Qian, Keji Mao
ICIC (25)6
2025 An Intelligent Detection Method for Safety Equipment Non-Compliance in High-Altitude Power Grid Operation
Changquan He, Kaikai Chi, Keji Mao
ICIC (2)4
2025 MeFacialNet: High-Accuracy Micro-expression Recognition on Facial Optical Flow
abstract
Expression recognition has witnessed substantial advancements in recent years due to the incorporation of the optical flow models. However, it still faces challenges in accurately capturing the nuances of facial micro-expressions, primarily due to insufficient micro-expression features and inadequate flow features for precise expression estimation. In this paper, we propose a high accuracy Micro-expression recognition based on Facial optical flow(MeFacialNet), which combines a facial semantic-aware encoder to extract micro-facial features, and a dynamical expression-aware decoder that decomposed into head and facial components to accurately estimate facial expression. Firstly, it innovatively incorporates a micro-expression recognition module for feature enhancement. Furthermore, a dynamically parameterized weights learning mechanism is proposed for the final expression prediction, addressing the issue of hand-crafted weights being unsuitable for real-world applications. Experimental results demonstrate that our method improves accuracy by 18.9% (from 0.132 to 0.107) compared to DecFlow, the best-performing method to date, and outperforms other mainstream models by at least 24% on the FFN dataset.
Shenggang Qian, Pinyi Chen, Runzhe Zhang, Zhihu Zhou, Keji Mao
IJCNN5
2025 ST-LCNet: A Dual-Stream Transformer for Interpretable EEG-Based Anxiety Diagnosis
abstract
Anxiety disorder is a common mental health condition, and accurate diagnosis is crucial for patient treatment and prognosis. Electroencephalography (EEG), as a non-invasive neurophysiological diagnostic tool, directly reflects the functional state of the brain, offering a promising approach for the objective diagnosis of anxiety disorder. However, EEG signals exhibit high-dimensionality, nonlinearity, and non-stationarity, making it challenging for traditional methods to effectively capture their complex spatiotemporal patterns. To address this issue, this paper proposes a novel spatiotemporal feature extraction network, ST-LCNet (Sparse-channel Temporal Transformer Network). The network optimizes the input features using a spatio-spectral attention mechanism, adopts a dual-stream Transformer architecture to model the global dependencies in both the temporal and spatial dimensions, and integrates local feature extraction modules to capture multi-scale features. Experimental results show that ST-LCNet achieves a classification accuracy of 97.07% in anxiety disorder diagnosis, significantly outperforming existing methods. Furthermore, this study also constructs a multilevel interpretability framework, providing intuitive and reliable reference for clinical diagnosis through the visualization of attention distributions across key brain regions and frequency bands.
Kaichen Shen, Ruiji Xu, Keji Mao
IJCNN5
2025 Multidimensional Speech Feature Extraction for Depression Detection using MDCF-Net
abstract
Depression, a mental health illness that affects more than 350 million people worldwide, frequently lacks obvious diagnostic signs, making precise and efficient detection difficult. In this paper, we present MDCF-Net, a unique Speech Emotion Recognition (SER) system that uses multidimensional convolutional neural networks to automatically extract emotional aspects from speech, hence aiding in depression identification. Our model uses both 1D and 2D convolutions to improve the extraction of temporal, spectral, and multi-channel properties from Mel-Frequency Cepstral Coefficients (MFCC). In addition, we use a Cross Attention Transformer and Global Average Pooling (GAP) to improve emotion classification. When tested on several emotional speech datasets, MDCF-Net achieves a stunning 98.51% accuracy on the ESD_Chinese dataset, beating previous algorithms in emotion recognition. Our approach represents a promising development in real-time mental health monitoring via speech analysis.
Ligang Ren, Yan Ling, Tianxiang He, Juntian Du, Ruiji Xu, Hengtan Zhang, Keji Mao
ISCAS7
2024 Optimizing Future Predictions in Children's Health: Implementing OGPA-enhanced Deep Learning for Precise Child Height Forecasting in the Social Media Age
abstract
In recent years, the issue of children's height has garnered widespread attention on social media. Social media platforms serve as pivotal communication channels for parents, doctors, educators, and researchers, with children's height, a crucial health indicator, becoming one of the hot topics. Inaccurate prediction methods might mislead the public, resulting in parents harboring erroneous expectations regarding their children's future height. Such misplaced expectations may culminate in unwarranted worries or pressure. In pursuit of devising an accurate height prediction model, this paper thoroughly leverages a substantial data sample obtained from the physical health examinations of primary and secondary school students in Zhejiang Province, as well as continuous observation samples provided by the Zhejiang Provincial Bone Age Research Center, to delve deeply into the issue of height prediction in children and adolescents. In this research, we have developed a lightweight neural net-work model suitable for particle swarm optimization to predict children’s stage-wise height. When the difference between the actual and predicted values is within ±2cm, the prediction accuracy for boys reached 86.67%, and for girls, it was 85.32%, with an RMSE of 1.3503.
Yantao Shao, Tianxiang He, Kai Fang 0001, Wei Wang 0077, Keji Mao
CSCWD6
2024 Depression Detection Based on Multilevel Semantic Features
Xingda Yao, Lingkang Ying, Tianxiang He, Ligang Ren, Ruiji Xu, Keji Mao
ICANN (8)6
2024 A ROI Extraction Method for Wrist Imaging Applied in Smart Bone-Age Assessment System
abstract
Bone Age (BA) is reckoned to be closely associated with the growth and development of teenagers, whose assessment highly depends on the accurate extraction of the reference bone from the carpal bone. Being uncertain in its proportion and irregular in its shape, wrong judgment and poor average extraction accuracy of the reference bone will no doubt lower the accuracy of Bone Age Assessment (BAA). In recent years, machine learning and data mining are widely embraced in smart healthcare systems. Using these two instruments, this article aims to tackle the aforementioned problems by proposing a Region of Interest (ROI) extraction method for wrist X-ray images based on optimized YOLO model. The method combines Deformable convolution-focus (Dc-focus), Coordinate attention (Ca) module, Feature level expansion, and Efficient Intersection over Union (EIoU) loss all together as YOLO-DCFE. With the improvement, the model can better extract the features of irregular reference bone and reduce the potential misdiscrimination between the reference bone and other similarly shaped reference bones, improving the detection accuracy. We select 10041 images taken by professional medical cameras as the dataset to test the performance of YOLO-DCFE. Statistics show the advantages of YOLO-DCFE in detection speed and high accuracy. The detection accuracy of all ROIs is 99.8%, which is higher than other models. Meanwhile, YOLO-DCFE is the fastest of all comparison models, with the Frames Per Second (FPS) reaching 16.
Jinfeng Xu 0003, Jianan Wu, Kunxiu Wu, Keji Mao, Kai Fang 0001
IEEE J. Biomed. Health Informatics6
2023 Multi-branch feature learning based speech emotion recognition using SCAR-NET
abstract
Speech emotion recognition (SER) is an active research area in affective computing. Recognizing emotions from speech signals helps to assess human behaviour, which has promising applications in the area of human-computer interaction. The performance of deep learning-based SER methods relies heavily on feature learning. In this paper, we propose SCAR-NET, an improved convolutional neural network, to extract emotional features from speech signals and implement classification. This work includes two main parts: First, we extract spectral, temporal, and spectral-temporal correlation features through three parallel paths; and then split-convolve-aggregate residual blocks are designed for multi-branch deep feature learning. The features are refined by global average pooling (GAP) and pass through a softmax classifier to generate predictions for different emotions. We also conduct a series of experiments to evaluate the robustness and effectiveness of SCAR-NET which can achieve 96.45%, 83.13%, and 89.93% accuracy on the speech emotion datasets EMO-DB, SAVEE, and RAVDESS. These results show the outperformance of SCAR-NET.
Keji Mao, Ligang Ren, Jiefan Qiu, Guanglin Dai
Connect. Sci.1
2023 Hamate classification method based on feature-enhanced residual network and probabilistic joint judgment
abstract
Abstract Accurate analysis of the maturity grade of the reference bone in the wrist is critical for bone age assessment. The main difference between the maturity grades of the reference bone in the wrist X‐ray image is the difference in the texture and morphological characteristics of the bone, and the difference in the characteristics between adjacent levels is small, which brings great challenges in bone age assessment. The hamate (the wrist bone in line with the 4th and 5th fingers) is one of the most important reference bones in the standards of skeletal maturity of the hand and wrist for the Chinese (CHN) method. Aiming at the problem of hamate maturity level evaluation, a method of hamate maturity level classification based on feature enhanced residual network and probability joint judgment is proposed. In this method, we propose the enhanced characteristic residual network (ECR‐Net) is proposed to enhance the feature extraction capability of the network and improve the loss function to reduce the impact of cross‐grade errors on the accuracy of bone age assessment. On this basis, multiple convolutional neural networks to make joint probabilistic judgments are relied on to obtain the final hamate maturity grade. The proposed ECR‐Net achieves a macro accuracy of 96.92% on the hamate maturity evaluation, and there are no errors across two levels. Also, this approach achieves a macro accuracy of 97.15% when the probabilistic joint judgment method is employed while avoiding errors that span more than two levels. The network model and method proposed in this paper have good accuracy and practicability for the automatic identification of the hamate maturity level, which is of great significance for the accurate assessment of bone age.
Weilong Ding 0001, Ze-yong Zong, Keji Mao
IET Image Process.4
2023 Trinity-Yolo: High-precision logo detection in the real world
abstract
Abstract Logo detection has a wide range of applications in the multimedia field, such as video advertising research, brand awareness monitoring and analysis, trademark infringement detection, autonomous driving and intelligent transportation. Compared with other types of images, logo images in the real world have greater diversity in appearance and more complex backgrounds. Therefore, identifying logos from images is a challenge. A strong baseline method Trinity‐Yolo, is proposed, which incorporates attention mechanism, stripe pooling and weighted boxes fusion (WBF) into the state‐of‐the‐art Yolov4 framework for large‐scale logo detection. The attention mechanism improves the feature extraction ability of the deep detection model, the stripe pooling expands the field of view of the model and the weighted boxes fusion enables the model to obtain excellent corrections when outputting the prediction boxes. Trinity‐Yolo can solve the problems of lack of training data, multi‐scale objects and inconsistent bounding‐box regression. On the dataset LogoDet‐3K, the average performance of Trinity‐Yolo is 3% higher than that of Yolov4. Compared with other deep detection models, the performance of Trinity‐Yolo is improved more. The experimental performance on other existing datasets verifies the effectiveness of this method.
Keji Mao, Runhui Jin, Kaiyan Chen, Jiafa Mao, Guanglin Dai
IET Image Process.1
2023 IDRes: Identity-Based Respiration Monitoring System for Digital Twins Enabled Healthcare
abstract
Currently, powerful and ubiquitous mobile devices provide an opportunity to map physical conditions to cyberspace and realize Digital Twins enabled Healthcare (DTeH). Especially, the impact of the COVID-19 epidemic renders it necessary to keep an eye on the changing trend of respiration. Long-term respiration monitoring helps to assess personal health status and thus becomes an important issue in DTeH. However, previous mobile device-assistant methods mostly implement the monitoring via short-time detection in a best-effort way and with less consideration of identity recognition, the only mean to bind physical vital signs into personal profiles in digital twins space. Thus, it is necessary to introduce the identification to complete string multiple short-time detections and form long-term personal monitoring. To this end, we propose IDRes, an identity-based respiration monitoring system for DTeH. This system employs mobile devices to generate a high-frequency sonar signal to complete respiration detection and identity recognition. As well as it also estimates the respiration rate by tracking the phase change of the sonar signal and recognizes identity via the Doppler frequency shift of the signal to capture characteristics of chest movement. Moreover, via band-pass filtering to remove the low-frequency voice component of the received signals, the usage of the high-frequency sonar signal also enhances security at the physical level. At last, we conduct a series of experiments under different conditions. Experimental results illustrate that IDRes achieves the mean detection error of 0.49bpm with over 93.3% recognition accuracy, and manifest that IDRes can satisfy the requirements of mapping the accurate vital sign data to the personal profile of DTeH.
Kai Fang 0001, Jiefan Qiu, Tingting Wang 0006, Kailu Zheng, Liyao Xing, Keji Mao, Kaikai Chi
IEEE J. Sel. Areas Commun.6
2022 A Multitarget Interested Region Extraction Method for Wrist X-Ray Images Based on Optimized AlexNet and Two-Class Combined Model
abstract
Bone age assessment based on X-ray Images can accurately determine the actual bone age of adolescents. Accurate extraction of the key regions of interest (ROIs) in X-ray images is required to accurately assess bone age. However, existing ROI extraction methods can only extract a few targets and have poor extraction accuracy. Thus, the strict demands for imaging in the medical field via these methods are difficult. In this article, we propose a multitarget interested region extraction method for wrist X-ray images based on optimized AlexNet and two-class combined model, named OATC, which can simultaneously extract multiple ROIs with high accuracy. Specifically, the square-wave scanning algorithm was implemented to obtain the bounding box size of each bone ROI according to the shape information of the wrist. Then, the optimized AlexNet was used to obtain the key point coordinates of each bone ROI. Bone ROIs could be extracted by combining key point coordinates with the bone bounding box size. Finally, the two-classification model was combined to improve the accuracy of ROI extraction. Experiments conducted on our wrist X-ray image dataset showed that OTAC has a fast convergence speed and small deviation. The average accuracy of extracting 14 ROIs reached 95.57%, which are 7.76% and 4.68% higher than that of VGG16 and AlexNet, respectively.
Kai Fang 0001, Xiaolong Zhou 0001, Keji Mao
IEEE Trans. Comput. Soc. Syst.5
2021 A fast calibration algorithm for Non-Dispersive Infrared single channel carbon dioxide sensor based on deep learning
Keji Mao, Runhui Jin, Kai Fang 0001
Comput. Commun.1
2021 Classification of hand-wrist maturity level based on similarity matching
abstract
Abstract Judging the maturity level of each hand‐wrist reference bone is the core issue in bone age assessment. Relying on the superiority of convolutional neural networks in feature representation, deep learning is widely studied for the automatic bone age assessment. However, an efficient but complex deep learning network requests a large dataset with bone‐maturity‐level labels for training, restricting its large‐scale application in bone maturity classification. For this reason, we transform the bone‐maturity‐level classification problem into the similarity matching problem. Also, we propose a general structure based on Siamese network by merging two inputs into a two‐channel input and introducing a dual attention mechanism, to create an Attentional Two‐Channel Network (ATC‐Net). This paper takes the intermediate phalanges III as an example to assess the performance of the similarity matching method and the ATC‐Net. Experiments show that our method can perform better on small datasets, which effectively makes up for the data shortage problem. The ATC‐Net used for classification significantly reduces the evaluation time compared with other classical networks. It reduces the time of assessing one sample by about 49% as compared to VGG‐16. And more importantly, it achieves the highest classification accuracy of 92.74% among all investigated networks.
Keji Mao, Minhao Wang, Ruiji Xu
IET Image Process.1
2020 Energy provision minimisation in large-scale wireless powered communication networks with throughput demand
abstract
So far, the research of wireless powered communication networks (WPCNs) mainly considers the scenarios with a single radio‐frequency (RF) energy transmitter (ET) and a single sink. However, in practice, there are many applications where multiple ETs and sinks need to be deployed. This study focuses on large‐scale WPCNs having multiple RF ETs and sinks. Specifically, the authors aim to minimise the total energy provision by optimising ETs' transmit powers with the node‐throughput demand and sum‐throughput demand, respectively. For the node‐throughput demand case, they firstly formulate it to be a convex optimisation problem, then transform it to be a linear programming (LP) problem, and finally present a distributed algorithm to obtain the optimal solution. For the sum‐throughput demand case, they firstly formulate it to be a non‐linear optimisation problem, then prove its convexity and finally propose an efficient dual subgradient algorithm to obtain the optimal solution. Simulation results demonstrate that compared to the sum‐throughput demand, imposing the node‐throughput demand can effectively alleviate the throughput unfairness at the cost of increased energy provision; the proposed optimal algorithms can substantially decrease the total energy provision of ETs; the energy provision reduction percentage achieved by their schemes increases as the number of ETs increases.
Haijiang Ge, Zhanwei Yu, Kaikai Chi, Keji Mao, Qike Shao
IET Commun.4
2020 SAR multi-target interactive motion recognition based on convolutional neural networks
abstract
Synthetic aperture radar (SAR) multi‐target interactive motion recognition classifies the type of interactive motion and generates descriptions of the interactive motions at the semantic level by considering the relevance of multi‐target motions. A method for SAR multi‐target interactive motion recognition is proposed, which includes moving target detection, target type recognition, interactive motion feature extraction, and multi‐target interactive motion type recognition. Wavelet thresholding denoising combined with a convolutional neural network (CNN) is proposed for target type recognition. The method performs wavelet thresholding denoising on SAR target images and then uses an eight‐layer CNN named EilNet to achieve target recognition. After target type recognition, a multi‐target interactive motion type recognition method is proposed. A motion feature matrix is constructed for recognition and a four‐layer CNN named FolNet is designed to perform interactive motion type recognition. A motion simulation dataset based on the MSTAR dataset is built, which includes four kinds of interactive motions by two moving targets. The experimental results show that the recognition performance of the authors’ Wavelet + EilNet method for target type recognition and FolNet for multi‐target interactive motion type recognition are both better than other methods. Thus, the proposed method is an effective method for SAR multi‐target interactive motion recognition.
Ruohong Huan, Luoqi Ge, Chaojie Xie, Kaikai Chi, Keji Mao
IET Image Process.6
2019 ARU-Net: Research and Application for Wrist Reference Bone Segmentation
abstract
Segmenting reference bones from radiographs of the hand is important for bone age assessment. Due to the influence of the irregular shapes and the adjacent positions of the wrist reference bones, it is difficult for the expert to accurately estimate the mature indication of the wrist reference bones in the figures. How to precisely segment the reference bones automatically from the radiographs is a challenge. For this is problem, an improved U-Net, Attention Residual U-Net (ARU-Net) proposed in this paper. Firstly, we extract the reference bone region of interest (ROI) by faster region-based convolutional neural networks(R-CNN). Then, the pre-processed ROI is fed into ARU-Net for segmentation. On the basis of traditional U-Net, ARU-Net adds residual mapping and attention mechanism, which improves the utilization rate of features and the accuracy of reference bone segmentation. Finally, a post-processing method including the flood fill algorithm and the morphological operation is used to eliminate jagged edges and holes in the segmented result. The hamate is one of the most difficult reference bones to segment in the wrist. This paper takes it as an example to assess the performance of ARU-Net. Experiments show that compared with Fully Convolutional Neural Network (FCN), U-Net and ResUnet, the accuracy and F1 scores of ARU-Net are higher. Its accuracy rate is 96.41%, and F1 score is 0.9529. The post-processing method can further improve the result. Finally, the accuracy rate reaches 96.51%, and the F1 score reaches 0.9544. ARU-Net can precisely segment the reference bone, which facilitates the expert to assess its mature indication, so as to accurately evaluate the bone age.
Xiannian Zhou, Minhao Wang, Jiefan Qiu, Keji Mao
WiMob6
2005 The Research on Fuzzy Data Mining Applied on Browser Records
Qingzhang Chen, Jianghong Han, Yungang Lai, Wenxiu He, Keji Mao
ADMA5