Jiuwen Cao

dblp:62/5220 · DBLP profile ↗
← Back
101ranked-venue papers
18as first author
52since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 9 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 13 since 2021Systems, architecture and hardware · 19 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 11 since 2021Computer networks · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author
YearPublicationVenuePosition
2026 Accelerating Audio-driven 3D Facial Animation Training with a Compressive Sensing Framework
Hongzhen Chen, Lei Sun 0006, Jiuwen Cao, Zhiping Lin 0001
ISCAS3
2026 Discrete high-gain observer based epileptic seizure prediction by single-lead ECG
Dinghan Hu, Tiejia Jiang, Jiuwen Cao
ISCAS4
2026 Multiple EEG channel attention based multi-task model for epileptiform activity quantification
Dinghan Hu, Zirun Jiang, Chenzhi Jin, Zuonian Xie, Feng Gao 0018, Jiuwen Cao
Neurocomputing6
2026 Developing a knowledge-guided federated graph attention learning network with a diffusion module to diagnose Alzheimer's disease
Xuegang Song, Kaixiang Shu, Peng Yang 0011, Cheng Zhao 0003, Feng Zhou 0003, Alejandro F. Frangi, Jiuwen Cao, Xiaohua Xiao, Shuqiang Wang, Tianfu Wang 0001, Bai Ying Lei
Medical Image Anal.7
2026 Edge feature enhancement: Generating adversarial edge perturbations for perterm infant movement recognition
Xianfu Bao, Jiuwen Cao
Neural Networks4
2026 Exploring non-local spatial-angular correlations with a hybrid Mamba-Transformer framework for light field super-resolution
Haosong Liu, Xiancheng Zhu, Huanqiang Zeng, Jianqing Zhu, Jiuwen Cao, Junhui Hou
Pattern Recognit.5
2026 CL4CEA: A Clinical-Knowledge-Informed Augmentation for Contrastive Learning on Childhood Epilepsy Analysis
abstract
Although existing contrastive learning models utilizing conventional data augmentation achieve modest performance in EEG-based epilepsy analysis, such models risk disrupting EEG semantic consistency during pre-training and compromise the robustness of downstream tasks. To this end, we propose a clinical-knowledge-informed augmentation for contrastive learning on childhood epilepsy analysis (CL4CEA) in this paper. Our method comprises two core components: firstly, inspired by clinical knowledge, we introduce a montage conversion-based plug-and-play augmentation strategy for preserving EEG semantic consistency in contrastive learning. Secondly, a channel-aware adaptive fusion block is employed to integrate features from both temporal and frequency domains, enabling the model to capture discriminative representations of childhood EEG signals, and thus enhancing performance in downstream tasks. Through pre-training on an EEG dataset with more than 1,000 hours childhood EEG recording, and performance fine-tuning, the developed CL4CEA model can achieve promising performance on 3 downstream tasks from 3 medical centers in childhood epilepsy analysis, including onset detection, seizure type classification, and hypoxic-ischaemic encephalopathy (HIE) grading. Comparative experiments with state-of-the-art methods and systematic ablation studies demonstrate the superiority of our proposed model.
Yuanmeng Feng, Dinghan Hu, Tiejia Jiang, Jiuwen Cao
IEEE Signal Process. Lett.5
2026 Summarize Before Glimpse: Brain-Inspired Non-Autoregressive Scene Text Recognizer
abstract
The language modeling paradigm for scene text recognition (STR) has demonstrated impressive universal capabilities across extensive STR scenarios. However, existing methods still encounter challenges in effectively handling text images with irregular shapes and diverse appearances (e.g., curve, artistic, multi-oriented) due to the absence of contextual information during initial decoding. In this work, inspired by the principle of ‘forest before trees’ in human visual perception, we introduce NASTR, a non-autoregressive scene text recognizer capable of endowing global-aware for the attentional decoder. Specifically, we design a global-to-local attention procedure, simulating the mechanism of globally holistic visual signal processing preceding locally detailed response in the human brain visual system. This is achieved by leveraging the global image information queries to condition the generation of glimpse vectors at each decoding time step. This procedure empowers the NASTR model to achieve on-par performance with its state-of-the-art autoregressive counterparts, while operating in a fully parallel manner. Moreover, we propose multiple optional and flexible encoding constraint components to alleviate the representation quality degradation issue caused by the global image information queries in handling STR tasks with multilingual and in multi-domains. These components constrain the global image features from the perspective of global structural, global semantic, and linguistic knowledge. Extensive experimental results demonstrate that NASTR consistently outperforms existing methods on both Chinese and English STR benchmarks. Our source code, trained models, and logs are available at https://github.com/ML-HDU/NASTR.
Tianlei Wang, Zhiping Lin 0001, Jiuwen Cao
IEEE Trans. Circuits Syst. Video Technol.4
2026 Sleep Staging Algorithm Incorporating Multi-Threhold Neighborhood Polar Pattern Statistics
abstract
Sleep plays a vital role in human life, and its quality has a direct impact on overall health. Sleep staging is a crucial process and a key indicator used to evaluate sleep quality. This paper proposes a sleep staging method based on statistical mode of multi-threshold neighborhood extreme (SMNE). We employ the discrete wavelet transform (DWT) in conjunction with a data enhancement algorithm to preproces the EEG signals, improving their quality through a combined approach based on signal-to-noise ratio(SNR) assessment and signal overlap analysis. Then, the extremes of EEG signal are classified into 5 distinct states. Multi-threshold is applied to differences in 5-state extreme value matrix to define and extract patterns. These patterns are then statistically encoded. The extracted codes, representing the various patterns, are input into 5 weighting layers. Then, the extracted codes are fed into grey wolf optimization (GWO) to denote threshold for SNR and SMNE feature. Finally, the features derived from this process are input into a Random Forest (RF) classifier. The SleepEDFx dataset shows accuracy of 94.6%, kappa coefficient of 0.92 and F1-score of 89.3%. For the SleepEDF-20 dataset, the corresponding metrics are accuracy of 96.3%, kappa of 0.94 and F1-score of 94.2%. Meanwhile, the ISRUC-Sleep dataset achieves accuracy of 88.5%, kappa value of 0.83 and F1-score of 86.5%.
Qinqin Liu, Duanpo Wu, Guangsheng Wu, Pierre-Paul Vidal, Jiuwen Cao, Danping Wang
IEEE J. Biomed. Health Informatics6
2025 Preterm infant limb movement recognition with graph and convolution fusion network
Xianfu Bao, Jiuwen Cao
Pattern Recognit.5
2025 Mutual Information Driven Representation Learning for Cross-Subject Seizure Detection
abstract
Developing a generalizable model across subjects is crucial for the practical application of Electroencephalogram (EEG) based seizure detection model. However, inter-subject variability poses a challenge to the accurate identification of epileptic EEG, and applications often require recalibration and training of the base model using individual labeled data. To overcome this limitation, we propose a cross-subject transfer learning algorithm based on mutual information decomposition driven representation learning (MIDRL). The algorithm first introduces the structured state space sequence model to capture the long-term dependencies of epileptic EEG, and the residual module is used to mine the deep information between channels. Additionally, the mutual information estimation is employed to decompose the middle layer features of the network into domain-invariant representations and domain-specific representations, with the dynamically learnable weight updating mechanism to adaptively balance the learning tasks associated with the two representations. Finally, to address the problem of target samples being easily confused near the classification boundary, the minimum class confusion loss is introduced to reduce the class correlation predicted by the classifier. Experimental results demonstrate that the proposed algorithm effectively retains patterns of seizure region and exhibits strong performance for cross-subject seizure detection.
Dinghan Hu, Xiaonan Cui, Tiejia Jiang, Jiuwen Cao
IEEE Signal Process. Lett.5
2025 Fuzzy Multivariate Variational Mode Decomposition With Applications in EEG Analysis
abstract
This article introduces a novel extension of the multivariate variational mode decomposition (MVMD) method, termed fuzzy MVMD (FMVMD), designed to enhance alignment information extraction. In contrast to MVMD, FMVMD focuses on capturing finer alignment details by leveraging fuzzy clustering techniques. The proposed FMVMD algorithm proceeds through the following steps: First, FMVMD employs a modified clustering algorithm, termed fuzzy C-means (FCM), to categorize submodes within each channel into fuzzy clusters based on their contribution to common center frequencies. Second, a variational optimization model is formulated, extending the principles of MVMD to accommodate the fuzzy clustering approach used in FMVMD. Finally, an optimization technique called the alternating direction method of multipliers is employed to derive the optimal solution for the FMVMD model. Experimental results show that FMVMD achieves a 41% and 28% improvement in center frequency alignment performance compared to MVMD when using two and three fuzzy clusters, respectively, and a 13% improvement compared to GMVMD with the same number of clusters. Under a 25 dB SNR condition, FMVMD demonstrates a noise resistance improvement of 44% and 24% compared to MVMD with two and three fuzzy clusters, respectively, and a 37% improvement compared to GMVMD. Validation using EEG data in the forms of bipolar leads and common average reference confirms the effectiveness of FMVMD, achieving consistently favorable results.
Hongkai Tang, Xun Yang 0001, Yixuan Yuan, Pierre-Paul Vidal, Danping Wang, Jiuwen Cao, Duanpo Wu
IEEE Trans. Fuzzy Syst.6
2025 Combining Autoregressive and Non-Autoregressive Models for Ship License Plate Recognition
abstract
Ship license plate recognition (SLPR) is a fundamental visual task in intelligent waterway transportation systems that aims to transcribe ship name text images into editable text strings. Previous works construct recognizers with complicated training procedures and additional corpora to boost performance, limiting their efficiency in practical application scenarios. This paper proposes a novel ship license plate recognizer by combining autoregressive (AR) and non-autoregressive (NAR) decoding mechanisms. By adaptively incorporating the dual-branch character representations, our method explicitly sidesteps the dependence on an external language model and is adapted to the weak semantic correlation characteristics of ship name text images. Furthermore, we introduce a confidence-based dynamic reweighting strategy that includes character-and instance-level granularities. This encourages the model to learn more from hard or challenging samples. Experimental results conducted on two SLPR benchmarks demonstrate the effectiveness of our method, showing that it is competitive and outperforms previous approaches. Our study comprehensively explores the relative importance of linguistic and visual cues in SLPR and demonstrates the advantages of the dynamically fused model with dual-branch autoregressive and non-autoregressive over single-branch recognition.
Tianlei Wang, Jiuwen Cao
IEEE Trans. Intell. Transp. Syst.3
2025 Multiple artifacts detection based on channel masking and multi-feature domain semi-supervised network
Dinghan Hu, Feng Gao 0018, Xiaohui Lou, Zuonian Xie, Jiuwen Cao
J. Supercomput.6
2025 EAViz: a user-friendly deep learning-based epilepsy analysis visualizer using multimodal data
Ze Xia, Dinghan Hu, Tiejia Jiang, Shuangpeng Zhu, Xiaohui Lou, Jiuwen Cao
J. Supercomput.6
2025 On the Behavior of Contrastive Regularization in Improving Chinese Text Recognizer
abstract
The dense representation space in Chinese scene text recognition (STR) makes discriminating between categories highly challenging, because of the large candidate category set. Mainstream STR methods have achieved remarkable advancements by leveraging linguistic knowledge to implicitly address this challenge. In this paper, inspired by the correlation between recognizer performance and the distributional properties of character representations, as well as the inherent consistency between this correlation and supervised contrastive learning (SupCon), we thoroughly investigate how to integrate SupCon with an STR model to alleviate this challenge, and elucidate some dynamic behaviors underlying the performance improvements. Specifically, we analyze the SupCon-STR models instantiated with different projectors and evaluate their distributional properties through metrics, including intra-class compactness, inter-class separability, and feature redundancy, while assessing performances that involve in-domain accuracy and cross-domain recognition generalization. The main results reveal how the temperature$\tau$and projectors affect the representation distribution, and highlight that suitable intra-class compactness and sufficient inter-class separability are key factors for delivering competitive performances in both in-domain and cross-domain STR scenarios. Moreover, these results also provide valuable insights into the design of SupCon-STR architectures for diverse resource constraints. Taking existing Chinese STR models as baselines, and combining SupCon-STR with them, the average improvements in cross-domain recognition performance are over 5% across 7 testing datasets. A new state-of-the-art accuracy of 77.19% on the ChineseScenebenchmark is also established.
Tianlei Wang, Huanqiang Zeng, Jiuwen Cao
IEEE Trans. Multim.4
2024 Audio-Visual Cross-Modal Generation with Multimodal Variational Generative Model
abstract
Audio and Visual are two important visual modalities in video content understanding. However, the absence of one modality may be observed in practical applications due to the real environmental factors, which leads to the information loss. Therefore, audio and visual fusion is focused on using the shared and complementary information between modalities to recover the missing modalities from the available data modalities. In this paper, an Adversarial Hierarchical Variational Auto-Encoder (Adv-HVAE) model is proposed to solve this problem of modality data loss. A multimodal representation is first learned using a hierarchical Variational Autoencoder (VAE) model that enables the generation of missing modal data under any subset of available modalities. Also to obtain a more robust multimodal representation, a feature generation network is utilized to approximate the latent distribution of missing modalities. Finally, the adversarial training network is shown to be effective in improving the data quality generated through the Adv-HVAE framework. Experimental results demonstrate that Adv-HVAE achieves best generation results on two benchmark datasets, avMNIST and Sub-URMP.
Zhubin Xu, Tianlei Wang, Dinghan Hu, Huanqiang Zeng, Jiuwen Cao
ISCAS6
2024 Automatic EEG-based Spike Ripples Detection with Multi-band Frequency Analysis
abstract
Spike ripples in electroencephalogram (EEG) have been considered as a more promising biomarker for epilepsy analysis than using spikes. Almost all existing spike ripples detection concentrates in the high frequency band (80-500Hz) without considering its co-occurrence spikes in low frequency band (1-70Hz). In this paper, a novel EEG-based spike ripples detection algorithm combining both low and high frequency is proposed. For the low frequency band, the energy histogram is derived by the nonlinear energy operator (NLEO). When the average energy exceeds a pre-set threshold, the average duration of the monotonically decreasing segment (ADDS) and the signal filtered by smooth nonlinear energy operator (SNEO) are further calculated. The enhanced K-means algorithm is used for candidate spikes selection. For the high frequency band, the peak distribution is generated to select high frequency oscillations (HFOs). Then, 21 significant features are extracted from HFOs and a quadratic kernel support vector machine (SVM) is trained for candidate ripples selection. If both the candidate ripple and spike are in the same frame, it is considered as a spike ripple. Finally, feature selection based on Max-Relevance and Min-Redundancy (mRMR) is studied to enhance overall performance. The proposed algorithm is compared with two related methods on EEGs of 6 subjects, which can achieve a convincing performance with an average of 91.35% precision, 93.88% recall, 92.56% F1score, and 96.62% BA, respectively.
Sihan Zhou, Dinghan Hu, Feng Gao 0018, Tiejia Jiang, Jiuwen Cao
ISCAS5
2024 A highly efficient ADMM-based algorithm for outlier-robust regression with Huber loss
Tianlei Wang, Xiaoping Lai, Jiuwen Cao
Appl. Intell.3
2024 An End-to-End Vision-Based Seizure Detection With a Guided Spatial Attention Module for Patient Detection
abstract
Video recording has been extensively studied for seizure detection and classification due to its convenience of collection. Most existing vision-based studies generally followed a two-stage scheme of first object detection and then action recognition to detect seizures for better real-world application. However, all of these approaches are two-stage not end-to-end, which may make the model locally optimal. Besides, the object detection algorithms applied in existing methods often suffer heavy computational burden, leading to slow inference speed and high hardware support. All these issues can seriously hinder the practical application and deployment of the model. Therefore, we proposed a novel end-to-end model in this paper, which could simultaneously achieve patient detection and seizure detection. The amount of parameters and computations in the conventional object detection branch can be reduced by innovatively exploring the idea of using a spatial attention module instead of object detection networks for patient detection. However, based on a toy example, we found that relying solely on a spatial attention module without guidance is not reliable, despite its high performance in seizure detection. Therefore, a guided spatial attention module (GSAM) is proposed in this paper. An extra regression loss function is used for guiding the learning of GSAM. In addition, the hard shrinkage operation is applied on the generated spatial attention heatmap (SAH), making the generated SAH closer to the real object detection with a faster model convergence. Besides, a temporal attention module is used to reduce the amount of parameters and computations, as well as to fuse the temporal information well. Experiments show that our method has less parameters and faster running speed than competing methods, yet better performance on seizure detection. The proposed GSAM with high performance could well replace the object detection algorithm for patient detection.
Dinghan Hu, Jiuwen Cao, Tiejia Jiang, Feng Gao 0018
IEEE Internet Things J.3
2024 A Novel Seizure Detection Method Based on the Feature Fusion of Multimodal Physiological Signals
abstract
Seizure detection is traditionally done using video/electroencephalography monitoring, but for out-of-hospital patients, this method is costly. In recent years, portable device to detect seizures gains attention. In this paper, multimodal signals collected by portable devices are studied, and a seizure detection algorithm is proposed based on adaptive multi-bit local differential ternary pattern (MLDTP). This algorithm is used for detecting seizure period and inter-seizure period. Traditional local binary pattern has certain limitations in describing one-dimensional time series signals. It can only describe two types of structures in signals: Rising structure and falling structure, making the signal patterns overly monotonous and not conducive to classification tasks. To address this issue, this paper introduces two additional structures, slowly rising structure and slowly falling structure, into the signal description using MLDTP method. This method constructs multi-bit neighboring relationships of the signals, and adaptively selects the optimal MLDTP parameters for different modalities using the Archimedes optimization algorithm (AOA). Additionally, this paper extensively discusses a multimodal signal fusion strategy, mapping features of different modal signals to the same feature space through the MLDTP algorithm to achieve information complementarity. Long-term recorded data from 18 patients were collected using the wearable device Biovital P1, with 13 cases from the Children’s Hospital affiliated with Children’s Hospital, Zhejiang University School of Medicine, and 5 cases from the fourth Affiliated Hospital of Anhui Medical University. The dataset underwent five-fold cross-validation, resulting in average accuracy, precision, sensitivity and F1 score of 96.81%, 98.55%, 95.24% and 96.87%, respectively.
Duanpo Wu, Pierre-Paul Vidal, Danping Wang, Yixuan Yuan, Jiuwen Cao, Tiejia Jiang
IEEE Internet Things J.6
2024 Global and multi-partition local network analysis of scalp EEG in West syndrome before and after treatment
Lishan Liu, Duanpo Wu, Yixuan Yuan, Danping Wang, Tiejia Jiang, Jiuwen Cao, Yuansheng Xu
Neural Networks8
2024 M-DDC: MRI based demyelinative diseases classification with U-Net segmentation and convolutional network
Deyang Zhou, Tianlei Wang, Shaonong Wei, Feng Gao 0018, Xiaoping Lai, Jiuwen Cao
Neural Networks7
2024 Matrix randomized autoencoder
Tianlei Wang, Jiuwen Cao, Wandong Zhang, Badong Chen
Pattern Recognit.3
2024 Auxiliary Label Classification Based Multi-Label Limb Movement Recognition of Preterm Infant
abstract
Limb movement recognition of preterm infants (PI-LMR) in neonatal intensive care units (NICUs) is important for infant health monitoring. However, little attention has been paid to intelligent PI-LMR. Due to the weak correlation among limb movements of preterm infants, the various limb movement combinations and imbalanced data distributions are the main challenges of PI-LMR. To address these issues, a novel multi-label limb movement recognition (MLLMR) algorithm with a dual-branch structure and multi-label fusion loss is proposed. The various movement combinations can be decomposed into limbs thanks to multi-label learning. Particularly, the multi-label fusion loss consisting of the binary cross entropy (BCE) and the pairwise ranking loss (PRL) is proposed to optimize the probabilities to the ground truth labels and the ranking between positive and negative labels, simultaneously. The weighted fusion loss is further developed to address the imbalanced label distributions. Subsequently, an auxiliary task for the classification of zero-, single- and multi-limb movements is constructed to constrain the feature space of primary task for better multi-label learning. Experiments on real clinical preterm infants video dataset from Jiaxing Maternity and Child Health Care Hospital are conducted and the results demonstrate the effectiveness of the proposed algorithm.
Hongliang Lei, Tianlei Wang, Xianfu Bao, Jiuwen Cao
IEEE Trans. Circuits Syst. Video Technol.5
2024 Fully Convolutional Network-Based Fast UAV Detection in Pulse Doppler Radar
abstract
With the popularity of drones, how to conduct effective and fast detection of unmanned aerial vehicle (UAV) to prevent unauthorized flying becomes a hot topic. Based on statistical theory, traditional constant false alarm rate (CFAR) works well on data with uniform background. But for low-slow-small UAV, it is prone to miss detection. In recent years, data-driven deep learning method is proved to have better performance than CFAR. However, the use of sliding window to convert complex detection task into simple classification task leads to low efficiency. In this paper, we propose a fast detection method that applies a fully convolutional network on the whole range-Doppler map. To achieve comparable accuracy to our previous work, the network is firstly designed on the principle that the effective receptive field of unit in the feature map for prediction is close to the size of the sliding window. And the best bifurcation position of classification and regression is searched. Then, considering the imbalance of positive and negative samples, a new scheme to create GT data is designed to expand the positive samples, and random sampling of negative samples is adopted further. Lastly, a post processing mechanism combining probability thresholding and minimum deviation positioning is developed for accurate location of target. Comparison with existing methods on the experimental data shows that the proposed method can increase the detection speed by up to 47 times while maintain a promising accuracy.
Jiangmin Tian, Jiuwen Cao
IEEE Trans. Geosci. Remote. Sens.3
2024 M$^{3}$ANet: Multi-Modal and Multi-Attention Fusion Network for Ship License Plate Recognition
abstract
Shiplicense plate recognition (SLPR) plays an important role in intelligent waterway management, but few attention has been paid to SLPR in scene text recognition (STR) community. Inspired by various outstanding achievements on STR, combined the intrinsic properties of SLPR, we propose aMulti-Modal andMulti-Attention dynamic fusion network (M$^{3}$ANet) for SLPR in this article. Specifically, the visual-language joint modeling for SLPR is developed and the channel-spatial-self attention dynamic fusion mechanism is proposed for accuracy boosting. Explicitly fusing linguistic information extracted from ship name related corpus improves the adaptability of the recognition model to occlusion, background confusion, blur, etc., which is integrated with vision features to establish a multi-modal recognition network. Gated fully fusion is utilized to fuse visual features re-weighted by multi-attention components, inducing flexible compatibility with multiple types of decoders and more refined recognition decoder inputs. Additionally, to comprehensively mine spatially salient text regions in ship license plate images, we investigate the grouped spatial attention. Extensive experiments empirically demonstrate the effectiveness of M$^{3}$ANet and superior performance (93.80% with regular images, while 90.34% with irregular images) on two benchmarks.
Tianlei Wang, Jiangmin Tian, Jiuwen Cao
IEEE Trans. Multim.5
2023 Advanced License Plate Detector in Low-Quality Images with Smooth Regression Constraint
Jiefu Yu, Tianlei Wang, Jiangmin Tian, Fangyong Xu, Jiuwen Cao
PRCV (2)6
2023 LAC-GAN: Lesion attention conditional GAN for Ultra-widefield image synthesis
Haijun Lei, Zhihui Tian, Hai Xie, Benjian Zhao, Xianlu Zeng, Jiuwen Cao, Weixin Liu 0002, Shuqiang Wang, Bai Ying Lei
Neural Networks6
2023 Multichannel Matrix Randomized Autoencoder
Tianlei Wang, Jiuwen Cao
Neural Process. Lett.3
2023 A Transformer-Based End-to-End Automatic Speech Recognition Algorithm
abstract
End-to-End (E2E) automatic speech recognition (ASR) becomes popular recent years and has been widely used in many applications. However, current ASR algorithms are usually less effective when applied in specific applications with terminologies such as medical and economic fields. To address this issue, we propose a powerful Transformer based ASR decoding method for beam searching, called soft beam pruning algorithm (SBPA). SBPA can dynamically adjust the width of beam search. Meanwhile, a prefix module (PM) is added to access the contextual information and avoid removing professional words in the beam search. Combining SBPA and PM, the proposed ASR can achieve promising recognition performance on professional terminologies. To verify the effectiveness, experiments are conducted on real-world conversation data with medical terminology. It is shown that the proposed ASR achieved significant performance on both professional and regular words.
Fang Dong 0003, Yiyang Qian, Tianlei Wang, Jiuwen Cao
IEEE Signal Process. Lett.5
2023 Efficient ADMM-Based Algorithm for Regularized Minimax Approximation
abstract
Minimax approximations have found many applications but are lack of efficient solution algorithms for large-scale problems. Based on the alternating direction method of multipliers (ADMM) for convex optimization, this letter presents an efficient scalarwise algorithm for a regularized minimax approximation problem. The ADMM-based algorithm is then applied in the minimax design of two-dimensional (2-D) digital filters and the training of randomized neural networks for regression on a realworld benchmark dataset. Experimental results demonstrate the fast convergence rate and low computational complexity of the proposed algorithm, as well as the good approximation/prediction performance of the learned approximation model.
Xuanyue Shentu, Xiaoping Lai, Tianlei Wang, Jiuwen Cao
IEEE Signal Process. Lett.4
2023 Ship License Plate Super-Resolution in the Wild
abstract
Ship license plate (SLP) recognition plays an important role in ship supervision and harbour management. Practically, low-resolution (LR) SLP images are illegible and challenging to SLP recognition. Most existing super-resolution (SR) methods are not suitable for real-world LR SLP images with over smoothed reconstructions. To alleviate these deficiencies, in this letter, we propose a parallel enhanced SR generative adversarial network (PESRGAN) for SLP images. A novel degradation model is developed to construct a more feasible LR dataset. A parallel SR convolutional neural network (SRCNN) module based on ESRGAN is proposed for feature extraction. To characterize the difference between text foreground and background, a new gradient loss is developed in PESRGAN to sharpen the character boundary. Comparisons to many state-of-the-art (SOTA) SR methods are presented to show the effectiveness of the proposed algorithm.
Huahua Wu, Jiagui Chen, Tianlei Wang, Xiaoping Lai, Jiuwen Cao
IEEE Signal Process. Lett.5
2023 Multicenter and Multichannel Pooling GCN for Early AD Diagnosis Based on Dual-Modality Fused Brain Network
abstract
For significant memory concern (SMC) and mild cognitive impairment (MCI), their classification performance is limited by confounding features, diverse imaging protocols, and limited sample size. To address the above limitations, we introduce a dual-modality fused brain connectivity network combining resting-state functional magnetic resonance imaging (fMRI) and diffusion tensor imaging (DTI), and propose three mechanisms in the current graph convolutional network (GCN) to improve classifier performance. First, we introduce a DTI-strength penalty term for constructing functional connectivity networks. Stronger structural connectivity and bigger structural strength diversity between groups provide a higher opportunity for retaining connectivity information. Second, a multi-center attention graph with each node representing a subject is proposed to consider the influence of data source, gender, acquisition equipment, and disease status of those training samples in GCN. The attention mechanism captures their different impacts on edge weights. Third, we propose a multi-channel mechanism to improve filter performance, assigning different filters to features based on feature statistics. Applying those nodes with low-quality features to perform convolution would also deteriorate filter performance. Therefore, we further propose a pooling mechanism, which introduces the disease status information of those training samples to evaluate the quality of nodes. Finally, we obtain the final classification results by inputting the multi-center attention graph into the multi-channel pooling GCN. The proposed method is tested on three datasets (i.e., an ADNI 2 dataset, an ADNI 3 dataset, and an in-house dataset). Experimental results indicate that the proposed method is effective and superior to other related algorithms, with a mean classification accuracy of 93.05% in our binary classification tasks. Our code is available at: https://github.com/Xuegang-S.
Xuegang Song, Feng Zhou 0003, Alejandro F. Frangi, Jiuwen Cao, Xiaohua Xiao, Tianfu Wang 0001, Bai Ying Lei
IEEE Trans. Medical Imaging4
2023 DDistill-SR: Reparameterized Dynamic Distillation Network for Lightweight Image Super-Resolution
abstract
Recent research on deep convolutional neural networks (CNNs) has provided a significant performance boost on efficient super-resolution (SR) tasks by trading off the performance and applicability. However, most existing methods focus on subtracting feature processing consumption to reduce the parameters and calculations without refining the immediate features, which leads to inadequate information in the restoration. In this paper, we propose a lightweight network termed DDistill-SR, which significantly improves the SR quality by capturing and reusing more helpful information in a static-dynamic feature distillation manner. Specifically, we propose a plug-in reparameterized dynamic unit (RDU) to promote the performance and inference cost trade-off. During the training phase, the RDU learns to linearly combine multiple reparameterizable blocks by analyzing varied input statistics to enhance layer-level representation. In the inference phase, the RDU is equally converted to simple dynamic convolutions that explicitly capture robust dynamic and static feature maps. Then, the information distillation block is constructed by several RDUs to enforce hierarchical refinement and selective fusion of spatial context information. Furthermore, we propose a dynamic distillation fusion (DDF) module to enable dynamic signals aggregation and communication between hierarchical modules to further improve performance. Empirical results show that our DDistill-SR outperforms the baselines and achieves state-of-the-art results on most super-resolution domains with much fewer parameters and less computational overhead.
Yan Wang 0086, Tongtong Su, Yusen Li, Jiuwen Cao, Gang Wang 0001, Xiaoguang Liu 0001
IEEE Trans. Multim.4
2023 An Accelerated Maximally Split ADMM for a Class of Generalized Ridge Regression
abstract
Ridge regression (RR) has been commonly used in machine learning, but is facing computational challenges in big data applications. To meet the challenges, this article develops a highly parallel new algorithm, i.e., an accelerated maximally split alternating direction method of multipliers (A-MS-ADMM), for a class of generalized RR (GRR) that allows different regularization factors for different regression coefficients. Linear convergence of the new algorithm along with its convergence ratio is established. Optimal parameters of the algorithm for the GRR with a particular set of regularization factors are derived, and a selection scheme of the algorithm parameters for the GRR with general regularization factors is also discussed. The new algorithm is then applied in the training of single-layer feedforward neural networks. Experimental results on performance validation on real-world benchmark datasets for regression and classification and comparisons with existing methods demonstrate the fast convergence, low computational complexity, and high parallelism of the new algorithm.
Xiaoping Lai, Jiuwen Cao, Zhiping Lin 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 Clustering-Guided Pairwise Metric Triplet Loss for Person Reidentification
abstract
Most of the loss functions proposed for person reidentification (Re-ID) are expected to be easy to deploy, efficiently improve network performance, and will not introduce redundant parameters. This study proposes a no-parameter and generic clustering-guided pairwise metric triplet (CPM-Triplet) loss based on the hard sample mining triplet loss for the metric learning loss. CPM-Triplet loss deploys two metrics: 1) the Euclidean metric and 2) the cosine metric, to complementarily improve the metric learning of the model. Paralleled to the Euclidean metric, the cosine metric quantifies the sample similarity in a different way to the Euclidean metric, which takes a different perspective to explore the distribution of samples. But the pairwise metric mainly improves the precision between dissimilar samples of the same label and could not solve the problem of excessive outliers. Therefore, the clustering-guided correction term was deployed to apply to all samples with the same label to mine the similarity in the samples, while weakening the influence of outliers in CPM-Triplet loss. Experiments conducted on four benchmark data sets show that the combination of the CPM-Triplet loss and the widely used Bag-of-Tricks baseline generally outperforms the baseline and numerous state-of-the-art methods studied in this article. The source code would be available athttps://github.com/weiyu-zeng/CPM-Triplet-loss.
Weiyu Zeng, Tianlei Wang, Jiuwen Cao, Huanqiang Zeng
IEEE Internet Things J.3
2022 Vibration-based hypervelocity impact identification and localization
abstract
Hypervelocity impact (HVI) vibration source identification and localization have found wide applications in many fields, such as manned spacecraft protection and machine tool collision damage detection and localization. In this paper, we study the synchrosqueezed transform (SST) algorithm and the texture color distribution (TCD) based HVI source identification and localization using impact images. The extracted SST and TCD image features are fused for HVI image representation. To achieve more accurate detection and localization, the optimal selective stitching features OS SST+TCD are obtained by correlating and evaluating the similarity between the sample label and each dimension of the features. Popular conventional classification and regression models are merged by voting and stacking to achieve the final detection and localization. To demonstrate the effectiveness of the proposed algorithm, the HVI data recorded from three kinds of high-speed bullet striking on an aluminum alloy plate is used for experimentation. The experimental results show that the proposed HVI identification and localization algorithm is more accurate than other algorithms. Finally, based on sensor distribution, an accurate four-circle centroid localization algorithm is developed for HVI source coordinate localization.
Jiao Bao, Lifu Liu, Jiuwen Cao
Frontiers Inf. Technol. Electron. Eng.3
2022 3D residual-attention-deep-network-based childhood epilepsy syndrome classification
Yuanmeng Feng, Xiaonan Cui, Tianlei Wang, Tiejia Jiang, Feng Gao 0018, Jiuwen Cao
Knowl. Based Syst.7
2022 Longitudinal study of early mild cognitive impairment via similarity-constrained group learning and self-attention based SBi-LSTM
Bai Ying Lei, Yanwu Xu 0001, Guanghui Yue 0001, Jiuwen Cao, Huoyou Hu, Shuangzhi Yu, Peng Yang 0011, Tianfu Wang 0001, Yali Qiu, Xiaohua Xiao, Shuqiang Wang
Knowl. Based Syst.6
2022 Deep feature fusion based childhood epilepsy syndrome classification from electroencephalogram
Xiaonan Cui, Dinghan Hu, Jiuwen Cao, Xiaoping Lai, Tianlei Wang, Tiejia Jiang, Feng Gao 0018
Neural Networks4
2022 Scalp EEG functional connection and brain network in infants with West syndrome
Yuanmeng Feng, Tianlei Wang, Jiuwen Cao, Duanpo Wu, Tiejia Jiang, Feng Gao 0018
Neural Networks4
2022 Deep Learning-Based UAV Detection in Pulse-Doppler Radar
abstract
With the popularity of unmanned aerial vehicles (UAVs), how to conduct automatic and effective detection to prevent unauthorized flying has become an important issue. The conventional constant false alarm rate (CFAR) detector based on radar signal has shown advantages in moving target detection. However, the CFAR-based detectors are strongly dependent on some manual experience, such as the ambient noise distribution estimation and the detection windows’ size selection, and usually suffered poor performance on small UAV detection due to the weak signal. Inspired by the success of deep learning (DL) on natural scene object detection, this article tries to explore a DL-based method for UAV detection in pulse-Doppler radar. Concretely, we propose a convolutional neural network (CNN) with two heads: one for the classification of the input range-Doppler map patch into target present or target absent and the other for the regression of offset between the target and the patch center. Then, based on the output of the network, a nonmaximum suppression (NMS) mechanism composed of probability-based initial recognition, distribution density-based recognition, and voting-based regression is developed to reduce false alarms as well as control the false alarms. Finally, experiments on both simulated data and real data are carried out, and it is shown that the proposed method can locate the target more accurately and achieve a much lower false alarm rate at a comparable detection rate than CFAR.
Jiangmin Tian, Jiuwen Cao
IEEE Trans. Geosci. Remote. Sens.3
2022 Screen Content Video Quality Assessment Model Using Hybrid Spatiotemporal Features
abstract
In this paper, a full-reference video quality assessment (VQA) model is designed for the perceptual quality assessment of the screen content videos (SCVs), called the hybrid spatiotemporal feature-based model (HSFM). The SCVs are of hybrid structure including screen and natural scenes, which are perceived by the human visual system (HVS) with different visual effects. With this consideration, the three dimensional Laplacian of Gaussian (3D-LOG) filter and three dimensional Natural Scene Statistics (3D-NSS) are exploited to extract the screen and natural spatiotemporal features, based on the reference and distorted SCV sequences separately. The similarities of these extracted features are then computed independently, followed by generating the distorted screen and natural quality scores for screen and natural scenes. After that, an adaptive screen and natural quality fusion scheme through the local video activity is developed to combine them for arriving at the final VQA score of the distorted SCV under evaluation. The experimental results on the Screen Content Video Database (SCVD) and Compressed Screen Content Video Quality (CSCVQ) databases have shown that the proposed HSFM is more in line with the perceptual quality assessment of the SCVs perceived by the HVS, compared with a variety of classic and latest IQA/VQA models.
Huanqiang Zeng, Hailiang Huang 0002, Junhui Hou, Jiuwen Cao, Yongtao Wang, Kai-Kuang Ma
IEEE Trans. Image Process.4
2022 SLPR: A Deep Learning Based Chinese Ship License Plate Recognition Framework
abstract
Automatic ship license plate recognition (SLPR) for ship identification is of great significance to waterway shipping management. But few attention has been paid to SLPR in the past. In this paper, a novel cascaded Chinese SLPR framework consisting of the quadrangle-based ship license plate detection (QSLPD) algorithm and the rectification-based text recognition network (RTRNet) is developed. Concretely, in QSLPD algorithm, detection is performed based on the pyramid feature fusion architecture ameliorated by the proposed variable receptive field feature enhancement strategy and three task-specific output heads. In addition, a new loss function combining the dice coefficient and cross entropy is explored in the proposed SLPR which can generate significant improvement over the baseline. In RTRNet, regions of interest (RoIs) extraction and irregular text line rectification based on the vertices information predicted by QSLPD are performed before text recognition. Data augmentation are also applied to cope with the problem of limited text recognizer training data and the extremely imbalance distribution of corpus. Extensive experiments are carried out to demonstrate the reliability of the proposed cascaded SLPR framework, that can achieve the highest F-measure of 87.78% and 76.59% with IoU and TIoU metric on the collected dataset, surpasses many existing advanced methods.
Jiuwen Cao, Tianlei Wang, Huahua Wu, Jiangmin Tian, Fangyong Xu
IEEE Trans. Intell. Transp. Syst.2
2021 Cascading Scene and Viewpoint Feature Learning for Pedestrian Gender Recognition
abstract
Pedestrian gender recognition plays an important role in smart city. To effectively improve the pedestrian gender recognition performance, a new method, called cascading scene and viewpoint feature learning (CSVFL), is proposed in this article. The novelty of the proposed CSVFL lies on the joint consideration of two crucial challenges in pedestrian gender recognition, namely, scene and viewpoint variation. For that, the proposed CSVFL starts with the scene transfer (ST) scheme, followed by the viewpoint adaptation (VA) scheme in a cascading manner. Specifically, the ST scheme exploits the key pedestrian segmentation network to extract the key pedestrian masks for the subsequent key pedestrian transfer generative adversarial network, with the goal of encouraging the input pedestrian image to have the similar style to the target scene while preserving the image details of the key pedestrian as much as possible. Afterward, the obtained scene-transferred pedestrian images are fed to train the deep feature learning network with the VA scheme, in which each neuron will be enabled/disabled for different viewpoints depending on whether it has contribution on the corresponding viewpoint. Extensive experiments conducted on the commonly used pedestrian attribute data sets have demonstrated that the proposed CSVFL approach outperforms multiple recently reported pedestrian gender recognition methods.
Huanqiang Zeng, Jianqing Zhu, Jiuwen Cao, Yongtao Wang, Kai-Kuang Ma
IEEE Internet Things J.4
2021 Graph convolution network with similarity awareness and adaptive calibration for disease-induced deterioration prediction
Xuegang Song, Feng Zhou 0003, Alejandro F. Frangi, Jiuwen Cao, Xiaohua Xiao, Tianfu Wang 0001, Bai Ying Lei
Medical Image Anal.4
2021 Cross-attention multi-branch network for fundus diseases classification using SLO images
Hai Xie, Xianlu Zeng, Haijun Lei, Jie Du 0001, Jiuwen Cao, Tianfu Wang 0001, Bai Ying Lei
Medical Image Anal.7
2021 Robust High-Order Manifold Constrained Low Rank Representation for Subspace Clustering
abstract
Due to the effectiveness in learning the subspace structures, low-rank representation (LRR) and its variations have been widely applied in various fields, such as computer vision and pattern recognition. However, in real applications, it is a challenge to handle the complex noises. To address this problem, we propose a novel robust LRR method based on kernel risk-sensitive loss (KRSL) with high-order manifold constraint, called RHLRR, in which the KRSL is introduced to deal with the noises and the multiple hypergraph regularization term is used as a high order manifold constraint to effectively capture the locality, similarity and the intrinsic geometric information in data. Besides, an iterative algorithm based on the half-quadratic (HQ) and the accelerated block coordinate update (BCU) is developed. The experimental results demonstrate that the proposed method can outperform other state-of-the-art LRR variants.
Lei Xing 0003, Badong Chen, Jianji Wang 0001, Shaoyi Du, Jiuwen Cao
IEEE Trans. Circuits Syst. Video Technol.5
2021 Unsupervised Eye Blink Artifact Detection From EEG With Gaussian Mixture Model
abstract
Eye blink is one of the most common artifacts in electroencephalogram (EEG) and significantly affects the performance of the EEG related applications, such as epilepsy recognition, spike detection, encephalitis diagnosis, etc. To achieve an accurate and efficient eye blink detection, a novel unsupervised learning algorithm based on a hybrid thresholding followed with a Gaussian mixture model (GMM) is presented in this paper. The EEG signal is priliminarily screened by a cascaded thresholding method built on the distributions of signal amplitude, amplitude displacement, as well as the cross channel correlation. Then, the channel correlation of the two frontal electrodes (FP1, FP2), the fractal dimension, and the mean of amplitude difference between FP1 and FP2, are extracted to characterize the filtered EEGs. The GMM trained on these features is applied for the eye blink detection. The performance of the proposed algorithm is studied on two EEG datasets collected by the Temple University Hospital (TUH) and the Children's Hospital, Zhejiang University School of Medicine (CHZU), where the datasets are recorded from epilepsy and encephalitis patients, and contain a lot of eye blink artifacts. Experimental results show that the proposed algorithm can achieve the highest detection precision and F1 score over the state-of-the-art methods.
Jiuwen Cao, Dinghan Hu, Fang Dong 0003, Tiejia Jiang, Weidong Gao 0006, Feng Gao 0018
IEEE J. Biomed. Health Informatics1
2021 Maximum Correntropy Criterion-Based Hierarchical One-Class Classification
abstract
Due to the effectiveness of anomaly/outlier detection, one-class algorithms have been extensively studied in the past. The representatives include the shallow-structure methods and deep networks, such as the one-class support vector machine (OC-SVM), one-class extreme learning machine (OC-ELM), deep support vector data description (Deep SVDD), and multilayer OC-ELM (ML-OCELM/MK-OCELM). However, existing algorithms are generally built on the minimum mean-square-error (mse) criterion, which is robust to the Gaussian noises but less effective in dealing with large outliers. To alleviate this deficiency, a robust maximum correntropy criterion (MCC)-based OC-ELM (MC-OCELM) is first proposed and then further extended to a hierarchical network to enhance its capability in characterizing complex and large data (named HC-OCELM). The gradient derivation combining with a fixed-point iterative updation scheme is adopted for the output weight optimization. Experiments on many benchmark data sets are conducted for effectiveness validation. Comparisons to many state-of-the-art approaches are provided for the superiority demonstration.
Jiuwen Cao, Haozhen Dai, Bai Ying Lei, Chun Yin, Huanqiang Zeng, Anton Kummert
IEEE Trans. Neural Networks Learn. Syst.1
2021 Hierarchical One-Class Classifier With Within-Class Scatter-Based Autoencoders
abstract
Autoencoding is a vital branch of representation learning in deep neural networks (DNNs). The extreme learning machine-based autoencoder (ELM-AE) has been recently developed and has gained popularity for its fast learning speed and ease of implementation. However, the ELM-AE uses random hidden node parameters without tuning, which may generate meaningless encoded features. In this brief, we first propose a within-class scatter information constraint-based AE (WSI-AE) that minimizes both the reconstruction error and the within-class scatter of the encoded features. We then build stacked WSI-AEs into a one-class classification (OCC) algorithm based on the hierarchical regularized least-squared method. The effectiveness of our approach was experimentally demonstrated in comparisons with several state-of-the-art AEs and OCC algorithms. The evaluations were performed on several benchmark data sets.
Tianlei Wang, Jiuwen Cao, Xiaoping Lai, Q. M. Jonathan Wu
IEEE Trans. Neural Networks Learn. Syst.2
2020 Affine Transformation Based Hierarchical Extreme Learning Machine
abstract
Recently, the signal hidden layer feedforward network (SLFN) based extreme learning machine (ELM) has been extended to a hierarchical learning framework (HELM). Although the HELM shows better generalization performance with lower computational complexity than many deep neural networks (DNNs), it is found that as the layer increases, the input distribution of each layer may move to the saturated regime of the non-linear activation function, which affects the generalization performance. However, few attentions have been paid to the data distribution normalization to address this issue. Thus, in this paper, an affine transformation (AT) inputs based activation function layer is introduced to normalize the data distribution and a novel AT based HELM (AT-HELM) is developed. The proposed AT-HELM can adapt the activation function inputs to the distribution of each layer and obtains better generalization performance. Experiments on 29 benchmark datasets are carried out to demonstrate the superiority of AT-HELM.
Rongzhi Ma, Jiuwen Cao, Tianlei Wang, Xiaoping Lai
ISCAS2
2020 MAM: Mixed Attention Module with Random Disruption Augmentation for Image Classification
abstract
Visual attention mechanism extracts features by establishing a mask branch characterizing the distribution regions we are interested in. With the thought that eliminating useless features helps in the process of extracting significant features, in this paper we design a novel visual attention mechanism structure called the mixed attention module (MAM), which is capable of distilling significant patterns as well as suppressing unimportant features. The generic MAM constructed by the min-pooling and the tanh function is embedded into a mixed attention network for image classification. Moreover, a new image enhancement method named the random disruption (RandomDisrupt) is developed by slicing an image and then scrambling it. RandomDisrupt destroys the positional information between the features and the structural information of the object to accomplish data augmentation. The combination of the MAM and RandomDisrupt has improved the image classification performance significantly. Benchmark experiments are conducted to show the superiority of the proposed algorithm over a state-of-the-art method. Source code would be available at https://github.com/weiyu-zeng/MAM.
Weiyu Zeng, Jiuwen Cao, Xiaoping Lai, Zhiping Lin 0001
ISCAS2
2020 Integrating Similarity Awareness and Adaptive Calibration in Graph Convolution Network to Predict Disease
Xuegang Song, Alejandro F. Frangi, Xiaohua Xiao, Jiuwen Cao, Tianfu Wang 0001, Bai Ying Lei
MICCAI (7)4
2020 Multi-scale Enhanced Graph Convolutional Network for Early Mild Cognitive Impairment Detection
Shuangzhi Yu, Shuqiang Wang, Xiaohua Xiao, Jiuwen Cao, Guanghui Yue 0001, Tianfu Wang 0001, Yanwu Xu 0001, Bai Ying Lei
MICCAI (7)4
2020 Imbalanced learning algorithm based intelligent abnormal electricity consumption detection
Hongyun Qin, Houpan Zhou, Jiuwen Cao
Neurocomputing3
2020 Self-weighted adaptive structure learning for ASD diagnosis via multi-template multi-center representation
Fanglin Huang, Ee-Leng Tan, Peng Yang 0011, Le Ou-Yang, Jiuwen Cao, Tianfu Wang 0001, Bai Ying Lei
Medical Image Anal.6
2020 Self-calibrated brain network estimation and joint non-convex multi-task learning for identification of early Alzheimer's disease
Bai Ying Lei, Nina Cheng, Alejandro F. Frangi, Ee-Leng Tan, Jiuwen Cao, Peng Yang 0011, Ahmed El-Azab, Jie Du 0001, Yanwu Xu 0001, Tianfu Wang 0001
Medical Image Anal.5
2020 Joint Pyramid Feature Representation Network for Vehicle Re-identification
Xiangwei Lin, Huanqiang Zeng, Jinhui Hou, Jiuwen Cao, Jianqing Zhu, Jing Chen 0001
Mob. Networks Appl.4
2020 Heart sound classification based on improved MFCC features and convolutional recurrent neural networks
Muqing Deng, Tingting Meng, Jiuwen Cao, Shimin Wang, Jing Zhang 0037, Huijie Fan
Neural Networks3
2020 Regularized correntropy criterion based semi-supervised ELM
Jie Yang 0051, Jiuwen Cao, Tianlei Wang, Anke Xue, Badong Chen
Neural Networks2
2020 A Maximally Split and Relaxed ADMM for Regularized Extreme Learning Machines
abstract
One of the salient features of the extreme learning machine (ELM) is its fast learning speed. However, in a big data environment, the ELM still suffers from an overly heavy computational load due to the high dimensionality and the large amount of data. Using the alternating direction method of multipliers (ADMM), a convex model fitting problem can be split into a set of concurrently executable subproblems, each with just a subset of model coefficients. By maximally splitting across the coefficients and incorporating a novel relaxation technique, a maximally split and relaxed ADMM (MS-RADMM), along with a scalarwise implementation, is developed for the regularized ELM (RELM). The convergence conditions and the convergence rate of the MS-RADMM are established, which exhibits linear convergence with a smaller convergence ratio than the unrelaxed maximally split ADMM. The optimal parameter values of the MS-RADMM are obtained and a fast parameter selection scheme is provided. Experiments on ten benchmark classification data sets are conducted, the results of which demonstrate the fast convergence and parallelism of the MS-RADMM. Complexity comparisons with the matrix-inversion-based method in terms of the numbers of multiplication and addition operations, the computation time and the number of memory cells are provided for performance evaluation of the MS-RADMM.
Xiaoping Lai, Jiuwen Cao, Xiaofeng Huang, Tianlei Wang, Zhiping Lin 0001
IEEE Trans. Neural Networks Learn. Syst.2
2019 Minimax Magnitude Response Approximation of Pole-radius Constrained IIR Digital Filters
abstract
Design of infinite impulse response (IIR) digital filters to approximate some desired magnitude-frequency response is a classical research topic in signal processing. When a pole radius constraint is imposed, however, the problem becomes challenging and few solution methods are available. This paper converts the magnitude-response approximation problem into another problem that approximates the desired magnitude response and an accompanied phase response simultaneously. By iteratively updating the accompanied phase response, a solution to the original magnitude-response approximation problem can be obtained. A striking feature of the proposed method is that the pole radius constraint can be easily incorporated in the problem. Simulations and comparisons demonstrate the effectiveness of the method.
Xiaoping Lai, Jiuwen Cao, Zhiping Lin 0001
ICASSP2
2019 An Enhanced Hierarchical Extreme Learning Machine with Random Sparse Matrix Based Autoencoder
abstract
Recently, by employing the stacked extreme learning machine (ELM) based autoencoders (ELM-AE) and sparse AEs (SAE), multilayer ELM (ML-ELM) and hierarchical ELM (H-ELM) has been developed. Compared to the conventional stacked AEs, the ML-ELM and H-ELM usually achieve better generalization performance with a significantly reduced training time. However, the ℓ1-norm based SAE may suffer the overfitting problem and it is unable to provide analytical solution leading to long training time for big data. To alleviate these deficiencies, we propose an enhanced H-ELM (EH-ELM) with a novel random sparse matrix based AE (SMA) in this paper. The contributions are in two aspects, 1) utilizing the random sparse matrix, the sparse features can be obtained; 2) the proposed SMA can provide an analytical solution so that the high computational complexity issue in SAE can be addressed. Experimental results on benchmark datasets show that the proposed EH-ELM achieves a higher recognition rate and a faster training speed than H-ELM and ML-ELM.
Tianlei Wang, Xiaoping Lai, Jiuwen Cao, Chi-Man Vong, Badong Chen
ICASSP3
2019 View-Invariant Gait Recognition Based on Deterministic Learning and Knowledge Fusion
abstract
Deformation of gait silhouettes caused by different view angles heavily affects the performance of gait recognition. In this paper, a new method based on deterministic learning and knowledge fusion is proposed to eliminate the effect of view angle for efficient view-invariant gait recognition. First, the binarized walking silhouettes are characterized with three kinds of time-varying width parameters. The nonlinear dynamics underlying different individuals' width parameters is effectively approximated by radial basis function (RBF) neural networks through deterministic learning algorithm. The extracted gait dynamics captures the spatio-temporal characteristics of human walking, represents the dynamics of gait motion, and is shown to be insensitive to the variance across various view angles. The learned knowledge of gait dynamics is stored in constant RBF networks and used as the gait pattern. Second, in order to handle the problem of view change no matter the variation is small or large, the learned knowledge of gait dynamics from different views is fused by constructing a deep convolutional and recurrent neural network (CRNN) model for later human identification task. This knowledge fusion strategy can take advantage of the encoded local characteristics extracted from the CNN and the long-term dependencies captured by the RNN. Experimental results show that promising recognition accuracy can be achieved.
Muqing Deng, Jiuwen Cao, Xiaoreng Feng
IJCNN3
2019 Maximum correntropy adaptation approach for robust compressive sensing reconstruction
Yicong He, Fei Wang 0008, Jiuwen Cao, Badong Chen
Inf. Sci.4
2019 Radar emitter identification with bispectrum and hierarchical extreme learning machine
Ru Cao, Jiuwen Cao, Jian-Ping Mei, Chun Yin, Xuegang Huang
Multim. Tools Appl.2
2019 Urban noise recognition with convolutional neural network
Jiuwen Cao, Chun Yin, Danping Wang, Pierre-Paul Vidal
Multim. Tools Appl.1
2019 Extreme learning machine model for water network management
Ahmed M. A. Sattar, Ömer Faruk Ertugrul, Bahram Gharabaghi, Edward A. McBean, Jiuwen Cao
Neural Comput. Appl.5
2019 Multilayer one-class extreme learning machine
Haozhen Dai, Jiuwen Cao, Tianlei Wang, Muqing Deng, Zhi-Xin Yang 0001
Neural Networks2
2019 Extreme Learning Machine With Affine Transformation Inputs in an Activation Function
abstract
The extreme learning machine (ELM) has attracted much attention over the past decade due to its fast learning speed and convincing generalization performance. However, there still remains a practical issue to be approached when applying the ELM: the randomly generated hidden node parameters without tuning can lead to the hidden node outputs being nonuniformly distributed, thus giving rise to poor generalization performance. To address this deficiency, a novel activation function with an affine transformation (AT) on its input is introduced into the ELM, which leads to an improved ELM algorithm that is referred to as an AT-ELM in this paper. The scaling and translation parameters of the AT activation function are computed based on the maximum entropy principle in such a way that the hidden layer outputs approximately obey a uniform distribution. Application of the AT-ELM algorithm in nonlinear function regression shows its robustness to the range scaling of the network inputs. Experiments on nonlinear function regression, real-world data set classification, and benchmark image recognition demonstrate better performance for the AT-ELM compared with the original ELM, the regularized ELM, and the kernel ELM. Recognition results on benchmark image data sets also reveal that the AT-ELM outperforms several other state-of-the-art algorithms in general.
Jiuwen Cao, Kai Zhang 0008, Hongwei Yong, Xiaoping Lai, Badong Chen, Zhiping Lin 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 A Novel Relaxed ADMM with Highly Parallel Implementation for Extreme Learning Machine
abstract
One of the most attractive features of the extreme learning machine (ELM) is its fast speed of learning. In a big data environment, however, ELM may still suffer an overly-heavy computational issue. This paper presents a novel relaxed alternating direction method of multipliers (ADMM) for convex model fitting problems with a focus on a highly parallel implementation for least-squares problems arising from neural network training by ELM. Convergence results and computational complexity of the relaxed ADMM for least-squares problems are given, and comparisons with existing methods are also provided.
Xiaoping Lai, Jiuwen Cao, Zhiping Lin 0001
ISCAS2
2018 Deep CNNs for microscopic image classification by exploiting transfer learning and feature concatenation
abstract
Deep convolutional neural networks (CNNs) have become one of the state-of-the-art methods for image classification in various domains. For biomedical image classification where the number of training images is generally limited, transfer learning using CNNs is often applied. Such technique extracts generic image features from nature image datasets and these features can be directly adopted for feature extraction in smaller datasets. In this paper, we propose a novel deep neural network architecture based on transfer learning for microscopic image classification. In our proposed network, we concatenate the features extracted from three pretrained deep CNNs. The concatenated features are then used to train two fully-connected layers to perform classification. In the experiments on both the 2D-Hela and the PAP-smear datasets, our proposed network architecture produces significant performance gains comparing to the neural network structure that uses only features extracted from single CNN and several traditional classification methods.
Long D. Nguyen, Dongyun Lin, Zhiping Lin 0001, Jiuwen Cao
ISCAS4
2018 Research on crack detection applications of improved PCNN algorithm in moi nondestructive test method
Yuhua Cheng 0001, Lulu Tian, Chun Yin, Xuegang Huang, Jiuwen Cao, Libing Bai
Neurocomputing5
2018 Kernel based online learning for imbalance multiclass classification
Shuya Ding, Bilal Mirza, Zhiping Lin 0001, Jiuwen Cao, Xiaoping Lai, Tam V. Nguyen 0002, Jose Sepulveda
Neurocomputing4
2018 Design of optimal lighting control strategy based on multi-variable fractional-order extremum seeking method
Chun Yin, Xuegang Huang, Sara Dadras, Yuhua Cheng 0001, Jiuwen Cao, Hadi Malek, Jun Mei
Inf. Sci.5
2018 Mixture correntropy for robust learning
Badong Chen, Jiuwen Cao, Harry Qin
Pattern Recognit.5
2018 Postboosting Using Extended G-Mean for Online Sequential Multiclass Imbalance Learning
abstract
In this paper, a novel learning method called postboosting using extended G-mean (PBG) is proposed for online sequential multiclass imbalance learning (OS-MIL) in neural networks. PBG is effective due to three reasons. 1) Through postadjusting a classification boundary under extended G-mean, the challenging issue of imbalanced class distribution for sequentially arriving multiclass data can be effectively resolved. 2) A newly derived update rule for online sequential learning is proposed, which produces a high G-mean for current model and simultaneously possesses almost the same information of its previous models. 3) A dynamic adjustment mechanism provided by extended G-mean is valid to deal with the unresolved challenging dense-majority problem and two dynamic changing issues, namely, dynamic changing data scarcity (DCDS) and dynamic changing data diversity (DCDD). Compared to other OS-MIL methods, PBG is highly effective on resolving DCDS, while PBG is the only method to resolve dense-majority and DCDD. Furthermore, PBG can directly and effectively handle unscaled data stream. Experiments have been conducted for PBG and two popular OS-MIL methods for neural networks under massive binary and multiclass data sets. Through the analyses of experimental results, PBG is shown to outperform the other compared methods on all data sets in various aspects including the issues of data scarcity, dense-majority, DCDS, DCDD, and unscaled data.
Chi-Man Vong, Jie Du 0001, Chiman Wong, Jiuwen Cao
IEEE Trans. Neural Networks Learn. Syst.4
2018 Kernel-Based Multilayer Extreme Learning Machines for Representation Learning
abstract
Recently, multilayer extreme learning machine (ML-ELM) was applied to stacked autoencoder (SAE) for representation learning. In contrast to traditional SAE, the training time of ML-ELM is significantly reduced from hours to seconds with high accuracy. However, ML-ELM suffers from several drawbacks: 1) manual tuning on the number of hidden nodes in every layer is an uncertain factor to training time and generalization; 2) random projection of input weights and bias in every layer of ML-ELM leads to suboptimal model generalization; 3) the pseudoinverse solution for output weights in every layer incurs relatively large reconstruction error; and 4) the storage and execution time for transformation matrices in representation learning are proportional to the number of hidden layers. Inspired by kernel learning, a kernel version of ML-ELM is developed, namely, multilayer kernel ELM (ML-KELM), whose contributions are: 1) elimination of manual tuning on the number of hidden nodes in every layer; 2) no random projection mechanism so as to obtain optimal model generalization; 3) exact inverse solution for output weights is guaranteed under invertible kernel matrix, resulting to smaller reconstruction error; and 4) all transformation matrices are unified into two matrices only, so that storage can be reduced and may shorten model execution time. Benchmark data sets of different sizes have been employed for the evaluation of ML-KELM. Experimental results have verified the contributions of the proposed ML-KELM. The improvement in accuracy over benchmark data sets is up to 7%.
Chiman Wong, Chi-Man Vong, Pak-Kin Wong 0001, Jiuwen Cao
IEEE Trans. Neural Networks Learn. Syst.4
2017 Design of orthogonal filterbanks with rational coefficients using Grobner bases
abstract
Filterbanks are widely used in many applications. In the literature, most filterbanks with irrational value coefficients were designed, while in practice, filters and filterbanks are implemented using rational coefficients. This paper introduces a powerful mathematical tool called Gröbner basis (GB)into the design of orthogonal filterbanks. Specifically, we use GB to find good rational lattice parameters for Daubechies length 6 (Db6) filterbank. We show that there does not exist any Db6 filterbank with rational coefficients and also meeting the first two vanishing moments (VM's). Moreover, we obtain several good Db6 filterbanks which meet the first VM exactly and the second VM approximately.
Nhu Y. Le, Zhiping Lin 0001, David B. H. Tay, Li Xu 0004, Jiuwen Cao
ISCAS5
2017 Document image binarization via optimized hybrid thresholding
abstract
Document image binarization is a crucial step towards optical character recognition and analysis. One common way to achieve image binarization is thresholding. Thresholding methods can be divided into global and local ones in terms of the regional information used in obtaining the threshold values. Both methods have their respective drawbacks. Global methods can not adapt to background variations while local methods have the problem of local widow size determination. Hybrid methods that combines both local and global thresholds and can alleviating these drawbacks. In this paper, the hybrid threshold method is utilized and trade-off between the local and global contents is determined using variational optimization. The proposed algorithm is tested on (H-)DIBCO benchmarks and has shown superior performance to one state-of-the-art document image binarization method.
Yunfeng Liang, Zhiping Lin 0001, Lei Sun 0006, Jiuwen Cao
ISCAS4
2017 LLC encoded BoW features and softmax regression for microscopic image classification
abstract
This paper proposes a method based on the bag-of-words (BoW) and the softmax regression for microscopic image classification. Essentially, the locality-constrained linear coding (LLC) is adopted for local feature encoding. Compared with the traditionally adopted vector quantization (VQ) in the BoW framework, the LLC encodes local structures of microscopic images with lower quantization errors and generates a sparse image representation. This enables the use of linear classifiers with low computational complexity. A softmax regression classifier is then adopted to address the multi-categorical classification task where the confidence of categorical prediction is quantified by posterior probabilities. Compared with other linear classifiers (such as the linear SVM) which only assign labels to images, such probabilistic outputs provide extra quantitative information to analyze misclassified images. Our experiments on the 2D-Hela and the PAP smear data sets show significant performance improvement of the proposed method comparing with competing methods using different features and classifiers under the BoW framework.
Dongyun Lin, Zhiping Lin 0001, Lei Sun 0006, Kar-Ann Toh, Jiuwen Cao
ISCAS5
2017 Acoustic vector sensor: reviews and future perspectives
abstract
Acoustic vector sensor (AVS) has been recently researched and developed for acoustic wave capturing and signal processing. Conventional array generally employs spatially displayed sensors for signal enhancement, source localisation, target tracking, etc. However, the large size usually limits its implementations on some portable devices. AVS which generally includes one omni‐directional sensor and three orthogonally co‐located directional sensors has been recently introduced. An AVS is able to provide the four‐dimensional information of sound field in space: the acoustic pressure and its three‐dimensional particle velocities. A compact assembled AVS could be as small as a match head and the weight can be <50 g. Benefits from these properties, AVS tends to be more attractive for exploitation and commercialisation than conventional sensor array. To have a well understanding of the research progress on AVS, an overview on its recent developments is first given in this study. Then, discussions of challenges on AVS and extensions on its possible future prospects are presented.
Jiuwen Cao, Xiaoping Lai
IET Signal Process.1
2017 Excavation equipment classification based on improved MFCC features and ELM
Jiuwen Cao, Tuo Zhao, Jianzhong Wang 0003, Ruirong Wang, Yun Chen 0008
Neurocomputing1
2017 Excavation Equipment Recognition Based on Novel Acoustic Statistical Features
abstract
Excavation equipment recognition attracts increasing attentions in recent years due to its significance in underground pipeline network protection and civil construction management. In this paper, a novel classification algorithm based on acoustics processing is proposed for four representative excavation equipments. New acoustic statistical features, namely, the short frame energy ratio, concentration of spectrum amplitude ratio, truncated energy range, and interval of pulse are first developed to characterize acoustic signals. Then, probability density distributions of these acoustic features are analyzed and a novel classifier is presented. Experiments on real recorded acoustics of the four excavation devices are conducted to demonstrate the effectiveness of the proposed algorithm. Comparisons with two popular machine learning methods, support vector machine and extreme learning machine, combined with the popular linear prediction cepstral coefficients are provided to show the generalization capability of our method. A real surveillance system using our algorithm is developed and installed in a metro construction site for real-time recognition performance validation.
Jiuwen Cao, Jianzhong Wang 0003, Ruirong Wang
IEEE Trans. Cybern.1
2016 MVDR beamformer analysis of acoustic vector sensor with single directional interference
abstract
The minimum variance distortionless response beamformer (MVDR) using the direction-of-arrival (DOA) of the source signal as the steering vector for an acoustic vector sensor (AVS) s proposed. The interference reduction performance of the MVDR beamformer under a scenario of one directional interference and different sensor noises is analyzed. The output signal-to-noise ratio (oSNR) and array gain of a single AVS are presented. The theoretical upper and lower bounds of oSNR and array gain for an AVS under the scenario of one directional interference and background noise are given. Simulations are provided to support the analysis.
Jiuwen Cao, Xiaoping Lai
ISCAS1
2016 A matrix-based algorithm for the CLS design of centrally symmetric 2-D FIR filters
abstract
An efficient algorithm is presented in this paper for the constrained least-squares design of centrally symmetric two-dimensional (2-D) finite impulse response filters. The problem is basically a quadratic programming with positive-definite quadratic cost and linear constraint functions. This paper formulates both the cost and constraint functions in terms of two coefficient matrices of small size and then presents an efficient algorithm to solve the quadratic programming directly for the two coefficient matrices rather than vectorizing it first as in conventional methods. Design example and comparisons demonstrate the effectiveness and high efficiency of the proposed algorithm.
Xiaoying Hong, Ruijie Zhao 0002, Xiaoping Lai, Jiuwen Cao
ISCAS4
2016 Landmark recognition with compact BoW histogram and ensemble ELM
Jiuwen Cao, Tao Chen 0003, Jiayuan Fan 0001
Multim. Tools Appl.1
2016 Ensemble based extreme learning machine for cross-modality face matching
Yi Jin 0001, Jiuwen Cao, Ruicong Zhi
Multim. Tools Appl.2
2016 Extreme learning machine and adaptive sparse representation for image classification
Jiuwen Cao, Kai Zhang 0008, Minxia Luo, Chun Yin, Xiaoping Lai
Neural Networks1
2015 Performance bound of multiple hypotheses classification in compressed sensing
abstract
Compressed sensing (CS) has been widely researched in the past decade due to its important contributions in sparse signal processing. In this paper, we study the problem of multiple hypotheses classification with sparse signals in compressed sensing. The performance of classifying sparse signals reconstructed with the underdetermined linear measurements under Gaussian random noise is considered. With the prior knowledge of the support set of a sparse signal, the theoretical classification bound with the recovered signal based on the oracle estimator and the restricted isometry property (RIP) of the sampling matrix is developed. The effectiveness of the proposed theoretical bound is demonstrated by the simulations results obtained by four representative reconstruction algorithms in CS.
Jiuwen Cao, Zhiping Lin 0001
ISCAS1
2015 Voting based weighted online sequential extreme learning machine for imbalance multi-class classification
abstract
In this paper, a voting based weighted online sequential extreme learning machine (VWOS-ELM) is proposed for class imbalance learning (CIL). VWOS-ELM is the first sequential classifier that can tackle the class imbalance problem in multi-class data streams. Utilizing WOS-ELM and the recently proposed voting based online sequential extreme learning machine (VOS-ELM) method, VWOS-ELM adapts better to newly received data than the original WOS-ELM method. Experimental results show that VWOS-ELM outperforms both the WOS-ELM and the recent meta-cognitive extreme learning machine methods. It also achieves similar performance to that of ensemble of subset OS-ELM (ESOS-ELM) but using fewer independent classifiers.
Bilal Mirza, Zhiping Lin 0001, Jiuwen Cao, Xiaoping Lai
ISCAS3
2014 Bayesian signal detection with compressed measurements
Jiuwen Cao, Zhiping Lin 0001
Inf. Sci.1
2013 Voting base online sequential extreme learning machine for multi-class classification
abstract
In this paper, we propose a voting based online sequential extreme learning machine (VOS-ELM) for single hidden layer feedforward networks (SLFNs) to perform the online sequential multi-class classification. Utilizing the recent voting based extreme learning machine (V-ELM) and the online sequential extreme learning machine (OS-ELM), the newly developed VOS-ELM is able to classify online sequences by learning data one-by-one or chunk-by-chunk with fixed or varying chunk size and to reach a higher classification accuracy than the original OS-ELM. Simulations on several real world classification datasets show that VOS-ELM outperforms OS-ELM as well as several state-of-the-art online sequential algorithms.
Jiuwen Cao, Zhiping Lin 0001, Guang-Bin Huang
ISCAS1
2012 The detection bound of the probability of error in compressed sensing using Bayesian approach
abstract
In this paper, we consider the theoretical bound of the probability of error in compressed sensing (CS) with the Bayesian approach. In the detection problem, the signal is sparse and is reconstructed from a compressed measurement vector. Utilizing the oracle estimator in CS, we provide a theoretical bound of the probability of error when the noise in CS is white Gaussian noise (WGN). We show that without any additional information in CS, the probability of error obtained using the signal reconstructed by four recovery algorithms: the basis pursuit denoising (BPDN) algorithm, the Dantzig selector, the orthogonal matching pursuit (OMP) method and the compressive sampling matching pursuit (CoSaMP) algorithm is always larger than the derived theoretical bound. Simulation results demonstrate the effectiveness of our result.
Jiuwen Cao, Zhiping Lin 0001
ISCAS1
2012 Voting based extreme learning machine
Jiuwen Cao, Zhiping Lin 0001, Guang-Bin Huang, Nan Liu 0003
Inf. Sci.1
2012 Self-Adaptive Evolutionary Extreme Learning Machine
Jiuwen Cao, Zhiping Lin 0001, Guang-Bin Huang
Neural Process. Lett.1
2012 An Intelligent Scoring System and Its Application to Cardiac Arrest Prediction
abstract
Traditional risk score prediction is based on vital signs and clinical assessment. In this paper, we present an intelligent scoring system for the prediction of cardiac arrest within 72 h. The patient population is represented by a set of feature vectors, from which risk scores are derived based on geometric distance calculation and support vector machine. Each feature vector is a combination of heart rate variability (HRV) parameters and vital signs. Performance evaluation is conducted on the leave-one-out cross-validation framework, and receiver operating characteristic, sensitivity, specificity, positive predictive value, and negative predictive value are reported. Experimental results reveal that the proposed scoring system not only achieves satisfactory performance on determining the risk of cardiac arrest within 72 h but also has the ability to generate continuous risk scores rather than a simple binary decision by a traditional classifier. Furthermore, the proposed scoring system works well for both balanced and imbalanced datasets, and the combination of HRV parameters and vital signs shows superiority in prediction to using HRV parameters only or vital signs only.
Nan Liu 0003, Zhiping Lin 0001, Jiuwen Cao, Zhixiong Koh, Tongtong Zhang, Guang-Bin Huang, Wee Ser, Marcus Eng Hock Ong
IEEE Trans. Inf. Technol. Biomed.3
2011 Composite Function Wavelet Neural Networks with Differential Evolution and Extreme Learning Machine
Jiuwen Cao, Zhiping Lin 0001, Guang-Bin Huang
Neural Process. Lett.1
2010 Composite function wavelet neural networks with extreme learning machine
Jiuwen Cao, Zhiping Lin 0001, Guang-Bin Huang
Neurocomputing1