EDBT 2026 Demo / reviewers in the wild / expert
Tianlei Wang
dblp:132/5409
· DBLP profile ↗
43ranked-venue papers
10as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Summarize Before Glimpse: Brain-Inspired Non-Autoregressive Scene Text RecognizerabstractThe language modeling paradigm for scene text recognition (STR) has demonstrated impressive universal capabilities across extensive STR scenarios. However, existing methods still encounter challenges in effectively handling text images with irregular shapes and diverse appearances (e.g., curve, artistic, multi-oriented) due to the absence of contextual information during initial decoding. In this work, inspired by the principle of ‘forest before trees’ in human visual perception, we introduce NASTR, a non-autoregressive scene text recognizer capable of endowing global-aware for the attentional decoder. Specifically, we design a global-to-local attention procedure, simulating the mechanism of globally holistic visual signal processing preceding locally detailed response in the human brain visual system. This is achieved by leveraging the global image information queries to condition the generation of glimpse vectors at each decoding time step. This procedure empowers the NASTR model to achieve on-par performance with its state-of-the-art autoregressive counterparts, while operating in a fully parallel manner. Moreover, we propose multiple optional and flexible encoding constraint components to alleviate the representation quality degradation issue caused by the global image information queries in handling STR tasks with multilingual and in multi-domains. These components constrain the global image features from the perspective of global structural, global semantic, and linguistic knowledge. Extensive experimental results demonstrate that NASTR consistently outperforms existing methods on both Chinese and English STR benchmarks. Our source code, trained models, and logs are available at https://github.com/ML-HDU/NASTR. Tianlei Wang, Zhiping Lin 0001, Jiuwen Cao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | BiPC: Bidirectional Probability Calibration for Unsupervised Domain Adaption
Wenlve Zhou, Zhiheng Zhou 0001, Junyuan Shang, Chang Niu, Xiyuan Tao, Tianlei Wang |
Expert Syst. Appl. | 7 |
| 2025 | F2MCANet: joint frequency fusion and multi-channels attention neural network for surface defect recognition
Tianlei Wang, Zeliang Li, Yikui Zhai, Jiajie Tian |
Soft Comput. | 1 |
| 2025 | CLIP-Vision Guided Few-Shot Metal Surface Defect RecognitionabstractMetal surface defect recognition (MSDR) based on deep learning encounters the challenge of few-shot expert-labeled data. In this study, we proposed a CLIP-vision guided self supervised learning (CVGSSL) framework for representation learning of unlabeled data, completing MSDR using few-shot labeled data. This framework initially generates rich and diverse representation information through multiple CLIP-Vs to ensure effective SSL pretraining, followed by the design of an MLP-adapter to distill knowledge and adapt these representations to recognition tasks. In addition, we constructed a self-constrained loss to address the inherent problem of intraclass and interclass distance ambiguity that causes the representation to fall into an equivocal decision margin. Following label-free pretraining of CVGSSL, the downstream model adapts to one-shot to four-shot defect recognition tasks through fine-tuning. Experimental results demonstrate that CVGSSL outperforms state-of-the-art SSL methods across three public metal surface defect datasets, with the efficacy of the approach validated through extensive ablation experiments. Tianlei Wang, Zeliang Li, Ying Xu 0005, Yikui Zhai, Xiaofen Xing, Kailing Guo, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Combining Autoregressive and Non-Autoregressive Models for Ship License Plate RecognitionabstractShip license plate recognition (SLPR) is a fundamental visual task in intelligent waterway transportation systems that aims to transcribe ship name text images into editable text strings. Previous works construct recognizers with complicated training procedures and additional corpora to boost performance, limiting their efficiency in practical application scenarios. This paper proposes a novel ship license plate recognizer by combining autoregressive (AR) and non-autoregressive (NAR) decoding mechanisms. By adaptively incorporating the dual-branch character representations, our method explicitly sidesteps the dependence on an external language model and is adapted to the weak semantic correlation characteristics of ship name text images. Furthermore, we introduce a confidence-based dynamic reweighting strategy that includes character-and instance-level granularities. This encourages the model to learn more from hard or challenging samples. Experimental results conducted on two SLPR benchmarks demonstrate the effectiveness of our method, showing that it is competitive and outperforms previous approaches. Our study comprehensively explores the relative importance of linguistic and visual cues in SLPR and demonstrates the advantages of the dynamically fused model with dual-branch autoregressive and non-autoregressive over single-branch recognition. Tianlei Wang, Jiuwen Cao |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | On the Behavior of Contrastive Regularization in Improving Chinese Text RecognizerabstractThe dense representation space in Chinese scene text recognition (STR) makes discriminating between categories highly challenging, because of the large candidate category set. Mainstream STR methods have achieved remarkable advancements by leveraging linguistic knowledge to implicitly address this challenge. In this paper, inspired by the correlation between recognizer performance and the distributional properties of character representations, as well as the inherent consistency between this correlation and supervised contrastive learning (SupCon), we thoroughly investigate how to integrate SupCon with an STR model to alleviate this challenge, and elucidate some dynamic behaviors underlying the performance improvements. Specifically, we analyze the SupCon-STR models instantiated with different projectors and evaluate their distributional properties through metrics, including intra-class compactness, inter-class separability, and feature redundancy, while assessing performances that involve in-domain accuracy and cross-domain recognition generalization. The main results reveal how the temperature$\tau$and projectors affect the representation distribution, and highlight that suitable intra-class compactness and sufficient inter-class separability are key factors for delivering competitive performances in both in-domain and cross-domain STR scenarios. Moreover, these results also provide valuable insights into the design of SupCon-STR architectures for diverse resource constraints. Taking existing Chinese STR models as baselines, and combining SupCon-STR with them, the average improvements in cross-domain recognition performance are over 5% across 7 testing datasets. A new state-of-the-art accuracy of 77.19% on the ChineseScenebenchmark is also established. Tianlei Wang, Huanqiang Zeng, Jiuwen Cao |
IEEE Trans. Multim. | 2 |
| 2024 | Attention-Based Deep Neural Network for Point Cloud LearningabstractPoint cloud learning is a vital task in the field of computer vision. Due to the irregularity and disorderliness of point clouds, learning their features has been a challenging issue for a long time. Attention mechanisms have been widely applied in deep learning and have achieved good results in point cloud learning tasks. However, existing models mainly focus on local features, and the stacking of numerous attention layers requires significant computational resources. To address these problems, this study first introduces a framework that considers both local and global features in a parallel manner. Based on this framework, we propose two practical implementations, PCAttn and PCTrans. The PCAttn explores the effectiveness of the channel and spatial attention mechanism, while the PCTrans introduces the self-attention mechanism to the point cloud learning. These models are tested on two publicly available benchmark datasets. Experimental results demonstrate that the proposed methods exhibit high performance and computational efficiency. PCAttn achieves 93.2% overall accuracy on the ModelNet40 dataset and 84.6% overall accuracy on the ScanObjectNN dataset with only 0.61M parameters and 3.4G floating point operations. PCTrans achieves 93.1% overall accuracy on the ModelNet40 dataset and 83.3% overall accuracy on the ScanObjectNN dataset with only 0.68M parameters and 0.35G floating point operations. Tianlei Wang, Ma Luo, Hong Qu 0002 |
IJCNN | 1 |
| 2024 | Audio-Visual Cross-Modal Generation with Multimodal Variational Generative ModelabstractAudio and Visual are two important visual modalities in video content understanding. However, the absence of one modality may be observed in practical applications due to the real environmental factors, which leads to the information loss. Therefore, audio and visual fusion is focused on using the shared and complementary information between modalities to recover the missing modalities from the available data modalities. In this paper, an Adversarial Hierarchical Variational Auto-Encoder (Adv-HVAE) model is proposed to solve this problem of modality data loss. A multimodal representation is first learned using a hierarchical Variational Autoencoder (VAE) model that enables the generation of missing modal data under any subset of available modalities. Also to obtain a more robust multimodal representation, a feature generation network is utilized to approximate the latent distribution of missing modalities. Finally, the adversarial training network is shown to be effective in improving the data quality generated through the Adv-HVAE framework. Experimental results demonstrate that Adv-HVAE achieves best generation results on two benchmark datasets, avMNIST and Sub-URMP. Zhubin Xu, Tianlei Wang, Dinghan Hu, Huanqiang Zeng, Jiuwen Cao |
ISCAS | 2 |
| 2024 | Reinforcement Learning-Based Adaptive Search Strategy Combination AlgorithmabstractIntegrating multiple search operators to leverage the characteristics of different search operators is one of the common methods to enhance the performance of evolutionary algorithms. Most of these combination algorithms adopt a fixed framework, which makes it difficult to adaptively select the appropriate search operator during the algorithm iteration process, resulting in poor performance on certain problems. To overcome the aforementioned shortcomings of combination algorithms, an adaptive search strategy combination algorithm based on reinforcement learning and neighborhood search (RLASCA) is proposed. During the algorithm iteration process, a reinforcement learning-based adaptive search operator selection method (RLAS) is designed, enabling the algorithm to adaptively select the appropriate search operator based on the state of the individual. Furthermore, to address the issue of premature convergence, a neighborhood search strategy based on differential evolution (NSDE) is designed to increase the diversity of the population. To verify the effectiveness of the proposed algorithm, a comprehensive testing was conducted using the CEC2017 test suite and two constrained engineering design problems. The algorithm was compared with four basic algorithms that compose RLASCA and four classical combination algorithms. The experimental results indicate that RLASCA outperforms the other eight algorithms in terms of convergence speed and accuracy. Sicheng Wan, Jiajie Tian, Tianlei Wang, Zhiyong Hong |
ISPA | 5 |
| 2024 | Nonlinear Control of Crane Systems Based on Intelligence Computing of Disturbance Observer Under Mismatched DisturbanceabstractA control method based on disturbance observer is proposed to address the positioning drift issue of the gantry crane under unmatched disturbances, aiming at precise positioning of the trolley and effective payload swing angle attenuation in a variable-length two-dimensional lifting system. The control scheme comprises a combination of a proportional-derivative controller and a disturbance observer. Firstly, the design of the coupling function in this paper is based on the observation of the natural characteristics of the dynamical model. It cleverly designs the coupling function, which not only helps to effectively reduce the swing angle but also prevents the problem of positioning drift under non-matching disturbances, subsequently, design a controller by integrating intelligent computing with control theory. Secondly, the designed disturbance observer observes and compensates for the system output to achieve disturbance suppression. Thirdly, a thorough stability analysis has been conducted to demonstrate that the proposed controller satisfies the desired conditions. Finally, simulation results illustrate that the pro-posed method success-fully addresses the limitations of existing methods and exhibits superior performance and robustness. Tianlei Wang, Yiyan Wang, Jiajie Tian, Zhiyong Hong |
ISPA | 1 |
| 2024 | A highly efficient ADMM-based algorithm for outlier-robust regression with Huber loss
Tianlei Wang, Xiaoping Lai, Jiuwen Cao |
Appl. Intell. | 1 |
| 2024 | Class-agnostic counting and localization with feature augmentation and scale-adaptive aggregation
Yuhui Du, Hong Qu 0002, Tianlei Wang, Fan Zhang 0068, Mingsheng Fu, Wenyu Chen 0001 |
Knowl. Based Syst. | 4 |
| 2024 | An adaptive multi-level-sets active contour model based on block search
Zhiheng Zhou 0001, Guoqi Liu, Tianlei Wang |
Multim. Tools Appl. | 4 |
| 2024 | M-DDC: MRI based demyelinative diseases classification with U-Net segmentation and convolutional network
Deyang Zhou, Tianlei Wang, Shaonong Wei, Feng Gao 0018, Xiaoping Lai, Jiuwen Cao |
Neural Networks | 3 |
| 2024 | Within-Class Constraint Based Multi-task Autoencoder for One-Class ClassificationabstractAutoencoders (AEs) have attracted much attention in one-class classification (OCC) based unsupervised anomaly detection. The AEs aim to learn the unity features on targets without involving anomalies and thus the targets are expected to obtain smaller reconstruction errors than anomalies. However, AE-based OCC algorithms may suffer from the overgeneralization of AE and fail to detect anomalies that have similar distributions to target data. To address these issues, a novel within-class constraint based multi-task AE (WC-MTAE) is proposed in this paper. WC-MTAE consists of two different task: one for reconstruction and the other for the discrimination-based OCC task. In this way, the encoder is compelled by the OCC task to learn the more compact encoded feature distribution for targets when minimizing OCC loss. Meanwhile, the within-class scatter based penalty term is constructed to further regularize the encoded feature distribution. The aforementioned two improvements enable the unsupervised anomaly detection by the compact encoded features, thereby addressing the issue of the overgeneralization in AEs. Comparisons with several state-of-the-art (SOTA) algorithms on several non-image datasets and an image dataset CIFAR10 are provided where the WC-MTAE is conducted on 3 different network structures including the multilayer perception (MLP), LeNet-type convolution network and full convolution neural network. Extensive experiments demonstrate the superior performance of the proposed WC-MTAE. The source code would be available in future. Tianlei Wang, Wandong Zhang, Xiaoping Lai |
Neural Process. Lett. | 2 |
| 2024 | Matrix randomized autoencoder
Tianlei Wang, Jiuwen Cao, Wandong Zhang, Badong Chen |
Pattern Recognit. | 2 |
| 2024 | Auxiliary Label Classification Based Multi-Label Limb Movement Recognition of Preterm InfantabstractLimb movement recognition of preterm infants (PI-LMR) in neonatal intensive care units (NICUs) is important for infant health monitoring. However, little attention has been paid to intelligent PI-LMR. Due to the weak correlation among limb movements of preterm infants, the various limb movement combinations and imbalanced data distributions are the main challenges of PI-LMR. To address these issues, a novel multi-label limb movement recognition (MLLMR) algorithm with a dual-branch structure and multi-label fusion loss is proposed. The various movement combinations can be decomposed into limbs thanks to multi-label learning. Particularly, the multi-label fusion loss consisting of the binary cross entropy (BCE) and the pairwise ranking loss (PRL) is proposed to optimize the probabilities to the ground truth labels and the ranking between positive and negative labels, simultaneously. The weighted fusion loss is further developed to address the imbalanced label distributions. Subsequently, an auxiliary task for the classification of zero-, single- and multi-limb movements is constructed to constrain the feature space of primary task for better multi-label learning. Experiments on real clinical preterm infants video dataset from Jiaxing Maternity and Child Health Care Hospital are conducted and the results demonstrate the effectiveness of the proposed algorithm. Hongliang Lei, Tianlei Wang, Xianfu Bao, Jiuwen Cao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | ChainFrame: A Chain Framework for Point Cloud ClassificationabstractPoint cloud analysis is challenging due to its data structure. To capture the 3-D geometries, prior works mainly rely on exploring local geometric extractors. However, the human visual system suggested that both global and local features should be considered. In this article, we introduce a novel framework for point cloud classification, called ChainFrame, which takes the pair-wise global-local correlations into consideration within the intermediate scales hierarchically. The ChainFrame captures the global features that characterize the entire outline of the object. Simultaneously, the ChainFrame organizes the local features that incorporate the point itself and its neighboring region. With such a framework, our practical implementations (ChainMLP and ChainGraph) perform on par or even better than other methods. Evaluations on two popular datasets show the effectiveness and efficiency of our ChainFrame. ChainMLP and ChainGraph achieve$ 87.2%$and$87.6%$overall point-wise accuracy scores, respectively, on the real-world ScanObjectNN benchmark. Besides, ChainMLP delivers comparable performance on ModelNet40 with only 0.47 M parameters and 0.33 G floating point operations (FLOPs), which are much smaller than the prior methods. Tianlei Wang, Mingsheng Fu, Hong Qu 0002, Ma Luo |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Width-Adaptive CNN: Fast CU Partition Prediction for VVC Screen Content CodingabstractScreen content coding (SCC) in Versatile Video Coding (VVC) improves the coding efficiency of screen content videos (SCVs) significantly but results in high computational complexity due to the quad-tree plus multi-type tree (QTMT) structure of the coding unit (CU) partitioning. Therefore, we make the first attempt to reduce the encoding complexity from the perspective of CU partitioning for SCC in VVC. To this end, a fast CU partition prediction method is technically developed for VVC-SCC. First, to solve the problem of lacking sufficient SCC training data, SCVs are collected to establish a database containing CUs of various sizes and corresponding partition labels. Second, to determine the partition decision in advance, a novel WA-CNN model is proposed, which is capable of predicting two large CUs for VVC-SCC by adjusting the feature channels based on the size of input CU blocks. Finally, considering the imbalanced proportion of diverse partition decisions, a loss function with the weight that equalizes the contribution of imbalanced data is formulated to train the proposed WA-CNN model. Experimental results show that the proposed model reduces the SCC intra-encoding time by 35.65%${\sim }$38.31% with an average of 1.84%${\sim }$2.42% BDBR increase. Chao Jiao, Huanqiang Zeng, Jing Chen 0001, Chih-Hsien Hsia, Tianlei Wang, Kai-Kuang Ma |
IEEE Trans. Multim. | 5 |
| 2024 | M$^{3}$ANet: Multi-Modal and Multi-Attention Fusion Network for Ship License Plate RecognitionabstractShiplicense plate recognition (SLPR) plays an important role in intelligent waterway management, but few attention has been paid to SLPR in scene text recognition (STR) community. Inspired by various outstanding achievements on STR, combined the intrinsic properties of SLPR, we propose aMulti-Modal andMulti-Attention dynamic fusion network (M$^{3}$ANet) for SLPR in this article. Specifically, the visual-language joint modeling for SLPR is developed and the channel-spatial-self attention dynamic fusion mechanism is proposed for accuracy boosting. Explicitly fusing linguistic information extracted from ship name related corpus improves the adaptability of the recognition model to occlusion, background confusion, blur, etc., which is integrated with vision features to establish a multi-modal recognition network. Gated fully fusion is utilized to fuse visual features re-weighted by multi-attention components, inducing flexible compatibility with multiple types of decoders and more refined recognition decoder inputs. Additionally, to comprehensively mine spatially salient text regions in ship license plate images, we investigate the grouped spatial attention. Extensive experiments empirically demonstrate the effectiveness of M$^{3}$ANet and superior performance (93.80% with regular images, while 90.34% with irregular images) on two benchmarks. Tianlei Wang, Jiangmin Tian, Jiuwen Cao |
IEEE Trans. Multim. | 3 |
| 2024 | Multimodal Moore-Penrose Inverse-Based Recomputation Framework for Big Data AnalysisabstractMost multilayer Moore-Penrose inverse (MPI)-based neural networks, such as deep random vector functional link (RVFL), are structured with two separate stages: unsupervised feature encoding and supervised pattern classification. Once the unsupervised learning is finished, the latent encoding is fixed without supervised fine-tuning. However, in complex tasks such as handling the ImageNet dataset, there are often many more clues that can be directly encoded, while unsupervised learning, by definition, cannot know exactly what is useful for a certain task. There is a need to retrain the latent space representations in the supervised pattern classification stage to learn some clues that unsupervised learning has not yet been learned. In particular, the residual error in the output layer is pulled back to each hidden layer, and the parameters of the hidden layers are recalculated with MPI for more robust representations. In this article, a recomputation-based multilayer network using Moore-Penrose inverse (RML-MP) is developed. A sparse RML-MP (SRML-MP) model to boost the performance of RML-MP is then proposed. The experimental results with varying training samples (from 3k to 1.8 million) show that the proposed models provide higher Top-1 testing accuracy than most representation learning algorithms. For reproducibility, the source codes are available at https://github.com/W1AE/Retraining. Wandong Zhang, Yimin Yang 0001, Q. M. Jonathan Wu, Tianlei Wang, Hui Zhang 0023 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Advanced License Plate Detector in Low-Quality Images with Smooth Regression Constraint
Jiefu Yu, Tianlei Wang, Jiangmin Tian, Fangyong Xu, Jiuwen Cao |
PRCV (2) | 3 |
| 2023 | Study the correlation between the readme file of GitHub projects and their popularity
Tianlei Wang, Shaowei Wang 0002, Tse-Hsun (Peter) Chen |
J. Syst. Softw. | 1 |
| 2023 | Multichannel Matrix Randomized Autoencoder
Tianlei Wang, Jiuwen Cao |
Neural Process. Lett. | 2 |
| 2023 | A Transformer-Based End-to-End Automatic Speech Recognition AlgorithmabstractEnd-to-End (E2E) automatic speech recognition (ASR) becomes popular recent years and has been widely used in many applications. However, current ASR algorithms are usually less effective when applied in specific applications with terminologies such as medical and economic fields. To address this issue, we propose a powerful Transformer based ASR decoding method for beam searching, called soft beam pruning algorithm (SBPA). SBPA can dynamically adjust the width of beam search. Meanwhile, a prefix module (PM) is added to access the contextual information and avoid removing professional words in the beam search. Combining SBPA and PM, the proposed ASR can achieve promising recognition performance on professional terminologies. To verify the effectiveness, experiments are conducted on real-world conversation data with medical terminology. It is shown that the proposed ASR achieved significant performance on both professional and regular words. Fang Dong 0003, Yiyang Qian, Tianlei Wang, Jiuwen Cao |
IEEE Signal Process. Lett. | 3 |
| 2023 | Efficient ADMM-Based Algorithm for Regularized Minimax ApproximationabstractMinimax approximations have found many applications but are lack of efficient solution algorithms for large-scale problems. Based on the alternating direction method of multipliers (ADMM) for convex optimization, this letter presents an efficient scalarwise algorithm for a regularized minimax approximation problem. The ADMM-based algorithm is then applied in the minimax design of two-dimensional (2-D) digital filters and the training of randomized neural networks for regression on a realworld benchmark dataset. Experimental results demonstrate the fast convergence rate and low computational complexity of the proposed algorithm, as well as the good approximation/prediction performance of the learned approximation model. Xuanyue Shentu, Xiaoping Lai, Tianlei Wang, Jiuwen Cao |
IEEE Signal Process. Lett. | 3 |
| 2023 | Ship License Plate Super-Resolution in the WildabstractShip license plate (SLP) recognition plays an important role in ship supervision and harbour management. Practically, low-resolution (LR) SLP images are illegible and challenging to SLP recognition. Most existing super-resolution (SR) methods are not suitable for real-world LR SLP images with over smoothed reconstructions. To alleviate these deficiencies, in this letter, we propose a parallel enhanced SR generative adversarial network (PESRGAN) for SLP images. A novel degradation model is developed to construct a more feasible LR dataset. A parallel SR convolutional neural network (SRCNN) module based on ESRGAN is proposed for feature extraction. To characterize the difference between text foreground and background, a new gradient loss is developed in PESRGAN to sharpen the character boundary. Comparisons to many state-of-the-art (SOTA) SR methods are presented to show the effectiveness of the proposed algorithm. Huahua Wu, Jiagui Chen, Tianlei Wang, Xiaoping Lai, Jiuwen Cao |
IEEE Signal Process. Lett. | 3 |
| 2023 | 3D-Gradient Guided Rate Control Model for Screen Content Video CodingabstractCompared with natural videos,screen content videos(SCVs) have particular features, such as fruitful sharper edges, lots of computer-generated graphics and texts, a large amount of flat areas. New tools are adopted toHEVC extensions on Screen Content Coding(HEVC-SCC), the traditional video rate control methods for natural videos are not effective for SCVs. For that, a3D-gradient guided rate control modelfor SCV coding, named 3DG-RC, is proposed to allocate bitrate more efficiently serving for SCVs. By considering the particular spatial-temporal characteristics of SCVs, the spatial and temporal feature extraction scheme is developed by using 3D-gradient filter and performed on the SCV to extract the spatial and temporal features simultaneously for guiding the bit allocation. The spatial-temporal feature similarity between three original reference SCV frames and their reconstructed ones is used to estimate the encoding parameters of the current block and frame. Experimental results demonstrate that compared with the classical and state-of-the-art rate control methods for HEVC-SCC, the proposed 3DG-RC algorithm achieves significant bitrate mismatch reduction and coding efficiency improvement for HEVC-SCC. In specific, the proposed 3DG-RC model outperforms the rate control model in SCM-8.8 with over 41.33% and 37.95% BD-BR savings on average, forlow delay B(LDB) andrandom access(RA) coding structure, respectively. Jing Chen 0001, Huanqiang Zeng, Chih-Hsien Hsia, Tianlei Wang, Kai-Kuang Ma |
IEEE Trans. Multim. | 5 |
| 2022 | Clustering-Guided Pairwise Metric Triplet Loss for Person ReidentificationabstractMost of the loss functions proposed for person reidentification (Re-ID) are expected to be easy to deploy, efficiently improve network performance, and will not introduce redundant parameters. This study proposes a no-parameter and generic clustering-guided pairwise metric triplet (CPM-Triplet) loss based on the hard sample mining triplet loss for the metric learning loss. CPM-Triplet loss deploys two metrics: 1) the Euclidean metric and 2) the cosine metric, to complementarily improve the metric learning of the model. Paralleled to the Euclidean metric, the cosine metric quantifies the sample similarity in a different way to the Euclidean metric, which takes a different perspective to explore the distribution of samples. But the pairwise metric mainly improves the precision between dissimilar samples of the same label and could not solve the problem of excessive outliers. Therefore, the clustering-guided correction term was deployed to apply to all samples with the same label to mine the similarity in the samples, while weakening the influence of outliers in CPM-Triplet loss. Experiments conducted on four benchmark data sets show that the combination of the CPM-Triplet loss and the widely used Bag-of-Tricks baseline generally outperforms the baseline and numerous state-of-the-art methods studied in this article. The source code would be available athttps://github.com/weiyu-zeng/CPM-Triplet-loss. Weiyu Zeng, Tianlei Wang, Jiuwen Cao, Huanqiang Zeng |
IEEE Internet Things J. | 2 |
| 2022 | 3D residual-attention-deep-network-based childhood epilepsy syndrome classification
Yuanmeng Feng, Xiaonan Cui, Tianlei Wang, Tiejia Jiang, Feng Gao 0018, Jiuwen Cao |
Knowl. Based Syst. | 4 |
| 2022 | Deep feature fusion based childhood epilepsy syndrome classification from electroencephalogram
Xiaonan Cui, Dinghan Hu, Jiuwen Cao, Xiaoping Lai, Tianlei Wang, Tiejia Jiang, Feng Gao 0018 |
Neural Networks | 6 |
| 2022 | Scalp EEG functional connection and brain network in infants with West syndrome
Yuanmeng Feng, Tianlei Wang, Jiuwen Cao, Duanpo Wu, Tiejia Jiang, Feng Gao 0018 |
Neural Networks | 3 |
| 2022 | SLPR: A Deep Learning Based Chinese Ship License Plate Recognition FrameworkabstractAutomatic ship license plate recognition (SLPR) for ship identification is of great significance to waterway shipping management. But few attention has been paid to SLPR in the past. In this paper, a novel cascaded Chinese SLPR framework consisting of the quadrangle-based ship license plate detection (QSLPD) algorithm and the rectification-based text recognition network (RTRNet) is developed. Concretely, in QSLPD algorithm, detection is performed based on the pyramid feature fusion architecture ameliorated by the proposed variable receptive field feature enhancement strategy and three task-specific output heads. In addition, a new loss function combining the dice coefficient and cross entropy is explored in the proposed SLPR which can generate significant improvement over the baseline. In RTRNet, regions of interest (RoIs) extraction and irregular text line rectification based on the vertices information predicted by QSLPD are performed before text recognition. Data augmentation are also applied to cope with the problem of limited text recognizer training data and the extremely imbalance distribution of corpus. Extensive experiments are carried out to demonstrate the reliability of the proposed cascaded SLPR framework, that can achieve the highest F-measure of 87.78% and 76.59% with IoU and TIoU metric on the collected dataset, surpasses many existing advanced methods. Jiuwen Cao, Tianlei Wang, Huahua Wu, Jiangmin Tian, Fangyong Xu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Distributed Big Data Computing for Supporting Predictive Analytics of Service RequestsabstractIn the current era of big data, huge volumes of valuable data can be easily generated and collected at a rapid velocity from a wide variety of rich data sources. In recent years, the initiates of open data also led to the willingness of many government, researchers, and organizations to share their data and make them publicly accessible. An example of open big data is service request data. Analyzing these open big data can be for social good. For instance, by analyzing and mining data on non-emergency city service requests, the city could get an insight about its residents’ demand for services. By taking appropriate actions (e.g., adding more staff and/or services, providing more information regarding city services) could enhance the living condition of city residents. In this paper, we present a distributed big data mining system to analyze and mine big data on these non-emergency city service requests. Evaluation on an open big data from a North American city shows the effectiveness and practicality of our distributed big data system in mining these requests for city services and in supporting predictive analytics. Tianlei Wang, James D. Harvey, Carson K. Leung, Adam G. M. Pazdor, Animesh Singh Chauhan, Lihe Fan, Alfredo Cuzzocrea |
COMPSAC | 1 |
| 2021 | Hierarchical One-Class Classifier With Within-Class Scatter-Based AutoencodersabstractAutoencoding is a vital branch of representation learning in deep neural networks (DNNs). The extreme learning machine-based autoencoder (ELM-AE) has been recently developed and has gained popularity for its fast learning speed and ease of implementation. However, the ELM-AE uses random hidden node parameters without tuning, which may generate meaningless encoded features. In this brief, we first propose a within-class scatter information constraint-based AE (WSI-AE) that minimizes both the reconstruction error and the within-class scatter of the encoded features. We then build stacked WSI-AEs into a one-class classification (OCC) algorithm based on the hierarchical regularized least-squared method. The effectiveness of our approach was experimentally demonstrated in comparisons with several state-of-the-art AEs and OCC algorithms. The evaluations were performed on several benchmark data sets. Tianlei Wang, Jiuwen Cao, Xiaoping Lai, Q. M. Jonathan Wu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Affine Transformation Based Hierarchical Extreme Learning MachineabstractRecently, the signal hidden layer feedforward network (SLFN) based extreme learning machine (ELM) has been extended to a hierarchical learning framework (HELM). Although the HELM shows better generalization performance with lower computational complexity than many deep neural networks (DNNs), it is found that as the layer increases, the input distribution of each layer may move to the saturated regime of the non-linear activation function, which affects the generalization performance. However, few attentions have been paid to the data distribution normalization to address this issue. Thus, in this paper, an affine transformation (AT) inputs based activation function layer is introduced to normalize the data distribution and a novel AT based HELM (AT-HELM) is developed. The proposed AT-HELM can adapt the activation function inputs to the distribution of each layer and obtains better generalization performance. Experiments on 29 benchmark datasets are carried out to demonstrate the superiority of AT-HELM. Rongzhi Ma, Jiuwen Cao, Tianlei Wang, Xiaoping Lai |
ISCAS | 3 |
| 2020 | Common-specific feature learning for multi-source domain adaptationabstractMulti‐source domain adaptation (MDA) aims to leverage knowledge from multiple source domains to improve the classification performance on target domains. Different degrees of distribution discrepancies between every two domains pose a huge challenge to MDA tasks. Most works focus on extracting features shared by all domains, which is critical but not enough to reduce distribution discrepancies. In this paper, we propose a method named as common‐specific feature learning (CSFL). Constituting a framework of feature learning, CSFL explores a subspace where the combination of common and specific features makes learned representations comprehensive. Based on this framework, we conduct a metric learning method for learning a discriminative feature representation. Considering redundant information caused by source domains is likely to hurt the performance, we impose an effective low‐rank constraint to remove the redundant information. Further, we adopt structure consistent constraint to preserve the local structure in each domain. CSFL has obtained about 1–5% improvement of mean accuracy, compared to the state‐of‐the‐art shallow methods. Further, compared with 90.2% and 89.4% of the best baseline deep method, CSFL achieves mean accuracy of 90.8% and 89.7% on the Office‐31 and ImageCLEF‐DA datasets respectively. The encouraging results validate the effectiveness of our method. Chang Niu, Junyuan Shang, Zhiheng Zhou 0001, Junchu Huang, Tianlei Wang, Xiangwei Li |
IET Image Process. | 5 |
| 2020 | Non-contact heart rate detection by combining empirical mode decomposition and permutation entropy under non-cooperative face shake
Hongwei Yue, Xiaorong Li, Ken Cai, Huazhou Chen, Shufen Liang, Tianlei Wang |
Neurocomputing | 6 |
| 2020 | Regularized correntropy criterion based semi-supervised ELM
Jie Yang 0051, Jiuwen Cao, Tianlei Wang, Anke Xue, Badong Chen |
Neural Networks | 3 |
| 2020 | A Maximally Split and Relaxed ADMM for Regularized Extreme Learning MachinesabstractOne of the salient features of the extreme learning machine (ELM) is its fast learning speed. However, in a big data environment, the ELM still suffers from an overly heavy computational load due to the high dimensionality and the large amount of data. Using the alternating direction method of multipliers (ADMM), a convex model fitting problem can be split into a set of concurrently executable subproblems, each with just a subset of model coefficients. By maximally splitting across the coefficients and incorporating a novel relaxation technique, a maximally split and relaxed ADMM (MS-RADMM), along with a scalarwise implementation, is developed for the regularized ELM (RELM). The convergence conditions and the convergence rate of the MS-RADMM are established, which exhibits linear convergence with a smaller convergence ratio than the unrelaxed maximally split ADMM. The optimal parameter values of the MS-RADMM are obtained and a fast parameter selection scheme is provided. Experiments on ten benchmark classification data sets are conducted, the results of which demonstrate the fast convergence and parallelism of the MS-RADMM. Complexity comparisons with the matrix-inversion-based method in terms of the numbers of multiplication and addition operations, the computation time and the number of memory cells are provided for performance evaluation of the MS-RADMM. Xiaoping Lai, Jiuwen Cao, Xiaofeng Huang, Tianlei Wang, Zhiping Lin 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | An Enhanced Hierarchical Extreme Learning Machine with Random Sparse Matrix Based AutoencoderabstractRecently, by employing the stacked extreme learning machine (ELM) based autoencoders (ELM-AE) and sparse AEs (SAE), multilayer ELM (ML-ELM) and hierarchical ELM (H-ELM) has been developed. Compared to the conventional stacked AEs, the ML-ELM and H-ELM usually achieve better generalization performance with a significantly reduced training time. However, the ℓ1-norm based SAE may suffer the overfitting problem and it is unable to provide analytical solution leading to long training time for big data. To alleviate these deficiencies, we propose an enhanced H-ELM (EH-ELM) with a novel random sparse matrix based AE (SMA) in this paper. The contributions are in two aspects, 1) utilizing the random sparse matrix, the sparse features can be obtained; 2) the proposed SMA can provide an analytical solution so that the high computational complexity issue in SAE can be addressed. Experimental results on benchmark datasets show that the proposed EH-ELM achieves a higher recognition rate and a faster training speed than H-ELM and ML-ELM. Tianlei Wang, Xiaoping Lai, Jiuwen Cao, Chi-Man Vong, Badong Chen |
ICASSP | 1 |
| 2019 | Prior distribution-based statistical active contour model
Zhiheng Zhou 0001, Tianlei Wang, Ruzheng Zhao |
Multim. Tools Appl. | 3 |
| 2019 | Multilayer one-class extreme learning machine
Haozhen Dai, Jiuwen Cao, Tianlei Wang, Muqing Deng, Zhi-Xin Yang 0001 |
Neural Networks | 3 |