Jiajun Wen 0001

dblp:58/7955 · DBLP profile ↗
← Back
39ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0002-1688-2541ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BCE3S: Binary Cross-Entropy Based Tripartite Synergistic Learning for Long-Tailed Recognition
abstract
For long-tailed recognition (LTR) tasks, high intra-class compactness and inter-class separability in both head and tail classes, as well as balanced separability among all the classifier vectors, are preferred. The existing LTR methods based on cross-entropy (CE) loss not only struggle to learn features with desirable properties but also couple imbalanced classifier vectors in the denominator of its Softmax, amplifying the imbalance effects in LTR. In this paper, for the LTR, we propose a binary cross-entropy (BCE)-based tripartite synergistic learning, termed BCE3S, which consists of three components: (1) BCE-based joint learning optimizes both the classifier and sample features, which achieves better compactness and separability among features than the CE-based joint learning, by decoupling the metrics between feature and the imbalanced classifier vectors in multiple Sigmoid; (2) BCE-based contrastive learning further improves the intra-class compactness of features; (3) BCE-based uniform learning balances the separability among classifier vectors and interactively enhances the feature properties by combining with the joint learning. The extensive experiments show that the LTR model trained by BCE3S not only achieves higher compactness and separability among sample features, but also balances the classifier's separability, achieving SOTA performance on various long-tailed datasets such as CIFAR10-LT, CIFAR100-LT, ImageNet-LT, and iNaturalist2018.
Weijia Fan, Qiufu Li, Jiajun Wen 0001, Xiaoyang Peng
AAAI3
2026 AEPL: Adaptive empirical prototype learning with dynamic margins for deep face recognition
Weijia Fan, Zhixiang Cai, Chunsong Chen, Yanxi Liu 0004, Jiajun Wen 0001, Xi Jia, LinLin Shen, Jiancan Zhou, Qiufu Li
Pattern Anal. Appl.6
2025 DeeperForward: Enhanced Forward-Forward Training for Deeper and Better Performance
abstract
While backpropagation effectively trains models, it presents challenges related to bio-plausibility, resulting in high memory demands and limited parallelism. Recently, Hinton (2022) proposed the Forward-Forward (FF) algorithm for high-parallel local updates. FF leverages squared sums as the local update target, termed goodness, and decouples goodness by normalizing the vector length to extract new features. However, this design encounters issues with feature scaling and deactivated neurons, limiting its application mainly to shallow networks. This paper proposes a novel goodness design utilizing **layer normalization** and **mean goodness** to overcome these challenges, demonstrating performance improvements even in 17-layer CNNs. Experiments on CIFAR-10, MNIST, and Fashion-MNIST show significant advantages over existing FF-based algorithms, highlighting the potential of FF in deep models. Furthermore, the model parallel strategy is proposed to achieve highly efficient training based on the property of local updates.
Yang Zhang 0012, Weizhao He, Jiajun Wen 0001, LinLin Shen, Weicheng Xie 0001
ICLR4
2025 MC3D-AD: A Unified Geometry-aware Reconstruction Model for Multi-category 3D Anomaly Detection
abstract
3D Anomaly Detection (AD) is a promising means of controlling the quality of manufactured products. However, existing methods typically require carefully training a task-specific model for each category independently, leading to high cost, low efficiency, and weak generalization. This study presents a novel unified model for Multi-Category 3D Anomaly Detection (MC3D-AD) that aims to utilize both local and global geometry-aware information to reconstruct normal representations of all categories. First, to learn robust and generalized features of different categories, we propose an adaptive geometry-aware masked attention module that extracts geometry variation information to guide mask attention. Then, we introduce a local geometry-aware encoder reinforced by the improved mask attention to encode group-level feature tokens. Finally, we design a global query decoder that utilizes point cloud position embeddings to improve the decoding process and reconstruction ability. This leads to local and global geometry-aware reconstructed feature tokens for the 3D AD task. MC3D-AD is evaluated on two publicly available Real3D-AD and Anomaly-ShapeNet datasets, and exhibits significant superiority over current state-of-the-art single-category methods, achieving 3.1% and 9.3% improvement in object-level AUROC over Real3D-AD and Anomaly-ShapeNet, respectively. The code is available at https://github.com/iCAN-SZU/MC3D-AD.
Jiayi Cheng, Can Gao, Jie Zhou 0009, Jiajun Wen 0001, Jinbao Wang 0001
IJCAI4
2025 Self-inferring incomplete multi-view clustering
abstract
Abstract With the advantage of exploiting complementary and consensus information across multiple views, techniques for Multi‐view Clustering have attracted increasing attention in recent years. However, it is common that data on some views is not completed in real‐world applications, which brings the challenge of partial mapping between the views. To explore the information hidden in the local geometric structure and recover missing instances through mining the information hidden in existing instances, a self‐inferring incomplete multi‐view clustering algorithm is proposed. Firstly, the incomplete multi‐view data is replenished directly and exploited as variables for inferring the missing instances. And then, a feature graph constraint is united in consensus learning. Besides, a similarity graph learning method is imposed to preserve the local manifold structure. At last, the inferred instances are filled in the missing instances for learning better consensus representation in the iterative process. Extensive experiment results show that this method can improve the clustering performance compared with the state‐of‐the‐art methods.
Junjun Fan, Zeqi Ma, Jiajun Wen 0001, Zhihui Lai 0001, Weicheng Xie 0001, Wai Keung Wong
IET Comput. Vis.3
2025 Characteristic discriminative prototype network with detailed interpretation for classification
Jiajun Wen 0001, Heng Kong, Zhihui Lai 0001, Zhijie Zhu
Pattern Recognit.1
2025 READ3D-Net: Residual Autoencoder and GAN-Based 3-D Convolutional Network for Anomaly Detection
abstract
Video anomaly detection (VAD) is of great importance for a variety of real-time applications in video surveillance. Most deep learning-based anomaly detection algorithms adopt a one-class learning scheme to train a classifier using only normal data to distinguish between normal and abnormal events during the test phase. However, these methods, whether they are reconstruction or prediction models, commonly face the challenge of the model’s overly strong representation capability, which leads to excessive fitting of abnormal events and thus limits the performance of the model in diverse scenarios. To address these challenges, this work develops a novel residual autoencoder and generative adversarial network-based 3-D convolutional network, called READ3D-net, for anomaly detection. An adaptive multimodal pseudoanomaly generator is developed to simulate and generate diverse pseudoanomalies, aiming to enhance the model’s ability to extract the features with regard to “abnormal” behaviors while reducing the interference of background on the detection performance. In addition, residual structures are incorporated into the design of a reconstruction-based autoencoder model to enhance its feature extraction and discriminative capabilities. To further improve the model’s reconstruction ability on normal data, a dynamic generative adversarial strategy is proposed for effective feature learning. Extensive experiments conducted on public benchmark datasets demonstrate that the proposed model is more competitive than the state-of-the-art methods in video anomaly detection tasks, fully validating the effectiveness and practicality of the proposed approach.
Yuwu Lu, Yinsheng Liu, Jiajun Wen 0001, Yang Zhang 0012, Yingyi Liang, Zhihui Lai 0001, LinLin Shen
IEEE Trans. Ind. Informatics3
2024 Robust weighted fuzzy margin-based feature selection with three-way decision
Zhenxi Chen, Can Gao, Jie Zhou 0009, Jiajun Wen 0001
Int. J. Approx. Reason.5
2024 LPRR: Locality preserving robust regression based jointly sparse feature extraction
Jiajun Wen 0001, Zhihui Lai 0001, Jie Zhou 0009, Heng Kong
Inf. Sci.2
2024 Multi-view robust regression for feature extraction
Zhihui Lai 0001, Foping Chen, Jiajun Wen 0001
Pattern Recognit.3
2024 Heterogeneous domain adaptation via incremental discriminative knowledge consistency
Yuwu Lu, Dewei Lin, Jiajun Wen 0001, LinLin Shen, Xuelong Li 0001, Zhenkun Wen
Pattern Recognit.3
2024 Correlation-Guided Distribution and Geometry Alignments for Heterogeneous Domain Adaptation
abstract
In this paper, we present a novel approach named correlation-guided distribution and geometry alignments (CDGA) for heterogeneous domain adaptation. Unlike existing methods that typically combine feature alignment and domain alignment into a single objective function, our proposed CDGA separates the two alignments into distinct steps. The two adaptation steps are: paired canonical correlation analysis (PCCA) and distribution and geometry alignments (DGA). In the PCCA step, CDGA focuses on maximizing the within-category correlation between source and target samples to produce the dimension-aligned feature representations for the next adaptation step. In the DGA step, CDGA is responsible for learning a classifier that incorporates both distribution and geometry alignments. Furthermore, during this step, the highly confident pseudo labeled samples are carefully selected for the next iteration of PCCA, establishing a beneficial coupling between PCCA and DGA to improve the adaptation performance in an iterative manner. Experimental results on various visual cross-domain benchmarks demonstrate that CDGA achieves remarkable performance compared to the existing shallow heterogeneous domain adaptation methods and even exhibits superiority over the state-of-the-art neural network-based approaches.
Wai Keung Wong, Dewei Lin, Yuwu Lu, Jiajun Wen 0001, Zhihui Lai 0001, Xuelong Li 0001
IEEE Trans. Multim.4
2023 Multi Task-Based Facial Expression Synthesis with Supervision Learning and Feature Disentanglement of Image Style
abstract
Image-to-Image synthesis paradigms have been widely used for facial expression synthesis. However, current generators are apt to either produce artifacts for largely posed and non-aligned faces or unduly change the identity information like AdaIN-based generator. In this work, we suggest to use image style feature to surrogate the expression cues in the generator, and propose a multi-task learning paradigm to explore this style information via the supervision learning and feature disentanglement. While the supervision learning can make the encoded style specifically represent the expression cues and enable the generator to produce correct expression, the feature disentanglement of content and style cues enables the generator to better preserve the identity information in expression synthesis. Experimental results show that the proposed algorithm can well reduce the artifacts for the synthesis of posed and non-aligned expressions, and achieves competitive performances in terms of FID, PNSR and classification accuracy, compared with four publicly available GANs. The code and pre-trained models are available at https://github.com/lumanxi236/MTSS.
Wenya Lu, Zhibin Peng, Weicheng Xie 0001, Jiajun Wen 0001, Zhihui Lai 0001, LinLin Shen
ICIP5
2023 CATrack: Convolution and Attention Feature Fusion for Visual Object Tracking
Longkun Zhang, Jiajun Wen 0001, Zichen Dai, Rouyi Zhou, Zhihui Lai 0001
PRCV (9)2
2023 Three-way decision-based co-detection for outliers
Xiaofeng Tan 0001, Can Gao, Jie Zhou 0009, Jiajun Wen 0001
Int. J. Approx. Reason.4
2023 Enhanced robust spatial feature selection and correlation filter learning for UAV tracking
Jiajun Wen 0001, Hong-Lin Chu, Zhihui Lai 0001, Tianyang Xu 0001, LinLin Shen
Neural Networks1
2023 Multiview Jointly Sparse Discriminant Common Subspace Learning
Zhihui Lai 0001, Jie Zhou 0009, Jiajun Wen 0001, Heng Kong
Pattern Recognit.4
2023 Cross-domain structure learning for visual data recognition
Yuwu Lu, Xingping Luo, Jiajun Wen 0001, Zhihui Lai 0001, Xuelong Li 0001
Pattern Recognit.3
2023 Maximal Margin Support Vector Machine for Feature Representation and Classification
abstract
High-dimensional small sample size data, which may lead to singularity in computation, are becoming increasingly common in the field of pattern recognition. Moreover, it is still an open problem how to extract the most suitable low-dimensional features for the support vector machine (SVM) and simultaneously avoid singularity so as to enhance the SVM's performance. To address these problems, this article designs a novel framework that integrates the discriminative feature extraction and sparse feature selection into the support vector framework to make full use of the classifiers' characteristics to find the optimal/maximal classification margin. As such, the extracted low-dimensional features from high-dimensional data are more suitable for SVM to obtain good performance. Thus, a novel algorithm, called the maximal margin SVM (MSVM), is proposed to achieve this goal. An alternatively iterative learning strategy is adopted in MSVM to learn the optimal discriminative sparse subspace and the corresponding support vectors. The mechanism and the essence of the designed MSVM are revealed. The computational complexity and convergence are also analyzed and validated. Experimental results on some well-known databases (including breastmnist, pneumoniamnist, colon-cancer, etc.) show the great potential of MSVM against classical discriminant analysis methods and SVM-related methods, and the codes can be available on https://www.scholat.com/laizhihui.
Zhihui Lai 0001, Xi Chen 0096, Junhong Zhang, Heng Kong, Jiajun Wen 0001
IEEE Trans. Cybern.5
2022 Uncertainty-Guided Pixel Contrastive Learning for Semi-Supervised Medical Image Segmentation
abstract
Recently, contrastive learning has shown great potential in medical image segmentation. Due to the lack of expert annotations, however, it is challenging to apply contrastive learning in semi-supervised scenes. To solve this problem, we propose a novel uncertainty-guided pixel contrastive learning method for semi-supervised medical image segmentation. Specifically, we construct an uncertainty map for each unlabeled image and then remove the uncertainty region in the uncertainty map to reduce the possibility of noise sampling. The uncertainty map is determined by a well-designed consistency learning mechanism, which generates comprehensive predictions for unlabeled data by encouraging consistent network outputs from two different decoders. In addition, we suggest that the effective global representations learned by an image encoder should be equivariant to different geometric transformations. To this end, we construct an equivariant contrastive loss to strengthen global representation learning ability of the encoder. Extensive experiments conducted on popular medical image benchmarks demonstrate that the proposed method achieves better segmentation performance than the state-of-the-art methods.
Jianglin Lu, Zhihui Lai 0001, Jiajun Wen 0001, Heng Kong
IJCAI4
2022 Correlated Matching and Structure Learning for Unsupervised Domain Adaptation
Xingping Luo, Yuwu Lu, Jiajun Wen 0001, Zhihui Lai 0001
PRCV (1)3
2022 Ghost shuffle lightweight pose network with effective feature representation and learning for human pose estimation
abstract
Abstract Despite their success, existing human pose estimation approaches mostly have complex architectures, high cost, and lack of lightweight modules. To address this problem, this paper proposes a Ghost Shuffle Lightweight Pose Network (GSLPN) with a more lightweight and efficient network architecture than the popular Lightweight Pose Network. First, in order to condense the scale of the network while maintaining its performance, we stack two lightweight modules, depthwise convolution and the Ghost module, to build our initial prototype bottleneck. Then, we impose a channel shuffle operation on the prototype bottleneck to shuffle the sequence of the feature maps for constructing Ghost Shuffle Bottleneck (Ghost Shuffle Bottleneck) with effective feature representation so as to develop a GSLPN. Second, a lightweight, efficient parallel attention mechanism, Lightweight Pose Parallel Attention, is proposed to improve keypoint locating accuracy. An experiment validating the proposed method showed that GSLPN achieved competitive performance with a smaller model size and less computational complexity than state‐of‐the‐art methods, indicating that the GSLPN is a superior approach for human pose estimation.
Senquan Yang, Jiajun Wen 0001, Junjun Fan
IET Comput. Vis.2
2022 Deep asymmetric hashing with dual semantic regression and class structure quantization
Jianglin Lu, Jie Zhou 0009, Mengfan Yan, Jiajun Wen 0001
Inf. Sci.5
2021 Three-way decision with co-training for partially labeled data
Can Gao, Jie Zhou 0009, Duoqian Miao 0001, Jiajun Wen 0001, Xiaodong Yue 0002
Inf. Sci.4
2019 Granular maximum decision entropy-based monotonic uncertainty measure for attribute reduction
Can Gao, Zhihui Lai 0001, Jie Zhou 0009, Jiajun Wen 0001, Wai Keung Wong
Int. J. Approx. Reason.4
2019 A general moving detection method using dual-target nonparametric background model
Zuofeng Zhong, Jiajun Wen 0001, Bob Zhang 0001, Yong Xu 0001
Knowl. Based Syst.2
2019 Binary sparse signal recovery algorithms based on logic observation
Xiao-Li Hu, Jiajun Wen 0001, Zhihui Lai 0001, Wai Keung Wong, LinLin Shen
Pattern Recognit.2
2019 Generalized Robust Regression for Jointly Sparse Subspace Learning
abstract
Ridge regression is widely used in multiple variable data analysis. However, in very high-dimensional cases such as image feature extraction and recognition, conventional ridge regression or its extensions have the small-class problem, that is, the number of the projections obtained by ridge regression is limited by the number of the classes. In this paper, we proposed a novel method called generalized robust regression (GRR) for jointly sparse subspace learning which can address the problem. GRR not only imposes L2,1-norm penalty on both loss function and regularization term to guarantee the joint sparsity and the robustness to outliers for effective feature selection, but also utilizes L2,1-norm as the measurement to take the intrinsic local geometric structure of the data into consideration to improve the performance. Moreover, by incorporating the elastic factor on the loss function, GRR can enhance the robustness to obtain more projections for feature selection or classification. To obtain the optimal solution of GRR, an iterative algorithm was proposed and the convergence was also proved. Experiments on six wellknown data sets demonstrate the merits of the proposed method. The result indicates that GRR is a robust and efficient regression method for face recognition.
Zhihui Lai 0001, Dongmei Mo, Jiajun Wen 0001, LinLin Shen, Wai Keung Wong
IEEE Trans. Circuits Syst. Video Technol.3
2018 Robust jointly sparse embedding for dimensionality reduction
Zhihui Lai 0001, Yudong Chen 0002, Dongmei Mo, Jiajun Wen 0001, Heng Kong
Neurocomputing4
2018 On uniqueness of sparse signal recovery
Xiao-Li Hu, Jiajun Wen 0001, Wai Keung Wong, Le Tong, Jinrong Cui
Signal Process.2
2017 Directional Gaussian Model for Automatic Speeding Event Detection
abstract
This paper proposes a velocity learning method based on the directional Gaussian model to detect speeding events in surveillance scenarios. The proposed method is of an uncalibrated type, but yet has considered the influences of projective transformation on estimating the motion velocities in an image plane, which is convenient and feasible to use in real application. We have analyzed the theory that the velocity in the image plane varies due to the changes of the direction of the moving object with constant velocity in a real world plane. With the support of this theory, we propose to learn the velocities calculated on a certain position in different direction bins to tolerate the effect of projective transformation on velocity modeling. To facilitate the whole framework, two key issues have to be addressed. First, we have designed an improved Fisher model to optimize the direction bins, which reflect the major moving directions in a scenario. Second, we have adopted a weighted sampling strategy and surface fitting to solve the lack of sample problem during the learning process. Experiments conducted on real surveillance videos show that the proposed method can obtain competitive results compared with the state-of-the-art methods.
Jiajun Wen 0001, Zhihui Lai 0001, Zhong Ming 0001, Wai Keung Wong, Zuofeng Zhong
IEEE Trans. Inf. Forensics Secur.1
2017 Low-Rank Embedding for Robust Image Feature Extraction
abstract
Robustness to noises, outliers, and corruptions is an important issue in linear dimensionality reduction. Since the sample-specific corruptions and outliers exist, the class-special structure or the local geometric structure is destroyed, and thus, many existing methods, including the popular manifold learning- based linear dimensionality methods, fail to achieve good performance in recognition tasks. In this paper, we focus on the unsupervised robust linear dimensionality reduction on corrupted data by introducing the robust low-rank representation (LRR). Thus, a robust linear dimensionality reduction technique termed low-rank embedding (LRE) is proposed in this paper, which provides a robust image representation to uncover the potential relationship among the images to reduce the negative influence from the occlusion and corruption so as to enhance the algorithm's robustness in image feature extraction. LRE searches the optimal LRR and optimal subspace simultaneously. The model of LRE can be solved by alternatively iterating the argument Lagrangian multiplier method and the eigendecomposition. The theoretical analysis, including convergence analysis and computational complexity, of the algorithms is presented. Experiments on some well-known databases with different corruptions show that LRE is superior to the previous methods of feature extraction, and therefore, it indicates the robustness of the proposed method. The code of this paper can be downloaded from http://www.scholat.com/laizhihui.
Wai Keung Wong, Zhihui Lai 0001, Jiajun Wen 0001, Xiaozhao Fang, Yuwu Lu
IEEE Trans. Image Process.3
2016 The L2, 1-norm-based unsupervised optimal feature selection with applications to action recognition
Jiajun Wen 0001, Zhihui Lai 0001, Yinwei Zhan, Jinrong Cui
Pattern Recognit.1
2015 Discriminant non-negative graph embedding for face recognition
Jinrong Cui, Jiajun Wen 0001, Li Bin
Neurocomputing2
2015 Appearance-based bidirectional representation for palmprint recognition
Jinrong Cui, Jiajun Wen 0001, Zizhu Fan
Multim. Tools Appl.2
2015 Erratum to: Appearance-based bidirectional representation for palmprint recognition
Jinrong Cui, Jiajun Wen 0001, Zizhu Fan
Multim. Tools Appl.2
2015 Joint Tensor Feature Analysis For Visual Object Recognition
abstract
Tensor-based object recognition has been widely studied in the past several years. This paper focuses on the issue of joint feature selection from the tensor data and proposes a novel method called joint tensor feature analysis (JTFA) for tensor feature extraction and recognition. In order to obtain a set of jointly sparse projections for tensor feature extraction, we define the modified within-class tensor scatter value and the modified between-class tensor scatter value for regression. The k-mode optimization technique and the L(2,1)-norm jointly sparse regression are combined together to compute the optimal solutions. The convergent analysis, computational complexity analysis and the essence of the proposed method/model are also presented. It is interesting to show that the proposed method is very similar to singular value decomposition on the scatter matrix but with sparsity constraint on the right singular value matrix or eigen-decomposition on the scatter matrix with sparse manner. Experimental results on some tensor datasets indicate that JTFA outperforms some well-known tensor feature extraction and selection algorithms.
Wai Keung Wong, Zhihui Lai 0001, Yong Xu 0001, Jiajun Wen 0001, Chu Po Ho
IEEE Trans. Cybern.4
2014 Optimal Feature Selection for Robust Classification via l2, 1-Norms Regularization
abstract
This paper aims to explore the optimal feature selection with dimensionality reduction and jointly sparse representation scheme for classification. The proposed method is called Optimal Feature Selection Classification (OFSC). Our model simultaneously learns an orthogonal subspace for jointly sparse feature selection and representation via l2,1-norms regularization. To solve the proposed model, an alternately iterative algorithm is proposed to optimize both the jointly sparse projection matrix and representation matrix. Experimental results on three public face datasets and one action dataset validate the quick convergence of our algorithm and show that the proposed method is more competitive than the state-of-the-art methods.
Jiajun Wen 0001, Zhihui Lai 0001, Wai Keung Wong, Jinrong Cui, Minghua Wan
ICPR1
2014 Joint Video Frame Set Division and Low-Rank Decomposition for Background Subtraction
abstract
The recently proposed robust principle component analysis (RPCA) has been successfully applied in background subtraction. However, low-rank decomposition makes sense on the condition that the foreground pixels (sparsity patterns) are uniformly located at the scene, which is not realistic in real-world applications. To overcome this limitation, we reconstruct the input video frames and aim to make the foreground pixels not only sparse in space but also sparse in time. Therefore, we propose a joint video frame set division and RPCA-based method for background subtraction. In addition, we use the motion as a priori knowledge which has not been considered in the current subspace-based methods. The proposed method consists of two phases. In the first phase, we propose a lower bound-based within-class maximum division method to divide the video frame set into several subsets. In this way, the successive frames are assigned to different subsets in which the foregrounds are located at the scene randomly. In the second phase, we augment each subset using the frames with a small quantity of motion. To evaluate the proposed method, the experiments are conducted on real-world and public datasets. The comparisons with the state-of-the-art background subtraction methods validate the superiority of our method.
Jiajun Wen 0001, Yong Xu 0001, Jinhui Tang 0001, Yinwei Zhan, Zhihui Lai 0001, Xiao-Tang Guo
IEEE Trans. Circuits Syst. Video Technol.1