VLDB 2026 Research / reviewers in the wild / expert
Jun Wang 0071
dblp:125/8189-71
· DBLP profile ↗
32ranked-venue papers
4as first author
28since 2021 · last 2026
0000-0002-3676-1958ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 15 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Security and privacy · 9 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Cooperative Formation Navigation of Multiple Unmanned Ground Vehicles in Unknown and Complex EnvironmentsabstractThe cooperative formation navigation of multiple unmanned ground vehicles (UGVs) is critical for enhancing the efficiency and safety of intelligent transportation and industrial internet of things systems. However, existing methods remain limited in unknown and complex environments due to insufficient adaptability, a lack of effective path recovery mechanisms, and inflexible formation adjustment, which may lead to motion failure. To address these challenges, this paper proposes an adaptive cooperative formation navigation method based on dynamic tree structure (ACFN-DTree), aiming to achieve feasible path exploration, dynamic decision recovery, and formation self-adaptation. First, a dynamic tree is constructed in obstacle-free regions based on real-time perception, preserving critical environmental information while identifying feasible directions. Then, sub-goals are selected via a spatial cost function and backtracking mechanism, guiding UGVs toward the global goal or returning to feasible branches when deadlocks occur. Furthermore, the formation compression coefficient is optimized based on the effective formation width, allowing for adaptive formation configuration adjustments. Extensive numerical and virtual simulation experiments demonstrate that ACFN-DTree exhibits significant advantages in terms of path length, formation maintenance, and path smoothness level. These results verify the effectiveness of the proposed method in enabling multiple UGVs to achieve safe navigation and adaptive formation adjustment in unknown and complex environments. Jun Wang 0071, Baolei Wu, Miaoying Hong, Ning Bian, Shaoze You, Yongqiang Qi |
IEEE Internet Things J. | 2 |
| 2026 | MsMemoryGAN: A Multiscale Memory GAN for Palm-Vein Adversarial PurificationabstractDeep neural networks have recently achieved promising performance in the vein recognition task and have shown an increasing application trend. However, they are prone to adversarial attacks by adding imperceptible perturbations to the input, resulting in incorrect recognition. To address this issue, we propose a novel defense model named MsMemoryGAN, which aims to filter the perturbations from adversarial samples before recognition. First, we design a multiscale memory autoencoder (MsMemoryAE) to achieve high-quality reconstruction, where the memory module (MM) within it is capable of learning the detailed patterns of normal samples at different scales. Second, to overcome the limitations of handcrafted similarity metrics, we propose an MM with learnable similarity (LSMM), which retrieves the most relevant memory items to purify the input feature. Finally, the perceptual loss and adversarial loss are integrated with the pixel loss to further enhance the quality of the reconstructed image. During the training phase, the MsMemoryGAN learns to reconstruct the input by merely using fewer prototypical elements of the normal patterns recorded in the memory. At the testing stage, given an adversarial sample, the MsMemoryGAN retrieves its most relevant normal patterns in MMs for reconstruction. Perturbations in the adversarial sample are usually not reconstructed well, resulting in adversarial purification. We conduct extensive experiments on two public vein datasets under different adversarial attack methods to evaluate the performance of the proposed approach. The experimental results show that our approach removes a wide variety of adversarial perturbations, allowing vein classifiers to achieve the highest recognition accuracy. Huafeng Qin, Yuming Fu 0001, Huiyan Zhang 0001, Mounim A. El-Yacoubi, Xinbo Gao 0001, Qun Song 0007, Jun Wang 0071 |
IEEE Trans. Cybern. | 7 |
| 2026 | Learning Suspected Anomalies from Event Prompts for Video Anomaly DetectionabstractMost models for Weakly Supervised Video Anomaly Detection (WS-VAD) rely on multiple instance learning, aiming to distinguish normal and abnormal snippets without specifying the type of anomaly. However, the ambiguous nature of anomaly definitions across contexts may introduce inaccuracy in discriminating abnormal and normal events. To show the model what is anomalous, a novel framework is proposed to guide the learning of suspected anomalies from event prompts. Given a textual prompt dictionary of potential anomaly events and the captions generated from anomaly videos, the semantic anomaly similarity between them could be calculated to identify the suspected events for each video snippet. It enables a new multi-prompt learning process to constrain the visual-semantic features across all videos, as well as provides a new way to label pseudo anomalies for self-training. To demonstrate its effectiveness, comprehensive experiments and detailed ablation studies are conducted on four datasets, namely XD-Violence, UCF-Crime, TAD, and ShanghaiTech. Our proposed model outperforms most state-of-the-art methods in terms of AP or AUC (86.5%, 90.4%, 94.4%, and 97.4%). Furthermore, it shows promising performance in open-set and cross-dataset cases. The data, code, and models can be found at: https://github.com/shiwoaz/lap . Chenchen Tao, Xiaohao Peng, Chong Wang 0001, Jiafei Wu, Puning Zhao, Jun Wang 0071, Jiangbo Qian |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Rethinking Early-Fusion Strategies for Improved Multimodal Image SegmentationabstractRGB and thermal image fusion have great potential to exhibit improved semantic segmentation in low-illumination conditions. Existing methods typically employ a two-branch encoder framework for multimodal feature extraction and design complicated feature fusion strategies to achieve feature extraction and fusion for multimodal semantic segmentation. However, these methods require massive parameter updates and computational effort during the feature extraction and fusion. To address this issue, we propose a novel multimodal fusion network (EFNet) based on an early fusion strategy and a simple but effective feature clustering for training efficient RGB-T semantic segmentation. In addition, we also propose a lightweight and efficient multi-scale feature aggregation decoder based on Euclidean distance. We validate the effectiveness of our method on different datasets and outperform previous state-of-the-art methods with lower parameters and computation. Zhengwen Shen, Yulian Li, Yuchen Weng, Jun Wang 0071 |
ICASSP | 5 |
| 2025 | A Mutual Distillation Learning Framework for Multimodal Biometric Recognition with Uncertain Missing ModalityabstractCurrently, multimodal biometric recognition technology is receiving increasing attention. Most existing multimodal biometric recognition techniques require complete multimodal data during both the training and testing phases. However, due to various limitations, it is challenging to obtain complete and high-quality multimodal biometric data. To address this problem, we proposed a Mutual Distillation Framework (MDF) for palmprint and palmvein based multimodal biometric recognition with uncertain missing modality. Specifically, we firstly design an intra-inter modal data missing augmentation mechanism to generate heterogeneous missing samples. Moreover, a mutual knowledge distillation module is introduced to establish correlation of beneficial semantic information across different missing scenarios. In particular, we incorporate a Sample Semantic Discrepancy Guided mechanism within the mutual distillation framework, which calculates the semantic discrepancy between samples under corresponding heterogeneous missing patterns to encourage the model to focus on samples containing more complementary information. Experimental results demonstrate that the proposed model outperforms state-of-the-art incomplete multimodal learning models across three multimodal biometric benchmark datasets. Shuangtian Jiang, Hai Yuan, Jun Wang 0071, Zaiyu Pan |
IJCB | 4 |
| 2025 | LMBR-Net: A Lightweight Multimodal Biometric Recognition Network via Joint Progressive Dynamic SparsityabstractMultimodal biometric recognition technology has shown great potential in enhancing both recognition accuracy and robustness. However, challenges persist in controlling network parameter size and preserving fine-grained features. While existing lightweight methods excel in unimodal tasks, they often struggle to effectively balance parameter scale with fusion performance in multimodal scenarios. In this paper, we propose a lightweight multimodal recognition network based on joint progressive dynamic sparsity (LMBR-Net). Our method introduces the Joint Progressive Dynamic Sparsity Module (JPDSM), which dynamically reduces parameter redundancy while preserving complementary information across modalities. Additionally, the Lightweight Collaborative Sparse Feature Fusion Module (LCSFM) effectively captures global channel information, improving inter-modality fusion and ensuring consistency in the feature space while retaining fine-grained details across modalities. Experimental results demonstrate that the proposed network achieves superior recognition accuracy on multiple public datasets, maintaining a parameter count of approximately 3.64M, thus achieving an effective balance between network complexity and recognition performance. Jun Wang 0071, Penghao Jin, Zaiyu Pan |
IJCB | 1 |
| 2025 | Rethinking Early-Fusion Strategy for Palmprint and Palm Vein Fusion RecognitionabstractMultimodal biometric recognition has shown great potential in identity authentication tasks. Currently, most existing multimodal biometric recognition methods mainly employ multiple separate feature extraction networks for multimodal biometrics, leading to nearly multiple times the inference time compared to a single-stream feature extraction network. This increased inference time has hindered the widespread employment of multimodal biometric recognition in embedded devices for autonomous systems. To this end, we rethink the early-fusion strategy for palmprint and palm vein fusion recognition. First, an Image-Level Adaptive Fusion module (ILAF) is proposed to achieve the dynamic fusion of palmprint and palm vein images in pixel space, and then the fused images are inputted into the single-stream feature extraction network for identity recognition. Second, Modality-Aware Knowledge Distillation (MAKD) is presented, which includes PalmVein-Aware Knowledge Distillation (PVKD) and PalmPrint-Aware Knowledge Distillation (PPKD). This approach enables the single-stream feature extraction network to learn more modality-specific information of different modalities, thereby enhancing the discriminative ability of multimodal fusion representation. Experimental results on four multimodal biometric benchmark datasets demonstrate that our proposed model outperforms state-of-the-art multimodal biometric recognition models. Penghao Jin, Jun Wang 0071, Zaiyu Pan |
IJCB | 3 |
| 2025 | Pixel-Level Adaptive Refinement Framework with Knowledge Distillation for Weakly Supervised Semantic SegmentationabstractClass activation maps (CAMs) generated from image-level labels often yield incomplete object activation and inaccurate boundaries, posing challenges for semantic segmentation networks dependent on image-level supervision. To address this, we propose a weakly supervised semantic segmentation method based on pixel-level adaptive refinement within a knowledge distillation framework. Specifically, we designed a pixel-level adaptive refinement module that iteratively updates the CAMs through an adaptive refining kernel, enhancing the local consistency of the activation and producing more complete and accurate semantic masks. Additionally, we developed a local assisted enhancement network that transmits local detail information of channel and spatial to the global CAMs via knowledge distillation while cleverly mitigating the interference of background noise, achieving more refined boundaries. The model using the segmentation labels produced by our method achieved superior performance on the PASCAL VOC 2012 and MS COCO 2014 datasets and is even competitive with models that rely on stronger supervision. Yulian Li, Xinfang Qin, Zhengwen Shen, Shuyu Han, Jun Wang 0071 |
ICME | 5 |
| 2025 | Dynamic interaction and router selection network for multi-modality biometric recognition
Hai Yuan, Zaiyu Pan, Zhengwen Shen, Jun Wang 0071 |
Knowl. Based Syst. | 6 |
| 2025 | Hierarchical Cross-Modal Image Generation for Multimodal Biometric Recognition With Missing ModalityabstractMultimodal biometric recognition has shown great potential in identity authentication tasks and has attracted increasing interest recently. Currently, most existing multimodal biometric recognition algorithms require test samples with complete multimodal data. However, it often encounters the problem of missing modality data and thus suffers severe performance degradation in practical scenarios. To this end, we proposed a hierarchical cross-modal image generation for palmprint and palmvein based multimodal biometric recognition with missing modality. First, a hierarchical cross-modal image generation model is designed to achieve the pixel alignment of different modalities and reconstruct the image information of missing modality. Specifically, a cross-modal texture transfer network is utilized to implement the texture style transformation between different modalities, and then a cross-modal structure generation network is proposed to establish the correlation mapping of structural information between different modalities. Second, multimodal dynamic sparse feature fusion model is presented to obtain more discriminative and reliable representations, which can also enhance the robustness of our proposed model to dynamic changes in image quality of different modalities. The proposed model is evaluated on three multimodal biometric benchmark datasets, and experimental results demonstrate that our proposed model outperforms recent mainstream incomplete multimodal learning models. Zaiyu Pan, Shuangtian Jiang, Hai Yuan, Jun Wang 0071 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Cooperative Control of the UGV Formation in Complex and Dynamic EnvironmentsabstractThe cooperative control of the unmanned ground vehicle (UGV) formation is crucial for improving the efficiency and safety of intelligent transportation systems. Considering their complex and dynamic operating environments, the scene-adaptive switching control algorithm is proposed to address this challenge. Firstly, for the obstacle avoidance of a single UGV within a formation, a dynamic window approach that can adjust the weight of the formation function in real time is presented. By minimizing the deviation between the desired formation position and the actual position, a single UGV can quickly recover the original formation after completing local obstacle avoidance. For the cooperative obstacle avoidance of multiple UGVs, a multi-layer hierarchical obstacle avoidance strategy is designed, which is based on the constraints such as obstacle threat level and density. Furthermore, an improved consensus algorithm is used to achieve rapid reconfiguration of the formation, and a target factor is introduced into the leader-follower algorithm to ensure that the UGVs reach the target position smoothly. Finally, the simulation and experimental results show that the multi-vehicle system can dynamically avoid obstacles, quickly maintain formation, and reconfigure formation. Miaoying Hong, Baolei Wu, Yongqiang Qi, Jun Wang 0071 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Synergizing Global and Local Knowledge via Dynamic Focus Mechanism for Low-Light Image Enhancement
Shuyu Han, Zhengwen Shen, Yulian Li, Zaiyu Pan, Jun Wang 0071 |
PRCV (9) | 5 |
| 2024 | HEFANet: hierarchical efficient fusion and aggregation segmentation network for enhanced rgb-thermal urban scene parsing
Zhengwen Shen, Zaiyu Pan, Yuchen Weng, Yulian Li, Jiangyu Wang, Jun Wang 0071 |
Appl. Intell. | 6 |
| 2024 | A multi-stage LSTM federated forecasting method for multi-loads under multi-time scales
Xianfang Song, Jun Wang 0071, Yong Zhang 0016, Xiaoyan Sun 0002 |
Expert Syst. Appl. | 3 |
| 2024 | Adversarial Learning-Based Data Augmentation for Palm-Vein IdentificationabstractPalm-vein identification is a highly secure pattern biometrics that has become an active research area in recent years. Despite the recent progress in deep neural networks (DNNs) for vein identification, existing solutions for feature representation continue to lack robustness due to the limited training samples. To address this limitation, data augmentation approaches, including Generative Adversarial Networks (GANs), have been investigated, but these schemes suffer from the following issues. First, it is practically unfeasible to use all the generated samples for classifier training due to the limited storage space and computation resources. Further, some of these generated samples may be non-representative or ineffective, seriously compromising models’ generalization capabilities. Second, the augmented dataset is fed to the target classifier repeatedly, resulting in overfitting after substantial training epochs. To tackle the above problems, we propose AdveinAU, an Adversarial vein AUtomatic AUgmentation approach that generates challenging samples to train a more robust vein classifier for palm-vein identification by alternatively optimizing the vein classifier and a set of latent variables. First, we consider a conditional deep convolution generative adversarial net (cDCGAN) to learn the distribution of real data and the generated data, and then a latent variable from the latent variable space is mapped to the sample space. Second, we combine the trained generator with the vein classifier to constitute AdveinAU, where the input sets of the generator and the classifier are alternatively updated by adversarial training. Specifically, a latent variable set is learned to increase the training loss of a target network through generating adversarial samples, while the classifier learns more robust features from harder examples to improve the generalization. To avoid collapsing inherent meanings of images, an exponential moving average (EMA) teacher andcosinesimilarity are employed for regularization to reduce the search space. Unlike previous works where GANs synthesize new realistic images, our model aims to search a latent variable set, based on which the generator can produce challenging samples along with the training process to improve the classifier’s performance. Finally, we conduct extensive experiments on three public palm-vein datasets to evaluate the performance of AdveinAU, and the experimental results demonstrate that the proposed AdveinAU is capable of generating harder samples to improve the performance of the vein classifier. Huafeng Qin, Haofei Xi, Yantao Li 0001, Mounim A. El-Yacoubi, Jun Wang 0071, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Attention BLSTM-Based Temporal-Spatial Vein Transformer for Multi-View Finger-Vein RecognitionabstractFinger-vein biometrics has recently gained significant attention due to its robust privacy and high security features. Despite notable advancements, most existing methods focus on extracting features from a 2-dimensional (2D) image projected from 3D vein vessels with a single view. However, recognition based on a single view is prone to errors due to variations in finger positioning, especially those caused by finger roll movements, which can degrade recognition performance. To address this challenge, we propose ABLSTM-TSVT, an Attention Bidirectional LSTM-based Temporal-Spatial Vein Transformer for multi-view finger-vein recognition. First, we enhance LSTM with an attention mechanism to create an attention LSTM for extracting temporal features. We further improve this by introducing a local attention module, which learns temporal dependencies between a patch (token) and its adjacent patches across multiple views, integrating it with the attention LSTM to form a temporal attention module. Second, we develop a spatial attention module that captures the spatial dependencies of patches within an image. Finally, merging the temporal and the spatial attention modules, we create our temporal-spatial transformer model, which effectively represents features from multi-view images. Experimental results on two multi-view datasets demonstrate that our approach outperforms state-of-the-art approaches in enhancing identification accuracy and reducing verification errors in vein classifiers. Huafeng Qin, Zhipeng Xiong, Yantao Li 0001, Mounim A. El-Yacoubi, Jun Wang 0071 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Enlightening the Student in Knowledge DistillationabstractKnowledge distillation is a common method of model compression, which uses large models (teacher networks) to guide the training of small models (student networks). However, the student may find a hard time absorbing the knowledge from a sophisticated teacher due to the capacity and confidence gaps between them. To address this issue, a new knowledge distillation and refinement (KDrefine) framework is proposed to enlighten the student by expending and refining its network structure. In addition, a confidence refinement strategy is utilized to generate adaptive soften logits for efficient distillation. The experiments show that the proposed framework outperforms state-of-the-art methods on both CIFAR-100 and Tiny-ImageNet datasets. The code is available at https://github.com/YujieZheng99/KDrefine. Chong Wang 0001, Yi Chen 0001, Jiangbo Qian, Jun Wang 0071, Jiafei Wu |
ICASSP | 5 |
| 2023 | Synthetic Feature Assessment for Zero-Shot Object DetectionabstractZero-shot object detection aims to simultaneously identify and localize classes that were not presented during training. Many generative model-based methods have shown promising performance by synthesizing the visual features of unseen classes from semantic embeddings. However, these synthetic features are inevitably of varied quality, which may be far from the ground truth. It degrades the performance of trained unseen classifier. Instead of tweaking the generative model, a new idea of feature quality assessment is proposed to utilize both the good and bad features to optimize the classifier in the right direction. Moreover, contrastive learning is also introduced to enhance the feature uniqueness between unseen and seen classes, which helps the feature assessment implicitly. To demonstrate the effectiveness of the proposed algorithm, comprehensive experiments are conducted on the MS COCO dataset and PASCAL VOC dataset, the state-of-the-art performance is achieved. Our code is available at: https://github.com/Dai1029/SFA-ZSD. Xinmiao Dai, Chong Wang 0001, Haohe Li, Sunqi Lin, Li Dong 0006, Jiafei Wu, Jun Wang 0071 |
ICME | 7 |
| 2023 | Cross-domain collaborative learning for single image deraining
Zaiyu Pan, Jun Wang 0071, Zhengwen Shen, Shuyu Han, Jihong Zhu 0001 |
Expert Syst. Appl. | 2 |
| 2023 | A federated feature selection algorithm based on particle swarm optimization under privacy protection
Ying Hu 0006, Yong Zhang 0016, Xiao Zhi Gao 0001, Dun-Wei Gong, Xianfang Song, Yinan Guo 0001, Jun Wang 0071 |
Knowl. Based Syst. | 7 |
| 2023 | Disentangled Representation and Enhancement Network for Vein RecognitionabstractRecently, vein recognition has been paid more considerable attention in biometric recognition fields. In the process of vein image acquisition, due to the influence of external factors such as illumination change, the texture information of vein images with the same identity information may change, which enhances the difference of intra-class and extremely degrades the performance of vein recognition systems. To address this problem, we proposed a Disentangled Representation and Enhancement Network for vein recognition called DRE-Net. First, robust vein shape masks are obtained by the designed vein segmentation algorithms, which are utilized as label information in DRE-Net to acquire the shape features of vein images. Second, a disentangled representation network which contains two encoders and two decoders is designed to disentangle the texture features and shape features of vein images. Besides, a Multi-Scale Attention Residual Block (MSARB) is presented to better mine the vein network information and enhance the representation ability of DRE-Net for vein images. Finally, a Weight-Guided Feature Enhancement Module (WGFEM) is presented to obtain more discriminative representations for vein recognition by reducing the importance of texture features and increasing the importance of shape features, which decreases the difference of intra-class caused by illumination change. Extensive experiments have been carried out on three benchmark vein databases, and the experimental results demonstrate that our proposed model outperforms the state-of-the-art methods. Zaiyu Pan, Jun Wang 0071, Zhengwen Shen, Shuyu Han |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Small-Area Finger Vein RecognitionabstractRecently, finger vein sensors have been embedded in all kinds of electronic devices for personal identification, such as intelligent door locks and attendance machines. The embedded sensors are generally small, thus capturing only part of the finger vein. However, prior studies have focused on near-full finger vein recognition, without considering the partial finger vein image caused by the small imaging window of the finger vein sensor. This paper aims to study personal identification based on partial finger vein images, known as small-area finger vein recognition. The effect of the small-area finger vein on recognition performance is first analyzed by cutting out the local part from the near-full finger vein image to model a small-area finger vein image. Second, a small-area finger vein database is built using a commercial finger vein imaging device, in which the vein pattern from approximately one-third of one adult finger is captured. To explore more discriminative information from small-area finger vein images, we propose a locality-constrained consistent dictionary learning (LCDL) method to fuse multiple features for small-area finger vein recognition. Finally, the proposed method is evaluated on the self-built small-area finger vein database and four synthetic small-area finger vein databases. Experimental results show the promising recognition performance of the proposed method. Lu Yang 0005, Gongping Yang 0001, Jun Wang 0071, Yilong Yin |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Balanced Stripe-Wise Pruning In The FilterabstractNeural network pruning offers a promising prospect to compress and accelerate modern deep convolution networks. The stripe-wise pruning method with a finer granularity than traditional methods has become the focus of research. Inspired by the previous work, a new balanced stripe-wise pruning, including the balanced pruning strategy and dynamic pruning threshold, is proposed to achieve higher performance. Specifically, the survived inter-filter stripes and the intra-filter stripes are redistributed by a balanced pruning strategy. Meanwhile the dynamic pruning threshold method makes survival rates will be further balanced across all layers. Comprehensive experiments are conducted on two public datasets (CIFAR-10 and TinyImageNet-200) for different models (ResNet and VGG). The experimental results show that the proposed model is capable of reducing the most parameters, yet achieving the highest accuracy. Our code is available at: https://github.com/ajdt1111/BSWP. Zheng Huo, Chong Wang 0001, Jun Wang 0071, Jiafei Wu |
ICASSP | 5 |
| 2022 | Novel Instance Mining with Pseudo-Margin Evaluation for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) enables the detector to recognize novel objects only using limited training samples, which could greatly alleviate model’s dependency on data. Most existing methods include two training stages, namely base training and fine-tuning. However, the unlabeled novel instances in the base set were untouched in previous works, which can be re-used to enhance the FSOD performance. Thus, a new instance mining model is proposed in this paper to excavate the novel samples from the base set. The detector is thus fine-tuned again by these additional free novel instances. Meanwhile, a novel pseudo-margin evaluation algorithm is designed to address the quality problem of pseudo-labels brought by those new novel instances. The experimental results on MS-COCO dataset show the effectiveness of the proposed model, which does not require any additional training samples or parameters. Our code is available at: https://github.com/liuweijie19980216/NimPme. Chong Wang 0001, Shenghao Yu, Chenchen Tao, Jun Wang 0071, Jiafei Wu |
ICASSP | 5 |
| 2022 | Pruning Dynamic Group Convolution with Static SubstituteabstractDeep learning networks are gradually deployed in edge applications, such as phones and cameras, which has a restriction of the computational resources. Thus, to improve the computational efficiency, numerous types of group convolution-based frameworks have been studied, including the dynamic group convolution (DGC). However, it is worth noting that sometimes the same channels are selected in different dynamic heads in DGC, which violates its original intention. It indicates the number of dynamic heads exceeds what is really needed. In this paper, a novel model by pruning dynamic group convolution with static substitute is proposed. Specifically, those channels pruned in the dynamic head can be reselected in the static head of the group convolution in the proposed model. In addition, a new polarization regularization is introduced to prune more useless channels with less accuracy loss. The experiment results on two image classification benchmarks (CIFAR-I00 and TinylmageNet-200) show promising performance. Jieyong Che, Chong Wang 0001, Xinmiao Dai, Jun Wang 0071, Jiafei Wu |
ICME | 5 |
| 2022 | CTFusion: Convolutions Integrate with Transformers for Multi-modal Image Fusion
Zhengwen Shen, Jun Wang 0071, Zaiyu Pan, Jiangyu Wang, Yulian Li |
PRCV (1) | 2 |
| 2021 | Multi-Scale Deep Representation Aggregation for Vein RecognitionabstractThe recent success of Deep Convolutional Neural Network (DCNN) for various computer vision tasks such as image recognition has already demonstrated its robust feature representation ability. However, the limitation of training database on small scale vein recognition tasks restricts its performance because the recognition result of DCNN depends heavily on the number of trainsets. This motivates the design of a Multi-Scale Deep Representation Aggregation (MSDRA) model based on a pre-trained DCNN for vein recognition. First, the multi-scale feature maps are extracted by a pre-trained DCNN model. Second, a local mean threshold approach is designed to preliminarily remove the noisy information of multi-scale feature maps and generate the selected feature maps. Third, we propose an Unsupervised Vein Information Mining (UVIM) method to localize vein information of selected feature maps for generating a binary vein information mask, and then the vein information mask is utilized to keep useful deep representation and discard the background information. Finally, the discriminative multi-scale deep representations, which are generated by using the vein information mask to aggregate multi-scale feature maps, are concatenated into the final compact feature vectors, and then a Support Vector Machine (SVM) is introduced for final recognition. Our proposed model outperforms the state-of-the-art methods on two benchmark vein databases. Moreover, an additional experiment using the subset of PolyU Palmprint database illustrates the system's generalization ability and robustness. Zaiyu Pan, Jun Wang 0071, Guoqing Wang 0001, Jihong Zhu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | A Multilayer Pyramid Network Based on Learning for Vehicle Logo RecognitionabstractIn this paper, we present a novel learning-based scheme for vehicle logo recognition (VLR). This scheme is termed Multilayer Pyramid Network Based on Learning (MLPNL) and is based on the principle that considering multiple resolutions is helpful for extracting valuable features that benefit the final recognition performance. The innovations of this scheme include (1) a multilayer pyramid network, with pixel difference matrices (PDMs) as its input and output and feature parameters mapping one PDM to another; (2) an objective function and a corresponding optimization method designed to facilitate the learning of the feature parameters of the proposed multilayer pyramid network; and (3) a multi-codebook-based encoding method that makes best use of the features extracted from PDMs corresponding to different resolutions. Extensive experiments conducted with an open dataset, HFUT-VL, demonstrate that the proposed MLPNL scheme outperforms state-of-the-art handcrafted descriptors and non-deep-learning-based learning methods when fewer training samples exist. Experiments conducted with a benchmark dataset, XMU, demonstrate that MLPNL outperforms existing state-of-the-art VLR methods. Experiments conducted both on HFUT-VL and XMU demonstrate that MLPNL is faster than most deep-learning-based learning methods while maintaining nearly the same recognition rate. Code has been made available at:https://github.com/HFUT-CV/MLPNL. Jun Wang 0071, Hai Min, Wei Jia 0001, Jun Yu 0001, Chang Wen Chen |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Bimodal Vein Data Mining via Cross-Selected-Domain Knowledge TransferabstractRecent success in large-scale image recognition challenge (i.e., ImageNet) fully demonstrates the capability of deep neural network (DNN) in learning complex and semantic representation, and this also motivates the generation of transfer learning model, which fine-tunes state-of-the-art DNN models with other small-scale databases for better performance. Driven by such an idea, a task-specific DNN model fine-tuned from VGG-face is constructed for both gender and identity recognition with hand vein information. Unlike the traditional transfer learning models, which fine-tune directly from source to target, we leverage the coarse-to-fine scheme to train the task-specific models in a step-aware way, such that the inherent correlation between the neighboring databases could serve as initialization base to relieve the problem of over-fitting, which is inevitable with the small-scaled hand vein database, and also speed up the convergence. Besides, the task-driven network training idea, which involves joint optimization of linear regression classifier and network parameters, is also adopted during training of each model to obtain more discriminative representation for specified tasks. Instead of adopting the trained linear regression classifier for gender and identity classification, the large margin distribution machine (LDM) is introduced to ensure the discriminative and generalization performance of the model simultaneously, and it should be noted that before feeding the gender feature vector into the LDM, a supervised feature selection step is incorporated to improve the classification performance by discarding the redundant feature and highlighting the important ones for gender classification. Rigorous experiments using the lab-made database are conducted to demonstrate the effectiveness and feasibility of the proposed model. What is more, additional experiment with a subset of the PolyU database illustrates its generalization ability and robustness. Jun Wang 0071, Guoqing Wang 0001, Mei Zhou |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | Quality-Specific Hand Vein Recognition SystemabstractVein images generally appear darker with low contrast, which require contrast enhancement during preprocessing to design satisfactory hand vein recognition system. However, the modification introduced by contrast enhancement (CE) is reported to bring side effects through pixel intensity distribution adjustments. Furthermore, the inevitable results of fake vein generation or information loss occur and make nearly all vein recognition systems unconvinced. In this paper, a “CE-free” quality-specific vein recognition system is proposed, and three improvements are involved. First, a high-quality lab-vein capturing device is designed to solve the problem of low contrast from the view of hardware improvement. Then, a high quality lab-made database is established. Second, CFISH score, a fast and effective measurement for vein image quality evaluation, is proposed to obtain quality index of lab-made vein images. Then, unsupervised $K$ -means with optimized initialization and convergence condition is designed with the quality index to obtain the grouping results of the database, namely, low quality (LQ) and high quality (HQ). Finally, discriminative local binary pattern (DLBP) is adopted as the basis for feature extraction. For the HQ image, DLBP is adopted directly for feature extraction, and for the LQ one. CE_DLBP could be utilized for discriminative feature extraction for LQ images. Based on the lab-made database, rigorous experiments are conducted to demonstrate the effectiveness and feasibility of the proposed system. What is more, an additional experiment with PolyU database illustrates its generalization ability and robustness. Jun Wang 0071, Guoqing Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2010 | Cooperative co-evolutionary algorithm in satellite imaging scheduling of cooperative multiple centersabstractIn this paper, satellite imaging scheduling of cooperative multiple centers, an example of multi-agent systems, is discussed with a proposal of a novel Cooperative CoEvolutionary Planning Algorithm (CCEPA). Considering the numbers of the multi-center and the characteristics of targets to be observed, CCEPA decomposes the tasks into smaller components and evolves multiple solutions in the form of cooperative subpopulations. At the same time, for such evolutionary algorithm based on scheduling, a novel fixed-length binary encoding mechanism for tasks assigned to each center is also proposed. Incorporated with various features like archiving, dynamic sharing, the CCEPA is capable of maintaining archive diversity in the evolution and distributing the solutions uniformly along the Pareto front. Simulation and analysis show that the proposed algorithm can solve the problem effectively. Chong Wang 0001, Ning Jing, Jun Li 0020, Jun Wang 0071, Hao Chen 0046 |
IEEE Congress on Evolutionary Computation | 4 |
| 2007 | A multi-objective imaging scheduling approach for earth observing satellitesabstractEOSs (Earth Observing Satellites) circle the earth to take shotswhich are requested by customers. To make replete use of resourcesof EOSs, it is required to deal with the problem of united imagingscheduling of EOSs in a given scheduling horizon, which is acomplicated multi-objective combinatorial optimization problem. Inthis paper, we construct a mathematical model for the problem byabstracting imaging constraints of different EOSs. Then we propose anovel multi-objective EOSs imaging scheduling method, which is basedon the Strength Pareto Evolutionary Algorithm 2. The specialencoding technique and imaging constraint control are applied toguarantee feasibility of solutions. The approach is tested upon fourreal application problems of CBERS EOSs series. From the results, itis confirmed that the proposed approach is effective in solvingmulti-objective EOSs imaging scheduling problems. Jun Wang 0071, Ning Jing, Jun Li 0020, Zhong Hui Chen |
GECCO | 1 |