Wei Jia 0001

dblp:43/1191-1 · DBLP profile ↗
← Back
120ranked-venue papers
13as first author
67since 2021 · last 2026
0000-0001-5628-6237ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 4 first-author · 42 since 2021Artificial intelligence and machine learning · 52 · 6 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 11 · 4 first-author · 7 since 2021Security and privacy · 5 · 1 first-author · 3 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Event-Guided Scene Text Image Super-Resolution
abstract
Scene text image super-resolution aims to enhance text legibility by recovering high-resolution text images from low-resolution inputs. However, maintaining fine details such as text strokes, edges, and textual accuracy remains challenging, particularly in low-light environments and high-speed motion scenarios, where degradation is more severe. Event cameras, with their high temporal resolution and ability to capture intensity changes, offer a promising solution for restoring lost fine details and mitigating degradation in these challenging conditions. In this paper, we propose EvTSR, the first framework that integrates Event data for scene Text image Super-Resolution. The core of EvTSR is the dual-stream frequency boost (DSFB) mechanism, which separates image features into high- and low-frequency components. High-frequency details like edges and strokes are enhanced using event data via the event-guided high-frequency (EGH) mechanism, while low-frequency components, responsible for global structure, are refined using the Text-Guided Low-frequency (TGL) mechanism with a pre-trained text recognizer, ensuring textual coherence. To further improve cross-modal integration, we introduce the cross-modal fusion (CMF) mechanism, which effectively aligns event and image features, enabling robust information fusion. Extensive experiments demonstrate that EvTSR achieves superior performance over existing methods.
Zihan Qi, Zeyu Xiao 0002, Haoyi Zhao, Yang Zhao 0002, Feng Xue 0002, Wei Jia 0001
AAAI6
2026 Bidirectional Counterfactual Distillation for Review-Based Recommendation
abstract
Review-based recommendation methods typically integrate multiple behaviors, including interactions, reviews, and ratings, to model user preferences. To effectively extract preference signals from diverse behaviors, some studies train multiple student models to capture distinct behavioral patterns, and leverage online distillation to facilitate collaborative learning among them. However, we argue that these techniques suffer from bias contamination from rating distributions and feature homogenization during cross-behavior knowledge transfer: (1) Rating distribution bias, arising from non-uniform historical ratings, propagates across behaviors through distillation, contaminating the true preference representations of other behaviors. (2) Static distillation strategies often lead to homogenized behavioral features, hindering the learning of behavior-specific preferences. To address these issues, we propose a novel Bidirectional Counterfactual Distillation (BiCoD) framework for review-based recommendation. In BiCoD, we first design an adversarial counterfactual distillation module to suppress the impact of non-uniform rating distributions on distillation, thereby preventing it from contaminating the user's true preference representations across behaviors. Subsequently, we introduce a stage-aware bidirectional distillation strategy to enhance the distinctiveness of behavioral features, facilitating the effective learning of behavior-specific preferences. Extensive experiments on five real-world datasets validate the effectiveness and superiority of the proposed framework.
Sheng Sang, Shujie Li 0002, Shuaiyang Li 0001, Kang Liu 0024, Wei Jia 0001, Dan Guo 0001, Feng Xue 0002
AAAI6
2026 LSAP-PV: High-Fidelity Palm Vein Image Synthesis via Layered Spectral Absorption Projection-Guided Diffusion Model
abstract
Palm vein recognition has emerged as a promising biometric technology, yet its development remains constrained by the scarcity of large-scale publicly available datasets. Several methods of palm vein image generation have been proposed to address this issue. These methods usually focus on the anatomical realism of palm vein patterns, but overlook the biophysical correlation between identities and vein patterns, particularly in simulating identity-specific vein contrast. To tackle this limitation, we propose a novel biophysics-driven synthesis method. Our method constructs a 3D palm vascular tree via established modeling method. Then, a projection model is proposed to map the 3D tree into 2D space to derive palm vein patterns. The projection model is based on skin spectral absorption and simulates the natural attenuation of light passing through the skin using a layer integration method. For different identities, we sample different skin parameters, resulting in varying degrees of attenuation. This method effectively simulates the variation in vein contrast across different identities. Furthermore, we introduce a conditional diffusion model that uses the projected patterns as identity conditions to generate palm vein images. To the best of our knowledge, this is the first palm vein generation method based on the diffusion model. Experimental results demonstrate that our method not only outperforms existing methods, but also enables a recognition model trained on our synthetic data to achieve superior performance compared to a model trained on real-world data at a scale of 2,000 IDs under an open-set protocol with a TAR@FAR=1:1 of 1e-4.
Sheng Shang, Chenglong Zhao, Jianlong Jin, Yang Zhao 0002, Shouhong Ding, Wei Jia 0001
AAAI9
2026 LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition
abstract
Visual Speech Recognition (VSR), commonly known as lipreading, enables the recognition of spoken text by analyzing lip visual features. Due to the subtlety of lip movements, its recognition is much harder than other motion recognition tasks. Existing VSR models face the challenge of viseme ambiguity when processing phonemes with similar pronunciations—multiple phonemes share similar viseme features, leading to a notable drop in lipreading accuracy. To address this issue, this study proposes a Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition(LinProVSR) framework. First, an ambiguous sample set is constructed based on linguistic knowledge to provide supervisory signals for the model's training. Then, a Progressive Contrastive Disambiguation Network (PCDN) is designed, which progressively enhances the model's ability to capture the subtle viseme differences corresponding to similar phonemes through viseme-phoneme contrastive disambiguation in the encoding stage and text contrastive disambiguation in the decoding stage. Furthermore, we pioneer the Ambiguous Word Error Rate (AWER) metric specifically for evaluating recognition of phonetically ambiguous text, and verify the effectiveness of the proposed method on multiple public datasets, achieving a significant breakthrough especially in distinguishing visually similar phonemes.
Feng Xue 0002, Baochao Zhu, Wei Jia 0001, Shujie Li 0002, Yu Li 0053, Shengeng Tang, Dan Guo 0001
AAAI3
2026 Parameter-efficient transfer for CLIP-based text-to-person retrieval
Hai Min, Guanghui Zhan, Chunxiao Fan 0002, Yang Zhao 0002, Wei Jia 0001
Signal Process. Image Commun.5
2026 BITMNet: A Degradation-Aware Mixture-of-Experts Framework for Blind Inverse Tone Mapping
Wenyou Zhang, Yang Zhao 0002, Fangxing Zhang, Yuan Chen 0012, Zhao Zhang 0001, Wei Jia 0001
IEEE Signal Process. Lett.6
2026 Event-Based Dynamic Turbulence Mitigation
abstract
Atmospheric turbulence induces coupled spatio-temporal distortions, including blur, geometric deformation, and temporal jitter, which severely degrade image quality. We propose EvTurM, a practical framework leveraging event camera data for dynamic turbulence mitigation with precise motion cues and stable temporal modeling. Leveraging the high temporal resolution and dynamic range of events, EvTurM achieves robust restoration under diverse turbulence conditions. EvTurM comprises two key modules: (1) the event-aware modality enhancement module, which uses event-derived motion to enrich RGB features and recover structural details, and (2) the bidirectional modality calibration module, which jointly aligns RGB and event features in forward and backward propagation to reduce misalignment and enhance temporal consistency. Extensive experiments show EvTurM consistently surpasses existing methods and achieves superior performance.
Haoyi Zhao, Zeyu Xiao 0002, Zihan Qi, Yang Zhao 0002, Wei Jia 0001
IEEE Signal Process. Lett.5
2026 Joint Resolution and Rendering Artifacts Removal for Cloud Gaming Image
abstract
With the rapid development of the cloud gaming industry, low-quality rendering and rescaling strategies are commonly employed to mitigate the high costs of cloud-based computation and bandwidth. As a result, client-side images often contain artifacts such as mixed aliasing and resolution distortions, which cannot be effectively handled by current super-resolution models. In response, this paper proposes a cloud gaming image enhancement (CGIE) model to tackle both rendering and resolution degradations. Initially, this paper builds a large dataset by rendering and rescaling paired data with different qualities from collected 3D game scenes. Subsequently, a lightweight dual-branch enhancement network is designed, which consists of a high-frequency branch primarily focused on detail enhancement and a sampling-space branch aimed at enlarging the receptive field and perceiving multi-scale aliasing artifacts. Experimental results demonstrate the superior anti-aliasing and image enhancement performance of the proposed method across various real-world cloud games and even mobile games. The dataset and codes are available at https://github.com/YCheno/CGIE/.
Yang Zhao 0002, Yuan Chen 0012, Lin Li 0053, Wei Jia 0001, Ronggang Wang
IEEE Trans. Circuits Syst. Video Technol.5
2026 Wiener-Deconvolution-Driven Event-Based Deblurring for Low-Light Imaging
abstract
We address event-based deblurring for low-light imaging, where conventional frames suffer severe blur, noise and saturation, while events capture sharp high-frequency contrast changes with microsecond latency that can guide the recovery of lost structures. Existing event-based reconstruction methods neither explicitly model low-light noise and saturation nor enforce precise alignment between events and frames, which limits cross-modal fusion and deblurring quality. We propose the Wiener-Deconvolution-Driven Event-Based Deblurring Network (WiED-Net), which embeds the Wiener deconvolution into a deep architecture so that the physical imaging model and noise statistics are encoded in the frequency domain and high-frequency recovery is stabilized on noise dominated night data. WiED-Net adopts a two stage design. The first stage applies Wiener deconvolution in both image and feature spaces to suppress noise, recover saturated regions and reduce ringing, assisted by an eventguided cross-modal feature fusion (ECFF) module for accurate alignment. The second stage uses a multi-scale fusion module to integrate the complementary event and image branches. Training is constrained by a set of losses, including a tailored blur kernel loss that provides closed-loop regularization from physical priors. Together, these designs enable WiED-Net to recover fine details while robustly suppressing artifacts and noise, and to achieve superior quantitative and qualitative performance, achieving superior quantitative and qualitative performance with a notable improvement of 1.97 dB in PSNR and 5% in SSIM over the previous state-of-the-art methods in low-light deblurring. Code will be available at https://github.com/zhuzifeng38/WiED-Net.
Zeyu Xiao 0002, Jianlong Jin, Feng Xue 0002, Yu Liu 0023, Zhao Zhang 0001, Wei Jia 0001
IEEE Trans. Circuits Syst. Video Technol.7
2026 Learning Corruption-Invariant Components and Cross-Modal Correspondence for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised Visible-Infrared Person Re-Identification (US-VI-ReID) has great potential prospects because it does not require label information. However, corrupted pedestrian images collected due to corruption factors in real-world scenarios (e.g., noise, blur, and weather changes) largely limit the scalability of US-VI-ReID. In this paper, we explore the robustness of US-VI-ReID for the first time and propose a Multi-Granularity Spatial-Frequency Prototype Learning (MSPL) framework. The framework mainly consists of Multi-Channel Soft Augmentation (MSA), Robust Frequency Domain Feature Learning (RFL) module and Cross-modal Spatial-Frequency Prototype Matching (CSPM). Specifically, the MSA alleviates the sensitivity of model to color and abnormal samples through rich channel combinations and soft erasing. Subsequently, the RFL performs deep global filtering and amplitude attention compensated InstanceNorm to complete frequency and style modulation, concentrating on degradation-robust frequency content. Finally, the CSPM is designed to achieve multi-granularity prototype contrastive learning on cluster level and view level, then conduct cross-modal matching of multi-granularity spatial-frequency prototypes, thus establishing robust label association. With the above modules, our proposed framework can learn corruption-invariant feature components and generate robust cross-modal correspondence from unlabeled cross-modal images. Extensive experiments demonstrate that our MSPL outperforms other state-of-the-art methods by a large margin on the challenging SYSU-MM01-C and RegDB-C, while maintaining competitive on the SYSU-MM01 and RegDB.
Rui Sun 0004, Guoxi Huang, Jingjing Wu 0001, Wei Jia 0001
IEEE Trans. Inf. Forensics Secur.6
2026 Learning Dual Modality Interactions for Event-Based Motion Deblurring
abstract
Event cameras hold great potential for motion deblurring because they capture motion information with microsecond precision, offering robustness to motion blur. However, the limited interaction between RGB frames and event streams presents a significant challenge, preventing the full utilization of the event cameras' unique advantages. To address this, we proposeDual frame-eventInteraction and introduce a multi-scaleNetwork structure, DuInt-Net. DuInt-Net aims to tackle two key challenges: (1) enhancing the representational and interaction capabilities between RGB frames and event streams, and (2) adaptively selecting richer visual features for improved motion deblurring. We introduce an event-frame joint interaction module that consists of three branches: a base branch, a global awareness attention branch, and a local enhancement attention branch. The base branch processes essential pixel-level features that retain the original structural information. The global branch integrates event data to improve large-scale motion understanding, while the local branch uses large-kernel convolutions to refine fine-grained details in RGB frames. For superior reconstruction performance, we also propose the event-guided multi-scale fusion attention module, which effectively combines local visual information and global frame-event relationships. Extensive experiments demonstrate that DuInt-Net achieves superior performance, both quantitatively and qualitatively, showcasing its superior motion deblurring capabilities.
Zeyu Xiao 0002, Zhuoyuan Li 0001, Yang Zhao 0002, Yu Liu 0023, Zhao Zhang 0001, Wei Jia 0001
IEEE Trans. Multim.6
2026 Learning Multilayer Feature Projection for Homogeneous and Heterogeneous Palmprint Recognition
abstract
Owing to its remarkable convenience, weak invasiveness, and strong private security, palmprint recognition has become one of the most promising biometric methods and has attracted increasing attention in both academia and industry. Although considerable recognition performance has been achieved by existing palmprint learning methods, they generally require the use of substantial labeled datasets and involve substantial computational overhead for feature learning. In this article, we propose a novel multilayer projection learning (MLPL) method to achieve efficient palmprint feature learning and recognition. First, we transform the palmprint images into their direction-specific representations by computing the difference in the multiple directional responses. Then, we learn three layers of feature projections for robust feature learning, including low-rank projection for image noise decoupling, feature projection for discriminative feature exploration, and quantization projection for information preservation during feature encoding. With multilayer feature projections, palmprint images can be transformed into discriminative feature representations through a single-step process for efficient palmprint recognition. Moreover, we extend the proposed MLPL, referred to as E-MLPL, by minimizing the representation discrepancy between heterogeneous palmprint images to make it applicable for heterogeneous palmprint recognition. The results obtained from five widely adopted databases confirm the superior performance of the proposed method in terms of both accuracy and efficiency.
Lunke Fei, Kaiting Huang, Shuping Zhao, Qi Zhu 0001, Bob Zhang 0001, Wei Jia 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2026 Deep Learning in Palmprint Recognition: A Comprehensive Survey
abstract
Palmprint recognition has emerged as a prominent biometric technology, widely applied in diverse scenarios. Traditional handcrafted methods for palmprint recognition often fall short in representation capability, as they heavily depend on researchers’ prior knowledge. Deep learning (DL) has been introduced to address this limitation, leveraging its remarkable successes across various domains. While existing surveys focus narrowly on specific tasks within palmprint recognition—often grounded in traditional methodologies—there remains a significant gap in comprehensive research exploring DL-based approaches across all facets of palmprint recognition. This article bridges that gap by thoroughly reviewing recent advancements in DL-powered palmprint recognition. This article systematically examines progress across key tasks, including region-of-interest (ROI) segmentation, feature extraction, and security and privacy-oriented challenges. Beyond highlighting these advancements, this article identifies current challenges and uncovers promising opportunities for future research. By consolidating state-of-the-art progress, this review serves as a valuable resource for researchers, enabling them to stay abreast of cutting-edge technologies and drive innovation in palmprint recognition.
Chengrui Gao, Ziyuan Yang 0001, Wei Jia 0001, Lu Leng, Bob Zhang 0001, Andrew Beng Jin Teoh
IEEE Trans. Syst. Man Cybern. Syst.3
2025 PVTree: Realistic and Controllable Palm Vein Generation for Recognition Tasks
abstract
Palm vein recognition is an emerging biometric technology that offers enhanced security and privacy. However, acquiring sufficient palm vein data for training deep learning-based recognition models is challenging due to the high costs of data collection and privacy protection constraints. This has led to a growing interest in generating pseudo-palm vein data using generative models. Existing methods, however, often produce unrealistic palm vein patterns or struggle with controlling identity and style attributes. To address these issues, we propose a novel palm vein generation framework named PVTree. First, the palm vein identity is defined by a complex and authentic 3D palm vascular tree, created using an improved Constrained Constructive Optimization (CCO) algorithm. Second, palm vein patterns of the same identity are generated by projecting the same 3D vascular tree into 2D images from different views and converting them into realistic images using a generative model. As a result, PVTree satisfies the need for both identity consistency and intra-class diversity. Extensive experiments conducted on several publicly available datasets demonstrate that our proposed palm vein generation method surpasses existing methods and achieves a higher TAR@FAR=1e-4 under the 1:1 Open-set protocol. To the best of our knowledge, this is the first time that the performance of a recognition model trained on synthetic palm vein data exceeds that of the recognition model trained on real data, which indicates that palm vein image generation research has a promising future.
Sheng Shang, Chenglong Zhao, Jianlong Jin, Rizen Guo, Shouhong Ding, Yunsheng Wu, Yang Zhao 0002, Wei Jia 0001
AAAI10
2025 Occlusion-Embedded Hybrid Transformer for Light Field Super-Resolution
abstract
Transformer-based networks have set new benchmarks in light field super-resolution (SR), but adapting them to capture both global and local spatial-angular correlations efficiently remains challenging. Moreover, many methods fail to account for geometric details like occlusions, leading to performance drops. To tackle these issues, we introduce OHT. This hybrid network leverages occlusion maps through an occlusion-embedded mix layer. It combines the strengths of convolutional networks and Transformers via spatial-angular separable convolution (SASep-Conv) and angular self-attention (ASA). SASep-Conv offers a lightweight alternative to 3D convolution for capturing spatial-angular correlations, while the ASA mechanism applies 3D self-attention across the angular dimension. These designs allow OHT to capture global angular correlations effectively. Extensive experiments on multiple datasets demonstrate OHT's superior performance.
Zeyu Xiao 0002, Zhuoyuan Li 0001, Wei Jia 0001
AAAI3
2025 Diff-Palm: Realistic Palmprint Generation with Polynomial Creases and Intra-Class Variation Controllable Diffusion Models
abstract
Palmprint recognition is significantly limited by the lack of large-scale publicly available datasets. Previous methods have adopted Bézier curves to simulate the palm creases, which then serve as input for conditional GANs to generate realistic palmprints. However, without employing real data fine-tuning, the performance of the recognition model trained on these synthetic datasets would drastically decline, indicating a large gap between generated and real palmprints. This is primarily due to the utilization of an inaccurate palm crease representation and challenges in balancing intra-class variation with identity consistency. To address this, we introduce a polynomial-based palm crease representation that provides a new palm crease generation mechanism more closely aligned with the real distribution. We also propose the palm creases conditioned diffusion model with a novel intra-class variation control method. By applying our proposed K-step noise-sharing sampling, we are able to synthesize palmprint datasets with large intra-class variation and high identity consistency. Experimental results show that, for the first time, recognition models trained solely on our synthetic datasets, without any fine-tuning, outperform those trained on real datasets. Furthermore, our approach achieves superior recognition performance as the number of generated identities increases.
Jianlong Jin, Chenglong Zhao, Sheng Shang, Jianqing Xu, Shaoming Wang, Yang Zhao 0002, Shouhong Ding, Wei Jia 0001, Yunsheng Wu
CVPR10
2025 Unified Adversarial Augmentation for Improving Palmprint Recognition
Jianlong Jin, Chenglong Zhao, Sheng Shang, Yang Zhao 0002, Shouhong Ding, Wei Jia 0001, Yunsheng Wu
ICCV9
2025 Multi-Layer Gaussian Splatting for Single-Image Feed-Forward Spatial Scene Reconstruction
abstract
Recently, 3D Gaussian Splatting (3DGS) has achieved remarkable results in 3D reconstruction and view synthesis tasks. However, single-view feed-forward 3DGS still faces significant challenges. Current state-of-the-art (SOTA) single-view 3DGS methods typically employ a small number of layers (1-2 layers) with Gaussian Splatting (GS) representations at the same resolution as the input image to address the irregularity of GS data. However, such shallow and uniform GS primitive distributions is difficult to represent occluded regions and important spatial details. Inspired by multi-plane images, this paper proposes a Multi-Layer Gaussian Splatting (MLGS) representation, which consists of shallow base GS layers for visible content and multiple occlusion GS layers dedicated to reconstructing occluded regions. The proposed MLGS representation explicitly decouples the learning processes of visible and occluded content while enhancing occlusion prediction through the following components. First, spatial stratification of GS is achieved by estimating the depth distribution range of GS primitives across different layers, forcing GS to learn spatial content reconstruction at different depths. Second, a mask-guided mechanism is proposed to effectively isolate occlusion regions and guide inpainting using spatially context-aware features. Finally, a gated convolution block is designed to dynamically modulate feature fusion to enhance reconstruction fidelity. With separate loss supervision for base and occlusion layers, MLGS enables geometrically plausible scene completion. Experiments on RealEstate10K, KITTI, and NYUv2 datasets demonstrate that the proposed method achieves SOTA performance for single-image spatial scene reconstruction.
Shanding Diao, Yang Zhao 0002, Yuan Chen 0012, Zhao Zhang 0001, Wei Jia 0001, Ronggang Wang
ACM Multimedia5
2025 Cross-domain facial expression recognition: Bi-Directional Fusion of Active and Stable Information
Jiaqiu Ai, Weibao Xue, Wei Jia 0001
Eng. Appl. Artif. Intell.6
2025 Rethinking Contemporary Deep Learning Techniques for Error Correction in Biometric Data
Yen-Lung Lai, Xingbo Dong, Zhe Jin 0001, Wei Jia 0001, Massimo Tistarelli, Xuejun Li 0001
Int. J. Comput. Vis.4
2025 Local Texture Pattern Estimation for Image Detail Super-Resolution
abstract
In the image super-resolution (SR) field, recovering missing high-frequency textures has always been an important goal. However, deep SR networks based on pixel-level constraints tend to focus on stable edge details and cannot effectively restore random high-frequency textures. It was not until the emergence of the generative adversarial network (GAN) that GAN-based SR models achieved realistic texture restoration and quickly became the mainstream method for texture SR. However, GAN-based SR models still have some drawbacks, such as relying on a large number of parameters and generating fake textures that are inconsistent with ground truth. Inspired by traditional texture analysis research, this paper proposes a novel SR network based on local texture pattern estimation (LTPE), which can restore fine high-frequency texture details without GAN. A differentiable local texture operator is first designed to extract local texture structures, and a texture enhancement branch is used to predict the high-resolution local texture distribution based on the LTPE. Then, the predicted high-resolution texture structure map can be used as a reference for the texture fusion SR branch to obtain high-quality texture reconstruction. Finally, $L_{1}$L1 loss and Gram loss are simultaneously used to optimize the network. Experimental results demonstrate that the proposed method can effectively recover high-frequency texture without using GAN structures. In addition, the restored high-frequency details are constrained by local texture distribution, thereby reducing significant errors in texture generation.
Yang Zhao 0002, Yuan Chen 0012, Nannan Li 0001, Wei Jia 0001, Ronggang Wang
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 MulFS-CAP: Multimodal Fusion-Supervised Cross-Modality Alignment Perception for Unregistered Infrared-Visible Image Fusion
abstract
In this study, we propose Multimodal Fusion-supervised Cross-modality Alignment Perception (MulFS-CAP), a novel framework for single-stage fusion of unregistered infrared-visible images. Traditional two-stage methods depend on explicit registration algorithms to align source images spatially, often adding complexity. In contrast, MulFS-CAP seamlessly blends implicit registration with fusion, simplifying the process and enhancing suitability for practical applications. MulFS-CAP utilizes a shared shallow feature encoder to merge unregistered infrared-visible images in a single stage. To address the specific requirements of feature-level alignment and fusion, we develop a consistent feature learning approach via a learnable modality dictionary. This dictionary provides complementary information for unimodal features, thereby maintaining consistency between individual and fused multimodal features. As a result, MulFS-CAP effectively reduces the impact of modality variance on cross-modality feature alignment, allowing for simultaneous registration and fusion. Additionally, in MulFS-CAP, we advance a novel cross-modality alignment approach, creating a correlation matrix to detail pixel relationships between source images. This matrix aids in aligning features across infrared and visible images, further refining the fusion process. The above designs make MulFS-CAP more lightweight, effective and explicit registration-free. Experimental results from different datasets demonstrate the effectiveness of our proposed method and its superiority over the state-of-the-art two-stage methods.
Huafeng Li 0001, Zengyi Yang, Wei Jia 0001, Zhengtao Yu 0001, Yu Liu 0023
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Joint Finger Valley Points-Free ROI Detection and Recurrent Layer Aggregation for Palmprint Recognition in Open Environment
abstract
Cooperative palmprint recognition, pivotal for civilian and commercial uses, stands as the most essential and broadly demanded branch in biometrics. These applications, often tied to financial transactions, require high accuracy in recognition. Currently, research in palmprint recognition primarily aims to enhance accuracy, with relatively few studies addressing the automatic and flexible palm region of interest (ROI) extraction (PROIE) suitable for complex scenes. Particularly, the intricate conditions of open environment, alongside the constraint of human finger skeletal extension limiting the visibility of Finger Valley Points (FVPs), render conventional FVPs-based PROIE methods ineffective. In response to this challenge, we propose an FVPs-Free Adaptive ROI Detection (FFARD) approach, which utilizes cross-dataset hand shape semantic transfer (CHSST) combined with the constrained palm inscribed circle search, delivering exceptional hand segmentation and precise PROIE. Furthermore, a Recurrent Layer Aggregation-based Neural Network (RLANN) is proposed to learn discriminative feature representation for high recognition accuracy in both open-set and closed-set modes. The Angular Center Proximity Loss (ACPLoss) is designed to enhance intra-class compactness and inter-class discrepancy between learned palmprint features. Overall, the combined FFARD and RLANN methods are proposed to address the challenges of palmprint recognition in open environment, collectively referred to as RDRLA. Experimental results on four palmprint benchmarks HIT-NIST-V1, IITD, MPD and BJTU_PalmV2 show the superiority of the proposed method RDRLA over the state-of-the-art (SOTA) competitors. The code of the proposed method is available athttps://github.com/godfatherwang2/RDRLA.
Tingting Chai, Ru Li 0002, Wei Jia 0001, Xiangqian Wu 0002
IEEE Trans. Inf. Forensics Secur.4
2025 An Active Multi-Target Domain Adaptation Strategy: Progressive Class Prototype Rectification
abstract
Compared to single-source to single-target (1S1T) domain adaptation, single-source to multi-target (1SmT) domain adaptation is more practical but also more challenging. In 1SmT scenarios, the significant differences in feature distributions between various target domains increase the difficulty for models to adapt to multiple domains. Moreover, 1SmT requires effective transfer to each target domain while maintaining performance in the source domain, demanding higher generalization capabilities from the model. In 1S1T scenarios, active domain adaptation methods improve generalization by incorporating a few target domain samples, but these methods are rarely applied in 1SmT due to potential sampling bias and outlier interference. To address this, we propose Progressive Prototype Refinement (PPR), an active multi-target domain adaptation method combining 1SmT with active learning to enhance cross-domain knowledge transfer. Specifically, an uncertainty assessment strategy is used to select representative samples from multiple target domains, forming a candidate set for model training. Based on the Lindeberg--Levy central limit theorem, we sample from a Gaussian distribution using corrected prototype statistics to augment the classifier's feature input, allowing the model to learn transitional information between domains. Finally, a mapping matrix is used for cross-domain alignment, addressing incomplete class coverage and outlier interference. Extensive experiments on multiple benchmark datasets demonstrate PPR's superior performance, with a 6.35% improvement on the PACS dataset and a 17.32% improvement on the Remote Sensing dataset.
Jiaqiu Ai, Le Wu 0001, Dan Guo 0001, Wei Jia 0001, Richang Hong
IEEE Trans. Multim.5
2024 Stereo Vision Conversion from Planar Videos Based on Temporal Multiplane Images
abstract
With the rapid development of 3D movie and light-field displays, there is a growing demand for stereo videos. However, generating high-quality stereo videos from planar videos remains a challenging task. Traditional depth-image-based rendering techniques struggle to effectively handle the problem of occlusion exposure, which occurs when the occluded contents become visible in other views. Recently, the single-view multiplane images (MPI) representation has shown promising performance for planar video stereoscopy. However, the MPI still lacks real details that are occluded in the current frame, resulting in blurry artifacts in occlusion exposure regions. In fact, planar videos can leverage complementary information from adjacent frames to predict a more complete scene representation for the current frame. Therefore, this paper extends the MPI from still frames to the temporal domain, introducing the temporal MPI (TMPI). By extracting complementary information from adjacent frames based on optical flow guidance, obscured regions in the current frame can be effectively repaired. Additionally, a new module called masked optical flow warping (MOFW) is introduced to improve the propagation of pixels along optical flow trajectories. Experimental results demonstrate that the proposed method can generate high-quality stereoscopic or light-field videos from a single view and reproduce better occluded details than other state-of-the-art (SOTA) methods. https://github.com/Dio3ding/TMPI
Shanding Diao, Yuan Chen 0012, Yang Zhao 0002, Wei Jia 0001, Zhao Zhang 0001, Ronggang Wang
AAAI4
2024 PCE-Palm: Palm Crease Energy Based Two-Stage Realistic Pseudo-Palmprint Generation
abstract
The lack of large-scale data seriously hinders the development of palmprint recognition. Recent approaches address this issue by generating large-scale realistic pseudo palmprints from Bézier curves. However, the significant difference between Bézier curves and real palmprints limits their effectiveness. In this paper, we divide the Bézier-Real difference into creases and texture differences, thus reducing the generation difficulty. We introduce a new palm crease energy (PCE) domain as a bridge from Bézier curves to real palmprints and propose a two-stage generation model. The first stage generates PCE images (realistic creases) from Bézier curves, and the second stage outputs realistic palmprints (realistic texture) with PCE images as input. In addition, we also design a lightweight plug-and-play line feature enhancement block to facilitate domain transfer and improve recognition performance. Extensive experimental results demonstrate that the proposed method surpasses state-of-the-art methods. Under extremely few data settings like 40 IDs (only 2.5% of the total training set), our model achieves a 29% improvement over RPG-Palm and outperforms ArcFace with 100% training set by more than 6% in terms of TAR@FAR=1e-6.
Jianlong Jin, Chenglong Zhao, Shouhong Ding, Yang Zhao 0002, Wei Jia 0001
AAAI9
2024 Blind Video Bit-Depth Expansion
abstract
With the rapid development of high-bit-depth display devices, bit-depth expansion (BDE) algorithms that extend low-bit-depth images to high-bit-depth images have received increasing attention. Due to the sensitivity of bit-depth distortions to tiny numerical changes in the least significant bits, the nuanced degradation differences in the training process may lead to varying degradation data distributions, causing the trained models to overfit specific types of degradations. This paper focuses on the problem of blind video BDE, proposing a degradation prediction and embedding framework, and designing a video BDE network based on a recurrent structure and dual-frame alignment fusion. Experimental results demonstrate that the proposed model can outperform some state-of-the-art (SOTA) models in terms of banding artifact removal and color correction, avoiding overfitting to specific degradations and obtaining better generalization ability across multiple datasets. https://github.com/duanpanjun/BVBDE
Panjun Duan, Yang Zhao 0002, Yuan Chen 0012, Wei Jia 0001, Zhao Zhang 0001, Ronggang Wang
ACM Multimedia4
2024 Night-time vehicle model recognition based on domain adaptation
Weixiao Chen, Fengxin Chen, Wei Jia 0001, Qiang Lu 0002
Multim. Tools Appl.4
2024 AMGNet: Aligned Multilevel Gabor Convolution Network for Palmprint Recognition
abstract
Palmprint recognition has seen significant advancements and garnered considerable attention recently. However, deep learning methods have yet to effectively incorporate insights from traditional approaches to extract palmprint-specific features. Moreover, intra-class spatial variation problems, which degrade the recognition performance, have not been adequately addressed. To tackle these limitations, this study proposes an Aligned Multilevel Gabor Convolution Network (AMGNet) to identify the informative and salient aspects of the palmprints. The network unifies a multilevel Gabor feature fusion branch with a spatial alignment branch, enabling the joint mining of aligned multilevel features specific to palmprints. Within the feature fusion branch, we incorporate two specialized Gabor convolution modules: one targets the principal lines of the palm, while the other focuses on the wrinkles, augmenting the discriminative power of the acquired features. To enhance the model’s robustness against within-class variations, we design a spatial alignment branch that specifically enables the rectification of palmprints’ spatial positions. In conjunction with this, we introduce a novel direction-based CosAngle loss function to facilitate geometric alignment among samples from same palms while spatially distancing those from different palms. Furthermore, we construct a palmprint database consisting of 3, 000 palms from 1, 500 individuals to explore large-scale population potential. Extensive experimental results on six benchmark datasets demonstrate that our proposed method outperforms other popular approaches in palmprint recognition tasks.
Wei Jia 0001, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Toward Individual Tone Preference in Underwater Image Enhancement
abstract
Underwater images often suffer from severe color distortion due to the challenging imaging environment. Underwater image enhancement (UIE) techniques have been developed to recover clear images, laying the foundation for various underwater research. However, existing UIE methods tend to produce fixed results without considering individual preferences for different color tones. And there is no dataset with ground truth (GT) in different tones. Therefore, we came up with the possibility of using the currently popular multimodal methods to control the color tone of enhanced images. This article proposes a method for generating underwater enhanced images with cold, warm, and normal tones using multimodal information supervision (MM-UIE). First, we leverage the relationship between text prompts and images to supervise the generation of cold or warm images. In addition, we introduce a 6-D color operator, which not only enhances the tone control of underwater images but also serves as a bridge between different tone images. Finally, we also found that multimodal supervision methods can not only control the color tone of underwater images but also improve the quality of underwater image generation. Experimental results demonstrate the superior performance of our method compared to state-of-the-art (SOTA) techniques. Our codes will be publicly available athttps://github.com/perseveranceLX/MM-UIE.
Yang Zhao 0002, Kaichen Chi, Zhao Zhang 0001, Wei Jia 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Manifold-Based Incomplete Multi-View Clustering via Bi-Consistency Guidance
abstract
Incomplete multi-view clustering primarily focuses on dividing unlabeled data into corresponding categories with missing instances, and has received intensive attention due to its superiority in real applications. Considering the influence of incomplete data, the existing methods mostly attempt to recover data by adding extra terms. However, for the unsupervised methods, a simple recovery strategy will cause errors and outlying value accumulations, which will affect the performance of the methods. Broadly, the previous methods have not taken the effectiveness of recovered instances into consideration, or cannot flexibly balance the discrepancies between recovered data and original data. To address these problems, we propose a novel method termed Manifold-based Incomplete Multi-view clustering via Bi-consistency guidance (MIMB), which flexibly recovers incomplete data among various views, and attempts to achieve biconsistency guidance via reverse regularization. In particular, MIMB adds reconstruction terms to representation learning by recovering missing instances, which dynamically examines the latent consensus representation. Moreover, to preserve the consistency information among multiple views, MIMB implements a biconsistency guidance strategy with reverse regularization of the consensus representation and proposes a manifold embedding measure for exploring the hidden structure of the recovered data. Notably, MIMB aims to balance the importance of different views, and introduces an adaptive weight term for each view. Finally, an optimization algorithm with an alternating iteration optimization strategy is designed for final clustering. Extensive experimental results on 6 benchmark datasets are provided to confirm that MIMB can significantly obtain superior results as compared with several state-of-the-art baselines.
Huibing Wang, Mingze Yao, Yawei Chen, Yunqiu Xu, Haipeng Liu 0004, Wei Jia 0001, Xianping Fu, Yang Wang 0023
IEEE Trans. Multim.6
2024 Depth Matters: Spatial Proximity-Based Gaze Cone Generation for Gaze Following in Wild
abstract
Gaze following aims to predict where a person is looking in a scene. Existing methods tend to prioritize traditional 2D RGB visual cues or require burdensome prior knowledge and extra expensive datasets annotated in 3D coordinate systems to train specialized modules to enhance scene modeling. In this work, we introduce a novel framework deployed on a simple ResNet backbone, which exclusively uses image and depth maps to mimic human visual preferences and realize 3D-like depth perception. We first leverage depth maps to formulate spatial-based proximity information regarding the objects with the target person. This process sharpens the focus of the gaze cone on the specific region of interest pertaining to the target while diminishing the impact of surrounding distractions. To capture the diverse dependence of scene context on the saliency gaze cone, we then introduce a learnable grid-level regularized attention that anticipates coarse-grained regions of interest, thereby refining the mapping of the saliency feature to pixel-level heatmaps. This allows our model to better account for individual differences when predicting others’ gaze locations. Finally, we employ the KL-divergence loss to super the grid-level regularized attention, which combines the gaze direction, heatmap regression, and in/out classification losses, providing comprehensive supervision for model optimization. Experimental results on two publicly available datasets demonstrate the comparable performance of our model with less help of modal information. Quantitative visualization results further validate the interpretability of our method. The source code will be available at https://github.com/VUT-HFUT/DepthMatters .
Kun Li 0008, Zhun Zhong, Wei Jia 0001, Bin Hu 0001, Xun Yang 0001, Meng Wang 0001, Dan Guo 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 A Novel Hybrid Fusion Combining Palmprint and Palm Vein for Large-Scale Palm-Based Recognition
abstract
Palmprint and palm vein are emerging as unique biometric traits for identity authentication, each with its own advantages and limitations. Using these two traits jointly promises to enhance the discriminative and anti-spoofing capabilities. While existing research often combines these traits in parallel, such approaches lead to unnecessary increase in response time. Moreover, large-scale palm-based recognition, as a great potential task, raises higher requirements in accuracy and time efficiency. Nevertheless, few efforts have been dedicated to either data establishment or method investigation for this task. To this end, a large-scale palm-based multimodal dataset covering$ 20\,000 $palms is proposed, far larger than any of its kind. We also propose a hybrid fusion method to leverage these two diverse features. Our method employs a two-stage recognition process. First, a dual likelihood ratio test for coarse recognition is designed to assign palms into imposter certainty, genuine certainty or uncertainty classes. The coarse recognition narrows down the number of possible identities accurately using one trait, palm vein, consuming less recognition time. Then, in fine recognition, an adaptive weighted fusion of palmprint and palm vein is proposed to delicately rerecognize the uncertainty subsets that are in doubt in the coarse recognition, resulting in a more discriminative capacities. Experimental results confirm the effectiveness of our method, showing improved recognition performance with high-time efficiency.
Wei Jia 0001, Junan Chen 0001, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Toward Large-Scale Palmprint Image Analysis by a Rich Orientation Code
abstract
Palmprint recognition has gained considerable attention in recent years, accompanied by significant progress. However, large-scale palmprint recognition, which holds great potential for extensive civilian applications like university access control, remains underexplored. In this work, we propose the CUHK-T dataset, the largest palmprint dataset to date, containing over10k individuals’ palmprints. As the data volume expands, the heightened complexity, such as similar principal lines from different palms, necessitates a recognition method capable of extracting more discriminative features. Motivated by the rich palm lines distributed in palmprint, including not only nonintersecting, but also intersecting line segments, we model intersecting lines, investigate their properties, and propose a novel and explainable palmprint recognition method. The model treats the nonintersecting line segment as a special case, allowing for the extraction of orientation information from both types of line segments. In addition to the rich orientation information of the intersecting lines, the extracted feature accounts for the relative width of these lines. These advancements enable the extracted rich orientation code to be more discriminative and representative for palmprint. We then present a bitwise similarity measurement for efficiently and effectively comparing two rich orientation codes. Our extensive experiments and evaluations with popular palmprint recognition algorithms demonstrate the effectiveness and superior performance of our method on the large-scale dataset. These results also serve as a foundational baseline, facilitating the advancement of further research in the domain of large-scale palmprint recognition.
Wei Jia 0001, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Dense Hybrid Attention Network for Palmprint Image Super-Resolution
abstract
Palmprint has attracted increasing attention for biometric recognition in recent years due to its outstanding reliability, user-friendliness and hygiene. However, existing palmprint recognition methods usually require high-quality palmprint images with clear texture and line patterns; however, in practical applications palmprint images are usually of low quality. In this study, we propose a dense hybrid attention (DHA) network for palmprint image super-resolution (SR) by recovering the clear palmprint-specific characteristics. The proposed DHA network first obtains the high-dimensional shallow representation via a single convolution layer, and then jointly learns the local and global palmprint-specific features via parallel convolutional neural network (CNN)-and transformer-based branches. Particularly, we develop two enhanced spatial and channel attention (CA) modules to adaptively emphasize the local position-specific characteristics of palmprints, such that the SR palmprint images can be well recovered with clear texture and edge characteristics. Experimental results on three publicly used palmprint databases clearly show the effectiveness of the proposed method for palmprint image SR.
Yao Wang 0012, Lunke Fei, Shuping Zhao, Qi Zhu 0001, Jie Wen 0001, Wei Jia 0001, Imad Rida
IEEE Trans. Syst. Man Cybern. Syst.6
2023 RPG-Palm: Realistic Pseudo-data Generation for Palmprint Recognition
abstract
Palmprint recently shows great potential in recognition applications as it is a privacy-friendly and stable biometric. However, the lack of large-scale public palmprint datasets limits further research and development of palmprint recognition. In this paper, we propose a novel realistic pseudo-palmprint generation (RPG) model to synthesize palmprints with massive identities. We first introduce a conditional modulation generator to improve the intra-class diversity. Then an identity-aware loss is proposed to ensure identity consistency against unpaired training. We further improve the Bézier palm creases generation strategy to guarantee identity independence. Extensive experimental results demonstrate that synthetic pretraining significantly boosts the recognition model performance. For example, our model improves the state-of-the-art BézierPalm by more than 5% and 14% in terms of TAR@FAR=1e-6 under the 1 : 1 and 1 : 3 Open-set protocol. When accessing only 10% of the real training data, our method still outperforms ArcFace with 100% real training data, indicating that we are closer to real-data-free palmprint recognition.
Jianlong Jin, Huaen Li, Kai Zhao 0012, Shouhong Ding, Yang Zhao 0002, Wei Jia 0001
ICCV10
2023 Cross-view Resolution and Frame Rate Joint Enhancement for Binocular Video
abstract
With the popular of stereo video and free-viewpoint video, binocular and multi-view video enhancement has attracted increasing attention. Current binocular video enhancement methods mainly focus on stereo super-resolution. In this paper, we tend to discuss a new binocular video resolution and frame-rate enhancement scenario to fully utilize the cross-view complementary information. Specifically, one view is captured with high resolution (HR) and low frame-rate (LFR), while the other viewpoint records low resolution (LR) and high frame-rate (HFR) video. Then, a binocular video joint enhancement network, which adopts dual-branch structure with cross-view guidance, is proposed to jointly reconstruct HR and HFR stereo videos. The proposed framework can reduce the capture, storage, compression, and transmission cost of normal HR and HFR stereo videos. Compared with single-view super-resolution and video frame interpolation techniques, the proposed method can recover more realistic HR details and intermediate motion by using cross-view reference. Experimental results on stereo video datasets demonstrate the effectiveness of the proposed joint resolution and frame-rate enhancement framework.
Panda Pan, Yang Zhao 0002, Yuan Chen 0012, Wei Jia 0001, Zhao Zhang 0001, Ronggang Wang
ACM Multimedia4
2023 A Video Face Recognition Leveraging Temporal Information Based on Vision Transformer
Hui Zhang 0039, Jiewen Yang, Xingbo Dong, Xingguo Lv, Wei Jia 0001, Zhe Jin 0001, Xuejun Li 0001
PRCV (5)5
2023 Hybrid feature enhancement network for few-shot semantic segmentation
Hai Min, Yemao Zhang, Yang Zhao 0002, Wei Jia 0001, Ying-Ke Lei, Chunxiao Fan 0002
Pattern Recognit.4
2023 Fast Blind Decontouring Network
abstract
Contouring artifacts usually appear in large and smooth flat areas, which are caused by many widely used processes such as bit-depth expansion, compression, image sharpening and contrast enhancement. Unfortunately, recent decontouring methods were mainly designed for specific and non-blind degradations, which significantly reduces the generalization ability of these methods when applied to complex and various real-world false contours. Therefore, this paper explores the blind decontouring problem by proposing a blind decontouring network (BDCN). Instead of directly training a decontouring network with mixed degradations, the proposed model consists of two independent modules, i.e., a flat region detection module (FDM) and a decontouring module (DCM). The FDM is designed to extract flat region masks robust to various false contours, which can preserve texture details from global smoothing. Then, the task of DCM becomes simply smoothing different contouring artifacts. Both the FDM and DCM are designed with a lightweight architecture and reparameterization strategy. Experimental results on both synthetic and real-world contouring artifacts demonstrate the effectiveness and generalization of the proposed method.
Yang Zhao 0002, Wei Jia 0001, Yuan Chen 0012, Ronggang Wang
IEEE Trans. Circuits Syst. Video Technol.2
2023 Learning Deep Blind Quality Assessment for Cartoon Images
abstract
Although the cartoon industry has developed rapidly in recent years, few studies pay special attention to cartoon image quality assessment (IQA). Unfortunately, applying blind natural IQA algorithms directly to cartoons often leads to inconsistent results with subjective visual perception. Hence, this brief proposes a blind cartoon IQA method based on convolutional neural networks (CNNs). Note that training a robust CNN depends on manually labeled training sets. However, for a large number of cartoon images, it is very time-consuming and costly to manually generate enough mean opinion scores (MOSs). Therefore, this brief first proposes a full reference (FR) cartoon IQA metric based on cartoon-texture decomposition and then uses the estimated FR index to guide the no-reference IQA network. Moreover, in order to improve the robustness of the proposed network, a large-scale dataset is established in the training stage, and a stochastic degradation strategy is presented, which randomly implements different degradations with random parameters. Experimental results on both synthetic and real-world cartoon image datasets demonstrate the effectiveness and robustness of the proposed method.
Yuan Chen 0012, Yang Zhao 0002, Wei Jia 0001, Xiaoping Liu 0003
IEEE Trans. Neural Networks Learn. Syst.4
2023 Toward Efficient Palmprint Feature Extraction by Learning a Single-Layer Convolution Network
abstract
In this article, we propose a collaborative palmprint-specific binary feature learning method and a compact network consisting of a single convolution layer for efficient palmprint feature extraction. Unlike most existing palmprint feature learning methods, such as deep-learning, which usually ignore the inherent characteristics of palmprints and learn features from raw pixels of a massive number of labeled samples, palmprint-specific information, such as the direction and edge of patterns, is characterized by forming two kinds of ordinal measure vectors (OMVs). Then, collaborative binary feature codes are jointly learned by projecting double OMVs into complementary feature spaces in an unsupervised manner. Furthermore, the elements of feature projection functions are integrated into OMV extraction filters to obtain a collection of cascaded convolution templates that form a single-layer convolution network (SLCN) to efficiently obtain the binary feature codes of a new palmprint image within a single-stage convolution operation. Particularly, our proposed method can easily be extended to a general version that can efficiently perform feature extraction with more than two types of OMVs. Experimental results on five benchmark databases show that our proposed method achieves very promising feature extraction efficiency for palmprint recognition.
Lunke Fei, Shuping Zhao, Wei Jia 0001, Bob Zhang 0001, Jie Wen 0001, Yong Xu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Learning Unified Binary Feature Codes for Cross-Illumination Palmprint Recognition
Wei Jia 0001, Lunke Fei, Shuping Zhao, Shuyi Li 0003, Jie Wen 0001, Jinrong Cui
CGI1
2022 Depth Estimation by Combining Binocular Stereo and Monocular Structured-Light
abstract
It is well known that the passive stereo system cannot adapt well to weak texture objects, e.g., white walls. However, these weak texture targets are very common in indoor environments. In this paper, we present a novel stereo system, which consists of two cameras (an RGB camera and an IR camera) and an IR speckle projector. The RGB camera is used both for depth estimation and texture acquisition. The IR camera and the speckle projector can form a monocular structured-light (MSL) subsystem, while the two cameras can form a binocular stereo subsystem. The depth map generated by the MSL subsystem can provide external guidance for the stereo matching networks, which can improve the matching accuracy significantly. In order to verify the effectiveness of the proposed system, we build a prototype and collect a test dataset in indoor scenes. The evaluation results show that the Bad 2.0 error of the proposed system is 28.2% of the passive stereo system when the network RAFT is used. The dataset and trained models are available at https://github.com/YuhuaXu/MonoStereoFusion.
Yuhua Xu 0006, Yushan Yu, Wei Jia 0001, Zhaobi Chu, Yulan Guo
CVPR4
2022 BézierPalm: A Free Lunch for Palmprint Recognition
Kai Zhao 0012, Chuhan Zhou, Shouhong Ding, Wei Jia 0001, Wei Shen 0002
ECCV (13)8
2022 Dual-Rank Attention Module for Fine-Grained Vehicle Model Recognition
Wen Cai, Wenjia Zhu, Longdao Xu, Qiang Lu 0002, Wei Jia 0001
PRCV (1)7
2022 An iterative solution for improving the generalization ability of unsupervised skeleton motion retargeting
Shujie Li 0002, Wei Jia 0001, Yang Zhao 0002, Liping Zheng
Comput. Graph.3
2022 Cartoon Image Processing: A Survey
Yang Zhao 0002, Diya Ren, Yuan Chen 0012, Wei Jia 0001, Ronggang Wang, Xiaoping Liu 0003
Int. J. Comput. Vis.4
2022 Touchless palmprint recognition based on 3D Gabor template and block feature refinement
Zhaoqun Li, Jinxing Li 0003, Wei Jia 0001, David Zhang 0001
Knowl. Based Syst.5
2022 EEPNet: An efficient and effective convolutional neural network for palmprint recognition
Wei Jia 0001, Yang Zhao 0002, Shujie Li 0002, Hai Min
Pattern Recognit. Lett.1
2022 Representation and Reinforcement Learning for Task Scheduling in Edge Computing
abstract
Recently, many deep reinforcement learning (DRL)-based task scheduling algorithms have been widely used in edge computing (EC) to reduce energy consumption. Unlike the existing algorithms considering fixed and fewer edge nodes (servers) and tasks, in this article, a representation model with a DRL based algorithm is proposed to adapt the dynamic change of nodes and tasks and solve the dimensional disaster in DRL caused by a massive scale. Specifically, 1) we apply the representation learning models to describe the different nodes and tasks in EC, i.e., nodes and tasks are mapped to corresponding vector sub-spaces to reduce the dimensions and store the vector space efficiently. 2) With the space after dimensionality reduction, a DRL-based algorithm is employed to learn the vector representations of nodes and tasks and make scheduling decisions. 3) The experiments are conducted with the real-world data set, and the results show that the proposed representation model with DRL-based algorithm outperforms the baselines 18.04 and 9.94 percent on average regarding energy consumption and service level agreement violation (SLAV), respectively.
Zhiqing Tang, Wei Jia 0001, Wenmian Yang, Yongjian You
IEEE Trans. Big Data2
2022 Embedding Pose Information for Multiview Vehicle Model Recognition
abstract
Vehicle model recognition is a typical fine-grained classification task that has a wide range of application prospects in safe cities and constitutes a research hotspot in the field of computer vision. Vehicles in images can appear at various angles, resulting in large differences in appearance. The existence of “multiviews” renders vehicle model recognition challenging. Recent research on vehicle model recognition has not fully explored the pose information of vehicles in different images, resulting in low model performance. In this study, we use vehicle pose information to solve the multiview vehicle model recognition (MV-VMR) problem and design a convolutional neural network (CNN) model with embedded vehicle pose information, known as the embedding pose CNN (EP-CNN). The proposed model includes two subnetworks: the pose estimation subnetwork (PE-SubNet) and vehicle model classification subnetwork (VMC-SubNet). PE-SubNet extracts the vehicle pose information, including the pose features and vehicle viewpoint. In VMC-SubNet, considering the scale variation of vehicles, an improved squeeze-and-excitation (SE) block, named the MultiSE block is implemented. We embed the vehicle viewpoint into the MultiSE block, which reweighs each channel such that the extracted features elicit different responses to different viewpoints. Subsequently, the pose features and classification features are integrated for classification. Experiments are conducted on the benchmark CompCars web-nature and Stanford Cars datasets. The results demonstrate that the proposed EP-CNN method can achieve higher recognition accuracy than most classic CNN models and several state-of-the-art fine-grained vehicle model classification algorithms. Code has been made available at:https://github.com/HFUT-CV/EP-CNN.
Yuanzi Fu, Wei Jia 0001, Jun Yu 0001, Zhisheng Yan
IEEE Trans. Circuits Syst. Video Technol.4
2022 Rethinking Deinterlacing for Early Interlaced Videos
abstract
In recent years, high-definition restoration of early videos have received much attention. Real-world interlaced videos usually contain various degradations mixed with interlacing artifacts, such as noises and compression artifacts. Unfortunately, traditional deinterlacing methods only focus on the inverse process of interlacing scanning, and cannot remove these complex and complicated artifacts. Hence, this paper proposes an image deinterlacing network (DIN), which is specifically designed for joint removal of interlacing mixed with other artifacts. The DIN is composed of two stages,i.e., a cooperative vertical interpolation stage for splitting and fully using the information of adjacent fields, and a field-merging stage to perceive movements and suppress ghost artifacts. Experimental results demonstrate the effectiveness of the proposed DIN on both synthetic and real-world test sets.
Yang Zhao 0002, Wei Jia 0001, Ronggang Wang
IEEE Trans. Circuits Syst. Video Technol.2
2022 Multiframe Joint Enhancement for Early Interlaced Videos
abstract
Early interlaced videos usually contain multiple and interlacing and complex compression artifacts, which significantly reduce the visual quality. Although the high-definition reconstruction technology for early videos has made great progress in recent years, related research on deinterlacing is still lacking. Traditional methods mainly focus on simple interlacing mechanism, and cannot deal with the complex artifacts in real-world early videos. Recent interlaced video reconstruction deep deinterlacing models only focus on single frame, while neglecting important temporal information. Therefore, this paper proposes a multiframe deinterlacing network joint enhancement network for early interlaced videos that consists of three modules, i.e., spatial vertical interpolation module, temporal alignment and fusion module, and final refinement module. The proposed method can effectively remove the complex artifacts in early videos by using temporal redundancy of multi-fields. Experimental results demonstrate that the proposed method can recover high quality results for both synthetic dataset and real-world early interlaced videos. At the same time, the method also won the first place in the MSU Deinterlacer Benchmark. The code is available at: https://github.com/anymyb/MFDIN.
Yang Zhao 0002, Yanbo Ma, Yuan Chen 0012, Wei Jia 0001, Ronggang Wang, Xiaoping Liu 0003
IEEE Trans. Image Process.4
2022 Audio Matters in Video Super-Resolution by Implicit Semantic Guidance
abstract
Video super-resolution (VSR) aims to use multiple consecutive low-resolution frames to recover the corresponding high-resolution frames. However, existing VSR methods only consider videos as image sequences, ignoring another essential timing informationaudio, while in fact, there is a semantic link between audio and vision, and extensive studies have shown that audio can provide supervisory information in visual networks. Meanwhile, the addition of semantic priors has been proven to be effective in super-resolution (SR) tasks, but a pretrained segmentation network is required to obtain semantic segmentation maps. By contrast, audio as the information contained in the video itself can be directly used. Therefore, in this study, we propose a novel and pluggable multiscale audiovisual fusion (MS-AVF) module to enhance VSR performance by exploiting the relevant audio information, which can be regarded as implicit semantic guidance compared with the kind of explicit segmentation priors. Specifically, we first fuse audiovisual features on the semantic feature maps of different granularities of the target frames, and then through a top-down multiscale fusion approach, feedback high-level semantics to the underlying global visual features layer by layer, thereby providing effective audio implicit semantic guidance for VSR. Experimental results show that audio can further improve the VSR effect. Moreover, by visualizing the learned attention mask, the proposed end-to-end model can automatically learn potential audiovisual semantic links, especially improving the accuracy and effectiveness of the SR of sound sources and their surrounding regions.
Meibin Qi, Yang Zhao 0002, Wei Jia 0001, Ronggang Wang
IEEE Trans. Multim.5
2021 Compact Double Attention Module Embedded CNN for Palmprint Recognition
Yongmin Zheng, Lunke Fei, Wei Jia 0001, Jie Wen 0001, Shaohua Teng, Imad Rida
CGI3
2021 Bilateral Grid Learning for Stereo Matching Networks
abstract
Real-time performance of stereo matching networks is important for many applications, such as automatic driving, robot navigation and augmented reality (AR). Although significant progress has been made in stereo matching networks in recent years, it is still challenging to balance real-time performance and accuracy. In this paper, we present a novel edge-preserving cost volume upsampling module based on the slicing operation in the learned bilateral grid. The slicing layer is parameter-free, which allows us to obtain a high quality cost volume of high resolution from a low-resolution cost volume under the guide of the learned guidance map efficiently. The proposed cost volume upsampling module can be seamlessly embedded into many existing stereo matching networks, such as GCNet, PSMNet, and GANet. The resulting networks are accelerated several times while maintaining comparable accuracy. Furthermore, we design a real-time network (named BGNet) based on this module, which outperforms existing published real-time deep stereo matching networks, as well as some complex networks on the KITTI stereo datasets. The code is available at https://github.com/YuhuaXu/BGNet.
Yuhua Xu 0006, Wei Jia 0001, Yulan Guo
CVPR4
2021 Deep Multi-loss Hashing Network for Palmprint Retrieval and Recognition
abstract
With the wide application of biometrics technology, the scale of biometrics databases is increasing rapidly. In this situation, fast retrieval technology is more and more necessary for large-scale biometrics retrieval and recognition. Palmprint recognition is one of the emerging biometrics technologies. However, the research on fast palmprint retrieval algorithm is still preliminary. Hashing is one of the most popular image retrieval technologies due to its fast speed and low storage cost. In this paper, we propose a new deep palmprint hashing method, which integrates classification loss, pairing loss and quantization loss in a unified deep learning framework. Experimental results show that the proposed deep multi-loss hashing method has better performance for palmprint recognition and retrieval than other existing classic hashing methods.
Wei Jia 0001, Shuwei Huang, Lunke Fei, Yang Zhao 0002, Hai Min
IJCB1
2021 Discrete semantic embedding hashing for scalable cross-modal retrieval
abstract
Cross-modal hashing has attracted much attention for cross-modal retrieval and achieved promising performance due to its powerful capacity. Some existing cross-modal hashing methods construct pairwise similarities to represent the relationship of heterogeneous data, which require much computation time and storage space, making them unscalable for large-scale retrieval tasks. In this paper, we propose a novel supervised Discrete Semantic Embedding Hashing (DSEH) for cross-modal retrieval. Specifically, we first learn the common representation of heterogeneous data by embedding the semantic labels into a collective matrix factorization, such that both intra- and inter-modality similarities can be well captured. Then, we learn the hash codes in the discrete space based on the learned common representation via an orthogonal rotation technique. Moreover, we learn the multi-modal hash functions that can efficiently convert out-of-sample instances into unified hash codes. Extensive experimental results on three widely used benchmark databases demonstrate the superiority of the proposed DSEH compared with previous state-of-the-arts.
Lunke Fei, Wei Jia 0001, Shuping Zhao, Jie Wen 0001, Shaohua Teng, Wei Zhang 0005
SMC3
2021 Real-time automatic helmet detection of motorcyclists in urban traffic using improved YOLOv5 detector
abstract
Abstract In traffic accidents, motorcycle accidents are the main cause of casualties, especially in developing countries. The main cause of fatal injuries in motorcycle accidents is that motorcycle riders or passengers do not wear helmets. In this paper, an automatic helmet detection of motorcyclists method based on deep learning is presented. The method consists of two steps. The first step uses the improved YOLOv5 detector to detect motorcycles (including motorcyclists) from video surveillance. The second step takes the motorcycles detected in the previous step as input and continues to use the improved YOLOv5 detector to detect whether the motorcyclists wear helmets. The improvement of the YOLOv5 detector includes the fusion of triplet attention and the use of soft‐NMS instead of NMS. A new motorcycle helmet dataset (HFUT‐MH) is being proposed, which is larger and more comprehensive than the existing dataset derived from multiple traffic monitoring in Chinese cities. Finally, the proposed method is verified by experiments and compared with other state‐of‐the‐art methods. Our method achieves mAP of 97.7%, F1‐score of 92.7% and frames per second (FPS) of 63, which outperforms other state‐of‐the‐art detection methods.
Wei Jia 0001, Shiquan Xu, Yang Zhao 0002, Hai Min, Shujie Li 0002
IET Image Process.1
2021 A survey on dorsal hand vein biometrics
Wei Jia 0001, Bob Zhang 0001, Yang Zhao 0002, Lunke Fei, Wenxiong Kang, Di Huang 0001, Guodong Guo
Pattern Recognit.1
2021 Lighter but Efficient Bit-Depth Expansion Network
abstract
With the development of display technology, bit-depth expansion (BDE) has emerged as a basic process to display low-bit-depth image and video resources on high-bit-depth monitors. Most current BDE methods are based on traditional algorithms, and the few existing methods based on deep neural networks still suffer from loss of pixel-level details or from high computational cost. This paper proposes a lightweight but efficient BDE network that can effectively improve the capacity of shallow network by introducing a residual-block-in-residual-block structure. Furthermore, the proposed network adopts residual network architecture and dilated convolution to balance the preservation of pixel-level information and the expansion of the receptive field. Hence, the proposed method can also totally remove significant artifacts from very low-bit-depth images. Experimental results demonstrate that the proposed method can achieve performance comparable to or even better than that of some state-of-the-art methods while having much lighter architecture and fewer parameters.
Yang Zhao 0002, Ronggang Wang, Yuan Chen 0012, Wei Jia 0001, Xiaoping Liu 0003, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 AdvKin: Adversarial Convolutional Network for Kinship Verification
abstract
Kinship verification in the wild is an interesting and challenging problem. The goal of kinship verification is to determine whether a pair of faces are blood relatives or not. Most previous methods for kinship verification can be divided as handcrafted features-based shallow learning methods and convolutional neural network (CNN)-based deep-learning methods. Nevertheless, these methods are still facing the challenging task of recognizing kinship cues from facial images. The reason is that the family ID information and the distribution difference of pairwise kin-faces are rarely considered in kinship verification tasks. To this end, a family ID-based adversarial convolutional network (AdvKin) method focused on discriminative Kin features is proposed for both small-scale and large-scale kinship verification in this article. The merits of this article are four-fold: 1) for kin-relation discovery, a simple yet effective self-adversarial mechanism based on a negative maximum mean discrepancy (NMMD) loss is formulated as attacks in the first fully connected layer; 2) a pairwise contrastive loss and family ID-based softmax loss are jointly formulated in the second and third fully connected layer, respectively, for supervised training; 3) a two-stream network architecture with residual connections is proposed in AdvKin; and 4) for more fine-grained deep kin-feature augmentation, an ensemble of patch-wise AdvKin networks is proposed (E-AdvKin). Extensive experiments on 4 small-scale benchmark KinFace datasets and 1 large-scale families in the wild (FIW) dataset from the first Large-Scale Kinship Recognition Data Challenge, show the superiority of our proposed AdvKin model over other state-of-the-art approaches.
Lei Zhang 0038, Qingyan Duan, David Zhang 0001, Wei Jia 0001, Xizhao Wang
IEEE Trans. Cybern.4
2021 Accurate 3-D Reconstruction Under IoT Environments and Its Applications to Augmented Reality
abstract
With the remarkable development of sensor devices and the Internet of Things (IoT), today's researchers can easily know what changes have taken place in the real world by acquiring a 3-D model. Conversely, a large amount of image data promotes the development of perceptual computing technology. In this article, we focus on modeling 3-D scenes from the multisource image data obtained from the IoT with cameras. Although great progress has been made in 3-D reconstruction, it is still challenging to recover the 3-D model from IoT data because the captured images are usually noisy, incomplete, varying scale, and with repetitive structures or features. In this article, we propose an accurate 3-D reconstruction method under IoT environments for perceptual computing of the scene. This method consists of sparse, dense, and surface reconstruction processes, which can gradually recover high-quality geometric models from the image data and efficiently deal with various repetitive structures. By analyzing the reconstructed model, we can detect the changes of scenes. We evaluate the proposed method on the benchmark data sets (i.e., tanks and temples) and publicly available data sets(in which samples usually contain repeated structures, lighting change, and different scales). Experimental results show that the proposed method outperforms the state-of-the-art methods according to the standard evaluation metric. We also use our method to enhance the real scenes with virtual objects, thus producing promising results.
Mingwei Cao, Liping Zheng, Wei Jia 0001, Huimin Lu 0001, Xiaoping Liu 0003
IEEE Trans. Ind. Informatics3
2021 Joint 3D Reconstruction and Object Tracking for Traffic Video Analysis Under IoV Environment
abstract
Benefits from artificial intelligence and the Internet of Vehicles (IoV), Management of modern transportation have great progress, especially in urban areas. However, traditional traffic video analysis and visualization are usually conducted in offsite and textural environments, i.e., text and number, which do not promote user's sensorial perception and interaction. Thus, the problem that how to use modern novel techniques to analyze traffic video for improving intelligent transportation is so emergency. In this paper, we introduce a joint 3D reconstruction and object tracking approach to traffic video analysis under the IoV environment, which is an integrative framework and consists of 3D reconstruction, object detection, and visual tracking. The 3D reconstruction system is connected to the Internet of Vehicles and integrated into the system to retrieve image data for recovering the 3D model of vehicles, and then, visualizing vehicle trajectories in real-time by augmented reality. And the system can also locate the vehicle's position in real-time. The experiments in both laboratory and practice show great feedback, which will effectively contribute to intelligent transportation.
Mingwei Cao, Liping Zheng, Wei Jia 0001, Xiaoping Liu 0003
IEEE Trans. Intell. Transp. Syst.3
2021 A Multilayer Pyramid Network Based on Learning for Vehicle Logo Recognition
abstract
In this paper, we present a novel learning-based scheme for vehicle logo recognition (VLR). This scheme is termed Multilayer Pyramid Network Based on Learning (MLPNL) and is based on the principle that considering multiple resolutions is helpful for extracting valuable features that benefit the final recognition performance. The innovations of this scheme include (1) a multilayer pyramid network, with pixel difference matrices (PDMs) as its input and output and feature parameters mapping one PDM to another; (2) an objective function and a corresponding optimization method designed to facilitate the learning of the feature parameters of the proposed multilayer pyramid network; and (3) a multi-codebook-based encoding method that makes best use of the features extracted from PDMs corresponding to different resolutions. Extensive experiments conducted with an open dataset, HFUT-VL, demonstrate that the proposed MLPNL scheme outperforms state-of-the-art handcrafted descriptors and non-deep-learning-based learning methods when fewer training samples exist. Experiments conducted with a benchmark dataset, XMU, demonstrate that MLPNL outperforms existing state-of-the-art VLR methods. Experiments conducted both on HFUT-VL and XMU demonstrate that MLPNL is faster than most deep-learning-based learning methods while maintaining nearly the same recognition rate. Code has been made available at:https://github.com/HFUT-CV/MLPNL.
Jun Wang 0071, Hai Min, Wei Jia 0001, Jun Yu 0001, Chang Wen Chen
IEEE Trans. Intell. Transp. Syst.5
2021 Learning Compact Multifeature Codes for Palmprint Recognition From a Single Training Image per Palm
abstract
In this article, we propose a multifeature learning method to jointly learn compact multifeature codes (LCMFCs) for palmprint recognition with a single training sample per palm. Unlike most existing hand-crafted methods that extract single-type features from raw pixels, we first form the multi-type data vectors such as the direction-data, and texture-data to completely sample the multiple information of a palmprint image. Then, we learn the discriminative multifeatures from multi-type data vectors by maximizing the inter-palm distance, and minimizing the energy loss between the learned codes, and the original data. Moreover, our LCMFC method adaptively learns the optimal weights of multi-type features to jointly learn the compact multifeature codes. Finally, we cluster the nonoverlapping blockwise histograms of the compact multifeature codes into a feature vector for palmprint representation. Extensive experimental results on six benchmark palmprint databases are presented to show the effectiveness of the proposed method.
Lunke Fei, Bob Zhang 0001, Lin Zhang 0014, Wei Jia 0001, Jie Wen 0001, Jigang Wu
IEEE Trans. Multim.4
2020 Estimated Exposure Guided Reconstruction Model for Low-Light Image Enhancement
Xiaona Liu, Yang Zhao 0002, Yuan Chen 0012, Wei Jia 0001, Ronggang Wang, Xiaoping Liu 0003
PRCV (1)4
2020 Real-time video stabilization via camera path correction and its applications to augmented reality on edge devices
Mingwei Cao, Liping Zheng, Wei Jia 0001, Xiaoping Liu 0003
Comput. Commun.3
2020 Adversarial-learning-based image-to-image transformation: A survey
Yuan Chen 0012, Yang Zhao 0002, Wei Jia 0001, Xiaoping Liu 0003
Neurocomputing3
2020 Constructing big panorama from video sequence based on deep local feature
Mingwei Cao, Liping Zheng, Wei Jia 0001, Xiaoping Liu 0003
Image Vis. Comput.3
2020 CAM: A fine-grained vehicle model recognition method based on visual attention model
Longdao Xu, Wei Jia 0001, Wenjia Zhu, Yunxiang Fu, Qiang Lu 0002
Image Vis. Comput.3
2020 Blind Quality Assessment for Cartoon Images
abstract
Current blind image quality assessment (BIQA) algorithms are mainly designed for natural images. Unfortunately, cartoon and cartoon-like images are quite different from natural images. Hence, recent BIQA methods are not very robust to cartoon images. In this paper, we propose a specific BIQA algorithm designed for cartoon images, which consists of the following terms. First, a cartoon image is divided into edge areas and nonedge areas via a Tchebichef moment (TM)-based process. Second, a multiorder sharpness statistic term is used to measure the quality of the edges, and a sharpness statistic prior model of high-quality (HQ) cartoon images is built. Finally, a local encoding statistic term is adopted to describe the textural complexity in the nonedge areas, and a texture statistic prior model is also established. The experimental results on the cartoon image datasets demonstrate that the proposed method can accurately evaluate the visual quality of cartoon images and is more suitable for cartoon scenarios than some traditional BIQA algorithms.
Yuan Chen 0012, Yang Zhao 0002, Shujie Li 0002, Wangmeng Zuo, Wei Jia 0001, Xiaoping Liu 0003
IEEE Trans. Circuits Syst. Video Technol.5
2020 Local Discriminant Direction Binary Pattern for Palmprint Representation and Recognition
abstract
Direction-based methods are the most powerful and popular palmprint recognition methods. However, there is no existing work that completely analyzes the essential differences among different direction-based methods and explores the most discriminant direction representation of a palmprint. In this paper, we attempt to establish the connection between the direction feature extraction model and the discriminability of direction features, and we propose a novel exponential and Gaussian fusion model (EGM) to characterize the discriminative power of different directions. The EGM can provide us with a new insight into the optimal direction feature selection of palmprints. Moreover, we propose a local discriminant direction binary pattern (LDDBP) to completely represent the direction features of a palmprint. Guided by the EGM, the most discriminant directions can be exploited to form the LDDBP-based descriptor for palmprint representation and recognition. Extensive experiment results conducted on four widely used palmprint databases demonstrate the superiority of the proposed LDDBP method over the state-of-the-art direction-based methods.
Lunke Fei, Bob Zhang 0001, Yong Xu 0001, Di Huang 0001, Wei Jia 0001, Jie Wen 0001
IEEE Trans. Circuits Syst. Video Technol.5
2019 Precision direction and compact surface type representation for 3D palmprint identification
Lunke Fei, Bob Zhang 0001, Yong Xu 0001, Wei Jia 0001, Jie Wen 0001, Jigang Wu
Pattern Recognit.4
2019 From Noise to Feature: Exploiting Intensity Distribution as a Novel Soft Biometric Trait for Finger Vein Recognition
abstract
Most finger vein feature extraction algorithms achieve satisfactory performance due to their texture representation abilities, despite simultaneously ignoring the intensity distribution that is formed by the finger tissue, and in some cases, processing it as background noise. In this paper, we exploit this kind of “noise” as a novel soft biometric trait for achieving better finger vein recognition performance. First, a detailed analysis of the finger vein imaging principle and the characteristics of the image are presented to show that the intensity distribution that is formed by the finger tissue in the background can be extracted as a soft biometric trait for recognition. Then, two finger vein background layer extraction algorithms and three soft biometric trait extraction algorithms are proposed for intensity distribution feature extraction. Finally, a hybrid matching strategy is proposed to solve the issue of dimension difference between the primary and soft biometric traits on the score level. A series of rigorous contrast experiments on three open-access databases demonstrate that our proposed method is feasible and effective for finger vein recognition.
Wenxiong Kang, Wei Jia 0001
IEEE Trans. Inf. Forensics Secur.4
2019 Deep Reconstruction of Least Significant Bits for Bit-Depth Expansion
abstract
Bit-depth expansion (BDE) is important for displaying a low bit-depth image in a high bit-depth monitor. Current BDE algorithms often utilize traditional methods to fill the missing least significant bits and suffer from multiple kinds of perceivable artifacts. In this paper, we present a deep residual network-based method for BDE. Based on the different properties of flat and non-flat areas, two channels are proposed to reconstruct these two kinds of areas, respectively. Moreover, a simple yet efficient local adaptive adjustment preprocessing is presented in the flat-area-channel. By combining the benefits of both the traditional debanding strategy and network-based reconstruction, the proposed method can further promote the subjective quality of the flat area. Experimental results on several image sets demonstrate that the proposed BDE network can obtain favorable visual quality as well as decent quantitative performance.
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Wangmeng Zuo, Xiaoping Liu 0003, Wen Gao 0001
IEEE Trans. Image Process.3
2019 Learning Discriminant Direction Binary Palmprint Descriptor
abstract
Palmprint directions have been proved to be one of the most effective features for palmprint recognition. However, most existing direction-based palmprint descriptors are hand-craft designed and require strong prior knowledge. In this paper, we propose a discriminant direction binary code (DDBC) learning method for palmprint recognition. Specifically, for each palmprint image, we first calculate the convolutions of the direction-based templates and palmprint and form the informative convolution difference vectors by computing the convolution difference between the neighboring directions. Then, we propose a simple yet effective model to learn feature mapping functions that can project these convolution difference vectors into DDBCs. For all training samples: (1) the variance of the learned binary codes is maximized; (2) the intra-class distance of the binary codes is minimized; and (3) the inter-class distance of the binary codes is maximized. Finally, we cluster the block-wise histograms of DDBC forming the discriminant direction binary palmprint descriptor for palmprint recognition. The experimental results on four challenging contactless palmprint databases clearly demonstrate the effectiveness of the proposed method.
Lunke Fei, Bob Zhang 0001, Yong Xu 0001, Zhenhua Guo 0001, Jie Wen 0001, Wei Jia 0001
IEEE Trans. Image Process.6
2019 Feature Extraction Methods for Palmprint Recognition: A Survey and Evaluation
abstract
Palmprint processes a number of unique features for reliable personal recognition. However, different types of palmprint images contain different dominant features. Instead, only some features of the palmprint are visible in a palmprint image, whereas the other features may not be notable. For example, the low-resolution palmprint image has visible principal lines and wrinkles. By contrast, the high-resolution palmprint image contains clear ridge patterns and minutiae points. In addition, the three dimensional (3-D) palmprint image possesses curvatures of the palmprint surface. So far, there is no work to summarize the feature extraction of different types of palmprint images. In this paper, we have an aim to completely study the feature extraction and recognition of palmprint. We propose to use a unified framework to classify palmprint images into four categories: (1) the contact-based; (2) contactless; (3) high-resolution; and (4) 3-D palmprint images. Then, we analyze the motivations and theories of the representative extraction and matching methods for different types of palmprint images. Finally, we compare and test the state-of-the-art methods via the widely used palmprint databases, and point out some potential directions for future research.
Lunke Fei, Guangming Lu 0002, Wei Jia 0001, Shaohua Teng, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2018 An effective local regional model based on salient fitting for image segmentation
Hai Min, Wei Jia 0001, Yang Zhao 0002, Yue-Tong Luo
Neurocomputing3
2018 Local patch encoding-based method for single image super-resolution
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Jianchao Yang, Wenmin Wang 0001, Wen Gao 0001
Inf. Sci.3
2018 Evaluation of Local Features for Structure from Motion
Mingwei Cao, Wei Jia 0001, Yujie Li 0001, Zhihan Lyu, Liping Zheng, Xiaoping Liu 0003
Multim. Tools Appl.3
2018 A polynomial piecewise constant approximation method based on dual constraint relaxation for segmenting images with intensity inhomogeneity
Hai Min, Wei Jia 0001, Yang Zhao 0002, Yue-Tong Luo
Pattern Recognit.2
2018 Finger Vein Presentation Attack Detection Using Total Variation Decomposition
abstract
Finger vein recognition is an emerging biometric technique for personal authentication that has garnered considerable attention in the past decade. Although shown to be effective, recent studies have revealed that finger vein biometrics is also vulnerable to presentation attacks, i.e., printed versions of authorized individual finger vein images can be used to gain access to facilities or services. In this paper, given that both blurriness and the noise distribution are slightly different between real and forged finger vein images, we propose an efficient and robust method for detecting presentation attacks that use forged finger vein images (print artifacts). First, we use total variation regularization to decompose original finger vein images into structure and noise components, which represent the degrees of blurriness and the noise distribution. Second, a block local binary pattern descriptor is used to encode both structure and noise information in the decomposed components. Finally, we use a cascaded support vector machine model for classification, by which finger vein presentation attacks can be effectively detected. To evaluate the performance of our approach, we constructed a new finger vein presentation attack database. Extensive experimental results gleaned from the two finger vein presentation attack databases and a palm vein presentation attack database show that our method clearly outperforms state-of-the-art methods.
Xinwei Qiu, Wenxiong Kang, Senping Tian, Wei Jia 0001, Zhixing Huang
IEEE Trans. Inf. Forensics Secur.4
2018 LATE: A Level-Set Method Based on Local Approximation of Taylor Expansion for Segmenting Intensity Inhomogeneous Images
abstract
Intensity inhomogeneity is common in real-world images and inevitably leads to many difficulties for accurate image segmentation. Numerous level-set methods have been proposed to segment images with intensity inhomogeneity. However, most of these methods are based on linear approximation, such as locally weighted mean, which may cause problems when handling images with severe intensity inhomogeneities. In this paper, we view segmentation of such images as a nonconvex optimization problem, since the intensity variation in such an image follows a nonlinear distribution. Then, we propose a novel level-set method named local approximation of Taylor expansion (LATE), which is a nonlinear approximation method to solve the nonconvex optimization problem. In LATE, we use the statistical information of the local region as a fidelity term and the differentials of intensity inhomogeneity as an adjusting term to model the approximation function. In particular, since the first-order differential is represented by the variation degree of intensity inhomogeneity, LATE can improve the approximation quality and enhance the local intensity contrast of images with severe intensity inhomogeneity. Moreover, LATE solves the optimization of function fitting by relaxing the constraint condition. In addition, LATE can be viewed as a constraint relaxation of classical methods, such as the region-scalable fitting model and the local intensity clustering model. Finally, the level-set energy functional is constructed based on the Taylor expansion approximation. To validate the effectiveness of our method, we conduct thorough experiments on synthetic and real images. Experimental results show that the proposed method clearly outperforms other solutions in comparison.
Hai Min, Wei Jia 0001, Yang Zhao 0002, Wangmeng Zuo, Haibin Ling, Yue-Tong Luo
IEEE Trans. Image Process.2
2018 F-SVM: Combination of Feature Transformation and SVM Learning via Convex Relaxation
abstract
The generalization error bound of the support vector machine (SVM) depends on the ratio of the radius and margin. However, conventional SVM only considers the maximization of the margin but ignores the minimization of the radius, which restricts its performance when applied to joint learning of feature transformation and the SVM classifier. Although several approaches have been proposed to integrate the radius and margin information, most of them either require the form of the transformation matrix to be diagonal, or are nonconvex and computationally expensive. In this paper, we suggest a novel approximation for the radius of the minimum enclosing ball in feature space, and then propose a convex radius-margin-based SVM model for joint learning of feature transformation and the SVM classifier, i.e., F-SVM. A generalized block coordinate descent method is adopted to solve the F-SVM model, where the feature transformation is updated via the gradient descent and the classifier is updated by employing the existing SVM solver. By incorporating with kernel principal component analysis, F-SVM is further extended for joint learning of nonlinear transformation and the classifier. F-SVM can also be incorporated with deep convolutional networks to improve image classification performance. Experiments on the UCI, LFW, MNIST, CIFAR-10, CIFAR-100, and Caltech101 data sets demonstrate the effectiveness of F-SVM.
Xiaohe Wu, Wangmeng Zuo, Liang Lin 0004, Wei Jia 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2017 Iterative projection reconstruction for fast and efficient image upsampling
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Wenmin Wang 0001, Wen Gao 0001
Neurocomputing3
2017 Robust bundle adjustment for large-scale structure from motion
Mingwei Cao, Wei Jia 0001, Shanglin Li, Xiaoping Liu 0003
Multim. Tools Appl.3
2017 Erratum to "Local line directional pattern for palmprint recognition" [Pattern Recognit. 50(2016) 26-44]
Yue-Tong Luo, Lan-Ying Zhao, Bob Zhang 0001, Wei Jia 0001, Feng Xue 0002, Yihai Zhu, Bing-Qing Xu
Pattern Recognit.4
2017 Palmprint Recognition Based on Complete Direction Representation
abstract
Direction information serves as one of the most important features for palmprint recognition. In the past decade, many effective direction representation (DR)-based methods have been proposed and achieved promising recognition performance. However, due to an incomplete understanding for DR, these methods only extract DR in one direction level and one scale. Hence, they did not fully utilize all potentials of DR. In addition, most researchers only focused on the DR extraction in spatial coding domain, and rarely considered the methods in frequency domain. In this paper, we propose a general framework for DR-based method named complete DR (CDR), which reveals DR by a comprehensive and complete way. Different from traditional methods, CDR emphasizes the use of direction information with strategies of multi-scale, multi-direction level, multi-region, as well as feature selection or learning. This way, CDR subsumes previous methods as special cases. Moreover, thanks to its new insight, CDR can guide the design of new DR-based methods toward better performance. Motived this way, we propose a novel palmprint recognition algorithm in frequency domain. First, we extract CDR using multi-scale modified finite radon transformation. Then, an effective correlation filter, namely, band-limited phase-only correlation, is explored for pattern matching. To remove feature redundancy, the sequential forward selection method is used to select a small number of CDR images. Finally, the matching scores obtained from different selected features are integrated using score-level-fusion. Experiments demonstrate that our method can achieve better recognition accuracy than the other state-of-the-art methods. More importantly, it has fast matching speed, making it quite suitable for the large-scale identification applications.
Wei Jia 0001, Bob Zhang 0001, Yihai Zhu, Yang Zhao 0002, Wangmeng Zuo, Haibin Ling
IEEE Trans. Image Process.1
2016 A novel dual minimization based level set method for image segmentation
Hai Min, De-Shuang Huang, Wei Jia 0001
Neurocomputing4
2016 Design of digital camouflage by recursive overlapping of pattern templates
Feng Xue 0002, Yue-Tong Luo, Wei Jia 0001
Neurocomputing4
2016 Camouflage performance analysis and evaluation framework based on features fusion
Feng Xue 0002, Chengxi Yong, Yue-Tong Luo, Wei Jia 0001
Multim. Tools Appl.6
2016 Local line directional pattern for palmprint recognition
Yue-Tong Luo, Lan-Ying Zhao, Bob Zhang 0001, Wei Jia 0001, Feng Xue 0002, Yihai Zhu, Bing-Qing Xu
Pattern Recognit.4
2015 An Intensity-Texture model based level set method for image segmentation
Hai Min, Wei Jia 0001, Yang Zhao 0002, Rong-Xiang Hu, Yue-Tong Luo, Feng Xue 0002
Pattern Recognit.2
2014 Angular Pattern and Binary Angular Pattern for Shape Retrieval
abstract
In this paper, we propose two novel shape descriptors, angular pattern (AP) and binary angular pattern (BAP), and a multiscale integration of them for shape retrieval. Both AP and BAP are intrinsically invariant to scale and rotation. More importantly, being global shape descriptors, the proposed shape descriptors are computationally very efficient, while possessing similar discriminability as state-of-the-art local descriptors. As a result, the proposed approach is attractive for real world shape retrieval applications. The experiments on the widely used MPEG-7 and TARI-1000 data sets demonstrate the effectiveness of the proposed method in comparison with existing methods.
Rong-Xiang Hu, Wei Jia 0001, Haibin Ling, Yang Zhao 0002, Jie Gui
IEEE Trans. Image Process.2
2014 Histogram of Oriented Lines for Palmprint Recognition
abstract
Subspace learning methods are very sensitive to the illumination, translation, and rotation variances in image recognition. Thus, they have not obtained promising performance for palmprint recognition so far. In this paper, we propose a new descriptor of palmprint named histogram of oriented lines (HOL), which is a variant of histogram of oriented gradients (HOG). HOL is not very sensitive to changes of illumination, and has the robustness against small transformations because slight translations and rotations make small histogram value changes. Based on HOL, even some simple subspace learning methods can achieve high recognition rates.
Wei Jia 0001, Rong-Xiang Hu, Ying-Ke Lei, Yang Zhao 0002, Jie Gui
IEEE Trans. Syst. Man Cybern. Syst.1
2013 Completed robust local binary pattern for texture classification
Yang Zhao 0002, Wei Jia 0001, Rong-Xiang Hu, Hai Min
Neurocomputing2
2012 Newborn footprint recognition using orientation feature
Wei Jia 0001, Hai-Yang Cai, Jie Gui, Rong-Xiang Hu, Ying-Ke Lei
Neural Comput. Appl.1
2012 Discriminant sparse neighborhood preserving embedding for face recognition
Jie Gui, Zhenan Sun, Wei Jia 0001, Rong-Xiang Hu, Ying-Ke Lei, Shuiwang Ji
Pattern Recognit.3
2012 Perceptually motivated morphological strategies for shape retrieval
Rong-Xiang Hu, Wei Jia 0001, Yang Zhao 0002, Jie Gui
Pattern Recognit.2
2012 Hand shape recognition based on coherent distance shape contexts
Rong-Xiang Hu, Wei Jia 0001, David Zhang 0001, Jie Gui, Liang-Tu Song
Pattern Recognit.2
2012 Robust Classification Method of Tumor Subtype by Using Correlation Filters
abstract
Tumor classification based on gene expression profiles, which is of great benefit to the accurate diagnosis and personalized treatment for different types of tumor, has drawn a great attention in recent years. This paper proposes a novel tumor classification method based on correlation filters to identify the overall pattern of tumor subtype hidden in differentially expressed genes. Concretely, two correlation filters, i.e., Minimum Average Correlation Energy (MACE) and Optimal Tradeoff Synthetic Discriminant Function (OTSDF), are introduced to determine whether a test sample matches the templates synthesized for each subclass. The experiments on six publicly available datasets indicate that the proposed method is robust to noise, and can more effectively avoid the effects of dimensionality curse. Compared with many model-based methods, the correlation filter based method can achieve better performance when balanced training sets are exploited to synthesize the templates. Particularly, the proposed method can detect the similarity of overall pattern while ignoring small mismatches between test sample and the synthesized template. And it performs well even if only few training samples are available. More importantly, the experimental results can be visually represented, which is helpful for the further analysis of results.
Shu-Ling Wang, Yihai Zhu, Wei Jia 0001, De-Shuang Huang
IEEE ACM Trans. Comput. Biol. Bioinform.3
2012 Multiscale Distance Matrix for Fast Plant Leaf Recognition
abstract
In this brief, we propose a novel contour-based shape descriptor, called the multiscale distance matrix, to capture the shape geometry while being invariant to translation, rotation, scaling, and bilateral symmetry. The descriptor is further combined with a dimensionality reduction to improve its discriminative power. The proposed method avoids the time-consuming pointwise matching encountered in most of the previously used shape recognition algorithms. It is therefore fast and suitable for real-time applications. We applied the proposed method to the task of plan leaf recognition with experiments on two data sets, the Swedish Leaf data set and the ICL Leaf data set. The experimental results clearly demonstrate the effectiveness and efficiency of the proposed descriptor.
Rong-Xiang Hu, Wei Jia 0001, Haibin Ling, De-Shuang Huang
IEEE Trans. Image Process.2
2012 Completed Local Binary Count for Rotation Invariant Texture Classification
abstract
In this brief, a novel local descriptor, named local binary count (LBC), is proposed for rotation invariant texture classification. The proposed LBC can extract the local binary grayscale difference information, and totally abandon the local binary structural information. Although the LBC codes do not represent visual microstructure, the statistics of LBC features can represent the local texture effectively. In addition, a completed LBC (CLBC) is also proposed to enhance the performance of texture classification. Experimental results obtained from three databases demonstrate that the proposed CLBC can achieve comparable accurate classification rates with completed local binary pattern.
Yang Zhao 0002, De-Shuang Huang, Wei Jia 0001
IEEE Trans. Image Process.3
2011 Robust Gait Recognition Using Gait Energy Image and Band-Limited Phase-Only Correlation
Wei Jia 0001, Huanglin Zeng
ICIC (1)2
2011 Orthogonal local spline discriminant projection with application to face recognition
Ying-Ke Lei, Zhiguo Ding 0004, Rong-Xiang Hu, Shanwen Zhang, Wei Jia 0001
Pattern Recognit. Lett.5
2010 Newborn Footprint Recognition Using Subspace Learning Methods
Wei Jia 0001, Jie Gui, Rong-Xiang Hu, Ying-Ke Lei, Xue-Yang Xiao
ICIC (1)1
2010 Locality preserving discriminant projections for face and palmprint recognition
Jie Gui, Wei Jia 0001, Shu-Ling Wang, De-Shuang Huang
Neurocomputing2
2010 Maximum margin criterion with tensor representation
Rong-Xiang Hu, Wei Jia 0001, De-Shuang Huang, Ying-Ke Lei
Neurocomputing2
2009 Gait Recognition Using Hough Transform and Principal Component Analysis
Ling-Feng Liu, Wei Jia 0001, Yihai Zhu
ICIC (1)2
2009 Survey of Gait Recognition
Ling-Feng Liu, Wei Jia 0001, Yihai Zhu
ICIC (2)2
2009 Palmprint Recognition Using Band-Limited Phase-Only Correlation and Different Representations
Yihai Zhu, Wei Jia 0001, Ling-Feng Liu
ICIC (1)2
2009 Fast Palmprint Retrieval Using Principal Lines
abstract
In this paper, we propose a novel palmprint retrieval scheme based on principal lines. In the proposed scheme, the principal lines are firstly extracted by the modified finite radon transform. And then, a lot of key points located in three principal lines i.e. heart line, life line and head line are detected. Finally, the palmprints are retrieved by several key points' position and direction. The results of experiments conducted on PolyU palmprint database show that the proposed scheme is feasible and has high accurate retrieval rate with fast speed.
Wei Jia 0001, Yihai Zhu, Ling-Feng Liu, De-Shuang Huang
SMC1
2009 Efficient discovery of abundant post-translational modifications and spectral pairs using peptide mass and retention time differences
abstract
BACKGROUND: Peptide identification via tandem mass spectrometry is the basic task of current proteomics research. Due to the complexity of mass spectra, the majority of mass spectra cannot be interpreted at present. The existence of unexpected or unknown protein post-translational modifications is a major reason. RESULTS: This paper describes an efficient and sequence database-independent approach to detecting abundant post-translational modifications in high-accuracy peptide mass spectra. The approach is based on the observation that the spectra of a modified peptide and its unmodified counterpart are correlated with each other in their peptide masses and retention time. Frequently occurring peptide mass differences in a data set imply possible modifications, while small and consistent retention time differences provide orthogonal supporting evidence. We propose to use a bivariate Gaussian mixture model to discriminate modification-related spectral pairs from random ones. Due to the use of two-dimensional information, accurate modification masses and confident spectral pairs can be determined as well as the quantitative influences of modifications on peptide retention time. CONCLUSION: Experiments on two glycoprotein data sets demonstrate that our method can effectively detect abundant modifications and spectral pairs. By including the discovered modifications into database search or by propagating peptide assignments between paired spectra, an average of 10% more spectra are interpreted.
Wei Jia 0001, Zhuang Lu, Zuofei Yuan, Hao Chi, Liyun Xiu, Leheng Wang, Ruixiang Sun 0001, Wen Gao 0001, Xiaohong Qian, Simin He 0001
BMC Bioinform.2
2008 Palmprint identification based on directional representation
abstract
In this paper, we propose a novel approach for palmprint identification, which contains two interesting components. Firstly, we propose the directional representation for appearance based approaches. The new representation is robust to drastic illumination changes and preserves important discriminative information for classification. We then generate virtual samples to enlarge the training set to compensate for matching errors caused by large rotations and translations. Based on these two strategies, the recognition performance of representative appearance based approaches can be improved significantly. Secondly, in order to improve the robustness of palmprint identification, we propose a fusion method combining proposed method and orientation based approaches e.g. Competitive Code, which can obtain very low Equal Error Rates.
Wei Jia 0001, De-Shuang Huang, Dacheng Tao, David Zhang 0001
SMC1
2008 Palmprint verification based on principal lines
De-Shuang Huang, Wei Jia 0001, David Zhang 0001
Pattern Recognit.2
2008 Palmprint verification based on robust line orientation code
Wei Jia 0001, De-Shuang Huang, David Zhang 0001
Pattern Recognit.1
2007 Palmprint Verification Based on Robust Orientation Code
abstract
In palmprint recognition field, orientation based approaches are thought to achieve the best results in terms of recognition rates. In this paper, we propose a novel orientation based scheme, in which three strategies, the modified finite Radon transform, enlarged training set and pixel to area matching, have been designed to further improve its performance. The experimental results of verification conducted on Hong Kong Polytechnic University Palmprint Database show that our approach has higher recognition rates and faster processing speed.
Wei Jia 0001, De-Shuang Huang
IJCNN1
2007 Palmprint recognition with 2DPCA+PCA based on modular neural networks
Zhong-Qiu Zhao, De-Shuang Huang, Wei Jia 0001
Neurocomputing3