EDBT 2026 Demo / reviewers in the wild / expert
Wenxiong Kang
dblp:75/1559
· DBLP profile ↗
103ranked-venue papers
7as first author
76since 2021 · last 2026
0000-0001-9023-7252ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 1 first-author · 46 since 2021Artificial intelligence and machine learning · 42 · 1 first-author · 30 since 2021Security and privacy · 24 · 3 first-author · 19 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pb4U-GNet: Resolution-Adaptive Garment Simulation via Propagation-before-Update Graph NetworkabstractGarment simulation is fundamental to various applications in computer vision and graphics, from virtual try-on to digital human modelling. However, conventional physics-based methods remain computationally expensive, hindering their application in time-sensitive scenarios. While graph neural networks (GNNs) offer promising acceleration, existing approaches exhibit poor cross-resolution generalisation, demonstrating significant performance degradation on higher-resolution meshes beyond the training distribution. This stems from two key factors: (1) existing GNNs employ fixed message-passing depth that fails to adapt information aggregation to mesh density variation, and (2) vertex-wise displacement magnitudes are inherently resolution-dependent in garment simulation. To address these issues, we introduce Propagation-before-Update Graph Network (Pb4U-GNet), a resolution-adaptive framework that decouples message propagation from feature updates. Pb4U-GNet incorporates two key mechanisms: (1) dynamic propagation depth control, adjusting message-passing iterations based on mesh resolution, and (2) geometry-aware update scaling, which scales predictions according to local mesh characteristics. Extensive experiments show that even trained solely on low-resolution meshes, Pb4U-GNet exhibits strong generalisability across diverse mesh resolutions, addressing a fundamental challenge in neural garment simulation. Aoran Liu, Kun Hu 0008, Clinton Mo, Qiuxia Wu, Wenxiong Kang, Zhiyong Wang 0001 |
AAAI | 5 |
| 2026 | Knowledge-guided lightweight vision transformer with circular relative positional encoding for condition identification of industrial rotary kilns
Hao Wang 0234, Wenxiong Kang, Xiaojun Liang, Cao Liu, Chunhua Yang 0001, Weihua Gui 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Enhancing GNN learning with node augmentation
Maria Marrium, Arif Mahmood, Muhammad Haris Khan, M. Saad Shakeel, Wenxiong Kang |
Neural Networks | 5 |
| 2026 | 4DStyleGaussian: Generalizable 4D style transfer with Gaussian splatting
Wanlin Liang, Wenxiong Kang |
Pattern Recognit. | 5 |
| 2026 | LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding
Guocan Zhao, Wenxiong Kang |
Pattern Recognit. | 4 |
| 2026 | HKNet: Rethinking palm vein recognition framework with hybrid key set score
Tianming Xie, Wenxiong Kang |
Pattern Recognit. | 2 |
| 2026 | MCCENet: Multimodal Contrastive Learning Channel-Exchanging Networks for Palm Multimodal AuthenticationabstractA straightforward method for multimodal palm-based authentication is to integrate palm shape into the system, which enhances reliability, security, and accuracy compared to unimodal methods. However, most existing methods rely on handcrafted feature extraction, which fails to fully exploit palm shape information. Moreover, there have been limited attempts to apply deep learning-based methods in this field. This paper explores a deep multimodal fusion method of palm vein (PV) and palm shape (PS) for authentication called multimodal contrastive learning channel-exchanging networks (MCCENet) to better utilize palm shape contour information. Specifically, we observe that the discriminative palm shape contour information is primarily captured in the shallow layers of the model, while the deeper layers tend to focus on irrelevant local high-level semantics. Based on this, we design hierarchical feature fusion (HFF), a module that enables inter-modal channel exchange at shallow layers. Further, we introduce a multimodal contrastive learning loss to align features across modalities, enhancing their representational embeddings. Extensive experiments across eight widely-used public datasets demonstrate that MCCENet achieves state-of-the-art performance in all cases. Junqin Huang, Dacan Luo, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | Enhancing Perceptron Constancy for Real-World Dynamic Hand Gesture AuthenticationabstractDynamic hand gesture authentication (DHGA) has emerged as a promising biometric technology, offering enhanced theoretical security over conventional unimodal systems by combining both physiological and behavioral characteristics. Existing DHGA research predominantly focuses on controlled lab conditions, therefore showing low generalizability to uncontrolled application conditions. To bridge this gap, we propose a novel Skeleton-assistant Standardization and Authentication Framework (SSAF) that incorporates a generic data preprocessing method before authentication. First, we introduce a Geometry- Environment Standardization (GE-Stan) method to standardize five primary geometric and environmental factors inducing data distribution discrepancy, significantly improving robustness across different sessions and scenarios. Notably, the GE-Stan method can be applied to most existing algorithms and brings substantial improvement. Second, we design an Appearance and Motion Network (AM-Net) to fully leverage standardized video and skeleton data. It decouples appearance and motion features using specialized representation and processing strategies. Therefore, our SSAF achieves a flexible balance between accuracy and efficiency, enabling up to 3.6× efficiency boost with only minor accuracy trade-offs. Finally, to support real-world evaluation, we also contribute a new challenging dataset, SCUT-RealDHGA, captured under uncontrolled practical conditions with diverse backgrounds and illuminations. Extensive experiments across three DHGA datasets demonstrate that SSAF outperforms existing methods in terms of accuracy, efficiency, and robustness. The code and dataset are available at https://github.com/SCUTBIP-Lab/SSAF. Xilai Wang, Wenwei Song, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | Enhanced Geometry and Semantics for Camera-Based 3D Semantic Scene CompletionabstractGiving machines the ability to infer the complete 3D geometry and semantics of complex scenes is crucial for many downstream tasks, such as decision-making and planning. Vision-centric Semantic Scene Completion (SSC) has emerged as a trendy 3D perception paradigm due to its compatibility with task properties, low cost, and rich visual cues. Despite impressive results, current approaches inevitably suffer from problems such as depth errors or depth ambiguities during the 2D-to-3D transformation process. To overcome these limitations, in this paper, we first introduce an Optical Flow-Guided (OFG) DepthNet that leverages the strengths of pretrained depth estimation models, while incorporating optical flow images to improve depth prediction accuracy in regions with significant depth changes. Then, we propose a depth ambiguity-mitigated feature lifting strategy that implements deformable cross-attention in 3D pixel space to avoid depth ambiguities caused by the projection process from 3D to 2D and further enhances the effectiveness of feature updating through the utilization of prior mask indices. Moreover, we customize two subnetworks: a residual voxel network and a sparse UNet, to enhance the network's geometric prediction capabilities and ensure consistent semantic reasoning across varying scales. By doing so, our method achieves performance improvements over state-of-the-art methods on the SemanticKITTI, SSCBench-KITTI-360 and Occ3D-nuScene benchmarks. Haihong Xiao, Wenxiong Kang, Yulan Guo, Hao Liu 0061, Ying He 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Geometry-Aware 3D Gaussian Representation for Real-Time Rendering of Large-Scale ScenesabstractExisting NeRF-based methods for reconstructing large-scale scenes face challenges in visual quality and rendering speed due to spectral biases and extensive sampling requirements. Recent 3DGS-based methods for real-time rendering of 3D objects and small scenes outperform NeRF, but several issues persist when extending these techniques to large-scale scenes. These include robust rendering in weak-texture areas, effective densification under memory constraints, finer detail rendering, and achieving natural lighting transitions. To address these, we introduce a geometry-aware 3DGS method for efficient real-time rendering of large scenes. First, we propose a geometry-guided anchor point initialization that reduces noise away from structural surfaces and generates new points in weak texture areas, particularly for large-scale datasets. We also present a structure-aware joint densification strategy combining surface- and curvature-based densification, ensuring Gaussian points are near structural surfaces and increasing density in low-curvature areas. Additionally, we propose a hash grid-assisted, viewpoint-sensitive feature enhancement scheme to improve detail rendering and natural lighting transitions. Our method achieves superior rendering quality compared to state-of-the-art methods while maintaining reasonable memory usage. Extensive experiments across 16 scenes, including 11 from five public datasets (MatrixCity-Aerial, Mill-19, Tanks & Temples, WHU, and UrbanScene3D) and five self-collected scenes from SCUT-CA and plateau regions, demonstrate its generalization capability. Codes are available athttps://github.com/SCUT-BIP-Lab/Geo_gs. Haihong Xiao, Jianan Zou, Shuai Xing, Wenxiong Kang |
IEEE Trans. Multim. | 5 |
| 2025 | PhysMamba: Synergistic State Space Duality Model for Remote Physiological Measurement
Zhixin Yan, Shangru Yi, Wenxiong Kang |
ICANN (4) | 6 |
| 2025 | OnePV: A Novel One-Stage Palm Vein Recognition Method Based on Oriented Object DetectionabstractIn the conventional two-stage palm vein (PV) recognition the feature extraction is followed by region of interest (ROI) extraction, thus the recognition performance heavily depends on the robustness of the ROI extraction method, and the feature extractor lacks angular modeling capabilities. To address this, the relationship between oriented object detection and palm vein recognition is extensively studied in this paper. We propose a rotation-aware, one-stage PV recognition algorithm, where the rotated bounding box (RBB) of ROI is introduced as supervision information to guide the feature extractor and we annotate the ROI RBB of five public PV datasets. To the best of our knowledge, this is the first work that applies oriented object detection to one-stage PV recognition. Extensive experiments demonstrate that our method can achieve one-stage, high-accuracy PV recognition system and get remarkable results on five annotated datasets. Haoheng Lin, Runzhang Chen, Dacan Luo, Wenxiong Kang |
IJCB | 4 |
| 2025 | M3DHMR: Monocular 3D Hand Mesh RecoveryabstractMonocular 3D hand mesh recovery is challenging due to high degrees of freedom of hands, 2D-to-3D ambiguity and self-occlusion. Most existing methods are either inefficient or less straightforward for predicting the position of 3D mesh vertices. Thus, we propose a new pipeline called Monocular 3D Hand Mesh Recovery (M3DHMR) to directly estimate the positions of hand mesh vertices. M3DHMR provides 2D cues for 3D tasks from a single image and uses a new spiral decoder consist of several Dynamic Spiral Convolution (DSC) Layers and a Region of Interest (ROI) Layer. On the one hand, DSC Layers adaptively adjust the weights based on the vertex positions and extract the vertex features in both spatial and channel dimensions. On the other hand, ROI Layer utilizes the physical information and refines mesh vertices in each predefined hand region separately. Extensive experiments on popular dataset FreiHAND demonstrate that M3DHMR significantly outperforms state-of-the-art real-time methods. The code is available at https://github.com/Jackson-coder/M3DHMR. Yihong Lin, Xianjia Wu, Xilai Wang, Jianqiao Hu, Songju Lei, Xiandong Li, Wenxiong Kang |
IJCB | 7 |
| 2025 | Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025abstractHuman identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which includes significant variations in clothing, carried objects, and view angles. No training data is provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, reducing the risk of overfitting and ensuring fair evaluation of cross-domain generalization. Although the previous two competitions (HID 2023 and HID 2024) already utilized this dataset, HID 2025 aimed explicitly to explore whether algorithmic improvements could surpass the accuracy limits observed previously. Despite these heightened challenges, participants again demonstrated significant advancements, with the highest accuracy reaching 94.2%, setting a new benchmark for this dataset. We also analyze key technical trends and outline potential directions for future research on gait recognition. Jingzhe Ma, Jianlong Yu, Zunxiao Xu, Xue Cheng, Zepeng Wang 0002, Kazuki Osamura, Rujie Liu, Narishige Abe, Shunli Zhang 0005, Haojun Xie, Weiming Wu, Wenxiong Kang, Qingshuo Gao, Jiaming Xiong, Xianye Ben, Lei Chen 0095, Lichen Song, Junjian Cui, Haijun Xiong, Junhao Lu, Bin Feng 0001, Baoquan Zhao, Ke Xu 0001, Yongzhen Huang, Liang Wang 0001, Manuel J. Marín-Jiménez, Md. Atiqur Rahman Ahad, Shiqi Yu 0001 |
IJCB | 19 |
| 2025 | SCDFormer: Spatial and Channel Denoising Transformer for Human Pose Estimation Using Millimeter-Wave RadarabstractThe millimeter-wave radar-based human pose estimation technology has attracted significant attention due to its cost-effectiveness and non-intrusive nature. However, the existing methods focus on modeling spatial dependence but overlook channel connections, failing to highlight important channels. In addition, the millimeter-wave point cloud is noisy, which implies some point-pair connections are irrelevant and channel connections are noisy. To address the above issues, we propose two basic units named Spatial Denoising Self-Attention Layer (SDAL) and Channel Denoising Self-Attention Layer (CDAL). SDAL filters out the irrelevant point-pair connections while modeling spatial dependence. CDAL emphasizes the important channels and filters out the noisy channel connections. Based on the SDAL and CDAL, we propose a new network named Spatial and Channel Denoising Transformer (SCDFormer). Our SCDFormer can not only filter out the irrelevant point-pair connections, but also emphasize important channels and filter out the noisy channel connections. Experiments demonstrate that our SCDFormer achieves state-of-the-art on the Mars, mRI and MiliPoint datasets. Qiuxia Wu, Panpan Cai, Wenxiong Kang |
IJCB | 4 |
| 2025 | MoTeNet: Motion-Temporal Network for Dynamic Hand Gesture Recognition on Point CloudsabstractAs deep learning techniques are increasingly applied to gesture recognition, point cloud-based dynamic gesture recognition methods have attracted significant attention. However, many existing approaches overlook per-point motion features and the overall temporal dynamics of gestures, thereby failing to fully exploit motion and temporal cues embedded in point cloud sequences. To address this issue, we propose Motion-Temporal Network (MoTeNet), a novel framework for dynamic gesture recognition on point clouds, which preserves spatial structural information while modeling both global frame-level temporal changes and per-point motion. MoTeNet extracts spatial geometric features through a Hierarchical Graph Convolution (HGC) module, and incorporates a Motion Feature Encoding Module (MFEM) to encode point-level motion features across adjacent frames. This approach provides fine-grained dynamic information for subsequent temporal modeling. Furthermore, an Adaptive Temporal Feature Fusion (ATFF) module integrates convolutional neural networks (CNNs) and transformers to adaptively fuse short-term and long-term temporal dependencies, enabling comprehensive modeling of the dynamic evolution of gestures. Experimental results demonstrate that MoTeNet achieves state-of-the-art performance on point cloud gesture recognition benchmarks, including SHREC’17, DHG, and NVGesture, with ablation studies further validating the effectiveness of the proposed framework. Qiuxia Wu, Xinran Xie, Sangni Xu, Wenxiong Kang |
IJCB | 4 |
| 2025 | Study of Finger Biometrics on Finger Semantic Segmentation and Finger Shape AuthenticationabstractIn hand-based biometrics, fingerprint, finger vein, finger knuckle print, palm print, palm vein, dorsal hand vein, and hand shape are the traits that are getting much attention. However, finger shape (FS), a forgettable trait, has not been studied specifically for identification purposes. In this work, we explore this content as a complement to the hand-based biometrics. Firstly, we annotate the FS on a publicly available finger vein dataset as the ground truth for finger semantic segmentation. Then we explore the finger semantic segmentation task on the annotated data and propose a lightweight network, namely FinSeg-Net (finger segmentation network). Finally, we conduct the FS authentication experiment based on four matching methods; experimental results show that the FS traits can achieve identity authentication. This work is the first study for FS biometrics specifically, and built the first FS dataset, which will be accessed via: https://github.com/SCUT-BIP-Lab/FinSeg. Junduan Huang, Dacan Luo, Weili Yang, Jiahui Pan 0003, Wenxiong Kang |
ICME | 5 |
| 2025 | GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion TransformerabstractSpeech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have achieved notable progress, they often suffer from modality inconsistencies, particularly misalignment between audio and mesh, leading to reduced motion diversity and lip-sync accuracy. To address this, we propose GLDiTalker, a novel speech-driven 3D facial animation model based on a Graph Latent Diffusion Transformer. GLDiTalker resolves modality misalignment by diffusing signals within a quantized spatiotemporal latent space. It employs a two-stage training pipeline: the Graph-Enhanced Quantized Space Learning Stage ensures lip-sync accuracy, while the Space-Time Powered Latent Diffusion Stage enhances motion diversity. Together, these stages enable GLDiTalker to generate realistic, temporally stable 3D facial animations. Extensive evaluations on standard benchmarks demonstrate that GLDiTalker outperforms existing methods, achieving superior results in both lip-sync accuracy and motion diversity. Yihong Lin, Zhaoxin Fan, Xianjia Wu, Lingyu Xiong, Xiandong Li, Wenxiong Kang, Songju Lei |
IJCAI | 6 |
| 2025 | Diffusion-Guided Graph Data AugmentationabstractGraph Neural Networks (GNNs) have achieved remarkable success in a wide range of applications. However, when trained on limited or low-diversity datasets, GNNs are prone to overfitting and memorization, which impacts their generalization. To address this, graph data augmentation (GDA) has become a crucial task to enhance the performance and generalization of GNNs.
Traditional GDA methods employ simple transformations that result in limited performance gains. Although recent diffusion-based augmentation methods offer improved results, they are sparse, task-specific, and constrained by class labels. In this work, we propose a more general and effective diffusion-based GDA framework that is task-agnostic and label-free.
For better training stability and reduced computational cost, we employ a graph variational auto-encoder (GVAE) to learn a compact latent graph representation. A diffusion model is used in the learned latent space to generate both consistent and diverse augmentations.
For a fixed augmentation budget, our algorithm selects a subset of samples that would benefit the most from the augmentation.
To further improve performance, we also perform test-time augmentation, leveraged by the label-free nature of our method.
Thanks to the efficient utilization of GVAE and latent diffusion, our algorithm significantly enhances machine learning safety measures, including calibration, robustness to corruptions, and prediction consistency. Moreover, our method has shown improved robustness against four types of adversarial attacks and achieves better generalization performance.
To demonstrate the effectiveness of the proposed method, we compare it with 30 existing methods on 12 benchmark datasets across node classification, link prediction, and graph classification in various learning settings, including semi-supervised, supervised, and long-tailed data distributions.
The code will soon be made publicly available. Maria Marrium, Arif Mahmood, Muhammad Haris Khan, M. Saad Shakeel, Wenxiong Kang |
NeurIPS | 5 |
| 2025 | Improving 3D Finger Traits Recognition via Generalizable Neural Rendering
Junduan Huang, Yuer Ma, Wenxiong Kang |
Int. J. Comput. Vis. | 5 |
| 2025 | Normalized-Full-Palmar-Hand: Toward More Accurate Hand-Based Multimodal BiometricsabstractHand-based multimodal biometrics have attracted significant attention due to their high security and performance. However, existing methods fail to adequately decouple various hand biometric traits, limiting the extraction of unique features. Moreover, effective feature extraction for multiple hand traits remains a challenge. To address these issues, we propose a novel method for the precise decoupling of hand multimodal features called 'Normalized-Full-Palmar-Hand' and construct an authentication system based on this method. First, we propose HSANet, which accurately segments various hand regions with diverse backgrounds based on low-level details and high-level semantic information. Next, we establish two hand multimodal biometric databases with HSANet: SCUT Normalized-Full-Palmar-Hand Database Version 1 (SCUT_NFPH_v1) and Version 2 (SCUT_NFPH_v2). These databases include full hand images, semantic masks, and images of various hand biometric traits obtained from the same individual at the same scale, totaling 157,500 images. Third, we propose the Full Palmar Hand Authentication Network framework (FPHandNet) to extract unique features of multiple hand biometric traits. Finally, extensive experimental results, performed via the publicly available CASIA, IITD, COEP databases, and our proposed databases, validate the effectiveness of our methods. Yitao Qiao, Wenxiong Kang, Dacan Luo, Junduan Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Mirror-Based Full-View Finger Vein Authentication With Illumination AdaptationabstractFull-view finger vein (FV) biometrics systems capture multiple FV images of the presented finger ensuring that the entire surface of the finger is covered. Existing full-view FV systems suffer from three common problems: large device size, high cost for multi-camera system, and sub-optimal illumination in the recorded FV images. To address the problem of device size, we propose a novel Mirror-based Full-view FV (MFFV) capture device. The MFFV device has a compact size by using mirror-reflection approach. We reduce the cost of the device by using low-cost components, in particular, consumer-grade cameras. To address the problems of lower-quality images captured by such cameras and obtain optimally illuminated FV images, we propose a two-step approach. The first step is a Multi-illumination Intensities FV (MIFV) capture strategy, which capture the FV image set with varying illumination intensities. In the second step, a FV illumination adaptation (FVIA) algorithm is proposed to select the optimally illuminated FV image from the MIFV image set. Using the proposed MFFV device, we collect a comprehensive dataset, namely MFFV dataset, along with reproducible baseline FV authentication results for both single-view and full-view FV. Our experimental results demonstrate that the MIFV capture strategy as well as the FVIA algorithm can effectively improve the authentication performance, and that the full-view FV authentication is significantly superior than the single-view FV authentication. The source-code and dataset for reproducing our experimental results are publicly available. The code and the license for MFFV-N dataset can be accessed at:https://github.com/SCUT-BIP-Lab/MFFV. Junduan Huang, Sushil Bhattacharjee, Sébastien Marcel, Wenxiong Kang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Multiscale Super-Images for Dynamic Hand Gesture AuthenticationabstractThe dynamic hand gesture is an emerging biometric trait that has attracted the attention of researchers due to its rich physiological and behavioral characteristics. The previous studies primarily focused on extracting and utilizing the physiological characteristics, while ignoring the rich behavioral characteristics contained in hand gesture movements. The dynamic hand gesture authentication performance will be improved if behavioral characteristics can be effectively extracted and fused with physiological characteristics for authentication. In addition, existing methods still suffer from insufficient feature extraction capabilities and low efficiency in extracting behavioral characteristics from complex dynamic hand gestures. To address these issues, this paper first proposes multiscale dynamic hand gesture (MDHG) super-images to represent the behavioral characteristics of hand gestures, containing sufficient local and global motion cues. Furthermore, for the super-images, this paper proposes a two-stream network consisting of a spatiotemporal feature extraction backbone and an identity-aggregation module to fully extract and fuse the physiological and behavioral characteristics of hand gestures, which significantly improves the accuracy of dynamic hand gesture authentication. Extensive experiments on two benchmark datasets, SCUT-DHGA and HandLogin, show that our method achieves superior performance with fewer parameters and FLOPs than other networks, validating the effectiveness, generalizability, and security of our proposed method. Zenan Lin, Wenwei Song, Wenxiong Kang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Semantic Scene Completion via Semantic-Aware Guidance and Interactive Refinement TransformerabstractPredicting per-voxel occupancy status and corresponding semantic labels in 3D scenes is pivotal to 3D intelligent perception in autonomous driving. In this paper, we propose a novel semantic scene completion framework that can generate complete 3D volumetric semantics from a single image at a low cost. To the best of our knowledge, this is the first endeavor specifically aimed at mitigating the negative impacts of incorrect voxel query proposals caused by erroneous depth estimates and enhancing interactions for positive ones in camera-based semantic scene completion tasks. Specifically, we present a straightforward yet effective Semantic-aware Guided (SAG) module, which seamlessly integrates with task-related semantic priors to facilitate effective interactions between image features and voxel query proposals in a plug-and-play manner. Furthermore, we introduce a set of learnable object queries to better perceive objects within the scene. Building on this, we propose an Interactive Refinement Transformer (IRT) block, which iteratively updates voxel query proposals to enhance the perception of semantics and objects within the scene by leveraging the interaction between object queries and voxel queries through query-to-query cross-attention. Extensive experiments demonstrate that our method outperforms existing state-of-the-art approaches, achieving overall improvements of 0.30 and 2.74 in mIoU metric on the SemanticKITTI and SSCBench-KITTI-360 validation datasets, respectively, while also showing superior performance in the aspect of small object generation. Haihong Xiao, Wenxiong Kang, Hao Liu 0061, Yuqiong Li, Ying He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Study of Full-View Finger Vein Biometrics on Redundancy Analysis and Dynamic Feature ExtractionabstractAs a biometric trait drawing increasing attention, finger vein (FV) has been studied from many perspectives. One promising new direction in FV biometrics research is full-view FV biometrics, where multiple images, covering the entire surface of the presented finger, are captured. Full-view FV biometrics presents two main problems: increased computational load, and low performance-to-cost ratio for some views/regions. Both problems are related to the inherent redundancy in vascular information available in full-view FV images. In this work, we address this redundancy issue in full-view FV biometrics. Firstly, we propose a straightforward FV redundancy analysis (FVRA) method for quantifying the information redundancy in FV images. Our analysis shows that the redundancy ratio of full-view FV images is up to 83%-87%. Then, we propose a novel feature extraction model, named FV dynamic Transformer (FDT), whose architecture is configured based on the redundancy analysis results. The FDT focuses on both local (single-view) information as well as global (full view) information at different processing stages. Both stages provide the advantage of de-redundancy and noise avoidance. Additionally, the end-to-end architecture simplifies the full-view FV biometrics pipeline by enabling the direct, simultaneous processing of multiple input images, thus consolidating multiple steps into one. A series of rigorous experiments is conducted to evaluate the effectiveness of the proposed methods. Experimental results show that the proposed FDT achieves state of the art authentication performance on the MFFV-N dataset, yielding an EER of 0.97% on the development set and an HTER of 1.84% on the test set under the balanced protocol and EER criterion. The cross-domain generalization capability of FDT is also demonstrated on the LFMB-3DFB dataset, where it achieves an EER of 7.24% and an HTER of 7.34% under the same protocol and criterion. Code for the proposed methods can be access via: https://github.com/SCUT-BIP-Lab/FDT. Junduan Huang, Sushil Bhattacharjee, Sébastien Marcel, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | RSNet: Region-Specific Network for Contactless Palm Vein AuthenticationabstractMore palm features, such as veins and shapes obtained from an enlarged contactless palm vein region of interest (ROI), have been shown to improve recognition performance. However, a few efforts have been made to adequately utilize these features for mining identity information. To address this issue, we propose a Region-Specific Network (RSNet) for contactless palm vein authentication. Our RSNet is a dual-branch structure for global and local feature extraction. Firstly, a Region-based Local feature Enhancement Block (RLEB) is proposed at the local branch to extract region-specific features. In the RLEB, the intermediate feature maps are divided into three asymmetrical patches based on the physiological characteristics of palm vein and palm shape for extracting diversified features, enhancing the local feature representation. Then, a Multi-scale Aggregation Block (MAB) is proposed that efficiently aggregates multi-scale features at a more granular level. Furthermore, to guide the global and local branches in learning complementary feature aspects, a difference loss is introduced to apply a soft subspace orthogonality constraint between the global and local vectors during training. The global branch is designed to assist the learning process of local features, without being adopted for inference. Extensive experiments have demonstrated the effectiveness and superiority of our method, and the RSNet achieves new State-Of-The-Art (SOTA) authentication performance on seven public contactless palm vein databases in the open-set scenario. Dacan Luo, Junduan Huang, Weili Yang, M. Saad Shakeel, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | A Random-Binding-Based Bio-Hashing Template Protection Method for Palm Vein RecognitionabstractTo mitigate the risk of data breaches, an increasing number of biometric recognition systems are introducing encryption biometric template protection methods and directly matching in the encrypted domain. Depending on the approach to key management, prevailing biometric template protection strategies can be categorized into declarative and distributive methods. The former are challenged by complexities and vulnerabilities linked to key loss, while the latter are compromised by fixed mapping rules that may expose personal information. We present a biometric template protection method that combines random-fixed factors to handle these challenges, thereby protecting the user’s biometric privacy. Firstly, we introduce a random activation factor generation module that extracts scaling and offset factors from the user’s biometric data. This module randomly binds factors to different positions in each authentication process, rendering distance-dependent bitwise cracking algorithms ineffective. Secondly, we propose a fixed multi-branch mapping module that enhances feature expression and minimizes information loss post-encryption. We also develop a trainable min-max hash method, optimized using an improved approximate contrastive loss. Employing palm veins as a case study, we conducted experiments across five datasets, where our method outperformed other encrypted domain methods and showed competitive advantages over mainstream non-encrypted methods. Moreover, we have demonstrated that our method ensures robust performance while meeting essential security requirements of irreversibility, unlinkability, and revocability. Tianming Xie, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | SyncLearnNet: Generalized Epileptic Seizure Detection Network Based on Brain SignalsabstractEpilepsy is a prevalent neurological disorder with significant detrimental effects on health. Accurate seizure detection is crucial for the precise diagnosis and effective treatment of epilepsy. Brain signals is widely recognized as a reliable clinical tool for diagnosing and evaluating severity of seizures. Traditionally, medical researchers have relied on visual inspection to identify and locate seizures and epileptogenic areas. However, manual analysis of brain data is both subjective and time-consuming. In recent years, there has been a surge in studies focusing on automatic seizure detection algorithms based on brain signals, driven by the advancements in artificial intelligence and digital brain signal technology. Nevertheless, in tackling this task, many of these studies have neglected to leverage the rich implicit information of samples to extract comprehensive feature representation for enhancing model performance. To address this gap, we propose a generalized model called SyncLearnNet for seizure detection based on brain signals. SyncLearnNet incorporates VariaScan and BatchAttention modules designed to fully utilize both intra-sample and inter-sample information, thereby improving feature discrimination without requiring additional data. Furthermore, the introduction of CurriClassifier aims to enhance the model's generalization performance. Experiments conducted on a public human seizure dataset CHB-MIT and a self-built animal seizure dataset comprising data from five rats demonstrated this method outperforms existing seizure detection methods in terms of generalization performance. Yuer Ma, Jiaoyang Wang, Wenxiong Kang |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Zero-Shot Text-Driven Dynamic Neural Radiance Fields StylizationabstractText-driven style transfer for Neural Radiance Fields (NeRFs) is an emerging research topic that leverages text descriptions instead of reference style images to apply style transfer. However, existing methods for stylizing NeRFs predominantly struggle to extend to 4D dynamic scenes, due to NeRFs' inherent limitation to static environments. Moreover, these current methods require training for each specific text input, which limits them to a single style description and significantly hampers generalizability and applications. In this paper, we introduce a novel approach to zero-shot text-driven 4D style transfer that adopts text inputs into the CLIP's style space with a canonical feature volume. Specifically, using geometric priors from pre-trained dynamic Neural Radiance Fields, we train a canonical feature volume by rendering feature maps under the supervision of a pre-trained VGG encoder. Then we utilize CLIP's multi-modal embedding to connect the text descriptions with style images and learn a canonical style transformation matrix in CLIP's feature space. Experiments show that our method achieves zero-shot text-driven style transfer for dynamic neural radiance fields and maintains good multi-view and cross-time consistency. Wanlin Liang, Wanshui Gan, Wenxiong Kang |
IEEE Trans. Multim. | 4 |
| 2025 | $\mathrm{Tri^{2}plane}$: Advancing Neural Implicit Surface Reconstruction for Indoor ScenesabstractReconstructing 3D indoor scenes presents significant challenges, requiring models capable of inferring both planar surfaces and intricate details. Although recent methods can generate complete surfaces, they often struggle to simultaneously reconstruct low-texture regions and high-frequency details due to non-local effects. In this paper, we introduce a novel triangle-based triplane representation, named (tri$^{2}$plane), specifically designed to account for the diverse spatial feature distribution and information density of indoor environments. Our method begins by projecting point clouds onto three orthogonal planes, followed by 2D Delaunay triangulation. This representation enables adaptive encoding of low-texture and high-frequency regions by employing triangles of variable sizes. Moreover, we develop a dual tri$^{2}$plane framework that incorporates both geometric and semantic information, significantly enhancing the reconstruction quality. We combine these key modules and evaluate our method on benchmark indoor scene datasets. The results unequivocally demonstrate the superiority of our proposed method over the state-of-the-art Occ-SDF. Specifically, our method achieves significant improvements over Occ-SDF, with margins of 1.3, 1.7, and 2.3in F-score on the ScanNet, Tanks & Temples, and Replica datasets, respectively. To facilitate further research, we will make our code publicly available. Haihong Xiao, Wenxiong Kang |
IEEE Trans. Multim. | 3 |
| 2024 | FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction TextabstractRearrangement operations form the crux of interactions between humans and their environment. The ability to generate natural, fluid sequences of this operation is of essential value in AR/VR and CG. Bridging a gap in the field, our study introduces FAVOR: a novel dataset for Full-body AR-driven Virtual Object Rearrangement that uniquely employs motion capture systems and AR eyeglasses. Comprising 3k diverse motion rearrangement sequences and 7.17 million interaction data frames, this dataset breaks new ground in research data. We also present a pipeline FAVORITE for producing digital human rearrangement motion sequences guided by instructions. Experimental results, both qualitative and quantitative, suggest that this dataset and pipeline deliver high-quality motion sequences. Our dataset, code, and appendix are available at https://kailinli.github.io/FAVOR. Kailin Li 0001, Lixin Yang 0001, Zenan Lin, Jian Xu 0027, Xinyu Zhan 0001, Yifei Zhao 0003, Pengxiang Zhu, Wenxiong Kang, Kejian Wu, Cewu Lu |
AAAI | 8 |
| 2024 | Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial AnimationabstractSpeech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip synchronization while neglecting to model the subject-specific speaking style, often resulting in unrealistic facial animations. To the best of our knowledge, this work makes the first attempt to explore the coupled information between the speaking style and the semantic content in facial motions. Specifically, we introduce an innovative speaking style disentanglement method, which enables arbitrary-subject speaking style encoding and leads to a more realistic synthesis of speech-driven facial animations. Subsequently, we propose a novel framework called Mimic to learn disentangled representations of the speaking style and content from facial motions by building two latent spaces for style and content, respectively. Moreover, to facilitate disentangled representation learning, we introduce four well-designed constraints: an auxiliary style classifier, an auxiliary inverse classifier, a content contrastive loss, and a pair of latent cycle losses, which can effectively contribute to the construction of the identity-related style space and semantic-related content space. Extensive qualitative and quantitative experiments conducted on three publicly available datasets demonstrate that our approach outperforms state-of-the-art methods and is capable of capturing diverse speaking styles for speech-driven 3D facial animation. The source code and supplementary video are publicly available at: https://zeqing-wang.github.io/Mimic/ Zeqing Wang, Keze Wang, Tianshui Chen, Haifeng Zeng, Wenxiong Kang |
AAAI | 8 |
| 2024 | MRGait: A Multi-range feature learning framework for Cross-View Gait Recognition
M. Saad Shakeel, Kun Liu 0029, Xiaochuan Liao, Wenxiong Kang |
MMAsia | 4 |
| 2024 | T2QRM: Text-Driven Quadruped Robot Motion Generation
Kun Hu 0008, Zhiyong Wang 0001, Wenxiong Kang |
MMAsia | 6 |
| 2024 | L3AM: Linear Adaptive Additive Angular Margin Loss for Video-Based Hand Gesture Authentication
Wenwei Song, Wenxiong Kang, Adams Wai-Kin Kong, Yitao Qiao |
Int. J. Comput. Vis. | 2 |
| 2024 | Efficient disentangled representation learning for multi-modal finger biometrics
Weili Yang, Junduan Huang, Dacan Luo, Wenxiong Kang |
Pattern Recognit. | 4 |
| 2024 | EA-MVSNet: Learning Error-Awareness for Enhanced Multi-View StereoabstractMulti-view stereo (MVS) aims to reconstruct the dense 3D geometry of a scene by processing and relating images captured from different viewpoints. Despite impressive successes, most existing techniques simply supervise cost volumes or depth maps through conventional classification or regression methods, thereby inadequately exploring the depth representation’s full potential. Moreover, reconstructing areas with occlusions or weak textures continues to be a long-standing challenge within MVS. Another critical issue, frequently neglected, is the potential inaccuracy of ground truth depths, as evidenced in datasets like DTU. To address these problems, we introduce EA-MVSNet, an innovative error-aware MVS framework designed to enhance depth prediction. The key contributions of this work include three parts: (1) We present a novel error-aware depth representation that enhances depth prediction accuracy through error-aware learning, thereby improving reconstruction quality. (2) We develop a Deformable Feature Pyramid Network (DFPN), meticulously designed to augment reconstruction details in occluded and texture-deficient areas. (3) We introduce a cross-view consistency guidance module into the learning process, effectively mitigating the detrimental effects of ground truth depth inaccuracies and fostering faster convergence. Comprehensive experiments on the DTU dataset and Tanks and Temples dataset validate the superiority of our EA-MVSNet. Compared to the preceding UniMVSNet, EA-MVSNet achieves a notable 7.6% decrease in overall reconstruction error on the DTU dataset, and boosts the mean F-score by 3.0% and 4.1% in the intermediate and advanced groups of the Tanks and Temples dataset, respectively, surpassing most recent state-of-the-art methods. Wencong Gu, Haihong Xiao, Xueyan Zhao, Wenxiong Kang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Hand Gesture Authentication by Discovering Fine-Grained Spatiotemporal Identity CharacteristicsabstractDynamic hand gesture is an emerging and promising biometric trait containing both physiological and behavioral characteristics. Possessing the two kinds of characteristics makes dynamic hand gesture have more identity information enabling more accurate and secure authentication theoretically, but also poses a challenge of efficient fine-grained spatiotemporal feature extraction. This challenge involves a seemingly paradoxical problem that high-frame-rate videos are required for behavioral characteristic analysis, but they can also introduce high computational costs. To mitigate this issue, we propose a Frequency Spatiotemporal Attention Network (FSTA-Net) with a focus on satisfying the high-performance and low-computation requirements of authentication systems. The FSTA-Net is established with a two-stage identity characteristic analysis paradigm for short- and long-term modeling. Specifically, considering that models prefer to analyze physiological characteristics which are relatively straightforward to understand, we first design a Behavior Enhanced (BE) module to emphasize hand motions and reduce redundant information to facilitate local identity feature distillation in the first stage. We then present a Frequency Spatiotemporal Attention (FSTA) module to summarize global identity features with decent FLOPs and GPU memory occupation in the second stage. Incorporating the BE and FSTA modules enables them to complement each other’s strengths, resulting in a clear-cut improvement in equal error rate and running speed. Extensive experiments on the SCUT-DHGA dataset demonstrate the superiority of the FSTA-Net. The code is available athttps://github.com/SCUT-BIP-Lab/FSTA-Net. Wenwei Song, Wenxiong Kang, Liang Lin 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Adaptive Positive Sample Selection and Dynamic Soft Label Assignment for Keypoint DetectionabstractPose estimation plays a crucial role in human-centered vision applications. Some recent efforts achieved pose estimation by keypoints detection. Drawing inspiration from object detection, they treated keypoints as objects and achieved unbiased estimation through implementation of classification and regression heads. However, they still failed to achieve satisfactory performance for detecting heavily occluded keypoints and required elaborate and unavoidable post-processing steps. With a thorough exploration of keypoints’ characteristics, we have developed a novel Adaptive positive Sample selection and dynamic soft Label Assignment (ASLA) scheme tailored for keypoint detection. Specifically, we select positive samples for each keypoint according to the summation distance from the sample coordinates and their predicted coordinates to their corresponding ground truth (GT) in the training phase. For occluded keypoints, the positive samples defined by our method may fall in the semantically relevant regions of pedestrians, rather than the spatially adjacent regions of obstructions, significantly improving their localization performance. Meanwhile, we dynamically assign classification labels to these positive samples based on the distance between their predicted coordinates and their corresponding GT, which ensures that high quality positive samples are assigned with high classification labels. Benefiting from the practical design of our ASLA, the post-processing step is not essential; however, the simple vector-level post-processing would be the icing on the cake. Finally, we extensively evaluate our ASLA performance on two popular human pose estimation benchmarks, COCO and MPII, and comprehensive experiments show that our ASLA significantly outperforms state-of-the-art algorithms. Our code and models will be available athttps://github.com/SCUT-BIP-Lab/ASLA. Wenxiao Tang, M. Saad Shakeel, Wenxiong Kang, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Point Cloud Completion via Self-Projected View Augmentation and Implicit Field ConstraintabstractRecent advances in point cloud completion make it possible to simultaneously recover complete shapes and fine details from partial point clouds captured by professional 3D devices, such as Lidar, or consumer cameras, such as iPhones. Despite significant progress, the potential utilization of self-projected views from partial inputs and the effective reduction of noise in generated point clouds remain under-explored. In this paper, we propose a novel point cloud completion method that leverages self-projected view augmentation and implicit field constraints. Specifically, we introduce a cross-view augmentation (CVA) module and a cross-modal fusion (CMF) module to enhance information interaction and integration at the image and modality levels, respectively. We also propose a bidirection-aware refinement block to improve detail and completeness by considering both complete-to-partial detail perception and partial-to-complete structure perception paths. Additionally, we address the issue of noise reduction from the perspective of implicit field constraints. We evaluate our method on several baseline datasets, including PCN, ShapeNet55/34 and KITTI (car). Extensive experiments demonstrate that our method outperforms state-of-the-art methods, achieving improvements of 0.11 CD-$\ell _{1}$, 0.015 DCD and 0.009 F-score on the standard PCN test set. Furthermore, our approach effectively reduces noise in the generated point clouds, showcasing its promising potential for practical applications. Haihong Xiao, Ying He 0001, Hao Liu 0061, Wenxiong Kang, Yuqiong Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Learning an Augmented RGB Representation for Dynamic Hand Gesture AuthenticationabstractDynamic hand gesture authentication aims to recognize users’ identity through the characteristics of their hand gestures. How to extract favorable features for verification is the key to success. Cross-modal knowledge distillation is an intuitive approach that can introduce additional modality information in the training phase to enhance the target modality representation, improving model performance without incurring additional computation in the inference phase. However, most previous cross-modal knowledge distillation methods directly transfer information from one modality to another one without considering the modality gap. In this paper, we propose a novel translation mechanism in cross-modal knowledge distillation that can effectively mitigate the modality gap and utilize the information from the additional modality to enhance the target modality representation. In order to better transfer modality information, we propose a novel modality fusion-enhanced non-local (MFENL) module, which can fuse the multi-modal information from the teacher network and enhance the fused features based on the modality input into the student network. We use cascaded MFENL modules as the translator based on the proposed cross-modal knowledge distillation method to learn an enhanced RGB representation for dynamic hand gesture authentication. Extensive experiments on the SCUT-DHGA dataset demonstrate that our method has compelling advantages over the state-of-the-art methods. The code is available athttps://github.com/SCUT-BIP-Lab/TranslationCKD. Huilong Xie, Wenwei Song, Wenxiong Kang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | RobustMVS: Single Domain Generalized Deep Multi-View StereoabstractDespite the impressive performance of Multi-view Stereo (MVS) approaches given plenty of training samples, the performance degradation when generalizing to unseen domains has not been clearly explored yet. In this work, we focus on the domain generalization problem in MVS. To evaluate the generalization results, we build a novel MVS domain generalization benchmark including synthetic and real-world datasets. In contrast to conventional domain generalization benchmarks, we consider a more realistic but challenging scenario, where only one source domain is available for training. The MVS problem can be analogized back to the feature matching task, and maintaining robust feature consistency among views is an important factor for improving generalization performance. To address the domain generalization problem in MVS, we propose a novel MVS framework, namely RobustMVS1. A Depth-Clustering-guided Whitening (DCW) loss is further introduced to preserve the feature consistency among different views, which decorrelates multi-view features from viewpoint-specific style information based on geometric priors from depth maps. The experimental results further show that our method achieves superior performance on the domain generalization benchmark2. Baigui Sun, Xuansong Xie, Wenxiong Kang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Palm Vein Recognition Under Unconstrained and Weak-Cooperative ConditionsabstractContactless palm vein has attracted significant attention for its high security, stability, and user-friendliness. However, current contactless palm vein recognition predominantly relies on databases collected from platforms with spatial and temporal constrained design, which inadequately reflect relaxed palm vein imaging circumstances. This paper proposes a novel manner called on-the-fly palm vein that frees the user’s palm from spatial and temporal constraints, enabling palm vein recognition under unconstrained and weak-cooperative conditions. Firstly, Designing efficient and user-friendly palm vein imaging and authentication via two dynamic palm motions is proposed, resulting in an on-the-fly palm vein recognition platform. Next, a large-scale and challenging palm vein database, SCUT Palm Vein Database Version 1 (SCUT_PV_v1), is constructed. It is the first palm vein database with images collected under unconstrained and weak-cooperative conditions, encompassing a wider range of palm pose variations, grayscale variations, and lower-quality images. Finally, a lightweight and efficient Adaptive Margin Palm Vein Authentication Network (AMPVNet) is proposed as a baseline for the SCUT_PV_v1, where a vein pattern-specific convolutional neural network (CNN) is designed to extract features and a tailored online data augmentation method, combining Random Perspective Transformation (RPT) with Random Grayscale Adjustment (RGA), is proposed to enrich the diversify of out-of-plane palm pose and grayscale variations. Extensive experimental results demonstrate the effectiveness of our proposed methods. As the first work for palm vein recognition under unconstrained and weak-cooperation conditions, the AMPVNet achieves a promising accuracy and computation result while maintaining robustness to palm pose and grayscale variations. The SCUT_PV_ v1 database will be public at https://github.com/SCUT-BIP-Lab/SCUT_PV_v1. Dacan Luo, Yitao Qiao, Di Xie, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Robust and Accurate Hand Gesture Authentication With Cross-Modality Local-Global Behavior AnalysisabstractObtaining robust fine-grained behavioral features is critical for dynamic hand gesture authentication. However, behavioral characteristics are abstract and complex, making them more difficult to capture than physiological characteristics. Moreover, various illumination and backgrounds in practical applications pose additional challenges to existing methods because commonly used RGB videos are sensitive to them. To overcome this robustness limitation, we propose a two-stream CNN-based cross-modality local-global network (CMLG-Net) with two complementary modules to enhance the discriminability and robustness of behavioral features. First, we introduce a temporal scale pyramid (TSP) module consisting of multiple parallel convolution subbranches with different temporal kernel sizes to capture the fine-grained local motion cues at various temporal scales. Second, a cross-modality temporal non-local (CMTNL) module is devised to simultaneously aggregate the global temporal features and cross-modality features with an attention mechanism. Through the complementary combination of the TSP and CMTNL modules, our CMLG-Net obtains a comprehensive and robust behavioral representation that contains both multi-scale (short- and long-term) and multimodal (RGB-D) behavioral information. Extensive experiments are conducted on the largest dataset, SCUT-DHGA, and a simulated practical dataset, SCUT-DHGA-br, to demonstrate the effectiveness of CMLG-Net in exploiting fine-grained behavioral features and complementary multimodal information. Finally, it achieves stat-of-the-art performance with the lowest ERR of 0.497% and 4.848% in two challenging evaluation protocols and shows significant superiority in robustness under practical scenes with unsatisfactory illumination and backgrounds. The code is available athttps://github.com/SCUT-BIP-Lab/CMLG-Net. Wenxiong Kang, Wenwei Song |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Instance-Aware Monocular 3D Semantic Scene CompletionabstractWe study outdoor 3D scene understanding, a challenging task demanding the intelligent system to infer both geometry and semantics from a single-view image – a critical skill for autonomous vehicles to navigate in the real 3D world. Towards this end, we present an instance-aware monocular semantic scene completion framework. To the best of our knowledge, this is the first endeavor specifically targeting the challenge of instance perception in the camera-based semantic scene completion task. Our method consists of two stages. In stage I, we design a region-based VQ-VAE network, providing an effective solution for 3D occupancy prediction. In stage II, we first introduce an instance-aware attention module, explicitly incorporating instance-level cues captured from mask images to enhance the instance features in RGB images. Then we leverage the deformable cross-attention to aggregate image features corresponding to each voxel query and utilize the deformable self-attention to refine query proposals. We combine these key ingredients and evaluate our method on two challenging datasets, namely SemanticKITTI and SSCBench-KITTI-360. The results unequivocally demonstrate the superiority of our proposed method over the state-of-the-art VoxFormer-S. Specifically, our method surpasses VoxFormer-S by 0.22 IoU and 0.72 mIoU on the validation set and achieves an impressive improvement of 3.04 IoU and 1.06 mIoU on the SSCBench-KITTI-360 validation set. Meanwhile, our approach ensures accurate perception of critical instances, thereby exhibiting its exceptional performance and potential for practical deployment. Haihong Xiao, Wenxiong Kang, Yuqiong Li |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Dual Masked Modeling for Weakly-Supervised Temporal Boundary DiscoveryabstractDiscovering temporal boundary is critical for untrimmed video tasks, such as temporal sentence grounding and action detection. Due to the labor-intensive boundary annotations, the recent studies focus on the weakly-supervised setting, with only sentences or action tags in the training videos. However, how to align temporal boundaries and textual descriptions is problematic in most weakly-supervised approaches. To alleviate this difficulty, we propose a novel Dual Masked Modeling (DM2) framework, which can effectively enhance clip-text alignment to boost temporal boundary discovery, by cross-modal masked modeling in the dual fashion. Specifically, we introduce two coupled reconstruction branches, i.e., Clip-Aware Masked Text Modeling (C-MTM), and Text-Aware Masked Clip Modeling (T-MCM), after generating a temporal proposal of the underlying clip. In C-MTM, we recover the masked text with visual assistance of the clip proposal. In T-MCM, we recover the masked clip proposal with lingual assistance of the text. Via such complementary reconstruction supervision, our DM2 can cooperatively exploit robust matching between the video clip and the referred text, allowing to unify grounding and localization in a concise manner. Finally, we perform extensive experiments on the popular temporal benchmarks, i.e., Charades-STA, ActivityNet Captions, ActivityNet-v1.3 and THUMOS-14. Our DM2 achieves state-of-the-art for both weakly-supervised temporal grounding and localization. Codes and models will be released afterward. Yuer Ma, Yi Liu 0081, Limin Wang 0002, Wenxiong Kang, Yu Qiao 0001, Yali Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | PointDC: Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel ClusteringabstractSemantic segmentation of point clouds usually requires exhausting efforts of human annotations, hence it attracts wide attention to the challenging topic of learning from unlabeled or weaker forms of annotations. In this paper, we take the first attempt for fully unsupervised semantic segmentation of point clouds, which aims to delineate semantically meaningful objects without any form of annotations. Previous works of unsupervised pipeline on 2D images fails in this task of point clouds, due to: 1) Clustering Ambiguity caused by limited magnitude of data and imbalanced class distribution; 2) Irregularity Ambiguity caused by the irregular sparsity of point cloud. Therefore, we propose a novel framework, PointDC, which is comprised of two steps that handle the aforementioned problems respectively: Cross-Modal Distillation (CMD) and Super-Voxel Clustering (SVC). In the first stage of CMD, multi-view visual features are back-projected to the 3D space and aggregated to a unified point feature to distill the training of the point representation. In the second stage of SVC, the point features are aggregated to super-voxels and then fed to the iterative clustering process for excavating semantic classes. PointDC1yields a significant improvement over the prior state-of-the-art unsupervised methods, on both the ScanNet-v2 (+18.4 mIoU) and S3DIS (+11.5 mIoU) semantic segmentation benchmarks. Zisheng Chen, Haihong Xiao, Baigui Sun, Xuansong Xie, Wenxiong Kang |
ICCV | 8 |
| 2023 | Ar3dHands: A Dataset and Baseline for Real-Time 3D Hand Pose Estimation from Binocular Distorted Images
Mengting Gan, Yihong Lin, Xingyan Liu, Wenwei Song, Wenxiong Kang |
ICIG (1) | 6 |
| 2023 | HuMoMM: A Multi-Modal Dataset and Benchmark for Human Motion Analysis
Ming Zeng 0012, Wenxiong Kang, Feiqi Deng |
ICIG (1) | 4 |
| 2023 | CostFormer: Cost Transformer for Cost Aggregation in Multi-view StereoabstractThe core of Multi-view Stereo(MVS) is the matching process among reference and source pixels. Cost aggregation plays a significant role in this process, while previous methods focus on handling it via CNNs. This may inherit the natural limitation of CNNs that fail to discriminate repetitive or incorrect matches due to limited local receptive fields. To handle the issue, we aim to involve Transformer into cost aggregation. However, another problem may occur due to the quadratically growing computational complexity caused by Transformer, resulting in memory overflow and inference latency. In this paper, we overcome these limits with an efficient Transformer-based cost aggregation network, namely CostFormer. The Residual Depth-Aware Cost Transformer(RDACT) is proposed to aggregate long-range features on cost volume via self-attention mechanisms along the depth and spatial dimensions. Furthermore, Residual Regression Transformer(RRT) is proposed to enhance spatial attention. The proposed method is a universal plug-in to improve learning-based MVS methods. Yang Liu 0356, Baigui Sun, Wenxiong Kang, Xuansong Xie |
IJCAI | 6 |
| 2023 | Semi-supervised Deep Multi-view StereoabstractSignificant progress has been witnessed in learning-based Multi-view Stereo (MVS) under supervised and unsupervised settings. To combine their respective merits in accuracy and completeness, meantime reducing the demand for expensive labeled data, this paper explores the problem of learning-based MVS in a semi-supervised setting that only a tiny part of the MVS data is attached with dense depth ground truth. However, due to huge variation of scenarios and flexible settings in views, it may break the basic assumption in classic semi-supervised learning, that unlabeled data and labeled data share the same label space and data distribution, named as semi-supervised distribution-gap ambiguity in the MVS problem. To handle these issues, we propose a novel semi-supervised distribution-augmented MVS framework, namely SDA-MVS. For the simple case that the basic assumption works in MVS data, consistency regularization encourages the model predictions to be consistent between original sample and randomly augmented sample. For further troublesome case that the basic assumption is conflicted in MVS data, we propose a novel style consistency loss to alleviate the negative effect caused by the distribution gap. The visual style of unlabeled sample is transferred to labeled sample to shrink the gap, and the model prediction of generated sample is further supervised with the label in original labeled sample. The experimental results in semi-supervised settings of multiple MVS datasets show the superior performance of the proposed method. With the same settings in backbone network, our proposed SDA-MVS outperforms its fully-supervised and unsupervised baselines. Yang Liu 0356, Haihong Xiao, Baigui Sun, Xuansong Xie, Wenxiong Kang |
ACM Multimedia | 8 |
| 2023 | LF-LVS: Label-Free Left Ventricular Segmentation for Transthoracic Echocardiogram
Qing Kang, Wenxiao Tang, Wenxiong Kang |
PRCV (13) | 4 |
| 2023 | Random hand gesture authentication via efficient Temporal Segment Set Network
Yihong Lin, Wenwei Song, Wenxiong Kang |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | Learning image blind denoisers without explicit noise modeling
Lin Nie, Junfan Lin, Wenxiong Kang, Yukai Shi |
Multim. Tools Appl. | 3 |
| 2023 | Distinguishing and Matching-Aware Unsupervised Point Cloud CompletionabstractReal-scanned point clouds are often incomplete due to occlusion, light reflection and limitations of sensor resolution, which impedes the related progress of downstream tasks, e.g., shape classification and object detection. Although there has been impressive research progress on the point cloud completion topic, they rely on the premise of extensive paired training data. However, collecting complete point clouds in some specified scenarios is labor-intensive and even impractical. To mitigate this problem, we propose DMNet, a distinguishing and matching-aware unsupervised point cloud completion network. Our work belongs to the group of unsupervised completion methods but goes beyond previous studies. Firstly, we propose a distinguishing-aware feature extractor to learn discriminable semantic information for different instances, simultaneously enhancing the robust invariant representation under noise disturbances. Secondly, we design a hierarchy-aware hyperbolic decoder to recover the complete geometry of point clouds, which not only can capture the implicit hierarchical relationships in data but also has an explicit extended nature. Finally, we develop a matching-aware refiner to eliminate noise points via aligning the topology structure of the input and predicted partial point clouds. Extensive experiments on MVP, Completion3D and KITTI datasets prove the effectiveness of our method, which performs favorably over state-of-the-art methods both quantitatively and qualitatively. Haihong Xiao, Yuqiong Li, Wenxiong Kang, Qiuxia Wu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | FVFSNet: Frequency-Spatial Coupling Network for Finger Vein AuthenticationabstractFinger vein biometrics is becoming an important source of human authentication due to its advantages in terms of liveness detection, high security, and user convenience. Although there exist a lot of deep learning-based methods for finger vein authentication, they only extract features from finger vein images in the spatial domain and may lose some important information that is present in other domains, such as the frequency domain. Motivated by this conjecture and the remarkable performance of image feature extraction in the frequency domain, this work explores a method capable of extracting finger vein features in both the spatial and frequency domains. Therefore, the features extracted from different domains can complement each other. In addition, we propose a novel frequency-spatial coupling network (FVFSNet) for finger vein authentication. FVFSNet is mainly composed of three parts: (1) the frequency domain processing module (FDPM), (2) the spatial domain processing module (SDPM), and (3) the frequency-spatial coupling module (FSCM). The FDPM is used to extract the finger vein features present in the frequency domain, which is mainly composed of the frequency-spatial domain transformation and the frequency domain convolution layer. The SDPM is used to extract the finger vein features present in the spatial domain, which is mainly composed of convolution layers with an efficient design. The FSCM is used to couple the features extracted from the FDPM and SDPM, which is mainly composed of the channel and spatial attention mechanisms. To validate our conjecture and the performances of FVFSNet, extensive experiments are conducted on nine commonly used publicly available finger vein datasets. Experimental results show that the frequency domain constitutional neural network has a surprising effect on finger vein authentication, and the proposed FVFSNet achieves the state-of-the-art performance with the advantages of lightweight and low computational cost. Junduan Huang, An Zheng, M. Saad Shakeel, Weili Yang, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Depthwise Temporal Non-Local Network for Faster and Better Dynamic Hand Gesture AuthenticationabstractDynamic hand gesture is an emerging and promising biometric trait. It contains both physiological and behavioral characteristics, which on the one hand can theoretically make authentication systems more accurate and more secure, and on the other hand can increase the difficulty of model design because it is essentially a fine-grained video understanding task. For authentication systems, equal error rate (EER) and real-time performance are two vital metrics. Current video understanding-based hand gesture authentication methods mainly focus on lowering the EER while neglecting to reduce the computational cost. In this paper, we propose a 2D CNN-based depthwise temporal non-local network (DwTNL-Net) that can take into account both EER and running efficiency. To enable the DwTNL-Net with spatiotemporal information processing capability, we design a temporal sharpening (TS) module and a DwTNL module for short- and long-term identity feature modeling, respectively. The TS module can assist the backbone in local behavioral characteristic understanding and can simultaneously remove redundant information and highlight behavioral cues while retaining sufficient physiological characteristics. In contrast, the DwTNL module focuses on summarizing global information and discovering stable patterns, which are finally used for local information enhancement. The complementary combination of our TS and DwTNL modules makes DwTNL-Net achieve substantial performance improvements. Extensive experiments on the SCUT-DHGA dataset and sufficient statistical analyses fully demonstrate the superiority and efficiency of our DwTNL-Net. The code is available at https://github.com/SCUT-BIP-Lab/DwTNL-Net. Wenwei Song, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Target Category Agnostic Knowledge Distillation With Frequency-Domain SupervisionabstractExisting knowledge distillation approaches require task-related data to train portable student networks for satisfactory performance. Nevertheless, in real-world applications, a majority of the data is unavailable over concerns about personal privacy, commercial confidentiality, etc. To mitigate the difficulty of acquiring the target dataset, the mainstream knowledge distillation methods generate training samples from the teacher network. However, the generated images still differ from the authentic ones, which limit the student network performance. To solve this issue, we propose a convenient and cost-negligible method to build datasets for training student networks. Specifically, we crawl the data from the web without knowing any category of the target dataset, named irrelevant category crawler data (ICCD). To prevent the performance collapse due to the data distribution gap between our ICCD and the target dataset, we propose a pseudo-classification strategy and frequency-domain supervision (PCFS) for target category agnostic knowledge distillation with ICCD. The pseudo-classification strategy classifies ICCD into different pseudo-categories by the teacher network, and uniformly but randomly samples images from each pseudo-category to construct the pseudo-target dataset. Furthermore, we transform the feature maps from the spatial domain to the frequency domain, and utilize the high- and low-frequency signals of the teacher network to impose strong constraints on the student network. Extensive experiments conducted on various test sets demonstrate the effectiveness of our proposed PCFS, which outperforms existing data-free methods and achieves comparable performance to those using the target training set. Code is available athttps://github.com/SCUT-BIP-Lab/PCFS-DFKD. Wenxiao Tang, M. Saad Shakeel, Zisheng Chen, Wenxiong Kang |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | DDAD: Detachable Crowd Density Estimation Assisted Pedestrian DetectionabstractDetecting pedestrians is a challenging computer vision task, especially in the intelligent transportation system. Mainstream pedestrian detection methods purely utilize information of bounding boxes, which overlooks the role of other valuable attributes (e.g., head, head-shoulders, and keypoints) of pedestrians and leads to sub-optimal solutions. Some works leveraged these valuable attributes with a minor performance improvement at the expense of increased computational complexity during the inference phase. To alleviate this dilemma, we propose a simple yet effective method, namely Detachable crowd Density estimation Assisted pedestrian Detection (DDAD), which leverages the crowd density attributes to assist pedestrian detection in the real-world scenes (e.g., crowded scenes and small-scale pedestrian scenes). The advantage of the crowd density estimation is that it allows the network to focus more on the human head and the small-scale pedestrians, which improves the features representation of pedestrians heavily occluded or far from cameras. Our DDAD works on a principle of multi-task learning and can be seamlessly applied to both one-stage and two-stage pedestrian detectors by equipping them with an extra detachable branch of crowd density estimation. The equipped crowd density estimation branch is trained with the annotations derived from the existing pedestrian bounding box annotations, occurring no extra annotation cost. Moreover, it can be removed during the inference phase without sacrificing the inference speed. Extensive experiments conducted on two challenging datasets, i.e., CrowdHuman and CityPersons, demonstrate that our proposed DDAD achieves a significant improvement upon the state-of-the-art methods. Code is available at https://github.com/SCUT-BIP-Lab/ DDAD. Wenxiao Tang, Kun Liu 0029, M. Saad Shakeel, Wenxiong Kang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | On the Importance of Different Frequency Bins for Speaker VerificationabstractThe majority of modern speaker verification systems take spectral analysis-based features as input, which contains multiple frequency bins. Naturally, there would be a question of whether all different frequency bins contribute equally to the speaker verification system performance? In this paper, we propose the frequency reweighting layer (FRL) to automatically learn and balance the importance of different frequency bins. This new layer can be freely inserted into the original speaker embedding learner once or multiple times at different layers, with an ignorable number of new parameters. Based on the proposed novel architecture, a set of experiments are designed and carried out on the VoxCeleb1 dataset, which not only achieves superior performance but also exhibits an interesting weight distribution – the lower frequencies matter more. Aiwen Deng, Shuai Wang 0016, Wenxiong Kang, Feiqi Deng |
ICASSP | 3 |
| 2022 | Semantic Center Guided Windows Attention Fusion Framework for Food Recognition
Yongxin Zhou 0003, Wenxiong Kang, Zeng Ming |
PRCV (2) | 4 |
| 2022 | Endowing rotation invariance for 3D finger shape and vein verification
Weili Yang, Qiuxia Wu, Wenxiong Kang |
Frontiers Comput. Sci. | 4 |
| 2022 | Unconstrained face sketch synthesis via perception-adaptive network and a new benchmark
Lin Nie, Lingbo Liu, Zhengtao Wu, Wenxiong Kang |
Neurocomputing | 4 |
| 2022 | Multi-scale attention guided network for end-to-end face alignment and recognition
M. Saad Shakeel, Wenxiong Kang, Arif Mahmood |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | Multi-label image recognition with attentive transformer-localizer module
Lin Nie, Tianshui Chen, Zhouxia Wang, Wenxiong Kang, Liang Lin 0004 |
Multim. Tools Appl. | 4 |
| 2022 | Learning upper patch attention using dual-branch training strategy for masked face recognition
M. Saad Shakeel, Wenxiong Kang |
Pattern Recognit. | 5 |
| 2022 | Study on Reflection-Based Imaging Finger Vein RecognitionabstractFinger vein modality plays an important role in biometrics due to its stability and security. However, existing state-of-the-art finger vein recognition systems adopt the transmission-based imaging mode with a sealed design, which requires redundant space in the imaging device and results in an uncomfortable user experience. Consequently, we design a reflection-based imaging device with an open structure to reduce the device volume and improve the portability, as well as the user experience. However, an open structure of the device inevitably introduces extra illumination variation to the image, which may deteriorate the performance of the system. In this paper, we propose Domain Adaptation Finger Vein Network (DAFVN) to narrow the domain shift between different illumination data domains and extract illumination-invariant features from finger vein images, improving the robustness to illumination variations. To evaluate the performance of DAFVN and remedy the lack of a publicly open reflection-based finger vein database, we use the self-made device to construct the first large-scale reflection-based finger vein database, namely SCUT Reflective Imaging Finger Vein database (SCUT-RIFV). It includes 32,064 images from 167 subjects with five different illumination conditions. Abundant experiments implemented on the SCUT-RIFV database indicate that the proposed method can effectively alleviate the influence of illumination variation on the reflection-based finger vein recognition system. Zejun Zhang 0008, Fei Zhong, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Self-supervised Multi-view Stereo via Effective Co-Segmentation and Data-AugmentationabstractRecent studies have witnessed that self-supervised methods based on view synthesis obtain clear progress on multi-view stereo (MVS). However, existing methods rely on the assumption that the corresponding points among different views share the same color, which may not always be true in practice. This may lead to unreliable self-supervised signal and harm the final reconstruction performance. To address the issue, we propose a framework integrated with more reliable supervision guided by semantic co-segmentation and data-augmentation. Specially, we excavate mutual semantic from multi-view images to guide the semantic consistency. And we devise effective data-augmentation mechanism which ensures the transformation robustness by treating the prediction of regular samples as pseudo ground truth to regularize the prediction of augmented samples. Experimental results on DTU dataset show that our proposed methods achieve the state-of-the-art performance among unsupervised methods, and even compete on par with supervised methods. Furthermore, extensive experiments on Tanks&Temples dataset demonstrate the effective generalization ability of the proposed method. Yu Qiao 0001, Wenxiong Kang, Qiuxia Wu |
AAAI | 4 |
| 2021 | Learning Discriminative Speaker Embedding by Improving Aggregation Strategy and Loss Function for Speaker VerificationabstractThe embedding-based speaker verification (SV) technology has witnessed significant progress due to the advances of deep convolutional neural networks (DCNN). However, how to improve the discrimination of speaker embedding in the open world SV task is still the focus of current research in the community. In this paper, we improve the discriminative power of speaker embedding from three-fold: (1) NeXtVLAD is introduced to aggregate frame-level features, which decomposes the high-dimensional frame-level features into a group of low-dimensional vectors before applying VLAD aggregation. (2) A multi-scale aggregation strategy (MSA) assembled with NeXtVLAD is designed with the purpose of fully extract speaker information from the frame-level feature in different hidden layers of DCNN. (3) A mutually complementary assembling loss function is proposed to train the model, which consists of a prototypical loss and a marginal-based softmax loss. Extensive experiments have been conducted on the VoxCeleb-1 dataset, and the experimental results show that our proposed system can obtain significant performance improvements compared with the baseline, and obtains new state-of-the-art results. The source code of this paper is available at https://github.com/LCF2764/Discriminative-Speaker-Embedding. Chengfang Luo, Aiwen Deng, Junhong Zhao, Wenxiong Kang |
IJCB | 6 |
| 2021 | TDS-Net: Towards Fast Dynamic Random Hand Gesture Authentication via Temporal Difference Symbiotic Neural NetworkabstractHand gesture is a new emerging biometric trait containing both physiological and behavioral characteristics. With the popularity of various cameras, and the rich identity features and contactless authentication mode embedded in gestures themselves, vision-based hand gesture authentication has great potential value. However, current hand gesture authentication methods heavily rely on defined gestures and require identical enrollment and verification gestures, which limits the user-friendliness and efficiency of authentication. It is arguably true that authentication in a simpler and faster way, without the need to remember gestures, will be more approachable. Thus, a fast dynamic random hand gesture authentication method is introduced, in which users can perform a random improvised gesture in both the enrollment and verification stage. To better utilize the physiological and behavioral characteristics of hand gestures, an efficient network named Temporal Difference Symbiotic Neural Network (TDS-Net) equipped with our designed behavioral energy-based feature fusion module (BE-Fusion module) is proposed. Extensive experiments on the SCUT-DHGA dataset demonstrate that TDS-Net outperforms the recent state-of-the-art methods. Wenwei Song, Wenxiong Kang, Linpu Fang, Chang Liu 0060, Xingyan Liu |
IJCB | 2 |
| 2021 | LFMB-3DFB: A Large-scale Finger Multi-Biometric Database and Benchmark for 3D Finger BiometricsabstractFinger contains several discriminative biometric traits, including fingerprint, finger vein, finger knuckle, and finger shape, which are complementary in identity information. However, in most current researches and practical applications, only a single or several traits are utilized, which are prone to unsatisfactory recognition performance and easy forgery. Our work is the first attempt to collect and study all biometric traits on the finger. Firstly, a novel multi-view, multi-spectral 3D finger imaging system is designed. To the best of our knowledge, it is the first biometric imaging system that can capture almost all finger-based traits. With this 3D finger imaging system, we scanned numerous fingers, acquiring their external skin images and internal vein images from 6 different views. Then 3D finger models with skin and vein textures are reconstructed by space carving, mesh regularization, and texture mapping algorithms. Secondly, we establish a benchmark dataset, namely the Large- scale Finger Multi-Biometric database and benchmark for 3D Finger Biometrics (LFMB-3DFB). LFMB-3DFB contains 695 fingers, and each finger is captured 10 times. Then, 6 finger skin images and 6 finger vein images are obtained for each acquisition, and final 83,400 images and 6,950 3D finger models are obtained. Besides, we designed a more rigorous and comprehensive evaluation protocol for both identification and verification tasks. Finally, we designed corresponding baselines for 2D finger traits recognition, multi-view finger traits recognition, 3D finger traits recognition, and score-level fusion. Rigorous experiments have been conducted to verify the significance and usefulness of the proposed LFMB-3DFB. Weili Yang, Zhuoming Chen, Junduan Huang, Wenxiong Kang |
IJCB | 5 |
| 2021 | Digging into Uncertainty in Self-supervised Multi-view StereoabstractSelf-supervised Multi-view stereo (MVS) with a pretext task of image reconstruction has achieved significant progress recently. However, previous methods are built upon intuitions, lacking comprehensive explanations about the effectiveness of the pretext task in self-supervised MVS. To this end, we propose to estimate epistemic uncertainty in self-supervised MVS, accounting for what the model ignores. Specially, the limitations can be categorized into two types: ambiguious supervision in foreground and invalid supervision in background. To address these issues, we propose a novel Uncertainty reduction Multi-view Stereo (U-MVS) framework for self-supervised learning. To alleviate ambiguous supervision in foreground, we involve extra correspondence prior with a flow-depth consistency loss. The dense 2D correspondence of optical flows is used to regularize the 3D stereo correspondence in MVS. To handle the invalid supervision in background, we use Monte-Carlo Dropout to acquire the uncertainty map and further filter the unreliable supervision signals on invalid regions. Extensive experiments on DTU and Tank&Temples benchmark show that our U-MVS framework1achieves the best performance among unsupervised MVS methods, with competitive performance with its supervised opponents. Yali Wang 0001, Wenxiong Kang, Baigui Sun, Hao Li 0030, Yu Qiao 0001 |
ICCV | 4 |
| 2021 | PVLNet: Parameterized-View-Learning neural network for 3D shape recognition
Lvequan Wang, Qiuxia Wu, Wenxiong Kang |
Comput. Graph. | 4 |
| 2021 | A survey on dorsal hand vein biometrics
Wei Jia 0001, Bob Zhang 0001, Yang Zhao 0002, Lunke Fei, Wenxiong Kang, Di Huang 0001, Guodong Guo |
Pattern Recognit. | 6 |
| 2021 | Dynamic-Hand-Gesture Authentication Dataset and BenchmarkabstractIn recent years, biometrics have received considerable attention for its reliability and usability. Dynamic-hand-gesture is one of the representative biometric modalities, with advantages of safety and template-replaceability, has huge potential value. However, due to the lack of large-scale dataset and comprehensive evaluation methods, few researches are intended to study the dynamic-hand-gesture authentication method. In this article, we introduce a new dataset SCUT-DHGA, which is the first large-scale Dynamic-Hand-Gestures-Authentication dataset. SCUT-DHGA contains 29,160 dynamic-hand-gesture video sequences and more than 1.86 million frames for both color and depth modalities acquired from 193 volunteers. Six kinds of dynamic-hand-gestures are carefully designed for researching two types of authentication tasks: gesture-predefined authentication and gesture-free authentication. To investigate the hypothesis that users' gestures would be variant after time-span, which will degrade the performance of a dynamic-hand-gesture authentication system, two separate sessions' data were acquired from 50 volunteers with an average interval of one week. Beside the SCUT-DHGA dataset, we also benchmark this dataset with our proposed DHGA-net. By releasing such a large-scale dataset and benchmark, we expect dynamic-hand-gesture authentication methods to gain further improvement and generalization. Chang Liu 0060, Xingyan Liu, Linpu Fang, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | Joint Input and Output Space Learning for Multi-Label Image ClassificationabstractMulti-label image classification aims to predict the labels associated with a given image. While most existing methods utilize unified image representations, extracting label-specific features through input space learning would improve the discriminative power of the learned features. On the other hand, most feature learning studies often ignore the learning in the output label space, although taking advantage of label correlations can boost the classification performance. In this paper, we propose a deep learning framework that incorporates flexible modules which can learn from both input and output spaces for multi-label image classification. For the input space learning, we devise a label-specific feature pooling method to refine convolutional features for obtaining features specific to each label. For the output space learning, we design a Two-Stream Graph Convolutional Network (TSGCN) to learn multi-label classifiers by mapping spatial object relationships and semantic label correlations. More specifically, we build object spatial graphs to characterize the spatial relationships among objects in an image, which supplements the label semantic graphs modelling the semantic label correlations. Experimental results on two popular benchmark datasets (i.e., Pascal VOC and MS-COCO) show that our proposed method achieves superior performance over the state-of-the-arts. Jiahao Xu 0002, Hongda Tian, Zhiyong Wang 0001, Yang Wang 0002, Wenxiong Kang, Fang Chen 0001 |
IEEE Trans. Multim. | 5 |
| 2020 | Universal-RCNN: Universal Object Detector via Transferable Graph R-CNNabstractThe dominant object detection approaches treat each dataset separately and fit towards a specific domain, which cannot adapt to other domains without extensive retraining. In this paper, we address the problem of designing a universal object detection model that exploits diverse category granularity from multiple domains and predict all kinds of categories in one system. Existing works treat this problem by integrating multiple detection branches upon one shared backbone network. However, this paradigm overlooks the crucial semantic correlations between multiple domains, such as categories hierarchy, visual similarity, and linguistic relationship. To address these drawbacks, we present a novel universal object detector called Universal-RCNN that incorporates graph transfer learning for propagating relevant semantic information across multiple datasets to reach semantic coherency. Specifically, we first generate a global semantic pool by integrating all high-level semantic representation of all the categories. Then an Intra-Domain Reasoning Module learns and propagates the sparse graph representation within one dataset guided by a spatial-aware GCN. Finally, an Inter-Domain Transfer Module is proposed to exploit diverse transfer dependencies across all domains and enhance the regional feature representation by attending and transferring semantic contexts globally. Extensive experiments demonstrate that the proposed method significantly outperforms multiple-branch models and achieves the state-of-the-art results on multiple object detection benchmarks (mAP: 49.1% on COCO). Hang Xu 0004, Linpu Fang, Xiaodan Liang, Wenxiong Kang, Zhenguo Li |
AAAI | 4 |
| 2020 | Dynamic Group Convolution for Accelerating Convolutional Neural Networks
Zhuo Su 0002, Linpu Fang, Wenxiong Kang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
ECCV (6) | 3 |
| 2020 | JGR-P2O: Joint Graph Reasoning Based Pixel-to-Offset Prediction Network for 3D Hand Pose Estimation from a Single Depth Image
Linpu Fang, Xingyan Liu, Li Liu 0002, Wenxiong Kang |
ECCV (6) | 5 |
| 2020 | Real-time hand posture recognition using hand geometric features and Fisher Vector
Linpu Fang, Ningxin Liang, Wenxiong Kang, Zhiyong Wang 0001, David Dagan Feng |
Signal Process. Image Commun. | 3 |
| 2020 | Correlation Filter Tracking via Distractor-Aware Learning and Multi-Anchor DetectionabstractCorrelation filter has demonstrated the power in object tracking, benefiting from its superior speed and competitive performance. However, existing correlation filter based trackers (CFTs) are fragile for some inherent defects caused by the boundary effect. To address this issue, we propose a novel correlation filter based tracking framework by integrating three highly collaborative components, including a fast target proposal module, a distractor-aware filter, and a correlation filter based refiner. Specifically, the target proposal aims at determining some target-like regions in contexts efficiently, which provides target-like patches to learn a distractor-aware filter and detect. Multi-region strategy enlarges space fields for learning and prediction. The filter learned from both target and distractors enhances its ability to identify background. Therefore, our method is capable of evaluating multiple candidates in wider context with less risk of drifting to distractors, namely multi-anchor detection. Besides, the proposed Proposal-Detect-Refine hierarchical searching process progressively achieves data alignment between testing and training samples, which benefits for reliable model prediction. A refiner is used to fine-tune positions after multi-anchor detection for lessening error accumulation and preventing model from drifting. Comprehensive experiments on five challenging datasets, i.e. OTB2013, OTB2015, VOT2017, VOT19, and TC128, demonstrate that the proposed method achieves superior performance against the state-of-the-art methods. Guochun Chen, Gengzheng Pan, Yongxin Zhou 0003, Wenxiong Kang, Junhui Hou, Feiqi Deng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Study of a Full-View 3D Finger Vein Verification TechniqueabstractFinger vein modality has unique advantages, allowing it to play an important role in biometrics. However, the approach to vein imaging and information acquisition typically adopted in current vein verification systems employs a monocular camera to acquire a single-view 2D vein image from only one side of the finger, which causes two problems: it acquires limited vein pattern information for verification, and it causes clear differences among samples of the same subject captured from different finger positions in contact-free mode. Both of these problems have adverse effects on system performance. In general, existing systems are more sensitive to positional variations of the finger, particularly those caused by pitch and roll movements. This concern remains a challenge despite considerable efforts to address it in recent years. To provide a fundamental solution to the above issues, we propose an entirely new system, which includes a software and hardware platform that collects a full-view of the vein pattern information from whole fingers with three cameras, a novel 3D reconstruction method to build the full-view 3D finger vein image, and a corresponding 3D finger vein feature extraction and matching strategy based on a lightweight convolutional neural network (CNN) with depthwise separable convolution. Experimental results demonstrate the potential of our proposed system and show that compared to the traditional single-view 2D mode of finger vein recognition, the new system both efficiently improves the recognition performance and simultaneously takes full advantage of additional valid information provided by the finger vein biometrics. Wenxiong Kang, Feiqi Deng |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2019 | Correlation filter tracker with siamese: A robust and real-time object tracking framework
Gengzheng Pan, Guochun Chen, Wenxiong Kang, Junhui Hou |
Neurocomputing | 3 |
| 2019 | Feature covariance matrix-based dynamic hand gesture recognition
Linpu Fang, Guile Wu, Wenxiong Kang, Qiuxia Wu, Zhiyong Wang 0001, David Dagan Feng |
Neural Comput. Appl. | 3 |
| 2019 | From Noise to Feature: Exploiting Intensity Distribution as a Novel Soft Biometric Trait for Finger Vein RecognitionabstractMost finger vein feature extraction algorithms achieve satisfactory performance due to their texture representation abilities, despite simultaneously ignoring the intensity distribution that is formed by the finger tissue, and in some cases, processing it as background noise. In this paper, we exploit this kind of “noise” as a novel soft biometric trait for achieving better finger vein recognition performance. First, a detailed analysis of the finger vein imaging principle and the characteristics of the image are presented to show that the intensity distribution that is formed by the finger tissue in the background can be extracted as a soft biometric trait for recognition. Then, two finger vein background layer extraction algorithms and three soft biometric trait extraction algorithms are proposed for intensity distribution feature extraction. Finally, a hybrid matching strategy is proposed to solve the issue of dimension difference between the primary and soft biometric traits on the score level. A series of rigorous contrast experiments on three open-access databases demonstrate that our proposed method is feasible and effective for finger vein recognition. Wenxiong Kang, Wei Jia 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | FV-Net: learning a finger-vein feature representation based on a CNNabstractFinger vein pattern has been proven to be an effective biometric for personal identification in recent years. Nevertheless, there remain challenges that need to be solved, such as finger-vein features that lack robustness and expressiveness. In this paper, we propose a deep convolutional neural network (CNN) model, named the Finger-vein Network (FV-Net), to learn the features representative of a finger vein that is more discriminative and robust than handcrafted features. Next, to address the issue of translation and rotation in vein imaging, we propose a template-like matching strategy while designing the top architecture of the FV-net to extract features with spatial information. Finally, the extensive experimental results show that our proposed method can achieve excellent performance on several public datasets. Wenxiong Kang, Yuxun Fang, Junhong Zhao, Feiqi Deng |
ICPR | 2 |
| 2018 | A novel finger vein verification system based on two-stream convolutional network learning
Yuxun Fang, Qiuxia Wu, Wenxiong Kang |
Neurocomputing | 3 |
| 2018 | Finger Vein Presentation Attack Detection Using Total Variation DecompositionabstractFinger vein recognition is an emerging biometric technique for personal authentication that has garnered considerable attention in the past decade. Although shown to be effective, recent studies have revealed that finger vein biometrics is also vulnerable to presentation attacks, i.e., printed versions of authorized individual finger vein images can be used to gain access to facilities or services. In this paper, given that both blurriness and the noise distribution are slightly different between real and forged finger vein images, we propose an efficient and robust method for detecting presentation attacks that use forged finger vein images (print artifacts). First, we use total variation regularization to decompose original finger vein images into structure and noise components, which represent the degrees of blurriness and the noise distribution. Second, a block local binary pattern descriptor is used to encode both structure and noise information in the decomposed components. Finally, we use a cascaded support vector machine model for classification, by which finger vein presentation attacks can be effectively detected. To evaluate the performance of our approach, we constructed a new finger vein presentation attack database. Extensive experimental results gleaned from the two finger vein presentation attack databases and a palm vein presentation attack database show that our method clearly outperforms state-of-the-art methods. Xinwei Qiu, Wenxiong Kang, Senping Tian, Wei Jia 0001, Zhixing Huang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Information Divergence-Based Matching Strategy for Online Signature VerificationabstractThe phenomenon of data clutter caused by intervariability (individual features) and intravariability (intrinsic noise of reference samples) is one of the most important reasons for performance degradation in online signature verification systems. To address this problem, we introduce information divergence in the field of signature verification to take full advantage of the information contained in all reference samples by shifting the distance measurement between the test and reference samples to a similarity measurement between two distributions (generated between reference samples as well as between the test sample and all reference samples) and simultaneously make full use of the spatial information of the reference samples. Based on that change, we propose a novel information divergence framework and provide some matcher instances that work within the proposed framework to effectively improve the performance of a signature verification system. Furthermore, to exploit the advantages of the new matching strategy, we propose a new dynamic time warping algorithm. In addition, we provide in-depth analysis of several distance normalizations and apply them to signature verification to reduce the adverse effect of “clutter” on signature data, which can effectively improve system performance. The experimental results on the MCYT-100 and SUSIG signature databases achieved equal error rates of 2.25% and 1.70% when 10 reference samples were used and 3.16% and 2.13% when 5 reference samples were used, respectively, illustrating the effectiveness of the proposed strategy in relation to other state-of-the-art strategies. Wenxiong Kang, Yuxun Fang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Real-Time Long-Term Tracking With Prediction-Detection-CorrectionabstractReal-time long-term visual tracking is one of the most challenging problems in computer vision due to various factors such as occlusion and motion ambiguity. To achieve robust long-term tracking, most state-of-the-art methods typically construct an online detector in each frame. However, they fail to achieve real-time performance due to high computational complexity. In this paper, we propose a novel real-time long-term tracking algorithm by exploiting a joint Prediction-Detection-Correction Tracking framework (PDCT). We utilize a superpixel optical flow to construct a predictor to estimate the target motion and internal scale variation. To locate the target at a finer level, we develop an improved kernelized correlation detector with an adaptive online learning rate and translation-scale parameters from the predictor. To refine the tracking result and redetect the target in the case of a tracking failure, we devise a corrector utilizing dual online SVMs with dense sampling and reliable history samples. The SVMs are trained with passive-aggressive learning and online retraining strategies. In addition, we employ a selection mechanism for the correlation responses to maintain reliable samples effectively. As a result, our proposed tracker is able to refine tracking results via the corrector and detector and maintains reliable tracking results for subsequent tracking. Extensive experiments on the widely used object tracking benchmark show that the proposed tracker is superior to state-of-the-art trackers in terms of both effectiveness and efficiency, and the integration of each component is effective under the PDCT framework. Ningxin Liang, Guile Wu, Wenxiong Kang, Zhiyong Wang 0001, David Dagan Feng |
IEEE Trans. Multim. | 3 |
| 2017 | Visual tracking utilizing robust complementary learner and adaptive refiner
Guile Wu, Wenxiong Kang, Zhiyong Wang 0001, David Dagan Feng |
Neurocomputing | 3 |
| 2017 | Exploiting superpixel and hybrid hash for kernel-based visual tracking
Guile Wu, Wenxiong Kang |
Pattern Recognit. | 2 |
| 2017 | Vision-Based Fingertip Tracking Utilizing Curvature Points Clustering and Hash Model RepresentationabstractFingertip tracking plays an increasingly important role in augmented reality and virtual-reality applications. However, existing approaches (either continuous-detection-based or separated-detection-tracking-based methods) cannot effectively learn temporal-spatial information or even separate fingertip tracking into unrelated stages, which causes poor real-time performance and incomplete tracking continuity. Moreover, due to the need for high-cost devices, the high degrees of freedom of the hand, and subtle differences among fingers, fingertip tracking remains a challenging task. To address these problems, we propose a novel tracking-combined-with-detection approach for vision-based fingertip tracking. By adopting clustering and geometric constraint analysis, we develop a curvature points clustering method for fingertip detection. Then, by exploiting the identified fingertip points for motion estimation with bidirectional optical flows and temporal-spatial probability calculation, the tracking stage is effectively integrated with the detection stage. To accurately locate the fingertip, we represent the fingertip model with a perceptual hash sequence and locate the fingertip by searching for the best-matching region. Extensive experimental results show the superiority of the proposed algorithm to commonly used and state-of-the-art methods and demonstrate its effectiveness and practicability. Guile Wu, Wenxiong Kang |
IEEE Trans. Multim. | 2 |
| 2016 | Real-time vehicle detection with foreground-based cascade classifierabstractThe strategy based on Haar‐like features and the cascade classifier for vehicle detection systems has captured growing attention for its effectiveness and robustness; however, such a vehicle detection strategy relies on exhaustive scanning of an entire image with different sizes sliding windows, which is tedious and inefficient, since a vehicle only occupies a small part of the whole scene. Therefore, the authors propose a real‐time vehicle detection algorithm which is based on the improved Haar‐like features and combines motion detection with a cascade of classifiers. They adopt a visual background extractor, accompanied by morphological processing, to obtain foregrounds. These foregrounds retain vehicle features and provide the positions within images where vehicles are most likely to be located. Subsequently, vehicle detection is performed only at these positions by using a cascade of classifiers instead of a single strong classifier, which is able to improve the detection performance. The authors’ algorithm has been successfully evaluated on the public datasets, which demonstrates its robustness and real‐time performance. Xiaobin Zhuang, Wenxiong Kang, Qiuxia Wu |
IET Image Process. | 2 |
| 2016 | Robust Fingertip Detection in a Complex EnvironmentabstractFingertip detection has a broad application in gesture recognition and finger tracking. It is also an important foundation of human-computer interaction systems. However, most algorithms are suitable for simple conditions with low accuracy because the hand is a nonrigid object, and its appearance model is complex. To address the challenging problem of accurately detecting fingertips in a complex environment, we propose a novel and robust fingertip detection algorithm in this paper. Unlike existing methods, our study requires no special device or mark, and users are free to move their hands. Via dense optical flow and a skin filter, we perform complete hand region segmentation in a complex environment. We find the maximum value of the local centroid distance outside the centroid circles and identify fingertips. Our algorithm performs favorably compared with common hand region segmentation and fingertip detection methods. Thorough experimentation proves that our proposed algorithm is effective and robust. Guile Wu, Wenxiong Kang |
IEEE Trans. Multim. | 2 |
| 2015 | Palm vein recognition based on multi-sampling and feature-level fusion
Xuekui Yan, Wenxiong Kang, Feiqi Deng, Qiuxia Wu |
Neurocomputing | 2 |
| 2015 | Fast Representation Based on a Double Orientation Histogram for Local Image DescriptorsabstractIn recent years, extensive research on local invariant features has been conducted, and many novel descriptors have been developed for different scenarios. Frequently, these descriptors can offer unique advantages (such as stability, precision, and speed) for select applications but perform unsatisfactorily from a comprehensive perspective. Consequently, a novel local image descriptor, named fast representation using a double orientation histogram (FRDOH), is developed in this paper based on existing descriptors. First, a region is divided using intensity order (as in the local intensity order pattern descriptor) to encode spatial information. Then, the discriminability of the descriptor is enhanced using our proposed double orientation histogram. Second, to further improve the discriminability of the descriptor, the Hellinger distance is used to balance the effects of large and small bins in the histogram for the similarity measure of features. Finally, a novel interpolation strategy known as rapidly cascaded interpolation is used to calculate the intensity of the neighboring points to reduce the computation time, while achieving high precision. The performance of the developed descriptor is evaluated via numerous experiments on the affine covariant feature data set of the Oxford data set, a subset of a 3D object data set, and a subset of the IIT Delhi Touchless Palmprint data set. These experiments demonstrate that the developed FRDOH descriptor outperforms the state-of-the-art descriptors in terms of comprehensive performance. Wenxiong Kang |
IEEE Trans. Image Process. | 1 |
| 2014 | A new descriptor resistant to affine transformation and monotonic intensity change
Zeyi Huang, Wenxiong Kang, Qiuxia Wu |
Comput. Vis. Image Underst. | 2 |
| 2014 | Contactless Palm Vein Recognition Using a Mutual Foreground-Based Local Binary PatternabstractLocal binary pattern (LBP) is popular for the texture representation owing to its discrimination ability and computational efficiency, but when used to describe the sparse texture in palm vein images, the discrimination ability is diluted, leading to lower performance, especially for contactless palm vein matching. In this paper, an improved mutual foreground LBP method is presented for achieving a better matching performance for contactless palm vein recognition. First, the normalized gradient-based maximal principal curvature algorithm and k -means method are utilized for texture extraction, which can effectively suppress noise and improve accuracy and robustness. Then, an LBP matching strategy was adopted for similarity measurements on the basis of extracted palm veins and their neighborhoods, which include the vast majority of useful distinctive information for identification while eliminating interference by excluding the background. To further improve the LBP performance, the matched pixel ratio was adopted to determine the best matching region (BMR). Finally, the matching score obtained in the process of finding the BMR was fused with results of LBP matching at the score level to further improve the identification performance. A series of rigorous contrast experiments using the palm vein data set in the CASIA multispectral palmprint image database were conducted. The obtained low equal error rate (0.267%) and comparisons with the most state-of-the-art approaches demonstrate that our method is feasible and effective for contactless palm vein recognition. Wenxiong Kang, Qiuxia Wu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2014 | Pose-Invariant Hand Shape Recognition Based on Finger GeometryabstractIn this paper, a pose-invariant hand shape recognition method based on the geometry of the fingers is proposed. Firstly, inspired by the segmentation method presented by Yoruk et al., we conduct a novel improvement on the segmentation for extracting the region of the fingers when the hand is in a natural pose. Secondly, Fourier descriptors and finger area functions are employed to extract the finger boundary curve features and region areas, respectively. Finally, score-level fusion based on a weighted sum is used to obtain matching results. Because the finger segmentation strategy and the feature extraction method are both rotation and translation invariant, the proposed method is more suitable for a naturally posed hand. Experiments using the Bogazici University Hand database show that the proposed method can achieve an equal error rate of 0.0369 for all data and 0.0273 for samples with an intragroup angle deviation of less than 45°. Thus, the proposed method is suitable for real-world applications. Wenxiong Kang, Qiuxia Wu |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2013 | Discriminative two-level feature selection for realistic human action recognition
Qiuxia Wu, Zhiyong Wang 0001, Feiqi Deng, Yong Xia 0001, Wenxiong Kang, David Dagan Feng |
J. Vis. Commun. Image Represent. | 5 |
| 2012 | Vein pattern extraction based on vectorgrams of maximal intra-neighbor difference
Wenxiong Kang |
Pattern Recognit. Lett. | 1 |
| 2010 | Direct gray-scale extraction of topographic features for vein recognition
Wenxiong Kang, Huasong Li, Feiqi Deng |
Sci. China Inf. Sci. | 1 |