VLDB 2026 Research / reviewers in the wild / expert
Wenwei Song
dblp:78/6464
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-7787-6604ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Perceptron Constancy for Real-World Dynamic Hand Gesture AuthenticationabstractDynamic hand gesture authentication (DHGA) has emerged as a promising biometric technology, offering enhanced theoretical security over conventional unimodal systems by combining both physiological and behavioral characteristics. Existing DHGA research predominantly focuses on controlled lab conditions, therefore showing low generalizability to uncontrolled application conditions. To bridge this gap, we propose a novel Skeleton-assistant Standardization and Authentication Framework (SSAF) that incorporates a generic data preprocessing method before authentication. First, we introduce a Geometry- Environment Standardization (GE-Stan) method to standardize five primary geometric and environmental factors inducing data distribution discrepancy, significantly improving robustness across different sessions and scenarios. Notably, the GE-Stan method can be applied to most existing algorithms and brings substantial improvement. Second, we design an Appearance and Motion Network (AM-Net) to fully leverage standardized video and skeleton data. It decouples appearance and motion features using specialized representation and processing strategies. Therefore, our SSAF achieves a flexible balance between accuracy and efficiency, enabling up to 3.6× efficiency boost with only minor accuracy trade-offs. Finally, to support real-world evaluation, we also contribute a new challenging dataset, SCUT-RealDHGA, captured under uncontrolled practical conditions with diverse backgrounds and illuminations. Extensive experiments across three DHGA datasets demonstrate that SSAF outperforms existing methods in terms of accuracy, efficiency, and robustness. The code and dataset are available at https://github.com/SCUTBIP-Lab/SSAF. Xilai Wang, Wenwei Song, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Multiscale Super-Images for Dynamic Hand Gesture AuthenticationabstractThe dynamic hand gesture is an emerging biometric trait that has attracted the attention of researchers due to its rich physiological and behavioral characteristics. The previous studies primarily focused on extracting and utilizing the physiological characteristics, while ignoring the rich behavioral characteristics contained in hand gesture movements. The dynamic hand gesture authentication performance will be improved if behavioral characteristics can be effectively extracted and fused with physiological characteristics for authentication. In addition, existing methods still suffer from insufficient feature extraction capabilities and low efficiency in extracting behavioral characteristics from complex dynamic hand gestures. To address these issues, this paper first proposes multiscale dynamic hand gesture (MDHG) super-images to represent the behavioral characteristics of hand gestures, containing sufficient local and global motion cues. Furthermore, for the super-images, this paper proposes a two-stream network consisting of a spatiotemporal feature extraction backbone and an identity-aggregation module to fully extract and fuse the physiological and behavioral characteristics of hand gestures, which significantly improves the accuracy of dynamic hand gesture authentication. Extensive experiments on two benchmark datasets, SCUT-DHGA and HandLogin, show that our method achieves superior performance with fewer parameters and FLOPs than other networks, validating the effectiveness, generalizability, and security of our proposed method. Zenan Lin, Wenwei Song, Wenxiong Kang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | L3AM: Linear Adaptive Additive Angular Margin Loss for Video-Based Hand Gesture Authentication
Wenwei Song, Wenxiong Kang, Adams Wai-Kin Kong, Yitao Qiao |
Int. J. Comput. Vis. | 1 |
| 2024 | Hand Gesture Authentication by Discovering Fine-Grained Spatiotemporal Identity CharacteristicsabstractDynamic hand gesture is an emerging and promising biometric trait containing both physiological and behavioral characteristics. Possessing the two kinds of characteristics makes dynamic hand gesture have more identity information enabling more accurate and secure authentication theoretically, but also poses a challenge of efficient fine-grained spatiotemporal feature extraction. This challenge involves a seemingly paradoxical problem that high-frame-rate videos are required for behavioral characteristic analysis, but they can also introduce high computational costs. To mitigate this issue, we propose a Frequency Spatiotemporal Attention Network (FSTA-Net) with a focus on satisfying the high-performance and low-computation requirements of authentication systems. The FSTA-Net is established with a two-stage identity characteristic analysis paradigm for short- and long-term modeling. Specifically, considering that models prefer to analyze physiological characteristics which are relatively straightforward to understand, we first design a Behavior Enhanced (BE) module to emphasize hand motions and reduce redundant information to facilitate local identity feature distillation in the first stage. We then present a Frequency Spatiotemporal Attention (FSTA) module to summarize global identity features with decent FLOPs and GPU memory occupation in the second stage. Incorporating the BE and FSTA modules enables them to complement each other’s strengths, resulting in a clear-cut improvement in equal error rate and running speed. Extensive experiments on the SCUT-DHGA dataset demonstrate the superiority of the FSTA-Net. The code is available athttps://github.com/SCUT-BIP-Lab/FSTA-Net. Wenwei Song, Wenxiong Kang, Liang Lin 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Learning an Augmented RGB Representation for Dynamic Hand Gesture AuthenticationabstractDynamic hand gesture authentication aims to recognize users’ identity through the characteristics of their hand gestures. How to extract favorable features for verification is the key to success. Cross-modal knowledge distillation is an intuitive approach that can introduce additional modality information in the training phase to enhance the target modality representation, improving model performance without incurring additional computation in the inference phase. However, most previous cross-modal knowledge distillation methods directly transfer information from one modality to another one without considering the modality gap. In this paper, we propose a novel translation mechanism in cross-modal knowledge distillation that can effectively mitigate the modality gap and utilize the information from the additional modality to enhance the target modality representation. In order to better transfer modality information, we propose a novel modality fusion-enhanced non-local (MFENL) module, which can fuse the multi-modal information from the teacher network and enhance the fused features based on the modality input into the student network. We use cascaded MFENL modules as the translator based on the proposed cross-modal knowledge distillation method to learn an enhanced RGB representation for dynamic hand gesture authentication. Extensive experiments on the SCUT-DHGA dataset demonstrate that our method has compelling advantages over the state-of-the-art methods. The code is available athttps://github.com/SCUT-BIP-Lab/TranslationCKD. Huilong Xie, Wenwei Song, Wenxiong Kang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Robust and Accurate Hand Gesture Authentication With Cross-Modality Local-Global Behavior AnalysisabstractObtaining robust fine-grained behavioral features is critical for dynamic hand gesture authentication. However, behavioral characteristics are abstract and complex, making them more difficult to capture than physiological characteristics. Moreover, various illumination and backgrounds in practical applications pose additional challenges to existing methods because commonly used RGB videos are sensitive to them. To overcome this robustness limitation, we propose a two-stream CNN-based cross-modality local-global network (CMLG-Net) with two complementary modules to enhance the discriminability and robustness of behavioral features. First, we introduce a temporal scale pyramid (TSP) module consisting of multiple parallel convolution subbranches with different temporal kernel sizes to capture the fine-grained local motion cues at various temporal scales. Second, a cross-modality temporal non-local (CMTNL) module is devised to simultaneously aggregate the global temporal features and cross-modality features with an attention mechanism. Through the complementary combination of the TSP and CMTNL modules, our CMLG-Net obtains a comprehensive and robust behavioral representation that contains both multi-scale (short- and long-term) and multimodal (RGB-D) behavioral information. Extensive experiments are conducted on the largest dataset, SCUT-DHGA, and a simulated practical dataset, SCUT-DHGA-br, to demonstrate the effectiveness of CMLG-Net in exploiting fine-grained behavioral features and complementary multimodal information. Finally, it achieves stat-of-the-art performance with the lowest ERR of 0.497% and 4.848% in two challenging evaluation protocols and shows significant superiority in robustness under practical scenes with unsatisfactory illumination and backgrounds. The code is available athttps://github.com/SCUT-BIP-Lab/CMLG-Net. Wenxiong Kang, Wenwei Song |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Ar3dHands: A Dataset and Baseline for Real-Time 3D Hand Pose Estimation from Binocular Distorted Images
Mengting Gan, Yihong Lin, Xingyan Liu, Wenwei Song, Wenxiong Kang |
ICIG (1) | 4 |
| 2023 | Random hand gesture authentication via efficient Temporal Segment Set Network
Yihong Lin, Wenwei Song, Wenxiong Kang |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Depthwise Temporal Non-Local Network for Faster and Better Dynamic Hand Gesture AuthenticationabstractDynamic hand gesture is an emerging and promising biometric trait. It contains both physiological and behavioral characteristics, which on the one hand can theoretically make authentication systems more accurate and more secure, and on the other hand can increase the difficulty of model design because it is essentially a fine-grained video understanding task. For authentication systems, equal error rate (EER) and real-time performance are two vital metrics. Current video understanding-based hand gesture authentication methods mainly focus on lowering the EER while neglecting to reduce the computational cost. In this paper, we propose a 2D CNN-based depthwise temporal non-local network (DwTNL-Net) that can take into account both EER and running efficiency. To enable the DwTNL-Net with spatiotemporal information processing capability, we design a temporal sharpening (TS) module and a DwTNL module for short- and long-term identity feature modeling, respectively. The TS module can assist the backbone in local behavioral characteristic understanding and can simultaneously remove redundant information and highlight behavioral cues while retaining sufficient physiological characteristics. In contrast, the DwTNL module focuses on summarizing global information and discovering stable patterns, which are finally used for local information enhancement. The complementary combination of our TS and DwTNL modules makes DwTNL-Net achieve substantial performance improvements. Extensive experiments on the SCUT-DHGA dataset and sufficient statistical analyses fully demonstrate the superiority and efficiency of our DwTNL-Net. The code is available at https://github.com/SCUT-BIP-Lab/DwTNL-Net. Wenwei Song, Wenxiong Kang |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | TDS-Net: Towards Fast Dynamic Random Hand Gesture Authentication via Temporal Difference Symbiotic Neural NetworkabstractHand gesture is a new emerging biometric trait containing both physiological and behavioral characteristics. With the popularity of various cameras, and the rich identity features and contactless authentication mode embedded in gestures themselves, vision-based hand gesture authentication has great potential value. However, current hand gesture authentication methods heavily rely on defined gestures and require identical enrollment and verification gestures, which limits the user-friendliness and efficiency of authentication. It is arguably true that authentication in a simpler and faster way, without the need to remember gestures, will be more approachable. Thus, a fast dynamic random hand gesture authentication method is introduced, in which users can perform a random improvised gesture in both the enrollment and verification stage. To better utilize the physiological and behavioral characteristics of hand gestures, an efficient network named Temporal Difference Symbiotic Neural Network (TDS-Net) equipped with our designed behavioral energy-based feature fusion module (BE-Fusion module) is proposed. Extensive experiments on the SCUT-DHGA dataset demonstrate that TDS-Net outperforms the recent state-of-the-art methods. Wenwei Song, Wenxiong Kang, Linpu Fang, Chang Liu 0060, Xingyan Liu |
IJCB | 1 |