VLDB 2026 Research / reviewers in the wild / expert
Zhuming Wang
dblp:88/5802
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-2230-5716ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VicKAM: Visual Conceptual Knowledge Guided Action Map for Weakly Supervised Group Activity RecognitionabstractMost of existing weakly supervised GAR methods are typically bottom-up, automatically mining key areas by the attention mechanism. Due to the lack of a semantic connection to individual actions, some regions associated with these actions may be omitted, potentially impacting performance. In fact, a group activity is a combination of multiple individual actions, and the prototype of a specific action can be obtained from visual representations of individuals performing it, denoted as visual conceptual knowledge. In this paper, we propose a Visual Conceptual Knowledge Guided Action Map framework. It uses prototypes to produce individual action maps that indicate the likelihood of actions occurring at different locations. In some scenarios, the spatial distribution of actions shows strong regularity, which we compile as A-A Maps to enhance individual action maps. The action maps are integrated with action semantic representations for group activity recognition. Extensive experiments on two public benchmarks, the Volleyball and the NBA datasets, demonstrate the effectiveness of our proposed method, even in cases of limited training data. Zhuming Wang, Yihao Zheng 0002, Jiarui Li 0002, Yaofei Wu, Yan Huang 0008, Zun Li 0001, Lifang Wu, Liang Wang 0001 |
ACM Multimedia | 1 |
| 2025 | Statistical Information Assisted Interaction Reasoning for skeleton-only group activity recognition
Zhuming Wang, Zun Li 0001, Yihao Zheng 0002, Lifang Wu |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Multi-scale motion-based relational reasoning for group activity recognition
Yihao Zheng 0002, Zhuming Wang, Lifang Wu, Zun Li 0001, Ye Xiang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Motion Attention-Guided Relational Reasoning for Weakly Supervised Group Activity RecognitionabstractThe existing attention-based label-free weakly supervised group activity recognition methods can automatically learn tokens related to the actors. And they have difficulties generating sufficiently diverse token embeddings. To address these issues, we automatically obtain the grayscale motion mask of all the moving objects based on the motion direction not the motion amplitude. A Motion-Guided Mask Generator module (MGMG) is proposed to estimate the attention region mask under the supervision of the grayscale motion mask. MGMG involves four parts. A correlation layer measures the relative displacement between two adjacent feature maps. A cosine attention mechanism is designed to reduce the module's sensitivity to feature amplitude changes. A mask generator is built to generate the attention region mask. And a specifically designed activation function is used to refine the attention region mask and to enhance its focus on actor motion regions. We also customize a normalized relative error loss function for MGMG module. This loss can address the value range mismatch problem for the estimated attention mask as well as the grayscale motion mask. Furthermore, a Motion Attention-Guided Relational Reasoning (MAGRR) framework is presented for the weakly supervised condition. It uses the MGMG module to estimate the attention region automatically, and a Spatial-temporal Aggregation Stack (SAS) module to activate the attention regions of the features at the spatial level, then transform them into multiple tokens, which are further captured by the attention mechanism for their temporal dependencies and interrelationships. MAGRR is experimented on the Collective Activity dataset and the Collective Activity Extension dataset, achieving state-of-the-art performance and competitive performance on the Volleyball and the NBA datasets. Yihao Zheng 0002, Zhuming Wang, Lifang Wu, Liang Wang 0001, Chang Wen Chen |
IEEE Trans. Image Process. | 2 |
| 2024 | Face Anti-Spoofing via Interaction Learning with Face Image Quality AlignmentabstractFace Anti-Spoofing is critical to secure face recognition systems from presentation attacks. Existing methods often suffer from performance degradation due to image quality issues, such as blurring, overexposure, or varied background, which cause distribution deviations of face images in the quality space, and hinder the learning of effective liveness features. In this paper, we propose a novel method that interactively co-reinforces the liveness and Face Quality representations for Face Anti-Spoofing (FQ-FAS). Specifically, to enhance the discrimination of face quality representation, FQ-FAS first designs a face quality learning module that naturally mitigates the interference from background. Subsequently, a quality-spoofing feature interaction module is devised to co-reinforce both liveness and face quality representations. Meanwhile, we propose a quality aware triplet loss to align the distribution of face images from two aspects: one is to pull the homogeneous face images with different quality together, while the other is to push the inhomogeneous samples with similar quality away in the feature space. In this way, FQ-FAS can learn reliable and discriminative representations for face anti-spoofing. Extensive intra-dataset and cross-dataset experiments clearly demonstrate that our method obtains better performance than previous state-of-the-art methods. Yongluo Liu, Zun Li 0001, Zhuming Wang, Lifang Wu |
FG | 4 |
| 2024 | Knowledge Augmented Relation Inference for Group Activity RecognitionabstractGroup activity recognition is a challenging task because it involves diverse individual actions and complex relations. Most existing methods enhance individual representation by introducing relation inference using appearance features. Some methods utilize extra knowledge, such as action labels, to enhance relation inference and refine the individual representation, but the knowledge they explored is simple and insufficient. In this paper, we propose a novel idea of knowledge concretization and further develop a Knowledge Augmented Relation Inference framework (KARI) for group activity recognition. Specifically, we first concretize knowledge from training data, and then represent them as Class-Class co-occurrence Map (C-C Map) and Class-Position distribution Map (C-P Map). On top of them, KARI explores concretized knowledge to integrate visual and semantic representation in a unified architecture for group activity recognition. Experimental results on two public datasets show that the proposed framework performs favorably compared with state-of-the-art approaches. Zhuming Wang, Zun Li 0001, Xianglong Lang, Yihao Zheng 0002, Lifang Wu, Liang Wang 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Dual-stream correlation exploration for face anti-Spoofing
Yongluo Liu, Lifang Wu, Zun Li 0001, Zhuming Wang |
Pattern Recognit. Lett. | 4 |
| 2023 | Active Spatial Positions Based Hierarchical Relation Inference for Group Activity RecognitionabstractGroup activity recognition aims to recognize behaviors characterized by multiple individuals within a scene. Existing schemes rely on individual relation inference and usually take the individuals as tokens. Essentially they select the most relevant region of the group activity from the entire image while filtering out irrelevant background noises. However, these schemes require individual bounding box labeling in both training and testing stages. Since individuals have usually been presented at one scale, multi-scale individuals cannot be combined in an effective way. In this paper, we present a novel end-to-end hierarchical relation inference framework based on active spatial positions for group activity recognition. This framework is designed to locate active spatial positions and use them as visual tokens to infer the relations for token embeddings. It requires individual bounding box labeling only in the training stage while automatically eliminating the background after locating active spatial positions from the entire scene. The hierarchical relations can be naturally inferred based on the visual tokens at different scales, contributing to further performance improvement. Experimental results demonstrate that the proposed framework is competitive against existing schemes that require more laboring and computation to generate labels in both the training and testing stage. Lifang Wu, Xianglong Lang, Ye Xiang, Chang Wen Chen, Zun Li 0001, Zhuming Wang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Improving Face Anti-spoofing via Advanced Multi-perspective Feature LearningabstractFace anti-spoofing (FAS) plays a vital role in securing face recognition systems. Previous approaches usually learn spoofing features from a single perspective, in which only universal cues shared by all attack types are explored. However, such single-perspective-based approaches ignore the differences among various attacks and commonness between certain attacks and bona fides, thus tending to neglect some non-universal cues that contain strong discernibility against certain types. As a result, when dealing with multiple types of attacks, the above approaches may suffer from the uncomprehensive representation of bona fides and spoof faces. In this work, we propose a novel Advanced Multi-Perspective Feature Learning network (AMPFL), in which multiple perspectives are adopted to learn discriminative features, to improve the performance of FAS. Specifically, the proposed network first learns universal cues and several perspective-specific cues from multiple perspectives, then aggregates the above features and further enhances them to perform face anti-spoofing. In this way, AMPFL obtains features that are difficult to be captured by single-perspective-based methods and provides more comprehensive information on bona fides and spoof faces, thus achieving better performance for FAS. Experimental results show that our AMPFL achieves promising results in public databases, and it effectively solves the issues of single-perspective-based approaches. Zhuming Wang, Yaowen Xu, Lifang Wu, Hu Han 0001, Zun Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Exploiting Non-uniform Inherent Cues to Improve Presentation Attack DetectionabstractFace anti-spoofing plays a vital role in face recognition systems. The existed deep learning approaches have effectively improved the performance of presentation attack detection (PAD). However, they learn a uniform feature for different types of presentation attacks, which ignore the diversity of the inherent cues presented in different spoofing types. As a result, they can not effectively represent the intrinsic difference between different spoof faces and live faces, and the performance drops on the cross-domain databases. In this paper, we introduce the inherent cues of different spoofing types by non-uniform learning as complements to uniform features. Two lightweight sub-networks are designed to learn inherent motion patterns from photo attacks and the inherent texture cues from video attacks. Furthermore, an element-wise weighting fusion strategy is proposed to integrate the non-uniform inherent cues and uniform features. Extensive experiments on four public databases demonstrate that our approach outperforms the state-of-the-art methods and achieves a superior performance of 3.7% ACER in the cross-domain Protocol 4 of the Oulu-NPU database. Code is available at https://github.com/BJUT-VIP/Non-uniform-cues. Yaowen Xu, Zhuming Wang, Hu Han 0001, Lifang Wu, Yongluo Liu |
IJCB | 2 |
| 2021 | Identity-constrained noise modeling with metric learning for face anti-spoofing
Yaowen Xu, Lifang Wu, Meng Jian, Wei-Shi Zheng 0001, Zhuming Wang |
Neurocomputing | 6 |