Zhuming Wang

dblp:88/5802 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-2230-5716ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 VicKAM: Visual Conceptual Knowledge Guided Action Map for Weakly Supervised Group Activity Recognition
abstract
Most of existing weakly supervised GAR methods are typically bottom-up, automatically mining key areas by the attention mechanism. Due to the lack of a semantic connection to individual actions, some regions associated with these actions may be omitted, potentially impacting performance. In fact, a group activity is a combination of multiple individual actions, and the prototype of a specific action can be obtained from visual representations of individuals performing it, denoted as visual conceptual knowledge. In this paper, we propose a Visual Conceptual Knowledge Guided Action Map framework. It uses prototypes to produce individual action maps that indicate the likelihood of actions occurring at different locations. In some scenarios, the spatial distribution of actions shows strong regularity, which we compile as A-A Maps to enhance individual action maps. The action maps are integrated with action semantic representations for group activity recognition. Extensive experiments on two public benchmarks, the Volleyball and the NBA datasets, demonstrate the effectiveness of our proposed method, even in cases of limited training data.
Zhuming Wang, Yihao Zheng 0002, Jiarui Li 0002, Yaofei Wu, Yan Huang 0008, Zun Li 0001, Lifang Wu, Liang Wang 0001
ACM Multimedia1
2025 Statistical Information Assisted Interaction Reasoning for skeleton-only group activity recognition
Zhuming Wang, Zun Li 0001, Yihao Zheng 0002, Lifang Wu
Eng. Appl. Artif. Intell.1
2025 Multi-scale motion-based relational reasoning for group activity recognition
Yihao Zheng 0002, Zhuming Wang, Lifang Wu, Zun Li 0001, Ye Xiang
Eng. Appl. Artif. Intell.2
2025 Motion Attention-Guided Relational Reasoning for Weakly Supervised Group Activity Recognition
abstract
The existing attention-based label-free weakly supervised group activity recognition methods can automatically learn tokens related to the actors. And they have difficulties generating sufficiently diverse token embeddings. To address these issues, we automatically obtain the grayscale motion mask of all the moving objects based on the motion direction not the motion amplitude. A Motion-Guided Mask Generator module (MGMG) is proposed to estimate the attention region mask under the supervision of the grayscale motion mask. MGMG involves four parts. A correlation layer measures the relative displacement between two adjacent feature maps. A cosine attention mechanism is designed to reduce the module's sensitivity to feature amplitude changes. A mask generator is built to generate the attention region mask. And a specifically designed activation function is used to refine the attention region mask and to enhance its focus on actor motion regions. We also customize a normalized relative error loss function for MGMG module. This loss can address the value range mismatch problem for the estimated attention mask as well as the grayscale motion mask. Furthermore, a Motion Attention-Guided Relational Reasoning (MAGRR) framework is presented for the weakly supervised condition. It uses the MGMG module to estimate the attention region automatically, and a Spatial-temporal Aggregation Stack (SAS) module to activate the attention regions of the features at the spatial level, then transform them into multiple tokens, which are further captured by the attention mechanism for their temporal dependencies and interrelationships. MAGRR is experimented on the Collective Activity dataset and the Collective Activity Extension dataset, achieving state-of-the-art performance and competitive performance on the Volleyball and the NBA datasets.
Yihao Zheng 0002, Zhuming Wang, Lifang Wu, Liang Wang 0001, Chang Wen Chen
IEEE Trans. Image Process.2
2024 Face Anti-Spoofing via Interaction Learning with Face Image Quality Alignment
abstract
Face Anti-Spoofing is critical to secure face recognition systems from presentation attacks. Existing methods often suffer from performance degradation due to image quality issues, such as blurring, overexposure, or varied background, which cause distribution deviations of face images in the quality space, and hinder the learning of effective liveness features. In this paper, we propose a novel method that interactively co-reinforces the liveness and Face Quality representations for Face Anti-Spoofing (FQ-FAS). Specifically, to enhance the discrimination of face quality representation, FQ-FAS first designs a face quality learning module that naturally mitigates the interference from background. Subsequently, a quality-spoofing feature interaction module is devised to co-reinforce both liveness and face quality representations. Meanwhile, we propose a quality aware triplet loss to align the distribution of face images from two aspects: one is to pull the homogeneous face images with different quality together, while the other is to push the inhomogeneous samples with similar quality away in the feature space. In this way, FQ-FAS can learn reliable and discriminative representations for face anti-spoofing. Extensive intra-dataset and cross-dataset experiments clearly demonstrate that our method obtains better performance than previous state-of-the-art methods.
Yongluo Liu, Zun Li 0001, Zhuming Wang, Lifang Wu
FG4
2024 Knowledge Augmented Relation Inference for Group Activity Recognition
abstract
Group activity recognition is a challenging task because it involves diverse individual actions and complex relations. Most existing methods enhance individual representation by introducing relation inference using appearance features. Some methods utilize extra knowledge, such as action labels, to enhance relation inference and refine the individual representation, but the knowledge they explored is simple and insufficient. In this paper, we propose a novel idea of knowledge concretization and further develop a Knowledge Augmented Relation Inference framework (KARI) for group activity recognition. Specifically, we first concretize knowledge from training data, and then represent them as Class-Class co-occurrence Map (C-C Map) and Class-Position distribution Map (C-P Map). On top of them, KARI explores concretized knowledge to integrate visual and semantic representation in a unified architecture for group activity recognition. Experimental results on two public datasets show that the proposed framework performs favorably compared with state-of-the-art approaches.
Zhuming Wang, Zun Li 0001, Xianglong Lang, Yihao Zheng 0002, Lifang Wu, Liang Wang 0001, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.1
2023 Dual-stream correlation exploration for face anti-Spoofing
Yongluo Liu, Lifang Wu, Zun Li 0001, Zhuming Wang
Pattern Recognit. Lett.4
2023 Active Spatial Positions Based Hierarchical Relation Inference for Group Activity Recognition
abstract
Group activity recognition aims to recognize behaviors characterized by multiple individuals within a scene. Existing schemes rely on individual relation inference and usually take the individuals as tokens. Essentially they select the most relevant region of the group activity from the entire image while filtering out irrelevant background noises. However, these schemes require individual bounding box labeling in both training and testing stages. Since individuals have usually been presented at one scale, multi-scale individuals cannot be combined in an effective way. In this paper, we present a novel end-to-end hierarchical relation inference framework based on active spatial positions for group activity recognition. This framework is designed to locate active spatial positions and use them as visual tokens to infer the relations for token embeddings. It requires individual bounding box labeling only in the training stage while automatically eliminating the background after locating active spatial positions from the entire scene. The hierarchical relations can be naturally inferred based on the visual tokens at different scales, contributing to further performance improvement. Experimental results demonstrate that the proposed framework is competitive against existing schemes that require more laboring and computation to generate labels in both the training and testing stage.
Lifang Wu, Xianglong Lang, Ye Xiang, Chang Wen Chen, Zun Li 0001, Zhuming Wang
IEEE Trans. Circuits Syst. Video Technol.6
2023 Improving Face Anti-spoofing via Advanced Multi-perspective Feature Learning
abstract
Face anti-spoofing (FAS) plays a vital role in securing face recognition systems. Previous approaches usually learn spoofing features from a single perspective, in which only universal cues shared by all attack types are explored. However, such single-perspective-based approaches ignore the differences among various attacks and commonness between certain attacks and bona fides, thus tending to neglect some non-universal cues that contain strong discernibility against certain types. As a result, when dealing with multiple types of attacks, the above approaches may suffer from the uncomprehensive representation of bona fides and spoof faces. In this work, we propose a novel Advanced Multi-Perspective Feature Learning network (AMPFL), in which multiple perspectives are adopted to learn discriminative features, to improve the performance of FAS. Specifically, the proposed network first learns universal cues and several perspective-specific cues from multiple perspectives, then aggregates the above features and further enhances them to perform face anti-spoofing. In this way, AMPFL obtains features that are difficult to be captured by single-perspective-based methods and provides more comprehensive information on bona fides and spoof faces, thus achieving better performance for FAS. Experimental results show that our AMPFL achieves promising results in public databases, and it effectively solves the issues of single-perspective-based approaches.
Zhuming Wang, Yaowen Xu, Lifang Wu, Hu Han 0001, Zun Li 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Exploiting Non-uniform Inherent Cues to Improve Presentation Attack Detection
abstract
Face anti-spoofing plays a vital role in face recognition systems. The existed deep learning approaches have effectively improved the performance of presentation attack detection (PAD). However, they learn a uniform feature for different types of presentation attacks, which ignore the diversity of the inherent cues presented in different spoofing types. As a result, they can not effectively represent the intrinsic difference between different spoof faces and live faces, and the performance drops on the cross-domain databases. In this paper, we introduce the inherent cues of different spoofing types by non-uniform learning as complements to uniform features. Two lightweight sub-networks are designed to learn inherent motion patterns from photo attacks and the inherent texture cues from video attacks. Furthermore, an element-wise weighting fusion strategy is proposed to integrate the non-uniform inherent cues and uniform features. Extensive experiments on four public databases demonstrate that our approach outperforms the state-of-the-art methods and achieves a superior performance of 3.7% ACER in the cross-domain Protocol 4 of the Oulu-NPU database. Code is available at https://github.com/BJUT-VIP/Non-uniform-cues.
Yaowen Xu, Zhuming Wang, Hu Han 0001, Lifang Wu, Yongluo Liu
IJCB2
2021 Identity-constrained noise modeling with metric learning for face anti-spoofing
Yaowen Xu, Lifang Wu, Meng Jian, Wei-Shi Zheng 0001, Zhuming Wang
Neurocomputing6