VLDB 2026 Research / reviewers in the wild / expert
Xiao Wang 0072
dblp:49/67-72
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2024
0000-0003-2105-7818ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Few-Shot Action Recognition via Multi-View Representation LearningabstractFew-shot action recognition aims to recognize novel action classes with limited labeled samples and has recently received increasing attention. The core objective of few-shot action recognition is to enhance the discriminability of feature representations. In this paper, we propose a novel multi-view representation learning network (MRLN) to model intra-video and inter-video relations for few-shot action recognition. Specifically, we first propose a spatial-aware aggregation refinement module (SARM), which mainly consists of a spatial-aware aggregation sub-module and a spatial-aware refinement sub-module to explore the spatial context of samples at the frame level. Then, we design a temporal-channel enhancement module (TCEM), which can capture the temporal-aware and channel-aware features of samples with the elaborately designed temporal-aware enhancement sub-module and channel-aware enhancement sub-module. Third, we introduce a cross-video relation module (CVRM), which can explore the relations across videos by utilizing the self-attention mechanism. Moreover, we design a prototype-centered mean absolute error loss to improve the feature learning capability of the proposed MRLN. Extensive experiments on four prevalent few-shot action recognition benchmarks show that the proposed MRLN can significantly outperform a variety of state-of-the-art few-shot action recognition methods. Especially, on the 5-way 1-shot setting, our MRLN respectively achieves 75.7%, 86.9%, 65.5% and 45.9% on the Kinetics, UCF101, HMDB51 and SSv2 datasets. Xiao Wang 0072, Yang Lu 0009, Wanchuan Yu, Yanwei Pang, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Cross-Modal Contrastive Learning Network for Few-Shot Action RecognitionabstractFew-shot action recognition aims to recognize new unseen categories with only a few labeled samples of each class. However, it still suffers from the limitation of inadequate data, which easily leads to the overfitting and low-generalization problems. Therefore, we propose a cross-modal contrastive learning network (CCLN), consisting of an adversarial branch and a contrastive branch, to perform effective few-shot action recognition. In the adversarial branch, we elaborately design a prototypical generative adversarial network (PGAN) to obtain synthesized samples for increasing training samples, which can mitigate the data scarcity problem and thereby alleviate the overfitting problem. When the training samples are limited, the obtained visual features are usually suboptimal for video understanding as they lack discriminative information. To address this issue, in the contrastive branch, we propose a cross-modal contrastive learning module (CCLM) to obtain discriminative feature representations of samples with the help of semantic information, which can enable the network to enhance the feature learning ability at the class-level. Moreover, since videos contain crucial sequences and ordering information, thus we introduce a spatial-temporal enhancement module (SEM) to model the spatial context within video frames and the temporal context across video frames. The experimental results show that the proposed CCLN outperforms the state-of-the-art few-shot action recognition methods on four challenging benchmarks, including Kinetics, UCF101, HMDB51 and SSv2. Xiao Wang 0072, Yan Yan 0001, Hai-Miao Hu, Bo Li 0006, Hanzi Wang |
IEEE Trans. Image Process. | 1 |
| 2023 | Task-Aware Dual-Representation Network for Few-Shot Action RecognitionabstractFew-shot action recognition has attracted increasing attention in recent years, but it remains challenging due to the intrinsic difficulty in learning transferable knowledge to generalize to novel classes by using a few labeled samples. Although some successful progress has been made, most few-shot action recognition methods commonly focus on the global characteristics of samples while ignoring the local characteristics of samples, which results in the weak generalization ability of the model. In this paper, we propose a task-aware dual-representation network (TADRNet) for few-shot action recognition, which learns how to adapt video representations to novel tasks in a meta-learning manner. It mainly includes a global relational graph subnetwork (GRG) and a fine-grained local representation subnetwork (FLR). Our method simultaneously considers both global and local characteristics of samples for few-shot action recognition. From a global perspective, we propose GRG to explore the relations across support-query sample pairs by using the relational graph neural network. To facilitate the few-shot visual learning, we propose a novel hybrid semantic attention module (HSA) for enhancing the discriminability of support and query features. From a local perspective, we utilize FLR to fully exploit the local characteristics of samples, which can improve the classification results obtained by GRG and thus guarantee high classification accuracy. Extensive experiments on four challenging benchmarks show that the proposed TADRNet significantly outperforms a variety of state-of-the-art few-shot action recognition methods. Xiao Wang 0072, Weirong Ye, Zhongang Qi, Guangge Wang, Ying Shan, Xiaohu Qie, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Embedding Adaptation Network with Transformer for Few-Shot Action Recognition
Rongrong Jin, Xiao Wang 0072, Guangge Wang, Yang Lu 0009, Hai-Miao Hu, Hanzi Wang |
ACML | 2 |
| 2022 | Egnet: A Novel Edge Guided Network for Instance SegmentationabstractEdge information plays a significant role in instance segmentation. However, many instance segmentation methods directly perform pixel-wise classification via fully convolutional networks, which may ignore object edges. In this paper, we propose a novel Edge Guided Network (EGNet), which exploits edge information to improve the mask accuracy, for instance segmentation. Specifically, we propose an edge branch to extract edge information. Then, we use edge information as guidance and fuse it with mask features, in order to enrich the mask features. Furthermore, we propose a Spatial Attention (SA) module and add it to the backbone of our EGNet, enabling the network to focus more on foreground objects. In addition, we incorporate a Semantic Enhancement (SE) module into the edge branch, aiming to obtain additional global context information. Experimental results on the COCO 2017 dataset show the effectiveness of the proposed EGNet. Kaiwen Du, Xiao Wang 0072, Yan Yan 0009, Yang Lu 0009, Hanzi Wang |
ICIP | 2 |
| 2022 | MDNet: Motion Distinction Network for Effective Action RecognitionabstractMotion information is critical for action recognition. Most existing methods perform motion enhancement only from the channel or spatiotemporal dimension. They may fail to achieve fine-grained motion modeling, leading to suboptimal performance. In this paper, we propose a novel motion distinction network (MDNet) to address this challenging problem. Specifically, we first propose a channel-wise motion enhancement (CME) module, which aims to emphasize motion-related channels by leveraging a channel-wise gating mechanism. Then, we propose a cascaded spatiotemporal enhancement (CSTE) module to enhance motion features along the spatiotemporal dimension. Moreover, we design a multi-attention fusion strategy to further refine the enhanced motion features in a moderate manner, enabling the network to focus on discriminative motion regions. The proposed two modules and fusion strategy are complementary for fine-grained motion enhancement. Extensive experiments are conducted on the challenging Something-Something V1 and Kinetics-400 datasets to show the effectiveness of the proposed MDNet. Rongrong Ji, Weirong Ye, Xiao Wang 0072, Yan Yan 0001, Hanzi Wang |
ICIP | 3 |
| 2022 | Visual Tempo Contrastive Learning for Few-Shot Action RecognitionabstractFew-shot action recognition aims to learn novel action classes with only a few annotated samples. This is a challenging problem because motion modeling is difficult, especially when a few training samples are available. Visual tempo, which is an essential variation factor of video semantics, characterizes the dynamic motion information in action videos. In this work, we propose a visual tempo contrastive learning framework (VTCL) to tackle the few-shot action recognition problem. Specifically, we propose a visual tempo encoding (VTE) module for visual tempo learning. The VTE module samples the same action instance at different frame sampling rates to obtain different visual tempo encoding vectors, which jointly form the features of each instance. To enhance the discriminability of visual tempo encoding vectors, we propose a visual tempo contrastive encoding (VTCE) loss to promote intra-class compactness and inter-class difference of visual tempo encoding vectors. Extensive experiments demonstrate that the proposed VTCL achieves promising results among the competing state-of-the-art methods on three few-shot action recognition datasets, including Kinetics, UCF101 and HMDB51. Guangge Wang, Weirong Ye, Xiao Wang 0072, Rongrong Ji, Hanzi Wang |
ICIP | 3 |
| 2022 | FastVOD-Net: A Real-Time and High-Accuracy Video Object DetectorabstractVideo object detection is a tough task due to the severe appearance degradation caused by rapid motion, sudden occlusion or rare poses. The great challenge facing video object detection is the simultaneous requirements on both accuracy and speed because the pursuit of one aspect usually causes significant expense to the other. Most existing methods mainly focus on improving detection accuracy with little attention to computationally efficient solutions, and thus they are impractical for many real-world applications. This motivates us to develop a real-time and high-accuracy video object detection method. In this paper, we propose a novel video object detector, called FastVOD-Net, which can yield highly accurate detection results at real-time speed. Specifically, we first develop a temporally-cascaded deformable alignment (TCDA) module to model the object displacements induced by video motion. Then, we introduce another two modules, namely spatially-refined temporal aggregation (SRTA) and attention-guided semantic distillation (AGSD), to improve the appearance feature of the currently processed frame and enhance the semantic representation of non-keyframes, respectively. For keyframe scheduling, we design an adaptive keyframe selection scheduler (AKSS) to adjust the keyframe interval online, making the keyframe usage more rational. On one hand, the characteristics of our FastVOD-Net enable it to sparsely perform expensive feature extraction, which significantly reduces the computational cost and thus guarantees real-time speed. On the other hand, the collaboration of the above tightly-coupled modules and adaptive keyframe scheduler makes FastVOD-Net fully exploit inter-frame temporal dependencies and thus guarantees high accuracy. Experiments on the ImageNet VID dataset show that our FastVOD-Net achieves 79.3% mAP at 29.6 fps or 81.2% mAP at 23.0 fps on an Nvidia RTX 2080 Ti GPU, which is the state-of-the-art performance in real time. Qiang Qi, Xiao Wang 0072, Tianxiang Hou, Yan Yan 0001, Hanzi Wang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Temporal Relation based Attentive Prototype Network for Few-shot Action RecognitionabstractFew-shot action recognition aims at recognizing novel action classes with only a small number of labeled video samples. We propose a temporal relation based attentive prototype network (TRAPN) for few-shot action recognition. Concretely, we tackle this challenging task from three aspects. Firstly, we propose a spatio-temporal motion enhancement (STME) module to highlight object motions in videos. The STME module utilizes cues from content displacements in videos to enhance the features in the motion-related regions. Secondly, we learn the core common action transformations by our temporal relation (TR) module, which captures the temporal relations at short-term and long-term time scales. The learned temporal relations are encoded into descriptors to constitute sample-level features. The abstract action transformations are described by multiple groups of temporal relation descriptors. Thirdly, a vanilla prototype for the support class (e.g., the mean of the support class) cannot fit well for different query samples. We generate an attentive prototype constructed from temporal relation descriptors of support samples, which gives more weight to discriminative samples. We evaluate our TRAPN on Kinetics, UCF101 and HMDB51 real-world few-shot datasets. Results show that our network achieves the state-of-the-art performance. Guangge Wang, Haihui Ye, Xiao Wang 0072, Weirong Ye, Hanzi Wang |
ACML | 3 |
| 2021 | Semantic-Guided Relation Propagation Network for Few-shot Action RecognitionabstractFew-shot action recognition has drawn growing attention as it can recognize novel action classes by using only a few labeled samples. In this paper, we propose a novel semantic-guided relation propagation network (SRPN), which leverages semantic information together with visual information for few-shot action recognition. Different from most previous works that neglect semantic information in the labeled data, our SRPN directly utilizes the semantic label as an additional supervisory signal to improve the generalization ability of the network. Besides, we treat the relation of each visual-semantic pair as a relational node, and we use a graph convolutional network to model and propagate such sample relations across visual-semantic pairs, including both intra-class commonality and inter-class uniqueness, to guide the relation propagation in the graph. However, since videos contain crucial sequences and ordering information, we propose a novel spatial-temporal difference module, which can facilitate the network to enhance the visual feature learning ability at both feature level and granular level for videos. Extensive experiments conducted on several challenging benchmarks demonstrate that our SRPN outperforms several state-of-the-art methods with a significant margin. Xiao Wang 0072, Weirong Ye, Zhongang Qi, Guangge Wang, Ying Shan, Hanzi Wang |
ACM Multimedia | 1 |