VLDB 2026 Research / reviewers in the wild / expert
Cong Wu 0006
dblp:30/10768-6
· DBLP profile ↗
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0001-9555-9445ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interactive image-to-video transfer learning
Cong Wu 0006, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler |
Neural Networks | 1 |
| 2025 | Adaptive Hyper-Graph Convolution Network for Skeleton-Based Human Action Recognition with Virtual ConnectionsabstractThe shared topology of human skeletons motivated the recent investigation of graph convolutional network (GCN) solutions for action recognition. However, most of the existing GCNs rely on the binary connection of two neighboring vertices (joints) formed by an edge (bone), overlooking the potential of constructing multi-vertex convolution structures. Although some studies have attempted to utilize hyper-graphs to represent the topology, they rely on a fixed construction strategy, which limits their adaptivity in uncovering the intricate latent relationships within the action. In this paper, we address this oversight and explore the merits of an adaptive hyper-graph convolutional network (Hyper-GCN) to achieve the aggregation of rich semantic information conveyed by skeleton vertices. In particular, our Hyper-GCN adaptively optimises the hyper-graphs during training, revealing the action-driven multi-vertex relations. Besides, virtual connections are often designed to support efficient feature aggregation, implicitly extending the spectrum of dependencies within the skeleton. By injecting virtual connections into hyper-graphs, the semantic clues of diverse action categories can be highlighted. The results of experiments conducted on the NTU-60, NTU-120, and NW-UCLA datasets demonstrate the merits of our Hyper-GCN, compared to the state-of-the-art methods. The code is available at https://github.com/6UOOON9/Hyper-GCN. Youwei Zhou, Tianyang Xu 0001, Cong Wu 0006, Xiaojun Wu 0001, Josef Kittler |
ICCV | 3 |
| 2025 | Dual attention focus network for few-shot skeleton-based action recognition
Chongben Tao, Cong Wu 0006, Tianyang Xu 0001, Xizhao Luo, Zufeng Zhang, Sai Xu |
Knowl. Based Syst. | 4 |
| 2025 | Adaptive pooling with dual-stage fusion for skeleton-based action recognition
Cong Wu 0006, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
Neural Networks | 1 |
| 2025 | I Know How You Move: Explicit Motion Estimation for Human Action RecognitionabstractEnabled by hierarchical convolutions and nonlinear mappings, recent action recognition studies have continuously boosted performance with spatiotemporal modelling. In general, motion clues are essential in video-oriented tasks, while existing approaches aggregate the spatial and temporal signatures via specially designed modules in the middle or output stages. To highlight the privilege provided by temporal motions, in this paper, we propose a simple but effectiveMOTion Estimator(MOTE) to generate the motion patterns from every single frame, avoiding complex dense-frame input. In particular, MOTE follows an encoder-decoder structure, which takes the short-term motion features generated by the pretrained dense-frame network as the learning target. The spatial information of a single frame is utilized to estimate the instantaneous motion appearance. It can support the expression of vulnerable regions, such as the ‘hand’ in ‘waving hands’, which would otherwise be suppressed in the feature maps as the ‘hand’ suffers from motion blur. The training process of MOTE is independent of the action recognition system. Therefore, the trained MOTE can be transplanted to the input-end of existing action recognition methods to provide instantaneous motion estimation as feature enhancement according to practical requirements. Our experiments performed on Something-Something V1, V2, Kinetics-400, and Diving48 verify the effectiveness of the proposed method. Xiaojun Wu 0001, Hui Li 0037, Tianyang Xu 0001, Cong Wu 0006 |
IEEE Trans. Multim. | 5 |
| 2024 | SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action RecognitionabstractContrastive learning has achieved great success in skeleton-based action recognition. However, most existing approaches encode the skeleton sequences as entangled spatiotemporal representations and confine the contrasts to the same level of representation. Instead, this paper introduces a novel contrastive learning framework, namely Spatiotemporal Clues Disentanglement Network (SCD-Net). Specifically, we integrate the decoupling module with a feature extractor to derive explicit clues from spatial and temporal domains respectively. As for the training of SCD-Net, with a constructed global anchor, we encourage the interaction between the anchor and extracted clues. Further, we propose a new masking strategy with structural constraints to strengthen the contextual associations, leveraging the latest development from masked image modelling into the proposed SCD-Net. We conduct extensive evaluations on the NTU-RGB+D (60&120) and PKU-MMD (I&II) datasets, covering various downstream tasks such as action recognition, action retrieval, transfer learning, and semi-supervised learning. The experimental results demonstrate the effectiveness of our method, which outperforms the existing state-of-the-art (SOTA) approaches significantly. Our code and supplementary material can be found at https://github.com/cong-wu/SCD-Net. Cong Wu 0006, Xiaojun Wu 0001, Josef Kittler, Tianyang Xu 0001, Muhammad Awais 0001, Zhenhua Feng 0001 |
AAAI | 1 |
| 2024 | Efficient Few-Shot Action Recognition via Multi-level Post-reasoning
Cong Wu 0006, Xiaojun Wu 0001, Linze Li 0002, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler |
ECCV (3) | 1 |
| 2024 | Spatio-Temporal Domain-Aware Network for Skeleton-Based Action Representation Learning
Jiannan Hu, Cong Wu 0006, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
ICPR (29) | 2 |
| 2024 | Scene adaptive mechanism for action recognition
Cong Wu 0006, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
Comput. Vis. Image Underst. | 1 |
| 2024 | Motion Complement and Temporal Multifocusing for Skeleton-Based Action RecognitionabstractModeling sequences with spatial-temporal graph convolutional networks has become a mainstream paradigm in skeleton-based action recognition. However, many existing methods adopt redundant or cluttered structures to mine the key action features, thus making it difficult to achieve a balanced or leading performance in accuracy and efficiency. In this paper, we propose a novel framework, referred to as Motion Complement and Temporal Multifocusing Network (MCTM-Net), to capture the relationships within skeleton sequences by means of an efficient decomposition of the spatiotemporal graph model. Specifically, for spatial modeling, we introduce a motion-related relational descriptor that extends the channel dimension so as to enhance the modeling of motion salient regions as a complement to the conventional physical adjacency relationships. An improved parameterized physical relationship model is also proposed to better fit the data characteristics. As for temporal modeling, we propose an efficient multi-focus temporal information acquisition strategy that aggregates the information from multiple temporal spans and adjacent regions. We conduct extensive experiments on multiple representative datasets, including NTU-RGB+D (60&120), Northwestern-UCLA, and UWA3D Multiview Activity II, to validate our innovations. The experimental results show the effectiveness of our method. The code will be available athttps://github.com/cong-wu/MCMT-Net. Cong Wu 0006, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Graph2Net: Perceptually-Enriched Graph Learning for Skeleton-Based Action RecognitionabstractSkeleton representation has attracted a great deal of attention recently as an extremely robust feature for human action recognition. However, its non-Euclidean structural characteristics raise new challenges for conventional solutions. Recent studies have shown that there is a native superiority in modeling spatiotemporal skeleton information with a Graph Convolutional Network (GCN). Nevertheless, the skeleton graph modeling normally focuses on the physical adjacency of the elements of the human skeleton sequence, which contrasts with the requirement to provide a perceptually meaningful representation. To address this problem, in this paper, we propose a perceptually-enriched graph learning method by introducing innovative features to spatial and temporal skeleton graph modeling. For the spatial information modeling, we incorporate a Local-Global Graph Convolutional Network (LG-GCN) that builds a multifaceted spatial perceptual representation. This helps to overcome the limitations caused by over-reliance on the spatial adjacency relationships in the skeleton. For temporal modeling, we present a Region-Aware Graph Convolutional Network (RA-GCN), which directly embeds the regional relationships conveyed by a skeleton sequence into a temporal graph model. This innovation mitigates the deficiency of the original skeleton graph models. In addition, we strengthened the ability of the proposed channel modeling methods to extract multi-scale representations. These innovations result in a lightweight graph convolutional model, referred to as Graph2Net, that simultaneously extends the spatial and temporal perceptual fields, and thus enhances the capacity of the graph model to represent skeleton sequences. We conduct extensive experiments on NTU-RGB+D 60&120, Northwestern-UCLA, and Kinetics-400 datasets to show that our results surpass the performance of several mainstream methods while limiting the model complexity and computational overhead. Cong Wu 0006, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Dynamic information enhancement for video classification
Rongchang Li 0001, Xiaojun Wu 0001, Cong Wu 0006, Tianyang Xu 0001, Josef Kittler |
Image Vis. Comput. | 3 |