VLDB 2026 Research / reviewers in the wild / expert
Chao Tang 0002
dblp:65/7773-2
· DBLP profile ↗
10ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-8934-9537ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic-static feature fusion and multi-level interaction reasoning for group activity recognition
Huajun Sun, Chao Tang 0002, Huosheng Hu, Wenjian Wang 0001, Fang Ren 0002, Anyang Tong |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Attention mechanism based multimodal feature fusion network for human action recognition
Chao Tang 0002, Huosheng Hu, Wenjian Wang 0001, Shuo Qiao, Anyang Tong |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | EMPC: Efficient multi-view parallel co-learning for semi-supervised action recognition
Anyang Tong, Chao Tang 0002, Wenjian Wang 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Skeleton-based human action recognition by fusing attention based three-stream convolutional neural network and SVM
Fang Ren 0002, Chao Tang 0002, Anyang Tong, Wenjian Wang 0001 |
Multim. Tools Appl. | 2 |
| 2024 | Public-Private Attributes-Based Variational Adversarial Network for Audio-Visual Cross-Modal MatchingabstractExisting audio-visual cross-modal matching methods focus on mitigating cross-modal heterogeneity but ignore the impact of intra-class discrepancy of the same identity in different scenarios, which might greatly limit the matching performance. To simultaneously handle both problems of intra-class discrepancy and cross-modal heterogeneity, we propose a novel public-private attributes-based variational adversarial network (P2VANet), which captures the consistency within and between classes, for audio-visual cross-modal matching. In particular,P2VANet first uses a variational auto-encoder, which captures the inherent global information in diverse scenarios from the hidden variable through reconstruction, to reduce the intra-class discrepancy. Then it integrates a public attributes guidance module to capture the consistency of audio and visual by supervision of the common high-level semantic information to mitigate cross-modal heterogeneity. In addition,P2VANet designs private attributes embedding module to enhance the discriminative features inherent in each class to decrease inter-class similarity. Extensive experiments on audio-visual cross-modal matching demonstrate the effectiveness of the proposed approach compared with the state-of-the-art methods. Aihua Zheng, Jiaxiang Wang 0001, Chao Tang 0002, Chenglong Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Semi-Supervised Action Recognition From Temporal Augmentation Using Curriculum LearningabstractSemi-supervised learning for video action recognition is a very challenging research area. Existing state-of-the-art methods perform data augmentation on the temporality of actions, which are combined with the mainstream consistency-based semi-supervised learning framework FixMatch for action recognition. However, these approaches have the following limitations: (1) data augmentation based on video clips lacks coarse-grained and fine-grained representations of actions in temporal sequences, and the models have difficulty understanding synonymous representations of actions in different motion phases. (2) Pseudo labeling selection based on the constant thresholds lacks a “make-up curriculum” for difficult actions, that results in the low utilization of unlabeled data corresponding to difficult actions. To address the above shortcomings, we propose a semi-supervised action recognition via the temporal augmentation using curriculum learning (TACL) algorithm. Compared to previous works, TACL explores different representations of the same semantics of actions in temporal sequences for video and uses the idea of curriculum learning (CL) to reduce the difficulty of the model training process. First, for different action expressions with the same semantics, we designed the temporal action augmentation (TAA) for videos to obtain coarse-grained and fine-grained action expressions based on constant-velocity and hetero-velocity methods, respectively. Second, we construct a temporal signal to constrain the model such that fine-grained action expressions containing different movement phases have the same prediction results, and achieve action consistency learning (ACL) by combining the label and pseudo-label signals. Finally, we propose action curriculum pseudo labeling (ACPL), a loosely and strictly parallel dynamic threshold evaluation algorithm for selecting and labeling unlabeled data. We evaluate TACL on three standard public datasets: UCF101, HMDB51, and Kinetics. The combined experiments show that TACL significantly improves the accuracy of models trained on a small amount of labeled data and better evaluates the learning effects for different actions. Anyang Tong, Chao Tang 0002, Wenjian Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Image Segmentation Based on Local Chan-Vese Model Combined with Fractional Order Derivative
Le Zou, Liang-Tu Song, Xiaofeng Wang 0009, Chao Tang 0002, Chen Zhang 0039 |
ICIC (1) | 5 |
| 2019 | Univariate Thiele Type Continued Fractions Rational Interpolation with Parameters
Le Zou, Liang-Tu Song, Xiaofeng Wang 0009, Qian-Jing Huang, Chao Tang 0002, Chen Zhang 0039 |
ICIC (3) | 6 |
| 2018 | Image Segmentation Based on Local Chan Vese Model by Employing Cosine Fitting Energy
Le Zou, Liang-Tu Song, Xiaofeng Wang 0009, Qiong Zhou, Chao Tang 0002, Chen Zhang 0039 |
PRCV (1) | 6 |
| 2017 | Hybrid level set method based on image diffusion
Xiaofeng Wang 0009, Le Zou, Li-Xiang Xu, Chao Tang 0002 |
Neurocomputing | 5 |