VLDB 2026 Research / reviewers in the wild / expert
Anyang Tong
dblp:313/0316
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-0960-1497ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic-static feature fusion and multi-level interaction reasoning for group activity recognition
Huajun Sun, Chao Tang 0002, Huosheng Hu, Wenjian Wang 0001, Fang Ren 0002, Anyang Tong |
J. Vis. Commun. Image Represent. | 6 |
| 2026 | Toward Trustworthy Dynamic Facial Expression Recognition via Information Bottleneck Modeling
Feng-Qi Cui, Anyang Tong, Jinyang Huang, Jie Zhang 0073, Meng Li 0006, Linsheng Huang, Dan Guo 0001, Meng Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | Boosting Semi-Supervised Learning With Entropy-Guided Adaptive Reward MaximizationabstractExisting semi-supervised learning (SSL) methods rely predominantly on pseudo-labeling and consistency regularization to leverage unlabeled data, demonstrating significant performance improvements. However, we pinpoint that these methods suffer from a confidence-for-weighting issue, overvaluing high-confidence pseudo-labels while undervaluing low-confidence yet informative samples that are critical for robust generalization. In this paper, we introduce EntropyMatch, an entropy-driven SSL framework that redefines sample importance through prediction entropy rather than confidence alone. EntropyMatch employs a bidirectional weighting strategy: upward exploitation exploits reliable hard samples to refine decision boundaries while downward exploration cautiously explores uncertain ones to reduce noise. Additionally, EntropyMatch features an adaptive training mechanism that aligns with model maturity, shifting focus from safe exploration to strategic exploitation as training progresses. Experiments on eight benchmarks across various SSL tasks-spanning image classification, facial expression recognition, and human action recognition-validate EntropyMatch's robustness and effectiveness. It consistently achieves state-of-the-art results, notably matching state-of-the-art LION's performance on RAF-DB with just half the labeled data, demonstrating superior data efficiency and generalization. Anyang Tong, Zenglin Shi, Zhun Zhong, Meng Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Hierarchical Matrix-Contrastive Bilateral Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) seeks to understand human sentiment by leveraging the correlations across multimodal data. Current approaches often employ contrastive learning and text-centric fusion methods to explore the sentiment mapping space, improve the ability to extract and integrate multimodal features, and capture modality correlations. However, these methods typically depend on complex sampling strategies to select predefined positive and negative samples and perform unidirectional fusion of other modalities aligned with the text. This process overlooks the collaborative information that could be shared between modalities, leading to a loss of valuable insights. To address these limitations, we propose Hierarchical MAtrix-Contrastive BiLateral FusiOn (HALO), which integrates two key components: Matrix-Aware Contrastive Learning (MACL) and Hierarchical Bilateral Fusion (HBF). Specifically, MACL uses two supervisory signals to sample positive and negative pairs within the same batch and assigns different weights according to the difficulty of samples, thereby enhancing the cross-modal discrimination ability of the model. In addition, HBF introduces a bilateral fusion method by guiding vision and audio fusion with text, while using vision and audio information to enhance the overall expressive ability of text. Extensive experiments on datasets MOSI and MOSEI demonstrate the effectiveness and superiority of HALO. Chaoxing Tang, Anyang Tong, Fei Wang 0073, Zhangling Duan |
ICMR | 2 |
| 2025 | Learning from Heterogeneity: Generalizing Dynamic Facial Expression Recognition via Distributionally Robust OptimizationabstractDynamic Facial Expression Recognition (DFER) plays a critical role in affective computing and human-computer interaction. Although existing methods achieve comparable performance, they inevitably suffer from performance degradation under sample heterogeneity caused by multi-source data and individual expression variability. To address these challenges, we propose a novel framework, called Heterogeneity-aware Distributional Framework (HDF), and design two plug-and-play modules to enhance time-frequency modeling and mitigate optimization imbalance caused by hard samples. Specifically, the Time-Frequency Distributional Attention Module (DAM) captures both temporal consistency and frequency robustness through a dual-branch attention design, improving tolerance to sequence inconsistency and visual style shifts. Then, based on gradient sensitivity and information bottleneck principles, an adaptive optimization module Distribution-aware Scaling Module (DSM) is introduced to dynamically balance classification and contrastive losses, enabling more stable and discriminative representation learning. Extensive experiments on two widely used datasets, DFEW and FERV39k, demonstrate that HDF significantly improves both recognition accuracy and robustness. Our method achieves superior weighted average recall (WAR) and unweighted average recall (UAR) while maintaining strong generalization across diverse and imbalanced scenarios. Codes are released at https://github.com/QIcita/HDF_DFER. Feng-Qi Cui, Anyang Tong, Jinyang Huang, Jie Zhang 0073, Dan Guo 0001, Zhi Liu 0002, Meng Wang 0001 |
ACM Multimedia | 2 |
| 2025 | Attention mechanism based multimodal feature fusion network for human action recognition
Chao Tang 0002, Huosheng Hu, Wenjian Wang 0001, Shuo Qiao, Anyang Tong |
J. Vis. Commun. Image Represent. | 6 |
| 2024 | EMPC: Efficient multi-view parallel co-learning for semi-supervised action recognition
Anyang Tong, Chao Tang 0002, Wenjian Wang 0001 |
Expert Syst. Appl. | 1 |
| 2024 | Skeleton-based human action recognition by fusing attention based three-stream convolutional neural network and SVM
Fang Ren 0002, Chao Tang 0002, Anyang Tong, Wenjian Wang 0001 |
Multim. Tools Appl. | 3 |
| 2023 | Semi-Supervised Action Recognition From Temporal Augmentation Using Curriculum LearningabstractSemi-supervised learning for video action recognition is a very challenging research area. Existing state-of-the-art methods perform data augmentation on the temporality of actions, which are combined with the mainstream consistency-based semi-supervised learning framework FixMatch for action recognition. However, these approaches have the following limitations: (1) data augmentation based on video clips lacks coarse-grained and fine-grained representations of actions in temporal sequences, and the models have difficulty understanding synonymous representations of actions in different motion phases. (2) Pseudo labeling selection based on the constant thresholds lacks a “make-up curriculum” for difficult actions, that results in the low utilization of unlabeled data corresponding to difficult actions. To address the above shortcomings, we propose a semi-supervised action recognition via the temporal augmentation using curriculum learning (TACL) algorithm. Compared to previous works, TACL explores different representations of the same semantics of actions in temporal sequences for video and uses the idea of curriculum learning (CL) to reduce the difficulty of the model training process. First, for different action expressions with the same semantics, we designed the temporal action augmentation (TAA) for videos to obtain coarse-grained and fine-grained action expressions based on constant-velocity and hetero-velocity methods, respectively. Second, we construct a temporal signal to constrain the model such that fine-grained action expressions containing different movement phases have the same prediction results, and achieve action consistency learning (ACL) by combining the label and pseudo-label signals. Finally, we propose action curriculum pseudo labeling (ACPL), a loosely and strictly parallel dynamic threshold evaluation algorithm for selecting and labeling unlabeled data. We evaluate TACL on three standard public datasets: UCF101, HMDB51, and Kinetics. The combined experiments show that TACL significantly improves the accuracy of models trained on a small amount of labeled data and better evaluates the learning effects for different actions. Anyang Tong, Chao Tang 0002, Wenjian Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |