Yunyao Mao

dblp:299/1533 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-9427-9086ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Leveraging Visual Captions for Enhanced Zero-Shot HOI Detection
abstract
Zero-shot Human-Object Interaction (HOI) detection aims to identify both seen and unseen HOI categories in an image. Most existing methods rely on semantic knowledge distilled from CLIP to find novel interactions but fail to fully exploit the powerful generalization ability of vision-language models, leading to impaired transferability. In this paper, we introduce a novel framework for zero-shot HOI detection. We first utilize vision-language models (VLMs) to generate visual captions from multiple perspectives, including humans, objects, and environments, to enhance interaction understanding. Then, we propose a multi-modal fusion encoder to fully leverage these visual captions. Additionally, to equip the HOI detector with a thorough consideration of contextual information in the image, we design a novel multi-branch HOI network that aggregates features at the instance, union, and global levels. Experiments on prevalent benchmarks demonstrate that our model achieves promising performance under a variety of zero-shot settings. The source codes are available at https://github.com/aqingcv/VC-HOI.
Yanqing Zeng, Yunyao Mao, Zhenbo Lu, Wengang Zhou 0001, Houqiang Li
ICASSP2
2025 Hyper-Connections
abstract
We present hyper-connections, a simple yet effective method that can serve as an alternative to residual connections. This approach specifically addresses common drawbacks observed in residual connection variants, such as the seesaw effect between gradient vanishing and representation collapse. Theoretically, hyper-connections allow the network to adjust the strength of connections between features at different depths and dynamically rearrange layers. We conduct experiments focusing on the pre-training of large language models, including dense and sparse models, where hyper-connections show significant performance improvements over residual connections. Additional experiments conducted on vision tasks also demonstrate similar improvements. We anticipate that this method will be broadly applicable and beneficial across a wide range of AI problems.
Defa Zhu, Hongzhi Huang, Zihao Huang 0009, Yutao Zeng, Yunyao Mao, Banggu Wu, Qiyang Min
ICLR5
2025 $\hbox {I}^2$MD: 3D Action Representation Learning with Inter- and Intra-Modal Mutual Distillation
Yunyao Mao, Jiajun Deng, Wengang Zhou 0001, Zhenbo Lu, Wanli Ouyang, Houqiang Li
Int. J. Comput. Vis.1
2025 Diffusion With Reinforcement Learning for Pedestrian Trajectory Prediction
abstract
The trajectories of pedestrian movements involve uncertainty, requiring a predictive probability model capable of modeling the underlying multimodality. To predict the trajectory of pedestrian, most existing methods try to learn the probability distribution of real pedestrian trajectories and then independently sample multiple times from this distribution to obtain a set of possible future paths. However, naively learning the distribution of real-world trajectories leads to sub-optimal results. In this paper, we design a model-agnostic reinforcement learning-based framework for pedestrian trajectory prediction. This framework models pedestrian trajectory generation as a denoising process, which is further formulated as a multi-step decision-making process. In our framework, we subtly design a reward function, which is used to optimize the diffusion model with policy-based reinforcement learning. We make evaluation on multiple benchmark datasets, including ETH/UCY and SDD datasets, where our approach achieves promising results. Our source code will be released at:https://github.com/ustc-yaojinchen/DRL-for-PTP
Jinchen Yao, Zhenbo Lu, Yunyao Mao, Wengang Zhou 0001, Houqiang Li
IEEE Trans. Intell. Transp. Syst.3
2024 Detect Any Shadow: Segment Anything for Video Shadow Detection
abstract
Segment anything model (SAM) has achieved great success in the field of natural image segmentation. Nevertheless, SAM tends to consider shadows as background and therefore does not perform segmentation on them. In this paper, we propose ShadowSAM, a simple yet effective framework for fine-tuning SAM to detect shadows. Besides, by combining it with long short-term attention mechanism, we extend its capability for efficient video shadow detection. Specifically, we first fine-tune SAM on ViSha training dataset by utilizing the bounding boxes obtained from the ground truth shadow mask. Then during the inference stage, we simulate user interaction by providing bounding boxes to detect a specific frame (e.g., the first frame). Subsequently, using the detected shadow mask as a prior, we employ a long short-term network to learn spatial correlations between distant frames and temporal consistency between adjacent frames, thereby achieving precise shadow information propagation across video frames. Extensive experimental results demonstrate the effectiveness of our method, with notable margin over the state-of-the-art approaches in terms of MAE and IoU metrics. Moreover, our method exhibits accelerated inference speed compared to previous video shadow detection approaches, validating the effectiveness and efficiency of our method. The source code is now publicly available athttps://github.com/harrytea/Detect-AnyShadow.
Wengang Zhou 0001, Yunyao Mao, Houqiang Li
IEEE Trans. Circuits Syst. Video Technol.3
2024 MASA: Motion-Aware Masked Autoencoder With Semantic Alignment for Sign Language Recognition
abstract
Sign language recognition (SLR) has long been plagued by insufficient model representation capabilities. Although current pre-training approaches have alleviated this dilemma to some extent and yielded promising performance by employing various pretext tasks on sign pose data, these methods still suffer from two primary limitations: i) Explicit motion information is usually disregarded in previous pretext tasks, leading to partial information loss and limited representation capability. ii) Previous methods focus on the local context of a sign pose sequence, without incorporating the guidance of the global meaning of lexical signs. To this end, we propose a Motion-Aware masked autoencoder with Semantic Alignment (MASA) that integrates rich motion cues and global semantic information in a self-supervised learning paradigm for SLR. Our framework contains two crucial components, i.e., a motion-aware masked autoencoder (MA) and a momentum semantic alignment module (SA). Specifically, in MA, we introduce an autoencoder architecture with a motion-aware masked strategy to reconstruct motion residuals of masked frames, thereby explicitly exploring dynamic motion cues among sign pose sequences. Moreover, in SA, we embed our framework with global semantic awareness by aligning the embeddings of different augmented samples from the input sequence in the shared latent space. In this way, our framework can simultaneously learn local motion cues and global semantic features for comprehensive sign language representation. Furthermore, we conduct extensive experiments to validate the effectiveness of our method, achieving new state-of-the-art performance on four public benchmarks. The source code are publicly available athttps://github.com/sakura/MASA.
Weichao Zhao, Hezhen Hu, Wengang Zhou 0001, Yunyao Mao, Min Wang 0019, Houqiang Li
IEEE Trans. Circuits Syst. Video Technol.4
2023 Masked Motion Predictors are Strong 3D Action Representation Learners
abstract
In 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating effective self-supervised pre-training strategies. In this work, we show that instead of following the prevalent pretext task to perform masked self-component reconstruction in human joints, explicit contextual motion modeling is key to the success of learning effective feature representation for 3D action recognition. Formally, we propose the Masked Motion Prediction (MAMP) framework. To be specific, the proposed MAMP takes as input the masked spatio-temporal skeleton sequence and predicts the corresponding temporal motion of the masked human joints. Considering the high temporal redundancy of the skeleton sequence, in our MAMP, the motion information also acts as an empirical semantic richness prior that guide the masking process, promoting better attention to semantically rich temporal regions. Extensive experiments on NTU-60, NTU-120, and PKU-MMD datasets show that the proposed MAMP pre-training substantially improves the performance of the adopted vanilla transformer, achieving state-of-the-art results without bells and whistles. The source code of our MAMP is available at https://github.com/maoyunyao/MAMP.
Yunyao Mao, Jiajun Deng, Wengang Zhou 0001, Yao Fang, Wanli Ouyang, Houqiang Li
ICCV1
2023 CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI Detection
abstract
Zero-shot Human-Object Interaction (HOI) detection aims to identify both seen and unseen HOI categories. A strong zero-shot HOI detector is supposed to be not only capable of discriminating novel interactions but also robust to positional distribution discrepancy between seen and unseen categories when locating human-object pairs. However, top-performing zero-shot HOI detectors rely on seen and predefined unseen categories to distill knowledge from CLIP and jointly locate human-object pairs without considering the potential positional distribution discrepancy, leading to impaired transferability. In this paper, we introduce CLIP4HOI, a novel framework for zero-shot HOI detection. CLIP4HOI is developed on the vision-language model CLIP and ameliorates the above issues in the following two aspects. First, to avoid the model from overfitting to the joint positional distribution of seen human-object pairs, we seek to tackle the problem of zero-shot HOI detection in a disentangled two-stage paradigm. To be specific, humans and objects are independently identified and all feasible human-object pairs are processed by Human-Object interactor for pairwise proposal generation. Second, to facilitate better transferability, the CLIP model is elaborately adapted into a fine-grained HOI classifier for proposal discrimination, avoiding data-sensitive knowledge distillation. Finally, experiments on prevalent benchmarks show that our CLIP4HOI outperforms previous approaches on both rare and unseen categories, and sets a series of state-of-the-art records under a variety of zero-shot settings.
Yunyao Mao, Jiajun Deng, Wengang Zhou 0001, Li Li 0040, Yao Fang, Houqiang Li
NeurIPS1
2022 CMT: Context-Matching-Guided Transformer for 3D Tracking in Point Clouds
Zhiyang Guo, Yunyao Mao, Wengang Zhou 0001, Min Wang 0019, Houqiang Li
ECCV (22)2
2022 CMD: Self-supervised 3D Action Representation Learning with Cross-Modal Mutual Distillation
Yunyao Mao, Wengang Zhou 0001, Zhenbo Lu, Jiajun Deng, Houqiang Li
ECCV (3)1
2021 Joint Inductive and Transductive Learning for Video Object Segmentation
abstract
Semi-supervised video object segmentation is a task of segmenting the target object in a video sequence given only a mask annotation in the first frame. The limited information available makes it an extremely challenging task. Most previous best-performing methods adopt matching-based transductive reasoning or online inductive learning. Nevertheless, they are either less discriminative for similar instances or insufficient in the utilization of spatio-temporal information. In this work, we propose to integrate transductive and inductive learning into a unified framework to exploit the complementarity between them for accurate and robust video object segmentation. The proposed approach consists of two functional branches. The transduction branch adopts a lightweight transformer architecture to aggregate rich spatio-temporal cues while the induction branch performs online inductive learning to obtain discriminative target information. To bridge these two diverse branches, a two-head label encoder is introduced to learn the suitable target prior for each of them. The generated mask encodings are further forced to be disentangled to better retain their complementarity. Extensive experiments on several prevalent benchmarks show that, without the need of synthetic training data, the proposed approach sets a series of new state-of-the-art records. Code is available at https://github.com/maoyunyao/JOINT.
Yunyao Mao, Ning Wang 0020, Wengang Zhou 0001, Houqiang Li
ICCV1