VLDB 2026 Research / reviewers in the wild / expert
Yanghao Zhou
dblp:227/8249
· DBLP profile ↗
17ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual SegmentationabstractReferring Audio-Visual Segmentation (Ref-AVS) aims to segment target objects in audible videos based on given reference expressions. Prior works typically rely on learning latent embeddings via multimodal fusion to prompt a tunable SAM/SAM2 decoder for segmentation, which requires strong pixel-level supervision and lacks interpretability. From a novel perspective of explicit reference understanding, we propose TGS-Agent, which decomposes the task into a Think-Ground-Segment process, mimicking the human reasoning procedure by first identifying the referred object through multimodal analysis, followed by coarse-grained grounding and precise segmentation. To this end, we first propose Ref-Thinker, a multimodal language model capable of reasoning over textual, visual, and auditory cues. We construct an instruction-tuning dataset with explicit object-aware think-answer chains for Ref-Thinker fine-tuning. The object description inferred by Ref-Thinker is used as an explicit prompt for Grounding-DINO and SAM2, which perform grounding and segmentation without relying on pixel-level supervision. Additionally, we introduce R2-AVSBench, a new benchmark with linguistically diverse and reasoning-intensive references for better evaluating model generalization. Our approach achieves state-of-the-art results on both standard Ref-AVSBench and proposed R2-AVSBench. Jinxing Zhou, Yanghao Zhou, Mingfei Han 0002, Tong Wang 0022, Xiaojun Chang, Hisham Cholakkal, Rao Muhammad Anwer |
AAAI | 2 |
| 2026 | CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event LocalizationabstractThe Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and more challenging weakly-supervised setting (W-DAVEL task), where only video-level event labels are provided and the temporal boundaries of each event are unknown. We address W-DAVEL by exploiting cross-modal salient anchors, which are defined as reliable timestamps that are well predicted under weak supervision and exhibit highly consistent event semantics across audio and visual modalities. Specifically, we propose a Mutual Event Agreement Evaluation module, which generates an agreement score by measuring the discrepancy between the predicted audio and visual event classes. Then, the agreement score is utilized in a Cross-modal Salient Anchor Identification module, which identifies the audio and visual anchor features through global-video and local temporal window identification mechanisms. The anchor features after multimodal integration are fed into an Anchor-based Temporal Propagation module to enhance event semantic encoding in the original temporal audio and visual features, facilitating better temporal localization under weak supervision. We establish benchmarks for W-DAVEL on both the UnAV-100 and ActivityNet1.3 datasets. Extensive experiments demonstrate that our method achieves state-of-the-art performance. Jinxing Zhou, Yanghao Zhou, Yuxin Mao, Zhangling Duan, Dan Guo 0001 |
AAAI | 3 |
| 2026 | MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video GenerationabstractYanghao Zhou, Haitian Li, Rexar Lin, Heyan Huang, Jinxing Zhou, Changsen Yuan, Tian Lan, Ziqin Zhou, Yudong Li, Jiajun Xu, Jingyun Liao, YiMing Cheng, Xuefeng Chen, Xian-Ling Mao, Yousheng Feng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yanghao Zhou, Haitian Li, Rexar Lin, Heyan Huang, Jinxing Zhou, Changsen Yuan, Tian Lan 0003, Ziqin Zhou, Jingyun Liao, YiMing Cheng, Xianling Mao, Yousheng Feng |
ACL (1) | 1 |
| 2026 | Selective and contrastive mechanism for distantly supervised relation extraction
Danjie Han, Heyan Huang, Shumin Shi, Cunhan Guo, Yanghao Zhou, Changsen Yuan |
Neurocomputing | 6 |
| 2026 | Towards unified scene understanding in audio-visual semantic segmentation
Danjie Han, Changsen Yuan, Xuan Zhao 0026, Cunhan Guo, Yanghao Zhou |
Knowl. Based Syst. | 6 |
| 2026 | Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual AdaptationabstractMainstream research in audio-visual learning has focused on designing task-specific expert models, primarily implemented through sophisticated multimodal fusion approaches. Recently, a few efforts have aimed to develop more task-independent or universal audiovisual embedding networks, encoding advanced representations for use in various audiovisual downstream tasks. This is typically achieved by fine-tuning large pretrained transformers, such as Swin-V2-L and HTS-AT, in a parameter-efficient manner through techniques such as tuning only a few adapter layers inserted into the pretrained transformer backbone. Although these methods are parameter-efficient, they suffer from significant training memory consumption due to gradient backpropagation through the deep transformer backbones, which limits accessibility for researchers with constrained computational resources. In this paper, we present Meta-Token Learning (Mettle), a simple and memory-efficient method for adapting large-scale pretrained transformer models to downstream audio-visual tasks. Instead of sequentially modifying the output feature distribution of the transformer backbone, Mettle utilizes a lightweight Layer-Centric Distillation (LCD) module to distill in parallel the intact audio or visual features embedded by each transformer layer into compact meta-tokens. This distillation process considers both pretrained knowledge preservation and task-specific adaptation. The obtained meta-tokens can be directly applied to classification tasks, such as audio-visual event localization and audio-visual video parsing. To further support fine-grained segmentation tasks, such as audio-visual segmentation, we introduce a Meta-Token Injection (MTI) module, which utilizes the audio and visual meta-tokens distilled from the top transformer layer to guide feature adaptation in earlier layers. Extensive experiments on multiple audiovisual benchmarks demonstrate that our method significantly reduces memory usage and training time while maintaining parameter efficiency and competitive accuracy. Jinxing Zhou, Zhihui Li 0001, Yongqiang Yu, Yanghao Zhou, Ruohao Guo, Guangyao Li 0001, Yuxin Mao, Mingfei Han 0002, Xiaojun Chang, Meng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Dynamic Model-Bank Test-Time Adaptation for Automatic Speech RecognitionabstractEnd-to-end automatic speech recognition (ASR) based on deep learning has achieved impressive progress in recent years.However, the performance of ASR foundation model often degrades significantly on out-of-domain data due to real-world domain shifts.Test-Time Adaptation (TTA) methods aim to mitigate this issue by adapting models during inference without access to source data.Despite recent progress, existing ASR TTA methods often struggle with instability under continual and long-term distribution shifts.To alleviate the risk of performance collapse due to error accumulation, we propose Dynamic Model-bank Single-Utterance Test-time Adaptation (DM-SUTA), a sustainable continual TTA framework based on adaptive ASR model ensembling.DMSUTA maintains a dynamic model bank, from which a subset of checkpoints is selected for each test sample based on confidence and uncertainty criteria.To preserve both model plasticity and long-term stability, DMSUTA actively manages the bank by filtering out potentially collapsed models.This design allows DMSUTA to continually adapt to evolving domain shifts in ASR test-time scenarios.Experiments on diverse, continuously shifting ASR TTA benchmarks show that DM-SUTA consistently outperforms existing continual TTA baselines, demonstrating superior robustness to domain shifts in ASR. Yanshuo Wang, Yanghao Zhou, Yukang Lin, Haoxing Chen |
EMNLP | 2 |
| 2025 | SignDiff: Diffusion Model for American Sign Language ProductionabstractIn this paper, we propose a dual-condition diffusion pre-training model named SignDIFF that can generate human sign language speakers from a skeleton pose. SignDiff has a novel Frame Reinforcement Network called FR-Net, similar to dense human pose estimation work, which enhances the correspondence between text lexical symbols and sign language dense pose frames, reduces the occurrence of multiple fingers in the diffusion model. In addition, we propose a new method for American Sign Language Production (ASLP), which can generate ASL skeletal pose videos from text input, integrating two new improved modules and a new loss function to improve the accuracy and quality of sign language skeletal posture and enhance the ability of the model to train on largescale data. We propose a simple baseline for ASL production and report the scores of 17.19 and 12.85 on BLEU-4 on the How2Sign dev/test sets. We evaluated our model on the previous mainstream dataset PHOENIX14T, and our method achieved the SOTA results. In addition, our image quality far exceeds all previous results by 10 percentage points in terms of SSIM. Sen Fang, Chunyu Sui, Yanghao Zhou, Hongbin Zhong, Yapeng Tian, Chen Chen 0001 |
FG | 3 |
| 2025 | AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
Yang Xiao 0019, Tianyi Peng, Yanghao Zhou, Rohan Kumar Das |
INTERSPEECH | 3 |
| 2025 | CARE: Contextual Augmentation with Retrieval Enhancement for Relation Extraction in Large Language Models
Danjie Han, Heyan Huang, Shumin Shi, Cunhan Guo, Yanghao Zhou, Changsen Yuan |
NLPCC (1) | 6 |
| 2024 | Leveraging Vote and Cooperate Mechanism for Brain Tumor SegmentationabstractThe task of brain tumor segmentation necessitates the processing of long Magnetic Resonance Imaging (MRI) sequences across multiple imaging modalities. Traditional convolutional neural networks often exhibit suboptimal ability of long-term memory, while Transformers demand substantial computational resources. To mitigate computational requirements and optimally utilize the information from various imaging modalities, we propose Mamba-based Vote and Cooperate Segmentation (VCSeg), for brain tumor segmentation. Features from the imaging sequences of each modality are extracted from multiple spatial orientations, and assigned different weights based on both modality and orientation considerations, enabling the model to adaptively learn modality-specific feature distribution for adult and child patient situation. The model employs a multi-modal cooperate module to further enhance the feature encoding capabilities. Additionally, deep supervision is applied to achieve progressive mask generation, thereby improving decoding quality. Experiments conducted on the BraTS2019-MEN and BraTS2023-PED datasets have demonstrated that VCSeg achieves state-of-the-art results, demonstrating the effectiveness of our advanced method. Cunhan Guo, Heyan Huang, Changsen Yuan, Yanghao Zhou |
BIBM | 4 |
| 2024 | Improving Inference via Rich Path Information for Dialogue Relation Extraction
Huizhe Su, Hang Yu 0006, Yanghao Zhou, Changsen Yuan, Shaorong Xie, Xiangfeng Luo |
NLPCC (5) | 3 |
| 2024 | Enhance audio-visual segmentation with hierarchical encoder and audio guidance
Cunhan Guo, Heyan Huang, Yanghao Zhou |
Neurocomputing | 3 |
| 2020 | Temporal Network Representation Learning via Historical Neighborhoods AggregationabstractNetwork embedding is an effective method to learn low-dimensional representations of nodes, which can be applied to various real-life applications such as visualization, node classification, and link prediction. Although significant progress has been made on this problem in recent years, several important challenges remain, such as how to properly capture temporal information in evolving networks. In practice, most networks are continually evolving. Some networks only add new edges or nodes such as authorship networks, while others support removal of nodes or edges such as internet data routing. If patterns exist in the changes of the network structure, we can better understand the relationships between nodes and the evolution of the network, which can be further leveraged to learn node representations with more meaningful information. In this paper, we propose the Embedding via Historical Neighborhoods Aggregation (EHNA) algorithm. More specifically, we first propose a temporal random walk that can identify relevant nodes in historical neighborhoods which have impact on edge formations. Then we apply a deep learning model which uses a custom attention mechanism to induce node embeddings that directly capture temporal information in the underlying feature representation. We perform extensive experiments on a range of real-world datasets, and the results demonstrate the effectiveness of our new approach in the network reconstruction task and the link prediction task. Shixun Huang, Zhifeng Bao, Guoliang Li 0001, Yanghao Zhou, J. Shane Culpepper |
ICDE | 4 |
| 2020 | Deep fractal residual network for fast and accurate single image super resolution
Yanghao Zhou, Jianfeng Dong |
Neurocomputing | 1 |
| 2019 | A novel test-cost-sensitive attribute reduction approach using the binary bat algorithm
Xiaolin Qin, Qian Zhou 0005, Yanghao Zhou, Ryszard Janicki, Wei Zhao 0061 |
Knowl. Based Syst. | 4 |
| 2019 | ITAA: An Intelligent Trajectory-driven Outdoor Advertising Deployment AssistantabstractIn this paper, we demonstrate an Intelligent Trajectory-driven outdoor Advertising deployment Assistant (ITAA), which assists users to find an optimal strategy for outdoor advertising (ad) deployment. The challenge is how to measure the influence to the moving trajectories of ads, and how to optimize the placement of ads among billboards that maximize the influence has been proven NP-hard. Therefore, we develop a framework based on two trajectory-driven influence models. ITAA is built upon this framework with a user-friendly UI. It serves both ad companies and their customers. We enhance the interpretability to improve the user's understanding of the influence of ads. The interactive function of ITAA is made interpretable and easy to engage. Yipeng Zhang 0002, Zhifeng Bao, Songsong Mo, Yuchen Li 0001, Yanghao Zhou |
Proc. VLDB Endow. | 5 |