EDBT 2026 Demo / reviewers in the wild / expert
Guangyao Li 0001
dblp:05/5839-1
· DBLP profile ↗
15ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-2179-8555ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video UnderstandingabstractLong video understanding (LVU) is challenging due to rich and complicated multimodal clues in long temporal range. Current methods adopt reasoning to improve the model's ability to analyze complex video clues in long videos via text-form reasoning. However, the existing literature suffers from the fact that the text-only reasoning under fixed video context may exacerbate hallucinations since detailed crucial clues are often ignored under limited video context length due to the temporal redundancy of long videos. To address this gap, we propose Video-TwG, a curriculum reinforced framework that employs a novel Think-with-Grounding paradigm, enabling video LLMs to actively decide when to perform on-demand grounding during interleaved text–video reasoning, selectively zooming into question-relevant clips only when necessary. Video-TwG can be trained end-to-end in a straightforward manner, without relying on complex auxiliary modules or heavily annotated reasoning traces. In detail, we design a Two-stage Reinforced Curriculum Strategy, where the model first learns think-with-grounding behavior on a small short-video GQA dataset with grounding labels, and then scales to diverse general QA data with videos of diverse domains to encourage generalization. Further, to handle complex think-with-grounding reasoning for various kinds of data, we propose the TwG-GRPO algorithm, which features the fine-grained grounding reward, self-confirmed pseudo reward, and accuracy-gated mechanism. Finally, we propose to construct a new TwG-51K dataset that facilitates training. Experiments on Video-MME, LongVideoBench, and MLVU show that Video-TwG consistently outperforms strong LVU baselines. Further ablation validates the necessity of our Two-stage Reinforced Curriculum Strategy and shows our TwG-GRPO better leverages diverse unlabeled data to improve grounding quality and reduce redundant groundings without sacrificing QA performance. https://github.com/hlchen23/Video-TwG Houlun Chen, Xin Wang 0019, Guangyao Li 0001, Yuwei Zhou, Jia Jia 0001, Wenwu Zhu 0001 |
SIGIR | 3 |
| 2026 | Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual AdaptationabstractMainstream research in audio-visual learning has focused on designing task-specific expert models, primarily implemented through sophisticated multimodal fusion approaches. Recently, a few efforts have aimed to develop more task-independent or universal audiovisual embedding networks, encoding advanced representations for use in various audiovisual downstream tasks. This is typically achieved by fine-tuning large pretrained transformers, such as Swin-V2-L and HTS-AT, in a parameter-efficient manner through techniques such as tuning only a few adapter layers inserted into the pretrained transformer backbone. Although these methods are parameter-efficient, they suffer from significant training memory consumption due to gradient backpropagation through the deep transformer backbones, which limits accessibility for researchers with constrained computational resources. In this paper, we present Meta-Token Learning (Mettle), a simple and memory-efficient method for adapting large-scale pretrained transformer models to downstream audio-visual tasks. Instead of sequentially modifying the output feature distribution of the transformer backbone, Mettle utilizes a lightweight Layer-Centric Distillation (LCD) module to distill in parallel the intact audio or visual features embedded by each transformer layer into compact meta-tokens. This distillation process considers both pretrained knowledge preservation and task-specific adaptation. The obtained meta-tokens can be directly applied to classification tasks, such as audio-visual event localization and audio-visual video parsing. To further support fine-grained segmentation tasks, such as audio-visual segmentation, we introduce a Meta-Token Injection (MTI) module, which utilizes the audio and visual meta-tokens distilled from the top transformer layer to guide feature adaptation in earlier layers. Extensive experiments on multiple audiovisual benchmarks demonstrate that our method significantly reduces memory usage and training time while maintaining parameter efficiency and competitive accuracy. Jinxing Zhou, Zhihui Li 0001, Yongqiang Yu, Yanghao Zhou, Ruohao Guo, Guangyao Li 0001, Yuxin Mao, Mingfei Han 0002, Xiaojun Chang, Meng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Sarcasm detection enhanced by multi-modal topics using denoising diffusion probabilistic models
Guangyao Li 0001, Buwen Liang |
Pattern Recognit. | 2 |
| 2025 | Crab: A Unified Audio-Visual Scene Understanding Model with Explicit CooperationabstractIn recent years, numerous tasks have been proposed to encourage model to develop specified capability in understanding audio-visual scene, primarily categorized into temporal localization, spatial localization, spatio-temporal reasoning, and pixel-level understanding. Instead, human possesses a unified understanding ability for diversified tasks. Therefore, designing an audio-visual model with general capability to unify these tasks is of great value. However, simply joint training for all tasks can lead to interference due to the heterogeneity of audiovisual data and complex relationship among tasks. We argue that this problem can be solved through explicit cooperation among tasks. To achieve this goal, we propose a unified learning method which achieves explicit inter-task cooperation from both the perspectives of data and model thoroughly. Specifically, considering the labels of existing datasets are simple words, we carefully refine these datasets and construct an Audio-Visual Unified Instruction-tuning dataset with Explicit reasoning process (AV-UIE), which clarifies the cooperative relationship among tasks. Subsequently, to facilitate concrete cooperation in learning stage, an interaction-aware LoRA structure with multiple LoRA heads is designed to learn different aspects of audiovisual data interaction. By unifying the explicit cooperation across the data and model aspect, our method not only surpasses existing unified audio-visual model on multiple tasks, but also outperforms most specialized models for certain tasks. Furthermore, we also visualize the process of explicit cooperation and surprisingly find that each LoRA head has certain audio-visual understanding ability. Code and dataset: https://github.com/GeWu-Lab/Crab Henghui Du, Guangyao Li 0001, Alan Zhao, Di Hu 0001 |
CVPR | 2 |
| 2025 | Audio-Visual Instance SegmentationabstractIn this paper, we propose a new multi-modal task, termed audio-visual instance segmentation (AVIS), which aims to simultaneously identify, segment and track individual sounding object instances in audible videos. To facilitate this research, we introduce a high-quality benchmark named AVISeg, containing over 90K instance masks from 26 semantic categories in 926 long videos. Additionally, we propose a strong baseline model for this task. Our model first localizes sound source within each frame, and condenses object-specific contexts into concise tokens. Then it builds long-range audio-visual dependencies between these tokens using window-based attention, and tracks sounding objects among the entire video sequences. Extensive experiments reveal that our method performs best on AVISeg, surpassing the existing methods from related tasks. We further conduct the evaluation on several multi-modal large models. Unfortunately, they exhibits subpar performance on instance-level sound source localization and temporal perception. We expect that AVIS will inspire the community towards a more comprehensive multi-modal understanding. Dataset and code is available at https://github.com/ruohaoguo/avis. Ruohao Guo, Xianghua Ying, Yaru Chen 0003, Dantong Niu, Guangyao Li 0001, Liao Qu, Yanyu Qi, Jinxing Zhou, Bowei Xing, Wenzhen Yue, Ji Shi 0003, Qixun Wang 0002, Peiliang Zhang, Buwen Liang |
CVPR | 5 |
| 2025 | PEDE: Enhance Multi-modal Sarcasm Detection in Videos via Prompted Emotion DistributionsabstractMulti-modal sarcasm detection is crucial for understanding human communications. A key aspect of multi-modal sarcasm detection is the analysis of emotion incongruity. However, the advancement of emotion analysis in video is hindered by the scarcity of labeled datasets, which are limited in both scale and diversity due to high human annotation cost. In this paper, to deal with different kinds of emotion distributions of open-topic in video data, we propose a simple yet remarkably effective method named Prompted Emotion Distribution Enhancement (PEDE). This method leverages large-scale pre-trained models to generate emotion distributions, thereby enriching the input features for sarcasm detection models. Then, intra- and inter-modality emotion graphs are constructed and a graph attention network (GAT) is used to learn emotion incongruity in input. Extensive experiments demonstrate that our approach can significantly enhance the performance of existing multi-modal sarcasm detection approaches on a sarcasm video dataset. Guangyao Li 0001, Buwen Liang |
ICASSP | 3 |
| 2025 | Improving Compositional Generalization in Cross-Embodiment Learning via Mixture of Disentangled Prototypes
Ren Wang 0011, Xin Wang 0019, Tongtong Feng, Xinyue Gong, Guangyao Li 0001, Yu-Wei Zhan, Qing Li 0046, Wenwu Zhu 0001 |
ACM Multimedia | 5 |
| 2024 | Prompting Segmentation with Sound Is Generalizable Audio-Visual Source LocalizerabstractNever having seen an object and heard its sound simultaneously, can the model still accurately localize its visual position from the input audio? In this work, we concentrate on the Audio-Visual Localization and Segmentation tasks but under the demanding zero-shot and few-shot scenarios. To achieve this goal, different from existing approaches that mostly employ the encoder-fusion-decoder paradigm to decode localization information from the fused audio-visual feature, we introduce the encoder-prompt-decoder paradigm, aiming to better fit the data scarcity and varying data distribution dilemmas with the help of abundant knowledge from pre-trained models. Specifically, we first propose to construct a Semantic-aware Audio Prompt (SAP) to help the visual foundation model focus on sounding objects, meanwhile, the semantic gap between the visual and audio modalities is also encouraged to shrink. Then, we develop a Correlation Adapter (ColA) to keep minimal training efforts as well as maintain adequate knowledge of the visual foundation model. By equipping with these means, extensive experiments demonstrate that this new paradigm outperforms other fusion-based methods in both the unseen class and cross-dataset settings. We hope that our work can further promote the generalization study of Audio-Visual Localization and Segmentation in practical application scenarios. Project page: https://github.com/GeWu-Lab/Generalizable-Audio-Visual-Segmentation Yaoting Wang, Weisong Liu, Guangyao Li 0001, Jian Ding 0001, Di Hu 0001 |
AAAI | 3 |
| 2024 | Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
Yaoting Wang, Peiwen Sun, Dongzhan Zhou, Guangyao Li 0001, Honggang Zhang 0002, Di Hu 0001 |
ECCV (74) | 4 |
| 2024 | CM-PIE: Cross-Modal Perception for Interactive-Enhanced Audio-Visual Video ParsingabstractAudio-visual video parsing is the task of categorizing a video with weak labels at the segment level, and predicting them as audible or visible events. Recent methods have leveraged the attention mechanism to capture the semantic correlations among the whole video across the audio-visual modalities. However, these approaches may overlook the importance of individual segments and their interrelations within a video, typically relying on a single modality when learning features. In this paper, we propose a novel interactive-enhanced cross-modal perception method (CM-PIE), which can learn fine-grained features by applying a segment-based attention module. In addition, a cross-modal aggregation block is introduced to jointly optimize the semantic representation of audio and visual signals by enhancing inter-modal interactions. Experimental results show that our model offers improved parsing performance on the Look, Listen, and Parse (LLP) dataset compared to other methods. Yaru Chen 0003, Ruohao Guo, Xubo Liu 0001, Peipei Wu, Guangyao Li 0001, Wenwu Wang 0001 |
ICASSP | 5 |
| 2024 | Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
Guangyao Li 0001, Henghui Du, Di Hu 0001 |
ACM Multimedia | 1 |
| 2023 | Multi-Scale Attention for Audio Question Answering
Guangyao Li 0001, Di Hu 0001 |
INTERSPEECH | 1 |
| 2023 | Progressive Spatio-temporal Perception for Audio-Visual Question AnsweringabstractAudio-Visual Question Answering (AVQA) task aims to answer questions about different visual objects, sounds, and their associations in videos. Such naturally multi-modal videos are composed of rich and complex dynamic audio-visual components, where most of which could be unrelated to the given questions, or even play as interference in answering the content of interest. Oppositely, only focusing on the question-aware audio-visual content could get rid of influence, meanwhile enabling the model to answer more efficiently. In this paper, we propose a Progressive Spatio-Temporal Perception Network (PSTP-Net), which contains three modules that progressively identify key spatio-temporal regions w.r.t. questions. Specifically, a temporal segment selection module is first introduced to select the most relevant audio-visual segments related to the given question. Then, a spatial region selection module is utilized to choose the most relevant regions associated with the question from the selected temporal segments. To further refine the selection of features, anaudio-guided visual attention module is employed to perceive the association between auido and selected spatial regions. Finally, the spatio-temporal features from these modules are integrated for answering the question. Extensive experimental results on the public MUSIC-AVQA and AVQA datasets provide compelling evidence of the effectiveness and efficiency of PSTP-Net. Guangyao Li 0001, Wenxuan Hou 0001, Di Hu 0001 |
ACM Multimedia | 1 |
| 2022 | Learning to Answer Questions in Dynamic Audio-Visual ScenariosabstractIn this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive multimodal understanding and spatio-temporal reasoning over audio-visual scenes. To benchmark this task and facilitate our study, we introduce a large-scale MUSIC-AVQA dataset, which contains more than 45K question-answer pairs covering 33 different question templates spanning over different modalities and question types. We develop several baselines and introduce a spatio-temporal grounded audio-visual network for the AVQA problem. Our results demonstrate that AVQA benefits from multisensory perception and our model outperforms recent A-, V-, and AVQA approaches. We believe that our built dataset has the potential to serve as testbed for evaluating and promoting progress in audio-visual scene understanding and spatio-temporal reasoning. Code and dataset: http://gewu-lab.github.io/MUSIC-AVQA/ Guangyao Li 0001, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, Di Hu 0001 |
CVPR | 1 |
| 2019 | Shellfish Detection Based on Fusion Attention Mechanism in End-to-End Network
Guangyao Li 0001, Chuyue Zhang, Yaodong Li |
PRCV (3) | 1 |