VLDB 2026 Research / reviewers in the wild / expert
Yanyu Qi
dblp:336/2131
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
0009-0008-9931-7855ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Segmentation and scene understanding · 79% Video understanding and tracking · 17% Vision and language · 4% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 100% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
instance segmentation |
0.9 | 1 | 2025 | Audio-Visual Instance Segmentation · CVPR 2025 |
Computer vision › Segmentation and scene understanding › audio-visual segmentation
audio-visual semantic segmentation |
0.8 | 1 | 2024 | Open-Vocabulary Audio-Visual Semantic Segmentation · ACM Multimedia 2024 |
Computer vision › Segmentation and scene understanding › image segmentation
co-segmentation |
0.8 | 1 | 2024 | UniTR: A Unified TRansformer-Based Framework for Co-Object and Multi-Modal Saliency Detection · IEEE Trans. Multim. 2024 |
Computer vision › Segmentation and scene understanding › saliency detection › salient object detection
multi-modal salient object detection |
0.8 | 1 | 2024 | UniTR: A Unified TRansformer-Based Framework for Co-Object and Multi-Modal Saliency Detection · IEEE Trans. Multim. 2024 |
Computer vision › Segmentation and scene understanding › saliency detection
salient object detection |
0.8 | 1 | 2024 | UniTR: A Unified TRansformer-Based Framework for Co-Object and Multi-Modal Saliency Detection · IEEE Trans. Multim. 2024 |
Multimedia analysis and retrieval › audio-visual learning
audio-visual saliency |
0.8 | 1 | 2024 | Instance-Level Panoramic Audio-Visual Saliency Detection and Ranking · ACM Multimedia 2024 |
Multimedia analysis and retrieval
audio-visual learning |
0.3 | 1 | 2025 | Audio-Visual Instance Segmentation · CVPR 2025 |
Computer vision › Segmentation and scene understanding › semantic segmentation
transformer-based segmentation |
0.2 | 1 | 2024 | UniTR: A Unified TRansformer-Based Framework for Co-Object and Multi-Modal Saliency Detection · IEEE Trans. Multim. 2024 |
Computer vision › Vision and language
vision-language pretraining |
0.2 | 1 | 2024 | Open-Vocabulary Audio-Visual Semantic Segmentation · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
window-based attention · 1.7multimodal fusion · 1.7audio-visual fusion · 1.5transformer · 0.8spatio-temporal object decoding · 0.8sound source localization · 0.8open-vocabulary classification · 0.8feature fusion · 0.8dual-stream decoding · 0.8distortion-aware decoding · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Audio-Visual Instance SegmentationabstractIn this paper, we propose a new multi-modal task, termed audio-visual instance segmentation (AVIS), which aims to simultaneously identify, segment and track individual sounding object instances in audible videos. To facilitate this research, we introduce a high-quality benchmark named AVISeg, containing over 90K instance masks from 26 semantic categories in 926 long videos. Additionally, we propose a strong baseline model for this task. Our model first localizes sound source within each frame, and condenses object-specific contexts into concise tokens. Then it builds long-range audio-visual dependencies between these tokens using window-based attention, and tracks sounding objects among the entire video sequences. Extensive experiments reveal that our method performs best on AVISeg, surpassing the existing methods from related tasks. We further conduct the evaluation on several multi-modal large models. Unfortunately, they exhibits subpar performance on instance-level sound source localization and temporal perception. We expect that AVIS will inspire the community towards a more comprehensive multi-modal understanding. Dataset and code is available at https://github.com/ruohaoguo/avis. Ruohao Guo, Xianghua Ying, Yaru Chen 0003, Dantong Niu, Guangyao Li 0001, Liao Qu, Yanyu Qi, Jinxing Zhou, Bowei Xing, Wenzhen Yue, Ji Shi 0003, Qixun Wang 0002, Peiliang Zhang, Buwen Liang |
CVPR | 7 |
| 2025 | SalienTR: A closer look at multi-modal transformer for RGB-T salient object detection
Ruohao Guo, Wenzhen Yue, Liao Qu, Yanyu Qi, Dantong Niu, Xianghua Ying |
Expert Syst. Appl. | 4 |
| 2025 | Plantformer: plant point cloud completion based on local-global feature aggregation and spatial context-aware transformer
Fei Li 0022, Yanyu Qi |
Neural Comput. Appl. | 3 |
| 2024 | Instance-Level Panoramic Audio-Visual Saliency Detection and RankingabstractPanoramic audio-visual saliency detection is to segment the most attention-attractive regions in 360° panoramic videos with sound. To meticulously delineate the detected salient regions and effectively model human attention shift, we extend this task to more fine-grained instance scenarios: identifying salient object instances and inferring their saliency ranks. In this paper, we propose the first instance-level framework that can simultaneously be applied to segmentation and ranking of multiple salient objects in panoramic videos. Specifically, it consists of a distortion-aware pixel decoder to overcome panoramic distortions, a sequential audio-visual fusion module to integrate audio-visual information, and a spatio-temporal object decoder to separate individual instances and predict their saliency scores. Moreover, owing to the absence of such annotations, we create the ground-truth saliency ranks for the PAVS10K benchmark. Extensive experiments demonstrate that our model is capable of achieving state-of-the-art performance on the PAVS10K for both saliency detection and ranking tasks. The code is available at https://github.com/ruohaoguo/pavsodr. Ruohao Guo, Dantong Niu, Liao Qu, Yanyu Qi, Ji Shi 0003, Wenzhen Yue, Bowei Xing, Taiyan Chen, Xianghua Ying |
ACM Multimedia | 4 |
| 2024 | Open-Vocabulary Audio-Visual Semantic SegmentationabstractAudio-visual semantic segmentation (AVSS) aims to segment and classify sounding objects in videos with acoustic cues. However, most approaches operate on the close-set assumption and only identify pre-defined categories from training data, lacking the generalization ability to detect novel categories in practical applications. In this paper, we introduce a new task: open-vocabulary audio-visual semantic segmentation, extending AVSS task to open-world scenarios beyond the annotated label space. This is a more challenging task that requires recognizing all categories, even those that have never been seen nor heard during training. Moreover, we propose the first open-vocabulary AVSS framework, OV-AVSS, which mainly consists of two parts: 1) a universal sound source localization module to perform audio-visual fusion and locate all potential sounding objects and 2) an open-vocabulary classification module to predict categories with the help of the prior knowledge from large-scale pre-trained vision-language models. To properly evaluate the open-vocabulary AVSS, we split zero-shot training and testing subsets based on the AVSBench-semantic benchmark, namely AVSBench-OV. Extensive experiments demonstrate the strong segmentation and zero-shot generalization ability of our model on all categories. On the AVSBench-OV dataset, OV-AVSS achieves 55.43% mIoU on base categories and 29.14% mIoU on novel categories, exceeding the state-of-the-art zero-shot method by 41.88%/20.61% and open-vocabulary method by 10.2%/11.6%. The code is available at https://github.com/ruohaoguo/ovavss. Ruohao Guo, Liao Qu, Dantong Niu, Yanyu Qi, Wenzhen Yue, Ji Shi 0003, Bowei Xing, Xianghua Ying |
ACM Multimedia | 4 |
| 2024 | Masked Visual Pre-training for RGB-D and RGB-T Salient Object Detection
Yanyu Qi, Ruohao Guo, Dantong Niu, Liao Qu |
PRCV (5) | 1 |
| 2024 | UniTR: A Unified TRansformer-Based Framework for Co-Object and Multi-Modal Saliency DetectionabstractRecent years have witnessed a growing interest in co-object segmentation and multi-modal salient object detection. Many efforts are devoted to segmenting co-existed objects among a group of images or detecting salient objects from different modalities. Albeit the appreciable performance achieved on respective benchmarks, each of these methods is limited to a specific task and cannot be generalized to other tasks. In this paper, we develop aUnifiedTRansformer-based framework, namelyUniTR, aiming at tackling the above tasks individually with a unified architecture. Specifically, a transformer module (CoFormer) is introduced to learn the consistency of relevant objects or complementarity from different modalities. To generate high-quality segmentation maps, we adopt a dual-stream decoding paradigm that allows the extracted consistent or complementary information to better guide mask prediction. Moreover, a feature fusion module (ZoomFormer) is designed to enhance backbone features and capture multi-granularity and multi-semantic information. Extensive experiments show that our UniTR performs well on17benchmarks, and surpasses existing state-of-the-art approaches. Ruohao Guo, Xianghua Ying, Yanyu Qi, Liao Qu |
IEEE Trans. Multim. | 3 |