VLDB 2026 Research / reviewers in the wild / expert
Zhuoyan Luo
dblp:348/6516
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0000-5260-8497ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Segmentation and scene understanding · 28% Vision and language · 28% Video understanding and tracking · 16% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
referring image segmentation |
2.5 | 3 | 2026 | Semantic-Assisted Object Clustering for Multi-Modal Referring Video Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2026 CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation · ICCV 2025 SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation · NeurIPS 2023 |
Computer vision › Video understanding and tracking
temporal modeling |
1.0 | 1 | 2026 | Semantic-Assisted Object Clustering for Multi-Modal Referring Video Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Computer vision › Vision and language
cross-modal alignment |
1.0 | 2 | 2026 | SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation · NeurIPS 2023 Semantic-Assisted Object Clustering for Multi-Modal Referring Video Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning › Generative modeling
autoregressive model |
0.9 | 1 | 2025 | End-to-End Vision Tokenizer Tuning · NeurIPS 2025 |
Computer vision › Segmentation and scene understanding › referring image segmentation
generalized referring expression segmentation |
0.9 | 1 | 2025 | CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation · ICCV 2025 |
Machine learning › Generative modeling
image tokenization |
0.9 | 1 | 2025 | Scalable Image Tokenization with Index Backpropagation Quantization · ICCV 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Scalable Image Tokenization with Index Backpropagation Quantization · ICCV 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | End-to-End Vision Tokenizer Tuning · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
visual tokenizer |
0.9 | 1 | 2025 | End-to-End Vision Tokenizer Tuning · NeurIPS 2025 |
Computer vision › Vision and language
moment retrieval |
0.8 | 1 | 2024 | Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection · CVPR 2024 |
Computer vision › Vision and language
video grounding |
0.8 | 1 | 2024 | Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection · CVPR 2024 |
Computer vision › Video understanding and tracking › video summarization
video highlight detection |
0.8 | 1 | 2024 | Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection · CVPR 2024 |
Computer vision › Video understanding and tracking
video object segmentation |
0.2 | 1 | 2023 | SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 2.4dynamic query fusion · 1.0attention mechanism · 1.0hierarchical decoding · 0.9end-to-end training · 0.9transformer · 0.8multimodal fusion · 0.8multimodal transformer · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-Assisted Object Clustering for Multi-Modal Referring Video SegmentationabstractThis paper concentrates on Multi-modal Referring Video Segmentation task, where a well optimized model is able to recognize and segment the target objects referred by the given guidance signals, e.g., language description. Early approaches model this task as a sequence prediction problem. The lack of a global view of video content leads to difficulties in effectively utilizing inter-frame relationships. Some recent works propose to perform temporal modeling with vanilla attention mechanism. However, the condensed visual representation tends to be messy about target information due to occlusion or motion blur. Unlimited non-local operation would spread such noise to all the sequences and interfere with the extraction of global representations. To address the above issue, we present Semantic-assisted Object Cluster network (SOC) and the improved SOC++ in this paper. Our method unifies temporally selective interaction and cross-modal alignment to achieve video-level understanding. In SOC++, a proxy-assisted multi-modal fusion module is introduced to perform preliminary bidirectional activation. Then a semantic integration module with progressive frame-to-video structure facilitates joint space learning across modalities and time steps. Considering that potential noisy visual embeddings would impair the overall representation of target objects in unconstrained inter-frame interactions, we propose to perform tendentious video aggregation through emphasizing the indicative role of the informative frames with lower entropy in this part. A multi-modal query contrastive supervision is also utilized to help construct well-aligned joint space at the video level. Moreover, to integrate the advantage of high-level video information and the low-level details of each frame, we introduce a dynamic query fusion module that performs joint updating of these embeddings. We conduct extensive experiments on popular referring video segmentation benchmarks, and our method outperforms state-of-the-art competitors on all benchmarks by a remarkable margin. Besides, the emphasis on temporal coherence enhances the segmentation stability and adaptability of our method in processing text expressions with temporal variations.. Yong Liu 0033, Zhuoyan Luo, Yicheng Xiao, Shuyan Li, Xiu Li 0001, Yujiu Yang 0001, Yansong Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation
Zhuoyan Luo, Yinghao Wu, Tianheng Cheng, Yong Liu 0033, Yicheng Xiao, Hongfa Wang, Yujiu Yang 0001 |
ICCV | 1 |
| 2025 | Scalable Image Tokenization with Index Backpropagation Quantization
Fengyuan Shi 0001, Zhuoyan Luo, Yixiao Ge, Yujiu Yang 0001, Ying Shan, Limin Wang 0002 |
ICCV | 2 |
| 2025 | End-to-End Vision Tokenizer TuningabstractExisting vision tokenization isolates the optimization of vision tokenizers from downstream training, implicitly assuming the visual tokens can generalize well across various tasks, e.g., image generation and visual question answering. The vision tokenizer optimized for low-level reconstruction is agnostic to downstream tasks requiring varied representations and semantics. This decoupled paradigm introduces a critical misalignment: The loss of the vision tokenization can be the representation bottleneck for target tasks. For example, errors in tokenizing text in a given image lead to poor results when recognizing or generating them. To address this, we propose ETT, an end-to-end vision tokenizer tuning approach that enables joint optimization between vision tokenization and target autoregressive tasks. Unlike prior autoregressive models that use only discrete indices from a frozen vision tokenizer, ETT leverages the visual embeddings of the tokenizer codebook, and optimizes the vision tokenizers end-to-end with both reconstruction and caption objectives. ETT can be seamlessly integrated into existing training pipelines with minimal architecture modifications. Our ETT is simple to implement and integrate, without the need to adjust the original codebooks or architectures of the employed large language models. Extensive experiments demonstrate that our proposed end-to-end vision tokenizer tuning unlocks significant performance gains, i.e., 2-6% for multimodal understanding and visual generation tasks compared to frozen tokenizer baselines, while preserving the original reconstruction capability. We hope this very simple and strong method can empower multimodal foundation models besides image generation and understanding. Wenxuan Wang 0002, Yufeng Cui, Haiwen Diao, Zhuoyan Luo, Huchuan Lu, Jing Liu 0001 |
NeurIPS | 5 |
| 2024 | Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight DetectionabstractVideo Moment Retrieval (MR) and Highlight Detection (HD) have attracted significant attention due to the growing demand for video analysis. Recent approaches treat MR and HD as similar video grounding problems and address them together with transformer-based architecture. However, we observe that the emphasis of MR and HD differs, with one necessitating the perception of local relationships and the other prioritizing the understanding of global contexts. Consequently, the lack of task-specific design will inevitably lead to limitations in associating the intrinsic specialty of two tasks. To tackle the issue, we propose a Unified Video COMprehension framework (UVCOM) to bridge the gap and jointly solve MR and HD effectively. By performing progressive integration on intra and inter-modality across multi-granularity, UVCOM achieves the comprehensive understanding in processing a video. Moreover, we present multi-aspect contrastive learning to consolidate the local relation modeling and global knowledge accumulation via well aligned multi-modal space. Extensive experiments on QVHighlights, Charades-STA, TACoS, YouTube Highlights and TVSum datasets demonstrate the effectiveness and rationality of UVCOM which outperforms the state-of-the-art methods by a remarkable margin. Code is available at https://github.com/EasonXiao-888/UVCOM. Yicheng Xiao, Zhuoyan Luo, Yong Liu 0033, Yue Ma 0016, Hengwei Bian, Yatai Ji, Yujiu Yang 0001, Xiu Li 0001 |
CVPR | 2 |
| 2023 | SOC: Semantic-Assisted Object Cluster for Referring Video Object SegmentationabstractThis paper studies referring video object segmentation (RVOS) by boosting video-level visual-linguistic alignment. Recent approaches model the RVOS task as a sequence prediction problem and perform multi-modal interaction as well as segmentation for each frame separately. However, the lack of a global view of video content leads to difficulties in effectively utilizing inter-frame relationships and understanding textual descriptions of object temporal variations. To address this issue, we propose Semantic-assisted Object Cluster (SOC), which aggregates video content and textual guidance for unified temporal modeling and cross-modal alignment. By associating a group of frame-level object embeddings with language tokens, SOC facilitates joint space learning across modalities and time steps. Moreover, we present multi-modal contrastive supervision to help construct well-aligned joint space at the video level. We conduct extensive experiments on popular RVOS benchmarks, and our method outperforms state-of-the-art competitors on all benchmarks by a remarkable margin. Besides, the emphasis on temporal coherence enhances the segmentation stability and adaptability of our method in processing text expressions with temporal variations. Code is available at https://github.com/RobertLuo1/NeurIPS2023_SOC. Zhuoyan Luo, Yicheng Xiao, Yong Liu 0033, Shuyan Li, Yansong Tang, Xiu Li 0001, Yujiu Yang 0001 |
NeurIPS | 1 |
| 2023 | FATE: a three-stage method for arithmetical exercise correction
Qipeng Zhu, Zhuoyan Luo, Shipeng Zhu, Zihang Xu, Hui Xue 0002 |
Neural Comput. Appl. | 2 |