VLDB 2026 Research / reviewers in the wild / expert
Junjie Zhou 0001
dblp:125/1074-1
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-5903-2806ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Vision and language · 23% Generative modeling · 23% Video understanding and tracking · 21% | |
| Computer graphics and multimedia
3 papers |
Visual content generation and editing · 52% Multimedia analysis and retrieval · 35% Image and video processing · 13% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
long video understanding |
1.7 | 2 | 2025 | Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding · CVPR 2025 MLVU: Benchmarking Multi-task Long Video Understanding · CVPR 2025 |
Information retrieval
multimodal retrieval |
1.6 | 2 | 2025 | Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval · ACL (1) 2025 VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval · ACL (1) 2024 |
Machine learning › Generative modeling
diffusion model |
1.5 | 2 | 2025 | OmniGen: Unified Image Generation · CVPR 2025 DocDiff: Document Enhancement via Residual Diffusion Models · ACM Multimedia 2023 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding · CVPR 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding · CVPR 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation |
0.9 | 1 | 2025 | MLVU: Benchmarking Multi-task Long Video Understanding · CVPR 2025 |
Machine learning › Efficient and distributed learning › model compression › token compression
visual token compression |
0.9 | 1 | 2025 | Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding · CVPR 2025 |
Multimedia analysis and retrieval › image retrieval
composed image retrieval |
0.9 | 1 | 2025 | MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval · ACL (1) 2025 |
Multimedia analysis and retrieval
cross-modal retrieval |
0.9 | 1 | 2025 | MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval · ACL (1) 2025 |
Visual content generation and editing
image editing |
0.9 | 1 | 2025 | OmniGen: Unified Image Generation · CVPR 2025 |
Visual content generation and editing › image generation › personalized image generation
subject-driven generation |
0.9 | 1 | 2025 | OmniGen: Unified Image Generation · CVPR 2025 |
Machine learning › Representation and self-supervised learning
text embedding |
0.8 | 1 | 2024 | VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval · ACL (1) 2024 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
visual token representation |
0.8 | 1 | 2024 | VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval · ACL (1) 2024 |
Information retrieval › multimodal retrieval
universal multimodal retrieval |
0.8 | 1 | 2024 | VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval · ACL (1) 2024 |
Machine learning › Generative modeling › diffusion model › diffusion model architecture
residual diffusion model |
0.7 | 1 | 2023 | DocDiff: Document Enhancement via Residual Diffusion Models · ACM Multimedia 2023 |
Image and video processing › image enhancement
document image enhancement |
0.7 | 1 | 2023 | DocDiff: Document Enhancement via Residual Diffusion Models · ACM Multimedia 2023 |
Computer vision › Vision and language
vision-language model |
0.5 | 2 | 2025 | MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval · ACL (1) 2025 VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval · ACL (1) 2024 |
Computer vision › Video understanding and tracking
video question answering |
0.3 | 1 | 2025 | MLVU: Benchmarking Multi-task Long Video Understanding · CVPR 2025 |
Information retrieval › distributed information retrieval
integrated search |
0.3 | 1 | 2025 | Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 3.3diffusion model · 3.1data synthesis · 1.7chain-of-thought · 1.7data generation · 1.5screenshot-based retrieval · 0.9multi-task evaluation · 0.9key-value sparsification · 0.9instruction fine-tuning · 0.9empirical study · 0.9curriculum learning · 0.9multi-stage training · 0.8residual refinement · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CPG: Contrastive Patch-Graph learning for 3D point cloud
Junjie Zhou 0001, Yingde Song, Chinwai Chiu, Yongping Xiong, Yuxin Luo, Siyang Song |
Pattern Recognit. | 1 |
| 2025 | MegaPairs: Massive Data Synthesis for Universal Multimodal RetrievalabstractDespite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massive synthetic dataset generated from this method. Our empirical analysis shows that MegaPairs generates high-quality data, enabling the multimodal retriever to significantly outperform the baseline model trained on 70\times more data from existing datasets. Moreover, since MegaPairs solely relies on general image corpora and open-source VLMs, it can be easily scaled up, enabling continuous improvements in retrieval performance. In this stage, we produced more than 26 million training instances and trained several models of varying sizes using this data. These new models achieve state-of-the-art zero-shot performance across 4 popular composed image retrieval (CIR) benchmarks and the highest overall performance on the 36 datasets provided by MMEB. They also demonstrate notable performance improvements with additional downstream fine-tuning. Our code, synthesized dataset, and pre-trained models are publicly available at https://github.com/VectorSpaceLab/MegaPairs. Junjie Zhou 0001, Yongping Xiong, Zheng Liu 0011, Shitao Xiao, Yueze Wang, Bo Zhao 0015, Chen Zhang 0013, Defu Lian |
ACL (1) | 1 |
| 2025 | Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information RetrievalabstractZheng Liu, Ze Liu, Zhengyang Liang, Junjie Zhou, Shitao Xiao, Chao Gao, Chen Jason Zhang, Defu Lian. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zheng Liu 0011, Zhengyang Liang, Junjie Zhou 0001, Shitao Xiao, Chen Zhang 0013, Defu Lian |
ACL (1) | 4 |
| 2025 | MLVU: Benchmarking Multi-task Long Video UnderstandingabstractThe evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in video types and evaluation tasks, and the inappropriateness for evaluating LVU performances. To address the above problems, we propose a new benchmark called MLVU (Multitask Long Video Understanding Benchmark) for the comprehensive and in-depth evaluation of LVU. MLVU presents the following critical values: 1) The substantial and flexible extension of video lengths, which enables the benchmark to evaluate LVU performance across a wide range of durations. 2) The inclusion of various video genres, such as movies, surveillance, egocentric videos, and cartoons, reflects the models’ LVU performances in different scenarios. 3) The development of diversified evaluation tasks, which enables a comprehensive examination of MLLMs’ key abilities in long-video understanding. The empirical study with 23 latest MLLMs reveals significant room for improvement in today’s technique, as all existing methods struggle with most of the evaluation tasks and exhibit severe performance degradation when handling longer videos. Additionally, it suggests that factors such as context length, image-understanding ability, and the choice of LLM backbone can play critical roles in future advancements. We anticipate that MLVU will advance the research of LVU by providing a comprehensive and in-depth analysis of MLLMs. The code and dataset can be accessed from https://github.com/JUNJIE99/MLVU. Junjie Zhou 0001, Bo Zhao 0015, Boya Wu, Zhengyang Liang, Shitao Xiao, Minghao Qin, Yongping Xiong, Tiejun Huang 0001, Zheng Liu 0011 |
CVPR | 1 |
| 2025 | Video-XL: Extra-Long Vision Language Model for Hour-Scale Video UnderstandingabstractLong video understanding poses a significant challenge for current Multi-modal Large Language Models (MLLMs). Notably, the MLLMs are constrained by their limited context lengths and the substantial costs while processing long videos. Although several existing methods attempt to reduce visual tokens, their strategies encounter severe bottleneck, restricting MLLMs’ ability to perceive fine-grained visual details. In this work, we propose Video-XL, a novel approach that leverages MLLMs’ inherent key-value (KV) sparsification capacity to condense the visual input. Specifically, we introduce a new special token, the Visual Summarization Token (VST), for each interval of the video, which summarizes the visual information within the interval as its associated KV. The VST module is trained by instruction fine-tuning, where two optimizing strategies are offered. 1. Curriculum learning, where VST learns to make small (easy) and large compression (hard) progressively. 2. Composite data curation, which integrates single-image, multi-image, and synthetic data to overcome the scarcity of long-video instruction data. The compression quality is further improved by dynamic compression, which customizes compression granularity based on the information density of different video intervals. Video-XL’s effectiveness is verified from three aspects. First, it achieves a superior long-video understanding capability, outperforming state-of-the-art models of comparable sizes across multiple popular benchmarks. Second, it effectively preserves video information, with minimal compression loss even at 16 × compression ratio. Third, it realizes outstanding cost-effectiveness, enabling high-quality processing of thousands of frames on a single A100 GPU. Zheng Liu 0011, Peitian Zhang, Minghao Qin, Junjie Zhou 0001, Zhengyang Liang, Tiejun Huang 0001, Bo Zhao 0015 |
CVPR | 5 |
| 2025 | OmniGen: Unified Image GenerationabstractThe emergence of Large Language Models (LLMs) has unified language generation tasks and revolutionized human-machine interaction. However, in the realm of image generation, a unified model capable of handling various tasks within a single framework remains largely unexplored. In this work, we introduce OmniGen, a new diffusion model for unified image generation. OmniGen is characterized by the following features: 1) Unification: OmniGen not only demonstrates text-to-image generation capabilities but also inherently supports various downstream tasks, such as image editing, subject-driven generation, and visual-conditional generation. 2) Simplicity: The architecture of OmniGen is highly simplified, eliminating the need for additional plugins. Moreover, compared to existing diffusion models, it is more user-friendly and can complete complex tasks end-to-end through instructions without the need for extra intermediate steps, greatly simplifying the image generation workflow. 3) Knowledge Transfer: Benefit from learning in a unified format, OmniGen effectively transfers knowledge across different tasks, manages unseen tasks and domains, and exhibits novel capabilities. We also explore the model’s reasoning capabilities and potential applications of the chain-of-thought mechanism. This work represents the first attempt at a general-purpose image generation model, and we will release our resources at https://github.com/VectorSpaceLab/OmniGen to foster future advancements. Shitao Xiao, Yueze Wang, Junjie Zhou 0001, Huaying Yuan, Xingrun Xing, Ruiran Yan, Shuting Wang 0002, Tiejun Huang 0001, Zheng Liu 0011 |
CVPR | 3 |
| 2025 | MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment RetrievalabstractAccurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or they focus solely on the end-to-end LVU performance, making them inappropriate for evaluating whether key moments can be accurately accessed. To address this challenge, we propose MomentSeeker, a novel benchmark for long-video moment retrieval (LVMR), distinguished by the following features. First, it is created based on long and diverse videos, averaging over 1,200 seconds in duration, and collected from various domains, e.g., movie, anomaly, egocentric, and sports. Second, it covers a variety of real-world scenarios in three levels: global-level, event-level, and object-level, covering common tasks like action recognition, object localization, causal reasoning, etc. Third, it incorporates rich forms of queries, including text-only queries, image-conditioned queries, and video-conditioned queries. On top of MomentSeeker, we conduct comprehensive experiments for both generation-based approaches (directly using MLLMs) and retrieval-based approaches (leveraging video retrievers). Our results reveal the significant challenges in long-video moment retrieval in terms of accuracy and efficiency, despite improvements from the latest long-video MLLMs and task-specific fine-tuning. We have publicly released MomentSeeker to facilitate future research in this area. Huaying Yuan, Jian Ni, Zheng Liu 0011, Yueze Wang, Junjie Zhou 0001, Zhengyang Liang, Bo Zhao 0015, Zhao Cao, Ji-Rong Wen, Zhicheng Dou |
NeurIPS | 5 |
| 2024 | VISTA: Visualized Text Embedding For Universal Multi-Modal RetrievalabstractMulti-modal retrieval becomes increasingly popular in practice.However, the existing retrievers are mostly text-oriented, which lack the capability to process visual information.Despite the presence of vision-language models like CLIP, the current methods are severely limited in representing the text-only and imageonly data.In this work, we present a new embedding model VISTA for universal multimodal retrieval.Our work brings forth threefold technical contributions.Firstly, we introduce a flexible architecture which extends a powerful text encoder with the image understanding capability by introducing visual token embeddings.Secondly, we develop two data generation strategies, which bring highquality composed image-text to facilitate the training of the embedding model.Thirdly, we introduce a multi-stage training algorithm, which first aligns the visual token embedding with the text encoder using massive weakly labeled data, and then develops multi-modal representation capability using the generated composed image-text data.In our experiments, VISTA achieves superior performances across a variety of multi-modal retrieval tasks in both zero-shot and supervised settings.Our model, data, and source code are available at https://github.com/FlagOpen/FlagEmbedding. Junjie Zhou 0001, Zheng Liu 0011, Shitao Xiao, Bo Zhao 0015, Yongping Xiong |
ACL (1) | 1 |
| 2024 | FAT: Field-Aware Transformer for Point Cloud Segmentation With Adaptive Attention FieldsabstractPoint cloud segmentation is crucial for various industrial applications, such as autonomous driving and robotics. Recent developments underscore the significant potential of transformer models in this field. However, existing attention mechanisms apply the same feature learning paradigm for all points equally, ignoring the considerable size differences among objects in a scene. To rectify this, we introduce the field-aware transformer (FAT), engineered to tailor effective receptive fields to objects of varying sizes. Our FAT achieves field-aware learning through two primary components: the multigranularity attention (MGA) scheme and the reattention module. The MGA scheme is proficient in aggregating tokens from distant areas while preserving multiscale features within each attention layer. The reattention module dynamically adjusts the attention scores to the fine- and coarse-grained features output by MGA for each point. Extensive experimental results underscore the effectiveness and efficiency of our FAT, which delivers state-of-the-art performance on both the stanford 3D indoor scene dataset (S3DIS) and ScanNetV2 datasets. Junjie Zhou 0001, Baolin Liu 0002, Yongping Xiong, Chinwai Chiu, Xiangyang Gong |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Fat: Field-Aware Transformer for 3D Point Cloud Semantic SegmentationabstractTransformer models have achieved promising performances in point cloud segmentation. However, most existing attention schemes provide the same feature learning paradigm for all points equally and overlook the enormous difference in size among scene objects. In this paper, we propose the Field-Aware Transformer (FAT) that adjusts the attentive receptive fields for objects of different sizes. Our FAT achieves field-aware learning via two steps: introduce multi-granularity features to each attention layer and allow each point to choose its attentive fields adaptively. It contains two key designs: the Multi-Granularity Attention (MGA) scheme and the Re-Attention module. Extensive experimental results demonstrate that FAT achieves state-of-the-art performances on S3DIS [1] and ScanNetV2 [2] datasets. Junjie Zhou 0001, Yongping Xiong, Chinwai Chiu, Xiangyang Gong |
ICIP | 1 |
| 2023 | DocDiff: Document Enhancement via Residual Diffusion ModelsabstractRemoving degradation from document images not only improves their visual quality and readability, but also enhances the performance of numerous automated document analysis and recognition tasks. However, existing regression-based methods optimized for pixel-level distortion reduction tend to suffer from significant loss of high-frequency information, leading to distorted and blurred text edges. To compensate for this major deficiency, we propose DocDiff, the first diffusion-based framework specifically designed for diverse challenging document enhancement problems, including document deblurring, denoising, and removal of watermarks and seals. DocDiff consists of two modules: the Coarse Predictor (CP), which is responsible for recovering the primary low-frequency content, and the High-Frequency Residual Refinement (HRR) module, which adopts the diffusion models to predict the residual (high-frequency information, including text edges), between the ground-truth and the CP-predicted image. DocDiff is a compact and computationally efficient model that benefits from a well-designed network architecture, an optimized training loss objective, and a deterministic sampling process with short time steps. Extensive experiments demonstrate that DocDiff achieves state-of-the-art (SOTA) performance on multiple benchmark datasets, and can significantly enhance the readability and recognizability of degraded document images. Furthermore, our proposed HRR module in pre-trained DocDiff is plug-and-play and ready-to-use, with only 4.17M parameters. It greatly sharpens the text edges generated by SOTA deblurring methods without additional joint training. Available codes: https://github.com/Royalvice/DocDiff https://github.com/Royalvice/DocDiff. Zongyuan Yang, Baolin Liu 0002, Yongping Xiong, Lan Yi, Guibin Wu, Junjie Zhou 0001 |
ACM Multimedia | 8 |