VLDB 2026 Research / reviewers in the wild / expert
Zuyan Liu
dblp:299/9349
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0009-0002-6943-3085ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Point2Seq: Quantized Serialization Encoding for Object Point Cloud Pretraining
Xumin Yu, Zuyan Liu, Jie Zhou 0001, Jiwen Lu |
Int. J. Comput. Vis. | 2 |
| 2026 | Efficient High-Order Spatial Interactions for Visual PerceptionabstractRecent progress in vision Transformers exhibits great success in various tasks driven by the new spatial modeling mechanism based on dot-product self-attention. In this paper, we show that the key ingredients behind the vision Transformers, namely input-adaptive, long-range and high-order spatial interactions, can also be efficiently implemented with a convolution-based framework. We present the Recursive Gated Convolution (${\mathit{g}}^{\mathit{n}}$gnConv) that performs high-order spatial interactions with gated convolutions and recursive designs. The new operation is highly flexible and customizable, which is compatible with various variants of convolution and extends the two-order interactions in self-attention to arbitrary orders without introducing significant extra computation. ${\mathit{g}}^{\mathit{n}}$gn Conv can serve as a plug-and-play module to improve various vision Transformers and convolution-based models. Based on the proposed operation, we construct a new family of generic vision backbones for various visual modalities and tasks, including HorNet and HorFPN for image recognition, Hor3D for point cloud analysis, and HorCLIP for vision-language modeling. For image recognition, we propose HorNet as a stronger visual encoder, where we conduct extensive experiments on ImageNet classification, COCO object detection, and ADE20K semantic segmentation. HorNet outperforms Swin Transformers and ConvNeXt by a significant margin with similar overall architecture and training configurations. HorNet also shows favorable scalability to more training data and larger model sizes. Apart from image encoders, we also show ${\mathit{g}}^{\mathit{n}}$gnConv can be applied to task-specific decoders and consistently improve dense prediction performance with less computation. For point cloud analysis, we design Hor3D, demonstrating the efficacy of high-order interactions for unstructured point cloud data through experiments on challenging 3D semantic segmentation tasks in S3DIS and ScanNet V2. In vision-language modeling, our proposed HorCLIP surpasses mainstream Vision Transformer and ConvNeXt architectures with shorter training schedules on ImageNet zero-shot classification and shows remarkably higher performance on vision-language dense representation tasks on COCO Panoptic datasets. Our results demonstrate that ${\mathit{g}}^{\mathit{n}}$gnConv with high-order spatial interactions can be a new basic operation for visual modeling that effectively combines the merits of both vision Transformers and CNNs. Zuyan Liu, Yongming Rao, Wenliang Zhao, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | ProtoComp++: Diverse Point Cloud Completion With Controllable PrototypeabstractPoint cloud completion aims to reconstruct the geometry of partial point clouds captured by various sensors. Traditionally, point cloud models are trained on synthetic datasets that feature limited categories and differ significantly from real-world scenarios. This gap often causes existing methods to struggle when faced with unfamiliar categories and severe incompleteness in real-world applications. In this paper, we propose PrototypeCompletion, a novel prototype-based approach for point cloud completion. The method begins by generating rough prototypes, which are then refined with additional geometric details to make the final prediction. We introduce two distinct approaches for integrating prototypes into the network: explicit prototypes and implicit prototypes. Our approach demonstrates strong generalization capabilities, allowing it to handle point cloud completion for a variety of unseen categories beyond the training data. We demonstrate that incorporating language prompts into the training of point cloud completion models significantly expands their applicability and enhances their performance in diverse point cloud completion tasks. Furthermore, we propose a new evaluation metric and a test benchmark based on ScanNet200 and KITTI, designed to assess the model's performance in real-world scenarios and foster future research in the field. Experimental results show that our method outperforms state-of-the-art models on the existing PCN and ShapeNet34 benchmarks and also excels in various real-world settings, handling different object categories and sensor types effectively. The code will be made publicly available. Xumin Yu, Zuyan Liu, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language ModelsabstractLarge Language Models (LLMs) demonstrate enhanced capabilities and reliability by reasoning more, evolving from Chain-of-Thought prompting to product-level solutions like OpenAI o1. Despite various efforts to improve LLM reasoning, high-quality long-chain reasoning data and optimized training pipelines still remain inadequately explored in vision-language tasks. In this paper, we present InsightV, an early effort to 1) scalably produce long and robust reasoning data for complex multi-modal tasks, and 2) an effective training pipeline to enhance the reasoning capabilities of multi-modal large language models (MLLMs). Specifically, to create long and structured reasoning data without human labor, we design a two-step pipeline with a progressive strategy to generate sufficiently long and diverse reasoning paths and a multi-granularity assessment method to ensure data quality. We observe that directly supervising MLLMs with such long and complex reasoning data will not yield ideal reasoning ability. To tackle this problem, we design a multi-agent system consisting of a reasoning agent dedicated to performing long-chain reasoning and a summary agent trained to judge and summarize reasoning results. We further incorporate an iterative DPO algorithm to enhance the reasoning agent’s generation stability and quality. Based on the popular LLaVA-NeXT model and our stronger base MLLM, we demonstrate significant performance gains across challenging multi-modal benchmarks requiring visual reasoning. Benefiting from our multi-agent system, Insight-V can also easily maintain or improve performance on perception-focused multi-modal tasks. Yuhao Dong, Zuyan Liu, Hai-Long Sun, Winston Hu, Yongming Rao, Ziwei Liu 0002 |
CVPR | 2 |
| 2025 | SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMsabstractMultimodal Large Language Models (MLLMs) are commonly derived by extending pre-trained Large Language Models (LLMs) with visual capabilities. In this work, we investigate how MLLMs process visual inputs by analyzing their attention mechanisms. We reveal a surprising sparsity phenomenon: only a small subset (approximately less than 5%) of attention heads in LLMs actively contribute to visual understanding, termed visual heads. To identify these heads efficiently, we design a training-free framework that quantifies head-level visual relevance through targeted response analysis. Building on this discovery, we introduce SparseMM, a KV-Cache optimization strategy that allocates asymmetric computation budgets to heads in LLMs based on their visual scores, leveraging the sparity of visual heads for accelerating the inference of MLLMs. Compared with prior KV-Cache acceleration methods that ignore the particularity of visual, SparseMM prioritizes stress and retaining visual semantics during decoding. Extensive evaluations across mainstream multimodal benchmarks demonstrate that SparseMM achieves superior accuracy-efficiency trade-offs. Notably, SparseMM delivers 1.38x real-time acceleration and 52% memory reduction during generation while maintaining performance parity on efficiency test. Our project is open sourced at https://github.com/CR400AF-A/SparseMM. Jiahui Wang 0001, Zuyan Liu, Yongming Rao, Jiwen Lu |
ICCV | 2 |
| 2025 | Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary ResolutionabstractVisual data comes in various forms, ranging from small icons of just a few pixels to long videos spanning hours. Existing multi-modal LLMs usually standardize these diverse visual inputs to fixed-resolution images or patches for visual encoders and yield similar numbers of tokens for LLMs. This approach is non-optimal for multimodal understanding and inefficient for processing inputs with long and short visual contents. To solve the problem, we propose Oryx, a unified multimodal architecture for the spatial-temporal understanding of images, videos, and multi-view 3D scenes. Oryx offers an on-demand solution to seamlessly and efficiently process visual inputs with arbitrary spatial sizes and temporal lengths through two core innovations: 1) a pre-trained OryxViT model that can encode images at any resolution into LLM-friendly visual representations; 2) a dynamic compressor module that supports 1x to 16x compression on visual tokens by request. These designs enable Oryx to accommodate extremely long visual contexts, such as videos, with lower resolution and high compression while maintaining high recognition precision for tasks like document understanding with native resolution and no compression. Beyond the architectural improvements, enhanced data curation and specialized training on long-context retrieval and spatial-aware data help Oryx achieve strong capabilities in image, video, and 3D multimodal understanding simultaneously. Zuyan Liu, Yuhao Dong, Ziwei Liu 0002, Winston Hu, Jiwen Lu, Yongming Rao |
ICLR | 1 |
| 2025 | Vision Generalist Model: A Survey
Ziyi Wang 0007, Yongming Rao, Shuofeng Sun, Xinrun Liu, Yi Wei 0003, Xumin Yu, Zuyan Liu, Hongmin Liu 0001, Jie Zhou 0001, Jiwen Lu |
Int. J. Comput. Vis. | 7 |
| 2024 | Efficient Inference of Vision Instruction-Following Models with Elastic Cache
Zuyan Liu, Benlin Liu, Jiahui Wang 0001, Yuhao Dong, Guangyi Chen 0002, Yongming Rao, Ranjay Krishna, Jiwen Lu |
ECCV (17) | 1 |
| 2023 | DiffSwap: High-Fidelity and Controllable Face Swapping via 3D-Aware Masked DiffusionabstractIn this paper, we propose DiffSwap, a diffusion model based framework for high-fidelity and controllable face swapping. Unlike previous work that relies on carefully designed network architectures and loss functions to fuse the information from the source and target faces, we reformulate the face swapping as a conditional inpainting task, performed by a powerful diffusion model guided by the desired face attributes (e.g., identity and landmarks). An important issue that makes it nontrivial to apply diffusion models to face swapping is that we cannot perform the time-consuming multi-step sampling to obtain the generated image during training. To overcome this, we propose a mid-point estimation method to efficiently recover a reasonable diffusion result of the swapped face with only 2 steps, which enables us to introduce identity constraints to improve the face swapping quality. Our framework enjoys several favorable properties more appealing than prior arts: 1) Controllable. Our method is based on conditional masked diffusion on the latent space, where the mask and the conditions can be fully controlled and customized. 2) High-fidelity. The formulation of conditional inpainting can fully exploit the generative ability of diffusion models and can preserve the background of target images with minimal artifacts. 3) Shape-preserving. The controllability of our method enables us to use 3D-aware landmarks as the condition during generation to preserve the shape of the source face. Extensive experiments on both FF++ and FFHQ demonstrate that our method can achieve state-of-the-art face swapping results both qualitatively and quantitatively. Wenliang Zhao, Yongming Rao, Weikang Shi, Zuyan Liu, Jie Zhou 0001, Jiwen Lu |
CVPR | 4 |
| 2023 | Unleashing Text-to-Image Diffusion Models for Visual PerceptionabstractDiffusion models (DMs) have become the new trend of generative models and have demonstrated a powerful ability of conditional synthesis. Among those, text-to-image diffusion models pre-trained on large-scale image-text pairs are highly controllable by customizable prompts. Unlike the unconditional generative models that focus on low-level attributes and details, text-to-image diffusion models contain more high-level knowledge thanks to the vision-language pre-training. In this paper, we propose VPD (Visual Perception with pre-trained Diffusion models), a new framework that exploits the semantic information of a pre-trained text-to-image diffusion model in visual perception tasks. Instead of using the pre-trained denoising autoencoder in a diffusion-based pipeline, we simply use it as a backbone and aim to study how to take full advantage of the learned knowledge. Specifically, we prompt the denoising decoder with proper textual inputs and refine the text features with an adapter, leading to a better alignment to the pre-trained stage and making the visual contents interact with the text prompts. We also propose to utilize the cross-attention maps between the visual features and the text features to provide explicit guidance. Compared with other pre-training methods, we show that vision-language pre-trained diffusion models can be faster adapted to downstream visual perception tasks using the proposed VPD. Extensive experiments on semantic segmentation, referring image segmentation, and depth estimation demonstrate the effectiveness of our method. Notably, VPD attains 0.254 RMSE on NYUv2 depth estimation and 73.3% oIoU on RefCOCO-val referring image segmentation, establishing new records on these two benchmarks. Code is available at https://github.com/wl-zhao/VPD. Wenliang Zhao, Yongming Rao, Zuyan Liu, Benlin Liu, Jie Zhou 0001, Jiwen Lu |
ICCV | 3 |
| 2023 | Dynamic Spatial Sparsification for Efficient Vision Transformers and Convolutional Neural NetworksabstractIn this paper, we present a new approach for model acceleration by exploiting spatial sparsity in visual data. We observe that the final prediction in vision Transformers is only based on a subset of the most informative regions, which is sufficient for accurate image recognition. Based on this observation, we propose a dynamic token sparsification framework to prune redundant tokens progressively and dynamically based on the input to accelerate vision Transformers. Specifically, we devise a lightweight prediction module to estimate the importance of each token given the current features. The module is added to different layers to prune redundant tokens hierarchically. While the framework is inspired by our observation of the sparse attention in vision Transformers, we find that the idea of adaptive and asymmetric computation can be a general solution for accelerating various architectures. We extend our method to hierarchical models including CNNs and hierarchical vision Transformers as well as more complex dense prediction tasks. To handle structured feature maps, we formulate a generic dynamic spatial sparsification framework with progressive sparsification and asymmetric computation for different spatial locations. By applying lightweight fast paths to less informative features and expressive slow paths to important locations, we can maintain the complete structure of feature maps while significantly reducing the overall computations. Extensive experiments on diverse modern architectures and different visual tasks demonstrate the effectiveness of our proposed framework. By hierarchically pruning 66% of the input tokens, our method greatly reduces 31% ∼ 35% FLOPs and improves the throughput by over 40% while the drop of accuracy is within 0.5% for various vision Transformers. By introducing asymmetric computation, a similar acceleration can be achieved on modern CNNs and Swin Transformers. Moreover, our method achieves promising results on more complex tasks including semantic segmentation and object detection. Our results clearly demonstrate that dynamic spatial sparsification offers a new and more effective dimension for model acceleration. Code is available at https://github.com/raoyongming/DynamicViT. Yongming Rao, Zuyan Liu, Wenliang Zhao, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | PoinTr: Diverse Point Cloud Completion with Geometry-Aware TransformersabstractPoint clouds captured in real-world applications are of-ten incomplete due to the limited sensor resolution, single viewpoint, and occlusion. Therefore, recovering the complete point clouds from partial ones becomes an indispensable task in many practical applications. In this paper, we present a new method that reformulates point cloud completion as a set-to-set translation problem and design a new model, called PoinTr that adopts a transformer encoder-decoder architecture for point cloud completion. By rep-resenting the point cloud as a set of unordered groups of points with position embeddings, we convert the point cloud to a sequence of point proxies and employ the transformers for point cloud generation. To facilitate transformers to better leverage the inductive bias about 3D geometric structures of point clouds, we further devise a geometry-aware block that models the local geometric relationships explicitly. The migration of transformers enables our model to better learn structural knowledge and preserve detailed information for point cloud completion. Furthermore, we propose two more challenging benchmarks with more diverse incomplete point clouds that can better reflect the real-world scenarios to promote future research. Experimental results show that our method outperforms state-of-the-art methods by a large margin on both the new bench-marks and the existing ones. Code is available at https://github.com/yuxumin/PoinTr. Xumin Yu, Yongming Rao, Ziyi Wang 0007, Zuyan Liu, Jiwen Lu, Jie Zhou 0001 |
ICCV | 4 |