VLDB 2026 Research / reviewers in the wild / expert
Zhipeng Bao
dblp:244/8798
· DBLP profile ↗
13ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Generative modeling · 25% Image recognition and object detection · 24% 3D vision · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 19 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
motion segmentation |
1.2 | 2 | 2023 | Object Discovery from Motion-Guided Tokens · CVPR 2023 Discovering Objects that Can Move · CVPR 2022 |
Computer vision › Image recognition and object detection
object discovery |
1.2 | 2 | 2023 | Object Discovery from Motion-Guided Tokens · CVPR 2023 Discovering Objects that Can Move · CVPR 2022 |
Computer vision › Image recognition and object detection › object discovery
unsupervised object discovery |
1.2 | 2 | 2023 | Object Discovery from Motion-Guided Tokens · CVPR 2023 Discovering Objects that Can Move · CVPR 2022 |
Computer vision › 3D vision
novel view synthesis |
1.2 | 2 | 2023 | Multi-task View Synthesis with Neural Radiance Fields · ICCV 2023 Bowtie Networks: Generative Modeling for Joint Few-Shot Recognition and Novel-View Synthesis · ICLR 2021 |
Machine learning › Generative modeling › diffusion model
diffusion-based perception |
0.9 | 1 | 2025 | Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models · ICLR 2025 |
Machine learning › Generative modeling
multimodal generation |
0.9 | 1 | 2025 | Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models · ICLR 2025 |
Computer vision › 3D vision
3d scene understanding |
0.8 | 1 | 2024 | Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning › vision foundation model
visual foundation model probing |
0.8 | 1 | 2024 | Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding · NeurIPS 2024 |
Machine learning › Learning paradigms › multi-task learning
multi-task scene understanding |
0.7 | 1 | 2023 | Multi-task View Synthesis with Neural Radiance Fields · ICCV 2023 |
Computer vision › 3D vision
neural radiance field |
0.7 | 1 | 2023 | Multi-task View Synthesis with Neural Radiance Fields · ICCV 2023 |
Machine learning › Generative modeling › generative model
multi-task generative modeling |
0.6 | 1 | 2022 | Generative Modeling for Multi-task Visual Learning · ICML 2022 |
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot image classification |
0.5 | 1 | 2021 | Bowtie Networks: Generative Modeling for Joint Few-Shot Recognition and Novel-View Synthesis · ICLR 2021 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.5 | 1 | 2021 | Bowtie Networks: Generative Modeling for Joint Few-Shot Recognition and Novel-View Synthesis · ICLR 2021 |
Bioinformatics and computational biology › structural biology
electron tomography |
0.4 | 1 | 2019 | A joint method for marker-free alignment of tilt series in electron tomography · Bioinform. 2019 |
Computer vision › Segmentation and scene understanding › image segmentation
scene segmentation |
0.2 | 1 | 2024 | Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
vector quantization |
0.2 | 1 | 2023 | Object Discovery from Motion-Guided Tokens · CVPR 2023 |
Machine learning › Representation and self-supervised learning › representation learning › neural network representation learning › deep representation learning
autoencoder representation learning |
0.2 | 1 | 2022 | Discovering Objects that Can Move · CVPR 2022 |
Bioinformatics and computational biology
structural biology |
0.1 | 1 | 2019 | A joint method for marker-free alignment of tilt series in electron tomography · Bioinform. 2019 |
Methods — techniques the papers use, named apart from their topics
autoencoder · 1.2self-improving learning · 0.9diffusion model · 0.9data augmentation · 0.9mixture-of-vision-expert · 0.8transformer decoder · 0.7motion-guided vector quantization · 0.7cross-view attention · 0.7cross-task attention · 0.7motion segmentation · 0.6landmark tracking · 0.4joint optimization · 0.4intensity-based alignment · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large Language Model-assisted Autonomous Vehicle Recovery from Immobilization
Zhipeng Bao, Qianwen Li |
IV | 1 |
| 2025 | ReferevErything: Towards Segmenting Everything we can Speak of in VideosabstractWe present REM, a framework for segmenting a wide range of concepts in video that can be described through natural language. Our method leverages the universal visual-language mapping learned by video diffusion models on Internet-scale data by fine-tuning them on small-scale Referring Object Segmentation datasets. Our key insight is to preserve the entirety of the generative model's architecture by shifting its objective from predicting noise to predicting mask latents. The resulting model can accurately segment rare and unseen objects, despite only being trained on a limited set of categories. Additionally, it can effortlessly generalize to non-object dynamic concepts, such as smoke or raindrops, as demonstrated in our new benchmark for Referring Video Process Segmentation (Ref-VPS). REM performs on par with the state-of-the-art on in-domain datasets, like Ref-DAVIS, while outperforming them by up to 12 IoU points out-of-domain, leveraging the power of generative pre-training. We also show that advancements in video generation directly improve segmentation. Anurag Bagchi, Zhipeng Bao, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert |
ICCV | 2 |
| 2025 | Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion ModelsabstractBeyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, most existing work treats diffusion models as a standalone component for perception tasks, employing them either solely for off-the-shelf data augmentation or as mere feature extractors. In contrast to these isolated and thus sub-optimal efforts, we introduce an integrated, versatile, diffusion-based framework, Diff-2-in-1, that can simultaneously handle both multi-modal data generation and dense visual perception, through a unique exploitation of the diffusion-denoising process. Within this framework, we further enhance discriminative visual perception via multi-modal generation, by utilizing the denoising network to create multi-modal data that mirror the distribution of the original training set. Importantly, Diff-2-in-1 optimizes the utilization of the created diverse and faithful data by leveraging a novel self-improving learning mechanism. Comprehensive experimental evaluations validate the effectiveness of our framework, showcasing consistent performance improvements across various discriminative backbones and high-quality multi-modal data generation characterized by both realism and usefulness. Our project website is available at https://zsh2000.github.io/diff-2-in-1.github.io/. Shuhong Zheng, Zhipeng Bao, Martial Hebert, Yu-Xiong Wang |
ICLR | 2 |
| 2024 | Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene UnderstandingabstractComplex 3D scene understanding has gained increasing attention, with scene encoding strategies built on top of visual foundation models playing a crucial role in this success. However, the optimal scene encoding strategies for various scenarios remain unclear, particularly compared to their image-based counterparts. To address this issue, we present the first comprehensive study that probes various visual encoding models for 3D scene understanding, identifying the strengths and limitations of each model across different scenarios. Our evaluation spans seven vision foundation encoders, including image, video, and 3D foundation models. We evaluate these models in four tasks: Vision-Language Scene Reasoning, Visual Grounding, Segmentation, and Registration, each focusing on different aspects of scene understanding. Our evaluation yields key intriguing findings: Unsupervised image foundation models demonstrate superior overall performance, video models excel in object-level tasks, diffusion models benefit geometric tasks, language-pretrained models show unexpected limitations in language-related tasks, and the mixture-of-vision-expert (MoVE) strategy leads to consistent performance improvement. These insights challenge some conventional understandings, provide novel perspectives on leveraging visual foundation models, and highlight the need for more flexible encoder selection in future vision-language and scene understanding tasks. Yunze Man, Shuhong Zheng, Zhipeng Bao, Martial Hebert, Liangyan Gui, Yu-Xiong Wang |
NeurIPS | 3 |
| 2023 | Object Discovery from Motion-Guided TokensabstractObject discovery – separating objects from the background without manual labels – is a fundamental open challenge in computer vision. Previous methods struggle to go beyond clustering of low-level cues, whether handcrafted (e.g., color, texture) or learned (e.g., from auto-encoders). In this work, we augment the auto-encoder representation learning framework with two key components: motion-guidance and mid-level feature tokenization. Although both have been separately investigated, we introduce a new transformer decoder showing that their benefits can compound thanks to motion-guided vector quantization. We show that our architecture effectively leverages the synergy between motion and tokenization, improving upon the state of the art on both synthetic and real datasets. Our approach enables the emergence of interpretable object-specific mid-level features, demonstrating the benefits of motion-guidance (no labeling) and quantization (interpretability, memory efficiency). Zhipeng Bao, Pavel Tokmakov, Yu-Xiong Wang, Adrien Gaidon, Martial Hebert |
CVPR | 1 |
| 2023 | Multi-task View Synthesis with Neural Radiance FieldsabstractMulti-task visual learning is a critical aspect of computer vision. Current research, however, predominantly concentrates on the multi-task dense prediction setting, which overlooks the intrinsic 3D world and its multi-view consistent structures, and lacks the capability for versatile imagination. In response to these limitations, we present a novel problem setting – multi-task view synthesis (MTVS), which reinterprets multi-task prediction as a set of novel-view synthesis tasks for multiple scene properties, including RGB. To tackle the MTVS problem, we propose MuvieNeRF, a framework that incorporates both multi-task and cross-view knowledge to simultaneously synthesize multiple scene properties. MuvieNeRF integrates two key modules, the Cross-Task Attention (CTA) and Cross-View Attention (CVA) modules, enabling the efficient use of information across multiple views and tasks. Extensive evaluation on both synthetic and realistic benchmarks demonstrates that MuvieNeRF is capable of simultaneously synthesizing different scene properties with promising visual quality, even outperforming conventional discriminative models in various settings. Notably, we show that MuvieNeRF exhibits universal applicability across a range of NeRF backbones. Our code is available at https://github.com/zsh2000/MuvieNeRF. Shuhong Zheng, Zhipeng Bao, Martial Hebert, Yu-Xiong Wang |
ICCV | 2 |
| 2023 | Glow Model-Based Latent Vector Optimization for Generative Image Steganography in Edge and Cloud Computing EnvironmentabstractIn edge and cloud computing environments, to protect and manage secret information, the sharing and recovery of each secret image are implemented by local servers. However, since existing Generative Image Steganography(GIS) schemes face issues such as low-quality image generation and small hiding capacity. This necessitates the generation of a large number of stego-images to meet the demands of information transmission, thereby imposing a excessive computational burden for those local servers, the above reasons make the existing GIS schemes not suitable for edge and cloud computing environments. To address the above issue, we propose a Latent Vector Optimization(LVO) scheme for GIS with high-quality image generation and large hiding capacity. In the proposed scheme, we introduce the concept of latent vector optimization, wherein the hiding probability of each element within the latent vector is computed based on its expected influence on the quality of the resulting stego-image. Furthermore, our LVO scheme employs an adaptive approach to identify the optimal locations for embedding information while considering a predefined hiding capacity. This adaptation involves giving priority to modifying elements in dimensions characterized by a low latent vector hiding probability, as guided by the characteristics of natural images. Simultaneously, the scheme hides the secret message within elements associated with a high latent vector hiding probability, thus achieving a large hiding capacity while minimizing any adverse effects on the stego-image quality. Compared with the existing GIS schemes, the proposed LVO scheme enhances security, provides high-quality image generation, a large data hiding capacity, meeting the requirements with fewer stego-images. This significantly reduces the communication and computational burden on local servers. Zhipeng Bao, Zhili Zhou 0001, Xutong Cui, Chengsheng Yuan 0001 |
ICPADS | 1 |
| 2023 | Beyond RGB: Scene-Property Synthesis with Neural Radiance FieldsabstractComprehensive 3D scene understanding, both geometrically and semantically, is important for real-world applications such as robot perception. Most of the existing work has focused on developing data-driven discriminative models for scene understanding. This paper provides a new approach to scene understanding, from a synthesis model perspective, by leveraging the recent progress on implicit scene representation and neural rendering. Building upon the great success of Neural Radiance Fields (NeRFs), we introduce Scene-Property Synthesis with NeRF (SS-NeRF) that is able to not only render photo-realistic RGB images from novel viewpoints, but also render various accurate scene properties (e.g., appearance, geometry, and semantics). By doing so, we facilitate addressing a variety of scene understanding tasks under a unified framework, including semantic segmentation, surface normal estimation, reshading, keypoint detection, and edge detection. Our SS-NeRF framework can be a powerful tool for bridging generative learning and discriminative learning, and thus be beneficial to the investigation of a wide range of interesting problems, such as studying task relationships within a synthesis paradigm, transferring knowledge to novel tasks, facilitating downstream discriminative tasks as ways of data augmentation, and serving as auto-labeller for data creation. Our code is available at https://github.com/zsh2000/SS-NeRF. Shuhong Zheng, Zhipeng Bao, Martial Hebert, Yu-Xiong Wang |
WACV | 3 |
| 2022 | Discovering Objects that Can MoveabstractThis paper studies the problem of object discovery - separating objects from the background without manual labels. Existing approaches utilize appearance cues, such as color, texture, and location, to group pixels into object-like regions. However, by relying on appearance alone, these methods fail to separate objects from the background in cluttered scenes. This is a fundamental limitation since the definition of an object is inherently ambiguous and context-dependent. To resolve this ambiguity, we choose to focus on dynamic objects - entities that can move independently in the world. We then scale the recent auto-encoder based frameworks for unsuper-vised object discovery from toy synthetic images to complex real-world scenes. To this end, we simplify their architecture, and augment the resulting model with a weak learning signal from general motion segmentation algorithms. Our experiments demonstrate that, despite only capturing a small subset of the objects that move, this signal is enough to generalize to segment both moving and static instances of dynamic objects. We show that our model scales to a newly collected, photo- realistic synthetic dataset with street driving scenarios. Additionally, we leverage ground truth segmentation and flow annotations in this dataset for thorough ablation and evaluation. Finally, our experiments on the real-world KITTI benchmark demonstrate that the proposed approach outperforms both heuristic- and learning-based methods by capitalizing on motion cues. Zhipeng Bao, Pavel Tokmakov, Allan Jabri, Yu-Xiong Wang, Adrien Gaidon, Martial Hebert |
CVPR | 1 |
| 2022 | Generative Modeling for Multi-task Visual LearningabstractGenerative modeling has recently shown great promise in computer vision, but it has mostly focused on synthesizing visually realistic images. In this paper, motivated by multi-task learning of shareable feature representations, we consider a novel problem of learning a shared generative model that is useful across various visual perception tasks. Correspondingly, we propose a general multi-task oriented generative modeling (MGM) framework, by coupling a discriminative multi-task network with a generative network. While it is challenging to synthesize both RGB images and pixel-level annotations in multi-task scenarios, our framework enables us to use synthesized images paired with only weak annotations (i.e., image-level scene labels) to facilitate multiple visual tasks. Experimental evaluation on challenging multi-task benchmarks, including NYUv2 and Taskonomy, demonstrates that our MGM framework improves the performance of all the tasks by large margins, consistently outperforming state-of-the-art multi-task approaches in different sample-size regimes. Zhipeng Bao, Martial Hebert, Yu-Xiong Wang |
ICML | 1 |
| 2021 | Bowtie Networks: Generative Modeling for Joint Few-Shot Recognition and Novel-View Synthesis
Zhipeng Bao, Yu-Xiong Wang, Martial Hebert |
ICLR | 1 |
| 2019 | Deep Learning-Based Strategy For Macromolecules Classification with Imbalanced Data from Cellular Electron CryotomographyabstractDeep learning model trained by imbalanced data may not work satisfactorily since it could be determined by major classes and thus may ignore the classes with small amount of data. In this paper, we apply deep learning based imbalanced data classification for the first time to cellular macromolecular complexes captured by Cryo-electron tomography (Cryo-ET). We adopt a range of strategies to cope with imbalanced data, including data sampling, bagging, boosting, Genetic Programming based method and. Particularly, inspired from Inception 3D network, we propose a multi-path CNN model combining focal loss and mixup on the Cryo-ET dataset to expand the dataset, where each path had its best performance corresponding to each type of data and let the network learn the combinations of the paths to improve the classification performance. In addition, extensive experiments have been conducted to show our proposed method is flexible enough to cope with different number of classes by adjusting the number of paths in our multi-path model. To our knowledge, this work is the first application of deep learning methods of dealing with imbalanced data to the internal tissue classification of cell macromolecular complexes, which opened up a new path for cell classification in the field of computational biology. Ziqian Luo, Zhipeng Bao, Min Xu 0009 |
IJCNN | 3 |
| 2019 | A joint method for marker-free alignment of tilt series in electron tomographyabstractMOTIVATION: Electron tomography (ET) is a widely used technology for 3D macro-molecular structure reconstruction. To obtain a satisfiable tomogram reconstruction, several key processes are involved, one of which is the calibration of projection parameters of the tilt series. Although fiducial marker-based alignment for tilt series has been well studied, marker-free alignment remains a challenge, which requires identifying and tracking the identical objects (landmarks) through different projections. However, the tracking of these landmarks is usually affected by the pixel density (intensity) change caused by the geometry difference in different views. The tracked landmarks will be used to determine the projection parameters. Meanwhile, different projection parameters will also affect the localization of landmarks. Currently, there is no alignment method that takes interrelationship between the projection parameters and the landmarks. RESULTS: Here, we propose a novel, joint method for marker-free alignment of tilt series in ET, by utilizing the information underlying the interrelationship between the projection model and the landmarks. The proposed method is the first joint solution that combines the extrinsic (track-based) alignment and the intrinsic (intensity-based) alignment, in which the localization of landmarks and projection parameters keep refining each other until convergence. This iterative approach makes our solution robust to different initial parameters and extreme geometric changes, which ensures a better reconstruction for marker-free ET. Comprehensive experimental results on three real datasets show that our new method achieved a significant improvement in alignment accuracy and reconstruction quality, compared to the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://github.com/icthrm/joint-marker-free-alignment. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Renmin Han, Zhipeng Bao, Tongxin Niu, Fa Zhang 0001, Min Xu 0009, Xin Gao 0001 |
Bioinform. | 2 |