EDBT 2026 Demo / reviewers in the wild / expert
Zezhou Cheng
dblp:176/1449
· DBLP profile ↗
17ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0001-7754-0871ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-Free Procedural 3D Shapes are Surprisingly Good TeachersabstractSelf-supervised learning has emerged as a promising approach for acquiring transferable 3D representations from unlabeled 3D point clouds. Unlike 2D images, which are widely accessible, acquiring 3D assets requires specialized expertise or professional 3D scanning equipment, making it difficult to scale and raising copyright concerns. To address these challenges, we propose learning 3D representations from procedural 3D programs that automatically generate 3D shapes using simple 3D primitives and augmentations. Remarkably, despite lacking semantic content, the 3D representations learned from the procedurally generated 3D shapes perform on par with state-of-the-art representations learned from semantically recognizable 3D models (e.g., airplanes) across various downstream 3D tasks, such as shape classification, part segmentation, masked point cloud completion, and both scene semantic and instance segmentation. We provide a detailed analysis on factors that make a good 3D procedural programs. Extensive experiments further suggest that current 3D self-supervised learning methods on point clouds do not rely on semantics of 3D shapes, shedding light on the nature of 3D representations learned. Xuweiyi Chen, Zezhou Cheng |
3DV | 2 |
| 2026 | Open Vocabulary Monocular 3D Object DetectionabstractWe propose and study open-vocabulary monocular 3D detection, a novel task that aims to detect objects of any categories in metric 3D space from a single RGB image. Existing 3D object detectors either rely on costly sensors such as LiDAR or multi-view setups, or remain confined to closed vocabularies settings with limited categories, restricting their applicability. We identify two key challenges in this new setting. First, the scarcity of 3D bounding box annotations limits the ability to train generalizable models. To reduce dependence on 3D supervision, we propose a framework that effectively integrates pretrained 2D and 3D vision foundation models. Second, missing labels and semantic ambiguities (e.g., table vs. desk) in existing datasets hinder reliable evaluation. To address this, we design a novel metric that captures model performance while mitigating annotation issues. Our approach achieves state-of-the-art results in zero-shot 3D detection of novel categories as well as in-domain detection on seen classes. We hope our method provides a strong baseline and our evaluation protocol establishes a reliable benchmark for future research. Xuweiyi Chen, Zezhou Cheng |
3DV | 5 |
| 2025 | Probing the Mid-level Vision Capabilities of Self-Supervised LearningabstractMid-level vision capabilities — such as generic object localization and 3D geometric understanding — are not only fundamental to human vision but are also crucial for many real-world applications of computer vision. These abilities emerge with minimal supervision during the early stages of human visual development. Despite their significance, current self-supervised learning (SSL) approaches are primarily designed and evaluated for high-level recognition tasks, leaving their mid-level vision capabilities largely unexamined.In this study, we introduce a suite of benchmark protocols to systematically assess mid-level vision capabilities and present a comprehensive, controlled evaluation of 22 prominent SSL models across 8 mid-level vision tasks. Our experiments reveal a weak correlation between mid-level and high-level task performance. We also identify several SSL methods with highly imbalanced performance across mid-level and high-level capabilities, as well as some that excel in both. Additionally, we investigate key factors contributing to mid-level vision performance, such as pretraining objectives and network architectures. Our study provides a holistic and timely view of what SSL models have learned, complementing existing research that primarily focuses on high-level vision tasks. We hope our findings guide future SSL research to benchmark models not only on high-level vision tasks but on mid-level as well. Xuweiyi Chen, Markus Marks, Zezhou Cheng |
CVPR | 3 |
| 2025 | Frame In-N-Out: Unbounded Controllable Image-to-Video GenerationabstractControllability, temporal coherence, and detail synthesis remain the most critical challenges in video generation. In this paper, we focus on a commonly used yet underexplored cinematic technique known as Frame In and Frame Out. Specifically, starting from image-to-video generation, users can control the objects in the image to naturally leave the scene or provide breaking new identity references to enter the scene, guided by a user-specified motion trajectory. To support this task, we introduce a new dataset that is curated semi-automatically, an efficient identity-preserving motion-controllable video Diffusion Transformer architecture, and a comprehensive evaluation protocol targeting this task. Our evaluation shows that our proposed approach significantly outperforms existing baselines. Xuweiyi Chen, Matheus Gadelha, Zezhou Cheng |
NeurIPS | 4 |
| 2025 | LabelAny3D: Label Any Object 3D in the WildabstractDetecting objects in 3D space from monocular input is crucial for applications ranging from robotics to scene understanding.
Despite advanced performance in the indoor and autonomous driving domains, existing monocular 3D detection models struggle with in-the-wild images due to the lack of 3D in-the-wild datasets and the challenges of 3D annotation. We introduce LabelAny3D, an analysis-by-synthesis framework that reconstructs holistic 3D scenes from 2D images to efficiently produce high-quality 3D bounding box annotations.
Built on this pipeline, we present COCO3D, a new benchmark for open-vocabulary monocular 3D detection, derived from the MS-COCO dataset and covering a wide range of object categories absent from existing 3D datasets. Experiments show that annotations generated by LabelAny3D improve monocular 3D detection performance across multiple benchmarks, outperforming prior auto-labeling approaches in quality. These results demonstrate the promise of foundation-model-driven annotation for scaling up 3D recognition in realistic, open-world settings. Radowan Mahmud Redoy, Sebastian G. Elbaum, Matthew B. Dwyer, Zezhou Cheng |
NeurIPS | 5 |
| 2024 | Machine Unlearning of Pre-trained Large Language ModelsabstractThis study investigates the concept of the 'right to be forgotten' within the context of large language models (LLMs).We explore machine unlearning as a pivotal solution, with a focus on pre-trained models-a notably under-researched area.Our research delineates a comprehensive framework for machine unlearning in pretrained LLMs, encompassing a critical analysis of seven diverse unlearning methods.Through rigorous evaluation using curated datasets from arXiv, books, and GitHub, we establish a robust benchmark for unlearning performance, demonstrating that these methods are over 10 5 times more computationally efficient than retraining.Our results show that integrating gradient ascent with gradient descent on in-distribution data improves hyperparameter robustness.We also provide detailed guidelines for efficient hyperparameter tuning in the unlearning process.Our findings advance the discourse on ethical AI practices, offering substantive insights into the mechanics of machine unlearning for pretrained LLMs and underscoring the potential for responsible AI development.1 Eli Chien, Minxin Du, Xinyao Niu, Tianhao Wang 0001, Zezhou Cheng, Xiang Yue |
ACL (1) | 6 |
| 2023 | LU-NeRF: Scene and Pose Estimation by Synchronizing Local Unposed NeRFsabstractA critical obstacle preventing NeRF models from being deployed broadly in the wild is their reliance on accurate camera poses. Consequently, there is growing interest in extending NeRF models to jointly optimize camera poses and scene representation, which offers an alternative to off-the-shelf SfM pipelines which have well-understood failure modes. Existing approaches for unposed NeRF operate under limiting assumptions, such as a prior pose distribution or coarse pose initialization, making them less effective in a general setting. In this work, we propose a novel approach, LU-NeRF, that jointly estimates camera poses and neural radiance fields with relaxed assumptions on pose configuration. Our approach operates in a local-to-global manner, where we first optimize over local subsets of the data, dubbed "mini-scenes." LU-NeRF estimates local pose and geometry for this challenging few-shot task. The mini-scene poses are brought into a global reference frame through a robust pose synchronization step, where a final global optimization of pose and scene can be performed. We show our LU-NeRF pipeline outperforms prior attempts at unposed NeRF without making restrictive assumptions on the pose prior. This allows us to operate in the general SE(3) pose setting, unlike the baselines. Our results also indicate our model can be complementary to feature-based SfM pipelines as it compares favorably to COLMAP on low-texture and low-resolution images. Zezhou Cheng, Carlos Esteves, Varun Jampani, Abhishek Kar, Subhransu Maji, Ameesh Makadia |
ICCV | 1 |
| 2022 | GANORCON: Are Generative Models Useful for Few-shot Segmentation?abstractAdvances in generative modeling based on GANs has motivated the community to find their use beyond image generation and editing tasks. In particular, several re-cent works have shown that GAN representations can be re-purposed for discriminative tasks such as part segmen-tation, especially when training data is limited. But how do these improvements stack-up against recent advances in self-supervised learning? Motivated by this we present an alternative approach based on contrastive learning and compare their performance on standard few-shot part seg-mentation benchmarks. Our experiments reveal that not only do the GAN-based approach offer no significant per-formance advantage, their multi-step training is complex, nearly an order-of-magnitude slower, and can introduce ad-ditional bias. These experiments suggest that the inductive biases of generative models, such as their ability to dis-entangle shape and texture, are well captured by standard feed-forward networks trained using contrastive learning. Oindrila Saha, Zezhou Cheng, Subhransu Maji |
CVPR | 2 |
| 2022 | Cross-modal 3D Shape Generation and Manipulation
Zezhou Cheng, Menglei Chai, Jian Ren 0005, Hsin-Ying Lee 0001, Kyle Olszewski, Zeng Huang, Subhransu Maji, Sergey Tulyakov |
ECCV (3) | 1 |
| 2022 | Improving Few-Shot Part Segmentation Using Coarse Supervision
Oindrila Saha, Zezhou Cheng, Subhransu Maji |
ECCV (30) | 2 |
| 2021 | A Realistic Evaluation of Semi-Supervised Learning for Fine-Grained ClassificationabstractWe evaluate the effectiveness of semi-supervised learning (SSL) on a realistic benchmark where data exhibits considerable class imbalance and contains images from novel classes. Our benchmark consists of two fine-grained classification datasets obtained by sampling classes from the Aves and Fungi taxonomy. We find that recently proposed SSL methods provide significant benefits, and can effectively use out-of-class data to improve performance when deep networks are trained from scratch. Yet their performance pales in comparison to a transfer learning baseline, an alternative approach for learning from a few examples. Furthermore, in the transfer setting, while existing SSL methods provide improvements, the presence of out-of-class is often detrimental. In this setting, standard fine-tuning followed by distillation-based self-training is the most robust. Our work suggests that semi-supervised learning with experts on realistic datasets may require different strategies than those currently prevalent in the literature. Jong-Chyi Su, Zezhou Cheng, Subhransu Maji |
CVPR | 2 |
| 2021 | On Equivariant and Invariant Learning of Object Landmark RepresentationsabstractGiven a collection of images, humans are able to discover landmarks by modeling the shared geometric structure across instances. This idea of geometric equivariance has been widely used for the unsupervised discovery of object landmark representations. In this paper, we develop a simple and effective approach by combining instance-discriminative and spatially-discriminative contrastive learning. We show that when a deep network is trained to be invariant to geometric and photometric transformations, representations emerge from its intermediate layers that are highly predictive of object landmarks. Stacking these across layers in a "hypercolumn" and projecting them using spatially-contrastive learning further improves their performance on matching and few-shot landmark regression tasks. We also present a unified view of existing equivariant and invariant representation learning approaches through the lens of contrastive learning, shedding light on the nature of invariances learned. Experiments on standard benchmarks for landmark learning, as well as a new challenging one we propose, show that the proposed approach surpasses prior state-of-the-art. Zezhou Cheng, Jong-Chyi Su, Subhransu Maji |
ICCV | 1 |
| 2020 | Detecting and Tracking Communal Bird Roosts in Weather Radar DataabstractThe US weather radar archive holds detailed information about biological phenomena in the atmosphere over the last 20 years. Communally roosting birds congregate in large numbers at nighttime roosting locations, and their morning exodus from the roost is often visible as a distinctive pattern in radar images. This paper describes a machine learning system to detect and track roost signatures in weather radar data. A significant challenge is that labels were collected opportunistically from previous research studies and there are systematic differences in labeling style. We contribute a latent-variable model and EM algorithm to learn a detection model together with models of labeling styles for individual annotators. By properly accounting for these variations we learn a significantly more accurate detector. The resulting system detects previously unknown roosting locations and provides comprehensive spatio-temporal data about roosts across the US. This data will provide biologists important information about the poorly understood phenomena of broad-scale habitat use and movements of communally roosting birds during the non-breeding season. Zezhou Cheng, Saadia Gabriel, Pankaj Bhambhani, Daniel Sheldon, Subhransu Maji, Andrew Laughlin, David Winkler |
AAAI | 1 |
| 2019 | A Bayesian Perspective on the Deep Image PriorabstractThe deep image prior was recently introduced as a prior for natural images. It represents images as the output of a convolutional network with random inputs. For “inference”, gradient descent is performed to adjust network parameters to make the output match observations. This approach yields good performance on a range of image reconstruction tasks. We show that the deep image prior is asymptotically equivalent to a stationary Gaussian process prior in the limit as the number of channels in each layer of the network goes to infinity, and derive the corresponding kernel. This informs a Bayesian approach to inference. We show that by conducting posterior inference using stochastic gradient Langevin dynamics we avoid the need for early stopping, which is a drawback of the current approach, and improve results for denoising and impainting tasks. We illustrate these intuitions on a number of 1D and 2D signal reconstruction tasks. Zezhou Cheng, Matheus Gadelha, Subhransu Maji, Daniel Sheldon |
CVPR | 1 |
| 2017 | Colorization Using Neural Network EnsembleabstractThis paper investigates into the colorization problem, which converts a grayscale image to a colorful version. This is a difficult problem and normally requires manual adjustment to achieve artifact-free quality. For instance, it normally requires human-labeled color scribbles on the grayscale target image or a careful selection of colorful reference images. The recent learning-based colorization techniques automatically colorize a grayscale image using a single neural network. Since different scenes usually have distinct color styles, it is difficult to accurately capture the color characteristics using a single neural network. We propose a mixture learning model representing the presence of sub-color-style within an overall image data set. We, therefore, ensemble multiple neural networks to obtain better color estimation performance than could be obtained from any of the constituent neural network alone. A two-step colorization strategy is utilized as an adaptive color style clustering followed by a neural network ensemble. To ensure artifact-free quality, a joint bilateral filtering-based post-processing step is proposed. Numerous experiments demonstrate that our method generates high-quality results comparable with state-of-the-art algorithms. Zezhou Cheng, Qingxiong Yang, Bin Sheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Image saliency detection based on rectangular-wave spectrum analysis
Bin Sheng 0001, Wen Wu 0001, Zezhou Cheng, Ruimin Shen |
Multim. Tools Appl. | 4 |
| 2015 | Deep ColorizationabstractThis paper investigates into the colorization problem which converts a grayscale image to a colorful version. This is a very difficult problem and normally requires manual adjustment to achieve artifact-free quality. For instance, it normally requires human-labelled color scribbles on the grayscale target image or a careful selection of colorful reference images (e.g., capturing the same scene in the grayscale target image). Unlike the previous methods, this paper aims at a high-quality fully-automatic colorization method. With the assumption of a perfect patch matching technique, the use of an extremely large-scale reference database (that contains sufficient color images) is the most reliable solution to the colorization problem. However, patch matching noise will increase with respect to the size of the reference database in practice. Inspired by the recent success in deep learning techniques which provide amazing modeling of large-scale data, this paper re-formulates the colorization problem so that deep learning techniques can be directly employed. To ensure artifact-free quality, a joint bilateral filtering based post-processing step is proposed. Numerous experiments demonstrate that our method outperforms the state-of-art algorithms both in terms of quality and speed. Zezhou Cheng, Qingxiong Yang, Bin Sheng 0001 |
ICCV | 1 |