Ziyao Zeng

dblp:306/8074 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2025
0009-0003-1803-7434ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ProtoDepth: Unsupervised Continual Depth Completion with Prototypes
abstract
We present ProtoDepth, a novel prototype-based approach for continual learning of unsupervised depth completion, the multimodal 3D reconstruction task of predicting dense depth maps from RGB images and sparse point clouds. The unsupervised learning paradigm is well-suited for continual learning, as ground truth is not needed. However, when training on new non-stationary distributions, depth completion models will catastrophically forget previously learned information. We address forgetting by learning prototype sets that adapt the latent features of a frozen pretrained model to new domains. Since the original weights are not modified, ProtoDepth does not forget when test-time domain identity is known. To extend ProtoDepth to the challenging setting where the test-time domain identity is withheld, we propose to learn domain descriptors that enable the model to select the appropriate prototype set for inference. We evaluate ProtoDepth on benchmark dataset sequences, where we reduce forgetting compared to baselines by 52.2% for indoor and 53.2% for outdoor to achieve the state of the art. Project Page: https://protodepth.github.io/.
Patrick Rim, Hyoungseob Park, Suchisrit Gangopadhyay, Ziyao Zeng, Younjoon Chung, Alex Wong 0001
CVPR4
2025 ETA: Energy-Based Test-Time Adaptation for Depth Completion
abstract
We propose a method for test-time adaptation of pretrained depth completion models. Depth completion models, trained on some ``source'' data, often predict erroneous outputs when transferred to ``target'' data captured in novel environmental conditions due to a covariate shift. The crux of our method lies in quantifying the likelihood of depth predictions belonging to the source data distribution. The challenge is in the lack of access to out-of-distribution (target) data prior to deployment. Hence, rather than making assumptions regarding the target distribution, we utilize adversarial perturbations as a mechanism to explore the data space. This enables us to train an energy model that scores local regions of depth predictions as in- or out-of-distribution. We update the parameters of pretrained depth completion models at test time to minimize energy, effectively aligning test-time predictions to those of the source distribution. We call our method ``Energy-based Test-time Adaptation'', or ETA for short. We evaluate our method across three indoor and three outdoor datasets, where ETA improve over the previous state-of-the-art method by an average of 6.94% for outdoors and 10.23% for indoors. Project Page: https://fuzzythecat.github.io/eta.
Younjoon Chung, Hyoungseob Park, Patrick Rim, Jihe He, Ziyao Zeng, Safa Cicek, Byung-Woo Hong, James S. Duncan, Alex Wong 0001
ICCV6
2025 Bug numbers matter: An empirical study of effort-aware defect prediction using class labels versus bug numbers
abstract
Abstract Previous research have utilized public software defect datasets such as NASA, RELINK, and SOFTLAB, which only contain class label information. Most effort‐aware defect prediction (EADP) studies are carried out around these datasets. However, EADP studies typically relying on predicted bug number (i.e., considering modules as effort) or density (i.e., considering lines of code as effort) for ranking software modules. To explore the impact of bug number information in constructing EADP models, we access the performance degradation of the best‐performing learning‐to‐rank methods when using class labels instead of bug numbers for training. The experimental results show that using class labels instead of bug numbers in building EADP models results in an decrease in the detected bugs when module is considering as effort. When effort is LOC, using class labels to construct EADP models can lead to a significant increase in the initial false alarms and a significant increase in the modules that need to be inspected. Therefore, we recommend not only the class labels but also the bug number information should be disclosed when publishing software defect datasets, in order to construct more accurate EADP models.
Peixin Yang, Ziyao Zeng, Yanjiao Zhang, Xin Wang 0114, Chuanxiang Ma
Softw. Pract. Exp.2
2024 Binding Touch to Everything: Learning Unified Multimodal Tactile Representations
abstract
Touch provides crucial information about the physical properties of the objects around us. Creating models that capture cross-modal associations between touch and other modalities, however, remains a challenging problem, due to wide variety of touch sensors and the intensive effort required to collect tactile data. We propose UniTouch, a unified model for vision-based touch sensors that connects their tactile signals to other modalities, including vision, language, and sound. We achieve this by aligning our tactile embeddings to pretrained image embeddings already associated with a variety of other modalities. We further propose learnable sensor-specific tokens, allowing the model to learn from a set of heterogeneous tactile sensors, all at the same time. UniTouch is capable of conducting various touch sensing tasks in a zero-shot setting, from robot grasping prediction to touch-based question answering. To the best of our knowledge, UniTouch is the first model to demonstrate these capabilities. Project Page: https://cfeng16.github.io/UniTouch/.
Fengyu Yang 0003, Hyoungseob Park, Daniel Wang 0005, Yiming Dou, Ziyao Zeng, Xien Chen, Rit Gangopadhyay, Andrew Owens, Alex Wong 0001
CVPR7
2024 WorDepth: Variational Language Prior for Monocular Depth Estimation
abstract
Three-dimensional (3D) reconstruction from a single image is an ill-posed problem with inherent ambiguities, i. e. scale. Predicting a 3D scene from text descriptionis) is similarly ill-posed, i. e. spatial arrangements of objects described. We investigate the question of whether two inher-ently ambiguous modalities can be used in conjunction to produce metric-scaled reconstructions. To test this, we fo-cus on monocular depth estimation, the problem of predicting a dense depth map from a single image, but with an additional text caption describing the scene. To this end, we begin by encoding the text caption as a mean and standard deviation; using a variational framework, we learn the distribution of the plausible metric reconstructions of 3D scenes corresponding to the text captions as a prior. To “select” a specific reconstruction or depth map, we encode the given image through a conditional sampler that samples from the latent space of the variational text encoder, which is then decoded to the output depth map. Our approach is trained alternatingly between the text and image branches: in one optimization step, we predict the mean and standard deviation from the text description and sample from a standard Gaussian, and in the other, we sample using a (image) conditional sampler. Once trained, we directly predict depth from the encoded text using the conditional sampler. We demonstrate our approach on indoor (NYUv2) and out-door (KITTI) scenarios, where we show that language can consistently improve performance in both. Code: https://github.com/Adonis-galaxy/WorDepth.
Ziyao Zeng, Daniel Wang 0005, Fengyu Yang 0003, Hyoungseob Park, Stefano Soatto, Dong Lao, Alex Wong 0001
CVPR1
2024 RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language Descriptions
abstract
We propose a method for metric-scale monocular depth estimation. Inferring depth from a single image is an ill-posed problem due to the loss of scale from perspective projection during the image formation process. Any scale chosen is a bias, typically stemming from training on a dataset; hence, existing works have instead opted to use relative (normalized, inverse) depth. Our goal is to recover metric-scaled depth maps through a linear transformation. The crux of our method lies in the observation that certain objects (e.g., cars, trees, street signs) are typically found or associated with certain types of scenes (e.g., outdoor). We explore whether language descriptions can be used to transform relative depth predictions to those in metric scale. Our method, RSA , takes as input a text caption describing objects present in an image and outputs the parameters of a linear transformation which can be applied globally to a relative depth map to yield metric-scaled depth predictions. We demonstrate our method on recent general-purpose monocular depth models on indoors (NYUv2, VOID) and outdoors (KITTI). When trained on multiple datasets, RSA can serve as a general alignment module in zero-shot settings. Our method improves over common practices in aligning relative to metric depth and results in predictions that are comparable to an upper bound of fitting relative depth to ground truth via a linear transformation. Code is available at: https://github.com/Adonis-galaxy/RSA.
Ziyao Zeng, Yangchao Wu, Hyoungseob Park, Daniel Wang 0005, Fengyu Yang 0003, Stefano Soatto, Dong Lao, Byung-Woo Hong, Alex Wong 0001
NeurIPS1
2023 iQuery: Instruments as Queries for Audio-Visual Sound Separation
abstract
Current audiovisual separation methods share a standard architecture design where an audio encoder-decoder network is fused with visual encoding features at the encoder bottleneck. This design confounds the learning of multimodal feature encoding with robust sound decoding for audio separation. To generalize to a new instrument, one must finetune the entire visual and audio network for all musical instruments. We re-formulate the visual-sound separation task and propose Instruments as Queries (iQuery) with a flexible query expansion mechanism. Our approach ensures cross-modal consistency and cross-instrument disentanglement. We utilize “visually named” queries to initiate the learning of audio queries and use cross-modal attention to remove potential sound source interference at the estimated waveforms. To generalize to a new instrument or event class, drawing inspiration from the text-prompt design, we insert additional queries as audio prompts while freezing the attention mechanism. Experimental results on three benchmarks demonstrate that our iQuery improves audiovisual sound source separation performance. Code is available at https://github.com/JiabenChen/iQuery.
Jiaben Chen, Renrui Zhang, Dongze Lian, Ziyao Zeng, Jianbo Shi
CVPR5
2023 PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world Learning
abstract
Large-scale pre-trained models have shown promising open-world performance for both vision and language tasks. However, their transferred capacity on 3D point clouds is still limited and only constrained to the classification task. In this paper, we first collaborate CLIP and GPT to be a unified 3D open-world learner, named as Point-CLIP V2, which fully unleashes their potential for zero-shot 3D classification, segmentation, and detection. To better align 3D data with the pre-trained language knowledge, Point-CLIP V2 contains two key designs. For the visual end, we prompt CLIP via a shape projection module to generate more realistic depth maps, narrowing the domain gap between projected point clouds with natural images. For the textual end, we prompt the GPT model to generate 3D-specific text as the input of CLIP’s textual encoder. Without any training in 3D domains, our approach significantly surpasses PointCLIP by +42.90%, +40.44%, and +28.75% accuracy on three datasets for zero-shot 3D classification. On top of that, V2 can be extended to few-shot 3D classification, zero-shot 3D part segmentation, and 3D object detection in a simple manner, demonstrating our generalization ability for unified 3D open-world learning. Code is available at https://github.com/yangyangyang127/PointCLIP_V2.
Renrui Zhang, Bowei He, Ziyao Zeng, Zipeng Qin, Shanghang Zhang, Peng Gao 0007
ICCV5
2023 DS-Point: A Dual-Scale 3D Framework for Point Cloud Understanding
abstract
Compared with grid-based 2D images, processing 3D point clouds is more challenging due to their irregular distribution and intricate spatial information. Most prior works introduce delicate designs on either local feature aggregators or global geometric architecture, but few combine two scales effectively. Therefore, to better incorporate the advantages of both local and global processing, we propose DS-Point, a dual-scale 3D framework for point cloud understanding. DS-Point firstly disentangles 3D features from channel dimension for concurrent dual-scale modeling, i.e., point-wise convolution for local fine-grained geometry parsing, and voxel-wise attention for global long-range spatial exploration. Upon that, an HF-fusion module is proposed to enhance the cross-modal interaction and thoroughly blend the dual-scale features. Then, with task-specific heads for different downstream tasks, DS-Point serves as an effective 3D framework for feature extraction. By the dual-scale paradigm, DS-Point achieves superior performance on multiple downstream tasks, e.g., 93.8% for shape classification on ModelNet40, 84.9% on ScanObjectNN, and 84.3% on ShapeNetPart.
Renrui Zhang, Ziyao Zeng, Borui Chen, Guangnan Zhang, Xilan Liu
SMC2
2022 Can Language Understand Depth?
abstract
Besides image classification, Contrastive Language-Image Pre-training (CLIP) has accomplished extraordinary success for a wide range of vision tasks, including object-level and 3D space understanding. However, it's still challenging to transfer semantic knowledge learned from CLIP into more intricate tasks of quantified targets, such as depth estimation with geometric information. In this paper, we propose to apply CLIP for zero-shot monocular depth estimation, named DepthCLIP. We found that the patches of input image could respond to a certain semantic distance token and then be projected to a quantified depth bin for coarse estimation. Without any training, our DepthCLIP surpasses existing unsupervised methods and even approaches the early fully-supervised networks. To our best knowledge, we are the first to conduct zero-shot adaptation from the semantic language knowledge to quantified downstream tasks and perform zero-shot monocular depth estimation. We hope our work could cast a light on the future research. The code is available at https://github.com/Adonis-galaxy/DepthCLIP.
Renrui Zhang, Ziyao Zeng
ACM Multimedia2