VLDB 2026 Research / reviewers in the wild / expert
Zehao Zhu
dblp:247/4146
· DBLP profile ↗
12ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SceneCrafter: Controllable Multi-View Driving Scene EditingabstractSimulation is crucial for developing and evaluating autonomous vehicle (AV) systems. Recent literature builds on a new generation of generative models to synthesize highly realistic images for full-stack simulation. However, purely synthetically generated scenes are not grounded in reality and have difficulty in inspiring confidence in the relevance of its outcomes. Editing models, on the other hand, leverage source scenes from real driving logs, and enable the simulation of different traffic layouts, behaviors, and operating conditions such as weather and time of day. While image editing is an established topic in computer vision, it presents fresh sets of challenges in driving simulation: (1) the need for cross-camera 3D consistency, (2) learning “empty street” priors from driving data with foreground occlusions, and (3) obtaining paired image tuples of varied editing conditions while preserving consistent layout and geometry. To address these challenges, we propose SceneCrafter, a versatile editor for realistic 3D-consistent manipulation of driving scenes captured from multiple cameras. We build on recent advancements in multi-view diffusion models, using a fully controllable framework that scales seamlessly to multi-modality conditions like weather, time of day, agent boxes and high-definition maps. To generate paired data for supervising the editing model, we propose a novel framework on top of Prompt-to-Prompt [15] to generate geometrically consistent synthetic paired data with global edits. We also introduce an alpha-blending framework to synthesize data with local edits, leveraging a model trained on empty street priors through novel masked training and multi-view repaint paradigm. SceneCrafter demonstrates powerful editing capabilities and achieves state-of-the-art realism, controllability, 3D consistency, and scene editing quality compared to existing baselines. Zehao Zhu, Yuliang Zou, Chiyu Max Jiang, Vincent Casser, Xiukun Huang, Zhenpei Yang, Ruiqi Gao, Leonidas J. Guibas, Mingxing Tan, Dragomir Anguelov |
CVPR | 1 |
| 2025 | Indoor FireRescue Radar: 4D Indoor Millimeter Wave Dataset and Analysis for Hazardous Environment PerceptionabstractFire-induced indoor environments, characterized by smoke, glare, and dimness, critically challenge rescue safety. While LiDAR and cameras suffer from signal attenuation, millimeter-wave (mmWave) radar exhibits robust imaging performance. Radar-based building mapping and object detection in indoor environments are thus required to facilitate situational awareness specified by firefighting standards. Prior radar datasets mostly focus on outdoor object detection and the few existing indoor datasets remain insufficient in several aspects: (1) lacking adverse scenario analysis; (2) lacking raw analog-to-digital converter (ADC) data for dense point cloud generation; and (3) lacking 3D object annotations for building layout understanding. This work introduces the Indoor FireRescue Radar (IFR) dataset, a novel large-scale multimodal benchmark for indoor situational awareness. It includes 27K frames of 4D radar point cloud, co-calibrated with LiDAR, RGB camera, and IMU streams, alongside 3D objects annotations across 10 buildings. This dataset also provides raw ADC data and sensor configuration metadata. We applied voxel-based and pillar-based object detectors to 4D radar-based indoor object detection. We also demonstrated the robustness of radar perception in fire-induced indoor environments by real smoke tests at a firefighter training facility. Dataset is available at: https://huggingface.co/datasets/yysd123/indoor_mmwave Kangkang Duan, Zehao Zhu, Zhengbo Zou |
IROS | 2 |
| 2025 | Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation ModelsabstractRecent advances in generative models have sparked exciting new possibilities in the field of autonomous vehicles. Specifically, video generation models are now being explored as controllable virtual testing environments. Simultaneously, end-to-end (E2E) driving models have emerged as a streamlined alternative to conventional modular autonomous driving systems, gaining popularity for their simplicity and scalability. However, the application of these techniques to simulation and planning raises important questions. First, while video generation models can generate increasingly realistic videos, can these videos faithfully adhere to the specified conditions and be realistic enough for E2E autonomous planner evaluation? Second, given that data is crucial for understanding and controlling E2E planners, how can we gain deeper insights into their biases and improve their ability to generalize to out-of-distribution scenarios? In this work, we bridge the gap between the driving models and generative world models (Drive&Gen) to address these questions. We propose novel statistical measures leveraging E2E drivers to evaluate the realism of generated videos. By exploiting the controllability of the video generation model, we conduct targeted experiments to investigate distribution gaps affecting E2E planner performance. Finally, we show that synthetic data produced by the video generation model offers a cost-effective alternative to real-world data collection. This synthetic data effectively improves E2E model generalization beyond existing Operational Design Domains, facilitating the expansion of autonomous vehicle services into new operational contexts. Zhenpei Yang, Yijing Bai, Yingwei Li 0002, Yuliang Zou, Abhijit Kundu, José Lezama, Luna Yue Huang, Zehao Zhu, Jyh-Jing Hwang, Dragomir Anguelov, Mingxing Tan, Chiyu Max Jiang |
IROS | 10 |
| 2025 | Subjective and Objective Quality-of-Experience Evaluation Study for Live Video StreamingabstractIn recent years, live video streaming has gained widespread popularity across various social media platforms. Quality of experience (QoE), which reflects end-users’ satisfaction and overall experience, plays a critical role for media service providers to optimize large-scale live compression and transmission strategies to achieve perceptually optimal rate-distortion trade-off. Although many QoE metrics for video-on-demand (VoD) have been proposed, there remain significant challenges in developing QoE metrics for live video streaming. To bridge this gap, we conduct a comprehensive study of subjective and objective QoE evaluations for live video streaming. For the subjective QoE study, we introduce the first live video streaming QoE dataset, TaoLive QoE, which consists of 42 source videos collected from real live broadcasts and 1, 155 corresponding distorted ones degraded due to a variety of streaming distortions, including conventional streaming distortions such as compression, stalling, as well as live streaming-specific distortions like frame skipping, variable frame rate, etc. Subsequently, a human study was conducted to derive subjective QoE scores of videos in the TaoLive QoE dataset. For the objective QoE study, we benchmark existing QoE models on the TaoLive QoE dataset as well as publicly available QoE datasets for VoD scenarios, highlighting that current models struggle to accurately assess video QoE, particularly for live content. Hence, we propose an end-to-end QoE evaluation model, Tao-QoE, which integrates multi-scale semantic features and optical flow-based motion features to predicting a retrospective QoE score, eliminating reliance on statistical quality of service (QoS) features. Extensive experiments demonstrate that Tao-QoE outperforms other models on the TaoLive QoE dataset and five publicly available QoE datasets, showcasing the effectiveness and feasibility of Tao-QoE. Zehao Zhu, Wei Sun 0029, Jun Jia, Jia Wang 0004, Guangtao Zhai |
VCIP | 1 |
| 2024 | ContactArt: Learning 3D Interaction Priors for Category-level Articulated Object and Hand Poses EstimationabstractWe propose a new dataset and a novel approach to learning hand-object interaction priors for hand and articulated object pose estimation. We first collect a dataset using visual teleoperation, where the human operator can directly play within a physical simulator to manipulate the articulated objects. We record the data and obtain free and accurate annotations on object poses and contact information from the simulator. Our system only requires an iPhone to record human hand motion, which can be easily scaled up and largely lower the costs of data and annotation collection. With this data, we learn 3D interaction priors including a discriminator (in a GAN) capturing the distribution of how object parts are arranged, and a diffusion model which generates the contact regions on articulated objects, guiding the hand pose estimation. Such structural and contact priors can easily transfer to real-world data with barely any domain gap. By using our data and learned priors, our method significantly improves the performance on joint hand and articulated object poses estimation over the existing state-of-the-art methods. The project is available at https://zehaozhu.github.io/ContactArt/. Zehao Zhu, Jiashun Wang, Yuzhe Qin, Deqing Sun, Varun Jampani, Xiaolong Wang 0004 |
3DV | 1 |
| 2024 | Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fieldsabstract3D scene representations have gained immense popularity in recent years. Methods that use Neural Radiance fields are versatile for traditional tasks such as novel view synthesis. In recent times, some work has emerged that aims to extend the functionality of NeRF beyond view synthesis, for semantically aware tasks such as editing and segmentation using 3D feature field distillation from 2D foundation models. However, these methods have two major limitations: (a) they are limited by the rendering speed of NeRF pipelines, and (b) implicitly represented feature fields suffer from continuity artifacts reducing feature quality. Recently, 3D Gaussian Splatting has shown state-of-the-art performance on real-time radiance field rendering. In this work, we go one step further: in addition to radiance field rendering, we enable 3D Gaussian splatting on arbitrary-dimension semantic features via 2D foundation model distillation. This translation is not straightforward: naively incorporating feature fields in the 3DGS framework encounters significant challenges, notably the disparities in spatial resolution and channel consistency between RGB images and feature maps. We propose architectural and training changes to efficiently avert this problem. Our proposed method is general, and our experiments showcase novel view semantic segmentation, language-guided editing and segment anything through learning feature fields from state-of-the-art 2D foundation models such as SAM and CLIP-LSeg. Across experiments, our distillation method is able to provide comparable or better results, while being significantly faster to both train and render. Additionally, to the best of our knowledge, we are the first method to enable point and bounding-box prompting for radiance field manipulation, by leveraging the SAM model. Project website at: https://feature-3dgs.github.io/. Shijie Zhou 0003, Sicheng Jiang, Zhiwen Fan, Zehao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang Wang, Achuta Kadambi |
CVPR | 5 |
| 2024 | FSGS: Real-Time Few-Shot View Synthesis Using Gaussian Splatting
Zehao Zhu, Zhiwen Fan, Yifan Jiang 0001, Zhangyang Wang |
ECCV (39) | 1 |
| 2024 | LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPSabstractRecent advances in real-time neural rendering using point-based techniques have enabled broader adoption of 3D representations. However, foundational approaches like 3D Gaussian Splatting impose substantial storage overhead, as Structure-from-Motion (SfM) points can grow to millions, often requiring gigabyte-level disk space for a single unbounded scene. This growth presents scalability challenges and hinders splatting efficiency. To address this, we introduce LightGaussian, a method for transforming 3D Gaussians into a more compact format. Inspired by Network Pruning, LightGaussian identifies Gaussians with minimal global significance on scene reconstruction, and applies a pruning and recovery process to reduce redundancy while preserving visual quality. Knowledge distillation and pseudo-view augmentation then transfer spherical harmonic coefficients to a lower degree, yielding compact representations. Gaussian Vector Quantization, based on each Gaussian’s global significance, further lowers bitwidth with minimal accuracy loss. LightGaussian achieves an average 15 times compression rate while boosting FPS from 144 to 237 within the 3D-GS framework, enabling efficient complex scene representation on the Mip-NeRF 360 and Tank & Temple datasets. The proposed Gaussian pruning approach is also adaptable to other 3D representations (e.g., Scaffold-GS), demonstrating strong generalization capabilities. Zhiwen Fan, Kairun Wen, Zehao Zhu, Dejia Xu, Zhangyang Wang |
NeurIPS | 4 |
| 2021 | Design of Raman microstructure fiber amplifier for 6GabstractThis A Raman fiber amplifier with multi-pump and GeO2doped microstructure fiber is designed to fulfil the requirements of 6G transmission system for optical communication network which can solve the problems of narrow band width, low output gain and uneven output gain of Raman fiber amplifier suitable for 6G system. In theory, using the fourth-order Runge-Kutta method is to solve the classical Raman coupled wave differential equations. In the structure, two pairs of pump beams with the same wavelength are combined to reduce the number of pumps and simplify the structure. At the same time, the method of cascading two sections of microstructured fiber is used to realize the signal gain amplification before and after compensation at the output end of the Raman amplifier. Finally, the average gain of the amplifier is up to 35.72dB in the bandwidth of 100nm, and the gain fluctuation is less than ±0.43dB. Jiamin Gong, Shutao Lei, Zehao Zhu |
WCNC | 6 |
| 2021 | Fine localization and distortion resistant detection of multi-class barcode in complex environments
Xiongkuo Min, Jun Jia, Zehao Zhu, Jia Wang 0004, Guangtao Zhai |
Multim. Tools Appl. | 4 |
| 2019 | A Robust Circular Two-Dimensional Barcode and Decoding MethodabstractWith the popularity of two-dimensional (2D) bar-codes, many image correction algorithms for two-dimensional barcodes have been proposed. However, limited by the form of matrix barcodes, the effectiveness of these correction algorithms is limited. Besides, circular 2D barcodes such as ShotCode have better ability to resist image distortion. But, due to the small information capacity, the research on circular 2D barcodes in recent years is limited. Therefore, we propose a robust circular 2D barcode in this paper. By adopting colour-coding and a new decoding method, it can not only achieve the basic information capacity, but also effectively enhance the ability to resist image distortion. In addition, we propose two indicators (maximum distortion ratio and maximum support opening angle) to measure the image distortion of two-dimensional barcodes on geometric bodies. Experiments verify the superiority of the new circular two-dimensional barcode and its decoding method. Fuwang Yi, Guangtao Zhai, Zehao Zhu |
PCS | 3 |
| 2019 | EMBDN: An Efficient Multiclass Barcode Detection Network for Complicated EnvironmentsabstractThis article presents a novel method for efficient barcodes detection in real and complicated environments using a convolutional neural network (CNN)-based model. The method is developed as a preprocess-module of existing decoders to enhance decoding rates. Our method is trained as an end-to-end model to determine accurate locations of four barcode vertexes. Our method consists of four modules: 1) base net module; 2) region proposals generator; 3) classification and regression module; and 4) distortion removal module. The feature of barcodes extracted from the base net is fed to the next module. Region proposals are generated and selected as region of interest (ROI). Then the ROI are forward propagated to the classification and regression module to determine the positions and shapes of the barcodes. Finally, the distortion removal module is used to remove the geometric distortion according to regression parameters acquired from the previous step. The accurate position and distorted barcodes shape can be determined and corrected by our method. We validate our method on a challenging large-scale dataset in experiments. Compared with the previous methods, our method provides an end-to-end solution to determine accurate locations of barcode vertexes, which shows an excellent performance on detection accuracy. In addition, our method can enhance decoding rate through distortion removal. Jun Jia, Guangtao Zhai, Zhongpai Gao, Zehao Zhu, Xiongkuo Min, Xiaokang Yang 0001, Guodong Guo |
IEEE Internet Things J. | 5 |