Zaiwei Zhang

dblp:186/4421 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 2
YearPublicationVenuePosition
2025 Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry Locality
abstract
Recent efforts recognize the power of scale in 3D learning (e.g. PTv3) and attention mechanisms (e.g. FlashAt-tention). However, current point cloud backbones fail to holistically unify geometric locality, attention mechanisms, and GPU architectures in one view. In this paper, we introduce Flash3D Transformer, which aligns geometric locality and GPU tiling through a principled locality mechanism based on Perfect Spatial Hashing (PSH). The common alignment with GPU tiling naturally fuses our PSH locality mechanism with FlashAttention at negligible extra cost. This mechanism affords flexible design choices throughout the backbone that result in superior downstream task results. Flash3D outperforms state-of-the-art PTv3 results on benchmark datasets, delivering a 2.25x speed increase and 2.4x memory efficiency boost. This efficiency enables scaling to wider attention scopes and larger models without additional overhead. Such scaling allows Flash3D to achieve even higher task accuracies than PTv3 under the same compute budget.
Gregory P. Meyer, Zaiwei Zhang, Eric M. Wolff, Paul Vernaza
CVPR3
2025 Uncertainty-Guided Enhancement on Driving Perception System Via Foundation Models
abstract
Multimodal foundation models offer promising advancements for enhancing driving perception systems, but their high computational and financial costs pose challenges. We develop a method that leverages foundation models to refine predictions from existing driving perception modelssuch as enhancing object classification accuracy-while minimizing the frequency of using these resource-intensive models. The method quantitatively characterizes uncertainties in the perception model's predictions and engages the foundation model only when these uncertainties exceed a pre-specified threshold. Specifically, it characterizes uncertainty by calibrating the perception model's confidence scores into theoretical lower bounds on the probability of correct predictions using conformal prediction. Then, it sends images to the foundation model and queries for refining the predictions only if the theoretical bound of the perception model's outcome is below the threshold. Additionally, we propose a temporal inference mechanism that enhances prediction accuracy by integrating historical predictions, leading to tighter theoretical bounds. The method demonstrates a 10 to 15 percent improvement in prediction accuracy and reduces the number of queries to the foundation model by 50 percent, based on quantitative evaluations from driving datasets.
Yunhao Yang, Zaiwei Zhang, Zhichao Lu, Ufuk Topcu, Ben Snyder
ICRA4
2023 Implicit Surface Contrastive Clustering for LiDAR Point Clouds
abstract
Self-supervised pretraining on large unlabeled datasets has shown tremendous success in improving the task performance of many 2D and small scale 3D computer vision tasks. However, the popular pretraining approaches have not been impactfully applied to outdoor LiDAR point cloud perception due to the latter's scene complexity and wide range. We propose a new self-supervised pretraining method ISCC with two novel pretext tasks for LiDAR point clouds. The first task uncovers semantic information by sorting local groups of points in the scene into a globally consistent set of semantically meaningful clusters using contrastive learning, complemented by a second task which reasons about precise surfaces of various parts of the scene through implicit surface reconstruction to learn geometric structures. We demonstrate their effectiveness through transfer learning on 3D object detection and semantic segmentation in real world LiDAR scenes. We further design an unsupervised semantic grouping task to show that our approach learns highly semantically meaningful features.
Zaiwei Zhang, Min Bai, Li Erran Li
CVPR1
2022 FvOR: Robust Joint Shape and Pose Optimization for Few-view Object Reconstruction
abstract
Reconstructing an accurate 3D object model from a few image observations remains a challenging problem in computer vision. State-of-the-art approaches typically assume accurate camera poses as input, which could be difficult to obtain in realistic settings. In this paper, we present FvOR, a learning-based object reconstruction method that predicts accurate 3D models given a few images with noisy input poses. The core of our approach is a fast and robust multi-view reconstruction algorithm to jointly refine 3D geometry and camera pose estimation using learnable neural network modules. We provide a thorough benchmark of state-of-the-art approaches for this problem on ShapeNet. Our approach achieves best-in-class results. It is also two orders of magnitude faster than the recent optimization-based approach IDR [67].
Zhenpei Yang, Zhile Ren, Miguel Ángel Bautista 0001, Zaiwei Zhang, Qi Shan, Qixing Huang
CVPR4
2022 Self-Supervised Pretraining for Large-Scale Point Clouds
abstract
Pretraining on large unlabeled datasets has been proven to improve the down-stream task performance on many computer vision tasks, such as 2D object detection and video classification. However, for large-scale 3D scenes, such as outdoor LiDAR point clouds, pretraining is not widely used. Due to the special data characteristics of large 3D point clouds, 2D pretraining frameworks tend to not generalize well. In this paper, we propose a new self-supervised pretraining method that targets large-scale 3D scenes. We pretrain commonly used point-based and voxel-based model architectures and show the transfer learning performance on 3D object detection and also semantic segmentation. We demonstrate the effectiveness of our approach on both dense 3D indoor point clouds and also sparse outdoor lidar point clouds.
Zaiwei Zhang, Min Bai, Li Erran Li
NeurIPS1
2021 Scene Synthesis via Uncertainty-Driven Attribute Synchronization
abstract
Developing deep neural networks to generate 3D scenes is a fundamental problem in neural synthesis with immediate applications in architectural CAD, computer graphics, as well as in generating virtual robot training environments. This task is challenging because 3D scenes exhibit diverse patterns, ranging from continuous ones, such as object sizes and the relative poses between pairs of shapes, to discrete patterns, such as occurrence and co-occurrence of objects with symmetrical relationships. This paper introduces a novel neural scene synthesis approach that can capture diverse feature patterns of 3D scenes. Our method combines the strength of both neural network-based and conventional scene synthesis approaches. We use the parametric prior distributions learned from training data, which provide uncertainties of object attributes and relative attributes, to regularize the outputs of feed-forward neural models. Moreover, instead of merely predicting a scene layout, our approach predicts an over-complete set of attributes. This methodology allows us to utilize the underlying consistency constraints among the predicted attributes to prune infeasible predictions. Experimental results show that our approach outperforms existing methods considerably. The generated 3D scenes interpolate the training data faithfully while preserving both continuous and discrete feature patterns.
Haitao Yang 0005, Zaiwei Zhang, Siming Yan, Chongyang Ma, Chandrajit L. Bajaj, Qixing Huang
ICCV2
2021 ARAPReg: An As-Rigid-As Possible Regularization Loss for Learning Deformable Shape Generators
abstract
This paper introduces an unsupervised loss for training parametric deformation shape generators. The key idea is to enforce the preservation of local rigidity among the generated shapes. Our approach builds on an approximation of the as-rigid-as possible (or ARAP) deformation energy. We show how to develop the unsupervised loss via a spectral decomposition of the Hessian of the ARAP energy. Our loss nicely decouples pose and shape variations through a robust norm. The loss admits simple closed-form expressions. It is easy to train and can be plugged into any standard generation models, e.g., variational auto-encoder (VAE) and auto-decoder (AD). Experimental results show that our approach outperforms existing shape generation approaches considerably on public benchmark datasets of various shape categories such as human, animal and bone. Our code and data are available at https://github.com/GitBoSun/ARAPReg.
Qixing Huang, Xiangru Huang, Zaiwei Zhang, Chandrajit L. Bajaj
ICCV4
2021 Self-Supervised Pretraining of 3D Features on any Point-Cloud
abstract
Pretraining on large labeled datasets is a prerequisite to achieve good performance in many computer vision tasks like image recognition, video understanding etc. However, pretraining is not widely used for 3D recognition tasks where state-of-the-art methods train models from scratch. A primary reason is the lack of large annotated datasets because 3D data labelling is time-consuming. Recent work shows that self-supervised learning is useful to pretrain models in 3D but requires multi-view data and point correspondences. We present a simple self-supervised pretraining method that can work with single-view depth scans acquired by varied sensors, without 3D registration and point correspondences. We pretrain standard point cloud and voxel based model architectures, and show that joint pretraining further improves performance. We evaluate our models on 9 benchmarks for object detection, semantic segmentation, and object classification, where they achieve state-of-the-art results. Most notably, we set a new state-of-the-art for object detection on ScanNet (69.0% mAP) and SUNRGBD (63.5% mAP). Our pretrained models are label efficient and improve performance for classes with few examples.
Zaiwei Zhang, Rohit Girdhar, Armand Joulin, Ishan Misra
ICCV1
2021 GAN-SRAF: Subresolution Assist Feature Generation Using Generative Adversarial Networks
abstract
As the integrated circuits (ICs) technology continues to scale, resolution enhancement techniques (RETs) are mandatory to obtain high manufacturing quality and yield. Among various RETs, subresolution assist feature (SRAF) generation is a key technique to improve the target pattern quality and lithographic process window. While model-based SRAF insertion techniques have demonstrated high accuracy, they usually suffer from high computational cost. Therefore, more efficient techniques that can achieve high accuracy while reducing runtime are in strong demand. In this article, we leverage the recent advancement in machine learning for image generation to tackle the SRAF insertion problem. In particular, we propose a new SRAF insertion framework, GAN-SRAF, which uses generative adversarial networks (GANs) to generate SRAFs directly for any given layout. Our proposed approach incorporates a novel layout to image encoding using multichannel heatmaps to preserve the layout information and facilitate layout reconstruction. Our experimental results demonstrate ~14.6× reduction in runtime when compared to the previous best machine learning approach for SRAF generation, and ~144× reduction compared to the model-based approach, while achieving comparable quality of results.
Mohamed Baker Alawieh, Yibo Lin, Zaiwei Zhang, Meng Li 0004, Qixing Huang, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 H3DNet: 3D Object Detection Using Hybrid Geometric Primitives
Zaiwei Zhang, Haitao Yang 0005, Qixing Huang
ECCV (12)1
2020 Deep Generative Modeling for Scene Synthesis via Hybrid Representations
abstract
We present a deep generative scene modeling technique for indoor environments. Our goal is to train a generative model using a feed-forward neural network that maps a prior distribution (e.g., a normal distribution) to the distribution of primary objects in indoor scenes. We introduce a 3D object arrangement representation that models the locations and orientations of objects, based on their size and shape attributes. Moreover, our scene representation is applicable for 3D objects with different multiplicities (repetition counts), selected from a database. We show a principled way to train this model by combining discriminative losses for both a 3D object arrangement representation and a 2D image-based representation. We demonstrate the effectiveness of our scene representation and the network training method on benchmark datasets. We also show the applications of this generative model in scene interpolation and scene completion.
Zaiwei Zhang, Zhenpei Yang, Chongyang Ma, Linjie Luo, Alexander G. Huth, Etienne Vouga, Qixing Huang
ACM Trans. Graph.1
2019 Path-Invariant Map Networks
abstract
Optimizing a network of maps among a collection of objects/domains (or map synchronization) is a central problem across computer vision and many other relevant fields. Compared to optimizing pairwise maps in isolation, the benefit of map synchronization is that there are natural constraints among a map network that can improve the quality of individual maps. While such self-supervision constraints are well-understood for undirected map networks (e.g., the cycle-consistency constraint), they are under-explored for directed map networks, which naturally arise when maps are given by parametric maps (e.g., a feed-forward neural network). In this paper, we study a natural self-supervision constraint for directed map networks called path-invariance, which enforces that composite maps along different paths between a fixed pair of source and target domains are identical. We introduce path-invariance bases for efficient encoding of the path-invariance constraint and present an algorithm that outputs a path-variance basis with polynomial time and space complexities. We demonstrate the effectiveness of our formulation on optimizing object correspondences, estimating dense image maps via neural networks, and 3D scene segmentation via map networks of diverse 3D representations. In particular, our approach only requires 8\% labeled data from ScanNet to achieve the same performance as training a single 3D semantic segmentation network with 30\% to 100\% labeled data.
Zaiwei Zhang, Zhenxiao Liang, Lemeng Wu, Xiaowei Zhou 0001, Qixing Huang
CVPR1
2019 GAN-SRAF: Sub-Resolution Assist Feature Generation Using Conditional Generative Adversarial Networks
abstract
As the integrated circuits (IC) technology continues to scale, resolution enhancement techniques (RETs) are mandatory to obtain high manufacturing quality and yield. Among various RETs, sub-resolution assist feature (SRAF) generation is a key technique to improve the target pattern quality and lithographic process window. While model-based SRAF insertion techniques have demonstrated high accuracy, they usually suffer from high computational cost. Therefore, more efficient techniques that can achieve high accuracy while reducing runtime are in strong demand. In this work, we leverage the recent advancement in machine learning for image generation to tackle the SRAF insertion problem. In particular, we propose a new SRAF insertion framework, GAN-SRAF, which uses conditional generative adversarial networks (CGANs) to generate SRAFs directly for any given layout. Our proposed approach incorporates a novel layout to image encoding using multi-channel heatmaps to preserve the layout information and facilitate layout reconstruction. Our experimental results demonstrate ~14.6× reduction in runtime when compared to the previous best machine learning approach for SRAF generation, and ~144× reduction compared to model-based approach, while achieving comparable quality of results.
Mohamed Baker Alawieh, Yibo Lin, Zaiwei Zhang, Meng Li 0004, Qixing Huang, David Z. Pan
DAC3
2017 Indoor Follow Me Drone
abstract
With the availability of inexpensive and powerful drones, it is possible to let drones automatically follow a user for video taping. This can not only reduce cost, but also support video taping in situations where otherwise not possible (e.g., during private moments or at inconvenient locations like indoor rock climbing). While there have been many follow-me drones on the market for outdoors, which rely on GPS, enabling indoor follow-me function is more challenging due to the lack of an effective approach to track users in indoor environments. To this end, we develop a holistic system that lets a mobile phone carried by a user accurately track the drone's relative location and control it to maintain a specified distance and orientation for automatic video taping. We develop a series of techniques to (i) track a drone's location using acoustic signals with sub-centimeter errors even under strong propeller noise from the drone and complicated multipath in indoor environments, and (ii) solve practical challenges in applying model predictive control (MPC) framework to control the drone. The latter consists of developing measurement-based flight models, designing measurement techniques to provide feedback to the controller, and predicting the user's movement. We implement our system on AR Drone 2.0 and Samsung S7. The extensive evaluation shows that our drone can follow a user effectively and maintain a specified following distance and orientation within 2-3 cm and 1-3 degree errors, respectively. The videos taped by the drone during flight are smooth according to the jerk metric.
Wenguang Mao, Zaiwei Zhang, Lili Qiu, Jian He 0002, Yuchen Cui, Sangki Yun
MobiSys2
2016 High-precision acoustic motion tracking: demo
abstract
Video games, virtual reality, augmented reality, and smart appliances all call for a new way for users to interact and control them. This paper develops high-preCision Acoustic Tracker (CAT), which aims to replace a traditional mouse and let a user control various devices by moving a smartphone in the air. At its heart lies a distributed Frequency Modulated Continuous Waveform (FMCW) that can accurately estimate the distance between a transmitter and a receiver that are separate and unsynchronized. We further develop an optimization framework to combine FMCW estimation with Doppler shifts to enhance the accuracy. We implement CAT on a mobile phone. The performance evaluation and user study show that our system achieves high tracking accuracy and ease of use using existing hardware.
Wenguang Mao, Jian He 0002, Huihuang Zheng, Zaiwei Zhang, Lili Qiu
MobiCom4