Zhenwei Miao

dblp:120/6993 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0002-5422-809XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Systems, architecture and hardware · 2
YearPublicationVenuePosition
2025 CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-Scale Reinforcement Learning in Autonomous Driving
abstract
Trajectory planning is vital for autonomous driving, ensuring safe and efficient navigation in complex environments. While recent learning-based methods, particularly reinforcement learning (RL), have shown promise in specific scenarios, RL planners struggle with training inefficiencies and managing large-scale, real-world driving scenarios. In this paper, we introduce CarPlanner, a Consistent auto-regressive Planner that uses RL to generate multi-modal trajectories. The auto-regressive structure enables efficient large-scale RL training, while the incorporation of consistency ensures stable policy learning by maintaining coherent temporal consistency across time steps. Moreover, Car-Planner employs a generation-selection framework with an expert-guided reward function and an invariant-view module, simplifying RL training and enhancing policy performance. Extensive analysis demonstrates that our proposed RL framework effectively addresses the challenges of training efficiency and performance enhancement, positioning CarPlanner as a promising solution for trajectory planning in autonomous driving. To the best of our knowledge, we are the first to demonstrate that the RL-based planner can surpass both IL-and rule-based state-of-the-arts (SOTAs) on the challenging large-scale real-world dataset nuPlan. Our proposed CarPlanner surpasses RL-, IL-, and rule-based SOTA approaches within this demanding dataset.
Dongkun Zhang, Qi Wang 0056, Rong Xiong, Zhenwei Miao, Yue Wang 0020
CVPR7
2024 LASIL: Learner-Aware Supervised Imitation Learning For Long-Term Microscopic Traffic Simulation
abstract
Microscopic traffic simulation plays a crucial role in transportation engineering by providing insights into in-dividual vehicle behavior and overall traffic flow. How-ever, creating a realistic simulator that accurately repli-cates human driving behaviors in various traffic conditions presents significant challenges. Traditional simulators relying on heuristic models often fail to deliver accurate simulations due to the complexity of real-world traffic environments. Due to the covariate shift issue, existing imitation learning-based simulators often fail to generate stable long-term simulations. In this paper, we propose a novel approach called learner-aware supervised imitation learning to address the covariate shift problem in multi-agent imi-tation learning. By leveraging a variational autoencoder simultaneously modeling the expert and learner state distribution, our approach augments expert states such that the augmented state is aware of learner state distribution. Our method, applied to urban traffic simulation, demon-strates significant improvements over existing state-of-the-art baselines in both short-term microscopic and long-term macroscopic realism when evaluated on the real-world dataset pNEUMA.
Zhenwei Miao, Weizi Li, Dayang Hao, Jia Pan 0001
CVPR2
2023 PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer
abstract
3D object detection in autonomous driving aims to reason “what” and “where” the objects of interest present in a 3D world. Following the conventional wisdom of previous 2D object detection, existing methods often adopt the canonical Cartesian coordinate system with perpendicular axis. However, we conjugate that this does not fit the nature of the ego car’s perspective, as each onboard camera perceives the world in shape of wedge intrinsic to the imaging geometry with radical (non perpendicular) axis. Hence, in this paper we advocate the exploitation of the Polar coordinate system and propose a new Polar Transformer (PolarFormer) for more accurate 3D object detection in the bird’s-eye-view (BEV) taking as input only multi-camera 2D images. Specifically, we design a cross-attention based Polar detection head without restriction to the shape of input structure to deal with irregular Polar grids. For tackling the unconstrained object scale variations along Polar’s distance dimension, we further introduce a multi-scale Polar representation learning strategy. As a result, our model can make best use of the Polar representation rasterized via attending to the corresponding image observation in a sequence-to-sequence fashion subject to the geometric constraints. Thorough experiments on the nuScenes dataset demonstrate that our PolarFormer outperforms significantly state-of-the-art 3D object detection alternatives.
Yanqin Jiang, Li Zhang 0040, Zhenwei Miao, Xiatian Zhu, Weiming Hu 0004, Yu-Gang Jiang 0001
AAAI3
2023 PUPS: Point Cloud Unified Panoptic Segmentation
abstract
Point cloud panoptic segmentation is a challenging task that seeks a holistic solution for both semantic and instance segmentation to predict groupings of coherent points. Previous approaches treat semantic and instance segmentation as surrogate tasks, and they either use clustering methods or bounding boxes to gather instance groupings with costly computation and hand-craft designs in the instance segmentation task. In this paper, we propose a simple but effective point cloud unified panoptic segmentation (PUPS) framework, which use a set of point-level classifiers to directly predict semantic and instance groupings in an end-to-end manner. To realize PUPS, we introduce bipartite matching to our training pipeline so that our classifiers are able to exclusively predict groupings of instances, getting rid of hand-crafted designs, e.g. anchors and Non-Maximum Suppression (NMS). In order to achieve better grouping results, we utilize a transformer decoder to iteratively refine the point classifiers and develop a context-aware CutMix augmentation to overcome the class imbalance problem. As a result, PUPS achieves 1st place on the leader board of SemanticKITTI panoptic segmentation task and state-of-the-art results on nuScenes.
Shihao Su, Jianyun Xu, Zhenwei Miao, Xin Zhan, Dayang Hao
AAAI4
2022 LIFT: Learning 4D LiDAR Image Fusion Transformer for 3D Object Detection
abstract
LiDAR and camera are two common sensors to collect data in time for 3D object detection under the autonomous driving context. Though the complementary information across sensors and time has great potential of benefiting 3D perception, taking full advantage of sequential cross-sensor data still remains challenging. In this paper, we propose a novel LiDAR Image Fusion Transformer (LIFT) to model the mutual interaction relationship of cross-sensor data over time. LIFT learns to align the input 4D sequential cross-sensor data to achieve multi-frame multi-modal information aggregation. To alleviate computational load, we project both point clouds and images into the bird-eye-view maps to compute sparse grid-wise self-attention. LIFT also benefits from a cross-sensor and cross-time data augmentation scheme. We evaluate the proposed approach on the challenging nuScenes and Waymo datasets, where our LIFT performs well over the state-of-the-art and strong baselines.
Yihan Zeng, Chunwei Wang, Zhenwei Miao, Xin Zhan, Dayang Hao, Chao Ma 0004
CVPR4
2022 SP-Net: Slowly Progressing Dynamic Inference Networks
Wenhu Zhang, Shihao Su, Hui Wang 0107, Zhenwei Miao, Xin Zhan, Xi Li 0001
ECCV (11)5
2022 INT: Towards Infinite-Frames 3D Detection with an Efficient Framework
Jianyun Xu, Zhenwei Miao, Hongyu Pan, Peihan Hao, Zhengyang Sun, Xin Zhan
ECCV (9)2
2022 DeepInteraction: 3D Object Detection via Modality Interaction
abstract
Existing top-performance 3D object detectors typically rely on the multi-modal fusion strategy. This design is however fundamentally restricted due to overlooking the modality-specific useful information and finally hampering the model performance. To address this limitation, in this work we introduce a novel modality interaction strategy where individual per-modality representations are learned and maintained throughout for enabling their unique characteristics to be exploited during object detection. To realize this proposed strategy, we design a DeepInteraction architecture characterized by a multi-modal representational interaction encoder and a multi-modal predictive interaction decoder. Experiments on the large-scale nuScenes dataset show that our proposed method surpasses all prior arts often by a large margin. Crucially, our method is ranked at the first position at the highly competitive nuScenes object detection leaderboard.
Zeyu Yang 0004, Zhenwei Miao, Xiatian Zhu, Li Zhang 0040
NeurIPS3
2022 Dual adversarial model: Exploring low-dimensional space features for point clouds generating and completing
Yuhang Zhang 0012, Zhenwei Miao, Tiebin Mi, Jie Li 0002, Robert C. Qiu
Comput. Vis. Image Underst.2
2021 PVGNet: A Bottom-Up One-Stage 3D Object Detector With Integrated Multi-Level Features
abstract
Quantization-based methods are widely used in LiDAR points 3D object detection for its efficiency in extracting context information. Unlike image where the context information is distributed evenly over the object, most LiDAR points are distributed along the object boundary, which means the boundary features are more critical in LiDAR points 3D detection. However, quantization inevitably introduces ambiguity during both the training and inference stages. To alleviate this problem, we propose a one-stage and voting-based 3D detector, named Point-Voxel-Grid Network (PVGNet). In particular, PVGNet extracts point, voxel and grid-level features in a unified backbone architecture and produces point-wise fusion features. It segments Li-DAR points into foreground and background, predicts a 3D bounding box for each foreground point, and performs group voting to get the final detection results. Moreover, we observe that instance-level point imbalance due to occlusion and observation distance also degrades the detection performance. A novel instance-aware focal loss is proposed to alleviate this problem and further improve the detection ability. We conduct experiments on the KITTI and Waymo datasets. Our proposed PVGNet outperforms previous state-of-the-art methods and ranks at the top of KITTI 3D/BEV detection leaderboards.
Zhenwei Miao, Jikai Chen, Hongyu Pan, Ruiwen Zhang, Peihan Hao, Xin Zhan
CVPR1
2017 Laplace gradient based Discriminative and Contrast Invertible descriptor
abstract
The performance of local descriptors such as SIFT drops under severe illumination changes. In this paper, we propose a Discriminative and Contrast Invertible (DCI) local feature descriptor. In order to increase the discriminative ability of the descriptor under illumination changes, a Laplace gradient based histogram is proposed. Moreover, a robust contrast flipping estimate is proposed based on the divergence of a local region. Experiments on fine-grained object recognition and retrieval applications demonstrate the superior performance of the DCI descriptor to others.
Zhenwei Miao, Kim-Hui Yap, Xudong Jiang 0001, Subbhuraam Sinduja
ICASSP1
2017 Lattice-Support repetitive local feature detection for visual search
Dipu Manandhar, Kim-Hui Yap, Zhenwei Miao, Lap-Pui Chau
Pattern Recognit. Lett.3
2016 Contrast Invariant Interest Point Detection by Zero-Norm LoG Filter
abstract
The Laplacian of Gaussian (LoG) filter is widely used in interest point detection. However, low-contrast image structures, though stable and significant, are often submerged by the high-contrast ones in the response image of the LoG filter, and hence are difficult to be detected. To solve this problem, we derive a generalized LoG filter, and propose a zero-norm LoG filter. The response of the zero-norm LoG filter is proportional to the weighted number of bright/dark pixels in a local region, which makes this filter be invariant to the image contrast. Based on the zero-norm LoG filter, we develop an interest point detector to extract local structures from images. Compared with the contrast dependent detectors, such as the popular scale invariant feature transform detector, the proposed detector is robust to illumination changes and abrupt variations of images. Experiments on benchmark databases demonstrate the superior performance of the proposed zero-norm LoG detector in terms of the repeatability and matching score of the detected points as well as the image recognition rate under different conditions.
Zhenwei Miao, Xudong Jiang 0001, Kim-Hui Yap
IEEE Trans. Image Process.1
2015 Hybrid feature-based wallpaper visual search
abstract
In this paper we propose a hybrid feature-based wallpaper visual search system. As opposed to conventional techniques that use global features to perform wallpaper search, this paper proposes to integrate local and global features to support both functions of recognition (identify the product ID of the query images) and retrieval (search wallpapers that are visually similar to the query images). An adaptive SIFT is designed to extract sufficient number of local features from both the query and reference images. The combination of the sparse and dense SIFT features results in a significant improvement of the recognition rate. Global features are further incorporated in the system for the visually similar image retrieval. A new query expansion is proposed to alleviate the problems caused by cluttered background, occlusion, scale change and illumination changes. Experiments on a dataset consisting of 2,208 reference images from 218 different designs show that the proposed method can achieve a recognition rate of more than 90%.
Kim-Hui Yap, Zhenwei Miao
ISCAS2
2015 Feature weighting in visual product recognition
abstract
Significant progress towards visual search has been made in the past two decades through the development of local invariant features. Among existing local feature detectors, the Scale Invariant Feature Transform (SIFT) is widely used since it is designed to be invariant to minimal illumination changes and certain geometric transformations. However, in practice, the recognition performance is still subject to actual condition. Some keypoints are more stable while others are less stable and can not be repeatedly detected. Besides, in visual object recognition where the foreground object is to be recognized while the background suppressed, the current scalable vocabulary tree (SVT) framework treats each descriptor as equally important, hence restricting its performance. This paper aims to study the effect of SIFT respect to illumination and geometric changes and develop a feature weighting algorithm to incorporate the stability of SIFT and saliency information into weighted scalable vocabulary tree (WSVT) based recognition. Experimental results on a commercial product database show the proposed feature weighting algorithm outperforms the baseline SVT recognition by 5%.
Kim-Hui Yap, Dajiang Zhang, Zhenwei Miao
ISCAS4
2014 Additive and exclusive noise suppression by iterative trimmed and truncated mean algorithm
Zhenwei Miao, Xudong Jiang 0001
Signal Process.1
2013 A vote of confidence based interest point detector
abstract
In this paper, a vote of confidence (VC) based detector is proposed to detect bright and dark regions from images. Whether a local region is bright or dark is voted by all the pixels in this region. Compared to the contrast based detectors, such as the popular SIFT detector, the VC detector is invariant to illumination change and robust to abrupt variations. Experiments are conducted on benchmark databases to verify the superior performance of the VC detector in terms of the repeatability and matching score. The proposed detector is also evaluated in the application of face recognition.
Zhenwei Miao, Xudong Jiang 0001
ICASSP1
2013 Interest point detection using rank order LoG filter
Zhenwei Miao, Xudong Jiang 0001
Pattern Recognit.1
2012 A novel rank order LoG filter for interest point detection
abstract
This paper proposes a novel non-linear filter, named rank order LoG (ROLG) filter, and a new interest point detector, named ROLG detector. The ROLG filter is a weighted rank order filter. It is used to detect image structures whose significant majority of pixels are brighter (or darker) than the significant majority of pixels in their corresponding surroundings. The ROLG detector is built on this filter. Compared to linear filter based detectors, the proposed rank order filter based detector is more robust to abrupt variations of images. Experiments on the benchmark databases demonstrate that the ROLG detector achieves superior performance compared to four state-of-the-art detectors. Evaluation experiments are also conducted on face recognition. The results further demonstrate that the ROLG detector has better performance compared to other detectors.
Zhenwei Miao, Xudong Jiang 0001
ICASSP1