Yao Li 0016

dblp:96/13-16 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-6063-3331ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Systems, architecture and hardware · 6 · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Standard-Compliant Joint Optimization of Rate-Distortion-Decoding-Complexity for Versatile Video Coding
Jiazhen Wang, Yao Li 0016, Xinmin Feng, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002
ISCAS4
2026 USTC-TD: A Test Dataset and Benchmark for Image and Video Coding in 2020s
abstract
Image/video coding has been a remarkable research area for both academia and industry for many years. Testing datasets, especially high-quality image/video datasets, are desirable for the justified evaluation of coding-related research, practical applications, and standardization activities. We put forward a test dataset, namely USTC-TD, which has been successfully adopted in the practical end-to-end image/video coding challenge ofIEEE International Conference on Visual Communications and Image Processing (VCIP)in 2022 and 2023. USTC-TD contains 40 images at 4K spatial resolution and 10 video sequences at 1080p spatial resolution, featuring various content due to the diverse environmental factors (e.g., scene type, texture, motion, view) and the designed imaging factors (e.g., illumination, lens, shadow). We quantitatively evaluate USTC-TD on different image/video features (spatial, temporal, color, lightness), and compare it with the previous image/video test datasets, which verifies its excellent compensation for the shortcomings of existing datasets. We also evaluate both classic standardized and recently learned image/video coding schemes on USTC-TD using objective quality metrics (PSNR, MS-SSIM, VMAF) and subjective quality metric (MOS), providing an extensive benchmark for these evaluated schemes. Based on the characteristics and specific design of the proposed test dataset, we analyze the benchmark performance and shed light on the future research and development of image/video coding. All the data are released online:https://esakak.github.io/USTC-TD.
Zhuoyuan Li 0001, Junqi Liao, Chuanbo Tang, Haotian Zhang 0009, Yifan Bian, Xihua Sheng, Xinmin Feng, Yao Li 0016, Changsheng Gao, Li Li 0040, Dong Liu 0002, Feng Wu 0005
IEEE Trans. Multim.9
2025 CELLmap: Enhancing LiDAR SLAM Through Elastic and Lightweight Spherical Map Representation
abstract
SLAM is a fundamental capability of unmanned systems, with LiDAR-based SLAM gaining widespread adoption due to its high precision. Current SLAM systems can achieve centimeter-level accuracy within a short period. However, there are still several challenges when dealing with largescale mapping tasks including significant storage requirements and difficulty of reusing the constructed maps. To address this, we first design an elastic and lightweight map representation called CELLmap, composed of several CELLS, each representing the local map at the corresponding location. Then, we design a general backend including CELL-based bidirectional registration module and loop closure detection module to improve global map consistency. Our experiments have demonstrated that CELLmap can represent the precise geometric structure of large-scale maps of KITTI dataset using only about 60 MB. Additionally, our general backend achieves up to a 26.88% improvement over various LiDAR odometry methods.
Yifan Duan, Yao Li 0016, Guoliang You, Xiaomeng Chu, Jianmin Ji, Yanyong Zhang
ICRA3
2025 Collaborative Decoder-side Motion Vector Refinement for Video Coding
abstract
Motion compensation prediction (MCP) is a key technology to reduce the temporal redundancy in video coding. Recently, in order to improve its efficiency, the decoder-side MCP schemes are gradually adopted in advanced video coding standards, especially the decoder-side motion vector refinement (DMVR). In DMVR, the bilateral matching scheme is used to refine the motion vector (MV) obtained from merge-based inter mode at the sub-block level, which assumes that the motion vector difference (MVD) in the two reference directions of a bi-prediction block has the symmetric property. Although the bilateral assumption can effectively reduce complexity without any extra signal, the fixed searching rules limit the motion vector accuracy. To address this limitation, we propose a collaborative decoder-side motion vector refinement (C-DMVR) framework. In C-DMVR, the sub-block-based collaborative mechanism is introduced to optimize the distortion calculation (CDC) and bilateral-based searching strategy (CSS) to avoid inaccurate searching results, respectively. In the CDC, the receptive field of the distortion function is enlarged with the collaboration of additional spatial neighbor information to assist the accurate decision of motion vector candidates. In the CSS, the coarse-to-fine candidate list derivation scheme is introduced to construct the accurate searching path with the collaboration of neighbor sub-blocks. The proposed method is implemented into the AOMedia Video 2 reference software, AV2. Experimental results show that the proposed method achieves on average 0.17%, and up to 0.60% BD-rate reduction compared to the AV2 anchor under the random access (RA) configuration, with a slight increase of time complexity on the encoding/decoding side.
Zhuoyuan Li 0001, Yao Li 0016, Li Li 0040, Houqiang Li
ISCAS3
2025 IVCA: Inter-relation-aware Video Complexity Analyzer
abstract
To address the real-time analysis requirements of video streaming applications, we propose an innovative inter-relation-aware video complexity analyzer (IVCA) to enhance the existing video complexity analyzer (VCA). The IVCA overcomes the limitations of the VCA by incorporating inter-frame relations, focusing on inter motion and reference structure. To begin with, we improve the accuracy of temporal features by integrating feature-domain motion estimation into the IVCA framework, which allows for a more nuanced understanding of motion across frames. Furthermore, inspired by the hierarchical reference structures utilized in modern codecs, we introduce layer-aware weights that effectively adjust the contributions of frame complexity across different layers, ensuring a more balanced representation of video characteristics. In addition, we broaden the analysis of temporal features by considering reference frames rather than relying solely on the preceding frame, thereby enriching the contextual understanding of video content. Experimental results demonstrate a significant enhancement in complexity estimation accuracy achieved by the IVCA, coupled with a negligible increase in time complexity, indicating its potential for real-time applications in video streaming scenarios. This advancement not only improves video processing efficiency but also paves the way for more sophisticated analytical tools in video technology.
Junqi Liao, Yao Li 0016, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002
ISCAS2
2025 Frequency Domain Intra Pattern Copy for JPEG XS Screen Content Coding
abstract
JPEG XS is a wavelet-based lightweight image coding standard that features low-complexity and low-latency. As currently there are no efficient intra-compensation prediction techniques conforming to these features, we propose a frequency domain intra-copy prediction framework named Intra Pattern Copy (IPC), to improve its coding efficiency on screen content. In IPC, prediction methods that leverage the diverse decomposed patterns of two-dimensional wavelets, including the directional and frequency characteristics, are proposed to achieve efficient predictions under low-complexity and low-latency constraints. Specifically, we perform in-band compensation predictions in a multi-band synchronized approach, with coefficients of similar pattern distributions predicted simultaneously. A coefficient grouping scheme is derived from the band characteristics to facilitate this compensation process. Based on the grouping scheme, a multi-band synchronized side information coding method is also proposed to code the pattern offset vector of coefficients. Moreover, pattern search schemes incorporating strict limitations on the search range and prediction block size are further developed. Simulation results on JPEG XS demonstrate that an average improvement of 0.75 dB and 1.99 dB in BD-PSNR can be achieved on screen content for two different wavelet decomposition configurations, respectively, with a moderate increase in complexity.
Yao Li 0016, Zhuoyuan Li 0001, Dong Liu 0002, Li Li 0040
IEEE Trans. Circuits Syst. Video Technol.1
2024 CalibFormer: A Transformer-based Automatic LiDAR-Camera Calibration Network
abstract
The fusion of LiDARs and cameras has been increasingly adopted in autonomous driving for perception tasks. The performance of such fusion-based algorithms largely depends on the accuracy of sensor calibration, which is challenging due to the difficulty of identifying common features across different data modalities. Previously, many calibration methods involved specific targets and/or manual intervention, which has proven to be cumbersome and costly. Learning-based online calibration methods have been proposed, but their performance is barely satisfactory in most cases. These methods usually suffer from issues such as sparse feature maps, unreliable cross-modality association, inaccurate calibration parameter regression, etc. In this paper, to address these issues, we propose CalibFormer, an end-to-end network for automatic LiDAR-camera calibration. We aggregate multiple layers of camera and LiDAR image features to achieve high-resolution representations. A multi-head correlation module is utilized to identify correlations between features more accurately. Lastly, we employ transformer architectures to estimate accurate calibration parameters from the correlation information. Our method achieved a mean translation error of 0.8751cm and a mean rotation error of 0.0562° on the KITTI dataset, surpassing existing state-of-the-art methods and demonstrating strong robustness, accuracy, and generalization capabilities.
Yao Li 0016, Chengzhen Meng, Jianmin Ji, Yanyong Zhang
ICRA2
2024 PhD Forum Abstract: Sensor Fusion for Vehicle-side and Roadside 3D Object Detection and Tracking
abstract
We focus on the sensor fusion for vehicle-side and roadside 3D object detection and tracking. Although quite a few sensor fusion algorithms have been proposed, some of which are top-ranked on various leaderboards, a systematic study on how to integrate three crucial sensors (LiDAR, camera and millimeter-wave Radar sensors) to develop effective multi-modal 3D object detection and tracking for vehicle-side perception is still missing. Therefore, we first study the three sensors’ strengths and weaknesses carefully, then compare several different fusion strategies to maximize their utility. Finally, based on the lessons learnt, we propose a simple yet effective multi-modal 3D object detection and tracking framework (namely EZFusion). Without fancy network modules, our proposed EZFusion makes remarkable improvements over the LiDAR-only baseline, and achieves comparable performance. For intelligent transportation, far-range perception with roadside sensors is vital. The main challenge of far-range perception is performing accurate object detection and tracking under far distances (e.g., > 150m) at a low cost. To cope with such challenges, deploying both millimeter wave Radars and high-definition cameras, and fusing their data has become a common practice. Towards this goal, the first question is to conduct the association on the 2D image plane or the BEV plane. We argue that the former is more suitable because the magnitude of location errors in the perspective projection points is smaller at far distances on the 2D plane, leading to more accurate association. Thus, we first project the Radar points to the 2D plane and then associate them with the camera-based 2D object locations. Subsequently, we map the camera-based object locations to the BEV plane through inverse projection mapping (IPM) with the corresponding depth information from the Radar data. Finally, we engage a BEV tracking module to generate target trajectories. Our system is capable of achieving an average location accuracy of 1.3m when we extend the detection range up to 500m.
Yao Li 0016
IPSN1
2024 CRPlace: Camera-Radar Fusion with BEV Representation for Place Recognition
abstract
The integration of complementary characteristics from camera and radar data has emerged as an effective approach in 3D object detection. However, such fusion-based methods remain unexplored for place recognition, an equally important task for autonomous systems. Given that place recognition relies on the similarity between a query scene and the corresponding candidate scene, the stationary background of a scene is expected to play a crucial role in the task. As such, current well-designed camera-radar fusion methods for 3D object detection can hardly take effect in place recognition because they mainly focus on dynamic foreground objects. In this paper, a background-attentive camera-radar fusion-based method, named CRPlace, is proposed to generate background-attentive global descriptors from multi-view images and radar point clouds for accurate place recognition. To extract stationary background features effectively, we design an adaptive module that generates the background-attentive mask by utilizing the camera BEV feature and radar dynamic points. With the guidance of a background mask, we devise a bidirectional cross-attention-based spatial fusion strategy to facilitate comprehensive spatial interaction between the background information of the camera BEV feature and the radar BEV feature. As the first camera-radar fusion-based place recognition network, CRPlace has been evaluated thoroughly on the nuScenes dataset. The results show that our algorithm outperforms a variety of baseline methods across a comprehensive set of metrics (recall@1 reaches 91.2%).
Shaowei Fu, Yifan Duan, Yao Li 0016, Chengzhen Meng, Jianmin Ji, Yanyong Zhang
IROS3
2024 RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric Strategies
abstract
The recent advances in query-based multi-camera 3D object detection are featured by initializing object queries in the 3D space, and then sampling features from perspective-view images to perform multi-round query refinement. In such a framework, query points near the same camera ray are likely to sample similar features from very close pixels, resulting in ambiguous query features and degraded detection accuracy. To this end, we introduce RayFormer, a camera-ray-inspired query-based 3D object detector that aligns the initialization and feature extraction of object queries with the optical characteristics of cameras. Specifically, RayFormer transforms perspective-view image features into bird's eye view (BEV) via the lift-splat-shoot method and segments the BEV map to sectors based on the camera rays. Object queries are uniformly and sparsely initialized along each camera ray, facilitating the projection of different queries onto different areas in the image to extract distinct features. Besides, we leverage the instance information of images to supplement the uniformly initialized object queries by further involving additional queries along the ray from 2D object detection boxes. To extract unique object-level features that cater to distinct queries, we design a ray sampling method that suitably organizes the distribution of feature sampling points on both images and bird's eye view. Extensive experiments are conducted on the nuScenes dataset to validate our proposed ray-inspired model design. The proposed RayFormer achieves 55.5% mAP and 63.3% NDS, respectively.
Xiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan, Yao Li 0016, Yanyong Zhang
ACM Multimedia5
2024 FARFusion V2: A Geometry-based Radar-Camera Fusion Method on the Ground for Roadside Far-Range 3D Object Detection
abstract
Fusing the data of millimeter-wave Radar sensors and high-definition cameras has emerged as a viable approach to achieving precise 3D object detection for roadside traffic surveillance. For roadside perception systems, earlier studies have pointed out that it is better to perform the fusion on the 2D image plane than on the BEV plane (which is popular for on-car perception systems), especially when the perception range is large (e.g., >150m). Image-plane fusion requires critical transformations, like perspective projection from the Radar's BEV to the camera's 2D plane and reverse IPM. However, real-world issues like uneven terrain and sensor movement degrade these transformations' precision, impacting fusion effectiveness. To alleviate these issues, we propose a geometry-based Radar-camera fusion method on the ground, namely FARFusion V2. Specifically, we extend the ground-plane assumption in FARFusion[20] to support arbitrary shapes by formulating the ground height as an implicit representation based on geometric transformations. By incorporating the ground information, we can enhance Radar data with target height measurements. Consequently, we can thus project the enhanced Radar data onto the 2D plane to obtain more accurate depth information, thereby assisting the IPM process. A real-time parameterized transformation parameters estimation module is further introduced to refine the view transformation processes. Moreover, considering various measurement noises across these two sensors, we introduce an uncertainty-based depth fusion strategy into the 2D fusion process to maximize the probability of obtaining the optimal depth value. Extensive experiments are conducted on our collected roadside OWL benchmark, demonstrating the excellent localization capacity of FARFusion V2 in far-range scenarios. Our method achieves an average location accuracy of 0.771m when we extend the detection range up to 500m.
Yao Li 0016, Jiajun Deng, Yingjie Wang 0004, Xiaomeng Chu, Jianmin Ji, Yanyong Zhang
ACM Multimedia1
2024 In-Loop Filtering via Trained Look-Up Tables
abstract
In-loop filtering (ILF) is a key technology in image/video coding for reducing the artifacts. Recently, neural network-based in-loop filtering methods achieve remarkable coding gains beyond the capability of advanced video coding standards, establishing themselves a promising candidate tool for future standards. However, the utilization of deep neural networks (DNN) brings high computational complexity and raises high demand of dedicated hardware, which is challenging to apply into general use. To address this limitation, we study an efficient in-loop filtering scheme by adopting look-up tables (LUTs). After training a DNN with a predefined reference range for in-loop filtering, we cache the output values of the DNN into a LUT via traversing all possible inputs. In the coding process, the filtered pixel is generated by locating the input pixels (to-be-filtered pixel and reference pixels) and interpolating between the cached values. To further enable larger reference range within the limited LUT storage, we introduce an enhanced indexing mechanism in the filtering process, and a clipping/finetuning mechanism in the training. The proposed method is implemented into the Versatile Video Coding (VVC) reference software, VTM-11.0. Experimental results show that the proposed method, with three different configurations, achieves on average 0.13%∼0.51%, and 0.10% ∼0.39% BD-rate reduction under the all-intra (AI) and random-access (RA) configurations respectively. The proposed method incurs only 1% ∼8% time increase, an additional computation of 0.13 ∼0.93 kMAC/pixel, and 164 ∼1148 KB storage cost for a single model. Our method has explored a new and more practical approach for neural network-based ILF.
Zhuoyuan Li 0001, Jiacheng Li 0004, Yao Li 0016, Li Li 0040, Dong Liu 0002, Feng Wu 0001
VCIP3
2024 Uniformly Accelerated Motion Model for Inter Prediction
abstract
Inter prediction is a key technology in video coding to reduce the temporal redundancy. In natural videos, there are usually moving objects with changing velocity, resulting in complex motion fields that are difficult to represent compactly. In Versatile Video Coding (VVC), existing inter prediction methods usually assume uniform speed motion between consecutive frames, which may not well handle the complex motion fields in the real world. To address these issues, we introduce a uniformly accelerated motion model (UAMM) to exploit velocity and acceleration of moving objects between the video frames, and further combine them to assist in the inter prediction methods to handle the motion change in the temporal domain. First, we review the theory of UAMM. Second, we propose UAMM-based parameter derivation and extrapolation schemes in the coding process. Third, we integrate the UAMM into existing inter prediction modes (Merge, MMVD, CIIP) to achieve higher prediction accuracy. The proposed method is implemented into the VVC reference software, VTM version 12.0. Experimental results show that the proposed method achieves up to 0.38% BD-rate reduction compared to the VTM anchor, under the Low-delay P configuration, with a slight increase of time complexity on the encoding/decoding side.
Zhuoyuan Li 0001, Yao Li 0016, Chuanbo Tang, Li Li 0040, Dong Liu 0002, Feng Wu 0001
VCIP2
2023 Bi-LRFusion: Bi-Directional LiDAR-Radar Fusion for 3D Dynamic Object Detection
abstract
LiDAR and Radar are two complementary sensing approaches in that LiDAR specializes in capturing an object's 3D shape while Radar provides longer detection ranges as well as velocity hints. Though seemingly natural, how to efficiently combine them for improved feature representation is still unclear. The main challenge arises from that Radar data are extremely sparse and lack height information. Therefore, directly integrating Radar features into LiDAR-centric detection networks is not optimal. In this work, we introduce a bi-directional LiDAR-Radar fusion framework, termed Bi-LRFusion, to tackle the challenges and improve 3D detection for dynamic objects. Technically, Bi-LRFusion involves two steps: first, it enriches Radar's local features by learning important details from the LiDAR branch to alleviate the problems caused by the absence of height information and extreme sparsity; second, it combines LiDAR features with the enhanced Radar features in a unified bird's-eye-view representation. We conduct extensive experiments on nuScenes and ORR datasets, and show that our Bi-LRFusion achieves state-of-the-art performance for detecting dynamic objects. Notably, Radar data in these two datasets have different formats, which demonstrates the generalizability of our method. Codes will be published.
Yingjie Wang 0005, Jiajun Deng, Yao Li 0016, Jinshui Hu, Cong Liu 0006, Yu Zhang 0086, Jianmin Ji, Wanli Ouyang, Yanyong Zhang
CVPR3
2023 CluB: Cluster Meets BEV for LiDAR-Based 3D Object Detection
abstract
Currently, LiDAR-based 3D detectors are broadly categorized into two groups, namely, BEV-based detectors and cluster-based detectors. BEV-based detectors capture the contextual information from the Bird's Eye View (BEV) and fill their center voxels via feature diffusion with a stack of convolution layers, which, however, weakens the capability of presenting an object with the center point. On the other hand, cluster-based detectors exploit the voting mechanism and aggregate the foreground points into object-centric clusters for further prediction. In this paper, we explore how to effectively combine these two complementary representations into a unified framework. Specifically, we propose a new 3D object detection framework, referred to as CluB, which incorporates an auxiliary cluster-based branch into the BEV-based detector by enriching the object representation at both feature and query levels. Technically, CluB is comprised of two steps. First, we construct a cluster feature diffusion module to establish the association between cluster features and BEV features in a subtle and adaptive fashion. Based on that, an imitation loss is introduced to distill object-centric knowledge from the cluster features to the BEV features. Second, we design a cluster query generation module to leverage the voting centers directly from the cluster branch, thus enriching the diversity of object queries. Meanwhile, a direction loss is employed to encourage a more accurate voting center for each cluster. Extensive experiments are conducted on Waymo and nuScenes datasets, and our CluB achieves state-of-the-art performance on both benchmarks.
Yingjie Wang 0005, Jiajun Deng, Yuenan Hou, Yao Li 0016, Yu Zhang 0086, Jianmin Ji, Wanli Ouyang, Yanyong Zhang
NeurIPS4
2023 TrajMatch: Toward Automatic Spatio-Temporal Calibration for Roadside LiDARs Through Trajectory Matching
abstract
Recently, deploying sensors such as LiDARs on the roadside to monitor the passing traffic and assist autonomous vehicle perception has become popular. However, unlike autonomous vehicle systems, roadside sensor systems involve sensors from different subsystems, resulting in a lack of synchronization in both time and space between the sensors. Calibration is a critical technology that enables the central server to fuse data generated by different location infrastructures, which vastly improves sensing range and detection robustness. Regrettably, existing calibration algorithms frequently assume that LiDARs have significant overlap or that temporal calibration has already been achieved. However, since these assumptions do not always hold in real-world scenarios, the calibration results obtained from existing algorithms are frequently unsatisfactory. In this paper, we propose TrajMatch - the first system that can automatically calibrate roadside LiDARs in both time and space. The main idea is to automatically calibrate the sensors based on the result of the detection/tracking task, rather than relying on extracting special features. Furthermore, we propose a novel mechanism for evaluating calibration parameters that align with our algorithm, and we demonstrate its effectiveness through experiments. This mechanism can also guide parameter iterations for multiple calibrations, further enhancing the accuracy and efficiency of our calibration method. Finally, to evaluate the performance of TrajMatch, we collected two datasets, one simulated dataset LiDARnet-sim 1.0 and one real-world dataset. The experimental results show that TrajMatch can achieve a spatial calibration error of less than$10cm$and a temporal calibration error of less than$1.5ms$.
Haojie Ren, Sha Zhang 0002, Sugang Li, Yao Li 0016, Xinchen Li, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang
IEEE Trans. Intell. Transp. Syst.4
2022 Global Homography Motion Compensation for Versatile Video Coding
abstract
In Versatile Video Coding (VVC), local affine motion compensation (LAMC) is adopted to handle complex motions, such as rotation and zooming. However, it is inefficient to use LAMC to handle the global motion due to the following two reasons. First, the use of LAMC may lead to some extra bit cost on the affine motion model parameters. Second, the precision of LAMC is restricted by the MV precision of the control points. Therefore, in this paper, we propose a global homography motion compensation (GHMC) framework to better characterize the global motion. For each coding block, an extra mode is added to perform motion compensation based on an 8-parameter global homography motion model. In addition, an extrapolation scheme is designed to derive the parameters from reference frames to save the bit cost for signaling them. The proposed framework is implemented into the VVC reference software VTM-6.0. Experimental results show that, on average, 0.69% and 0.66% BD-rate reduction is achieved under Low Delay P and Low Delay B configurations, respectively, for sequences with rich complex global motions.
Yao Li 0016, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002, Houqiang Li
VCIP1
2021 Neighbor-Vote: Improving Monocular 3D Object Detection through Neighbor Distance Voting
abstract
As cameras are increasingly deployed in new application domains such as autonomous driving, performing 3D object detection on monocular images becomes an important task for visual scene understanding. Recent advances on monocular 3D object detection mainly rely on the "pseudo-LiDAR'' generation, which performs monocular depth estimation and lifts the 2D pixels to pseudo 3D points. However, depth estimation from monocular images, due to its poor accuracy, leads to inevitable position shift of pseudo-LiDAR points within the object. Therefore, the predicted bounding boxes may suffer from inaccurate location and deformed shape. In this paper, we present a novel neighbor-voting method that incorporates neighbor predictions to ameliorate object detection from severely deformed pseudo-LiDAR point clouds. Specifically, each feature point around the object forms their own predictions, and then the "consensus'' is achieved through voting. In this way, we can effectively combine the neighbors' predictions with local prediction and achieve more accurate 3D detection. To further enlarge the difference between the foreground region of interest (ROI) pseudo-LiDAR points and the background points, we also encode the ROI prediction scores of 2D foreground pixels into the corresponding pseudo-LiDAR points. We conduct extensive experiments on the KITTI benchmark to validate the merits of our proposed method. Our results on the bird's eye view detection outperform the state-of-the-art performance, especially for the "hard" level detection. The code is available at https://github.com/cxmomo/Neighbor-Vote.
Xiaomeng Chu, Jiajun Deng, Yao Li 0016, Zhenxun Yuan, Yanyong Zhang, Jianmin Ji, Yu Zhang 0086
ACM Multimedia3