Lili Zhao 0001

dblp:54/5985-1 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0002-5182-7230ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 A Dynamic Point Cloud Dataset for MPEG Point Cloud Compression and Performance Analysis
abstract
Recent years witnessed the development in MPEG point cloud compression (PCC). However, the exploration of inter-frame coding may be impeded due to the lack of dynamic point clouds (point cloud sequences). To promote the development of PCC technology, we propose Dynamic3D , a dynamic 3D point cloud dataset with high-quality real-captured 3D persons and objects. There are several appealing properties: 1) Dynamic scenes: It contains five sequences and each sequence comprises 600 frames with temporal variation; 2) Complex content: instead of a single person or object in the existing dataset from MPEG, our established dataset contains multiple persons or both person and objects; 3) Realistic capture: the color industrial cameras and infrared cameras are used for data acquisition. This dataset provides the vast exploration space for PCC, especially the elimination of temporal redundancy. Extensive simulations are conducted on this dataset by using the reference software of MPEG G-PCC and V-PCC, i.e., (GeS-TM and TMC2), delivering observations, analysis and opportunities for the future research of PCC.
Lili Zhao 0001, Qian Yin 0002, Lancao Ren, Lei Yang 0063, Chuanmin Jia, Siwei Ma 0001
DCC1
2024 Towards Robust Visual Localization Using Multi-View Images and HD Vector Map
abstract
Robust and accurate localization is highly desired in intelligent driving and robotic navigation. Existing methods highly rely on feature maps and complex parameter tuning, while suffering from ineffective data association, heavy computation, high dependency on training data and low robustness. In this paper, we propose a high-robust and cost-effective visual localization system, which jointly exploits the semantic information of Bird’s-Eye-View (BEV) representation from multi-view images and the vectorized High Definition (HD) map. We formulate the visual localization as cross-modal data association issue and innovatively project the vectorized landmarks of HD map into BEV semantic map. Finally, the highly accurate vehicle’s pose can be estimated by pose optimization based on direct image alignment. Extensive simulations experimented on nuScenes dataset show that the proposed method can deliver robust and accurate localization results under various scenarios. In addition, the proposed system is convenient for large-scale deployment and has been tested on the commercial test car.
Lili Zhao 0001, Zhili Liu, Qian Yin 0002, Lei Yang 0063, Meng Guo 0007
ICIP1
2023 Learning Spatial-Temporal Embeddings for Sequential Point Cloud Frame Interpolation
abstract
A point cloud sequence is usually acquired at a low frame rate owing to the limitations from the sensing equipment. Consequently, the immersive experience of the virtual reality might be greatly degraded. To tackle this issue, a point cloud frame interpolation process can be used to increase the frame rate of the acquired point cloud sequence by generating new frames between the consecutive ones. However, it is still challenging for deep neural networks to synthesize high-fidelity point clouds, especially for those with complex geometric details and large motion. In this paper, a novel frame interpolation network is proposed, which jointly exploits the spatial features and flows. The key success of our method lies in the developed spatial-temporal feature propagation module and temporal-aware feature-to-point mapping module. The former effectively embeds the spatial features and scene flows into a spatial-temporal feature representation (STFR). The latter generates a much improved target frame from STFR. Extensive experimental results have demonstrated that our method has achieved the best performance in most cases.
Lili Zhao 0001, Zhuoqun Sun, Lancao Ren, Qian Yin 0002, Lei Yang 0063, Meng Guo 0007
ICIP1
2022 MVFI-Net: Motion-Aware Video Frame Interpolation Network
Xuhu Lin, Lili Zhao 0001
ACCV (3)2
2022 Rangeinet: Fast Lidar Point Cloud Temporal Interpolation
abstract
Due to the low scan rate of LiDAR sensors, LiDAR point cloud streams usually have a low frame rate, which is far below that of other sensors such as cameras. This could incur frame rate mismatch while conducting multi-sensor data fusion. LiDAR point cloud temporal interpolation aims to synthesize the non-existing intermediate frame between input frames to improve the frame rate of point clouds. However, the existing methods heavily depend on 3D scene flow or 2D flow estimation, which yield huge computational complexity and obstacles in real-time applications. To resolve this issue, we propose a fast and non-flow involved method, which analyzes the LiDAR point cloud by exploiting its corresponding 2D range images (RIs). Specifically, we develop a Siamese context extractor containing asymmetrical convolution kernels to learn the shape context and spatial feature of RIs, and the 3D space-time convolutions are introduced to precisely capture the temporal characteristics. Experimental results have clearly shown that our method is much faster than the state-of-the-art LiDAR point cloud temporal interpolation methods on various datasets, while delivering either comparable or superior frame interpolation performance.
Lili Zhao 0001, Xuhu Lin, Wenyi Wang 0005, Kai-Kuang Ma
ICASSP1
2022 Real-Time Scene-Aware LiDAR Point Cloud Compression Using Semantic Prior Representation
abstract
Existing LiDAR point cloud compression (PCC) methods tend to treat compression as afidelityissue, without sufficiently addressing itsmachine perceptionaspect. The latter issue is often encountered by the decoder agents that might aim to conduct scene-understanding related tasks only, such as computing the localization information. For tackling this challenge, a novel LiDAR PCC system is proposed to compress the point cloud geometry, which contains aback channelfor allowing the decoder to initiate such request to the encoder. The key success of our PCC method lies in our proposedsemantic prior representation(SPR) and its lossy encoding algorithm with variable precision to generate the final bitstream; the entire process is fast and achieves real-time performance. Note that our SPR is a compact and effective representation of three-dimensional (3D) input point clouds, and it consists oflabels, predictions, andresiduals. These information can be generated by first exploiting ascene-aware object segmentationto a set of 2D range images (frames) individually, which were generated from the 3D point clouds via a projection process. Based on the generated labels, the pixels associated with those moving objects are considered as noisy information and should be removed for not only saving bit budget on transmission but also, most importantly, improving the accuracy of localization computed at the decoder. Experimental results conducted on the commonly-used test dataset have shown that our proposed system outperforms the MPEG’s G-PCC (TMC13-v14.0) in a large bitrate range. In fact, the performance gap will become even larger when more and/or large moving objects are involved in the input point clouds.
Lili Zhao 0001, Kai-Kuang Ma, Zhili Liu, Qian Yin 0002
IEEE Trans. Circuits Syst. Video Technol.1
2021 An Unsupervised Optical Flow Estimation for Lidar Image Sequences
abstract
In recent years, the LiDAR images, as a 2D compact representation of 3D LiDAR point clouds, are widely applied in various tasks, e.g., 3D semantic segmentation, LiDAR point cloud compression (PCC). Among these works, the optical flow estimation for LiDAR image sequences has become a key issue, especially for the motion estimation of the inter prediction in PCC. However, the existing optical flow estimation models are likely to be unreliable for LiDAR images. In this work, we first propose a light-weight flow estimation model for LiDAR image sequences. The key novelty of our method lies in two aspects. One is that for the different characteristics (with the spatial-variation feature distribution) of the LiDAR images w.r.t. the normal color images, we introduce the attention mechanism into our model to improve the quality of the estimated flow. The other one is that to tackle the lack of large-scale LiDAR-image annotations, we present an unsupervised method, which directly minimizes the inconsistency between the reference image and the reconstructed image based on the estimated optical flow. Extensive experimental results have shown that our proposed model outperforms other mainstream models on the KITTI dataset, with much fewer parameters.
Xuezhou Guo, Xuhu Lin, Lili Zhao 0001, Zezhi Zhu
ICIP3
2021 Deep Inter Prediction via Reference Frame Interpolation for Blurry Video Coding
abstract
In High Efficiency Video Coding (HEVC), inter prediction is an important module for removing temporal redundancy. The accuracy of inter prediction is much affected by the similarity between the current and reference frames. However, for blurry videos, the performance of inter coding will be degraded by varying motion blur, which is derived from camera shake or the acceleration of objects in the scene. To address this problem, we propose to synthesize additional reference frame via the frame interpolation network. The synthesized reference frame is added into reference picture lists to supply more credible reference candidate, and the searching mechanism for motion candidates is changed accordingly. In addition, to make our interpolation network more robust to various inputs with different compression artifacts, we establish a new blurry video database to train our network. With the well-trained frame interpolation network, compared with the reference software HM-16.9, the proposed method achieves on average 1.55% BD-rate reduction under random access (RA) configuration for blurry videos, and also obtains on average 0.75% BD-rate reduction for common test sequences.
Zezhi Zhu, Lili Zhao 0001, Xuhu Lin, Xuezhou Guo
VCIP2
2019 PRED: A Parallel Network for Handling Multiple Degradations via Single Model in Single Image Super-Resolution
abstract
Existing SISR (single image super-resolution) methods mostly assume that a low-resolution (LR) image is bicubicly down-sampled from its high-resolution (HR) counterpart, which inevitably give rise to poor performance when the degradation is out of assumption. To address this issue, we propose a framework PRED (parallel residual and encoder-decoder network) with an innovative training strategy to enhance the robustness to multiple degradations. Consequently, the network can handle spatially variant degradations, which significantly improves the practicability of the proposed method. Extensive experimental results on real LR images show that the proposed method can not only produce favorable results on multiple degradations, but also reconstruct visually plausible HR images.
Guangyang Wu, Lili Zhao 0001, Wenyi Wang 0005, Liaoyuan Zeng
ICIP2
2019 Efficient Screen Content Coding Based on Convolutional Neural Network Guided by a Large-Scale Database
abstract
Screen content videos (SCVs) are becoming popular in many applications. Compared with natural content videos (NCVs), the SCVs have different characteristics. Therefore, the screen content coding (SCC) based on HEVC adopts some new coding tools (intra block copy and palette mode etc.) to improve coding efficiency, but these tools increase the computational complexity as well. In this paper, we propose to predict the CU partition of the SCVs by a convolutional neural network (CNN) which is trained by the large-scale database that we firstly established for screen content coding. The proposed approach is implemented in SCC reference software SCM-6.1. Experimental results show that our proposed approach can save 53.2% encoding time with 2.67% BD-rate increase on average in All Intra (AI) configurations.
Lili Zhao 0001, Zhiwen Wei, Weitong Cai, Wenyi Wang 0005, Liaoyuan Zeng
ICIP1
2018 High Efficient VR Video Coding Based on Auto Projection Selection Using Transferable Features
abstract
Given multiple texture projection methods from the sphere surface to the planar surface, this paper proposes an adaptive selection mode that automatically chooses the appropriate projection method to obtain high compression efficiency of the VR video. The video compression efficiency is inherently affected by the video content, which is closely related to the projection method in the case of VR video encoding. In order to represent the VR video content in a compact manner, a feature vector (transferable feature) for each frame is extracted by a Res-CNN which is pre-trained by a large scale data set for general classification. Afterwards, the relation between the feature and the optimal projection method is investigated by using PCA-KNN, which can project the initial feature vector to a subspace where the VR videos can be efficiently classified with low ambiguity. The experimental results show that the proposed method can select the appropriate projection method that generates the best BD rate.
Lili Zhao 0001, Wenyi Wang 0005, Rumin Zhang, Liaoyuan Zeng
VCIP1