Zhuangzi Li

dblp:226/2683 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
7since 2021 · last 2026
0000-0002-2565-6810ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author
YearPublicationVenuePosition
2026 Post-Processing Geometry Enhancement for G-PCC Compressed LiDAR via Cylindrical Densification
abstract
The geometry-based point cloud compression algorithm achieves efficient compression and transmission for LiDAR point clouds with high sparsity. However, the low-bitrate mode results in severe geometry compression artifacts, which involve both point reduction and coordinate offset. To the best of our knowledge, this is the first attempt to directly enhance the geometry quality for compressed LiDAR point cloud (CLGE) in a post-processing manner. Our proposed method consists of two branches: cylindrical densification and adaptive refinement. The former adopts a multi-scale sparse convolution framework to effectively extract spatial features in the cylindrical coordinate system and generate dense candidate points quickly. Large asymmetric sparse convolution kernels are also designed to capture the shapes of different regions and objects. The latter branch refines the candidate points through several MLP layers, which takes the neighborhood features between the candidate points and the input points into account. Finally, the designed ring-based farthest point resampling serves as an effective alternative for achieving the target number while maintaining the geometry distribution. Extensive experiments conducted on several datasets verify the effectiveness of our approach under different compression artifact levels. Furthermore, our method is easily extended to upsampling and is robust to noise. In addition to the geometry signal quality improvement, the point cloud enhanced by our proposed method alleviates the performance degradation in object detection task due to compression distortion.
Wang Liu 0001, Zhuangzi Li, Ge Li 0002, Siwei Ma 0001, Sam Kwong, Wei Gao 0003
IEEE Trans. Image Process.2
2025 SPU+: Dimension Folding for Semantic Point Cloud Upsampling
abstract
Semantic Point Cloud Upsampling (SPU) aims to reconstruct a high-resolution (dense) 3D point cloud from a low-resolution (sparse) one, ensuring that the upsampled point cloud is easily recognizable by downstream tasks. Conventional upsampling architectures typically represent point clouds using high-dimensional feature vectors. However, we observe a dimensional bottleneck, where simply increasing the feature dimensionality does not necessarily improve performance on semantic tasks. This insight motivates us to explore more effective feature representations within upsampling networks. In this paper, we propose a novel SPU method called SPU+, which introduces dimension folding as an alternative strategy for handling high-dimensional features. Specifically, SPU+ decomposes each high-dimensional feature into several g-dimensional packages, allowing interactions among packages within the feature space. Guided by the principle of maximizing feature diversity, we determine that setting the package dimension to 3 yields optimal performance. To enable convolutional operations over these 3D packages, we present a 3D Residual Graph Convolution Block (3D-RGCB) that achieves high computational efficiency. Based on 3D-RGCBs, we design an upsampling network that incorporates three structural modes: pre-mode, middle-mode, and end-mode. Additionally, for large-scale upsampling, we develop a scaling-and-shuffling strategy that adaptively adjusts the spatial size of each 3D package. Finally, we analyze the covering number of the 3D package representation and compare it to traditional high-dimensional feature representations. Experiments on publicly available datasets demonstrate not only the effectiveness of dimension folding but also the state-of-the-art performance achieved by SPU+. Code is available at: https://github.com/lizhuangzi/SPU_plus.
Zhuangzi Li, Thomas H. Li, Shan Liu 0001, Ge Li 0002
IEEE Trans. Image Process.1
2025 S4R: Rethinking Point Cloud Sampling via Guiding Upsampling-Aware Perception
abstract
Point cloud sampling aims to derive a sparse point cloud from a relatively dense point cloud, which is essential for efficient data transmission and storage. While existing deep sampling methods prioritize preserving the perception of sampled point clouds for downstream networks, few studies have critically examined the rationale behind this goal. Specifically, we observe that sampling can lead to a perceptual degradation phenomenon in many influential downstream networks, impairing their ability to effectively process sampled point clouds. We theoretically reveal the nature of the phenomenon and attempt to construct a novel sampling target by uniting upsampling and perceptual reconstruction. Accordingly, we propose a Maximum A Posteriori (MAP) sampling framework named Sample for Reconstruct (S4R), which impels the sampling stage to infer upsampling-guided perception. In S4R, we design very simple but effective sampling and upsampling networks using residual-based graph convolutions and incorporate a pseudo-residual connection to introduce prior knowledge. This architecture takes advantage of reconstruction properties and allows the sampling network to be trained in an unsupervised manner. Extensive experiments on classical networks demonstrates the excellent performance of S4R compared with the previous sampling schemes and reveals its advantages on different point cloud downstream tasks, i.e., classification, reconstruction and segmentation.
Zhuangzi Li, Shan Liu 0001, Wei Gao 0003, Guanbin Li, Ge Li 0002
IEEE Trans. Multim.1
2024 PointELM: Fast Point Cloud Classification Using Deep Random Mapping Based Extreme Learning Machines
abstract
Designing an influential and instructive deep network has become a prominent research in point cloud analysis. However, deep networks inevitably require extensive training time and are sensitive to variational data distribution. In this paper, we observe that a randomly initialized network exhibits discriminative abilities and show training deep networks is not necessary for point cloud classification. Specifically, we propose a Point Extreme Learning Machine (PointELM), which initially extracts infant features from point clouds using a randomly weighted network. Subsequently, a frequency-domain mapping is designed to enhance the infant features. Finally, an ELM classifier is adopted to categorize the enhanced features and generate the output predictions. We evaluate PointELM on classical networks and highlight three advantages: (1) PointELM does not require backpropagation, enabling extremely fast training. Yet PointELM can still achieve promising performance. For example, the classification accuracy of a DGCNN-based PointELM just lowers a trained DGCNN about 2.8% on ModelNet40. (2) PointELM is a flexible method that can conveniently transfer a trained network to a new dataset with minimal performance loss. (3) PointELM can rapidly alleviate the performance degradation caused by feeding sampled point clouds to trained networks, indicating strong potential on adjusting data with various distributions. Code is available at https://github.com/lizhuangzi/PointELM.
Zhuangzi Li, Shan Liu 0001, Ge Li 0002
ICME1
2023 Semantic Point Cloud Upsampling
abstract
Downsampled sparse point clouds are beneficial for data transmission and storage, but they are detrimental for semantic tasks due to information loss. In this paper, we examine an upsampling methodology that significantly reconstructs sparse clouds’ semantic representations. Specifically, we propose a novel semantic point cloud upsampling (SPU) framework for sparse point cloud classification. An SPU consists of two networks, i.e. an upsampling network and a classification network. They are skillfully unified to intensify semantic representations acting on the upsampling process. In the upsampling network, we first propose a novel graph aggregation convolution to construct hierarchical relations on sparse point clouds. To enhance stability and diversity during point upsampling, we then combine point shuffling and pre-interpolation technologies to build an enhanced upsampling module. Furthermore, we adopt the semantic prior information provided by a sparse point cloud to enhance its upsampling quality. The prior information is applied to an attention mechanism that can highlight key positions of the point cloud. We investigate different loss functions and conduct experiments on classical deep point networks, which effectively demonstrate the promising performance of our framework.
Zhuangzi Li, Ge Li 0002, Thomas H. Li, Shan Liu 0001, Wei Gao 0003
IEEE Trans. Multim.1
2021 Information-Growth Attention Network for Image Super-Resolution
abstract
It is generally known that a high-resolution (HR) image contains more productive information compared with its low-resolution (LR) versions, so image super-resolution (SR) satisfies an information-growth process. Considering the property, we attempt to exploit the growing information via a particular attention mechanism. In this paper, we propose a concise but effective Information-Growth Attention Network (IGAN) that shows the incremental information is beneficial for SR. Specifically, a novel information-growth attention is proposed. It aims to pay attention to features involving large information-growth capacity by assimilating the difference from current features to the former features within a network. We also illustrate its effectiveness contrasted by widely-used self-attention using entropy and generalization analysis. Furthermore, existing channel-wise attention generation modules (CAGMs) have large informational attenuation due to directly calculating global mean for feature maps. Therefore, we present an innovative CAGM that progressively decreases feature maps' sizes, leading to more adequate feature exploitation. Extensive experiments also demonstrate IGAN outperforms state-of-the-art attention-aware SR approaches.
Zhuangzi Li, Ge Li 0002, Thomas H. Li, Shan Liu 0001, Wei Gao 0003
ACM Multimedia1
2021 Video super-resolution based on a spatio-temporal matching network
Xiaobin Zhu 0001, Zhuangzi Li, Jungang Lou, Qing Shen 0005
Pattern Recognit.2
2020 Attention-aware perceptual enhancement nets for low-resolution image classification
Xiaobin Zhu 0001, Zhuangzi Li, Xianbo Li
Inf. Sci.2
2020 Attention-aware invertible hashing network with skip connections
Shanshan Li 0005, Qiang Cai 0001, Zhuangzi Li, Hai-Sheng Li 0002, Naiguang Zhang, Xiaoyu Zhang 0002
Pattern Recognit. Lett.3
2019 Residual Invertible Spatio-Temporal Network for Video Super-Resolution
abstract
Video super-resolution is a challenging task, which has attracted great attention in research and industry communities. In this paper, we propose a novel end-to-end architecture, called Residual Invertible Spatio-Temporal Network (RISTN) for video super-resolution. The RISTN can sufficiently exploit the spatial information from low-resolution to high-resolution, and effectively models the temporal consistency from consecutive video frames. Compared with existing recurrent convolutional network based approaches, RISTN is much deeper but more efficient. It consists of three major components: In the spatial component, a lightweight residual invertible block is designed to reduce information loss during feature transformation and provide robust feature representations. In the temporal component, a novel recurrent convolutional model with residual dense connections is proposed to construct deeper network and avoid feature degradation. In the reconstruction component, a new fusion method based on the sparse strategy is proposed to integrate the spatial and temporal features. Experiments on public benchmark datasets demonstrate that RISTN outperforms the state-ofthe-art methods.
Xiaobin Zhu 0001, Zhuangzi Li, Xiaoyu Zhang 0002
AAAI2
2019 Deep Super-Resolution Hashing Network for Low-Resolution Image Retrieval
Zhuangzi Li, Naiguang Zhang, Xiaobin Zhu 0001, Peng Li 0035
ICIG (3)2
2019 Attention-Aware Invertible Hashing Network
Shanshan Li 0005, Qiang Cai 0001, Zhuangzi Li, Hai-Sheng Li 0002, Naiguang Zhang, Jian Cao 0003
ICIG (3)3
2019 Representative Feature Matching Network for Image Retrieval
abstract
Recent convolutional neural network (CNNs) have shown promising performance on image retrieval due to the powerful feature extraction capability. However, the potential relations of feature maps are not effectively exploited in the before CNNs, resulting in inaccurate feature representations. To address this issue, we excavate feature channel-wise realtions by a matching strategy to adaptively highlight informative features. In this paper, we propose a novel representative feature matching network (RFMN) for image hashing retrieval. Specifically, we propose a novel representative feature matching block (RFMB) that can match feature maps with their representative one. So, the significance of each feature map can be exploited according to the matching similarity. In addition, we also present an innovative pooling layer based on the representative feature matching to build relations of pooled features with unpooled features, so as to highlight the pooled features retained more valuable information. Extensive experiments show that our approach can promote the average results of conventional residual network more than 2.6% on Cifar-10 and 1.4% on NUS-WIDE dataset, meanwhile achieve the state-of-the-art performance.
Zhuangzi Li, Naiguang Zhang, Lei Wang 0101
MMAsia1
2019 Multi-Scale Invertible Network for Image Super-Resolution
abstract
Deep convolutional neural networks (CNNs) based image super-resolution approaches have reached significant success in recent years. However, due to the information-discarded nature of CNN, they inevitably suffer from information loss during the feature embedding process, in which extracted intermediate features cannot effectively represent or reconstruct the input. As a result, the super-resolved image will have large deviations in image structure with its low-resolution version, leading to inaccurate representations in some local details. In this study, we address this problem by designing an end-to-end invertible architecture that can reversely represent low-resolution images in any feature embedding level. Specifically, we propose a novel image super-resolution method, named multi-scale invertible network (MSIN) to keep information lossless and introduce multi-scale learning in a unified framework. In MSIN, a novel multi-scale invertible stack is proposed, which adopts four parallel branches to respectively capture features with different scales and keeps balanced information-interaction by branch shifting. In addition, we employee global and hierarchical feature fusion to learn elaborate and comprehensive feature representations, in order to further benefit the quality of final image reconstruction. We show the reversibility of the proposed MSIN, and extensive experiments conducted on benchmark datasets demonstrate the state-of-the-art performance of our method.
Zhuangzi Li, Shanshan Li 0005, Naiguang Zhang, Lei Wang 0101
MMAsia1
2019 Deep convolutional representations and kernel extreme learning machines for image classification
Xiaobin Zhu 0001, Zhuangzi Li, Xiaoyu Zhang 0002, Peng Li 0035, Lei Wang 0101
Multim. Tools Appl.2
2018 Image Classification Using Convolutional Neural Networks and Kernel Extreme Learning Machines
abstract
We know that convolutional neural networks are good at learning invariant features, but not always optimal for classification. Contrarily, Kernel Extreme Learning Machines (KELMs) are good at approximating any target continuous function with extremely fast speed, but cannot learn complicated invariances. In this paper, we propose a novel image classification framework, in which KELM instead of Softmax function is adopted as a classifier in the convolutional neural network (CNN) architecture for promoting the performance of image classification. Experiments conducted on the publicly available datasets demonstrate the superior performance of the proposed method.
Zhuangzi Li, Xiaobin Zhu 0001, Lei Wang 0101, Peiyu Guo
ICIP1
2018 Generative Adversarial Image Super-Resolution Through Deep Dense Skip Connections
abstract
Abstract Recently, image super‐resolution works based on Convolutional Neural Networks (CNNs) and Generative Adversarial Nets (GANs) have shown promising performance. However, these methods tend to generate blurry and over‐smoothed super‐resolved (SR) images, due to the incomplete loss function and powerless architectures of networks. In this paper, a novel generative adversarial image super‐resolution through deep dense skip connections (GSR‐DDNet), is proposed to solve the above‐mentioned problems. It aims to take advantage of GAN's ability of modeling data distributions, so that GSR‐DDNet can select informative feature representation and model the mapping across the low‐quality and high‐quality images in an adversarial way. The pipeline of the proposed method consists of three main components: 1) The generator of a novel dense skip connection network with the deep structure for learning robust mapping function is proposed to generate SR images from low‐resolution images; 2) The feature extraction network based on VGG‐19 is adopted to capture high frequency feature maps for content loss; and 3) The discriminator with Wasserstein distance is adopted to identify the overall style of SR and ground‐truth images. Experiments conducted on four publicly available datasets demonstrate the superiority against the state‐of‐the‐art methods.
Xiaobin Zhu 0001, Zhuangzi Li, Xiaoyu Zhang 0002, Hai-Sheng Li 0002, Lei Wang 0101
Comput. Graph. Forum2