VLDB 2026 Research / reviewers in the wild / expert
Xiaolong Mao
dblp:54/7751
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
0009-0003-4354-5959ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PCAC-GAN: A Sparse-Tensor-Based Generative Adversarial Network for 3D Point Cloud Attribute CompressionabstractLearning-based methods have proven successful in compressing geometric information for point clouds. For attribute compression, however, they still lag behind non-learning-based methods such as the MPEG G-PCC standard. To bridge this gap, we propose a novel deep learning-based point cloud attribute compression method that uses a generative adversarial network (GAN) with sparse convolution layers. Our method also includes a module that adaptively selects the resolution of the voxels used to voxelize the input point cloud. Sparse vectors are used to represent the voxelized point cloud, and sparse convolutions process the sparse tensors, ensuring computational efficiency. To the best of our knowledge, this is the first application of GANs to compress point cloud attributes. Our experimental results show that our method outperforms existing learning-based techniques and rivals the latest G-PCC test model (TMC13v23) in terms of visual quality. Xiaolong Mao, Hui Yuan 0001, Xin Lu 0001, Raouf Hamzaoui, Wei Gao 0003 |
Comput. Vis. Media | 1 |
| 2025 | A novel lightweight model combined with convolutional neural network and transformer for gearbox fault diagnosis using infrared thermal images
Xiao Zhuang, Jian Ge 0003, Xiaolong Mao, Di Zhou 0006, Hongbing Yao, Weifang Sun, Lin Li 0052, Jiawei Xiang |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | SPAC: Sampling-Based Progressive Attribute Compression for Dense Point CloudsabstractWe propose an end-to-end attribute compression method for dense point clouds. The proposed method combines a frequency sampling module, an adaptive scale feature extraction module with geometry assistance, and a global hyperprior entropy model. The frequency sampling module uses a Hamming window and the Fast Fourier Transform to extract high-frequency components of the point cloud. The difference between the original point cloud and the sampled point cloud is divided into multiple sub-point clouds. These sub-point clouds are then partitioned using an octree, providing a structured input for feature extraction. The feature extraction module integrates adaptive convolutional layers and uses offset-attention to capture both local and global features. Then, a geometry-assisted attribute feature refinement module is used to refine the extracted attribute features. Finally, a global hyperprior model is introduced for entropy encoding. This model propagates hyperprior parameters from the deepest (base) layer to the other layers, further enhancing the encoding efficiency. At the decoder, a mirrored network is used to progressively restore features and reconstruct the color attribute through transposed convolutional layers. The proposed method encodes base layer information at a low bitrate and progressively adds enhancement layer information to improve reconstruction accuracy. Compared to the best anchor of the latest geometry-based point cloud compression (G-PCC) standard that was proposed by the Moving Picture Experts Group (MPEG), the proposed method can achieve an average Bjøntegaard delta bitrate of -24.58% for the Y component (resp. -21.23% for YUV components) on the MPEG Category Solid dataset and -22.48% for the Y component (resp. -17.19% for YUV components) on the MPEG Category Dense dataset. This is the first instance that a learning-based attribute codec outperforms the G-PCC standard on these datasets by following the common test conditions specified by MPEG. Our source code will be made publicly available on https://github.com/sduxlmao/SPAC. Xiaolong Mao, Hui Yuan 0001, Shiqi Jiang 0006, Raouf Hamzaoui, Sam Kwong |
IEEE Trans. Image Process. | 1 |
| 2025 | EdgeRegNet: Edge Feature-Based Multimodal Registration Network Between Images and LiDAR Point CloudsabstractCross-modal data registration has long been a critical task in computer vision, with extensive applications in autonomous driving and robotics. Accurate and robust registration methods are essential for aligning data from different modalities, forming the foundation for multimodal sensor data fusion and enhancing perception systems' accuracy and reliability. The registration task between 2D images captured by cameras and 3D point clouds captured by Light Detection and Ranging (LiDAR) sensors is usually treated as a visual pose estimation problem. High-dimensional feature similarities from different modalities are leveraged to identify pixel-point correspondences, followed by pose estimation techniques using least squares methods. However, existing approaches often resort to downsampling the original point cloud and image data due to computational constraints, inevitably leading to a loss in precision. Additionally, high-dimensional features extracted using different feature extractors from various modalities require specific techniques to mitigate cross-modal differences for effective matching. To address these challenges, we propose a method that uses edge information from the original point clouds and images for cross-modal registration. We retain crucial information from the original data by extracting edge points and pixels, enhancing registration accuracy while maintaining computational efficiency. The use of edge points and edge pixels allows us to introduce an attention-based feature exchange block to eliminate cross-modal disparities. Furthermore, we incorporate an optimal matching layer to improve correspondence identification. We validate the accuracy of our method on the KITTI and nuScenes datasets, demonstrating its state-of-the-art performance. Our code is publicly available on GitHub athttps://github.com/ESRSchao/EdgeRegNet. Yuanchao Yue, Hui Yuan 0001, Qinglong Miao, Xiaolong Mao, Raouf Hamzaoui, Peter Eisert |
IEEE Trans. Multim. | 4 |
| 2025 | CS-Net: Contribution-Based Sampling Network for Point Cloud SimplificationabstractPoint cloud sampling plays a crucial role in reducing computation costs and storage requirements for various vision tasks. Traditional sampling methods, such as farthest point sampling, lack task-specific information and, as a result, cannot guarantee optimal performance in specific applications. Learning-based methods train a network to sample the point cloud for the targeted downstream task. However, they do not guarantee that the sampled points are the most relevant ones. Moreover, they may result in duplicate sampled points, which requires completion of the sampled point cloud through post-processing techniques. To address these limitations, we propose a contribution-based sampling network (CS-Net), where the sampling operation is formulated as a Top-$k$k operation. To ensure that the network can be trained in an end-to-end way using gradient descent algorithms, we use a differentiable approximation to the Top-$k$k operation via entropy regularization of an optimal transport problem. Our network consists of a feature embedding module, a cascade attention module, and a contribution scoring module. The feature embedding module includes a specifically designed spatial pooling layer to reduce parameters while preserving important features. The cascade attention module combines the outputs of three skip connected offset attention layers to emphasize the attractive features and suppress less important ones. The contribution scoring module generates a contribution score for each point and guides the sampling process to prioritize the most important ones. Experiments on the ModelNet40 and PU147 showed that CS-Net achieved state-of-the-art performance in two semantic-based downstream tasks (classification and registration) and two reconstruction-based tasks (compression and surface reconstruction). CS-Net also achieved high average precision for objection detection on the KITTI LiDAR point cloud dataset, demonstrating its effectiveness in three-dimensional object detection. Chen Chen 0063, Hui Yuan 0001, Xiaolong Mao, Raouf Hamzaoui, Junhui Hou |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Enhancing Octree-Based Context Models for Point Cloud Geometry Compression With Attention-Based Child Node Number PredictionabstractIn point cloud geometry compression, most octree-based context models use the cross-entropy between the one-hot encoding of node occupancy and the probability distribution predicted by the context model as the loss. This approach converts the problem of predicting the number (a regression problem) and the position (a classification problem) of occupied child nodes into a 255-dimensional classification problem. As a result, it fails to accurately measure the difference between the one-hot encoding and the predicted probability distribution. We first analyze why the cross-entropy loss function fails to accurately measure the difference between the one-hot encoding and the predicted probability distribution. Then, we propose an attention-based child node number prediction (ACNP) module to enhance the context models. The proposed module can predict the number of occupied child nodes and map it into an 8-dimensional vector to assist the context model in predicting the probability distribution of the occupancy of the current node for efficient entropy coding. Experimental results demonstrate that the proposed module enhances the coding efficiency of octree-based context models. Hui Yuan 0001, Xiaolong Mao, Xin Lu 0001, Raouf Hamzaoui |
IEEE Signal Process. Lett. | 3 |
| 2023 | Fourier Series and Laplacian Noise-Based Quantization Error Compensation for End-to-End Learning-Based Image CompressionabstractQuantization is a core operation in lossy image compression. In the end-to-end learning-based image compression framework, quantization is conducted by a rounding operation during test, while it is replaced by additive uniform noise during training, leading to a mismatched problem between train and test. To address this problem, we propose a quantization error compensation method for the end-to-end learning-based image compression framework. The method uses Fourier series to approximate the periodic changes of the quantization error, and adds Laplacian noise to the quantized latent during test. The proposed method can be flexibly combined with different end-to-end learning-based image compression methods. Experimental results show that higher coding efficiency can be achieved by adding the proposed method with the state-of-the-art methods. Shiqi Jiang 0006, Hui Yuan 0001, Shuai Li 0005, Xiaolong Mao |
ICIP | 4 |