EDBT 2026 Demo / reviewers in the wild / expert
Liang Xie 0013
dblp:81/2806-13
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
0009-0001-0973-8402ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Perception Bias: A Training-Free Approach to Enhance LMM for Image Quality AssessmentabstractDespite the impressive performance of large multimodal models (LMMs) in high-level visual tasks, their capacity for image quality assessment (IQA) remains limited. One main reason is that LMMs are primarily trained for high-level tasks (e.g., image captioning), emphasizing unified image semantics extraction under varied quality. Such semantic-aware yet quality-insensitive perception bias inevitably leads to a heavy reliance on image semantics when those LMMs are forced for quality rating. In this paper, instead of retraining or tuning an LMM costly, we propose a training-free debiasing framework, in which the image quality prediction is rectified by mitigating the bias caused by image semantics. Specifically, we first explore several semantic-preserving distortions that can significantly degrade image quality while maintaining identifiable semantics. By applying these specific distortions to the query/test images, we ensure that the degraded images are recognized as poor quality while their semantics remain. During quality inference, both a query image and its corresponding degraded version are fed to the LMM along with a prompt indicating that the query image quality should be inferred under the condition that the degraded one is deemed poor quality. This prior condition effectively aligns the LMM’s quality perception, as all degraded images are consistently rated as poor quality, regardless of their semantic difference. Finally, the quality scores of the query image inferred under different prior conditions (degraded versions) are aggregated using a conditional probability model. Extensive experiments on various IQA datasets show that our debiasing framework could consistently enhance the LMM performance and the code will be publicly available. Baoliang Chen, Siyi Pan, Dongxu Wu, Liang Xie 0013, Xiangjie Sui, Lingyu Zhu 0006, Hanwei Zhu |
AAAI | 4 |
| 2026 | Temporal Quality Aggregation for VQA: Benchmark and Psychology-Inspired Model
Baoliang Chen, Changsheng Gao, Lingyu Zhu 0006, Liang Xie 0013, Hanwei Zhu, Zhijian Hao |
QoMEX | 4 |
| 2026 | Slide Deformable Transformer for High-Precision LiDAR Point Cloud CompressionabstractDynamic LiDAR point cloud compression with range images aims to reduce storage and transmission costs while preserving both spatial accuracy and temporal consistency across frames. Vision Transformers (ViTs) are commonly used for cross-frame dependency modeling. However, they suffer from feature misalignment under cross-frame displacement due to fixed patch partitioning, and their global attention across all patches is costly yet ineffective for local motions. High-precision sequences also face precision loss when 16-bit range data are quantized in a single channel. To address these limitations, we propose a Slide Deformable Transformer framework for high-precision dynamic LiDAR point cloud compression, termed SDT-PCC. At its core, the proposed SDT layer restricts attention to local sliding windows, capturing fine-grained correspondences across consecutive frames. It integrates deformable convolution into cross-frame attention to adaptively sample motion-offset locations, thereby enhancing temporal alignment and motion modeling. We also propose a Radix-Decomposition Multi-Channel Quantizer (RDMCQ), which decomposes range values into multiple channels and progressively refines precision across radix levels. Consequently, these designs can produce more temporally-coherent, accurate and stable reconstructions. Experiments on the SemanticKITTI dataset show that SDT-PCC achieves high efficiency in dynamic point cloud compression. The code is available on https://github.com/SYSU-SAIL/SDT-PCC. Haoran Li 0009, Lian Xu, Liang Xie 0013, Wei Gao 0003, Zhenwen Ren, Ge Li 0002, Yulan Guo |
IEEE Trans. Image Process. | 3 |
| 2025 | DPCSet: A Large-scale Dynamic Point Cloud Dataset for Compression and PerceptionabstractThe increasing demand for large-scale, high-quality datasets in dynamic point cloud compression (PCC) and human visual perception research underscores the limitations of existing datasets, which are often constrained by limited scale and insufficient dynamism, hindering algorithm validation and perceptual analysis in complex scenarios. To address this gap, we present DPCSet, a comprehensive dynamic point cloud dataset designed to support advanced research in PCC, human perception, and related domains. Comprising 100 dynamic object point clouds-the largest collection of its kind-DPCSet includes 200-frame sequences with geometry and attribute information, capturing diverse object types across real and virtual environments. Organized into seven superclasses, the dataset ensures broad scenario coverage. By rigorous selection, format conversion, quantization, DPCSet delivers standardized, high-precision point cloud data. Evaluation of multiple compression algorithms on a curated subset demonstrates DPCSet's efficacy in assessing trade-offs between compression efficiency and quality loss, positioning it as a potential benchmark for PCC. Furthermore, just noticeable distortion (JND) experiments on a compression-distorted subset reveal distinct perceptual characteristics of dynamic point clouds, offering valuable insights for perception-driven compression algorithms. The dataset is released at https://openi.pcl.ac.cn/gaowx/DPCSet. Wenxu Gao, Liang Xie 0013, Kangli Wang, Jingxuan Su, Changhao Peng, Wei Gao 0003 |
ACM Multimedia | 2 |
| 2025 | Poster: Automatically Generating High-Precision Simulated Road Networking in Traffic ScenarioabstractExisting lane-level simulation road network generation is labor-intensive, resource-demanding, and costly due to the need for large-scale data collection and manual post-editing. To overcome these limitations, we propose automatically generating high-precision simulated road networks in traffic scenario, an efficient and fully automated solution. Initially, real-world road street view data is collected through open-source street view map platforms, and a large-scale lane line dataset is constructed to provide a robust foundation for subsequent analysis. Next, an end-to-end lane line detection approach based on deep learning is designed, where a neural network model is trained to accurately detect the number and spatial distribution of lane lines in street view images, enabling automated extraction of lane information. Subsequently, by integrating coordinate transformation and map matching algorithms, the extracted lane information from street views is fused with the foundational road topology obtained from open-source map service platforms, resulting in the generation of a high-precision lane-level simulation road network. This method significantly reduces the costs associated with data collection and manual editing while enhancing the efficiency and accuracy of simulation road network generation. It provides reliable data support for urban traffic simulation, autonomous driving navigation, and the development of intelligent transportation systems, offering a novel technical pathway for the automated modeling of large-scale urban road networks. Liang Xie 0013, Wenke Huang 0002 |
MobiCom | 1 |
| 2025 | Poster: Efficient Geometry Compression and Communication for 3D Gaussian Splatting Point CloudsabstractAs dynamic 3D scene representations grow increasingly complex, the exponential expansion of 3D Gaussian data creates significant storage and transmission bottlenecks, resulting in excessive memory demands. To address this issue, we propose adopting the AVS PCRM reference software for efficient compression of Gaussian point cloud geometry data. The strategy deeply integrates the advanced encoding capabilities of AVS PCRM into the i3DV platform, forming technical complementarity with the original rate-distortion optimization mechanism based on binary hash tables. On one hand, the hash table efficiently caches inter-frame Gaussian point transformation relationships, which allows for high-fidelity transmission within a 40 Mbps bandwidth constraint. On the other hand, AVS PCRM performs precise compression on geometry data. Experiment demonstrate that the framework maintains the advantages of fast rendering and high-quality synthesis in 3D Gaussian technology while achieving significant 10%-25% bitrate savings on universal test dataset. It provides a superior rate-distortion tradeoff solution for the transmission and interaction of volumetric video. Liang Xie 0013, Luyang Tang, Wei Gao 0003 |
MobiCom | 1 |
| 2025 | Deep Learning-Based Point Cloud Compression: An In-Depth Survey and BenchmarkabstractWith the maturity of 3D capture technology, the explosive growth of point cloud data has burdened the storage and transmission process. Traditional hybrid point cloud compression (PCC) tools relying on handcrafted priors have limited compression performance and are increasingly weak in addressing the burden induced by data growth. Recently, deep learning-based PCC methods have been introduced to continue to push the PCC performance boundary. With the thriving of deep PCC, the community urgently demands a systematic overview to conclude the past progress and present future research directions. In this paper, we have a detailed review that covers popular point cloud datasets, algorithm evolution, benchmarking analysis, and future trends. Concretely, we first introduce several widely-used PCC datasets according to their major properties. Then the algorithm evolution of existing studies on deep PCC, including lossy ones and lossless ones proposed for various point cloud types, is reviewed. Apart from academic studies, we also investigate the development of relevant international standards (i.e., MPEG standards and JPEG standards). To help have an in-depth understanding of the advance of deep PCC, we select a representative set of methods and conduct extensive experiments on multiple datasets. Comprehensive benchmarking comparisons and analysis reveal the pros and cons of previous methods. Finally, based on the profound analysis, we highlight the challenges and future trends of deep learning-based PCC, paving the way for further study. Wei Gao 0003, Liang Xie 0013, Songlin Fan, Ge Li 0002, Shan Liu 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | PDNet: Parallel Dual-branch Network for Point Cloud Geometry Compression and AnalysisabstractIntegrating compression with analysis for point clouds poses a formidable challenge due to the inherent tension between the primary goals of compression for a compact representation and analysis for rich semantic retention. To alleviate this gap and maximize the practical requirements, we introduce a Parallel Dual-branch Network (PDNet) for lossy point cloud geometry compression, whose outputs are also analysis-friendly. The proposed method uses a novel Transformer-based encoder-decoder framework to incorporate local and global attention for point cloud latent representation computation. Specifically, the encoder comprises a Multi-scale Local-Global Feature extraction (MLGF) block to capture compact local and global latent features. The decoding and the hyper-prior modules employ a Transformer with No Position Embedding (TNPE) block and a Multilayer Perceptron (MLP) layer to reconstruct point clouds accurately. Furthermore, our method allows simultaneous point cloud analysis based on the compressed bitstream, such as point cloud classification. Experimental results demonstrate that our PDNet achieves nearly a 40% BD-Rate gain compared to G-PCC and other point-based compression counterparts. Besides, a 26% accuracy improvement in instance classification is observed compared to reconstructed point cloud classification. Liang Xie 0013, Wei Gao 0003, Songlin Fan, Zhaojian Yao |
DCC | 1 |
| 2024 | LearningPCC: A PyTorch Library for Learning-Based Point Cloud CompressionabstractThree-dimensional point cloud data is one of the most extensively used data representation today, favored in various fields for its realistic and lifelike visual effects. However, the substantial volume of data poses significant challenges for storage and transmission. To advance point cloud compression (PCC) technology, we develop a learning-based PCC algorithm library, namely LearningPCC. To our knowledge, this is the first comprehensive set of algorithms that is compatible with all types of point cloud data. This PyTorch library incorporates eleven learning-based algorithms that address both geometry and attribute compression of point cloud data. We categorize the existing methods into six main classes and thoroughly introduce and analyze the principles of these algorithms. Moreover, we conduct performance evaluations using point clouds with various densities, offering detailed test results on several compression metrics, such as RD curves, BD-BR gains, compression ratio improvements, and encoding times. We will provide researchers with convenient access to these methods, replicate codes, and experiment results. Our commitment includes maintaining and updating these algorithms to offer researchers the latest in compression technologies. Liang Xie 0013, Wei Gao 0003 |
ACM Multimedia | 1 |
| 2024 | PCHMVision: An Open-Source Library of Point Cloud Compression for Human and Machine VisionabstractIn today's era, three-dimensional point cloud data is not only voluminous but also widely applicable. Therefore, data compression has become a crucial step prior to processing. Although existing 3D point cloud compression techniques primarily focus on fidelity, in practical applications, the vast majority of compressed data serves machine perception tasks. Therefore, point cloud compression tailored for machine perception becomes particularly significant. To address this problem, we introduce an innovative point cloud compression algorithm library specifically designed for both machine and human perceptual requirements. This library represents the first collection of multi-perception point cloud compression algorithms on the PyTorch platform, integrating eleven advanced, learning-based algorithms. We category and analyze these algorithms in depth, according to different analysis tasks, to facilitate a better understanding and comparison. Moreover, we successfully replicate these algorithms and meticulously organize the pre-processing of point cloud data and the analysis networks for downstream tasks. Ultimately, we conduct experiments on multiple perceptual datasets for compression and analysis tasks, with results comprehensively summarized across various performance metrics. We will continue to update these algorithms to ease their adoption by researchers. Liang Xie 0013, Wei Gao 0003 |
ACM Multimedia | 1 |
| 2024 | ROI-Guided Point Cloud Geometry Compression Towards Human and Machine VisionabstractPoint cloud data is pivotal in applications like autonomous driving, virtual reality, and robotics. However, its substantial volume poses significant challenges in storage and transmission. In order to obtain a high compression ratio, crucial semantic details usually confront severe damage, leading to difficulties in guaranteeing the accuracy of downstream tasks. To tackle this problem, we are the first to introduce a novel Region of Interest (ROI)-guided Point Cloud Geometry Compression (RPCGC) method for human and machine vision. Our framework employs a dual-branch parallel structure, where the base layer encodes and decodes a simplified version of the point cloud, and the enhancement layer refines this by focusing on geometry details. Furthermore, the residual information of the enhancement layer undergoes refinement through an ROI prediction network. This network generates mask information, which is then incorporated into the residuals, serving as a strong supervision signal. Additionally, we intricately apply these mask details in the Rate-Distortion (RD) optimization process, with each point weighted in the distortion calculation. Our loss function includes RD loss and detection loss to better guide point cloud encoding for the machine. Experiment results demonstrate that RPCGC achieves exceptional compression performance and better detection accuracy (10% gain) than some learning-based compression methods at high bitrates in ScanNet and SUN RGB-D datasets. Liang Xie 0013, Wei Gao 0003, Huiming Zheng, Ge Li 0002 |
ACM Multimedia | 1 |
| 2022 | OpenPointCloud: An Open-Source Algorithm Library of Deep Learning Based Point Cloud CompressionabstractThis paper gives an overview of OpenPointCloud, the first open-source algorithm library containing outstanding deep learning methods on point cloud compression (PCC). We provide an introduction of our implementations, including 8 methods on lossless geometry PCC and lossy geometry PCC. Principles and contributions of these methods in our algorithm library are illustrated, which are also implemented with different deep learning programming frameworks, such as TensorFlow, Pytorch and TensorLayer. In order to systematically evaluate the performances of all these methods, we conduct a comprehensive benchmarking test. We provide analyses and comparisons of their performances according to their categories and draw constructive conclusions. This algorithm library has been released at https://git.openi.org.cn/OpenPointCloud. Wei Gao 0003, Ge Li 0002, Huiming Zheng, Yuyang Wu, Liang Xie 0013 |
ACM Multimedia | 6 |