EDBT 2026 Demo / reviewers in the wild / expert
Fan Liang 0001
dblp:19/7748-1
· DBLP profile ↗
22ranked-venue papers
0as first author
18since 2021 · last 2025
0000-0001-8724-7644ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 12 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fast Dynamic Point Cloud Geometry Compression via Resolution-Adaptive Context ModelabstractThe growing use of Dynamic Point Clouds (DPCs) requires efficient compression for high-quality, low-latency services. Recently, learning-based Point Cloud Compression (PCC) frameworks have made significant performance improvement. However, the high computational complexity of crossframe fusion of context poses a challenge for the application of dynamic PCC. Moreover, most of learning-based PCC frameworks need to train multiple models for various bitrates. We propose a resolution-adaptive dynamic PCC framework that addresses these limitations through: (1) A resolution-adaptive network architecture employing simple networks for high-resolution features and more complex networks for low-resolution features; (2) An inter-frame feature residual compression module with interpolation-based feature scaling for fine-grained rate control; (3) A two-step geometry mask compression module maintaining reconstruction quality across bitrates. Experiments demonstrate that our framework achieves faster processing than state-of-the-art frameworks while maintaining competitive rate-distortion performance and enabling single-model multi-rate operation. Fan Liang 0001 |
VCIP | 2 |
| 2025 | Hierarchical Distortion Learning for Fast Lossy Compression of Point CloudsabstractThe growth of 3D point cloud applications requires efficient compression techniques for high-quality and low-latency services. Recently, learning-based point cloud compression models have made significant progress. However, geometric distortion resulting from downsampling limits the feature depth within large-scale point clouds, thereby constraining the receptive field and suppressing the redundant removal. Moreover, the issues of computational efficiency and reconstruction quality still persist in the compression of large-scale point clouds. To address these challenges, we propose a hierarchical distortion learning framework for end-to-end lossy compression of point clouds. First, we design a feature residual compression module to efficiently transmit shallow semantics between the encoder and the decoder, which enables a lightweight design of our framework. Second, we introduce a geometry residual compression module to progressively complement the reconstruction distortion, avoiding the accumulation of geometric distortion. By integrating these two modules and employing sufficient downsampling processes, we develop a high-performance framework with a significantly enlarged receptive field and low computational cost. Extensive experiments demonstrate that our method achieves state-ofthe- art performance in geometry lossy compression, while delivering competitive performance in joint geometry and color lossy compression with fast running speed. Code is available athttps://github.com/pengpeng-yu/FastPCC. Pengpeng Yu, Ye Zhang 0037, Fan Liang 0001, Haoran Li 0009, Yulan Guo |
IEEE Trans. Multim. | 3 |
| 2024 | Efficient Point Cloud Attribute Compression Using Rich Parallelizable Context ModelabstractThe autoregressive context model has been proven effective in point cloud attribute compression. However, it suffers from unbearable decoding latency due to the limitations of serial decoding and the large scale of point clouds. In this paper, we propose a rich, parallelizable context model for point cloud attribute compression to speed up the decoding process. To further improve rate-distortion (RD) performance, we propose cross-coordinate and intra-coordinate attention modules to reduce the spatial redundancy of the latent representations. We validate our method on the large-scale Moving Picture Experts Group (MPEG) point cloud benchmarks, and demonstrate that our model achieves much lower decoding time than previous autoregression-based methods while maintaining similar RD performance. Ruishan Huang, Pengpeng Yu, Shaolin Liao, Fan Liang 0001 |
ICASSP | 4 |
| 2024 | Cross-Frame Integrated Prediction for Feature-Space Video CompressionabstractLearned video compression in the feature domain employs implicit motion compensation to acquire predicted features or compute feature residuals, effectively minimizing spatiotemporal redundancy in the reconstructed frames. In this paper, we propose a cross-frame integrated prediction (CIP) network for feature-space video compression. Leveraging multiple features in motion estimation and compensation, our approach enables more context-aware prediction. Specifically, on the one hand, we introduce a global feature extraction (GFE) module in motion estimation to extract the information of the current feature and multiple reference features, providing a high-quality offset map for deformable motion compensation. On the other hand, we utilize an attentional feature fusion (AFF) module for multiple predicted features in motion compensation, which is beneficial for preserving crucial details and adapting to diverse scenes and content. By passing the final predicted feature to the residual compression and frame reconstruction, we achieve a single end-to-end video compression framework avoiding laborious multi-stage training. Comprehensive experimental results show that the proposed method not only maintains a low number of model parameters but also achieves significant performance improvement in video compression tasks, especially in the case of high resolution and high bitrate. Hongxin Qiu, Zhidao Zhou, Fan Liang 0001 |
IJCNN | 4 |
| 2024 | PFR-VC: Learning-Based Video Compression Framework with Predicted Frame RefinementabstractLearning-based video compression has attracted more and more attention in recent years. Traditional video coding relies on block-based motion estimation and spatial frequency transformation. While these techniques can effectively compress videos, further enhancing the compression ratio becomes challenging. Introducing deep learning methods can overcome the limitations of manually designed algorithms. In this paper, we propose a learning-based video compression framework with Predicted Frame Refinement (PFR) to improve the compression efficiency. Firstly, a simple autoencoder is introduced to encode the motion information, eliminating the need for a complex optical-flow network. Then, we design a predicted frame refinement network with an attention feature fusion mechanism to generate predicted frames more suitable for extracting context. Finally, we introduce a context coding scheme to improve the compression ratio by jointly utilizing temporal prior and hyper prior. The entire network can be globally optimized and trained from scratch. The experimental result shows that the proposed compression framework outperforms previous methods. Our approach brings 31.2% more saved bit rate than x265 with veryslow preset. Our model also achieves a 7.1% gain in Multi-Scale Structural Similarity Index Measure (MS-SSIM) compared with the recent method proposed by Guo et al.(2023). Zhidao Zhou, Hongxin Qiu, Zhikai Liu, Wei Sun 0007, Fan Liang 0001 |
IJCNN | 5 |
| 2024 | Self-aware Cross-component Prediction Model Based on Template for Screen Content CodingabstractThe current video coding techniques in the field of reducing redundancy between luma and chroma components have limitations, as they often overlook cross-component correlations. Previous research has employed linear and multi-model linear models to capture cross-component correlations, which are not tailored for screen content sequences. To address this issue, this paper proposes a self-aware cross-component prediction method based on template for screen content coding. With the neighboring reference samples, four prediction models are derived, and chroma prediction values at the template are calculated with the models. The sum of absolute transformed difference (SATD) cost between chroma prediction values and chroma reconstruction values at the template is computed for each model. Subsequently, the model with the lowest SATD cost is determined to be the selected model, which is used to generate prediction values for the current chroma block. Notably, the selected model is adaptively determined at both the encoder and decoder sides consistently, without signaling a model index. Experimental results show that the proposed method achieves 0.73%, 1.62% and 1.75% bit-rate savings on Y, U and V components respectively over ECM 6.0, for class TGM (Text and Graphics with Motion) under All-Intra (AI) configuration. Hongxin Qiu, Zhikai Liu, Fan Liang 0001, Wei Sun 0007 |
ISCAS | 4 |
| 2024 | A Power-Law Transformation Approach for Template-Based Cross-Component PredictionabstractWhile current cross-component chroma prediction tools have achieved significant performance improvements, they still face challenges in handling the non-linear relationship between luma and chroma. The paper proposes a template-derived power-law cross-component prediction model, with the key advantage of improving the numerical distribution characteristics of pixels in an interpretable manner. It effectively compresses cross-component redundancy while avoiding overfitting. The model achieves BD-Rate gains of −0.03%, −1.33%, −1.38% under the All-Intra configuration. Zhikai Liu, Xin-Yi Cui, Wei Sun 0007, Fan Liang 0001 |
ISM | 5 |
| 2024 | PFT-ILF: In-loop Filter with Partition Feature Transform for Versatile Video CodingabstractThe new generation of video coding standards, Versatile Video Coding (VVC), integrates a range of loop filter mechanisms, notably the De-Blocking Filter (DBF), Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF). However, these traditional tools are handcrafted empirically and have limitations. Thus many CNN-based loop filters have been proposed to achieve better image quality. In this paper, we propose a novel network based on Coding Unit (CU) partition feature transformation, called PFT-ILF. Considering that the CU partition map contains image distortion information, our approach innovatively utilizes the CU partition map as prior information to guide filtering. And we design a Partition Feature Transform (PFT) layer, which uses partition features to generate modulation parameters pair for adjusting the features of several intermediate layers in the network. We integrate our proposed filter into the NNVC standard software VTM11.0_NNVC-4.0 and conduct ablation experiments. Under the all intra configuration, our method achieves Bjøntegaard-Delta Bit-Rate (BD-BR) reductions of 7.52%, 17.36%, and 18.65% for Y, U, and V components, respectively. Xin-Yi Cui, Zhikai Liu, Zhidao Zhou, Fan Liang 0001 |
VCIP | 5 |
| 2024 | Multi-stage Attention Network with Auxiliary Information Refinement for VVC In-loop FilteringabstractRecently, learning-based video compression techniques have brought significant performance improvements. However, most existing methods have not fully exploit the auxiliary information from the encoding process. To achieve better performance, we propose a multi-stage attention network with auxiliary information refinement for Versatile Video Coding (VVC) in-loop filtering. The proposed network consists of two branches: the main filter branch extracts reconstruction features, while the auxiliary information refinement branch processes prediction and partition. Specifically, the auxiliary information refinement aims to extract features of auxiliary information better to assist in removing compression artifacts. Lastly, we introduce an Auxiliary Information Attention Module (AAM) to fuse the information flow between the two branches. The proposed model is integrated into the NNVC standard software VTM11.0_NNVC-4.0 and tested under all intra configuration. Experimental results show that our method achieves -7.60%, -19.49%, and -20.50% Bjøntegaard-Delta Bit-Rate (BD-BR) improvements across the Y, U, and V components. Xin-Yi Cui, Zhidao Zhou, Zhikai Liu, Fan Liang 0001 |
VCIP | 5 |
| 2023 | Sparse Representation based Deep Residual Geometry Compression Network for Large-scale Point CloudsabstractThe increasing applications of 3D point clouds require efficient compression techniques to achieve high-quality and low-delay services. However, the computational efficiency and rate-distortion performance for large-scale dense point clouds are still challenging, and the phenomenon of reconstruction ability degradation also exists when the network is deep. To solve these challenges, we propose a novel fully end-to-end point cloud compression model based on sparse convolution. Specifically, we adopt a long-range-residual aided architecture to avoid the reconstruction degradation and high computational complexity of deep networks. Further, we propose a multi-scale geometry compression module to construct an end-to-end network that avoids the accumulation of reconstruction distortion during decoding. Experiments on the large-scale Moving Picture Experts Group (MPEG) PCC benchmarks show that our model outperforms the latest Video-based Point Cloud Compression (V-PCC) scheme in terms of lossy geometry compression by 50.4% in D1 BD-rate and 50.8% in D2 BD-rate, while maintaining affordable processing speed and memory consumption. Pengpeng Yu, Dian Zuo, Yueer Huang, Ruishan Huang, Hanyun Wang, Yulan Guo, Fan Liang 0001 |
ICME | 7 |
| 2023 | Perceptual Based Fast CU Partition Algorithm for VVC Intra CodingabstractThe introduction of quad-tree with nested multi-type tree (QTMT) brings a significant reduction in bitrate to Versatile Video Coding (VVC) because QTMT enables a more flexible coding unit (CU) partition based on image content. However, the extensive rate distortion optimization (RDO) process induces a massive increase in encoding time. To mitigate this burden, a perceptual-based fast CU partition algorithm for VVC intra coding is proposed in this paper. The just noticeable difference model (JND) is adopted to simulate the human visual system, and JND variance is employed for early partition termination and division mode selection by reflecting the perceptual texture consistency. Experimental results on VTM17.0 show that the proposed method can achieve 31.13% time complexity reduction, with 1.32% Bjontegaard delta bit rate (BDBR) increase. Xin-Yi Cui, Fan Liang 0001 |
TENCON | 2 |
| 2023 | Diverse Context Model for Large-Scale Dynamic Point Cloud CompressionabstractSufficient context is essential for modeling the geometric distribution of large-scale dynamic point clouds. However, previous methods gather the context without considering the distinctive characteristics of different contexts, which leads to suboptimal performances. In this paper, we propose an octree-based diverse context model that captures the large-scale context, local detailed context, and temporal context adaptively and separately. To effectively aggregate the large-scale context, we exploit large-range sibling and ancestor nodes with a dilated mask window. For the local detailed context, we aggregate adjacent encoded sibling nodes with a subsequent mask window. To incorporate temporal context, we propose a density network to take full advantage of the cross-frame information of dynamic point clouds. Experiments on LiDAR and dense object datasets show that our method saves 38.17% and 47.47% of bitrates compared to the MPEG G-PCC method, respectively. Dian Zuo, Pengpeng Yu, Ruishan Huang, Yueer Huang, Wei Sun 0007, Fan Liang 0001 |
VCIP | 6 |
| 2022 | An Optimization Algorithm for Color Table Coding of Palette for VVC Based on DPCM and CCLPabstractAmong the existing video applications, screen content videos occupy a large proportion. Therefore, how to effectively compress screen content according to its characteristics has been a hot topic in the field of video coding. Since palette mode is a core tool for screen content coding in the Versatile Video Coding (VVC) Standard, it is of great importance to optimize its performance. In this paper, we propose an optimization algorithm for palette color table coding. The algorithm improves the compression performance when encoding the palette color table by introducing two methods, namely differential pulse code modulation and cross-component linear prediction. Experimental results show that the algorithm is able to optimize the coding performance with almost no increase in encoding and decoding time. Minghong Mo, Fan Liang 0001, Jun Wang 0015 |
ISCAS | 2 |
| 2022 | TransPCC: Towards Deep Point Cloud Compression via TransformersabstractHigh-efficient point cloud compression (PCC) techniques are necessary for various 3D practical applications, such as autonomous driving, holographic transmission, virtual reality, etc. The sparsity and disorder nature make it challenging to design frameworks for point cloud compression. In this paper, we present a new model, called TransPCC that adopts a fully Transformer auto-encoder architecture for deep Point Cloud Compression. By taking the input point cloud as a set in continuous space with learnable position embeddings, we employ the self-attention layers and necessary point-wise operations for point cloud compression. The self-attention based architecture enables our model to better learn point-wise dependency information for point cloud compression. Experimental results show that our method outperforms state-of-the-art methods on large-scale point cloud dataset. Zujie Liang, Fan Liang 0001 |
ICMR | 2 |
| 2022 | An IBC Reference Block Enhancement Model Based on GAN for Screen Content Video Coding
Pengjian Yang, Jun Wang 0015, Guangyu Zhong, Pengyuan Zhang, Lai Zhang, Fan Liang 0001, Jianxin Yang |
MMM (2) | 6 |
| 2021 | A Lossless Intra Reference Block Recompression Scheme for Bandwidth Reduction in HEVC-IBCabstractThe reference frame memory accesses in inter prediction result in high DRAM bandwidth requirement and power consumption. This problem is more intensive by the adoption of intra block copy (IBC), a new coding tool in the screen content coding (SCC) extension to High Efficiency Video Coding (HEVC). In this paper, we propose a lossless recompression scheme that compresses the reference blocks in intra prediction, i.e., intra block copy, before storing them into DRAM to alleviate this problem. The proposal performs pixel-wise texture analysis with an edge-based adaptive prediction method yet no signaling for direction in bitstreams, thus achieves a high gain for compression. Experimental results demonstrate that the proposed scheme shows a 72% data reduction rate on average, which solves the memory bandwidth problem. Jun Wang 0015, Guangyu Zhong, Jian Cao 0005, Ren Mao, Fan Liang 0001 |
ISCAS | 6 |
| 2021 | Encounter CU Again: History-Based Complexity Reduction Strategy for VVC Intra-Frame EncoderabstractDue to the newly adopted Quad Tree with Nested Multi-Type Tree (QTMT) partitioning scheme in Versatile Video Coding (VVC), multiple partitioning combinations can lead to the same Coding Unit (CU) structure. In other words, a CU may be encoded more than once. Based on this feature, a history-based complexity reduction strategy is proposed to accelerate VVC intra-frame coding with extremely low coding losses.Firstly, analyses of the relationship between the 1stround CUs (encoded at the first time) and the following rounds CUs (encountered again and already analyzed in previous partitioning attempts) are provided. Correspondingly, some unnecessary partitioning types are identified and early terminated. Secondly, a hierarchical pruning algorithm is designed, where thresholds are adjusted adaptively in the 1stround and used for pruning in the following rounds. To our knowledge, it is the first attempt to apply this history-based feature to accelerate partitioning for VVC intra-frame coding.Results show that these strategies can achieve 20% encoding time saving (TS) with only 0.18% BDBR increase. In addition, there is a huge potential for High-Resolution videos (21% TS with only 0.1% BDBR increase for 4K sequences). Compared to other works, our method achieves a considerably high TS/BDBR ratio, which indicates a better tradeoff between coding efficiency and complexity. Jian Cao 0005, Yifan Jia 0005, Fan Liang 0001, Jun Wang 0015 |
MMSP | 3 |
| 2021 | An Optimization Algorithm for Color Table Generation of Palette Mode for VVCabstractPalette mode is one of core screen content coding (SCC) tools in Versatile Video Coding (VVC) Standard. Nevertheless, the existing oversimplified and unitary color table generation (initiation) method of palette mode is hard to effectively deal with the increasingly complex screen content. Aiming at this problem, we propose a color table generation algorithm, an encoder only optimization algorithm that is effective to suppress color cluster centers drifting and enhance the convergence of color clustering by introducing another distinct reference color table. Experimental results show that with this method added, the encoder can achieve a certain BD-rate saving in test conditions compared to VTM 10.0 with a marginal coding time increase. Minghong Mo, Fan Liang 0001, Jun Wang 0015 |
MMSP | 2 |
| 2020 | Texture-Based Fast CU Size Decision and Intra Mode Decision Algorithm for VVC
Jian Cao 0005, Na Tang, Jun Wang 0015, Fan Liang 0001 |
MMM (1) | 4 |
| 2020 | IBC-Mirror Mode for Screen Content Coding for the Next Generation Video Coding StandardsabstractThis paper proposes an IBC-Mirror mode for Screen Content Coding (SCC) for the next generation video coding standards, including Versatile Video Coding (VVC) and Audio Video Standard-3 in China (AVS3). It is the first time to take mirror characteristic into consideration for SCC in VVC/AVS3. Based on the translational motion model of Intra Block Copy (IBC) mode, the function of "horizontal and vertical flipping" is further added to reduce prediction error and improve coding efficiency. The proposed IBC-Mirror mode is implemented on the latest reference software, including VTM5.0 (VVC) and HPM-5.0 (AVS3). The simulations show that the proposed mode can achieve up to 1~2% (VVC) and 4~7% (AVS3) BD-rate saving for SCC test sequences. Drafts about the mode have been submitted to AVS meeting and investigated in SCC Core Experiments (CE). Jian Cao 0005, Zhengren Li, Fan Liang 0001, Jun Wang 0015 |
VCIP | 4 |
| 2019 | An Intra-Affine Current Picture Referencing Mode for Screen Content Coding in VVCabstractWith the rapid development of emerging applications, screen content coding (SCC) is playing a more and more important role. Intra block copy (IBC), as a new tool for SCC, is proved to be efficient when there are many repeating or similar areas within the same picture. However, IBC is based on translational motion model, which may not work well for some blocks with complicated movement content, such as rotation and zoom. In this paper, a new intra-affine current picture referencing mode is proposed. In this new mode, non-translational motion model (affine model) is introduced to intra prediction for SCC to improve coding efficiency. First, candidate set of initial block-affine vectors (BVAffis) is established by a newly designed method. Then, those initial BVAffis are updated through iterative searching algorithm. Moreover, compatibility checking is applied. Compared to VTM3.0, the proposed new mode can achieve 2.53%, 2.47%, and 2.48% BD-rate saving on average for Y, U and V respectively for SCC test sequences. The draft about the new mode (JVET-O0682) was submitted in the 15th JVET meeting. Jian Cao 0005, Zhengren Li, Fan Liang 0001, Jun Wang 0015 |
PCS | 3 |
| 2017 | Feature based inter prediction optimization for non-translational video coding in cloudabstractVisual features of images and video frames have become pervasive and maturely developed in extensive research fields such as computer vision and visual search. In more and more cases, the visual feature becomes necessary information which needs to be transmitted and stored at server side in cloud. Among visual features, the local feature descriptors extracted by SIFT can represent both translational and non-translational motion, such as orientation and zooming. On the other hand, only translational motion can be represented by the Motion Vector (MV) in current MV based block video coding standard. Inspired by these properties, a method that utilizes the available feature to optimize inter prediction video coding is proposed in this paper. In this method, the localization, orientation and scale parameters of matching features extracted by SIFT are delivered to inter prediction to provide non-translational motion estimation (ME) and optimized merge mode. Experimental results have shown that the proposed method can efficiently improve the coding performance according to the accurate feature-matching. Xuelin Shen, Jun Wang 0015, Peilin Chen 0001, Fan Liang 0001 |
VCIP | 5 |