VLDB 2026 Research / reviewers in the wild / expert
Pengpeng Yu
dblp:246/3608
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AIQViT: Architecture-Informed Post-Training Quantization for Vision TransformersabstractPost-training quantization (PTQ) has emerged as a promising solution for reducing the storage and computational cost of vision transformers (ViTs). Recent advances primarily target at crafting quantizers to deal with peculiar activations characterized by ViTs. However, most existing methods underestimate the information loss incurred by weight quantization, resulting in significant performance deterioration, particularly in low-bit cases. Furthermore, a common practice in quantizing post-Softmax activations of ViTs is to employ logarithmic transformations, which unfortunately prioritize less informative values around zero. This approach introduces additional redundancies, ultimately leading to suboptimal quantization efficacy. To handle these, this paper proposes an innovative PTQ method tailored for ViTs, termed AIQViT (Architecture-Informed Post-training Quantization for ViTs). First, we design an architecture-informed low-rank compensation mechanism, wherein learnable low-rank weights are introduced to compensate for the degradation caused by weight quantization. Second, we design a dynamic focusing quantizer to accommodate the unbalanced distribution of post-Softmax activations, which dynamically selects the most valuable interval for higher quantization resolution. Extensive experiments on five vision tasks, including image classification, object detection, instance segmentation, point cloud classification, and point cloud part segmentation, demonstrate the superiority of AIQViT over state-of-the-art PTQ methods. Runqing Jiang, Ye Zhang 0037, Longguang Wang, Pengpeng Yu, Yulan Guo |
AAAI | 4 |
| 2025 | Lightweight Spectral Super-Resolution Network for Hyperspectral Image CompressionabstractThe growing use of hyperspectral images demands efficient compression techniques to handle their extensive spectral data. However, current methods are constrained by their inability to adapt to high bit depth and effectively utilize the spectral characteristics, leading to suboptimal compression ratios. This paper presents a novel hyperspectral compression framework that employs a lightweight spectral super-resolution network to address these limitations. The proposed approach divides the hyperspectral image into two sub-images, comprising two distinct groups of bands: a base image consisting of anchor bands and a supplementary image comprising non-anchor bands. The base image is compressed losslessly using a conventional codec, thereby ensuring the preservation of essential information. In contrast, the supplementary image is compressed efficiently by overfitting a lightweight super-resolution network to predict the non-anchor bands during encoding. The optimized network parameters are encoded as side information to ensure high-quality spectral super-resolution during decoding. Experimental results on the ARAD hyperspectral image dataset demonstrate that our approach significantly outperforms state-of-the-art methods, effectively meeting the demand for efficient hyperspectral image compression while maintaining acceptable processing speeds. Wei Zhang 0072, Pengpeng Yu, Yueru Chen, Dingquan Li, Wen Gao 0001 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Hierarchical Distortion Learning for Fast Lossy Compression of Point CloudsabstractThe growth of 3D point cloud applications requires efficient compression techniques for high-quality and low-latency services. Recently, learning-based point cloud compression models have made significant progress. However, geometric distortion resulting from downsampling limits the feature depth within large-scale point clouds, thereby constraining the receptive field and suppressing the redundant removal. Moreover, the issues of computational efficiency and reconstruction quality still persist in the compression of large-scale point clouds. To address these challenges, we propose a hierarchical distortion learning framework for end-to-end lossy compression of point clouds. First, we design a feature residual compression module to efficiently transmit shallow semantics between the encoder and the decoder, which enables a lightweight design of our framework. Second, we introduce a geometry residual compression module to progressively complement the reconstruction distortion, avoiding the accumulation of geometric distortion. By integrating these two modules and employing sufficient downsampling processes, we develop a high-performance framework with a significantly enlarged receptive field and low computational cost. Extensive experiments demonstrate that our method achieves state-ofthe- art performance in geometry lossy compression, while delivering competitive performance in joint geometry and color lossy compression with fast running speed. Code is available athttps://github.com/pengpeng-yu/FastPCC. Pengpeng Yu, Ye Zhang 0037, Fan Liang 0001, Haoran Li 0009, Yulan Guo |
IEEE Trans. Multim. | 1 |
| 2025 | Energy-guided test-time adaptation for data shifts in multi-modal perception
Yun Pei 0001, Lingbo Liu, Runqing Jiang, Ye Zhang 0037, Pengpeng Yu, Liang Lin 0004, Yulan Guo |
Vis. Comput. | 5 |
| 2024 | Efficient Point Cloud Attribute Compression Using Rich Parallelizable Context ModelabstractThe autoregressive context model has been proven effective in point cloud attribute compression. However, it suffers from unbearable decoding latency due to the limitations of serial decoding and the large scale of point clouds. In this paper, we propose a rich, parallelizable context model for point cloud attribute compression to speed up the decoding process. To further improve rate-distortion (RD) performance, we propose cross-coordinate and intra-coordinate attention modules to reduce the spatial redundancy of the latent representations. We validate our method on the large-scale Moving Picture Experts Group (MPEG) point cloud benchmarks, and demonstrate that our model achieves much lower decoding time than previous autoregression-based methods while maintaining similar RD performance. Ruishan Huang, Pengpeng Yu, Shaolin Liao, Fan Liang 0001 |
ICASSP | 2 |
| 2023 | Sparse Representation based Deep Residual Geometry Compression Network for Large-scale Point CloudsabstractThe increasing applications of 3D point clouds require efficient compression techniques to achieve high-quality and low-delay services. However, the computational efficiency and rate-distortion performance for large-scale dense point clouds are still challenging, and the phenomenon of reconstruction ability degradation also exists when the network is deep. To solve these challenges, we propose a novel fully end-to-end point cloud compression model based on sparse convolution. Specifically, we adopt a long-range-residual aided architecture to avoid the reconstruction degradation and high computational complexity of deep networks. Further, we propose a multi-scale geometry compression module to construct an end-to-end network that avoids the accumulation of reconstruction distortion during decoding. Experiments on the large-scale Moving Picture Experts Group (MPEG) PCC benchmarks show that our model outperforms the latest Video-based Point Cloud Compression (V-PCC) scheme in terms of lossy geometry compression by 50.4% in D1 BD-rate and 50.8% in D2 BD-rate, while maintaining affordable processing speed and memory consumption. Pengpeng Yu, Dian Zuo, Yueer Huang, Ruishan Huang, Hanyun Wang, Yulan Guo, Fan Liang 0001 |
ICME | 1 |
| 2023 | Diverse Context Model for Large-Scale Dynamic Point Cloud CompressionabstractSufficient context is essential for modeling the geometric distribution of large-scale dynamic point clouds. However, previous methods gather the context without considering the distinctive characteristics of different contexts, which leads to suboptimal performances. In this paper, we propose an octree-based diverse context model that captures the large-scale context, local detailed context, and temporal context adaptively and separately. To effectively aggregate the large-scale context, we exploit large-range sibling and ancestor nodes with a dilated mask window. For the local detailed context, we aggregate adjacent encoded sibling nodes with a subsequent mask window. To incorporate temporal context, we propose a density network to take full advantage of the cross-frame information of dynamic point clouds. Experiments on LiDAR and dense object datasets show that our method saves 38.17% and 47.47% of bitrates compared to the MPEG G-PCC method, respectively. Dian Zuo, Pengpeng Yu, Ruishan Huang, Yueer Huang, Wei Sun 0007, Fan Liang 0001 |
VCIP | 2 |