Shiqi Jiang 0006

dblp:07/10820-6 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0001-6120-3843ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LPCM: Learning-Based Predictive Coding for LiDAR Point Cloud Compression
abstract
In recent years, LiDAR point clouds have been widely used in many applications. Since the data volume of LiDAR point clouds is very huge, efficient compression is necessary to reduce their storage and transmission costs. However, existing learning-based compression methods do not exploit the inherent angular resolution of LiDAR and ignore the significant differences in the correlation of geometry information at different bitrates. The predictive geometry coding method in the geometry-based point cloud compression (G-PCC) standard uses the inherent angular resolution to predict the azimuth angles. However, it only models a simple linear relationship between the azimuth angles of neighboring points. Moreover, it does not optimize the quantization parameters for residuals on each coordinate axis in the spherical coordinate system. To address these issues, we propose a learning-based predictive coding method (LPCM) with both high-bitrate and low-bitrate coding modes. LPCM converts point clouds into predictive trees using the spherical coordinate system. In high-bitrate coding mode, we use a lightweight Long-Short-Term Memory-based predictive (LSTM-P) module that captures long-term geometry correlations between different coordinates to efficiently predict and compress the elevation angles. In low-bitrate coding mode, where geometry correlation degrades, we introduce a variational radius compression (VRC) module to directly compress the point radii. Then, we analyze why the quantization of spherical coordinates differs from that of Cartesian coordinates and propose a differential evolution (DE)-based quantization parameter selection method, which improves rate-distortion performance without increasing coding time. Experimental results show that LPCM achieved a D1-PSNR BD-rate reduction of 21.2% compared with the G-PCC lossless octree-based coding mode on SemanticKITTI, and 5.6% compared with the PredGeom on Ford, using the latest G-PCC test model TMC13 v31.0.
Hui Yuan 0001, Shiqi Jiang 0006, Da Ai, Wei Zhang 0072, Raouf Hamzaoui
IEEE Trans. Image Process.3
2026 Inter-LPCM: Learning-Based Inter-Frame Predictive Coding for LiDAR Point Cloud Compression
abstract
Because LiDAR sensors acquire point clouds with a fixed angular resolution, the resulting data can be systematically parameterized and efficiently compressed in the spherical coordinate system. Traditional spherical coordinate-based point cloud compression methods have shown strong rate-distortion (RD) performance, with the predictive geometry coding (PredGeom) method in the geometry-based point cloud compression (G-PCC) standard being a prominent example. While PredGeom includes an inter-frame prediction mode, it relies on a simple linear model, which limits its ability to capture complex motion patterns or structural dependencies. On the other hand, existing learning-based compression methods in the spherical domain do not exploit inter-frame correlations to reduce geometry redundancy. To address these limitations, we propose a learning-based inter-frame predictive coding method (Inter-LPCM). For azimuth prediction, we use a delta coding strategy based on the predefined angular resolution. To improve compression for radii, we introduce an inter-frame radius predictive (Inter-RP) model that estimates the current point's radius using neighboring points from both the current frame and the registered reference frame. In addition, we design a lightweight attention-based prediction (LAEP) model to predict elevation angles by capturing long-range geometric correlations across different coordinates. For quantization, we propose an RD-optimized method to select the quantization steps in the spherical coordinate system. For entropy coding, we design distinct models for each spherical coordinate component. These models are adapted to the statistical priors of each coordinate, which enables more accurate probability estimation. Experimental results show that Inter-LPCM, in its best RD configuration, achieved a D1-PSNR BD-rate reduction of 26.1% compared with the G-PCC lossless octree-based coding mode on SemanticKITTI, and 8.3% compared with the inter-frame prediction mode of PredGeom on Ford, using the latest G-PCC test model TMC13 v31.0. Our source code is publicly available at https://github.com/SDUChangSun/Inter-LPCM.
Hui Yuan 0001, Shiqi Jiang 0006, Chongzhen Tian, Raouf Hamzaoui
IEEE Trans. Image Process.3
2025 OMR-Net+: A Frequency-Aware Feature Refinement and Entropy Modeling Method for Efficient Screen Content Image Compression
abstract
Screen content image (SCI) compression faces challenges due to distinct characteristics such as sharp edges and repetitive structures. Existing learned image compression methods encounter two key issues: 1) insufficient frequency-aware processing, and 2) suboptimal entropy modeling for mixed-frequency components. To this end, we propose OMR-Net+, a novel SCI compression method that incorporates frequency-aware feature characteristics, including a frequency-aware refinement network (FARN) and a frequency-aware entropy model (FAEM). The proposed FARN uses an invertible neural network to preserve critical high-frequency details and a transformer-based model to reduce redundancy in low-frequency features. Additionally, the proposed FAEM provides tailored conditional probability estimation based on a parallel context model for high- and low-frequency features, respectively, to improve both coding performance and computational efficiency. Experimental results on the SCID and SIQAD datasets show that OMR-Net+ significantly outperforms the previous OMR-Net and other state-of-the-art methods in rate-distortion performance, demonstrating its potential for efficient SCI compression.
Shiqi Jiang 0006, Ting Ren, Hui Yuan 0001, Junyan Huo, Xin Lu 0001
IEEE Signal Process. Lett.1
2025 SPAC: Sampling-Based Progressive Attribute Compression for Dense Point Clouds
abstract
We propose an end-to-end attribute compression method for dense point clouds. The proposed method combines a frequency sampling module, an adaptive scale feature extraction module with geometry assistance, and a global hyperprior entropy model. The frequency sampling module uses a Hamming window and the Fast Fourier Transform to extract high-frequency components of the point cloud. The difference between the original point cloud and the sampled point cloud is divided into multiple sub-point clouds. These sub-point clouds are then partitioned using an octree, providing a structured input for feature extraction. The feature extraction module integrates adaptive convolutional layers and uses offset-attention to capture both local and global features. Then, a geometry-assisted attribute feature refinement module is used to refine the extracted attribute features. Finally, a global hyperprior model is introduced for entropy encoding. This model propagates hyperprior parameters from the deepest (base) layer to the other layers, further enhancing the encoding efficiency. At the decoder, a mirrored network is used to progressively restore features and reconstruct the color attribute through transposed convolutional layers. The proposed method encodes base layer information at a low bitrate and progressively adds enhancement layer information to improve reconstruction accuracy. Compared to the best anchor of the latest geometry-based point cloud compression (G-PCC) standard that was proposed by the Moving Picture Experts Group (MPEG), the proposed method can achieve an average Bjøntegaard delta bitrate of -24.58% for the Y component (resp. -21.23% for YUV components) on the MPEG Category Solid dataset and -22.48% for the Y component (resp. -17.19% for YUV components) on the MPEG Category Dense dataset. This is the first instance that a learning-based attribute codec outperforms the G-PCC standard on these datasets by following the common test conditions specified by MPEG. Our source code will be made publicly available on https://github.com/sduxlmao/SPAC.
Xiaolong Mao, Hui Yuan 0001, Shiqi Jiang 0006, Raouf Hamzaoui, Sam Kwong
IEEE Trans. Image Process.4
2025 Global Spatial-Temporal Information-Based Residual ConvLSTM for Video Space-Time Super-Resolution
abstract
By converting low-frame-rate, low-resolution videos into high-frame-rate, high-resolution ones, space-time video super-resolution techniques can enhance visual experiences and facilitate more efficient information dissemination. We propose a convolutional neural network (CNN) for space-time video super-resolution, namely GIRNet. Our method combines long-term global information and short-term local information from the video to better extract complete and accurate spatial-temporal information. To generate highly accurate features and thus improve performance, the proposed network integrates a feature-level temporal interpolation module with deformable convolutions and a global spatial-temporal information-based residual convolutional long short-term memory (convLSTM) module. In the feature-level temporal interpolation module, we leverage deformable convolution, which adapts to deformations and scale variations of objects across different scene locations. This provides a more efficient solution than conventional convolution for extracting features from moving objects. Our network effectively uses forward and backward feature information to determine inter-frame offsets, leading to the direct generation of interpolated frame features. In the global spatial-temporal information-based residual convLSTM module, the first convLSTM is used to derive global spatial-temporal information from the input features, and the second convLSTM uses the previously computed global spatial-temporal information feature as its initial cell state. This second convLSTM adopts residual connections to preserve spatial information, thereby enhancing the output features. Experiments on the Vimeo90 K dataset show that the proposed method outperforms open source state-of-the-art techniques in peak signal-to-noise-ratio (by 1.45 dB, 1.14 dB, and 0.2 dB over STARnet, TMNet, and 3DAttGAN, respectively), structural similarity index(by 0.027, 0.023, and 0.006 over STARnet, TMNet, and 3DAttGAN, respectively), and visual quality.
Congrui Fu, Hui Yuan 0001, Shiqi Jiang 0006, Liquan Shen, Raouf Hamzaoui
IEEE Trans. Multim.3
2024 A Transformer-Based Intra Luma Enhancement for H.266/VVC
abstract
Intra prediction is essential in reducing spatial domain correlation in video coding. To improve intra prediction accuracy, we introduce a transformer-based quality enhancement method aiming atimproving the luma quality of reconstructed coding tree units (CTUs). Our transformer-based model, namely Enhanceformer, utilizes multi-head attention for comprehensive feature extraction across multiple stages, levels, and scales. By integrating the model into H.266/VVC codec, it not only improves the luma quality of the current reconstructed CTU, but also provides more accurate references for intra prediction of subsequent CTUs. Experimental results show that this method achieves average BD rate savings of 2.63%, 0.21% and 0.48% for Y, Cb and Cr components respectively in all intra configuration, outperforming H.266/Versatile Video Coding (VVC) anchor.
Wenrui Lv, Hui Yuan 0001, Congrui Fu, Shiqi Jiang 0006, Junyan Huo
PCS4
2024 OMR-NET: A Two-Stage Octave Multi-Scale Residual Network for Screen Content Image Compression
abstract
Screen content (SC) differs from natural scene (NS) with unique characteristics such as noise-free, repetitive patterns, and high contrast. Aiming at addressing the inadequacies of current learned image compression (LIC) methods for SC, we propose an improved two-stage octave convolutional residual blocks (IToRB) for high and low-frequency feature extraction and a cascaded two-stage multi-scale residual blocks (CTMSRB) for improved multi-scale learning and nonlinearity in SC. Additionally, we employ a window-based attention module (WAM) to capture pixel correlations, especially for high contrast regions in the image. We also construct a diverse SC image compression dataset (SDU-SCICD2K) for training, including text, charts, graphics, animation, movie, game and mixture of SC images and NS images. Experimental results show our method, more suited for SC than NS data, outperforms existing LIC methods in rate-distortion performance on SC images.
Shiqi Jiang 0006, Ting Ren, Congrui Fu, Shuai Li 0005, Hui Yuan 0001
IEEE Signal Process. Lett.1
2023 Fourier Series and Laplacian Noise-Based Quantization Error Compensation for End-to-End Learning-Based Image Compression
abstract
Quantization is a core operation in lossy image compression. In the end-to-end learning-based image compression framework, quantization is conducted by a rounding operation during test, while it is replaced by additive uniform noise during training, leading to a mismatched problem between train and test. To address this problem, we propose a quantization error compensation method for the end-to-end learning-based image compression framework. The method uses Fourier series to approximate the periodic changes of the quantization error, and adds Laplacian noise to the quantized latent during test. The proposed method can be flexibly combined with different end-to-end learning-based image compression methods. Experimental results show that higher coding efficiency can be achieved by adding the proposed method with the state-of-the-art methods.
Shiqi Jiang 0006, Hui Yuan 0001, Shuai Li 0005, Xiaolong Mao
ICIP1
2022 An Attention-Based Network for Single Image HDR Reconstruction
abstract
High dynamic range (HDR) imaging can represent a great range of real-world luminosity. In contrast, the traditional low dynamic range (LDR) imaging fails to represent a wide range of luminance since most digital cameras can capture a limited range of light intensity in a natural scene. Recent advances in deep learning allow reconstructing an HDR image from a single LDR image and surpass conventional methods performance. In this work, we propose a novel CNN for HDR image reconstruction based on residual learning and attention mechanism. The proposed network adopts an autoencoder structure with residual blocks trained in a fully end-to-end manner. Residual learning boosts the performance by optimizing the network to converge faster. Moreover, the attention mechanism allows the network to select and enhance meaningful features that will contribute to the reconstruction of the HDR image. In addition, we employ a contextual attention module to perform patch replacement on deep feature maps to help recover information in over-exposed areas. Extensive quantitative and qualitative experiments on public HDR datasets demonstrate the ability of our proposed method to effectively reconstruct a visually pleasing HDR image from a single LDR image and outperform existing approaches.
Mohamed Dafaallah, Hui Yuan 0001, Shiqi Jiang 0006
ISCAS3