VLDB 2026 Research / reviewers in the wild / expert
Shuhua Xiong
dblp:34/220
· DBLP profile ↗
29ranked-venue papers
0as first author
25since 2021 · last 2026
0009-0000-8848-5441ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 14 since 2021Artificial intelligence and machine learning · 13 · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Subjective and objective evaluation of visual security in perceptually encrypted images
Xiaodong Bi, Xiaohai He, Zeming Zhao, Haitao Wei, Shuhua Xiong, Zheng Liu 0002, Ray E. Sheriff |
Expert Syst. Appl. | 5 |
| 2026 | Enhancing arbitrary-scale super-resolution with scale-aware multiscale nonlocal feature extraction and local structure-adaptive upsampling
Honggang Chen, Shuhua Xiong, Xiaohai He |
Neurocomputing | 4 |
| 2026 | Blind visual security assessment using a simple parallel dual-stream network
Xiaodong Bi, Xiaohai He, Zeming Zhao, Shuhua Xiong, Honggang Chen, Ray E. Sheriff |
Signal Process. Image Commun. | 4 |
| 2026 | Degradation-Aware Contrastive Learning for Blind Image Quality AssessmentabstractImages affected by the same distortion type and level usually have consistent statistical characteristics, while different distortions exhibit significant discriminability. Inspired by this observation, this paper proposes a Degradation-aware Contrast Learning (DCL) framework to explicitly model degradation properties for Blind Image Quality Assessment (BIQA). First, a Latent Degradation Space (LDS) is constructed via self-supervised contrastive learning to effectively capture degradation features from distorted images. Then, deep semantic features are extracted using a pre-trained model and fused with the degradation features to provide complementary bias information. Finally, the fused features are mapped to perceptual quality scores by a regression model. The experimental results show that the proposed method outperforms existing state-of-the-art BIQA methods in terms of prediction accuracy, robustness, and generalization ability. Xiaodong Bi, Xiaohai He, Shuhua Xiong, Zheng Liu 0002, Ray E. Sheriff |
IEEE Signal Process. Lett. | 3 |
| 2026 | Human-Machine Vision Collaboration Based Rate Control Scheme for VVCabstractWith the widespread adoption of smart terminals, compressed video is increasingly utilized in the receiver for purposes beyond human vision. Conventional video coding standards are optimized primarily for human visual perception and often fail to accommodate the distinct requirements of machine vision. To simultaneously satisfy the perceptual needs and the analytical demands, we propose a novel rate control scheme based on Versatile Video Coding (VVC) for human-machine vision collaborative video coding. Specifically, we employ the You Only Look Once (YOLO) network to extract task-relevant features for machine vision and formulate a detection feature weight based on these features. Leveraging the feature weight and the spatial location information of Coding Tree Units (CTUs), we propose a region classification algorithm that partitions a frame into machine vision-sensitive region (MVSR) and machine vision non-sensitive region (MVNR). Subsequently, we develop an enhanced and refined bit allocation strategy that performs region-level and CTU-level bit allocation, thereby improving the precision and effectiveness of the rate control. Experimental results demonstrate that the scheme improves machine task detection accuracy while preserving perceptual quality for human observers, effectively meeting the dual encoding requirements of human and machine vision. Zeming Zhao, Xiaohai He, Xiaodong Bi, Shuhua Xiong |
IEEE Signal Process. Lett. | 5 |
| 2026 | Spatial-Temporal Correlation Information-Based Rate Control for Versatile Video CodingabstractAlthough lambda-domain-based rate control is widely used in video encoders, developing an efficient rate control scheme for Coding Tree Units (CTUs) under the rate-distortion (R-D) principle remains a significant challenge. In this paper, we propose a spatial-temporal correlation information-based rate control scheme for Versatile Video Coding (VVC), aiming to improve coding performance. We introduce a weight estimation network to establish a CTU-level bit allocation strategy that fully exploits spatial-temporal contextual information. Moreover, the CTU-level coding parameter λ is adaptively optimized based on a dependency factor derived from distortion dependency information in both the spatial and temporal domains. Experimental results demonstrate that, compared to the default VVC rate control, the proposed scheme achieves BD-Rate savings of 6.48%, 17.33% and 13.75% in terms of the Peak Signal-to-Noise Ratio (PSNR), the Multi-Scale Structural Similarity Index (MS-SSIM) and the Video Multimethod Assessment Fusion (VMAF), respectively, under the Low Delay_P (LDP) configuration in the VVC Test Model (VTM) 19.0. Furthermore, the proposed method outperforms other state-of-the-art rate control schemes. Zeming Zhao, Xiaohai He, Shuhua Xiong, Meng Wang 0017, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Efficient Coding Parameters Optimization for Rate Control in Versatile Video CodingabstractIn mainstream video encoders, rate control is crucial in scenarios with limited bandwidth. Within the existing R-$\lambda$model-based rate control scheme, the coding parameters ($\alpha$and$\beta$) are directly involved in the calculation of$\lambda$, and their values have a significant impact on the calculation results. By selecting precise and appropriate coding parameters, enhanced rate control and rate-distortion performance can be realized. However, during the mapping of target bits to$\lambda$,$\alpha$and$\beta$often inadequately consider the rate-distortion attributes and the content characteristics of the coding units. This paper introduces a parameter optimization algorithm for rate control in Versatile Video Coding (VVC), aimed at enhancing coding efficiency. Utilizing the actual coded contexts of the coded Coding Tree Unit (CTU) alongside pre-coding information, we establish a parameters relationship model to deliver better coding parameters according to the rate-distortion attributes of the current CTU. Furthermore, leveraging coded contexts from spatially adjacent CTUs and the feature complexity, we propose a spatial coupling strategy to further improve the preceding coding parameters, considering the content characteristics of CTU. The proposed coding parameter optimization algorithm is implemented in the rate control of the VVC test model (VTM). Experimental findings indicate that this optimization algorithm accomplishes BD-rate savings concerning Peak Signal-to-Noise Ratio (PSNR) as well as the Multiscale Structural Similarity Index Metric (MS-SSIM) across various configurations. In addition, a more stable buffer status and enhanced visual quality are visible, which highlights the benefits of the proposed algorithm. Zeming Zhao, Xiaohai He, Xiaodong Bi, Qizhi Teng, Shuhua Xiong |
IEEE Trans. Multim. | 5 |
| 2025 | Enhanced Attention Context Model for Learned Image CompressionabstractRecently, deep learning has witnessed encouraging advances in image compression. An accurate entropy model, which estimates the probability distribution of the latent representation and reduces the bits required for compressing an image, is one of the keys to the success of learned image compression methods. The latent representation presents potential correlations in local, non-local, and cross-channel contexts. However, most entropy models only consider partial correlations, leading to suboptimal entropy estimation. In this letter, we propose a novel enhanced attention context model (EACM) to make full use of various correlations between latent elements for accurate entropy estimation. The proposed EACM contains a local spatial attention block (LSAB), a local channel attention block (LCAB), a global spatial attention block (GSAB), and a global channel attention block (GCAB). LSAB, LCAB, GSAB, and GCAB are carefully designed to adaptively exploit local spatial, local channel, global spatial, and global channel correlations, respectively. The experimental results on benchmark datasets show that our image compression model with the proposed EACM outperforms several state-of-the-art methods quantitatively and qualitatively. Zhengxin Chen, Xiaohai He, Chao Ren 0002, Tingrong Zhang, Shuhua Xiong |
IEEE Signal Process. Lett. | 5 |
| 2025 | A compressed video quality enhancement algorithm based on CNN and transformer hybrid network
Xiaohai He, Shuhua Xiong, Haibo He, Honggang Chen |
J. Supercomput. | 3 |
| 2024 | Unsupervised Blind Image Deblurring Based on Self-EnhancementabstractSignificant progress in image deblurring has been achieved by deep learning methods, especially the remarkable performance of supervised models on paired synthetic data. However, real-world quality degradation is more complex than synthetic datasets, and acquiring paired data in real-world scenarios poses significant challenges. To address these challenges, we propose a novel unsupervised image deblurring framework based on self-enhancement. The framework progressively generates improved pseudo-sharp and blurry image pairs without the need for real paired datasets, and the generated image pairs with higher qualities can be used to enhance the performance of the reconstructor. To ensure the generated blurry images are closer to the real blurry images, we propose a novel re-degradation principal component consistency loss, which enforces the principal components of the generated low-quality images to be similar to those of re-degraded images from the original sharp ones. Furthermore, we introduce the self-enhancement strategy that significantly improves deblurring performance without increasing the computational complexity of network during inference. Through extensive experiments on multiple real-world blurry datasets, we demonstrate the superiority of our approach over other state-of-the-art unsupervised methods. Lufei Chen, Xiangpeng Tian, Shuhua Xiong, Yinjie Lei, Chao Ren 0002 |
CVPR | 3 |
| 2024 | Blind video quality assessment based on Spatio-Temporal Feature Resolver
Xiaodong Bi, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen, Ray E. Sheriff |
Neurocomputing | 3 |
| 2024 | Multi deep invariant feature learning for cross-resolution person re-identification
Weicheng Zhang, Shuhua Xiong, Xiaohai He, Honggang Chen |
Inf. Process. Manag. | 2 |
| 2024 | A channel-wise contextual module for learned intra video compression
Yanrui Zhan, Shuhua Xiong, Xiaohai He, Honggang Chen |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Fast CU partition strategy based on texture and neighboring partition information for Versatile Video Coding Intra Coding
Ruolan Yang, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen |
Multim. Tools Appl. | 3 |
| 2024 | Dual-stage feedback network for lightweight color image compression artifact reduction
Zhengxin Chen, Xiaohai He, Tingrong Zhang, Shuhua Xiong, Chao Ren 0002 |
Neural Networks | 4 |
| 2023 | Nonlocal-guided enhanced interaction spatial-temporal network for compressed video super-resolution
Junxiong Cheng, Shuhua Xiong, Xiaohai He, Chao Ren 0002, Tingrong Zhang, Honggang Chen |
Appl. Intell. | 2 |
| 2023 | Block-correlation-based intra prediction for VVC
Shuhua Xiong, Xiaohai He, Honggang Chen, Chao Ren 0002 |
Multim. Tools Appl. | 2 |
| 2023 | Efficient Rate Control in Versatile Video Coding With Adaptive Spatial-Temporal Bit Allocation and Parameter UpdatingabstractDespite the fact that Versatile Video Coding (VVC) has achieved superior coding performance, two major problems remain for the rate control (RC) model in VVC. First, the regions concerned by human eyes are not clear enough in the coded video due to the deviation between the target bit allocation strategy of the coding tree unit (CTU) in RC and the human visual attention mechanism (HVAM). Second, there are significant quality fluctuations in the coded video frames due to the inappropriate updating speed. To address the above problems, we propose an efficient rate control (ERC) model. Specifically, in order to make the coded video more consistent with the attention of human eyes, we extract texture and motion-based spatial-temporal information to guide the bit allocation at the CTU level. Furthermore, based on the quasi-Newton algorithm and bit error, we propose an adaptive parameter updating (APU) method with the proper updating speed to precisely control the bits per frame. The proposed ERC outperforms the default RC model of VVC Test Model (VTM) 9.1 by saving the average Bjøntegaard Delta Rate (BD-Rate) on full-frame video sequences by 3.60% and 4.94% under low delay P (LDP) and random access (RA) configurations respectively, with higher bitrate accuracy. Moreover, the Peak Signal-to-Noise Ratio (PSNR) and actual coded bits per frame in the video coded by the proposed ERC are more stable. Liqiang He, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | A nonlocal HEVC in-loop filter using CNN-based compression noise estimation
Weiheng Sun, Xiaohai He, Honggang Chen, Shuhua Xiong |
Appl. Intell. | 4 |
| 2022 | Deep dual-domain semi-blind network for compressed image quality enhancement
Jingbo He, Xiaohai He, Mozhi Zhang, Shuhua Xiong, Honggang Chen |
Knowl. Based Syst. | 4 |
| 2022 | A quality enhancement network with coding priors for constant bit rate video coding
Weiheng Sun, Xiaohai He, Chao Ren 0002, Shuhua Xiong, Honggang Chen |
Knowl. Based Syst. | 4 |
| 2022 | Sequential Enhancement for Compressed Video Using Deep Convolutional Generative Adversarial Network
Xiaohai He, Honggang Chen, Shuhua Xiong |
Neural Process. Lett. | 5 |
| 2022 | An Optimized Rate Control Algorithm in Versatile Video Coding for 360$^\circ$ VideosabstractToday, 360$°$video has become an integral part of people's lives. Despite the fact that the latest generation standard Versatile Video Coding (VVC) demonstrates a significant gain in encoding capacity over High Efficiency Video Coding (HEVC), it still has room for 360$°$video encoding improvements. To further enhance the applicability of 360$°$video coding, an optimized rate control (RC) algorithm in VVC for 360$°$video is proposed in this paper. We present an efficient extraction algorithm for obtaining the video's saliency feature. Furthermore, for the characteristics of 360$°$video, a partitioning algorithm is also proposed to divide a frame into demand and non-demand regions. Additionally, to achieve precise and rational RC, a Coding Tree Unit (CTU)-level bit allocation strategy is proposed based on the saliency feature for the above-mentioned regions. The experimental results show that the proposed RC algorithm can achieve 11.77$\%$bitrate savings and more accurate allocation compared with the default algorithm of VVC. Also, performance enhancement has been observed in comparison to the most advanced algorithm. Zeming Zhao, Xiaohai He, Shuhua Xiong, Liqiang He, Ray E. Sheriff |
IEEE Signal Process. Lett. | 3 |
| 2021 | An improved R-λ rate control model based on joint spatial-temporal domain information and HVS characteristics
Zeming Zhao, Shuhua Xiong, Weiheng Sun, Xiaohai He, Feiran Zhang |
Multim. Tools Appl. | 2 |
| 2021 | Enhanced wide-activated residual network for efficient and accurate image deblocking
Zhengxin Chen, Xiaohai He, Chao Ren 0002, Pradeep Karn, Shuhua Xiong |
Signal Process. Image Commun. | 5 |
| 2020 | A quality enhancement framework with noise distribution characteristics for high efficiency video coding
Weiheng Sun, Xiaohai He, Honggang Chen, Ray E. Sheriff, Shuhua Xiong |
Neurocomputing | 5 |
| 2020 | Reduction of JPEG compression artifacts based on DCT coefficients prediction
Mengdi Sun, Xiaohai He, Shuhua Xiong, Chao Ren 0002, Xinglong Li |
Neurocomputing | 3 |
| 2019 | Machine learning-based H.264/AVC to HEVC transcoding via motion information reuse and coding mode similarity analysisabstractHigh‐efficiency video coding (HEVC), which is the latest video coding standard, is expected to have a dominant position in the market in the near future. However, most video resources are now encoded using the H.264/AVC standard. Consequently, there is a growing need for fast H.264/AVC to HEVC transcoders to facilitate the migration to the updated standard. This paper proposes a fast H.264/AVC to HEVC transcoding scheme, which constructs a three‐level classifier using an optimised tree‐augmented Naive Bayesian approach to predict the HEVC coding unit depth. A feature selection method is then proposed to improve prediction accuracy. A motion vector (MV) calculation method is also proposed to reduce the complexity of MV prediction in HEVC by reusing MVs from H.264/AVC. Experimental results show that, compared with other state‐of‐the‐art transcoding algorithms, the proposed algorithm considerably reduces coding complexity while causing only negligible rate‐distortion degradation. Xiaohai He, Linbo Qing, Shan Su, Shuhua Xiong |
IET Image Process. | 5 |
| 2017 | Tree-structured Bayesian compressive sensing via generalised inverse Gaussian distributionabstractCompressive sensing (CS) implements signal sampling and compression simultaneously, which significantly alleviates the pressure on the sampling end. However, the reconstruction algorithm is an underdetermined linear inverse problem. To solve this problem, it is crucial to involve prior knowledge regarding the reconstructed signal. In this study, the compressibility of wavelet coefficients is utilised as prior knowledge. Moreover, a generalised inverse Gaussian (GIG) distribution is integrated in the context of tree‐structured Bayesian CS (TSBCS), which also imposes the persistence property between the successive levels. Finally, variational Bayesian inference is used to infer the posterior probability distribution of the model parameters. Due to the overall algorithm is based on TSBCS, the proposal is referred to as TSBCS via a GIG distribution (TSBCS‐GIG). Experimental results show that the authors’ proposed TSBCS‐GIG algorithm outperforms other well‐known algorithms in both peak signal‐to‐noise ratio and visual quality. Maojiao Wang, Xiaohai He, Linbo Qing, Shuhua Xiong |
IET Signal Process. | 4 |