Shuai Li 0005

dblp:57/2281-5 · DBLP profile ↗
← Back
57ranked-venue papers
12as first author
41since 2021 · last 2026
0000-0002-9938-0917ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 10 first-author · 31 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hierarchical Frequency-Guided Alignment Transformer for Compressed Video Quality Enhancement
abstract
During the video encoding process, the original spatial domain signal is first transformed into the frequency domain, followed by quantization and compression. As a result, the quality degradation in compressed videos primarily stems from distortions in the frequency domain information. However, existing video enhancement methods typically directly fuse information from adjacent frames in the spatial domain, making it difficult for models to effectively compensate for frequency domain distortions, which leads to suboptimal detail restoration. To address this issue, we propose a Hierarchical Frequency-Guided Alignment Transformer. Additionally, by analyzing the characteristics of the frequency domain, we find that different frequency bands exhibit both correlations and a certain degree of independence. Based on this, we introduce a Frequency-Aware Transformer module that employs a combination of independent and mixed processing to optimize information exchange across different frequency domains, effectively mitigating cross-interference from irrelevant information. Experimental results demonstrate that, compared to existing methods, our approach achieves state-of-the-art performance in objective metrics (PSNR/SSIM), perceptual quality (LPIPS), and subjective visual effects, while reducing model complexity.
Liuhan Peng, Shuai Li 0005, Yanbo Gao, Mao Ye 0001, Chong Lv
AAAI2
2026 Haze has many faces: Multi-domain haze style transfer for diverse haze removal
Cunchuan Huang, Shuai Li 0005, Xiang Chen 0015, Jianlei Liu, Dengwang Li
Pattern Recognit.2
2026 Deformable Feature Alignment and Refinement for moving infrared small target detection
Dengyan Luo, Yanping Xiang, Luping Ji, Shuai Li 0005, Mao Ye 0001
Pattern Recognit.5
2026 A Noise Constrained Diffusion (NC-Diffusion) Framework for High-Fidelity Image Compression
abstract
With the great success of diffusion models in image generation, diffusion-based image compression is attracting increasing interests. However, due to the random noise introduced in the diffusion learning, they usually produce reconstructions with deviation from the original images, leading to suboptimal compression results. To address this problem, in this paper, we propose a Noise Constrained Diffusion (NC-Diffusion) framework for high fidelity image compression. Unlike existing diffusion-based compression methods that add random Gaussian noise and direct the noise into the image space, the proposed NC-Diffusion formulates the quantization noise originally added in the learned image compression as the noise in the forward process of diffusion. Then a noise constrained diffusion process is constructed from the ground-truth image to the initial compression result generated with quantization noise. The NC-Diffusion overcomes the problem of noise mismatch between compression and diffusion, significantly improving the inference efficiency. In addition, an adaptive frequency-domain filtering module is developed to enhance the skip connections in the U-Net based diffusion architecture, in order to enhance high-frequency details. Moreover, a zero-shot sample-guided enhancement method is designed to further improve the fidelity of the image. Experiments on multiple benchmark datasets demonstrate that our method can achieve the best performance compared with existing methods.
Yanbo Gao, Shuai Li 0005, Hui Yuan 0001, Mao Ye 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 CWRNN-INVR: A Coupled WarpRNN Based Implicit Neural Video Representation
abstract
Implicit Neural Video Representation (INVR) has emerged as a novel approach for video representation and compression, using learnable grids and neural networks. Existing methods focus on developing new grid structures efficient for latent representation and neural network architectures with large representation capability, lacking the study on their roles in video representation. In this paper, the difference between INVR based on neural network and INVR based on grid is first investigated from the perspective of video information composition to specify their own advantages, i.e., neural network for general structure while grid for specific detail. Accordingly, an INVR based on mixed neural network and residual grid framework is proposed, where the neural network is used to represent the regular and structured information and the residual grid is used to represent the remaining irregular information in a video. A Coupled WarpRNN-based multi-scale motion representation and compensation module is specifically designed to explicitly represent the regular and structured information, thus terming our method as CWRNN-INVR. For the irregular information, a mixed residual grid is learned where the irregular appearance and motion information are represented together. The mixed residual grid can be combined with the coupled WarpRNN in a way that allows for network reuse. Experiments show that our method achieves the best reconstruction results compared with the existing methods, with an average PSNR of 33.73 dB on the UVG dataset under the 3M model and outperforms existing INVR methods in other downstream tasks. The code can be found athttps://github.com/yiyang-sdu/CWRNN-INVR.git.
Yanbo Gao, Shuai Li 0005, Jinglin Zhang 0001, Hui Yuan 0001, Mao Ye 0001, Xingyu Gao 0001
IEEE Trans. Multim.3
2026 Long-Short Match for Lost Control in UAV Multi-Object Tracking
abstract
Multi-Object Tracking (MOT) in Unmanned Aerial Vehicles (UAV) aims to continuously and stably detect and track objects in videos captured by UAVs. In existing MOT tracking-by-detection schemes, the tracker with a fixed step size is always employed, and a fixed length of past tracking information is input to the tracker to guide position prediction. However, the limited prediction range of a single-scale tracker leads to frequent tracking losses, and limited historical information also reduces tracking accuracy. To address these limitations, we propose a novel Long-Short Match (LSMTrack) tracking method. The key idea is to use long and short trackers and maintain a long-term motion state to improve tracking performance, thus reducing the likelihood of entering the lost status. To this end, a new Mamba-based tracker and a long-short match strategy are proposed. For long and short trackers, the same architecture is used based on Mamba. Unlike the previous Mamba-based approach, the proposed tracker maintains a long-term state while updating the state and making position predictions in each time step, so we call it a step Mamba tracker. Meanwhile, we devise a long-short match strategy at the inference stage to integrate long and short trackers, and design a lost control operation which updates the long-term states using historical state values. In this way, the matching probability and the inference efficiency are guaranteed. Experimental results on two UAV MOT datasets confirm the state-of-the-art performance. Specifically, the best results are achieved in terms of two popular MOTA and IDF1 tracking evaluation metrics.
Zi-Zhuang Zou, Mao Ye 0001, Luping Ji, Lihua Zhou, Song Tang 0001, Yan Gan, Shuai Li 0005
IEEE Trans. Multim.7
2025 MetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural Representation
abstract
This paper presents MetricGrids, a novel grid-based neural representation that combines elementary metric grids in various metric spaces to approximate complex nonlinear signals. While grid-based representations are widely adopted for their efficiency and scalability, the existing feature grids with linear indexing for continuous-space points can only provide degenerate linear latent space representations, and such representations cannot be adequately compensated to represent complex nonlinear signals by the following compact decoder. To address this problem while keeping the simplicity of a regular grid structure, our approach builds upon the standard grid-based paradigm by constructing multiple elementary metric grids as high-order terms to approximate complex nonlinearities, following the Taylor expansion principle. Furthermore, we enhance model compactness with hash encoding based on different sparsities of the grids to prevent detrimental hash collisions, and a high-order extrapolation decoder to reduce explicit grid storage requirements. experimental results on both 2D and 3D reconstructions demonstrate the superior fitting and rendering accuracy of the proposed method across diverse signal types, validating its robustness and generalizability. Code is available at https://github.com/wangshu31/MetricGrids.
Yanbo Gao, Shuai Li 0005, Chong Lv, Chuankun Li, Hui Yuan 0001, Jinglin Zhang 0001
CVPR3
2025 A Fourier priors-Guided Diffusion Model for Image Harmonization with Structure-Preservation and Illumination-Consistency
abstract
Image harmonization is a crucial computer vision task that adjusts the appearance of foreground regions in composite images to match the background. Existing methods face challenges in effectively separating illumination from structure and preserving content consistency. In this paper, we propose a Fourier priors-Guided Diffusion Model for Image Harmonization with Structure Preservation and Illumination Consistency (FGDIH), which leverages frequency domain information and diffusion models. The key insight is that illumination is primarily concentrated in low-frequency amplitude, while structure is preserved in phase. Leveraging this insight, a Background-guided Amplitude Transfer Module (BATM) is proposed to transfer illumination from background to foreground. Meanwhile, a Phase Preservation Module (PPM) is developed to keep the original structure. Furthermore, FGDIH introduces a dual-constraint training strategy that incorporates both content and frequency domain supervision, ensuring stable harmonization during inference. Extensive experiments on iHarmony4 datasets demonstrate that our algorithm achieves SOTA performance and delivers competitive visual quality.
Tianyou Wang, Yanbo Gao, Shuai Li 0005
ICME5
2025 EEPNet: Efficient Edge Pixel-based Matching Network for Cross-Modal Dynamic Registration between LiDAR and Camera
abstract
Multisensor fusion is essential for autonomous vehicles to accurately perceive, analyze, and plan their trajectories within complex environments. This typically involves the integration of data from LiDAR sensors and cameras, which necessitates high-precision and real-time registration. Current methods for registering LiDAR point clouds with images face significant challenges due to inherent modality differences and computational overhead. To address these issues, we propose an efficient edge pixel-based matching network (EEPNet), an advanced network that leverages reflectance maps obtained from point cloud projections to enhance registration accuracy. The introduction of point cloud projections substantially mitigates cross-modality differences at the network input level, while the inclusion of reflectance data improves performance in scenarios with limited spatial information of point cloud within the camera’s field of view. Furthermore, by employing edge pixels for feature matching and incorporating an efficient matching optimization layer, EEPNet markedly accelerates real-time registration tasks. Experimental validation demonstrates that EEPNet achieves superior accuracy and efficiency compared to state-of-the-art methods. Our contributions offer significant advancements in autonomous perception systems, paving the way for robust and efficient sensor fusion in real-world applications.
Yuanchao Yue, Hui Yuan 0001, Shuai Li 0005
ISCAS3
2025 Adaptive Depth-Converted-Scale Convolution for Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation (MDE) has received increasing interests in the last few years. The objects in the scene, including the object size and relationship among different objects, are the main clues to extract the scene structure. However, previous works lack the explicit handling of the changing sizes of the object due to the change of its depth. Especially in a monocular video, the size of the same object is continuously changed, resulting in size and depth ambiguity. To address this problem, we propose a Depth-converted-Scale Convolution (DcSConv) enhanced monocular depth estimation framework, by incorporating the prior relationship between the object depth and object scale to extract features from appropriate scales of the convolution receptive field. The proposed DcSConv focuses on the adaptive scale of the convolution filter instead of the local deformation of its shape. It establishes that the scale of the convolution filter matters no less (or even more in the evaluated task) than its local deformation. Moreover, a Depth-converted-Scale aware Fusion (DcS-F) is developed to adaptively fuse the DcSConv features and the conventional convolution features. Our DcSConv enhanced monocular depth estimation framework can be applied on top of existing CNN based methods as a plug-and-play module to enhance the conventional convolution block. Extensive experiments with different baselines have been conducted on the KITTI benchmark and our method achieves the best results with an improvement up to 11.6% in terms of SqRel reduction. Ablation study also validates the effectiveness of each proposed module.
Yanbo Gao, Huibin Bai, Huasong Zhou, Xingyu Gao 0001, Shuai Li 0005, Hui Yuan 0001, Wei Hua 0002, Tian Xie 0011
IEEE Trans. Circuits Syst. Video Technol.5
2025 Unsupervised Feature Enrichment and Fidelity Preservation Learning Framework for Skeleton-Based Action Recognition
abstract
Unsupervised skeleton-based action recognition has achieved remarkable progress recently. Existing unsupervised learning methods suffer from severe overfitting problem, and thus small networks are used, significantly reducing the representation capability. To address this problem, the overfitting mechanism behind the unsupervised learning for skeleton-based action recognition is first investigated. It is observed that skeleton is already a relatively high-level and low-dimension feature, but not in the same manifold as the features for action recognition. Simply applying the existing unsupervised learning method tends to produce features that discriminate the different samples rather than action classes, resulting in the overfitting problem. To address this problem, this paper proposes an Unsupervised spatial-temporal Feature Enrichment and Fidelity Preservation (U-FEFP) learning framework to generate rich distributed features that contain all the information of a skeleton sample. A spatial-temporal feature transformation subnetwork is developed using channel-wise topology refinement graph convolutional block and graph convolutional gated recurrent unit block as the basic feature extraction network. The unsupervised Bootstrap Your Own Latent-based learning is utilized to generate rich distributed features, and the unsupervised pretext task-based learning is employed to preserve the information contained in the skeleton. The two unsupervised learning ways are collaborated as U-FEFP to produce robust and discriminative representations. Experimental results on four widely used benchmarks, namely NTU-RGB+D-60, PKU-MMD, NTU-RGB+D-120 and AAV-Human dataset, demonstrate that the proposed U-FEFP obtains the best result compared with the state-of-the-art unsupervised learning methods.
Chuankun Li, Shuai Li 0005, Yanbo Gao, Xingyu Gao 0001, Ping Chen 0004, Wanqing Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Approximately Invertible Neural Network for Learned Image Compression
abstract
Learned image compression has attracted considerable interests in recent years. An analysis transform and a synthesis transform, which can be regarded as coupled transforms, are used to encode an image to latent feature and decode the feature after quantization to reconstruct the image. Inspired by the success of invertible neural networks in generative modeling, invertible modules can be used to construct the coupled analysis and synthesis transforms. Considering the noise introduced in the feature quantization invalidates the invertible process, this paper proposes an Approximately Invertible Neural Network (A-INN) framework for learned image compression. It formulates the rate-distortion optimization in lossy image compression when using INN with quantization, which differentiates from using INN for generative modelling. Generally speaking, A-INN can be used as the theoretical foundation for any INN based lossy compression method. Based on this formulation, A-INN with a progressive denoising module (PDM) is developed to effectively reduce the quantization noise in the decoding. Moreover, a Cascaded Feature Recovery Module (CFRM) is designed to learn high-dimensional feature recovery from low-dimensional ones to further reduce the noise in feature channel compression. In addition, a Frequency-enhanced Decomposition and Synthesis Module (FDSM) is developed by explicitly enhancing the high-frequency components in an image to address the loss of high-frequency information inherent in neural network based image compression, thereby enhancing the reconstructed image quality. Extensive experiments demonstrate that the proposed A-INN framework achieves better or comparable compression efficiency than the conventional image compression approach and state-of-the-art learned image compression methods.
Yanbo Gao, Shuai Li 0005, Chong Lv, Hui Yuan 0001, Mao Ye 0001
IEEE Trans. Image Process.2
2025 LiftFormer: Lifting and Frame Theory Based Monocular Depth Estimation Using Depth and Edge Oriented Subspace Representation
abstract
Monocular depth estimation (MDE) has attracted increasing interest in the past few years, owing to its important role in 3D vision. MDE is the estimation of a depth map from a monocular image/video to represent the 3D structure of a scene, which is a highly ill-posed problem. To solve this problem, in this paper, we propose a LiftFormer based on lifting theory topology, for constructing an intermediate subspace that bridges the image color features and depth values, and a subspace that enhances the depth prediction around edges. MDE is formulated by transforming the depth value prediction problem into depth-oriented geometric representation (DGR) subspace feature representation, thus bridging the learning from color values to geometric depth values. A DGR subspace is constructed based on frame theory by using linearly dependent vectors in accordance with depth bins to provide a redundant and robust representation. The image spatial features are transformed into the DGR subspace, where these features correspond directly to the depth values. Moreover, considering that edges usually present sharp changes in a depth map and tend to be erroneously predicted, an edge-aware representation (ER) subspace is constructed, where depth features are transformed and further used to enhance the local features around edges. The experimental results demonstrate that our LiftFormer achieves state-of-the-art performance on widely used datasets, and an ablation study validates the effectiveness of both proposed lifting modules in our LiftFormer.
Shuai Li 0005, Huibin Bai, Yanbo Gao, Chong Lv, Hui Yuan 0001, Chuankun Li, Wei Hua 0002, Tian Xie 0011
IEEE Trans. Multim.1
2024 Frequency-Domain Transformation-Based Dynamic Gesture Recognition with Skeleton
Chuankun Li, Shuai Li 0005, Wanqing Li 0001, Danyan Xie
PRCV (3)3
2024 Static graph convolution with learned temporal and channel-wise graph topology generation for skeleton-based action recognition
Chuankun Li, Shuai Li 0005, Yanbo Gao, Lijuan Zhou 0002, Wanqing Li 0001
Comput. Vis. Image Underst.2
2024 Compressed-SDR to HDR Video Reconstruction
abstract
The new generation of organic light emitting diode display is designed to enable the high dynamic range (HDR), going beyond the standard dynamic range (SDR) supported by the traditional display devices. However, a large quantity of videos are still of SDR format. Further, most pre-existing videos are compressed at varying degrees for minimizing the storage and traffic flow demands. To enable movie-going experience on new generation devices, converting the compressed SDR videos to the HDR format (i.e., compressed-SDR to HDR conversion) is in great demands. The key challenge with this new problem is how to solve the intrinsic many-to-many mapping issue. However, without constraining the solution space or simply imitating the inverse camera imaging pipeline in stages, existing SDR-to-HDR methods can not formulate the HDR video generation process explicitly. Besides, they ignore the fact that videos are often compressed. To address these challenges, in this work we propose a novel imaging knowledge-inspired parallel networks (termed as KPNet) for compressed-SDR to HDR (CSDR-to-HDR) video reconstruction. KPNet has two key designs: Knowledge-Inspired Block (KIB) and Information Fusion Module (IFM). Concretely, mathematically formulated using some priors with compressed videos, our conversion from a CSDR-to-HDR video reconstruction is conceptually divided into four synergistic parts: reducing compression artifacts, recovering missing details, adjusting imaging parameters, and reducing image noise. We approximate this process by a compact KIB. To capture richer details, we learn HDR representations with a set of KIBs connected in parallel and fused with the IFM. Extensive evaluations show that our KPNet achieves superior performance over the state-of-the-art methods.
Mao Ye 0001, Xiatian Zhu, Shuai Li 0005, Xue Li 0001, Ce Zhu
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Color and Geometric Contrastive Learning Based Intra-Frame Supervision for Self-Supervised Monocular Depth Estimation
abstract
In recent years, self-supervised monocular depth estimation has become popular due to its advantage in estimating the depth without the need of groundtruth depth labels. Instead, it takes an inter-frame supervision using depth based view synthesis to reconstruct temporal adjacent frames to indirectly supervise the generated depth. However, such supervision weakens the depth estimation at temporal incoherent regions containing small changes among consecutive frames. To overcome the above problem, we propose a color and geometric contrastive learning based intra-frame supervision framework to enhance self-supervised monocular depth estimation. Color-contrastive learning is proposed to guide the network to learn color invariant features considering color information is irrelevant to depth data. To improve the local details of the learned feature, a pixel-level contrastive learning is further used to optimize the learning. In view that the depth estimation, as a pixel-level task, is sensitive to the geometric transformation, geometric-contrastive learning is developed using an inverse geometric transformation to learn features that are equivariant to the geometric data augmentation. A local plane guidance layer (LPG) with contrastive learning is further used to decompose the geometric information and enhance the geometric contrastive learning. Experiments demonstrate that the proposed method achieves the best result compared to the state-of-the-art methods in all tested quality metrics, with the largest improvement of 22.8% over baseline Monodepth2 and 3.2% over Monovit, in terms of SqRel reduction.
Yanbo Gao, Xianye Wu, Shuai Li 0005, Chuankun Li
IEEE Signal Process. Lett.3
2024 OMR-NET: A Two-Stage Octave Multi-Scale Residual Network for Screen Content Image Compression
abstract
Screen content (SC) differs from natural scene (NS) with unique characteristics such as noise-free, repetitive patterns, and high contrast. Aiming at addressing the inadequacies of current learned image compression (LIC) methods for SC, we propose an improved two-stage octave convolutional residual blocks (IToRB) for high and low-frequency feature extraction and a cascaded two-stage multi-scale residual blocks (CTMSRB) for improved multi-scale learning and nonlinearity in SC. Additionally, we employ a window-based attention module (WAM) to capture pixel correlations, especially for high contrast regions in the image. We also construct a diverse SC image compression dataset (SDU-SCICD2K) for training, including text, charts, graphics, animation, movie, game and mixture of SC images and NS images. Experimental results show our method, more suited for SC than NS data, outperforms existing LIC methods in rate-distortion performance on SC images.
Shiqi Jiang 0006, Ting Ren, Congrui Fu, Shuai Li 0005, Hui Yuan 0001
IEEE Signal Process. Lett.4
2024 Aligned Intra Prediction and Hyper Scale Decoder Under Multistage Context Model for JPEG AI
abstract
Learning-based image compression has raised increasing interests in the last few years. Currently, Joint Photographic Experts Group (JPEG) is working on the standardization of learning-based image compression as JPEG AI. It adopts a deep neural network based encoder-decoder architecture with hyperprior based probability formulation for entropy coding. JPEG AI currently contains two coding profiles, including the Base Operating Point (BaseOP) and High Operating Point (HighOP). Among the various techniques developed in JPEG AI, Multistage Context Model (MCM) was adopted as the context model to perform intra prediction in HighOP. It transforms the spatially progressive context prediction into sub-image feature prediction among channels via feature down-shuffling. However, in this prediction process, sub-image features are not spatially aligned to each other, and directly using the neighboring sub-image features cannot provide accurate prediction. Moreover, the distributions of residual features generated by MCM are also not consistent with that of the hyper scale decoder, which is used to construct the probability model in the entropy coding of residual features, leading to suboptimal residual coding. To address the above problems, we propose an Aligned Intra Prediction (AIP) and Aligned Hyper Scale Decoder (AHSD) under Multistage Context Model for JPEG AI coding. AIP aligns the reference sub-image features to the to-be-predicted feature in MCM with an offset prediction network and deformable convolution. AHSD further generates hyper scale features with matched distributions to the residual features, in order to enhance the probability formulation in its entropy coding. Experimental results demonstrate that the proposed method improves the coding performance by 1.3% in terms of BD-rate saving over the JPEG AI reference software and the effectiveness of each module is verified in ablation study.
Shuai Li 0005, Yanbo Gao, Chuankun Li, Hui Yuan 0001
IEEE Signal Process. Lett.1
2024 Geometric Warping Error Aware Spatial-Temporal Enhancement for DIBR Oriented View Synthesis
Rui Peng 0009, Shuai Li 0005, Yanbo Gao, Chuankun Li
IEEE Signal Process. Lett.3
2024 PU-Mask: 3D Point Cloud Upsampling via an Implicit Virtual Mask
abstract
We present PU-Mask, a virtual mask-based network for 3D point cloud upsampling. Unlike existing upsampling methods, which treat point cloud upsampling as an “unconstrained generative” problem, we propose to address it from the perspective of “local filling”, i.e., we assume that the sparse input point cloud (i.e., the unmasked point set) is obtained by locally masking the original dense point cloud with virtual masks. Therefore, given the unmasked point set and virtual masks, our goal is to fill the point set hidden by the virtual masks. Specifically, because the masks do not actually exist, we first locate and form each virtual mask by a virtual mask generation module. Then, we propose a mask-guided transformer-style asymmetric auto-encoder (MTAA) to restore the upsampled features. Moreover, we introduce a second-order unfolding attention mechanism to enhance the interaction between the feature channels of MTAA. Next, we generate a coarse upsampled point cloud using a pooling technique that is specific to the virtual masks. Finally, we design a learnable pseudo Laplacian operator to calibrate the coarse upsampled point cloud and generate a refined upsampled point cloud. Extensive experiments demonstrate that PU-Mask is superior to the state-of-the-art methods. Our code will be made available at: https://github.com/liuhaoyun/PU-Mask.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Qi Liu 0029, Shuai Li 0005
IEEE Trans. Circuits Syst. Video Technol.5
2024 A Structure-Preserving and Illumination-Consistent Cycle Framework for Image Harmonization
abstract
Ina composite image, the foreground and background are filmed under different scenarios, such as different lighting conditions, causing inconsistency and reducing the overall realism of the image. Image harmonization aims to generate visually realistic composite images by adjusting the foreground to the background conditions while maintaining the structure. Existing methods focus on adjusting the foreground object by directly training the foreground generation network with the ground truth, neglecting the different roles of the illumination and structure of the foreground in image harmonization. Moreover, the use of background, except for providing illumination, is not thoroughly investigated in this task. In this paper, we propose a structure-preserving and illumination-consistent cycle (SP-IC cycle) framework for image harmonization by exploring the illumination and structure of both the foreground and background. It achieves image harmonization by specifically changing the illumination and keeping the structure instead of ambiguously changing the foreground. Then, an illumination-consistent foreground harmonization cycle is developed to change the foreground illumination, while a structure-preserving cycle is designed to keep the foreground structure. Background information is explored in both cycles to assist in decomposing the illumination and structure of the foreground. In addition, the proposed SP-IC cycle framework can be applied to any image harmonization method to further boost its performance. Experimental results demonstrate that our method achieves better harmonious image quality than state-of-the-art methods, especially on an illumination-varying dataset.
Qingjie Shi, Yanbo Gao, Shuai Li 0005, Wei Hua 0002, Tian Xie 0011
IEEE Trans. Multim.4
2023 Fourier Series and Laplacian Noise-Based Quantization Error Compensation for End-to-End Learning-Based Image Compression
abstract
Quantization is a core operation in lossy image compression. In the end-to-end learning-based image compression framework, quantization is conducted by a rounding operation during test, while it is replaced by additive uniform noise during training, leading to a mismatched problem between train and test. To address this problem, we propose a quantization error compensation method for the end-to-end learning-based image compression framework. The method uses Fourier series to approximate the periodic changes of the quantization error, and adds Laplacian noise to the quantized latent during test. The proposed method can be flexibly combined with different end-to-end learning-based image compression methods. Experimental results show that higher coding efficiency can be achieved by adding the proposed method with the state-of-the-art methods.
Shiqi Jiang 0006, Hui Yuan 0001, Shuai Li 0005, Xiaolong Mao
ICIP3
2023 Single-image HDR reconstruction by dual learning the camera imaging process
Lei She, Mao Ye 0001, Shuai Li 0005, Ce Zhu
Eng. Appl. Artif. Intell.3
2023 Improved Shift Graph Convolutional Network for Action Recognition With Skeleton
abstract
Shift graph convolutional network (Shift-GCN) achieves remarkable performance for skeleton based action recognition with lower computational complexity than other GCN based methods. However, the current Shift-GCN, with one spatial shift, a static mask and a local temporal convolution, cannot fully explore the spatial-temporal features among skeleton joints of different frames. In order to address these problems, an improved shift graph convolutional network (Ishift-GCN) is proposed in this letter. The Ishift-GCN consists of two parts including a bidirectional spatial shift graph convolution with a dynamic mask, and a multi-scale temporal shift graph convolution. The bidirectional spatial shift graph convolution exploits more spatial information among joints, and the dynamic mask with stronger generalization ability can learn different correlations among features of different joints for different actions. The multi-scale temporal shift graph convolution captures more temporal information by complementing the shifted features with multi-scale convolution. Furthermore, knowledge distillation is used to reduce computational complexity. Compared with Shift-GCN, the proposed Ishift-GCN achieves better results with less computation complexity on two widely used benchmarks, namely the NTU-RGB+D and UAV-Human dataset.
Chuankun Li, Shuai Li 0005, Yanbo Gao, Wanqing Li 0001
IEEE Signal Process. Lett.2
2023 Multi-Frame Compressed Video Quality Enhancement by Spatio-Temporal Information Balance
abstract
In recent years, the performance of multi-frame quality enhancement algorithms for compressed videos has been greatly improved compared with single-frame based algorithms. However, the existing methods mainly focus on mining the temporal information of multiple frames. The large number of reference frames reduces the exploration of spatial information, although the existing single-frame based algorithms for enhancement, denoising, and super-resolution demonstrate the significance of the spatial information. To address this problem, we propose a plug-and-play module called Spatio-temporal Information Balance (STIB) to adaptively balance the spatial and temporal information. In our method, we use a feature extractor to exploit richer spatial information, and use a refinement module to refine the aligned temporal information, to be more conducive to the fusion of spatio-temporal information. Finally, we use the deformable convolution based re-alignment module to do alignment and fusion in feature space for balancing the spatio-temporal information. Experiments show that our module can significantly improve the performance of the existing multi-frame based enhancement algorithms.
Zeyang Wang, Mao Ye 0001, Shuai Li 0005, Xue Li 0001
IEEE Signal Process. Lett.3
2023 Spatio-Temporal Detail Information Retrieval for Compressed Video Quality Enhancement
abstract
The past few years have witnessed the great success of multi-frame quality enhancement for compressed video. Although the existing methods based on deformable alignment have achieved the state-of-the-art performance, they do not pay enough attention to the recovery of detail information. In this work, we propose a Spatio-Temporal Detail Retrieval (STDR) method to promote the recovery of detail information. To alleviate the problem of inaccurate deformable offsets caused by the fixed receptive field, motivated by multi-task learning, we design a plug-and-play Multi-path Deformable Alignment (MDA) module to generate more accurate offsets by integrating the alignment features of different receptive fields, so that the temporal detail information can be better recovered. For the spatial detail information restoration, several residual dense blocks with channel attention layer are utilized in the reconstruction module to explore valuable high-frequency spatial information from the fused multi-path alignment features. Meanwhile, a complementary loss function based on the Pearson correlation coefficient is developed to ameliorate the over-smoothing shortcoming caused by pixel-wise mean square or absolute value loss. Experimental results demonstrate that the proposed STDR network achieves superior performance compared with the state-of-the-art methods in both quantitative and qualitative evaluations.
Dengyan Luo, Mao Ye 0001, Shuai Li 0005, Ce Zhu, Xue Li 0001
IEEE Trans. Multim.3
2022 A Multiscale Gradient-Backpropagation Optimization Framework for Deformable Convolution Based Compressed Video Enhancement
abstract
Deep learning based compressed video quality enhancement has raised lots of interest recently. To explore the information over multiple frames, deformable convolution has been used for temporal alignment. However, in the existing methods, the deformable convolution is used in a relatively naïve way, without differing the characteristics of offset and features, and their behavior in gradient backpropagation. In this paper, a multiscale gradient-backpropagation optimization framework is proposed for the deformable convolution based compressed video quality enhancement. By analyzing the gradient backpropagation mechanism of deformable convolution, a multi-scale deformable convolution alignment structure is developed to facilitate the gradient backpropagation at all scales. Moreover, a progressive offset prediction module is developed, which decouples the offset prediction from the feature up-sampling, thus reducing the noise flow over scales. Experimental results show that the proposed method achieves the state-of-the-art performance, with 25.6% BD-rate saving compared to the HEVC reference software (HM).
Yanbo Gao, Menghu Jia, Shuai Li 0005, Mao Ye 0001, Frédéric Dufaux
ICASSP3
2022 PU-Refiner: A Geometry Refiner with Adversarial Learning for Point Cloud Upsampling
abstract
We present PU-Refiner, a generative adversarial network for point cloud upsampling. The generator of our network includes a coarse feature expansion module to create coarse upsampled features, a geometry generation module to regress a coarse point cloud from the coarse upsampled features, and a progressive geometry refinement module to restore the dense point cloud in a coarse-to-fine fashion based on the coarse upsampled point cloud. The discriminator of our network helps the generator produce point clouds closer to the target distribution. It makes full use of multi-level features to improve its classification performance. Extensive experimental results show that PU-Refiner is superior to five state-of-the-art point cloud upsampling methods. Code: https://github.com/liuhaoyun/PU-Refiner.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Wei Gao 0003, Shuai Li 0005
ICASSP5
2022 Content Adaptive Compressed Screen Content Video Quality Enhancement
abstract
In recent years, with the rise of various online learning plat-forms and game live broadcasting industry, screen content video is explosively increasing. There is an urgent demand to reduce the inevitable compression artifacts produced by traditional lossy compression method. However, there do not exist any research on compressed Screen Content Video (SCV) quality enhancement. Since SCV frames always con-sist of two main types of contents with different characteris-tics, i.e., text and graphic, a Content Adaptive model based on Two branches (CAT) is proposed in this paper. For en-hancing graphic, we utilize temporal information by motion compensation, while for enhancing text, we explore spatial relevance in horizontal and vertical directions to recover sharp edges. Furthermore, a content adaptive block is used to select content related features for the collaborative enhancement of both contents. We build a large SCV dataset compressed by H.266/VVC Test Model (VTM 12.1). On this dataset, experi-mental results demonstrate that the proposed method achieves the state-of-the-art performance on 10 different kinds of SCV test sequences.
Mao Ye 0001, Yanbo Gao, Shuai Li 0005, Xue Li 0001
ICME4
2022 KUNet: Imaging Knowledge-Inspired Single HDR Image Reconstruction
abstract
Recently, with the rise of high dynamic range (HDR) display devices, there is a great demand to transfer traditional low dynamic range (LDR) images into HDR versions. The key to success is how to solve the many-to-many mapping problem. However, the existing approaches either do not consider constraining solution space or just simply imitate the inverse camera imaging pipeline in stages, without directly formulating the HDR image generation process. In this work, we address this problem by integrating LDR-to-HDR imaging knowledge into an UNet architecture, dubbed as Knowledge-inspired UNet (KUNet). The conversion from LDR-to-HDR image is mathematically formulated, and can be conceptually divided into recovering missing details, adjusting imaging parameters and reducing imaging noise. Accordingly, we develop a basic knowledge-inspired block (KIB) including three subnetworks corresponding to the three procedures in this HDR imaging process. The KIB blocks are cascaded in the similar way to the UNet to construct HDR image with rich global information. In addition, we also propose a knowledge inspired jump-connect structure to fit a dynamic range gap between HDR and LDR images. Experimental results demonstrate that the proposed KUNet achieves superior performance compared with the state-of-the-art methods. The code, dataset and appendix materials are available at https://github.com/wanghu178/KUNet.git.
Mao Ye 0001, Xiatian Zhu, Shuai Li 0005, Ce Zhu, Xue Li 0001
IJCAI4
2022 Recurrent Deformable Fusion for Compressed Video Artifact Reduction
abstract
The compressed video inevitably appears in compression artifacts, which seriously affect the Quality of Experience. The state-of-the-art methods employ deformable alignment to gather similar information from multiple neighborhood frames to enhance target frame quality. However, they always align multiple frames to the target frame simultaneously, which brings repetitive and useless information because of multiple and imperfect alignments. In this paper, we propose a recurrent deformable fusion method which considers the alignment quality distortion caused by time distance from the target frame. Specifically, a Deformable Alignment (DA) module aligns each pair of the target frame and an adjacent frame following the time line. At the same time, a Recurrent Fusion (RF) module integrates the current aligned feature with the previous fused feature. After that, the fused features are concatenated along the time line. Then, a Multi-Scale Attention Reconstruction (MSAR) module is proposed to gather useful information from the fused features. Compared with the previous multi-frame alignment approach, our method can avoid obtaining a lot of repetitive and useless information. Experiment results confirm that our method achieves state-of-the-art performance on the standard test sequences.
Liuhan Peng, Askar Hamdulla, Mao Ye 0001, Shuai Li 0005, Hongwei Guo 0001
ISCAS4
2022 Structure-Preserving Motion Estimation for Learned Video Compression
abstract
Following the conventional hybrid video coding framework, existing learned video compression methods rely on the decoded previous frame as the reference for motion estimation considering that it is available to the decoder. Diving into its essential advantage of strong representation capability with CNNs, however, we find this strategy is suboptimal due to two reasons: (1) Motion estimation based on the decoded (often distorted) frame would damage both the spatial structure of motion information inferred and the corresponding residual for each frame, making it difficult to be spatially encoded on the whole image basis using CNNs; (2) Typically, it would break the consistent nature across frames since the estimated motion information is no longer consistent with the movement in the original video due to the distortion in the decoded video, lowering the overall temporal coding efficiency. To overcome these problems, a novel asymmetric Structure-Preserving Motion Estimation (SPME) method is proposed, with the aim to fully explore the ignored original previous frame at the encoder side while complying with the decoded previous frame at the decoder side. Concretely, SPME estimates superior spatially structure-preserving and temporally consistent motion field by aggregating the motion prediction of both the original and the decoded reference frames w.r.t the current frame. Critically, our method can be universally applied to the existing feature prediction based video compression methods. Extensive experiments on several standard test datasets show that our SPME can significantly enhance the state-of-the-art methods.
Han Gao 0012, Jinzhong Cui, Mao Ye 0001, Shuai Li 0005, Xiatian Zhu
ACM Multimedia4
2022 Geometric Warping Error Aware CNN for DIBR Oriented View Synthesis
abstract
Depth Image based Rendering (DIBR) oriented view synthesis is an important virtual view generation technique. It warps the reference view images to the target viewpoint based on their depth maps, without requiring many available viewpoints. However, in the 3D warping process, pixels are warped to fractional pixel locations and then rounded (or interpolated) to integer pixels, resulting in geometric warping error and reducing the image quality. This resembles, to some extent, the image super-resolution problem, but with unfixed fractional pixel locations. To address this problem, we propose a geometric warping error aware CNN (GWEA) framework to enhance the DIBR oriented view synthesis. First, a deformable convolution based geometric warping error aware alignment (GWEA-DCA) module is developed, by taking advantage of the geometric warping error preserved in the DIBR module. The offset learned in the deformable convolution can account for the geometric warping error to facilitate the mapping from the fractional pixels to integer pixels. Moreover, in view that the pixels in the warped images are of different qualities due to the different strengths of warping errors, an attention enhanced view blending (GWEA-AttVB) module is further developed to adaptively fuse the pixels from different warped images. Finally, a partial convolution based hole filling and refinement module fills the remaining holes and improves the quality of the overall image. Experiments show that our model can synthesize higher-quality images than the existing methods, and ablation study is also conducted, validating the effectiveness of each proposed module.
Shuai Li 0005, Yanbo Gao, Mao Ye 0001
ACM Multimedia1
2022 Quality enhancement of compressed screen content video by cross-frame information fusion
Jiawang Huang, Jinzhong Cui, Mao Ye 0001, Shuai Li 0005
Neurocomputing4
2022 An explicit self-attention-based multimodality CNN in-loop filter for versatile video coding
Menghu Jia, Yanbo Gao, Shuai Li 0005, Jian Yue, Mao Ye 0001
Multim. Tools Appl.3
2022 End-to-end video compression for surveillance and conference videos
Shenhao Wang, Han Gao 0012, Mao Ye 0001, Shuai Li 0005
Multim. Tools Appl.5
2022 Coarse-to-Fine Spatio-Temporal Information Fusion for Compressed Video Quality Enhancement
abstract
With the successful application of deformable convolution in aligning different video frames, it has also been used in video compression artifact reduction. The existing methods based on deformable convolution only apply 2D convolutional layers to generate the features for predicting alignment offsets, which is inaccurate due to limited receptive field. In this letter, we propose a new end-to-end network called Coarse-to-Fine Spatio-Temporal Information Fusion (CF-STIF) for compressed video quality enhancement by predicting better offsets with a larger receptive field. Specifically, several 3D convolutional layers are first to roughly fuse the spatio-temporal information in the video sequence, and then a Multi-level Residual Fusion Module (MLRF) is developed to generate global and local fused fine features from different levels for predicting deformable offsets. Thanks to the inherent advantages of 3D convolution and multi-scale strategy, the receptive field is greatly increased in both spatial and temporal dimensions, so that information from neighboring frames can be efficiently aggregated. In the end, the enhanced frame is derived by the proposed reconstruction module (REModule). Both qualitative and quantitative experimental results show that the proposed CF-STIF performs better than the state-of-the-art approaches.
Dengyan Luo, Mao Ye 0001, Shuai Li 0005, Xue Li 0001
IEEE Signal Process. Lett.3
2022 AGVS: A New Change Detection Dataset for Airport Ground Video Surveillance
abstract
Change detection is the foundation of intelligent video surveillance of the airport ground. However, experiments have shown that change detection algorithms with good performance on traditional datasets (e.g., CDnet2014) perform poorly in airport ground surveillance. The reason is that traditional datasets focus on the diversity of scenarios, while the practical application requires robustness against various changes in a single scene. We posit that the solution to this problem is to establish a unique dataset for airport ground surveillance and develop specific algorithms for this scenario. In this paper, we present an Airport Ground Video Surveillance benchmark (AGVS) for change detection of the airport ground. AGVS includes 25 long videos, amounting to about 100000 frames and accurate ground truth for all frames. Each video contains multiple challenges specific to the airport ground (e.g., haze, camouflage, strip shape, shadow and illumination change, simultaneous multi-scale objects) and various appearance changes of the aircraft). Change detection ground truth is generated by manual annotation. The AGVS benchmark can be downloaded fromhttps://www.agvs-caac.com. Furthermore, we conduct a simple review of current change detection algorithms, both unsupervised or supervised, and then 21 state-of-the-art algorithms are tested and analyzed on the AGVS benchmark. Finally, we conclude with algorithm design principles of change detection for airport ground surveillance.
Xiang Zhang 0006, Shuai Li 0005, Celimuge Wu, Zhi Liu 0002
IEEE Trans. Intell. Transp. Syst.3
2021 Deep Marginal Fisher Analysis based CNN for Image Representation and Classification
abstract
Deep Convolutional Neural Networks (CNNs) have achieved great success in image classification. While conventional CNNs optimized with iterative gradient descent algorithms with large data have been widely used and investigated, there is also research focusing on learning CNNs with non-iterative optimization methods such as the principle component analysis network (PCANet). It is very simple and efficient but achieves competitive performance for some image classification tasks especially on tasks with only a small amount of data available. This paper further extends this line of research and proposes a deep Marginal Fisher Analysis (MFA) based CNN, termed as DMNet. It addresses the limitation of PCANet like CNNs when the samples do not follow Gaussian distribution, by using a local MFA for CNN filter optimization. It uses a graph embedding framework for convolution filter optimization by maximizing the inter-class discriminability among marginal points while minimizing intra-class distance. Cascaded MFA convolution layers can be used to construct a deep network. Moreover, a binary stochastic hashing is developed by randomly selecting features with a probability based on the importance of feature maps for binary hashing. Experimental results demonstrate that the proposed method achieves state-of-the-art result in non-iterative optimized CNN methods, and ablation studies have been conducted to verify the effectiveness of the proposed modules in our DMNet.
Jiajing Chai, Yanbo Gao, Shuai Li 0005
ACM Multimedia4
2021 Learning various length dependence by dual recurrent neural networks
Chenpeng Zhang, Shuai Li 0005, Mao Ye 0001, Ce Zhu, Xue Li 0001
Neurocomputing2
2020 Bidirectional Independently Recurrent Neural Network for Skeleton-Based Hand Gesture Recognition
abstract
Gestures are a common form of human communication and important for Human-Computer Interaction (HCI). In this paper, we propose a new approach for skeleton-based hand gesture recognition based on the Independently Recurrent Neural Network (IndRNN). First, a bidirectional IndRNN (Bi-IndRNN) is developed to extend the IndRNN with the capability of bidirectional processing. Then, a deep Bi-IndRNN network is constructed for gesture recognition, where, in addition to the joint coordinates, the temporal displacement of each joint is also used to enhance the input features. Experimental results demonstrate that the proposed method achieves the state-of-the-art performance on the widely used DHG dataset with an accuracy of 93.15% for the 14 gesture classes case and 91.13% for the 28 gesture classes case.
Shuai Li 0005, Longfei Zheng, Ce Zhu, Yanbo Gao
ISCAS1
2020 A Mixed Appearance-based and Coding Distortion-based CNN Fusion Approach for In-loop Filtering in Video Coding
abstract
With the success of the convolutional neural networks (CNNs) in image denoising and other computer vision tasks, CNNs have been investigated for in-loop filtering in video coding. Many existing methods directly use CNNs as powerful tools for filtering without much analysis on its effect. Considering the in-loop filters process the reconstructed video frames produced from a fixed line of video coding operations, the coding distortion in the reconstructed frames may share similar properties that can be learned by CNNs in addition to being a noisy image. Therefore, in this paper, we first categorize the CNN based filtering into two types of processes: appearance-based CNN filtering and coding distortion-based CNN filtering, and develop a two-stream CNN fusion framework accordingly. In the appearance-based CNN filtering, a CNN processes the reconstructed frame as a distorted image and extracts the global appearance information to restore the original image. In order to extract the global information, a CNN with pooling is used first to increase the receptive field and up-sampling is added in the late stage to produce pixel-level frame information. On the contrary, in the coding distortion-based filtering, a CNN processes the reconstructed frame as blocks with certain types of distortions by focusing on the local information to learn the coding distortion resulted by the fixed video coding pipeline. Finally, the appearance-based filtering stream and the coding distortion-based filtering stream are fused together to combine the two aspects of CNN filtering, and also the global and local information. To further reduce the complexity, the similar initial and last convolutional layers are shared over two streams to generate a mixed CNN. Experiments demonstrate that the proposed method achieves better performance than the existing CNN-based filtering methods, with 11.26% BD-rate saving under the All Intra configuration.
Jian Yue, Yanbo Gao, Shuai Li 0005, Menghu Jia
VCIP3
2019 A fully trainable network with RNN-based pooling
Shuai Li 0005, Wanqing Li 0001, Chris Cook, Ce Zhu, Yanbo Gao
Neurocomputing1
2019 Source Distortion Temporal Propagation Analysis for Random-Access Hierarchical Video Coding Optimization
abstract
Due to the widely used inter prediction in the current video coding standards, encoding units in different frames is of temporal dependency in that the rate-distortion optimization (RDO) of one unit may affect the coding performance of the following units in the temporal domain. To achieve optimal coding solution for a given video sequence, temporal dependency among units needs to be considered in the RDO process, which is known as the temporally dependent RDO (TD-RDO). The hierarchical coding structure (HCS) employed in the High Efficiency Video Coding (HEVC) standard further complicates this problem by grouping frames into different layers of varying coding strategies, leading to a more complex temporal relationship. In our earlier work, we addressed TD-RDO for the low delay HCS (LD-HCS), where only uni-prediction is considered. This paper aims to address more complicated TD-RDO under random access HCS (RA-HCS), where both uni-prediction and bi-prediction are considered, making the temporal relationship even more intricate. The temporal dependency introduced in the RA-HCS is thoroughly examined and an RA-based TD-RDO scheme is formulated for each layer by modeling temporal propagation of distortion under different prediction types. Based on the formulation, the global Lagrange multiplier can be obtained analytically. Moreover, the effect of random access point pictures is considered in the RA-based TD-RDO scheme. The proposed method can be simply realized by updating the Lagrange multiplier as in the independent RDO formulation or combined with adjusting quantization parameter (QP) for better results in terms of BD-rate saving. Experimental results show that under RA-HCS, the proposed method, by adapting the Lagrange multiplier only, can achieve about 2.2% bitrate savings in average. With multi-QP optimization, an average BD-rate gain of 5.2% can be obtained.
Yanbo Gao, Ce Zhu, Shuai Li 0005, Tianwu Yang
IEEE Trans. Circuits Syst. Video Technol.3
2018 Independently Recurrent Neural Network (IndRNN): Building a Longer and Deeper RNN
abstract
Recurrent neural networks (RNNs) have been widely used for processing sequential data. However, RNNs are commonly difficult to train due to the well-known gradient vanishing and exploding problems and hard to learn long-term patterns. Long short-term memory (LSTM) and gated recurrent unit (GRU) were developed to address these problems, but the use of hyperbolic tangent and the sigmoid action functions results in gradient decay over layers. Consequently, construction of an efficiently trainable deep network is challenging. In addition, all the neurons in an RNN layer are entangled together and their behaviour is hard to interpret. To address these problems, a new type of RNN, referred to as independently recurrent neural network (IndRNN), is proposed in this paper, where neurons in the same layer are independent of each other and they are connected across layers. We have shown that an IndRNN can be easily regulated to prevent the gradient exploding and vanishing problems while allowing the network to learn long-term dependencies. Moreover, an IndRNN can work with non-saturated activation functions such as relu (rectified linear unit) and be still trained robustly. Multiple IndRNNs can be stacked to construct a network that is deeper than the existing RNNs. Experimental results have shown that the proposed IndRNN is able to process very long sequences (over 5000 time steps), can be used to construct very deep networks (21 layers used in the experiment) and still be trained robustly. Better performances have been achieved on various tasks by using IndRNNs compared with the traditional RNN and LSTM.
Shuai Li 0005, Wanqing Li 0001, Chris Cook, Ce Zhu, Yanbo Gao
CVPR1
2018 A Fusion Framework for Camouflaged Moving Foreground Detection in the Wavelet Domain
abstract
Detecting camouflaged moving foreground objects has been known to be difficult due to the similarity between the foreground objects and the background. Conventional methods cannot distinguish the foreground from background due to the small differences between them and thus suffer from underdetection of the camouflaged foreground objects. In this paper, we present a fusion framework to address this problem in the wavelet domain. We first show that the small differences in the image domain can be highlighted in certain wavelet bands. Then the likelihood of each wavelet coefficient being foreground is estimated by formulating foreground and background models for each wavelet band. The proposed framework effectively aggregates the likelihoods from different wavelet bands based on the characteristics of the wavelet transform. Experimental results demonstrated that the proposed method significantly outperformed existing methods in detecting camouflaged foreground objects. Specifically, the average F-measure for the proposed algorithm was 0.87, compared to 0.71 to 0.8 for the other stateof- the-art methods.
Shuai Li 0005, Dinei A. F. Florêncio, Wanqing Li 0001, Yaqin Zhao, Chris Cook
IEEE Trans. Image Process.1
2018 Hole Filling With Multiple Reference Views in DIBR View Synthesis
abstract
Depth-image-based rendering (DIBR) oriented view synthesis has been widely employed in the current depth-based 3-D video systems by synthesizing a virtual view from an arbitrary viewpoint. However, holes may appear in the synthesized view due to disocclusion, thus significantly degrading the quality. Consequently, efforts have been made on developing effective and efficient hole-filling algorithms. Current hole-filling techniques generally extrapolate/interpolate the hole regions with the neighboring information based on an assumption that the texture pattern in the holes is similar to that of the neighboring background information. However, in many scenarios, especially of complex texture, the assumption may not hold. In other words, hole-filling techniques can only provide an estimation for a hole which may not be good enough or may even be erroneous considering a wide variety of complex scene of images. In this paper, we first examine the view interpolation with multiple reference views, demonstrating that the problem of emerging holes in a target virtual view can be greatly alleviated by making good use of other neighboring complementary views in addition to its two (commonly used) most neighboring primary views. The effects of using multiple views for view extrapolation in reducing holes are also investigated in this paper. In view of the 3D Video and ongoing free-viewpoint TV standardization, we propose a new view synthesis framework, which employs multiple views to synthesize output virtual views. Furthermore, a scheme of selective warping of complementary views is developed by efficiently locating a small number of useful pixels in the complementary views for hole reduction, to avoid full warping of additional complementary views thus lowering greatly the warping complexity. Experimental results show that the hole size based on two primary reference views may be reduced by up to about 70% with the help of two complementary reference views in the case of view interpolation, while the hole size based on one primary reference view may be reduced by about 27% with the help of one more complementary reference view in view extrapolation. Moreover, it is shown that by using one more pair of views in view interpolation and one more view in view extrapolation, 10% hole pixels may be reduced additionally.
Shuai Li 0005, Ce Zhu, Ming-Ting Sun
IEEE Trans. Multim.1
2017 Foreground detection in camouflaged scenes
abstract
Foreground detection has been widely studied for decades due to its importance in many practical applications. Most of the existing methods assume foreground and background show visually distinct characteristics and thus the foreground can be detected once a good background model is obtained. However, there are many situations where this is not the case. Of particular interest in video surveillance is the camouflage case. For example, an active attacker camouflages by intentionally wearing clothes that are visually similar to the background. In such cases, even given a decent background model, it is not trivial to detect foreground objects. This paper proposes a texture guided weighted voting (TGWV) method which can efficiently detect foreground objects in camouflaged scenes. The proposed method employs the stationary wavelet transform to decompose the image into frequency bands. We show that the small and hardly noticeable differences between foreground and background in the image domain can be effectively captured in certain wavelet frequency bands. To make the final foreground decision, a weighted voting scheme is developed based on intensity and texture of all the wavelet bands with weights carefully designed. Experimental results demonstrate that the proposed method achieves superior performance compared to the current state-of-the-art results.
Shuai Li 0005, Dinei A. F. Florêncio, Yaqin Zhao, Chris Cook, Wanqing Li 0001
ICIP1
2017 Temporally Dependent Rate-Distortion Optimization for Low-Delay Hierarchical Video Coding
abstract
Low-delay hierarchical coding structure (LD-HCS), as one of the most important components in the latest High Efficiency Video Coding (HEVC) standard, greatly improves coding performance. It groups consecutive P/B frames into different layers and encodes them with different quantization parameters (QPs) and reference mechanisms in such a way that temporal dependency among frames can be exploited. However, due to varying characteristics of video contents, temporal dependency among coding units differs significantly from each other in the same or different layers, while a fixed LD-HCS scheme cannot take full advantage of the dependency, leading to a substantial loss in coding performance. This paper addresses the temporally dependent rate distortion optimization (RDO) problem by attempting to exploit varying temporal dependency of different units. First, the temporal relationship of different frames under the LD-HCS is examined, and hierarchical temporal propagation chains are constructed to represent the temporal dependency among coding units in different frames. Then, a hierarchical temporally dependent RDO scheme is developed specifically for the LD-HCS based on a source distortion propagation model. Experimental results show that our proposed scheme can achieve 2.5% and 2.3% BD-rate gain in average compared with the HEVC codec under the same configuration of P and B frames, respectively, with a negligible increase in encoding time. Furthermore, coupled with QP adaption, our proposed method can achieve higher coding gains, e.g., with multi-QP optimization, about 5.4% and 5.0% BD-rate saving in average over the HEVC codec under the same setting of P and B frames, respectively.
Yanbo Gao, Ce Zhu, Shuai Li 0005, Tianwu Yang
IEEE Trans. Image Process.3
2016 Hierarchical temporal dependent rate-distortion optimization for low-delay coding
abstract
Hierarchical coding structure (HCS) is one of the most important components in High Efficiency Video Coding (HEVC) that improves the coding performance greatly, especially for Low-Delay (LD) coding. It groups frames into different layers and enc odes them with different quantization parameters (QP) and different reference mechanisms. Due to the extensively used inter-prediction, the coding of frames in different layers is highly dependent and an appropriate QP and reference selection scheme may significantly improve the performance by taking advantage of such temporal dependency. However in the current HEVC codec, a predefined HCS, such as the Low-Delay HCS (LD-HCS), is performed without considering the different characteristic of different video contents, thus leading to a suboptimal coding solution. In this paper, the hierarchical temporal relationship under LD-HCS is first investigated and a hierarchical temporal propagation chain is constructed to describe the temporal dependency among frames. Then a hierarchical temporal dependent rate-distortion optimization scheme is developed specifically for the LD-HCS in HEVC. Experiments results show that the proposed scheme achieves BD-rate saving of 2.9% and 2.8% in average against HEVC codec under LD-HCS of P and B frames, respectively, with a negligible increase in encoding time.
Yanbo Gao, Ce Zhu, Shuai Li 0005
ISCAS3
2016 Layer-based temporal dependent rate-distortion optimization in Random-Access hierarchical video coding
abstract
Rate-distortion optimization (RDO) plays an important part in improving the coding efficiency of High Efficiency Video Coding (HEVC), especially for the hierarchical coding structure defined in the Random-Access (RA) configuration, noted as Random-Access Hierarchical Video Coding (RA-HVC), where different frames are assigned to different temporal layers and further coded with different coding parameters. Due to the inter-frame prediction, coding result of one unit may affect the coding performance of the following temporally related units. Therefore, the temporal dependency among units needs to be considered in the coding process. However, the RDO process in the current video codec is performed without considering the varying temporal dependency, thus compromising the rate-distortion performance significantly. To address this problem, a layer-based temporal dependent RDO method is proposed in this paper where the temporal dependency among different frames in the same or different layers is examined. By reformulating the temporal dependent RDO for the RA-HVC, we show that it can be implemented in a way of simply refining the Lagrange multiplier. Experimental results show that the proposed method achieves, in average, about 1.4% BD-rate savings with a negligible increase in encoding time for the random-access configuration.
Yanbo Gao, Ce Zhu, Shuai Li 0005, Tianwu Yang
MMSP3
2016 Lagrangian Multiplier Adaptation for Rate-Distortion Optimization With Inter-Frame Dependency
abstract
Rate-distortion optimization (RDO) is widely used in video coding, which plays a critical role in enhancing the coding efficiency substantially. Currently, the RDO process is performed in a way that coding efficiency of each coding unit (CU) is maximized independently without considering the dependency among CUs. As we know, in the current hybrid video coding structure, spatial/temporal prediction techniques are extensively used, which introduce strong dependency among CUs. In this paper, we investigate RDO with inter-frame dependency, where the impact of coding performance of the current CU on that of the following frames is considered. Accordingly, an RDO scheme taking the inter-frame dependency into account is proposed by adapting the Lagrangian multiplier. The experimental results show that the proposed scheme can achieve about 3.22% and 3.19% BD-rate saving in average over the state-of-the-art High Efficiency Video Coding (HEVC) reference software HM15.0 in the low-delay $P$ (LDP) and low-delay $B$ (LDB) coding structures, respectively, with no extra encoding time. The proposed scheme can obtain a significantly higher coding gain than the multiple quantization parameter (MQP) (±3) optimization technique that would greatly increase the encoding time by a factor of about six. Coupled with MQP optimization, the proposed scheme can further achieve about 5.96% and 5.57% BD-rate savings in average over the HEVC and about 4.03% and 4.07% over the HEVC with MQP optimization, under the specified common test conditions for LDP and LDB coding structures, respectively.
Shuai Li 0005, Ce Zhu, Yanbo Gao, Yimin Zhou 0002, Frédéric Dufaux, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.1
2015 Inter-frame dependent rate-distortion optimization using lagrangian multiplier adaption
abstract
It is known that, in the current hybrid video coding structure, spatial and temporal prediction techniques are extensively used which introduce strong dependency among coding units. Such dependency poses a great challenge to perform a global rate-distortion optimization (RDO) when encoding a video sequence. RDO is usually performed in a way that coding efficiency of each coding unit is optimized independently without considering dependeny among coding units, leading to a suboptimal coding result for the whole sequence. In this paper, we investigate the inter-frame dependent RDO, where the impact of coding performance of the current coding unit on that of the following frames is considered. Accordingly, an inter-frame dependent rate-distortion optimization scheme is proposed and implemented on the newest video coding standard High Efficiency Video Coding (HEVC) platform. Experimental results show that the proposed scheme can achieve about 3.19% BD-rate saving in average over the state-of-the-art HEVC codec (HM15.0) in the low-delay B coding structure, with no extra encoding time. It obtains a significantly higher coding gain than the multiple QP (±3) optimization technique which would greatly increase the encoding time by a factor of about 6. Coupled with the multiple QP optimization, the proposed scheme can further achieve a higher BD-rate saving of 5.57% and 4.07% in average than the HEVC codec and the multiple QP optimization enabled HEVC codec, respectively.
Shuai Li 0005, Ce Zhu, Yanbo Gao, Yimin Zhou 0001, Frédéric Dufaux, Ming-Ting Sun
ICME1
2015 Depth Coding Based on Depth-Texture Motion and Structure Similarities
abstract
This paper addresses high performance depth coding in 3D video by making good use of its coded texture video counterpart. The relationship between the depth and its associated texture video in terms of coding mode and motion vector is carefully examined. Our statistical study suggests that the skip-coding mode and its associated motion vectors in the coded texture can be shared for depth coding by saving bit rate at the cost of little increase of distortion, which subsequently results in a nonsequential coding of the depth map. In this sense, coding/prediction of a block can be performed using the skip-coded blocks below and right, which are not available in the conventional sequential coding, thus producing the so-called omnidirectional blocks predicted in the intra-coding by making the best use of (at most) four neighboring blocks. Moreover, in view of the depth-texture structure similarity, a depth-texture cooperative clustering-based prediction method is proposed for cluster-based depth prediction in the intra-coding, which exploits the structure similarity for the current coding block and its neighboring pixels around the block. On the other hand, some large prediction errors may be present for the depth-texture misaligned pixels, which may greatly compromise the coding performance. To deal with these large residuals induced by the depth-texture misalignment, a simple yet effective detection and rectification approach is incorporated in the proposed depth coding scheme. Experimental results show that our proposed depth coding scheme achieves superior rate-distortion performance compared with other relevant coding methods.
Jianjun Lei 0001, Shuai Li 0005, Ce Zhu, Ming-Ting Sun, Chunping Hou
IEEE Trans. Circuits Syst. Video Technol.2
2014 Rate control of hierarchical B prediction structure for multi-view video coding
Jianjun Lei 0001, Meimin Wu, Shuai Li 0005, Chunping Hou
Multim. Tools Appl.4
2014 Pixel-Based Inter Prediction in Coded Texture Assisted Depth Coding
abstract
This letter presents a pixel-based motion estimation scheme assisted with the coded texture video for depth inter-prediction, in view of motion similarity between depth and texture video. The proposed scheme can achieve higher inter-prediction gain without transmitting any motion vector in the pixel-based motion estimation. Coupled with depth-texture structure similarity, the inter prediction method is further extended to an integrated prediction approach by making use of both intra and inter information. Experimental results show that our proposed method achieves superior rate-distortion performance.
Shuai Li 0005, Jianjun Lei 0001, Ce Zhu, Lu Yu 0003, Chunping Hou
IEEE Signal Process. Lett.1