Liquan Shen

dblp:97/4612 · DBLP profile ↗
← Back
152ranked-venue papers
27as first author
63since 2021 · last 2026
0000-0002-2148-6279ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 131 · 24 first-author · 53 since 2021Artificial intelligence and machine learning · 14 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 4 since 2021Computer networks · 3 · 2 first-authorSystems, architecture and hardware · 2Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Breaking Redundancy via 3D Sparse Geometry: 3D-aware Neural Compression for Multi-View Videos
Shiwei Wang 0005, Liquan Shen, Jimin Xiao, Zhaoyi Tian, Feifeng Wang, Xiangyu Hu 0003, Yao Zhu 0006, Guorui Feng
Int. J. Comput. Vis.2
2026 Feature Compression for Cloud-Edge Multimodal 3D Object Detection
abstract
Machine vision systems, which can efficiently manage extensive visual perception tasks, are becoming increasingly popular in industrial production and daily life. Due to the challenge of simultaneously obtaining accurate depth and texture information with a single sensor, multimodal data captured by cameras and LiDAR is commonly used to enhance performance. Additionally, cloud-edge cooperation has emerged as a novel computing approach to improve user experience and ensure data security in machine vision systems. This paper proposes a pioneering solution to address the feature compression problem in multimodal 3D object detection. Given a sparse tensor-based object detection network at the edge device, we introduce two modes to accommodate different application requirements: Transmission-Friendly Feature Compression (T-FFC) and Accuracy-Friendly Feature Compression (A-FFC). In T-FFC mode, only the output of the last layer of the network's backbone is transmitted from the edge device. The received feature is processed at the cloud device through a channel expansion module and two spatial upsampling modules to generate multi-scale features. In A-FFC mode, we expand upon the T-FFC mode by transmitting two additional types of features. These added features enable the cloud device to generate more accurate multi-scale features. Experimental results on the KITTI dataset using the VirConv-L detection network showed that T-FFC was able to compress the features by a factor of 4933 with less than a 3% reduction in detection performance. On the other hand, A-FFC compressed the features by a factor of about 733 with almost no degradation in detection performance. We also designed optional residual extraction and 3D object reconstruction modules to facilitate the reconstruction of detected objects. The reconstructed objects effectively reflected the shape, occlusion, and details of the original objects.
Chongzhen Tian, Hui Yuan 0001, Raouf Hamzaoui, Liquan Shen, Sam Kwong
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Generic feature extraction and compression for human and machine-oriented vision
Kunqiang Huang, Ping An 0001, Chao Yang 0021, Shipei Wang, Xinpeng Huang, Liquan Shen
Signal Process.6
2026 SCVQENet: Quality enhancement for compressed screen content video
Zhaoyi Tian, Shiwei Wang 0005, Feifeng Wang, Liquan Shen
Signal Process. Image Commun.5
2026 DBA-PCGC: Dual-Domain Boundary Aware for Task-Friendly Point Cloud Geometry Compression
abstract
Compressed point clouds are increasingly used in machine vision tasks, which rely on key semantic regions of the point cloud such as geometric details and structural boundaries. However, existing point cloud compression methods for machine vision lack explicit awareness of geometrically induced semantic boundaries, causing semantic ambiguity in certain boundary regions during compression, thereby degrading machine vision performance. To address this issue, we propose a Dual-domain Boundary Aware Point Cloud Geometry Compression (DBA-PCGC) method that explicitly preserves semantic geometric boundaries from complementary spatial and frequency perspectives, enabling beneficial for machine vision tasks. Specifically, a Structure Aware Transform Module (SATM) exploits Gram matrix traces on local graphs to capture structural variations and highlight high-variation boundary regions, while compactly encoding smooth areas. In parallel, a Frequency Aware Transform Module (FATM) applies Chebyshev high-pass filtering to enhance high-frequency components corresponding to semantic geometric boundaries and suppress redundant low-frequency content. Experimental results on point cloud machine vision tasks demonstrate that our method achieves superior performance compared with existing compression approaches.
Minjian Chen, Liquan Shen, Qi Teng, Shiwei Wang 0005, Feifeng Wang
IEEE Signal Process. Lett.2
2026 UCSMC: An Underwater Compressed Sensing With Measurement Compression Framework
abstract
Thriving ocean applications demand efficient underwater image compression over bandwidth-limited acoustic channels. Recent works combine compressed sensing with measurement compression to improve compression ratios. However, as underwater attenuation weakens structural cues, sampling methods tend to overlook structural information and yield poor reconstructions. Meanwhile, sampling leaves discrete measurements with weak intra-image correlations, making it difficult for entropy models within measurement compression to predict accurate probability distributions. In this paper, we propose an Underwater Compressed Sensing with Measurement Compression (UCSMC) framework including Sketch-Assisted Sampling (SAS) and Spatial-Dictionary-based Mixture Entropy Coding (SDMEC) for low-bit-rate reconstruction. Specifically, in sampling, we incorporate sketch with underwater priors to drive the sampling process, steering more measurements toward critical structural regions and ultimately improving reconstruction quality. Additionally, we introduce a learnable spatial dictionary storing per-location entropy statistical characteristics in the underwater domain, which indicates local estimation difficulty and guides adaptive attention allocation in the entropy model, thereby improving probability estimation accuracy. Experimental results show our method outperforms previous schemes in reconstruction quality and measurement compression efficiency.
Liquan Shen, Shiwei Wang 0005, Minjian Chen
IEEE Signal Process. Lett.2
2026 ESHIC: Efficient Learning-Based Scalable HDR Image Compression With Hybrid Structural-Tonal Prior Modeling
Liquan Shen, Zhaoyi Tian, Xiangyu Hu 0003, Feifeng Wang, Shiwei Wang 0005
IEEE Trans. Circuits Syst. Video Technol.2
2026 CoMasTRe+: Unleashing Disentangled Continual Segmentation With Mixture of Continual Adapters
abstract
Continual Semantic Segmentation (CSS) suffers from catastrophic forgetting, particularly challenging for traditional per-pixel methods. Our prior work, CoMasTRe (CVPR 2024), introduced a query-based approach leveraging objectness by disentangling CSS into objectness learning and class recognition stages. While effective, CoMasTRe exhibited performance limitations due to feature forgetting within its pixel decoder. This paper presents CoMasTRe+, an enhanced framework specifically designed to overcome this limitation. The core contribution is a novel plugin, the Mixture of Continual Adapters (MoCA), integrated into the pixel decoder. MoCA is a dynamic architecture that mitigates feature forgetting by learning task-specific expert adapters. Crucially, MoCA employs a task-aware routing strategy and a novel adaptive routing distillation objective, tailored for continual learning, to preserve specialized feature representations across sequential tasks. CoMasTRe+ further enhances the class decoder using MoCA for improved recognition and simplicity. We extensively evaluate CoMasTRe+ on PASCAL VOC and ADE20K for continual semantic and panoptic segmentation. Experiments demonstrate that CoMasTRe+ effectively addresses the identified feature forgetting issue, significantly outperforms the original CoMasTRe, and achieves state-of-the-art results compared to both per-pixel and query-based baselines.
Yizheng Gong, Siyue Yu, Liquan Shen, Jimin Xiao
IEEE Trans. Circuits Syst. Video Technol.3
2026 DSCVC: Deep Screen Content Video Compression
abstract
Different from natural videos, screen content videos (SCVs) often exhibit homogeneous regions, abrupt content changes, and high prevalence of repetitive patterns. Existing deep learning (DL)-based video compression methods inadequately address the unique characteristics of SCVs, resulting in suboptimal compression performance. Therefore, in this paper, a dedicated deep screen content video compression (DSCVC) framework is proposed based on the motion and content characteristics of SCVs, which includes superpixel-constrained a motion estimation (SCME) module and inter and intra context aggregation (I2CA) module. The SCME is designed to construct a superpixel-based representation of homogeneous regions, leveraging the global correlations among superpixels to effectively capture large-scale motions, which efficiently improves the compression performance. I2CA is developed to jointly utilize inter and intra contexts, which employs a gating mechanism for content-aware context fusion, dynamically aggregating more similar contexts within SCVs. This allows for flexible adaptation to both contiguous and abrupt content changes within SCVs. Furthermore, by leveraging both learnable window and pixel displacements, a displacement-guided window attention mechanism is implemented in I2CA for precise long range repetitive feature localization, thereby reducing redundancy caused by repetitive patterns. To the best of our knowledge, it is the first DL-based video compression framework specifically designed for SCVs. Extensive experimental results demonstrate that the proposed DSCVC significantly outperforms existing methods in terms of compression performance, achieving a bitrate saving of 26.82% compared to VVC and a bitrate saving of 12.30% compared to SOTA DL-based methods.
Feifeng Wang, Liquan Shen, Zhaoyi Tian, Shiwei Wang 0005, Qi Teng, Yao Zhu 0006, Chengtao Zhou
IEEE Trans. Circuits Syst. Video Technol.2
2026 Ensemble Strategy for Underwater Image Quality Assessment and Dataset Construction
abstract
Learning-based Underwater Image Enhancement (UIE) methods have made significant progress. Limited by the manual label selection process, the limited quantity and outdated label quality of UIE datasets have severely hindered the development of UIE society. The urgent demand for more and better paired training samples motivates us to propose Ensemble-Select (E-Select), a strategy that can serve as an alternative to fully manual annotation and can enable continuous expansion of dataset size and optimization of label quality. However, expanding size will encounter new images, optimizing labels will encounter new algorithms and the proposed strategy is required to maintain strong generalization in both scenarios, which overwhelms many IQA methods. This work improves generalization in two ways. First, three primary influencing factors and their interrelationships in quality assessment are systematically analyzed. Specifically, we first explore the interactive relationship between content and distortion perception, and further investigate the guiding value of aesthetic-aware features in image quality perception. Then, the proposed Distortion-Content Interaction Module (DCIM) enables the network to focus on perceptually important distortion features guided by content. Second, we investigate a multi-perspective quality evaluation framework based on the ensemble learning paradigm. Building upon the availability of numerous outstanding IQA works, we initially demonstrate their distinct excel regions and evaluation biases. Subsequently, we explore the ensemble of their results through the proposed Aesthetic-Guided Quality Regression module (AGQR), which generates dynamic quality regression layers and derives image-specific quality perception rules based on aesthetic features. We then construct the first expandable, updatable UIE dataset with the help of E-Select. We collect over 50k real underwater image pairs with optimal labels, covering diverse scenes and varied degradation characteristics. Unlike other datasets, our dataset can consistently expand the number of paired samples and maintain optimal labeling without requiring extensive human labor. Experiments show the facilitating effect of the newly constructed dataset on UIE and the SOTA performance of E-Select. Codes and datasets are available at URL.
Yihan Yu, Liquan Shen, Zhengyong Wang
IEEE Trans. Circuits Syst. Video Technol.2
2026 Sketch-Based Extreme Underwater Image Compression Network
abstract
Underwater applications such as exploration and salvage operations require capturing underwater images (UWIs) to evaluate attributes such as the shape and structural integrity of submerged targets. However, underwater image transmission faces significant challenges due to the limited wireless acoustic channel available in underwater communication systems. Existing image compression algorithms struggle with limited compression ratios, which leads to a loss of crucial structural information and poor reconstruction quality, making them unsuitable for underwater practical applications. To overcome these limitations, we propose a sparse Sketch-based Extreme Underwater Compression framework (SEUCN), which mainly includes two sub-networks: Sparse Sketch Generation Network (SSGN) and Underwater Prior-guided Reconstruction Network (UPRN). To reduce redundancy and ensure effective compression at extremely low bitrates, the SSGN is designed to generate a compression-friendly sparse structural sketch through two ways. Firstly, it focuses on extracting important structural information to support analysis tasks within the constraints of limited bitrates. Secondly, it incorporates an underwater imaging model to focus on learning critical texture information for visual reconstruction. To restore the information lost during compression and achieve high-quality reconstruction, UPRN is designed to enhance structure details, restore underwater style, and enrich texture information during the reconstruction of UWIs from the decoded sketches, by effectively integrating multiple sources of prior knowledge. Specially, considering the high similarity of semantics and texture across different UWIs with common targets, the Dictionary-guided Texture Recovery Module (DTRM) leverages a universal underwater multi-scale feature dictionary as texture prior knowledge to supplement missing texture details. Extensive experiments show that our SEUCN demonstrates outstanding performance in retaining significant structural information to assist underwater practical tasks, and achieves superior visual quality compared to existing methods.
Liquan Shen, Shiwei Wang 0005, Feifeng Wang
IEEE Trans. Circuits Syst. Video Technol.2
2026 Task-Driven Underwater Image Enhancement via Hierarchical Semantic Refinement
abstract
Underwater image enhancement (UIE) is crucial for robust marine exploration, yet existing methods prioritize perceptual quality while overlooking irreversible semantic corruption that impairs downstream tasks. Unlike terrestrial images, underwater semantics exhibit layer-specific degradations: shallow features suffer from color shifts and edge erosion, while deep features face semantic ambiguity. These distortions entangle with semantic content across feature hierarchies, where direct enhancement amplifies interference in downstream tasks. Even if distortions are removed, the damaged semantic structures cannot be fully recovered, making it imperative to further enhance corrupted content. To address these challenges, we propose a task-driven UIE framework that redefines enhancement as machine-interpretable semantic recovery rather than mere distortion removal. First, we introduce a multi-scale underwater distortion-aware generator to perceive distortions across feature levels and provide a prior for distortion removal. Second, leveraging this prior and the absence of clean underwater references, we propose a stable self-supervised disentanglement strategy to explicitly separate distortions from corrupted content through CLIP-based semantic constraints and identity consistency. Finally, to compensate for the irreversible semantic loss, we design a task-aware hierarchical enhancement module that refines shallow details via spatial-frequency fusion and strengthens deep semantics through multi-scale context aggregation, aligning results with machine vision requirements. Extensive experiments on segmentation, detection, and saliency tasks demonstrate the superiority of our method in restoring machine-friendly semantics from degraded underwater images. Our code is available at https://github.com/gemyumeng/HSRUIE.
Liquan Shen, Yihan Yu, Rui Le
IEEE Trans. Image Process.2
2025 High Dynamic Range Video Compression: A Large-Scale Benchmark Dataset and A Learned Bit-depth Scalable Compression Algorithm
abstract
Recently, learned video compression (LVC) is undergoing a period of rapid development. However, due to absence of large and high-quality high dynamic range (HDR) video training data, LVC on HDR video is still unexplored. In this paper, we are the first to collect a large-scale HDR video benchmark dataset, named HDRVD2K, featuring huge quantity, diverse scenes and multiple motion types. HDRVD2K fills gaps of video training data and facilitate the development of LVC on HDR videos. Based on HDRVD2K, we further propose the first learned bit-depth scalable video compression (LBSVC) network for HDR videos by effectively exploiting bit-depth redundancy between videos of multiple dynamic ranges. To achieve this, we first propose a compression-friendly bit-depth enhancement module (BEM) to effectively predict original HDR videos based on compressed tone-mapped low dynamic range (LDR) videos and dynamic range prior, instead of reducing redundancy only through spatio-temporal predictions. Our method greatly improves the reconstruction quality and compression performance on HDR videos. Extensive experiments demonstrate the effectiveness of HDRVD2K on learned HDR video compression and great compression performance of our proposed LB-SVC network. Code and dataset will be released in https://github.com/sdkinda/HDR-Learned-Video-Coding.
Zhaoyi Tian, Feifeng Wang, Shiwei Wang 0005, Yao Zhu 0006, Liquan Shen
CVPR6
2025 Spherical rotation for high efficiency ERP 360-degree video coding
Tianyu Hong, Guowei Teng, Ping An 0001, Liquan Shen
Multim. Syst.4
2025 Bézier Surface-Guided Sampling Space Constraint for Neural Point Cloud Geometry Compression
abstract
Implicit Neural Representations (INRs), capable of learning the mapping between sampling coordinates and occupancy statuses to effectively represent 3D content, are increasingly adopted in Point Cloud Geometry Compression (PCGC). Existing INR-based PCGC methods suffer from substantial sampling space redundancy, leading to limited compression performance on point clouds. To address this issue, we propose Bezier ´Surface-Guided Sampling Space Constraint (BSGSSC), a novel method that significantly reduces sampling space redundancy via Bezier surface guidance. The key contributions of our work are: ´ 1) Leveraging the inherent compactness and distortion robustness of Bezier surfaces, a Low Bitrate Robust Surface Approximation ´ (LBRSA) is presented to outline the surface of dense point clouds with low transmission overhead; 2) An Oriented Bounding Box-based Sampling Space Constraint (OBBSSC) is designed to effectively constrain sampling space using multiple Oriented Bounding Boxes generated by Bezier surfaces. Experimental ´ results show that our method can enhance INR-based PCGC by 0.89 dB in terms of Bjøntegaard Delta Peak Signal-to-Noise Rate (BP).
Qi Teng, Liquan Shen, Feifeng Wang, Minjian Chen
IEEE Signal Process. Lett.2
2025 Prior-Guided Dual-Reference Contrastive Learning for Underwater Object Detection
abstract
Underwater object detection (UOD) plays an important role in the exploitation of marine ecological resources. Different from terrestrial images, the complex underwater environment leads to significant degradation in underwater images, which brings great difficulty in accurate object detection. In recent years, many specially designed UOD methods have been proposed to improve the detection precision in two aspects based on underwater image characteristics: 1) Some UOD methods utilize underwater image enhancement (UIE) to alleviate degradation with the expectation of clean features. However, neither preprocessing nor cascade approaches are fully effective for detection-oriented enhancement, while the additional UIE network increases inference time. 2) Other UOD methods consider low visibility of objects, blurriness of small objects, and occlusion problems. However, the semantic complementarity between objects of the same category but different qualities and the background patterns of specific objects are ignored. Based on these two observations, we propose a novel framework for the UOD task, which performs feature enhancement in two ways. First, a group contrastive-based feature enhancement module (GCFEM) is proposed to bridge UIE and UOD. Specifically, multiple enhanced versions by UIEs are evaluated by the object detection precision evaluation pipeline. Then, group-based contrastive learning is introduced, which utilizes multiple groups of enhanced versions to guide the backbone in extracting detection-friendly features. Second, a prior-guided dual-reference feature enhancement module (PDFEM) is proposed to enhance the representation of objects further. Specifically, the explicit object-object relationship allows low-quality object regions to refer to high-quality ones, guided by a transmission map. At the same time, the implicit object-background relationship provides cues about the surroundings for the representation of the objects. Experimental results demonstrate that the proposed algorithm outperforms many state-of-the-art UOD methods on RUOD and URPC2020 datasets.
Liquan Shen, Zhengyong Wang
IEEE Trans. Circuits Syst. Video Technol.2
2025 DRLN: Disparity-Aware Rescaling Learning Network for Multi-View Video Coding Optimization
abstract
Efficient compression of multi-view video data is a critical challenge for various applications due to the large volume of data involved. Although multi-view video coding (MVC) has introduced inter-view prediction techniques to reduce video redundancies, further reduction can be achieved by encoding a subset of views at a lower resolution through asymmetric rescaling, achieving higher compression efficiency. However, existing network-based rescaling approaches are designed solely for single-viewpoint videos. These methods neglect inter-view characteristics inherent in multi-view videos, resulting in suboptimal performance. To address this issue, we first propose a Disparity-aware Rescaling Learning Network (DRLN) that integrates disparity-aware feature extraction and multi-resolution adaptive rescaling to enhance MVC efficiency by minimizing both self- and inter-view redundancies. On the one hand, during the encoding stage, our method leverages the non-local correlation of multi-view contexts and performs adaptive downscaling with an early-exit mechanism, resulting in substantial multi-view bitrate savings. On the other hand, during the decoding stage, a dynamic aggregation strategy is proposed to facilitate effective interaction with inter-view features, utilizing the inter-view and cross-scale information to reconstruct fine-grained multi-view videos. Extensive experiments show that our network achieves a significant 26.31% BD-Rate reduction compared to the 3D-HEVC standard baseline, offering state of-the-art coding performance.
Shiwei Wang 0005, Liquan Shen, Peiying Wu, Zhaoyi Tian, Feifeng Wang
IEEE Trans. Circuits Syst. Video Technol.2
2025 High Efficiency Wiener Filter-Based Point Cloud Quality Enhancement for MPEG G-PCC
abstract
Point clouds, which directly record the geometry and attributes of scenes or objects by a large number of points, are widely used in various applications such as virtual reality and immersive communication. However, due to the huge data volume and unstructured geometry, efficient compression of point clouds is very crucial. The Moving Picture Expert Group is establishing a geometry-based point cloud compression (G-PCC) standard for both static and dynamic point clouds in recent years. Although lossy compression of G-PCC can achieve a very high compression ratio, the reconstruction quality is relatively low, especially at low bitrates. To mitigate this problem, we propose a high efficiency Wiener filter that can be integrated into the encoder and decoder pipeline of G-PCC to improve the reconstruction quality as well as the rate-distortion performance for dynamic point clouds. Specifically, we first propose a basic Wiener filter, and then improve it by introducing coefficients inheritance and variance-based point classification for the Luma component. Besides, to reduce the complexity of the nearest neighbor search during the application of the Wiener filter, we also propose a Morton code-based fast nearest neighbor search algorithm for efficient calculation of filter coefficients. Experimental results demonstrate that the proposed method can achieve average Bjøntegaard delta rates of -6.1%, -7.3%, and -8.0% for Luma, Chroma Cb, and Chroma Cr components, respectively, under the condition of lossless-geometry-lossy-attributes configuration compared to the latest G-PCC encoding platform (i.e., geometry-based solid content test model version 7.0 release candidate 2) by consuming affordable computational complexity.
Yuxuan Wei, Hao Liu 0044, Liquan Shen, Hui Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Perceptual Quality Assessment of High-Dynamic-Range Image: A Benchmark Dataset and a No Reference Method
abstract
High dynamic range (HDR) imaging technology has received increasing attention in recent years, and HDR image quality assessment (IQA) metrics are indispensable during the capturing, processing and displaying of HDR images. However, existing HDR-IQA datasets and methods neglect complex distortions during the HDR image processing schemes, leading to limited generalization performance on practical application. In this work, to facilitate the development of HDR-IQA dataset, we present HDRQAD, a large-scale HDR Quality Assessment Dataset, which possesses diversified distortions during HDR imaging technologies, abundant scenes and considerable quantity. Specifically, the HDRQAD dataset contains 1409 HDR images, which are derived from source scenes with six types of distortions during the HDR imaging schemes. In contrast to existing datasets that contain only compression artifacts, the HDRQAD includes Under-exposure, Over-exposure, Motion blur and Ghosting in HDR images achieved with multi-exposure fusion technology, conversion artifacts in HDR images achieved with single image reconstruction technology and compression artifacts during the transmission of HDR images. Furthermore, during the process of constructing the dataset, we identified three key challenges in HDR-IQA tasks: 1) dynamic range variations, 2) HDR visual artifacts with large overall gap, 3) inter-regional non-uniform image quality. Based on these observations, we propose a new end-to-end network for HDR-IQA tasks, which consists of a Distortion-aware Representation Learning (DRL) module and an Inter-Regional Quality Interaction (IRQI) module. The DRL learns the representations of dynamic range variations and HDR visual artifacts, enhancing the reliability of prior information extraction. The IRQI captures inter-regional quality dependencies with interacting and fusing intermediate distortion features for more accurately predicting image quality. Extensive experiments prove the superiority of proposed HDRQAD and demonstrate that the proposed network achieves state-of-the-art performance. The Dataset and Code will be made publicly available at HDR-IQA-Dataset.
Liquan Shen, Zhaoyi Tian, Xiangyu Hu 0003, Shiwei Wang 0005
IEEE Trans. Circuits Syst. Video Technol.2
2025 Underwater Image Quality Assessment Using Feature Disentanglement and Dynamic Content-Distortion Guidance
abstract
Due to the complex underwater imaging process, underwater images contain a variety of unique distortions. While existing underwater image quality assessment (UIQA) methods have made progress by highlighting these distortions, they overlook the fact that image content also affects how distortions are perceived, as different content exhibits varying sensitivities to different types of distortions. Both the characteristics of the content itself and the properties of the distortions determine the quality of underwater images. Additionally, the intertwined nature of content and distortion features in underwater images complicates the accurate extraction of both. In this paper, we address these issues by comprehensively accounting for both content and distortion information and explicitly disentangling underwater image features into content and distortion components. To achieve this, we introduce a dynamic content-distortion guiding and feature disentanglement network (DysenNet), composed of three main components: the feature disentanglement sub-network (FDN), the dynamic content guidance module (DCM), and the dynamic distortion guidance module (DDM). Specifically, the FDN disentangles underwater features into content and distortion elements, allowing us to more clearly measure their respective contributions to image quality. The DCM generates dynamic multi-scale convolutional kernels tailored to the unique content of each image, enabling content-adaptive feature extraction for quality perception. The DDM, on the other hand, addresses both global and local underwater distortions by identifying distortion cues from both channel and spatial perspectives, focusing on regions and channels with severe degradation. Extensive experiments on UIQA datasets demonstrate the state-of-the-art performance of the proposed method.
Liquan Shen, Zhengyong Wang, Yihan Yu
IEEE Trans. Circuits Syst. Video Technol.2
2025 Unlabeled Samples Improve Few-Shot Underwater Acoustic Target Recognition
abstract
The complex underwater environments pose major challenges to acoustic target recognition with insufficient labeled samples and mismatched domain issues. Although current few-shot learning (FSL) methods could alleviate the data scarcity problem, yet they still suffer from limited application scenarios and poor performance. Here, this article proposes a novel domain-adapted few-shot underwater acoustic target recognition (UATR) method based on environmental feature adaptation and frequency feature contrast (EFFC) that fully leverages unlabeled samples in the target domain to enhance model fine-tuning performance. First, an environmental feature adaptation module is purposely designed to integrate the weighted statistics of unlabeled samples from the current domain to capture the environmental information, promoting rapid adaptation to new environments. Next, a frequency feature contrast module is specifically devised by utilizing contrastive-learning to understand the characteristics of frequency bands from augmented, unlabeled samples, enabling the model to learn the general feature representations of ship-radiated noise. The performance is comprehensively evaluated on three public datasets, ShipsEar, DeepShip, and QiandaoEar22, as well as our dataset, DanShip, which is publicly available athttps://github.com/NPU-416/DanShip. Experimental results demonstrate that the model achieves the highest accuracy rates of 72.91%, 85.67%, 98.67%, and 82.56% on the ShipsEar, DeepShip, QiandaoEar22, and DanShip datasets, respectively, outperforming the other FSL methods.
Zhuofan He, Jing Han 0008, Qunfei Zhang, Yangtao Xue, Liquan Shen
IEEE Trans. Geosci. Remote. Sens.5
2025 Mask-Aware Light Field De-Occlusion With Gated Feature Aggregation and Texture-Semantic Attention
abstract
A light field image records rich information of a scene from multiple views, thereby providing complementary information for occlusion removal. However, current occlusion removal methods have several issues: 1) inefficient exploitation of spatial and angular complementary information among views; 2) indistinguishable treatment of pixels from foreground occlusion and background; and 3) insufficient exploration of spatial detail supplementation. Therefore, in this article, we propose a mask-aware de-occlusion network (MANet). Specifically, MANet is a joint training network that integrates the occlusion mask predictor (OMP) and the occlusion remover (OR). First, OMP is proposed to provide the location of occluded regions for OR, as the occlusion removal task is ill-posed without occluded region localization. In OR, we introduce gated spatial-angular feature aggregation, which uses a soft gating mechanism to focus on spatial-angular interaction features in non-occluded regions, extracting effective aggregated features specific to the de-occlusion. Then, we design a complementary strategy to fully utilize spatial-angular information among views. Finally, we propose texture-semantic attention to improve the performance of detail generation. Experimental results demonstrate the superiority of MANet, with substantial improvements in both PSNR and SSIM metrics. Moreover, MANet stands out with an efficient parameter count of 2.4 M, making it a promising solution for real-world applications in public safety and security surveillance.
Jieyu Chen, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Liquan Shen
IEEE Trans. Multim.6
2025 Global Spatial-Temporal Information-Based Residual ConvLSTM for Video Space-Time Super-Resolution
abstract
By converting low-frame-rate, low-resolution videos into high-frame-rate, high-resolution ones, space-time video super-resolution techniques can enhance visual experiences and facilitate more efficient information dissemination. We propose a convolutional neural network (CNN) for space-time video super-resolution, namely GIRNet. Our method combines long-term global information and short-term local information from the video to better extract complete and accurate spatial-temporal information. To generate highly accurate features and thus improve performance, the proposed network integrates a feature-level temporal interpolation module with deformable convolutions and a global spatial-temporal information-based residual convolutional long short-term memory (convLSTM) module. In the feature-level temporal interpolation module, we leverage deformable convolution, which adapts to deformations and scale variations of objects across different scene locations. This provides a more efficient solution than conventional convolution for extracting features from moving objects. Our network effectively uses forward and backward feature information to determine inter-frame offsets, leading to the direct generation of interpolated frame features. In the global spatial-temporal information-based residual convLSTM module, the first convLSTM is used to derive global spatial-temporal information from the input features, and the second convLSTM uses the previously computed global spatial-temporal information feature as its initial cell state. This second convLSTM adopts residual connections to preserve spatial information, thereby enhancing the output features. Experiments on the Vimeo90 K dataset show that the proposed method outperforms open source state-of-the-art techniques in peak signal-to-noise-ratio (by 1.45 dB, 1.14 dB, and 0.2 dB over STARnet, TMNet, and 3DAttGAN, respectively), structural similarity index(by 0.027, 0.023, and 0.006 over STARnet, TMNet, and 3DAttGAN, respectively), and visual quality.
Congrui Fu, Hui Yuan 0001, Shiqi Jiang 0006, Liquan Shen, Raouf Hamzaoui
IEEE Trans. Multim.5
2025 Dual-Guided Video Frame Interpolation With Spatial-Temporal Global Attention
abstract
Video frame interpolation technology improves visual experience with the development of deep learning. However, capturing large motions while synthesizing fine texture details remains a challenging task. Regarding large motion scenarios, some pioneering Transformer-based methods primarily rely on local attention, which does not fully leverage the global receptive field advantage. To address this issue, this paper proposes to further broaden the receptive field of the Transformer to capture more correlations in the video frame interpolation task. Specifically, we propose a global self-attention mechanism in the form of spatial-temporal separation. Regarding texture details, since roughly enlarging the receptive field results in the loss of details, we propose to use large motion information in both feature and pixel spaces as a dual-guided prior to enhance detail synthesis. The separable attention mechanism and the straightforward frame synthesis design significantly enhance the resource efficiency of our model. Extensive experiments show that our method achieves state-of-the-art performance, effectively capturing large motions and preserving texture details.
Baojun Zhou, Xinpeng Huang, Gongyang Li, Chao Yang 0021, Liquan Shen, Ping An 0001
IEEE Trans. Multim.5
2024 Towards Real-World HDR Video Reconstruction: A Large-Scale Benchmark Dataset and A Two-Stage Alignment Network
abstract
As an important and practical way to obtain high dynamic range (HDR) video, HDR video reconstruction from sequences with alternating exposures is still less explored, mainly due to the lack of large-scale real-world datasets. Existing methods are mostly trained on synthetic datasets, which perform poorly in real scenes. In this work, to facilitate the development of real-world HDR video reconstruction, we present Real-HDRV, a large-scale real-world benchmark dataset for HDR video reconstruction, featuring various scenes, diverse motion patterns, and high-quality labels. Specifically, our dataset contains 500 LDRs-HDRs video pairs, comprising about 28,000 LDR frames and 4,000 HDR labels, covering daytime, nighttime, indoor, and outdoor scenes. To our best knowledge, our dataset is the largest real-world HDR video reconstruction dataset. Correspondingly, we propose an end-to-end network for HDR video reconstruction, where a novel two-stage strategy is designed to perform alignment sequentially. Specifically, the first stage performs global alignment with the adaptively estimated global offsets, reducing the difficulty of subsequent alignment. The second stage implicitly performs local alignment in a coarse-to-fine manner at the feature level using the adaptive separable convolution. Extensive experiments demonstrate that: (1) models trained on our dataset can achieve better performance on real scenes than those trained on synthetic datasets; (2) our method outperforms previous state-of-the-art methods. Our dataset is available at https://github.com/yungsyu99/Real-HDRV.
Yong Shu, Liquan Shen, Xiangyu Hu 0003
CVPR2
2024 EPL-UFLSID: Efficient Pseudo Labels-Driven Underwater Forward-Looking Sonar Images Object Detection
Liquan Shen
ACM Multimedia2
2024 Content adaptive spatial-temporal rescaling for video coding optimization
Chao Yang 0021, Siqian Qin, Ping An 0001, Xinpeng Huang, Liquan Shen
Expert Syst. Appl.5
2024 Learning-based CU partition prediction for fast panoramic video intra coding
Chao Yang 0021, Ping An 0001, Xinpeng Huang, Liquan Shen
Expert Syst. Appl.5
2024 EUICN: An Efficient Underwater Image Compression Network
abstract
Thriving ocean applications bring explosive growth of underwater images (UWIs), which urgently demand to be compressed efficiently for transmission in the narrow underwater acoustic channel. However, existing image compression networks achieve suboptimal performance on UWIs. More efficient UWI compression can be achieved by utilizing characteristics of UWIs: (1) Within an UWI, the details distribution is associated with the underwater imaging transmission map (T-map); (2) Different UWIs have higher correlation than terrestrial images because they often present gauzy-covered indistinct appearance and share some universal ocean objects that widely appear in different underwater scenes. This paper fully exploits the two characteristics in terms of quantization and entropy coding, two key components of image compression network. Specifically, we propose an efficient underwater image compression network (EUICN) including underwater T-map-based quantization (UTMQ) and mixture entropy coding (MEC). In which, UTMQ extracts the imaging features from T-map, which are integrated with latent features by a novel dual-spatial attention module (DSAM) to generate a feature reserved mask, adaptively reserving reasonable numbers of features for different regions. For more efficient entropy coding, MEC is designed, which includes three correlation information extraction modules (i.e., hyperprior, local and novel universal information) and a probability prediction module. Especially, the universal information extraction module utilizes a comprehensive underwater feature dictionary, which covers various universal ocean objects, to match with latent features to select the universal correlation features. After that, the probability prediction module is designed to consolidate the hyperprior, local, and universal information to predict more accurate probability of latent features. Extensive experiments show that our EUICN achieves better performance than SOTA learned and conventional codecs in terms of PSNR and MS-SSIM.
Liquan Shen, Zhaoyi Tian
IEEE Trans. Circuits Syst. Video Technol.2
2024 RUIESR: Realistic Underwater Image Enhancement and Super Resolution
abstract
Clear and high-resolution (HR) underwater images are indispensable in acquiring underwater information. However, existing underwater image enhancement and super-resolution (UIESR) networks achieve limited enhancement-super-resolution performance on real-world turbid low-resolution (LR) underwater images because (1) they assume that the resolution degradation is simple and known bicubic down-sampling, generating unrealistic training data for UIESR task; (2) they extract known priors from the underwater imaging model, which is meager to address complex UIESR problems caused by unknown mixed dual-degradation; and (3) they ignore the interaction between blurring and color casts in the RGB color space, leading to unsatisfactory correction results of two distortions. To address these issues, we propose a realistic UIESR network (RUIESR) consisting of three parts: a realistic LR image generation module (RLGM), a dual-degradation estimation module (DEM), and an enhancement and super-resolution module (ESRM). Firstly, RLGM aims to generate LR images obeying underwater LR image distribution by learning real LR properties from unpaired real LR-HR underwater images for training. Secondly, a contrast-driven learning strategy is proposed in the DEM to accurately estimate unknown dual-degradation priors that can aid the reconstruction task. Finally, ESRM is proposed to enhance textures and correct color casts, which includes a dual-branch structure to separate blurring and color casts distortions and utilizes specific priors for each distortion to assist reconstruction. Extensive experiments on real and synthetic underwater datasets show that the proposed RUIESR outperforms existing works regarding visual quality and quantitative metrics.
Yinyi Li, Liquan Shen, Zhengyong Wang, Lihao Zhuang
IEEE Trans. Circuits Syst. Video Technol.2
2024 DSCIC: Deep Screen Content Image Compression
abstract
Existing deep learning-based image compression methods overlook the unique properties of screen content images (SCIs), like limited color values and abundant repetitive patterns, leading to limited compression performance on SCIs. Therefore, a specialized framework, deep screen content image compression (DSCIC) is proposed, which contains a color context generator (CCG) and a region-based block aggregation (RBA) module. The CCG is designed to generate compression-friendly color contexts based on main color components, embedded in the encoder-decoder to remove color representation redundancy. Furthermore, to effectively reduce repetitive block redundancy in SCIs, the RBA captures repetitive patterns and enables adaptive aggregation in the latent space. It leverages region-based block matching and block content-aware aggregation to utilize repetitive features for further improving compression performance. Extensive experimental results demonstrate that the proposed DSCIC outperforms the most advanced traditional codec VVC-SCC, and is significantly superior to other learning-based image compression methods. Using VVC as the anchor, DSCIC exhibits further BD-Rate savings of 12.185% and 4.889% compared to VVC-SCC and the SOTA deep learning-based method, respectively.
Feifeng Wang, Liquan Shen, Qi Teng, Zhaoyi Tian
IEEE Trans. Circuits Syst. Video Technol.2
2024 Task-Friendly Underwater Image Enhancement for Machine Vision Applications
abstract
Underwater images are often affected by color cast and blurring, which degrade the performance of underwater machine vision tasks. While existing underwater image enhancement (UIE) methods have been proposed to improve image quality for human perception, their effectiveness in enhancing machine vision performance is limited. In this article, a novel unsupervised UIE framework based on disentangled representation (DR) is proposed, which is designed for machine vision tasks. Specifically, the proposed framework disentangles the underwater image into two parts in the latent space according to whether they are beneficial to machine vision tasks: the task-friendly content features and the task-unfriendly distortion features. In addition, a semantic-aware contrastive module (SACM) is employed to alleviate the impact of losing key information required for machine vision tasks using the strategy of contrastive learning. Furthermore, two branches on the features and images are incorporated into the enhancement network, which serve the purpose of delivering task-relevant information to the enhancement model and guide the network to generate task-friendly images. Evaluation of the proposed method is conducted on multiple underwater image datasets, and a comparison is made with state-of-the-art enhancement methods in terms of machine vision performance. The experimental results demonstrate that the proposed method surpasses existing approaches in improving the accuracy and robustness of machine vision tasks, including object detection, semantic segmentation, and saliency detection in underwater environments. Our code is available athttps://github.com/gemyumeng/TFUIE.
Liquan Shen, Zhengyong Wang
IEEE Trans. Geosci. Remote. Sens.2
2024 Blind Quality Enhancement for Compressed Video
abstract
Deep convolutional neural networks (CNNs) have achieved impressive success in enhancing the quality of compressed images/videos. These approaches mostly obtain the noise level in advance and train multiple architecture-identical models for enhancement on images/videos of known levels of noise. It largely hinders their practical applications where the noise level is unknown and resource is limited. To practically perform quality enhancement, we propose a novel blind quality enhancement framework for compressed video (BQEV), which utilizes a single network to conduct enhancement on videos compressed at various and unknown quality parameters (QPs). Since there exists feature similarity and difference among videos compressed at multiple QPs, BQEV utilizes this prior to efficiently handle enhancement on videos compressed at blind QPs, which consists of progressive feature extraction and QP-adaptive feature fusion subnets. They utilize temporal information and feature similarity to progressively extract valuable features and further employ the feature difference to conduct reasonable QP-adaptive feature fusion and quality enhancement, respectively. In the progressive feature extraction subnet, we first design a quality rank module to assign more attention to higher-quality frames for efficient utilization of temporal information, then propose a progressive extraction module to further extract features from different QPs. In the QP-adaptive feature fusion subnet, we develop a quality estimation module to guide reasonable feature fusion of these extracted progressive features for stable and promising enhancement results on multiple QPs. Experimental results demonstrate that BQEV achieves 0.31–0.69 dB PSNR improvement compared with videos compressed at various QPs, outperforming state-of-the-art approaches.
Liquan Shen, Liangwei Yu, Hao Yang 0008, Mai Xu
IEEE Trans. Multim.2
2024 Spatial-Temporal Inter-Layer Reference Frame Generation Network for Spatial SHVC
abstract
In the current spatial Scalable High Efficiency Video Coding (SHVC) standard, the main techniques involve exploiting the correlation between pixel values of different layers to achieve inter-layer prediction samples, allowing the enhancement layer (EL) to predict samples from the upsampled base layer (BL) frame and remove temporal redundancy. However, existing network-based methods cannot effectively handle multi-layer compressed images with different resolutions to generate reference frame in spatial SHVC. Meanwhile, spatial SHVC only uses traditional interpolation filters to upsample the BL frame for EL frame sample prediction, which cannot handle different structures and contents. Therefore, considering the high correlation of multi-scale distortion characteristics across different layers, this article proposes a spatial-temporal inter-layer reference frame generation network (ST-ILR) for spatial SHVC, which can generate a high-fidelity reference frame for efficient inter-prediction and insert it into the EL reference picture list. The proposed method consists of two modules: a multi-scale motion restoration (MMR) module and a guided multi-scale feature reconstruction (GMFR) module. The MMR model is designed to accurately predict the motion trend of the EL based on the BL motion information, while implicitly compensating for previous EL frames. This is achieved by dynamically modeling the current EL motion information from the BL, capturing compression downsampling differences of prior motion vectors across different layers. The GMFR module adaptively super-resolves compressed BL frames and selectively aggregates high-frequency information from aligned EL features to preserve precise spatial detail, fusing abundant features from different layers to achieve better ILR frame quality performance. Extensive experiments show that our network achieves a 13.6% BD-rate (Bjøntegaard Delta Rate) reduction in random access configuration compared to the SHVC baseline, which offers state-of-the-art coding performance.
Shiwei Wang 0005, Liquan Shen, Jingyue Liu 0002
IEEE Trans. Multim.2
2024 UIERL: Internal-External Representation Learning Network for Underwater Image Enhancement
abstract
Underwater image enhancement (UIE) is a meaningful but challenging task, and many learning-based UIE methods have been proposed in recent years. Although much progress has been made, these methods still have two issues: (1) There exists a significant region-wise quality difference in a single underwater image due to the underwater imaging process, especially in regions with different scene depths. However, existing methods neglect this internal characteristic of underwater images, resulting in inferior performance; (2) Due to the uniqueness of the acquisition approach, underwater image acquisition tools usually capture multiple images in the same or similar scenes. Thus, the underwater images to be enhanced in practical usage are highly correlated. However, when processing a single image, existing methods do not consider the rich external information provided by the related images. There is still room for improvement in their performance. Motivated by these two aspects, we propose a novel internal-external representation learning (UIERL) network to better perform UIE tasks with internal and external information, simultaneously. In the internal representation learning stage, a new depth-based region feature guidance network is designed, including a region segmentation module based on scene depth to sense regions with different quality levels, followed by a region-wise space encoder module. With performing region-wise feature learning for regions with different quality separately, the network provides an effective guidance for global features and thus guides intra-image differentiated enhancement. In the external representation learning stage, we first propose an external information extraction network to mine the rich external information in the related images. Then, internal and external features interact with each other via the proposed external-assist-internal module (external features are updated with the help of internal features) and internal-assist-external module (internal features are updated with the help of external features). In this way, our UIERL fully explores the rich internal and external information to better enhance a single image. All results show that our method can achieve state-of-the-art performance on five benchmarks.
Zhengyong Wang, Liquan Shen, Yihan Yu, Hui Yuan 0001
IEEE Trans. Multim.2
2024 Multi-Prior Driven Resolution Rescaling Blocks for Intra Frame Coding
abstract
Deep learning techniques are increasingly integrated into rescaling-based video compression frameworks and have shown great potential in improving compression efficiency. However, existing methods achieve limited performance because 1) they treat context priors generated by codec as independent sources of information, ignoring potential interactions between multiple priors in rescaling, which may not effectively facilitate compression; 2) they often employ a uniform sampling ratio across regions with varying content complexities, resulting in the loss of important information. To address the above two issues, this paper proposes a spatial multi-prior driven resolution rescaling framework for intra-frame coding, called MP-RRF, consisting of three sub-networks: a multi-prior driven network, a downscaling network, and an upscaling network. First, the multi-prior driven network employs complexity and similarity priors to smooth the unnecessarily complicated information while leveraging similarity and quality priors to produce high-fidelity complementary information. This interaction of complexity, similarity and quality priors ensures redundancy reduction and texture enhancement. Second, the downscaling network discriminatively processes components of different granularities to generate a compact, low-resolution image for encoding. The upscaling network aggregates a complementary set of contextual multi-scale features to reconstruct realistic details while combining variable receptive fields to suppress multi-scale compression artifacts and resampling noise. Extensive experiments show that our network achieves a significant 23.84% Bjøntegaard Delta Rate (BD-Rate) reduction under all-intra configuration compared to the codec anchor, offering the state-of-the-art coding performance.
Peiying Wu, Shiwei Wang 0005, Liquan Shen, Feifeng Wang, Zhaoyi Tian
IEEE Trans. Multim.3
2024 meTMQI: multi-task and exposure-prior learning for Tone-Mapped Quality Index
Mingxing Jiang, Liquan Shen, Xiangyu Hu 0003, Min Hu 0010, Ping An 0001, Tao Tian
Vis. Comput.2
2023 Multi-Modality Deep Network for Extreme Learned Image Compression
abstract
Image-based single-modality compression learning approaches have demonstrated exceptionally powerful encoding and decoding capabilities in the past few years , but suffer from blur and severe semantics loss at extremely low bitrates. To address this issue, we propose a multimodal machine learning method for text-guided image compression, in which the semantic information of text is used as prior information to guide image compression for better compression performance. We fully study the role of text description in different components of the codec, and demonstrate its effectiveness. In addition, we adopt the image-text attention module and image-request complement module to better fuse image and text features, and propose an improved multimodal semantic-consistent loss to produce semantically complete reconstructions. Extensive experiments, including a user study, prove that our method can obtain visually pleasing results at extremely low bitrates, and achieves a comparable or even better performance than state-of-the-art methods, even though these methods are at 2x to 4x bitrates of ours.
Xuhao Jiang, Weimin Tan, Tian Tan 0016, Bo Yan 0001, Liquan Shen
AAAI5
2023 RFD-ECNet: Extreme Underwater Image Compression with Reference to Feature Dictionary
abstract
Thriving underwater applications demand efficient extreme compression technology to realize the transmission of underwater images (UWIs) in very narrow underwater bandwidth. However, existing image compression methods achieve inferior performance on UWIs because they do not consider the characteristics of UWIs: (1) Multifarious underwater styles of color shift and distance-dependent clarity, caused by the unique underwater physical imaging; (2) Massive redundancy between different UWIs, caused by the fact that different UWIs contain several common ocean objects, which have plenty of similarities in structures and semantics. To remove redundancy among UWIs, we first construct an exhaustive underwater multi-scale feature dictionary to provide coarse-to-fine reference features for UWI compression. Subsequently, an extreme UWI compression network with reference to the feature dictionary (RFD-ECNet)1is creatively proposed, which utilizes feature match and reference feature variant to significantly remove redundancy among UWIs. To align the multifarious underwater styles and improve the accuracy of feature match, an underwater style normalized block (USNB) is proposed, which utilizes underwater physical priors extracted from the underwater physical imaging model to normalize the underwater styles of dictionary features toward the input. Moreover, a reference feature variant module (RFVM) is designed to adaptively morph the reference features, improving the similarity between the reference and input features. Experimental results on four UWI datasets show that our RFD-ECNet is the first work that achieves a significant BDrate saving of 31% over the most advanced VVC.
Liquan Shen, Peng Ye 0006, Guorui Feng, Zheyin Wang
ICCV2
2023 Multi-Modality Deep Network for JPEG Artifacts Reduction
abstract
In recent years, many convolutional neural network-based models are designed for JPEG artifacts reduction, and have achieved notable progress. However, few methods are suitable for extreme low-bitrate image compression artifacts reduction. The main challenge is that the highly compressed image loses too much information, resulting in reconstructing high-quality image difficultly. To address this issue, we propose a multimodal fusion learning method for text-guided JPEG artifacts reduction, in which the corresponding text description not only provides the potential prior information of the highly compressed image, but also serves as supplementary information to assist in image deblocking. We fuse image features and text semantic features from the global and local perspectives respectively, and design a contrastive loss built upon contrastive learning to produce visually pleasing results. Extensive experiments, including a user study, prove that our method can obtain better deblocking results compared to the state-of-the-art methods.
Xuhao Jiang, Weimin Tan, Chenxi Ma, Bo Yan 0001, Liquan Shen
IJCAI6
2023 Cuboid-Net: A multi-branch convolutional neural network for joint space-time video super resolution
abstract
Abstract The demand for high‐resolution videos has been consistently rising across various domains, propelled by continuous advancements in societal. Nonetheless, limitations in imaging and economic factors often result in obtaining low‐resolution images. The currently available space‐time video super‐resolution methods often fail to fully exploit the information existing within the spatio‐temporal domain. To address this problem, the issue is tackled by conceptualizing the input low‐resolution video as a cuboid structure. An innovative methodology called “Cuboid‐Net”, which incorporates a multi‐branch convolutional neural network, is introduced. Cuboid‐Net is designed to collectively enhance the spatial and temporal resolutions of videos, enabling the extraction of rich and meaningful information across both spatial and temporal dimensions. Specifically, the input video is taken as a cuboid to generate different directional slices as input for different branches of the network. The proposed network contains four modules, that is, a multi‐branch‐based hybrid feature extraction module, a multi‐branch‐based reconstruction module, a first‐stage quality enhancement module, and a second‐stage cross frame quality enhancement module for interpolated frames only. Experimental results demonstrate that the proposed method is not only effective for spatial and temporal super‐resolution of video but also for spatial and angular super‐resolution of light field.
Congrui Fu, Hui Yuan 0001, Hongji Xu, Hao Zhang 0211, Liquan Shen
IET Image Process.5
2023 TMSO-Net: Texture adaptive multi-scale observation for light field image depth estimation
Congrui Fu, Hui Yuan 0001, Hongji Xu, Hao Zhang 0211, Liquan Shen
J. Vis. Commun. Image Represent.5
2023 Maximum Block Energy Guided Robust Subspace Clustering
abstract
Subspace clustering is useful for clustering data points according to the underlying subspaces. Many methods have been presented in recent years, among which Sparse Subspace Clustering (SSC), Low-Rank Representation (LRR) and Least Squares Regression clustering (LSR) are three representative methods. These approaches achieve good results by assuming the structure of errors as a prior and removing errors in the original input space by modeling them in their objective functions. In this paper, we propose a novel method from an energy perspective to eliminate errors in the projected space rather than the input space. Since the block diagonal property can lead to correct clustering, we measure the correctness in terms of a block in the projected space with an energy function. A correct block corresponds to the subset of columns with the maximal energy. The energy of a block is defined based on the unary column, pairwise and high-order similarity of columns for each block. We relax the energy function of a block and approximate it by a constrained homogenous function. Moreover, we propose an efficient iterative algorithm to remove errors in the projected space. Both theoretical analysis and experiments show the superiority of our method over existing solutions to the clustering problem, especially when noise exists.
Yalan Qin, Xinpeng Zhang 0001, Liquan Shen, Guorui Feng
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Prior-Guided Contrastive Image Compression for Underwater Machine Vision
abstract
Machine analysis of underwater images is essential to most underwater applications. However, both the limitation of communication bandwidth and underwater degradation bring much difficulty to accurate machine recognition at the end system. Few existing underwater compression methods consider unique underwater prior knowledge to better serve for machine vision under low bit-rates. To address this problem, we propose a novel underwater image compression framework for machines, which utilizes underwater priors to contrastively enhance degraded features by contrastive learning and efficiently compress machine-friendly features under low bit-rates. A dataset is built to provide positive and negative samples for contrastive learning based on machine analysis performance. At the encoder side, a feature extractor and a feature encoder are employed to extract machine-related features and compress them into compact representations. To alleviate the effect of underwater degradation on machine vision, a prior-guided contrastive feature enhancement module is proposed to learn more machine-friendly features based on positive and negative samples from our dataset. Then a feature refinement block is designed to remove channel-wise redundancy and focus spatial-wise importance based on high similarity of machine-related features and characteristics of underwater images. More compact representations are obtained without degrading analysis performance under low bit-rates. At the decoder side, both machine-friendly features and image are reconstructed to support different types of analysis tasks. Experimental results demonstrate the superiority of our framework in machine vision tasks compared with traditional compression methods and learned-based methods. Besides, our method still preserves basic capability of human perception.
Zhengkai Fang, Liquan Shen, Zhengyong Wang, Yanliang Jin
IEEE Trans. Circuits Syst. Video Technol.2
2023 Extreme Underwater Image Compression Using Physical Priors
abstract
Underwater images (UWIs) require higher compression ratio than terrestrial images due to the limited bandwidth of underwater wireless acoustic channel. In many studies such as marine species, the foreground objects (FGOs) in UWIs need to be observed in detail, while the background only needs to be viewed in general. However, existing image compression methods achieve limited compression ratio and reconstruction quality, which cannot fulfill these practical applications since they do not consider the unique underwater physical priors. To overcome the limitation, we propose an underwater physical prior-based extreme compression network (PPECN) for UWIs compression, which includes an underwater physical prior-guided FGOs autoencoder (UPGAE) and a FGOs-assisted background generator (FG-BGGAN). Specifically, we design an underwater physical prior guidance structure that simulates the data flow in the underwater physical imaging process to adaptively adjust the distribution of received Gaussian features in the UPGAE to be more consistent with real UWIs. During the adjustment, some basic UWI properties can be reconstructed, which can improve the reconstruction quality and implicitly reduce bits through the end-to-end training. Furthermore, the background is generated from simple semantic map under the constraint of the perceptual consistency between background and FGOs, significantly saving coding bits and improving the perceptual quality of the generated background. Extensive experimental results on four underwater image datasets verify that, compared with state-of-the-art compression methods, our PPECN achieves both impressive improvement in the perceptual quality of the whole image and significant gain in the pixel fidelity of the FGOs at the similar low bitrate.
Liquan Shen, Kun Wang 0048
IEEE Trans. Circuits Syst. Video Technol.2
2023 UIALN: Enhancement for Underwater Image With Artificial Light
abstract
Since underwater images are seriously degraded due to the attenuation of light, artificial light (AL) is often used to assist photography in underwater. However, the normal underwater imaging process is changed by the AL. It is observed that the AL source typically alters the light condition to a large extent, resulting in non-uniform illumination of images. In addition, the color distortion of the area affected by AL is little because the AL close to the object suffers little attenuation. However, most existing underwater image enhancement algorithms ignore this phenomenon. In their results, the areas affected by AL tend to be over-enhanced or over-exposed and even affect the overall enhancement effect. To this end, we propose a novel underwater image enhancement algorithm (UIALN) based on luminance correction and AL area color self-guided restoration. The underwater image is converted into the LAB color space, where the AL and pseudo-blur effect on the L channel are removed based on the luminance correction network, and the color casts on the AB channels is removed by the guidance of the AL area. Specifically, a luminance correction network is first designed based on the retinex decomposition to correct luminance, where the uneven luminance caused by AL is corrected in the illumination layer decomposed by the L channel because AL is usually white light. After that, the AL area is detected by the difference between before and after luminance correction. Second, an AL area self-guidance network is designed to assist the restoration of the color channels AB. The color restoration module utilizes the internal characteristics of the image, where the characteristics of the AL area are utilized as the prior to make the color easy to be restored. In addition, to facilitate the training and testing of the algorithm, a method of synthetic underwater images with AL is proposed based on underwater image imaging model, and a new underwater image dataset with artificial light (UIDWAL) is provided. Experimental results show that our UIALN outperforms the existing state-of-the-art approaches for the enhancement of both synthetic and real underwater images with AL.
Kun Wang 0048, Liquan Shen, Zhengyong Wang, Qijie Zhao
IEEE Trans. Circuits Syst. Video Technol.3
2023 Generation-Based Joint Luminance-Chrominance Learning for Underwater Image Quality Assessment
abstract
Underwater enhanced images (UEIs) are affected by not only the color cast and haze effect due to light attenuation and scattering, but also the over-enhancement and texture distortion caused by enhancement algorithms. However, existing underwater image quality assessment (UIQA) methods mainly focus on the inherent distortion caused by underwater optical imaging, and ignore the widespread artificial distortion, which leads to poor performance in evaluating UEIs. In this paper, a novel mapping-based underwater image quality representation is proposed. We divide underwater enhanced images into different domains and utilize a feature vector to measure the distance from the raw image domain to each enhanced image domain. The length and direction of the vector are defined as the enhancement degree and enhancement direction of the image. We construct a best enhancement direction and map other vectors to this direction to obtain the corresponding quality representation. Based on this, a novel network, called generation-based joint luminance-chrominance underwater image quality evaluation (GLCQE), is proposed, which is mainly divided into three parts: bi-directional reference generation module (BRGM), chromatic distortion evaluation network (CDEN), and sharpness distortion evaluation network (SDEN). BRGM is designed to generate two reference images about the unenhanced and the optimal enhanced versions of input UEI. In addition, the distortions in the luminance and chrominance domains of the UEI are analyzed. The luminance and chrominance channels of images are separated and input to SDEN and CDEN respectively to detect different distortions. A multi-scale feature mapping module is proposed in CDEN and SDEN to extract the feature representation of quality in chrominance and luminance of these images respectively. Moreover, a parallel spatial attention module is designed to focus on distortions in structural space by utilizing the different receptive fields of the convolution layer, due to the diverse manifestations of structural loss in the image. Finally, the mapped features extracted by two collaborative networks help the model evaluate the quality of underwater images more accurately. Extensive experiments demonstrate the superiority of our model against other representative state-of-the-art models.
Zheyin Wang, Liquan Shen, Zhengyong Wang, Yanliang Jin
IEEE Trans. Circuits Syst. Video Technol.2
2023 UCSNet: Priors Guided Adaptive Compressive Sensing Framework for Underwater Images
abstract
Image acquisition and reconstruction play an important role in underwater detections and explorations. However, the limited underwater acoustic communication channels and narrow bandwidth resources will have a great impact on the performance of the traditional data acquisition methods, resulting in the loss of details and blur in reconstructed underwater images. Compressive sensing theory (CS) which can reconstruct images from fewer measurement than that required by Nyquist sampling law has been proved to have good effect on image sampling and reconstruction. Nevertheless, the existing CS methods are not suitable for underwater images because most of them are designed for on-land images which have huge differences from underwater images. In this paper, we propose a novel priors guided adaptive underwater compressive sensing framework, dubbed UCSNet, which can effectively sample and reconstruct underwater images under a fixed low sampling ratio. In particular, our framework is composed of three sub-networks: underwater priors extraction and guidance network (UEGN), sampling matrix generation network (SMGNet) and channel-wise reconstruction network (CWRNet). Specifically, inspired by the underwater imaging physical models, UEGN is designed to extract features of underwater priors information and combine them adaptively. UEGN also introduce the imaging process into CS task to make sampling and reconstruction consistent with underwater imaging characteristics. SMGNet uses underwater content degradation to assist the analysis of structural information to generate sampling matrices. Considering the monotony of color tones caused by light absorption in underwater images, CWRNet embedded with the channel-wise module (CWM) is designed to enforce the whole network to allocate different number of sampling points on luminance and chrominance channels respectively and make feature maps extracted from them complement with each other. Experimental results demonstrate that our proposed framework can achieve both PSNR and SSIM gains on underwater images reconstruction quality and have greater visual quality than other state-of-art methods under fixed sampling ratios.
Lihao Zhuang, Liquan Shen, Zhengyong Wang, Yinyi Li
IEEE Trans. Circuits Syst. Video Technol.2
2023 Cover Selection for Steganography Using Image Similarity
abstract
Existing cover selection methods for steganography mainly focus on embedding distortion of each image, but ignore the similarity between images. When the cover images are similar, a number of relevant samples are provided to steganalysis, which is disadvantageous to steganography. This paper proposes a new cover selection method to joint image similarity and embedding distortion. Due to the difference between steganography and other image processing tasks, e.g., image reconstruction, image recognition, we propose a customized method to calculate image similarity for steganography based on SVD (singular value decomposition). Importantly, the small singular values (instead of the large ones) are employed, since it is suitable for the properties of steganography. In addition, embedding distortion is calculated by the current distortion minimization framework. The obtained image similarity and embedding distortion are combined to form a new cover selection strategy. As a result, the properties of batch images can be fully used. Experimental results show that our scheme outperforms the state-of-the-art cover selection methods when they are checked by modern steganalytic tools.
Zichi Wang, Guorui Feng, Liquan Shen, Xinpeng Zhang 0001
IEEE Trans. Dependable Secur. Comput.3
2023 Underwater Forward-Looking Sonar Images Target Detection via Speckle Reduction and Scene Prior
abstract
Forward-looking sonar (FLS) imagery system plays a significant role in oceanic object recognition and detection since it can overcome the limitation of lighting conditions and reflect the real situation of the underwater environment. However, object detection algorithms for FLS images remain challenging for two main reasons: 1) the noise caused by the coherent characteristic of the scattering phenomenon impairs the detector capture of target information and 2) the scene prior based on the uneven target scale distribution is generally neglected, which leads to the detector generating redundant anchors and slows down detection efficiency. Confronting such challenges, this article characterizes the noise and the uneven target scale distribution in FLS images as multiplicative speckle noise and scene prior, respectively. Therefore, we propose a novel underwater FLS image detection network, namely UFIDNet, to further improve detection performance by considering speckle noise reduction and scene prior in FLS images. More specifically, a speckle reduction auxiliary branch (SRAB) is designed to introduce additional despeckled supervision information to encourage the feature extractor to produce clean features and share them with the detection pipeline during the training phase. In particular, the noise distribution of FLS images is excavated for synthetic dataset construction and despeckle network (DSN) design to obtain despeckled supervision images. In addition, a feature selection strategy (FSS) embedded in detection branch is designed to screen out feature levels that do not match the target size, thus significantly reducing the generation of redundant anchors and improving detection speed. Experimental results show that our UFIDNet achieves 70.5% and 47.3% average precision (AP), 81.3% and 54.6% average recall (AR) ($\text {AR}_{\text {max=10}}$), 27.0 and 26.1 FPS on two real FLS datasets, respectively, outperforming many state-of-the-art general detectors and sonar image detectors.
Hui Long, Liquan Shen, Zhengyong Wang
IEEE Trans. Geosci. Remote. Sens.2
2023 Domain Adaptation for Underwater Image Enhancement
abstract
Recently, learning-based algorithms have shown impressive performance in underwater image enhancement. Most of them resort to training on synthetic data and obtain outstanding performance. However, these deep methods ignore the significant domain gap between the synthetic and real data (i.e., inter-domain gap), and thus the models trained on synthetic data often fail to generalize well to real-world underwater scenarios. Moreover, the complex and changeable underwater environment also causes a great distribution gap among the real data itself (i.e., intra-domain gap). However, almost no research focuses on this problem and thus their techniques often produce visually unpleasing artifacts and color distortions on various real images. Motivated by these observations, we propose a novel Two-phase Underwater Domain Adaptation network (TUDA) to simultaneously minimize the inter-domain and intra-domain gap. Concretely, in the first phase, a new triple-alignment network is designed, including a translation part for enhancing realism of input images, followed by a task-oriented enhancement part. With performing image-level, feature-level and output-level adaptation in these two parts through jointly adversarial learning, the network can better build invariance across domains and thus bridging the inter-domain gap. In the second phase, an easy-hard classification of real data according to the assessed quality of enhanced images is performed, in which a new rank-based underwater quality assessment method is embedded. By leveraging implicit quality information learned from rankings, this method can more accurately assess the perceptual quality of enhanced images. Using pseudo labels from the easy part, an easy-hard adaptation technique is then conducted to effectively decrease the intra-domain gap between easy and hard samples. Extensive experimental results demonstrate that the proposed TUDA is significantly superior to existing works in terms of both visual quality and quantitative metrics.
Zhengyong Wang, Liquan Shen, Mai Xu, Mei Yu 0001, Kun Wang 0048
IEEE Trans. Image Process.2
2023 LA-HDR: Light Adaptive HDR Reconstruction Framework for Single LDR Image Considering Varied Light Conditions
abstract
The high dynamic range (HDR) image recovery from the low dynamic range (LDR) image aims to estimate HDR image by decompressing luminance range and enhancing details of the LDR input. In practical usages, when faced with the over-exposed, the under-exposed or the low-light images, the state-of-art prediction methods lack the capability for ideally handling them. Aiming for this, a light adaptation HDR recovery framework (LA-HDR) is proposed, which includes the multi-images generation for adaptive details amplification in different light ranges, and the following multi-details fusion. To create the multi-images, first, the designed bit-depth enhancement network (EnhanceNet) produces the high bit-depth result with enhanced contrast. This result can be furtherly processed by user-defined denoising method to refrain the low-light noise. Meanwhile, the proposed exposure bias network (EBNet) estimates the global exposure bias of the input for rectifying the mid-range details. With the enhanced result and the exposure bias, the designed transfer functions adaptively create three multi-images containing the enhanced details in different light ranges, and they are fused by the designed multi-images fusion network (FuseNet) for the final HDR prediction. The amplification and fusion scheme ensures robust HDR recovery under different light conditions, eliminating high-light recovery artifacts from previous methods. The proposed fusion masks generation (FMG) and the global feature embedding (GFE) modules inFuseNethelp eliminate the fusion artifacts. Experimental results show that LA-HDR acquires the best average performance under various light conditions, and it receives low influence from the input light conditions among the tested state-of-art HDR recovery methods.
Xiangyu Hu 0003, Liquan Shen, Mingxing Jiang, Ping An 0001
IEEE Trans. Multim.2
2022 Effective QTMT Partition Decision Algorithm for VVC Intercoding
abstract
Aiming at the problem of high complexity of Versatile Video Coding (VVC) inter coding, this paper proposes a QTMT partition decision algorithm based on a multi-level decision framework. Specifically, the multi-level decision framework decomposes the multi-mode partition decision problem into multiple independent single-mode partition decision problems and then adopts a classification method based on machine learning to predict result of each single mode partition. To design more efficient classification features, fast motion estimation on 4×4 blocks is first performed to construct a motion field, and features on global/local motion and global/local consistency of its corresponding residuals are designed to measure motion activity and texture homogeneity. Furthermore, a misclassification protection mechanism is designed to decrease influences of misclassification on coding performance loss. Experimental results show that the proposed effective QTMT partition decision algorithm achieves a computational complexity reduction more than 51%, while incurring 1.65% BDBR increase compared with that of the original coding in the test model of VVC(VTM).
Liquan Shen, Hao Yang 0008, Shiwei Wang 0005
MMSP1
2022 MW-GAN+ for Perceptual Quality Enhancement on Compressed Video
abstract
The great success of deep learning has boosted the fast development of video quality enhancement. However, existing methods mainly focus on enhancing the objective quality of compressed video, and ignore their perceptual quality that plays a key role in determining quality of experience (QoE) of videos. In this paper, we aim at enhancing the perceptual quality of compressed video. Our main observation is that perceptual quality enhancement mostly relies on recovering the high-frequency details with fine textures. Accordingly, we propose a novel generative adversarial network (GAN) based on multi-level wavelet packet transform (WPT), which is called multi-level wavelet-based GAN+ (MW-GAN+), to exploit high-frequency details for enhancing the perceptual quality of compressed video. In MW-GAN+, we first propose a multi-level wavelet pixel-adaptive (MWP) module to extract temporal information across video frames, such that frame similarity can be utilized in recovering high-frequency details. Then, a wavelet reconstruction network, consisting of wavelet-dense residual blocks (WDRB), is developed to recover high-frequency details in a multi-level manner for enhanced frame reconstruction. Finally, we develop a 3D discriminator to encourage temporal coherence with a 3D-CNN based architecture. Experimental results demonstrate the superiority of our method over state-of-the-art methods in enhancing the perceptual quality of compressed video. Our code is available athttps://github.com/IceClear/MW-GAN.
Jianyi Wang, Mai Xu, Xin Deng 0002, Liquan Shen, Yuhang Song 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Human Perceptual Quality Driven Underwater Image Enhancement Framework
abstract
Underwater images suffer from severe color casts, low contrast, and blurriness, which greatly degrade the visibility and color fidelity of underwater images. Recently, numerous underwater image enhancement (UIE) algorithms have been proposed. Existing synthetic datasets-based deep learning methods employ synthetic datasets to train UIE models. However, there is a gap between synthetic datasets and real underwater images, leading to poor generalization of synthetic datasets-based UIE methods. Besides, existing real datasets-based deep learning methods largely focus on minimizing the mean squared reconstruction error between UIE results and corresponding ground-truth on the real datasets, but do not take human visual perception into account. Thus, although they achieve high PSNR between UIE results and corresponding ground-truth obtained by user study on the real datasets, they often achieve unsatisfactory perceptual quality. To address these problems, we propose a Human Perceptual Quality Driven Underwater Image Enhancement Framework (HPQ-UIEF) to achieve better results in human perceptual quality and maintain satisfactory PSNR, which is trained on a real underwater enhancement quality assessment database (UEQAB). Specifically, an Underwater Image Quality Assessment Network (UIQAN) for UIE images is first proposed to assist UIE task, in which a novel depth map prior spatial attention block (DPPAB) is embedded into UIQAN. The DPPAB can adaptively recalibrate the quality-aware feature maps and model human visual attention in a data-driven manner. Then, the UIE model is proposed, in which the UIQAN is introduced as the loss function to optimize our UIE model in the direction of perceptual metrics. Moreover, since the confidence map acquired by UIQAN can effectively reflect the sensitivity of human perceptual of local area in an UIE image, the confidence map is introduced to our UIEF to help our UIEF to perceive the perceptually important regions. Thus, the confidence map is down-sampled and then concatenated into the decoder module of the UIE model, which can further improve the perceptual quality of the UIE results. Extensive experimental results show that the proposed HPQ-UIEF outperforms state-of-the-art UIE methods qualitatively and quantitatively.
Liquan Shen, Zheyin Wang, Kun Wang 0048, Zhengyong Wang
IEEE Trans. Geosci. Remote. Sens.3
2022 Fast Intra Mode Decision Algorithm for Versatile Video Coding
abstract
To achieve higher coding efficiency, the latest Versatile Video Coding (VVC) standard adopts a series of new intra coding techniques, including the quadtree plus multi-type tree (QTMT), intra sub-partitions (ISP) and intra block copy (IBC). However, this makes the intra coding more complicated, as VVC needs to traverse all prediction modes and partition types of QTMT to find the optimal combination. In this paper, we propose a fast algorithm for VVC from two aspects of mode selection and prediction terminating to reduce coding complexity. For the mode selection, adaptive mode pruning (AMP) is proposed to remove non-promising modes. First, since the newly introduced modes (IBC and ISP) are not effective for all blocks, learning-based classifiers are designed to remove them intelligently. Second, for normal modes, an ensemble decision strategy is proposed to sort the candidate modes and increase the probability of being the optimal mode for the first few candidates; thus, we can remove redundant candidates more efficiently. In terms of prediction terminating, we find that different optimal modes of current depth level lead to different termination probabilities of remaining intra predictions. Therefore, mode-dependent termination (MDT) is proposed to select an appropriate model through the optimal mode and terminate unnecessary intra predictions of remaining depth levels. The proposed algorithm is implemented on VVC test model, and simulation results show that it can achieve 51%$\sim$53% time savings with only 0.93%$\sim$1.08% BDBR increases.
Xinchao Dong, Liquan Shen, Mei Yu 0001, Hao Yang 0008
IEEE Trans. Multim.2
2022 Low Bitrate Light Field Compression With Geometry and Content Consistency
abstract
Light field imaging can simultaneously record the position and direction information of light rays; thus, digital refocusing and full depth-of-field extension — functions that are inaccessible for conventional images — can be achieved using the structural consistency of light field data. To meet the challenges of limited bandwidth and storage, such vast numbers of light field data must be compressed to a low bitrate. However, current compression solutions ignore the intrinsic consistency of light fields in pursuit of a low bitrate, thereby leading to the loss of light field capabilities. To solve this issue, this work focuses on structural consistency to achieve efficient light field compression with a low bitrate. The proposed light field compression method encodes the sparsely selected sub-aperture images (SAIs) and the disparity maps corresponding to the unselected SAIs. From the perspective of geometry consistency, the consistency of the initially estimated disparity maps is improved by using a color-guided refinement algorithm, thereby reducing the bitrate of the disparity maps. From the perspective of content consistency, the consistency of the SAI-transformed pseudo sequence is improved by the proposed content-similarity-based arrangement algorithm along with a specific prediction structure; thereby, the bitrate of the sparsely selected SAIs is reduced. The experimental results show that the proposed compression method can reduce the total bitrate while preserving good structural consistency.
Xinpeng Huang, Ping An 0001, Deyang Liu, Liquan Shen
IEEE Trans. Multim.5
2022 Objective Quality Assessment of Lenslet Light Field Image Based on Focus Stack
abstract
The large amount of complex scene information recorded by light field imaging has the potential for immersive media applications. Compression and reconstruction algorithms are crucial for the transmission, storage, and display of such massive data. Most of the existing quality evaluation indexes do not effectively account for light field characteristics. To accurately evaluate the distortions caused by compression and reconstruction algorithms, it is necessary to construct an image evaluation index that reflects the angular-spatial characteristics of the light field. This work proposes a full-reference light field image quality evaluation index that attempts to extract less information from the focus stack to accurately evaluate the entire light field quality. The proposed framework includes three specific steps. First, we construct a key refocused image extraction framework by the maximal spatial information contrast and the minimal angular information variation. Specifically, the gradient and phase congruency operators are used in the extraction framework. Second, a novel light field quality evaluation index is built based on the angular-spatial characteristics of the key refocused images. In detail, the features used in the key refocused image extraction framework and the chrominance feature are combined to construct the union feature. Third, the similarity of the union feature is pooled by the relevant visual saliency map to obtain the predicted score. Finally, the overall quality of the light field is measured by applying the proposed index to the key refocused images. The high efficiency and precision of the proposed method are shown by extensive comparison experiments.
Chunli Meng, Ping An 0001, Xinpeng Huang, Chao Yang 0021, Liquan Shen
IEEE Trans. Multim.5
2021 An optimized CNN-based quality assessment model for screen content image
Xuhao Jiang, Liquan Shen, Guorui Feng, Liangwei Yu, Ping An 0001
Signal Process. Image Commun.2
2021 Attenuation Coefficient Guided Two-Stage Network for Underwater Image Restoration
abstract
Underwater images suffer from severe color casts, low contrast and blurriness, which are caused by scattering and absorption when light propagates through water. However, existing deep learning methods treat the restoration process as a whole and do not fully consider the underwater physical distortion process. Thus, they cannot adequately tackle both absorption and scattering, leading to poor restoration results. To address this problem, we propose a novel two-stage network for underwater image restoration (UIR), which divides the restoration process into two parts viz. horizontal and vertical distortion restoration. In the first stage, a model-based network is proposed to handle horizontal distortion by directly embedding the underwater physical model into the network. The attenuation coefficient, as a feature representation in characterizing water type information, is first estimated to guide the accurate estimation of the parameters in the physical model. For the second stage, to tackle vertical distortion and reconstruct the clear underwater image, we put forth a novel attenuation coefficient prior attention block (ACPAB) to adaptively recalibrate the RGB channel-wise feature maps of the image suffering from the vertical distortion. Experiments on both synthetic dataset and real-world underwater images demonstrate that our method can effectively tackle scattering and absorption compared with several state-of-the-art methods.
Liquan Shen, Zhengyong Wang, Kun Wang 0048
IEEE Signal Process. Lett.2
2021 A Distortion-Aware Multi-Task Learning Framework for Fractional Interpolation in Video Coding
abstract
Motion-compensated prediction adopts fractional-pixel interpolation to obtain the best motion vector. Traditional fixed interpolation filters cannot handle various content and structures well, and existing convolutional neural network based methods cannot fully exploit the distortion characteristics for fractional interpolation. Therefore, this paper proposes a distortion-aware multi-task learning framework (DA-MLF) to perform fractional interpolation. First, a multi-task training framework is proposed to provide the distortion characteristics as complementary information for improving the performance of subsequent interpolation. Then, a uniform interpolation sub-network is proposed to accomplish fractional interpolation, which utilizes the feature fusion module to fuse abundant local features, and the distortion awareness module to capture the multi-scale information of compression artifacts. Furthermore, DA-MLF is integrated into High Efficiency Video Coding (HEVC) test model, and multiple experiments are performed to evaluate the effectiveness of our method. On HEVC testing sequences, DA-MLF achieves 5.0%, 4.0% and 1.7% BD-rate reduction on average compared to the HEVC baseline, under low-delay P, low-delay B and random-access configurations, respectively. The experimental results validate that our framework not only achieves the best interpolation performance but also has the lowest computational complexity compared with state-of-the-art methods.
Liangwei Yu, Liquan Shen, Hao Yang 0008, Xuhao Jiang, Bo Yan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Patch-Wise Spatial-Temporal Quality Enhancement for HEVC Compressed Video
abstract
Recently, many deep learning based researches are conducted to explore the potential quality improvement of compressed videos. These methods mostly utilize either the spatial or temporal information to perform frame-level video enhancement. However, they fail in combining different spatial-temporal information to adaptively utilize adjacent patches to enhance the current patch and achieve limited enhancement performance especially on scene-changing and strong-motion videos. To overcome these limitations, we propose a patch-wise spatial-temporal quality enhancement network which firstly extracts spatial and temporal features, then recalibrates and fuses the obtained spatial and temporal features. Specifically, we design a temporal and spatial-wise attention-based feature distillation structure to adaptively utilize the adjacent patches for distilling patch-wise temporal features. For adaptively enhancing different patch with spatial and temporal information, a channel and spatial-wise attention fusion block is proposed to achieve patch-wise recalibration and fusion of spatial and temporal features. Experimental results demonstrate our network achieves peak signal-to-noise ratio improvement, 0.55 - 0.69 dB compared with the compressed videos at different quantization parameters, outperforming state-of-the-art approach.
Liquan Shen, Liangwei Yu, Hao Yang 0008, Mai Xu
IEEE Trans. Image Process.2
2021 Blind Image Quality Assessment Based on Multi-scale KLT
abstract
Blind image quality assessment (BIQA) plays an important role in image services as independent of the reference image. Herein, the perceptual relevant feature design is the core of BIQA methods, but their performance is still not satisfied at present. In this work, we propose an unsupervised feature extraction approach for BIQA based on Karhunen-Loéve transform (KLT). Specifically, a normalization operation is firstly applied to the test image by calculating its mean subtracted contrast normalized (MSCN) coefficient. Then, KLT is employed as a data-driven feature extraction approach to extract image structural features, wherein kernels with different sizes are utilized to perform multi-scale analysis. Finally, generalized Gaussian distribution (GGD) is employed to model the KLT coefficients distribution in different spectral components as quality relevant features. Extensive experiments conducted on four widely utilized IQA databases have demonstrated that the proposed Multi-scale KLT (MsKLT) BIQA metric compares favorably with existing BIQA methods in terms of high accordance with human subjective scores on both common and uncommon distortion types.
Chao Yang 0021, Xinfeng Zhang 0001, Ping An 0001, Liquan Shen, C.-C. Jay Kuo
IEEE Trans. Multim.4
2020 No-reference screen content image quality assessment based on multi-region features
Xuhao Jiang, Liquan Shen, Liangwei Yu, Mingxing Jiang, Guorui Feng
Neurocomputing2
2020 Video intra prediction using convolutional encoder decoder network
Zhipeng Jin, Ping An 0001, Liquan Shen
Neurocomputing3
2020 Post-processing for intra coding through perceptual adversarial learning and progressive refinement
Zhipeng Jin, Ping An 0001, Chao Yang 0021, Liquan Shen
Neurocomputing4
2020 Screen content image quality assessment based on convolutional neural networks
Xuhao Jiang, Liquan Shen, Linru Zheng, Ping An 0001
J. Vis. Commun. Image Represent.2
2020 Background foreground boundary aware efficient motion search for surveillance videos
Tushar Shankar Shinde, Anil Kumar Tiwari, Weiyao Lin, Liquan Shen
Signal Process. Image Commun.4
2020 VRFCNN: Virtual Reference Frame Generation Network for Quality SHVC
abstract
For more efficient inter prediction in quality scalable high efficiency video coding (SHVC), a learning-based framework for generating a virtual reference frame (VRF) is proposed in this letter. In our method, reconstructed base layer (BL) and enhancement layer (EL) frames are employed to make the generated VRF and current EL frame as same as possible. To this end, this letter proposes a novel VRF generation convolutional neural network (VRFCNN) to jointly handle enhancement of corresponding BL frame and compensation of previous EL frame. Specifically, the VRFCNN consists of BL enhancement, EL compensation and feature fusion subnets. Previous EL frames are firstly compensated with the learned coarse flow between two adjacent BL frames and then adopted to provide interlayer information for corresponding BL enhancement. The learned finer flow between the EL and enhanced BL features is adopted to provide temporal information for previous EL compensation. For efficiently handling slow- and fast-motion videos, the enhanced BL and compensated EL features are fused to generate a VRF. Experimental results show that VRFCNN averagely achieves 11.8% BD-rate reduction under low delay P configuration, which outperforms other methods. The code of our VRFCNN approach is available at https://github.com/dq0309/VRFCNN.
Liquan Shen, Hao Yang 0008, Xinchao Dong, Mai Xu
IEEE Signal Process. Lett.2
2020 Recognition-Driven Compressed Image Generation Using Semantic-Prior Information
abstract
Image compression, as one of the fundamental low-level image processing tasks, is very essential for computer vision. Current image compression methods can maintain considerable visual quality even at relatively lower bit-rate, but pay little attention to their performance in high-level vision tasks, e.g, image recognition and semantic segmentation. In this letter, we aim to generate compressed images with similar visual quality as before, but with much higher recognition accuracy. To this end, we explore the significance of semantic-prior information in image compression and design a semantic-prior attention module to adaptively enhance the semantic-ware features. Moreover, a semantic perceptual loss combining semantic-prior map is employed to concentrate on the machine perceptual quality of semantic-ware content. Experiments on benchmarks demonstrate that the proposed algorithm has the effectiveness to improve recognition accuracy and maintain visual quality. In addition, performance improvement on different recognition networks and tasks shows the generality of our algorithm.
Qiang Wang 0036, Liquan Shen
IEEE Signal Process. Lett.2
2020 On Security Enhancement of Steganography via Generative Adversarial Image
abstract
Steganography plays an important role in information hiding. With the development of steganalysis, traditional steganography faces more detection threat. It is necessary to improve security of current steganographic methods. One effective way is to generate suitable covers for steganography, which can be achieved by adversarial learning. In this letter, we propose a new approach for quickly constructing high-quality adversarial images. Compared with original images, the generative adversarial images are more suitable for carrying secret information. According to the characteristics of steganography, we design a new loss function in adversarial attacks, which makes the adversarial images obtain the similar classification results before and after steganography. In addition, to further improve security of the adversarial images, we also make use of the zero-sum idea of generative adversarial networks. Experimental results show that the proposed method can significantly enhance security of steganography.
Lingchen Zhou, Guorui Feng, Liquan Shen, Xinpeng Zhang 0001
IEEE Signal Process. Lett.3
2020 Low-Complexity CTU Partition Structure Decision and Fast Intra Mode Decision for Versatile Video Coding
abstract
Quadtree with nested multi-type tree (QTMT) partition structure is an efficient improvement in versatile video coding (VVC) over the quadtree (QT) structure in the advanced high-efficiency video coding (HEVC) standard. With the exception of the recursive QT partition structure, recursive multi-type tree partition is applied to each leaf node, which generates more flexible block sizes. Besides, intra prediction modes are extended from 35 to 67 so as to satisfy various texture patterns. These newly developed techniques achieve high coding efficiency but also result in very high computational complexity. To tackle this problem, we propose a fast intra-coding algorithm consisting of low-complexity coding tree units (CTU) structure decision and fast intra mode decision in this paper. The contributions of the proposed algorithm lie in the following aspects: 1) the new block size and coding mode distribution features are first explored for a reasonable fast coding scheme; 2) a novel fast QTMT partition decision framework is developed, which can determine the partition decision on both QT and multi-type tree with a novel cascade decision structure; and 3) fast intra mode decision with gradient descent search is introduced, while the best initial search point and search step are also investigated in this paper. The simulation results show that the complexity reduction of the proposed algorithm is up to 70% compared to VVC reference software (VTM), and averagely 63% encoding time saving is achieved with 1.93% BDBR increasing. Such results demonstrate that our method yields a superior performance in terms of computational complexity and compression quality compared to the state-of-the-art methods.
Hao Yang 0008, Liquan Shen, Xinchao Dong, Ping An 0001, Gangyi Jiang
IEEE Trans. Circuits Syst. Video Technol.2
2020 Accelerate CTU Partition to Real Time for HEVC Encoding With Complexity Control
abstract
Recently, extensive approaches have been proposed for reducing the encoding complexity of high efficiency video coding, by predicting the coding tree unit partition using deep neural networks. However, these approaches cannot work in real time due to the complexity of the network architectures. In this paper, we propose a network pruning approach to accelerate a state-of-the-art deep neural network model, for real-time coding tree unit partition. Specifically, we first investigate the computational complexity throughout the network, and find that most calculations can be simplified by pruning the weight parameters. Considering that the number of weight parameters drastically differs by network layer and partition level, we design an adaptive pruning scheme by applying a well-suitable retention ratio of weight parameters to each layer at a level. The retention ratio indicates the ratio of weight parameters after and before pruning. By varying the retention ratios, we can obtain several accelerated network models with different levels of complexity. We further propose a complexity control algorithm by applying different accelerated models to different coding tree units, to ensure that the actual encoding complexity is close to a given target. To guarantee the rate-distortion performance, we model the complexity control algorithm as a convex optimization problem, and we can obtain a closed-form solution. Experimental results show that our approach can accelerate the original deep neural network model by 17-20 times, with little expense on the Bjøntegaard delta bit-rate. For complexity control, we achieve high control accuracy with a control error of less than 2% for most video sequences.
Tianyi Li 0004, Mai Xu, Xin Deng 0002, Liquan Shen
IEEE Trans. Image Process.4
2019 A Novel No-Reference Quality Assessment Model of Tone-Mapped HDR Image
abstract
Research on tone mapping operators (TMOs) attracts more attention recently, which can transform high dynamic range (HDR) images to low dynamic range (LDR) images for visualizing them on the common displays. In this paper, we propose a novel no-reference image quality assessment (IQA) model to evaluate the perceptual quality of tone-mapped images (TMIs). Specifically, local phase congruency (LPC) is first computed to evaluate the image sharpness and some statistical characteristics are extracted on the edge maps to measure the halo effect. Meanwhile, TMIs are transformed to opponent color (OC) space to gain the global image chromaticity and local image contract in the chromatic field. Finally, a regression module is learnt using support vector regression (SVR) to train the mapping function that maps all the features to subjective quality scores. The model shows admirable performance when tested on ESPL-LIVE HDR image database.
Liquan Shen, Mingxing Jiang, Linru Zheng, Ping An 0001
ICIP2
2019 A content-based rate control algorithm for screen content video coding
Liquan Shen, Hao Yang 0008, Ping An 0001
J. Vis. Commun. Image Represent.2
2019 Selective Ensemble Classification of Image Steganalysis Via Deep Q Network
abstract
Currently many important advances in digital media field have been achieved through the ensemble method. In image steganalysis, the application of classifiers has evolved from the early single classifier to the ensemble classifiers. The performance of the ensemble classifier is better than that of a single classifier, but the classifiers may have a certain degree of redundancy. Therefore, it is of great significance to study how to reduce the number of the ensemble classifiers under the premise of ensuring the classification performance. In this letter, we propose a selective ensemble method in image steganalysis based on deep Q network, which combines reinforcement learning with convolutional neural network and are seldom seen in ensemble pruning. This method improves the generalization performance of the model, and reduces the size of ensemble as well. The experimental results show that the proposed method has a certain degree of effect on the ensemble classification optimization of image steganalysis in both spatial and frequency domains.
Danni Ni, Guorui Feng, Liquan Shen, Xinpeng Zhang 0001
IEEE Signal Process. Lett.3
2019 Quality Enhancement Network via Multi-Reconstruction Recursive Residual Learning for Video Coding
abstract
Lossy compression algorithms introduce multiple compression artifacts that severely decrease visual quality. These compression artifacts are highly related to texture contents, and the hierarchical coding units decision structure also brings multi-scale similarity to these artifacts. Current loop filters fail to utilize these characteristics to comprehensively remove compression artifacts. To this end, this letter proposes a novel quality enhancement method by adopting a multi-reconstruction recurrent residual network (MRRN). In particular, a modified recursive residual structure is designed to capture the multi-scale similarity of compression artifact. To effectively enhance frames with uneven noise, a multi-reconstruction structure is proposed, which outputs images with different denoise ratios and adaptively fuses them. Experimental results show that the proposed MRRN can improve coding efficiency up to 15.1% compared with the original loop filter in high-efficiency video coding. Averagely, 6.7%, 7.8%, 7.6% BD-rate reduction is achieved for all intra, low-delay P, and low-delay B, respectively. Meanwhile, as a quality enhancement method performed at encoder side, MRRN also achieves a good balance between coding performance and computational complexity compared to the state-of-the-art methods.
Liangwei Yu, Liquan Shen, Hao Yang 0008, Ping An 0001
IEEE Signal Process. Lett.2
2019 SHVC CU Processing Aided by a Feedforward Neural Network
abstract
The development of multimedia and hardware technologies has led to a great number of industrial video applications, such as virtual reality, high-definition video surveillance, and remote monitoring. As complex communication environments and heterogeneous networks are common in industrial applications, industrial videos are required to support a diverse range of display resolutions and transmission channel capacities. Scalable high-efficiency video coding (SHVC) standards provide the tools to meet this requirement. However, it is highly computationally expensive. Coding complexity has a great impact on SHVC performance in industrial applications. Many of these applications are sensitive to time delay and have limited power. Thus, improvements are required to ensure the practical usability of SHVC encoders. In SHVC encoders, intra/interprediction of variable coding unit (CU) sizes is independently performed for the base and enhancement layers (ELs). There are many interlayer similarities that can be exploited to speed up the procedure for EL coding. In this paper, we propose a feedforward neural network aided model for CU size and mode decisions for SHVC, which utilizes base layer coding information and the coding data of spatiotemporal neighboring CUs to decide which CU sizes or prediction modes can be bypassed for certain EL CUs. Two feedforward neural network based learning models are built for CU classification, which are introduced in the procedures for CU size and mode decisions, respectively. According to the analysis from a large number of video sequences, the representative features are directly extracted from the coding information of previously coded neighboring CUs to avoid computational overheads. After the training is finished, these two models are designed and integrated to build classifiers. Then, two online classification approaches are designed for the CU size and mode decision procedures to classify each CU's type. Finally, different candidate CU sizes and prediction modes are adaptively assigned for each type of CU. This approach outperforms the state-of-the-art fast SHVC/high-efficiency video coding (HEVC) algorithms with approximately 19-42% coding time savings or better compression efficiency, which will be beneficial for the realization of real-time scalable video coding.
Liquan Shen, Guorui Feng, Ping An 0001
IEEE Trans. Ind. Informatics1
2019 A Modified Just Noticeable Depth Difference Model Built in Perceived Depth Space
abstract
This paper proposes a modified just noticeable depth difference (JNDD) (JNDiD) model in perceived depth space. The JNDiD model improves the accuracy of current JNDD models by also taking into account the blurriness caused by the change in accommodation when a 3-D object is perceived far from the screen. This change in accommodation is a result of the convergence-accommodation conflict, in which convergence plays a leading role. In the JNDiD model described in this paper, the JNDD threshold is the addition between a base threshold and an additional threshold. Adapted to different blurring effects in three depth regions, the additional threshold is defined as a three-piecewise linear function. The proposed model also attempts to separate the JNDD modeling from the display modeling and viewing conditions in perceived depth space, making it applicable for different types of displays given their specific model parameters. With the help of model parameter transfer functions, the JNDiD model in perceived depth space and the corresponding one in any specific stimulated depth space can be easily transformed to each other. Experimental results demonstrate the effectiveness and superiority of the JNDiD model in perceived depth space.
Ping An 0001, Liquan Shen, Kai Li 0016
IEEE Trans. Multim.3
2019 Content-Based Adaptive SHVC Mode Decision Algorithm
abstract
The scalable video coding extensions of the High Efficient Video Coding (HEVC) standard (SHVC) have adopted a new quadtree-structured coding unit (CU). The SHVC test model (SHM) needs to test seven intermode sizes and one intramode size at depth levels of “0,” “1,” “2,” and four intermode sizes and two intramode sizes at a depth level of “3” for interframe CUs. It checks all possible depth levels and prediction modes to find the one with the lowest rate distortion cost using the Lagrange multiplier method in the mode decision procedure to achieve high coding efficiency at the expense of computational complexity. Furthermore, it utilizes the conventional approach for the base layer (BL) and enhancement layer (EL) coding to support SNR/spatial scalable coding. Both the intralayer and interlayer predictions should be performed for each EL CU. Although there is a large amount of interlayer redundancy that can be exploited to speed up the EL encoding, the mode decision procedure is independently performed for the BL and the ELs. In this paper, we propose a content-adaptive mode decision algorithm to reduce the SHVC complexity at the ELs. When the major characteristics of the CUs, such as mode complexity and motion activity, can be estimated early and used for adjusting the mode decision procedure, unnecessary mode and CU size searches can be avoided. First, an experimental analysis is performed to study the interlayer and spatiotemporal correlations in the coding information and the interlevel correlations among the quadtree structures. Based on these correlations, three parameters, including the conditional probability of a SKIP/Merge mode, motion activity, and mode complexity, are defined to describe the video content and are further utilized to adaptively adjust the EL mode decision procedure. The experimental results show that the proposed algorithm can reduce the coding time for ELs by 62%-67% with less than a 1.5% Bjontegaard rate increase compared to the original SHVC encoder.
Liquan Shen, Guorui Feng
IEEE Trans. Multim.1
2019 No-Reference Quality Assessment for Screen Content Images Based on Hybrid Region Features Fusion
abstract
Research on screen content images (SCIs) attracts more attention as they are highly applied to image- and video-centric applications on mobile and other devices. It is important to develop an efficient image-quality assessment (IQA) method for SCIs because IQA can guide and optimize various image-processing methods for SCIs and improve user experience. In this paper, we propose a no-reference objective assessment model for SCIs including SCIs segmentation and the analysis of local and global perceptual feature representations. Since the human visual system is highly sensitive to sharp edges that are commonly encountered in SCIs, we utilize the variance of local standard deviation, which is a noise robust index to distinguish the sharp edge patches (SEPes) and non-SEPes of SCIs. For SEPes, we perform two kinds of feature extractions. First, the entropy and contrast features are extracted with a gray-level co-occurrence matrix, which are highly perceptive of microstructural change. Second, the local phase coherence is utilized to capture the loss in sharpness. Then, average pooling is adopted to fuse features obtained from all of the SEPes to represent the local features. We further combine local features with global features that are derived using the BRISQUE method as the hybrid region (HR)-based features. Finally, a regression module is learned using support vector regression to train the mapping function that maps HR-based features to subjective quality scores. Experimental results on the screen image-quality assessment database show that the proposed method can achieve better performance in visual-quality prediction for SCIs than the performance achieved by state-of-the-art methods.
Linru Zheng, Liquan Shen, Jianan Chen 0001, Ping An 0001, Jun Luo 0006
IEEE Trans. Multim.2
2019 Low-Complexity Scalable Extension of the High-Efficiency Video Coding (SHVC) Encoding System
abstract
The scalable extension of the high-efficiency video coding (SHVC) system adopts a hierarchical quadtree-based coding unit (CU) that is suitable for various texture and motion properties of videos. Currently, the test model of SHVC identifies the optimal CU size by performing an exhaustive quadtree depth-level search, which achieves a high compression efficiency at a heavy cost in terms of the computational complexity. However, many interactive multimedia applications, such as remote monitoring and video surveillance, which are sensitive to time delays, have insufficient computational power for coding high-definition (HD) and ultra-high-definition (UHD) videos. Therefore, it is important, yet challenging, to optimize the SHVC coding procedure and accelerate video coding. In this article, we propose a fast CU quadtree depth-level decision algorithm for inter-frames on enhancement layers that is based on an analysis of inter-layer, spatial, and temporal correlations. When motion/texture properties of coding regions can be identified early, a fast algorithm can be designed for adapting CU depth-level decision procedures to video contents and avoiding unnecessary computations during CU depth-level traversal. The proposed algorithm determines the motion activity level at the treeblock size of the hierarchical quadtree by utilizing motion vectors from its corresponding blocks at the base layer. Based on the motion activity level, neighboring encoded CUs that have larger correlations are preferentially selected to predict the optimal depth level of the current treeblock. Finally, two parameters, namely, the motion activity level and the predicted CU depth level, are used to identify a subset of candidate CU depth levels and adaptively optimize CU depth-level decision processes. The experimental results demonstrate that the proposed scheme can run approximately three times faster than the most recent SHVC reference software, with a negligible loss of compression efficiency. The proposed scheme is efficient for all types of scalable video sequences under various coding conditions and outperforms state-of-the-art fast SHVC and HEVC algorithms. Our scheme is a suitable candidate for interactive HD/UHD video applications that are expected to operate in real-time and power-constrained scenarios.
Liquan Shen, Ping An 0001, Guorui Feng
ACM Trans. Multim. Comput. Commun. Appl.1
2018 Quality Enhancement for Intra Frame Coding Via Cnns: An Adversarial Approach
abstract
Lossy compression is an indispensable technique in image/video processing, due to its highly desirable ability of reducing the huge data volume. However, lossy compression introduces complex compression artifacts. To reduce these artifacts, post-processing techniques have been extensively studied. In this paper, we propose a novel post-processing technique using multi-level progressive refinement network via an adversarial training approach, called MPRGAN, for artifacts reduction and coding efficiency improvement in intra frame coding. Furthermore, our network generates multi-level residues in one feed-forward pass through the progressive reconstruction. This coarse-to-fine work fashion, which makes our network have high flexibility, can make trade-off between enhanced quality and computational complexity. Thereby facilitates the resource-aware applications. Extensive evaluations on benchmark datasets verify the superiority of our proposed MPRGAN model over the latest state-of-the-art methods with fast deployment running speed.
Zhipeng Jin, Ping An 0001, Chao Yang 0021, Liquan Shen
ICASSP4
2018 View Synthesis for Light Field Coding Using Depth Estimation
abstract
Light Field (LF) image captured by plenoptic camera can record richer scenario information from our world. But a huge number of Sub-Aperture Images (SAIs) from LF image results in great challenges for coding LFI. Therefore, we choose a subset of SAIs as multi-view video, and encode them with their corresponding depth maps using Multi-view Video plus Depth (MVD) structure. Due to lack of depth maps for SAIs, we propose a cost function to determine the depth map for each SAI preliminarily based on horizontal and vertical Epipolar Plane Images (EPI), respectively. Then an SAI-guided depth enhancement algorithm is designed to optimize the estimated depth maps. Since those unselected SAIs have not been encoded but have been synthesized using the specific texture image and depth map, our LF image coding method can naturally achieve bitrates reduction dramatically with a good performance and outperform other algorithms significantly.
Xinpeng Huang, Ping An 0001, Liang Shan 0004, Liquan Shen
ICME5
2018 No-reference stereo image quality assessment by learning gradient dictionary-based color visual characteristics
abstract
In this paper, we propose a no-reference (NR) stereo image quality assessment metric by learning gradient dictionary-based color visual characteristics. To be specific, firstly, since human eyes are highly sensitive to the structure of images, the gradient magnitude (GM) and gradient orientation (GO) are extracted from left and right views of stereo image, meanwhile, the difference map is obtained. Considering the influence of color distortion, images are decomposed into RGB channels to be processed respectively, and we get the local gradient of the color image by adding up the RGB gradient vectors. Constructively, the gradient dictionary is generated, which is different from traditional image dictionary. All quality-aware features are extracted by joint sparse representation. Afterwards, to avoid over-fitting, the principal component analysis (PCA) is applied to optimize the quality-aware features. Finally, all features are fed into the trained support vector regression (SVR) model to predict the objective score. The experimental results show that the proposed metric always achieves high consistency with human subjective assessment for both symmetric and asymmetric distortions.
Jialu Yang, Ping An 0001, Jian Ma 0005, Kai Li 0016, Liquan Shen
ISCAS5
2018 Deep learning for steganalysis based on filter diversity selection
Guorui Feng, Liquan Shen, Jun Luo 0006
Sci. China Inf. Sci.3
2018 Hybrid linear weighted prediction and intra block copy based light field image coding
Deyang Liu, Ping An 0001, Liquan Shen
Multim. Tools Appl.4
2018 Scalable coding of 3D holoscopic image by using a sparse interlaced view image set and disparity map
Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen, Kai Li 0016
Multim. Tools Appl.5
2018 Efficient coding mode and partition decision for screen content intra coding
abstract
In order to tackle the exclusive characteristic in the screen content video, new coding tools such as intra block copy mode and palette mode are adopted in the screen content coding (SCC) based on high efficiency video coding (HEVC) standard. It results in high coding complexity for SCC and is hard to satisfy real time applications. An efficient intra coding algorithm for screen content video is proposed in this paper. Early skipping or early termination of coding unit (CU) partition and mode prediction is applied to speed up intra coding by exploiting the correlated information among luminance, gradient and neighboring CUs. Furthermore, the correlated information of hit rate and coding bit is revealed and properly used for fast CU depth decision. Experimental results demonstrate that the proposed algorithm obtains significant time saving with ignorable compression loss in comparison with several recent algorithms.
Huaping Liu 0002, Yameng Lin, Liquan Shen, Haibing Yin
Signal Process. Image Commun.4
2018 Joint binocular energy-contrast perception for quality assessment of stereoscopic images
Jian Ma 0005, Ping An 0001, Liquan Shen, Kai Li 0016
Signal Process. Image Commun.3
2018 Efficient screen content intra coding based on statistical learning
Hao Yang 0008, Liquan Shen, Ping An 0001
Signal Process. Image Commun.2
2018 Bivariate analysis of 3D structure for stereoscopic image quality assessment
Liquan Shen, Ping An 0001
Signal Process. Image Commun.2
2018 Naturalization Module in Neural Networks for Screen Content Image Quality Assessment
abstract
Deep learning approaches have demonstrated success in no-reference image quality assessment tasks. However, due to the specific properties of screen content images (SCIs), deep neural networks for SCI quality assessment are not as optimal as those designed for images depicting natural scenes. In order to tackle this discrepancy, a “naturalization” module composed of an upsampling layer and a convolutional layer is proposed to transform SCIs to have characteristics more similar to that of natural images. In addition, a new deep learning model architecture along with data augmentation techniques tailored to SCIs are implemented. The performance of the proposed approach is evaluated on the Screen Image Quality Assessment Database and Screen Content Image Database, and has shown to have superior performance to state-of-the-art methods in predicting the perceptual quality of SCIs.
Jianan Chen 0001, Liquan Shen, Linru Zheng, Xuhao Jiang
IEEE Signal Process. Lett.2
2018 Fast Intra Coding of High Dynamic Range Videos in SHVC
abstract
Compared with the conventional standard dynamic range (SDR) content, high dynamic range (HDR) content supplies viewers with more immersive experience by offering a much higher range of luminance. Most of current consumer devices cannot afford to this emerging technology, and content providers decide to create both an HDR version and an SDR version of the same video. In this letter, scalable high efficiency video coding (HEVC) scalable extension of HEVC (SHVC) serves as the coding framework where the base layer (BL) is an 8-b SDR version and the enhancement layer (EL) is a 12-b HDR version. Recently, many fast coding algorithms for SDR videos are proposed, and there is an urgent demand for fast coding algorithms for EL HDR videos. With the coding information of the BL SDR videos, this letter proposes a fast algorithm to reduce the complexity of intra coding for EL HDR videos. First, depth information of neighboring coding tree units (CTUs) in the HDR version and the colocated CTU in the SDR version is used for early coding unit (CU) depth determination. Moreover, four classifiers are trained to predict the CTU depth range. Two classifiers are trained for CTUs in frames with a high average luma, and another two classifiers are used for CTUs in frames with a low average luma. Experimental results show that the proposed algorithm achieves 43% encoding time saving on average, with only a 0.54% Bjøntegaard delta bit rate (BDBR) increase compared to the original SHVC test model.
Guoliang Fu, Liquan Shen, Hao Yang 0008, Xiangyu Hu 0003, Ping An 0001
IEEE Signal Process. Lett.2
2018 Efficient Intra Mode Selection for Depth-Map Coding Utilizing Spatiotemporal, Inter-Component and Inter-View Correlations in 3D-HEVC
abstract
3D-high efficiency video coding (HEVC) is developed for the compression of the multi-view video plus depth format, which is based on the latest generation of video coding standard, HEVC. It further adopts several new intra prediction modes, depth-modeling modes (DMMs) in intra candidate modes for a better representation of edges in depth maps, which introduces a drastic increase in the computational complexity. The procedure of depth intra mode decision together with DMMs and existing intra modes is a very time consuming part due to huge complexity of full rate distortion (RD) cost calculation. In this paper, a low complexity intra mode selection algorithm is proposed to reduce complexity of depth intra prediction in both intra-frames and inter-frames. An experimental analysis is first performed to study the inter-view correlation and the inter-component (texture video and its associated depth) correlation in intra coding information such as the intra mode and RD cost. All intra modes available in 3D-HEVC are classified into three activity classes assigned with different mode-weight factors, and the coding mode complexity of a coding unit (CU) is defined according to the intra mode information from available spatiotemporal, inter-view, and inter-component neighboring coded CUs. The coding mode complexity analysis is utilized to assign different candidate intra modes for different types of CUs. The optimal intra prediction mode and the RD cost value in current CU depth level are further used to skip unnecessary intra prediction sizes. Experimental results show that the proposed fast depth intra coding algorithm achieves 61% complexity reduction on intra prediction, while incurring a 0.2% Bjontegaard metric increase for coded and synthesized views compared to the test model of 3D-HEVC.
Liquan Shen, Kai Li 0016, Guorui Feng, Ping An 0001, Zhi Liu 0003
IEEE Trans. Image Process.1
2017 Coding of 3D holoscopic image by using spatial correlation of rendered view images
abstract
Holoscopic imaging is a prospective acquisition and display solution for providing natural and fatigue-free 3D visualization. However, large amount of data is required to represent the 3D holoscopic content. Therefore, efficient coding schemes for this particular type of image are needed. In this paper, an effective coding scheme is proposed by exploring the spatial correlation among the view images with different perspectives rendered from 3D holoscopic image. We utilize the interlaced view image to descript such spatial correlation. A linear prediction method is used on the interlaced view image instead of the original holoscopic image directly. Experimental results show that the proposed coding scheme performs better than HEVC intra standard and screen content coding extension of HEVC with around 2.41dB and 0.42 dB average quality improvement respectively.
Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen
ICASSP5
2017 Sparse Time-Varying Graphs for Slide Transition Detection in Lecture Videos
Zhijin Liu, Kai Li 0016, Liquan Shen, Ping An 0001
ICIG (1)3
2017 An efficient intra coding algorithm based on statistical learning for screen content coding
abstract
Screen content has different characteristics compared with natural content captured by cameras. To achieve more efficient compression, some new coding tools have been developed in the High Efficiency Video Coding (HEVC) Screen Content Coding (SCC) Extension, which also increase the computational complexity of encoder. In this paper, complexity analysis are first conducted to explore the distribution of complexities. Then, two classification trees, including early coding units (CU) partition tree (EPT) and CU content classification tree (CCT), are designed based on statistical characteristics and coding information. EPT is used to decide whether the CU skip the mode decision process of current depth level and CCT is used to classify the blocks into either natural blocks or screen blocks. Natural blocks will skip screen coding modes and screen blocks skip normal intra modes. Experimental results show the proposed algorithm can save 49% encoding time with 2.7% BD-rate increase on average for All Intra configuration under the SCC common test condition.
Hao Yang 0008, Liquan Shen, Ping An 0001
ICIP2
2017 CNN oriented fast QTBT partition algorithm for JVET intra coding
abstract
In this paper, a novel fast coding unit depth decision algorithm based on convolution neural network is presented for JVET future video coding. JVET employs quad-tree plus binary-tree (QTBT) block partitioning structure, which can support much more flexibility for coding units partition shapes, and improve the coding performance significantly than the HEVC standard. However, the flexible partitioning structure also introduces a tremendous computation complexity. To address this issue, we model the QTBT partition depth range as a multi-class classification problem, and try to predict the depth range of 32×32 block directly, rather than to judge split or not at each depth level. To the best of our knowledge, it is the first framework to formulate the QTBT partition range as a multi classification task, and optimized by an end-to-end learning model. For training optimization, we design an objective function consists of class penalty term and L2 HingeLoss function, which leverage the characteristics of category settings, can further boost the classification accuracy. Experimental results demonstrate the effectiveness of our proposed method, which can achieve 42.80% complexity reduction with only 0.65% Bjontegaard Delta bitrate (BD-rate) increase.
Zhipeng Jin, Ping An 0001, Liquan Shen, Chao Yang 0021
VCIP3
2017 SSIM-based binocular perceptual model for quality assessment of stereoscopic images
abstract
In this paper, we propose a novel full reference stereoscopic image quality assessment (FR-SIQA) metric by utilizing SSIM-based binocular perceptual model. The goal is to predict the perceptual quality of a stereoscopic image via jointly considering the qualities of cyclopean image and the difference image. Specifically, we first apply the contrast sensitivity filtering to both the reference and distorted stereo pairs. Constructively, a new cyclopean image is generated by considering binocular perceptual model and binocular rivalry simultaneously. Finally, the overall quality score of a testing stereoscopic image is predicted by combining the qualities of its cyclopean image and difference image. Experimental results show that the proposed metric achieves high consistency with human subjective assessment and outperforms several the state-of-the-art FR-SIQA methods.
Jian Ma 0005, Ping An 0001, Liquan Shen, Kai Li 0016, Jialu Yang
VCIP3
2017 Bivariate statistics and binocular energy induced stereo-pair quality evaluator
abstract
With the flourishment of 3D content, stereoscopic image quality assessment (SIQA) becomes an urgent issue in image processing field. In this paper, a new blind SIQA method is proposed based on bivariate natural scene statistics (NSS) model that is conducted on the binocular Gabor energy response. Specifically, Gabor responses of two views are combined using a binocular energy model. Then bivariate statistics of the jointly spatially adjacent responses of the fused Gabor energy are calculated for feature extraction. Experimental results on LIVE 3D Image Quality Database demonstrate the promising performance of the proposed method.
Liquan Shen, Ping An 0001
VCIP2
2017 A stereoscopic image quality assessment model based on independent component analysis and binocular fusion property
Xianqiu Geng, Liquan Shen, Kai Li 0016, Ping An 0001
Signal Process. Image Commun.2
2017 Bit allocation for 3D video coding based on lagrangian multiplier adjustment
Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen, Kai Li 0016
Signal Process. Image Commun.4
2017 Saliency Detection for Unconstrained Videos Using Superpixel-Level Graph and Spatiotemporal Propagation
abstract
This paper proposes an effective spatiotemporal saliency model for unconstrained videos with complicated motion and complex scenes. First, superpixel-level motion and color histograms as well as global motion histogram are extracted as the features for saliency measurement. Then a superpixel-level graph with the addition of a virtual background node representing the global motion is constructed, and an iterative motion saliency (MS) measurement method that utilizes the shortest path algorithm on the graph is exploited to reasonably generate MS maps. Temporal propagation of saliency in both forward and backward directions is performed using efficient operations on inter-frame similarity matrices to obtain the integrated temporal saliency maps with the better coherence. Finally, spatial propagation of saliency both locally and globally is performed via the use of intra-frame similarity matrices to obtain the spatiotemporal saliency maps with the even better quality. The experimental results on two video data sets with various unconstrained videos demonstrate that the proposed model consistently outperforms the state-of-the-art spatiotemporal saliency models on saliency detection performance.
Zhi Liu 0003, Linwei Ye, Guangling Sun, Liquan Shen
IEEE Trans. Circuits Syst. Video Technol.5
2017 Salient Object Segmentation via Effective Integration of Saliency and Objectness
abstract
This paper proposes an effective salient object segmentation method via the graph-based integration of saliency and objectness. Based on the superpixel segmentation result of the input image, a graph is built to represent superpixels using regular vertex, background seed vertex with the addition of a terminal vertex. The edge weights on the graph are defined by integrating the difference of appearance, saliency, and objectness between superpixels. Then, the object probability of each superpixel is measured by finding the shortest path from the corresponding vertex to the terminal vertex on the graph, and the resultant object probability map can generally better highlight salient objects and suppress background regions compared to both saliency map and objectness map. Finally, the object probability map is used to initialize salient object and background, and effectively incorporated into the framework of graph cut to obtain the final salient object segmentation result. Extensive experimental results on three public benchmark datasets show that the proposed method consistently improves the salient object segmentation performance and outperforms the state-of-the-art salient object segmentation methods. Furthermore, experimental results also demonstrate that the proposed graph-based integration method is more effective than other fusion schemes and robust to saliency maps generated using various saliency models.
Linwei Ye, Zhi Liu 0003, Liquan Shen, Cong Bai, Yang Wang 0003
IEEE Trans. Multim.4
2016 Depth map coding based on virtual view quality
abstract
Multi-view video plus depth (MVD) is a 3D video representation. In MVD, the depth map provides the scene distance information and is used to render the virtual view through Depth Image Based Rendering (DIBR) technique. The depth map coding error will induce distortion in the rendered virtual views. This paper proposes a mathematic model that can estimate the synthesized virtual view distortion induced by depth map compression, and the model is employed to the rate distortion optimization (RDO) in the depth map coding. Based on the rendered virtual view quality, a Lagrangian optimization adjustment scheme at Coding Unit (CU) level is proposed to improve the depth map encoding efficiency. Experimental results demonstrate that the proposed method can improve the BD-PSNR of virtual view for 0.62 dB, and the encoding complexity reduces compared with the view synthesis optimization (VSO) technique in the 3D-HEVC Test Model (HTM).
Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen
ICASSP4
2016 Using independent component analysis and binocular combination for stereoscopic image quality assessment
abstract
In this paper, a full reference stereoscopic image quality assessment (FR-SIQA) method is proposed based on independent component analysis (ICA) and binocular combination. Image features that reflect the responds of simple cells in the cortex are extracted by ICA-based algorithm. Both image feature similarity (IFS) and local luminance consistency (LLC) are calculated to measure the structure and brightness distortions, respectively. To simulate the binocular fusion properties, the energy of image features and the global relative luminance information are selected as the basic of binocular combination to fuse the right-left IFS and LLC into a final index. Experimental results demonstrate that the proposed algorithm achieves high consistency with subjective assessment on two public available 3D image quality assessment databases.
Xianqiu Geng, Liquan Shen, Ping An 0001, Zhi Liu 0003
VCIP2
2016 Just noticeable disparity difference model for 3D displays
abstract
Based on the related psychological and physiological advancement, a just noticeable disparity difference (JNDiD) model for 3D displays under the assumption that eyes converge at the virtual object is presented in this paper. Specifically, according to whether the surface of the virtual object is perceived to be blurred, the perceived depth space is divided into three regions. Furthermore, a three-phase linear function considering the accommodation convergence mismatch in different depth regions is employed to build the JNDiD model. Experimental results demonstrate the effectiveness and superiority of our method compared with the state-of-the-art just noticeable depth difference (JNDD) model for 3D displays.
Ping An 0001, Liquan Shen, Kai Li 0016, Nina Feng
VCIP3
2016 Parallax-aware local alignment for image stitching under large occlusion/disocclusion
abstract
This paper presents a parallax-aware image stitching approach under large occlusion/disocclusion. Different from previous research, we explore the image stitching issue in a parallax-aware perspective via local alignment. Specifically, we first label each feature point with a probability of being large parallax by developing a graph-based optimization framework. Afterwards, an integer programming model is built to pick out a group of feature matches free from parallax in a local region. Finally, by enforcing a stitching seam passing through such a locally aligned area, we are able to generate a high-quality stitching result under large parallax. Experimental results demonstrate the effectiveness and superiority of the proposed parallax-aware approach.
Yangxin Wang, Kai Li 0016, Ping An 0001, Liquan Shen, Xuemei Zou
VCIP4
2016 3D holoscopic image coding scheme using HEVC with Gaussian process regression
Deyang Liu, Ping An 0001, Chao Yang 0021, Liquan Shen
Signal Process. Image Commun.5
2016 Fast depth map coding based on virtual view quality
Chao Yang 0021, Ping An 0001, Liquan Shen, Nina Feng
Signal Process. Image Commun.3
2016 Saliency Detection Via Similar Image Retrieval
abstract
This letter proposes a novel saliency detection framework by propagating saliency of similar images retrieved from large and diverse Internet image collections to boost saliency detection performance effectively. For the input image, a group of similar images is retrieved based on the saliency weighted color histograms and the Gist descriptor from Internet image collections. Then, a pixel-level correspondence process between images is performed to guide the saliency propagation from the retrieved images. Both initial saliency map and correspondence saliency map are exploited to select the training samples by using the graph cut-based segmentation. Finally, the training samples are input into a set of weak classifiers to learn the boosted classifier for generating the boosted saliency map, which is integrated with the initial saliency map to generate the final saliency map. Experimental results on two public image datasets demonstrate that the proposed model can achieve the better saliency detection performance than the state-of-the-art single-image saliency models and co-saliency models.
Linwei Ye, Zhi Liu 0003, Xiaofei Zhou 0003, Liquan Shen, Jian Zhang 0002
IEEE Signal Process. Lett.4
2015 Virtual view distortion estimation for depth map coding
abstract
Multi-view video plus depth (MVD) format is a three-dimensional (3D) video representation. The depth map in MVD provides the scene geometry information and is used to render the virtual view through Depth Image Based Rendering (DIBR). In this paper, a virtual view distortion estimation function based on the characteristics of both texture image and depth map is proposed which can estimate virtual view distortion induced by depth map compression accurately, and the function is implemented to the Rate Distortion Optimization (RDO) in the depth map coding. Compared with the View Synthesis Optimization (VSO) in 3D-HEVC Test Model (HTM) reference software, the experimental results demonstrate that the proposed method can improve the BD-PSNR of virtual view for 0.26 dB on average, and the encoding time has reduced for 31% on average due to the low complexity of the proposed function.
Chao Yang 0021, Ping An 0001, Deyang Liu, Liquan Shen
VCIP4
2015 On the power allocation for hybrid DF and CF protocol with auxiliary parameter in fading relay channels
abstract
In fading channels, power allocation over channel state may bring a rate increment compared to the fixed constant power mode. Such a rate increment is referred to power allocation gain. It is expected that the power allocation gain varies for different relay protocols. In this paper, Decode-and-Forward (DF) and Compress-and-Forward (CF) protocols are considered. We first establish a general framework for relay power allocation of DF and CF over channel state in half-duplex relay channels and present the optimal solution for relay power allocation with auxiliary parameters, respectively. Then, we reconsider the power allocation problem for one hybrid scheme which always selects the better one between DF and CF and obtain a near optimal solution for the hybrid scheme by introducing an auxiliary rate function as well as avoiding the non-concave rate optimization problem. Simulation results show that the developed power allocation solutions bring significant rate gains in various fading relay channels compared to constant power allocation mode.
Zhengchuan Chen, Pingyi Fan, Dapeng Oliver Wu, Liquan Shen
WCNC4
2015 Spatiotemporal saliency detection based on superpixel-level trajectory
Zhi Liu 0003, Xiang Zhang 0006, Olivier Le Meur, Liquan Shen
Signal Process. Image Commun.5
2015 Fast TU size decision algorithm for HEVC encoders using Bayesian theorem detection
Liquan Shen, Zhaoyang Zhang 0002, Xinpeng Zhang 0001, Ping An 0001, Zhi Liu 0003
Signal Process. Image Commun.1
2015 Co-Saliency Detection via Co-Salient Object Discovery and Recovery
abstract
This letter proposes a novel co-saliency model to effectively discover and highlight co-salient objects in a set of images. Based on the gross similarity which combines color features and SIFT descriptors, some co-salient object regions are first discovered in each image as exemplars, which are exploited to generate the exemplar saliency maps with the use of single-image saliency model. Then both local recovery and global recovery of co-salient object regions are performed by propagating the exemplar saliency to the matched regions, and border connectivity is further exploited to generate the region-level co-saliency maps. Finally, the foci of attention area based pixel-level saliency derivation is used to generate the pixel-level co-saliency maps with even better quality. Experimental results on two benchmark datasets demonstrate that the proposed co-saliency model outperforms the state-of-the-art co-saliency models.
Linwei Ye, Zhi Liu 0003, Wanlei Zhao, Liquan Shen
IEEE Signal Process. Lett.5
2015 A 3D-HEVC Fast Mode Decision Algorithm for Real-Time Applications
abstract
3D High Efficiency Video Coding (3D-HEVC) is an extension of the HEVC standard for coding of multiview videos and depth maps. It inherits the same quadtree coding structure as HEVC for both components, which allows recursively splitting into four equal-sized coding units (CU). One of 11 different prediction modes is chosen to code a CU in inter-frames. Similar to the joint model of H.264/AVC, the mode decision process in HM (reference software of HEVC) is performed using all the possible depth levels and prediction modes to find the one with the least rate distortion cost using a Lagrange multiplier. Furthermore, both motion estimation and disparity estimation need to be performed in the encoding process of 3D-HEVC. Those tools achieve high coding efficiency, but lead to a significant computational complexity. In this article, we propose a fast mode decision algorithm for 3D-HEVC. Since multiview videos and their associated depth maps represent the same scene, at the same time instant, their prediction modes are closely linked. Furthermore, the prediction information of a CU at the depth level X is strongly related to that of its parent CU at the depth level X-1 in the quadtree coding structure of HEVC since two corresponding CUs from two neighboring depth levels share similar video characteristics. The proposed algorithm jointly exploits the inter-view coding mode correlation, the inter-component (texture-depth) correlation and the inter-level correlation in the quadtree structure of 3D-HEVC. Experimental results show that our algorithm saves 66% encoder runtime on average with only a 0.2% BD-Rate increase on coded views and 1.3% BD-Rate increase on synthesized views.
Liquan Shen, Ping An 0001, Zhaoyang Zhang 0002, Qianqian Hu, Zhengchuan Chen
ACM Trans. Multim. Comput. Commun. Appl.1
2014 Efficient depth coding in 3D video to minimize coding bitrate and complexity
Liquan Shen, Zhaoyang Zhang 0002
Multim. Tools Appl.1
2014 Compression of encrypted images with multi-layer decomposition
Xinpeng Zhang 0001, Guangling Sun, Liquan Shen, Chuan Qin 0001
Multim. Tools Appl.3
2014 Adaptive Inter-Mode Decision for HEVC Jointly Utilizing Inter-Level and Spatiotemporal Correlations
abstract
High Efficiency Video Coding (HEVC) adopts the quadtree structured coding unit (CU), which allows recursive splitting into four equally sized blocks. At each depth level, it enables SKIP mode, merge mode, inter 2N × 2N, inter 2N × N, inter N × 2N, inter 2N × nU, inter 2N × nD, inter nL x 2N, inter nR × 2N, inter N × N (only available for the smallest CU), intra 2N × 2N, and intra N × N (only available for the smallest CU) in inter-frames. Similar to H.264/AVC, the mode decision process in HEVC is performed using all the possible depth levels (or CU sizes) and prediction modes to find the one with the least rate distortion (RD) cost using Lagrange multiplier. This achieves the highest coding efficiency, but leads to a very high computational complexity. Since the optimal prediction mode is highly content dependent, it is not efficient to use all the modes. In this paper, we propose a fast inter-mode decision algorithm for HEVC by jointly using the inter-level correlation of quadtree structure and the spatiotemporal correlation. There exist strong correlations of the prediction mode, the motion vector and RD cost between different depth levels and between spatially temporally adjacent CUs. We statistically analyze the prediction mode distribution at each depth level and the coding information correlation among the adjacent CUs. Based on the analysis results, three adaptive inter-mode decision strategies are proposed including early SKIP mode decision, prediction size correlation-based mode decision and RD cost correlation-based mode decision. Experimental results show that the proposed overall algorithm can save 49%-52% computational complexity on average with negligible loss of coding efficiency, exhibiting applicability to various types of video sequences.
Liquan Shen, Zhaoyang Zhang 0002, Zhi Liu 0003
IEEE Trans. Circuits Syst. Video Technol.1
2014 Effective CU Size Decision for HEVC Intracoding
abstract
In high efficiency video coding (HEVC), the tree structured coding unit (CU) is adopted to allow recursive splitting into four equally sized blocks. At each depth level (or CU size), it enables up to 35 intraprediction modes, including a planar mode, a dc mode, and 33 directional modes. The intraprediction via exhaustive mode search exploited in the test model of HEVC (HM) effectively improves coding efficiency, but results in a very high computational complexity. In this paper, a fast CU size decision algorithm for HEVC intracoding is proposed to speed up the process by reducing the number of candidate CU sizes required to be checked for each treeblock. The novelty of the proposed algorithm lies in the following two aspects: 1) an early determination of CU size decision with adaptive thresholds is developed based on the texture homogeneity and 2) a novel bypass strategy for intraprediction on large CU size is proposed based on the combination of texture property and coding information from neighboring coded CUs. Experimental results show that the proposed effective CU size decision algorithm achieves a computational complexity reduction up to 67%, while incurring only 0.06-dB loss on peak signal-to-noise ratio or 1.08% increase on bit rate compared with that of the original coding in HM.
Liquan Shen, Zhaoyang Zhang 0002, Zhi Liu 0003
IEEE Trans. Image Process.1
2014 Compressing Encrypted Images With Auxiliary Information
abstract
This paper proposes a novel scheme of compressing encrypted images with auxiliary information. The content owner encrypts the original uncompressed images and also generates some auxiliary information, which will be used for data compression and image reconstruction. Then, the channel provider who cannot access the original content may compress the encrypted data by a quantization method with optimal parameters that are derived from a part of auxiliary information and a compression ratio-distortion criteria, and transmit the compressed data, which include an encrypted sub-image, the quantized data, the quantization parameters and another part of auxiliary information. At receiver side, the principal image content can be reconstructed using the compressed encrypted data and the secret key. Experimental result shows the ratio-distortion performance of the proposed scheme is better than that of previous techniques.
Xinpeng Zhang 0001, Yanli Ren, Liquan Shen, Zhenxing Qian, Guorui Feng
IEEE Trans. Multim.3
2013 A novel region merging based image segmentation approach for automatic object extraction
abstract
This paper presents a novel region merging based automatic image segmentation approach, which is applicable for object extraction. From an initial over-segmentation result, we exploit the regional histogram based similarity measure as merging criterion and merging order determination scheme with three priorities, to efficiently perform region merging, which is recorded using a binary partition tree (BPT). Based on the analysis of BPT, an appropriate subset of BPT nodes is selected to represent a meaningful image segmentation result and object extraction result. Experimental results demonstrate the better segmentation performance of our approach.
Lin Zha, Zhi Liu 0003, Shuhua Luo, Liquan Shen
ISCAS4
2013 Stretchability-aware block scaling for image retargeting
Huan Du, Zhi Liu 0003, Jianliang Jiang, Liquan Shen
J. Vis. Commun. Image Represent.4
2013 A novel H.264 rate control algorithm with consideration of visual attention
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002
Multim. Tools Appl.1
2013 An Effective CU Size Decision Method for HEVC Encoders
abstract
The emerging high efficiency video coding standard (HEVC) adopts the quadtree-structured coding unit (CU). Each CU allows recursive splitting into four equal sub-CUs. At each depth level (CU size), the test model of HEVC (HM) performs motion estimation (ME) with different sizes including 2N × 2N, 2N × N, N × 2N and N × N. ME process in HM is performed using all the possible depth levels and prediction modes to find the one with the least rate distortion (RD) cost using Lagrange multiplier. This achieves the highest coding efficiency but requires a very high computational complexity. In this paper, we propose a fast CU size decision algorithm for HM. Since the optimal depth level is highly content-dependent, it is not efficient to use all levels. We can determine CU depth range (including the minimum depth level and the maximum depth level) and skip some specific depth levels rarely used in the previous frame and neighboring CUs. Besides, the proposed algorithm also introduces early termination methods based on motion homogeneity checking, RD cost checking and SKIP mode checking to skip ME on unnecessary CU sizes. Experimental results demonstrate that the proposed algorithm can significantly reduce computational complexity while maintaining almost the same RD performance as the original HEVC encoder.
Liquan Shen, Zhi Liu 0003, Xinpeng Zhang 0001, Zhaoyang Zhang 0002
IEEE Trans. Multim.1
2012 Region Diversity Maximization for Salient Object Detection
abstract
Salient object detection is an important technique for many content-based applications, but it becomes a challenging work when handling the cluttered saliency maps, which cannot completely highlight salient object regions and cannot suppress background regions. In this letter, we propose a novel approach to detect salient object from saliency map without manually setting any parameters. Region diversity maximization is used as the objective function to direct the object detection, and the optimal window for locating the salient object is obtained using an efficient iterative search scheme. Experimental results on different saliency maps demonstrate the overall better detection performance and computational efficiency of our approach.
Zhi Liu 0003, Huan Du, Xiang Zhang 0006, Liquan Shen
IEEE Signal Process. Lett.5
2012 Content-Adaptive Motion Estimation Algorithm for Coarse-Grain SVC
abstract
A joint model of scalable video coding (SVC) uses exhaustive mode and motion searches to select the best prediction mode and motion vector for each macroblock (MB) with high coding efficiency at the cost of computational complexity. If major characteristics of a coding MB such as the complexity of the prediction mode and the motion property can be identified and used in adjusting motion estimation (ME), one can design an algorithm that can adapt coding parameters to the video content. This way, unnecessary mode and motion searches can be avoided. In this paper, we propose a content-adaptive ME for SVC, including analyses of mode complexity and motion property to assist mode and motion searches. An experimental analysis is performed to study interlayer and spatial correlations in the coding information. Based on the correlations, the motion and mode characteristics of the current MB are identified and utilized to adjust each step of ME at the enhancement layer including mode decision, search-range selection, and prediction direction selection. Experimental results show that the proposed algorithm can significantly reduce the computational complexity of SVC while maintaining nearly the same rate distortion performance as the original encoder.
Liquan Shen, Zhaoyang Zhang 0002
IEEE Trans. Image Process.1
2012 Unsupervised Salient Object Segmentation Based on Kernel Density Estimation and Two-Phase Graph Cut
abstract
In this paper, we propose an unsupervised salient object segmentation approach based on kernel density estimation (KDE) and two-phase graph cut. A set of KDE models are first constructed based on the pre-segmentation result of the input image, and then for each pixel, a set of likelihoods to fit all KDE models are calculated accordingly. The color saliency and spatial saliency of each KDE model are then evaluated based on its color distinctiveness and spatial distribution, and the pixel-wise saliency map is generated by integrating likelihood measures of pixels and saliency measures of KDE models. In the first phase of salient object segmentation, the saliency map based graph cut is exploited to obtain an initial segmentation result. In the second phase, the segmentation is further refined based on an iterative seed adjustment method, which efficiently utilizes the information of minimum cut generated using the KDE model based graph cut, and exploits a balancing weight update scheme for convergence of segmentation refinement. Experimental results on a dataset containing 1000 test images with ground truths demonstrate the better segmentation performance of our approach.
Zhi Liu 0003, Liquan Shen, Yinzhu Xue, King Ngi Ngan, Zhaoyang Zhang 0002
IEEE Trans. Multim.3
2011 An improved depth map estimation for coding and view synthesis
abstract
Inaccuracy depth estimation may influence on depth coding and virtual view rendering in the free-viewpoint television (FTV) system, an improved depth map estimation is proposed to solve the problem for coding and view synthesis. Firstly, check the consistency of initial depth, and the influence of initial miss-matches is minimized by introduction of an additional adaptive matching error selection that penalizes the unreliable matches. Then according to certain criteria, the multi-reference depth maps are merged into one disparity map to improve the quality of disparity map. Finally, a multilateral filtering is used to preserve details in the depth map and simultaneously smooth the depths in occluded areas at object boundary, less texture and discontinuity regions. Experimental results show a significant improvement of the initial input depth maps and coding efficiency, as well as a reduction of view synthesis artifacts.
Qiuwen Zhang, Ping An 0001, Liquan Shen, Zhaoyang Zhang 0002
ICIP4
2011 Unsupervised image segmentation based on analysis of binary partition tree for salient object extraction
Zhi Liu 0003, Liquan Shen, Zhaoyang Zhang 0002
Signal Process.2
2011 Low-Complexity Mode Decision for MVC
abstract
The finalized international standard for multiview video coding (MVC) is an extension of H.264. In the joint model of MVC, variable size motion estimation (ME) and disparity estimation (DE) are introduced to achieve the highest coding efficiency with the cost of very high computational complexity. A low complexity mode decision algorithm is proposed to reduce complexity of ME and DE. An experimental analysis is performed to study inter-view correlation in the coding information such as the prediction mode and rate-distortion (RD) cost. Based on the correlation, we propose four efficient mode decision techniques, including early SKIP mode decision, adaptive early termination, fast mode size decision, and selective intra coding in inter frame. Experimental results show that the proposed algorithm can significantly reduce computational complexity of MVC while maintaining almost the same RD performance.
Liquan Shen, Zhi Liu 0003, Ping An 0001, Zhaoyang Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2010 Nonparametric saliency detection using kernel density estimation
abstract
This paper proposes a nonparametric saliency model based on kernel density estimation (KDE) mainly aiming at content-based applications such as salient object segmentation. A set of KDE models are constructed on the basis of regions segmented using the mean shift algorithm. For each pixel, a set of color likelihood measures to all KDE models are calculated, and then the color saliency and spatial saliency of each KDE model are evaluated based on its color distinctiveness and spatial distribution. The final saliency map is generated by combining saliency measures of KDE models and color likelihood measures of pixels. Experimental results demonstrate the better saliency detection performance of our saliency model.
Zhi Liu 0003, Yinzhu Xue, Liquan Shen, Zhaoyang Zhang 0002
ICIP3
2010 An adaptive early termination of mode decision using inter-layer correlation in scalable video coding
abstract
The scalable video coding (SVC) standard adopts the variable size motion estimation (ME) to select the best coding mode for each macroblock (MB). Although this technique achieves the highest possible coding efficiency, it results in extremely large computation complexity which obstructs SVC from the practical application. In this paper, we propose an adaptive early termination of fast mode decision algorithm in SVC. It makes use of the coding information of spatial neighbor MBs and the corresponding MBs in base layer to early terminate the mode decision procedure. Experimental results show that the proposed fast mode decision algorithm can achieve computational saving up to 67% with no significant loss of rate distortion (RD) performance.
Liquan Shen, Zhi Liu 0003, Ping An 0001, Zhaoyang Zhang 0002
ICIP1
2010 Unsupervised salient object segmentation from color images
abstract
This paper proposes an efficient approach for unsupervised segmentation of salient objects from color images. A set of Gaussian models are first estimated based on a pre-segmentation result of the input image, and then for each pixel, a set of normalized color likelihood measures to each Gaussian model are calculated. The color saliency and spatial saliency of Gaussian models are exploited to generate the pixel-wise saliency map. By thresholding the saliency map, the pixels are classified into object seed pixels, background seed pixels and uncertain pixels to obtain the trimap. For each pixel, the probability belonging to salient object/background is evaluated using kernel density estimation, and the geodesic distances to salient object and background are calculated based on the object likelihood map. By comparing the two geodesic distances, uncertain pixels are finally classified into salient object or background. Experimental results demonstrate the better segmentation performance of the proposed approach.
Zhi Liu 0003, Liquan Shen, Zhaoyang Zhang 0002
VCIP3
2010 Rate control algorithm based on frame complexity estimation for MVC
abstract
Rate control has not been well studied for multi-view video coding (MVC). In this paper, we propose an efficient rate control algorithm for MVC by improving the quadratic rate-distortion (R-D) model, which reasonably allocate bit-rate among views based on correlation analysis. The proposed algorithm consists of four levels for rate bits control more accurately, of which the frame layer allocates bits according to frame complexity and temporal activity. Extensive experiments show that the proposed algorithm can efficiently implement bit allocation and rate control according to coding parameters.
Tao Yan 0003, Ping An 0001, Liquan Shen, Zhaoyang Zhang 0002
VCIP3
2010 Automatic segmentation of focused objects from images with low depth of field
Zhi Liu 0003, Liquan Shen, Zhongmin Han, Zhaoyang Zhang 0002
Pattern Recognit. Lett.3
2010 Early SKIP mode decision for MVC using inter-view correlation
Liquan Shen, Zhi Liu 0003, Tao Yan 0003, Zhaoyang Zhang 0002, Ping An 0001
Signal Process. Image Commun.1
2010 Efficient SKIP Mode Detection for Coarse Grain Quality Scalable Video Coding
abstract
Scalable video coding (SVC) was recently standardized by the Joint Video Team as an extension of H.264. In SVC, a computationally expensive exhaustive mode decision is employed to select the best coding mode for each macroblock (MB), which achieves a high coding efficiency. In order to reduce computational complexity, we propose an efficient SKIP mode detection approach for coarse grain quality SVC. It makes use of the coding information of spatial neighboring MBs and the co-located MB in base layer to predict the SKIP mode MB and early terminate its mode decision procedure. Experimental results show that the proposed early SKIP mode decision approach can achieve the average computational saving about 54% with almost no loss of rate distortion (RD) performance in the enhancement layer.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002
IEEE Signal Process. Lett.1
2010 View-Adaptive Motion Estimation and Disparity Estimation for Low Complexity Multiview Video Coding
abstract
The emerging international standard for multiview video coding (MVC) is an extension of H.264/advanced video coding. In the joint mode of MVC, both motion estimation (ME) and disparity estimation (DE) are included in the encoding process. This achieves the highest coding efficiency but requires a very high computational complexity. In this letter, we propose a fast ME and DE algorithm that adaptively utilizes the inter-view correlation. The coding mode complexity and the motion homogeneity of a macroblock (MB) are first analyzed according to the coding modes and motion vectors from the corresponding MBs in the neighbor views, which are located by means of global disparity vector. According to the coding mode complexity and the motion homogeneity, the proposed algorithm adjusts the search strategies for different types of MBs in order to perform a precise search according to video content. Experimental results demonstrate that the proposed algorithm can save 85% computational complexity on average, with negligible loss of coding efficiency.
Liquan Shen, Zhi Liu 0003, Tao Yan 0003, Zhaoyang Zhang 0002, Ping An 0001
IEEE Trans. Circuits Syst. Video Technol.1
2009 Fast mode decision for multiview video coding
abstract
In the draft of multi-view coding (MVC), variable size motion estimation and disparity estimation are employed to select the best coding mode for each macroblock. These techniques achieve the highest possible coding efficiency, but they result in extremely large computation complexity which obstructs MVC from practical application. This paper proposes a fast mode size decision algorithm for MVC in inter-frame coding. It makes use of the mode distribution correlation between neighbor views to deduct the executions of unnecessary modes. Experimental results show that the proposed fast mode decision algorithm reduces the computational complexity significantly with negligible coding efficiency.
Liquan Shen, Tao Yan 0003, Zhi Liu 0003, Zhaoyang Zhang 0002, Ping An 0001
ICIP1
2009 Frame-level bit allocation based on incremental PID algorithm and frame complexity estimation
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi
J. Vis. Commun. Image Represent.1
2009 Selective VS-MRF-ME and intra coding in H.264 based on spatiotemporal continuity of motion field
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi
Signal Process. Image Commun.1
2009 An Efficient Intermode Decision Algorithm Based on Motion Homogeneity for H.264/AVC
abstract
The latest video coding standard H.264/AVC significantly outperforms previous standards in terms of coding efficiency. H.264/AVC adopts variable block sizes ranging from 4 times 4 to 16 times 16 in inter frame coding, and achieves significant gain in coding efficiency compared to coding a macroblock (MB) using regular block size. However, this new feature causes extremely high computation complexity when rate-distortion optimization (RDO) is performed using the scheme of full mode decision. This paper presents an efficient intermode decision algorithm based on motion homogeneity evaluated on a normalized motion vector (MV) field, which is generated using MVs from motion estimation on the block size of 4 times 4. Three directional motion homogeneity measures derived from the normalized MV field are exploited to determine a subset of candidate intermodes for each MB, and unnecessary RDO calculations on other intermodes can be skipped. Experimental results demonstrate that our algorithm can reduce the entire encoding time about 40% on average, without any noticeable loss of coding efficiency.
Zhi Liu 0003, Liquan Shen, Zhaoyang Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.2
2008 Fast Inter Mode Decision Using Spatial Property of Motion Field
abstract
Variable size motion estimation with multiple reference frames has been adopted by the new video coding standard H.264. It can achieve significant coding efficiency compared to coding a macroblock (MB) in regular size with single reference frame. On the other hand, it causes high computational complexity of motion estimation at the encoder. Rate distortion optimized (RDO) decision is one powerful method to choose the best coding mode among all combinations of block sizes and reference frames, but it requires extremely high computation. In this paper, a fast inter mode decision is proposed to decide best prediction mode utilizing the spatial continuity of motion field, which is generated by motion vectors from 4times4 motion estimation. Motion continuity of each MB is decided based on the motion edge map detected by the Sobel operator. Based on the motion continuity of a MB, only a small number of block sizes are selected in motion estimation and RDO computation process. Simulation results show that our algorithm can save more than 50% computational complexity, with negligible loss of coding efficiency.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi
IEEE Trans. Multim.1
2007 A novel algorithm to fast mode decision with consideration about video texture in H.264
abstract
H.264 employs 7 different size block types for motion estimation that can significantly improve the coding performance compared with the previous video coding standards. However, H.264 requires extremely high computation with the R-D optimized decision since so many prediction modes are used. In this paper, a novel inter mode decision algorithm (NIMDA) is proposed that utilizes SADs of each 4X4 block and texture characteristic to reduce the candidate mode set after the 16 X16 prediction mode is tested. The simulation results show that the proposed algorithm reduces the entire encoding time by 64.52% with only negligible coding loss.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002
AICCSA1
2007 Video nature considerations for multi-frame selection algorithm in H.264
abstract
H.264 allows motion estimation performing on multiple reference frames. This new feature improves the prediction accuracy of inter-coding blocks significantly. However, the coding gain comes at the cost of a much higher computational complexity. The reference software JM adopts full search scheme, and the computational complexity of motion estimation increases linearly with the number of allowed reference frames. In fact, the reduction of prediction residues is highly dependent on the nature of sequences, not on the number of searched frames. In this paper, with consideration of video nature and the available information from previous searched reference frames, an adaptive multi-frame selection algorithm (AMFSA) is proposed to speed up the matching process for multiple reference frames in the H.264 video coding system. The proposed algorithm can effectively reduce 63.6% on average.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002
AICCSA1
2007 A Novel Video Object Tracking Approach Based on Kernel Density Estimation and Markov Random Field
abstract
In this paper, we propose a novel video object tracking approach based on kernel density estimation and Markov random field (MRF). The interested video objects are first segmented by the user, and a nonparametric model based on kernel density estimation is initialized for each video object and the remaining background, respectively. A temporal saliency map is also initialized for each object to memorize the temporal trajectory. Based on the probabilities evaluated on the non-parametric models, each pixel in the current frame is first classified into the corresponding video object or background using the maximum likelihood criterion. Starting from the initial classification result, a MRF model that combines spatial smoothness and temporal coherency is selectively exploited to generate more reliable video objects. The nonparametric model and the temporal saliency map for each video object are updated and propagated for the future tracking. Experimental results on several MPEG-4 test sequences demonstrate the good segmentation performance of our approach.
Zhi Liu 0003, Liquan Shen, Zhongmin Han, Zhaoyang Zhang 0002
ICIP (3)2
2007 An Adaptive and Fast H.264 Multi-Frame Selection Algorithm Based on Information from Previous Searches
abstract
The H.264 video coding standard adopts multiple reference frames for motion estimation. This new feature improves the prediction accuracy of inter-coding blocks significantly, but it results in a considerable increase in encoder complexity, mainly regarding to multi-frame selection and motion estimation. The reference software JM adopts the full search scheme, and the increased computation is in proportion to the number of searched reference frames. However, the reduction of prediction residues is highly dependent on the nature of sequences, not on the number of searched frames. In this paper, we propose an adaptive and fast multi-frame selection algorithm (AFMFS) based on motion vectors and SAD information coming from previous searches to adaptively terminate the procedure of multiple reference frames selection. Compared with the full search algorithm and the flexible multi-reference frame search criterion (FMRFSC), simulation results show that the proposed algorithm can save 55.28% and 40.23% computation cost on average, respectively, while it still maintains similar coding efficiency.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002, Xuli Shi
ICME1
2007 An adaptive and fast fractional pixel search algorithm in H.264
Liquan Shen, Zhaoyang Zhang 0002, Zhi Liu 0003, Wenjun Zhang 0001
Signal Process.1
2007 An Adaptive and Fast Multiframe Selection Algorithm for H.264 Video Coding
abstract
H.264 allows motion estimation performing on multiple reference frames. This new feature improves the prediction accuracy of inter-coding blocks significantly, but it is extremely computational intensive. The reference software JM adopts full search scheme, and the increased computation load is in proportion to the number of searched reference frames. However, the reduction of prediction residues is highly dependent on the content of sequences, not on the number of searched frames. In this letter, we propose an adaptive and fast multiframe selection algorithm (AFMFSA) to speed up the searching procedure for multiple reference frames. Simulation results show that the proposed algorithm can deduct 56.0%–74.2% computation load of motion estimation on average.
Liquan Shen, Zhi Liu 0003, Zhaoyang Zhang 0002
IEEE Signal Process. Lett.1