EDBT 2026 Demo / reviewers in the wild / expert
Xin Deng 0002
dblp:24/4856-2
· DBLP profile ↗
68ranked-venue papers
21as first author
52since 2021 · last 2026
0000-0002-4708-6572ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 17 first-author · 34 since 2021Artificial intelligence and machine learning · 26 · 8 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream TasksabstractIn recent years, the development of burst imaging technology has improved the capture and processing capabilities of visual data, enabling a wide range of applications. However, the redundancy in burst images leads to the increased storage and transmission demands, as well as reduced efficiency of downstream tasks. To address this, we propose a new task of Burst Image Quality Assessment (BuIQA), to evaluate the task-driven quality of each frame within a burst sequence, providing reasonable cues for burst image selection. Specifically, we establish the first benchmark dataset for BuIQA, consisting of 7,346 burst sequences with 45,827 images and 191,572 annotated quality scores for multiple downstream scenarios. Inspired by the data analysis, a unified BuIQA framework is proposed to achieve an efficient adaption for BuIQA under diverse downstream scenarios. Specifically, a task-driven prompt generation network is developed with heterogeneous knowledge distillation, to learn the priors of the downstream task. Then, the task-aware quality assessment network is introduced to assess the burst image quality based on the task prompt. Extensive experiments across 10 downstream scenarios demonstrate the impressive BuIQA performance of the proposed approach, outperforming the state-of-the-art. Furthermore, it can achieve 0.33 dB PSNR improvement in the downstream tasks of denoising and super-resolution, by applying our approach to select the high-quality burst frames. Xiaoye Liang, Lai Jiang 0004, Minglang Qiao, Yue Zhang 0082, Xin Deng 0002, Shengxi Li, Yufan Liu 0001, Mai Xu |
AAAI | 6 |
| 2026 | DASR+: Training Domain Distance Aware Network for Unsupervised Image Super-Resolution
Xiaorui Zhao, Yunxuan Wei, Xin Deng 0002, Yawei Li 0001, Radu Timofte, Hengjie Song, Shuhang Gu |
Int. J. Comput. Vis. | 3 |
| 2026 | Say the image: Auditory masking effect-driven invertible network for progressive image-in-audio steganography
Jinghang Song, Fangyuan Gao, Xin Deng 0002, Shengxi Li, Mai Xu |
J. Inf. Secur. Appl. | 3 |
| 2026 | AIRPNet: Adaptive Image Restoration With Privacy Protection in Steganographic DomainabstractCloud-based third-party multimedia services have become increasingly popular in last decade, however, they pose serious threats to users' privacy. To address this issue, in this paper, we propose a novel Adaptive Image Restoration network with Privacy protection, namely AIRPNet, which first attempts to perform image restoration in steganographic domain. Compared with existing methods, our method has significant advantages in invisibility, security and flexibility. Specifically, we first propose a wavelet lifting-based Adaptive Invertible Hiding (AIH) module to conceal the low-quality (LQ) secret image into a stego image. Then, instead of performing single type of restoration on the secret image, an adaptive secure restoration (ASR) module is developed to deal with multiple image degradations on the stego image. Finally, a high-quality (HQ) secret image can be extracted from the restored stego image. Here, since the secret image remains hidden throughout the whole image restoration process, the privacy of users can be greatly protected. The framework can be flexibly extended to multiple image restoration, which can restore multiple secret images from the same stego image. Experimental results on various datasets demonstrate that our AIRPNet outperforms existing methods in terms of restoration accuracy, invisibility and security on different image restoration tasks. Fangyuan Gao, Xin Deng 0002, Junjie Huang 0001, Mai Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | DeepELIC: Deep encrypted lossy image compression network via compressive sensing unfolding
Fangyuan Gao, Yufan Deng, Xin Deng 0002, Zhenyu Guan 0002, Mai Xu |
Pattern Recognit. | 3 |
| 2026 | Compressed image super-resolution based on invertible degradation and restoration
Mai Xu, Lai Jiang 0004, Xin Deng 0002, Yue Zhang 0082, Yufan Liu 0001 |
Pattern Recognit. | 4 |
| 2026 | A novel geometry-aware spatio-temporal network for multi-view video feature learning
Yue Zhang 0082, Mai Xu, Lai Jiang 0004, Xin Deng 0002, Si Liu 0001 |
Pattern Recognit. | 4 |
| 2026 | A Novel Visible-Infrared Image Compression Framework for High-Value Target ProtectionabstractThe joint compression of visible-infrared images is crucial for military and surveillance applications. The challenge lies in the protection of high-value targets (HVT) while maintaining high compression efficiency. This paper proposes a novel dual-stream compression framework that effectively addresses this challenge. In our framework, the sensitive HVT infrared signatures are concealed within the visible image stream, while residual infrared image is encoded separately. This dual-stream compression framework introduces three key innovations. 1) HVT protection: The HVT information is physically isolated and hidden within public visible images through a dedicated concealment stream; 2) Key-conditioned reconstruction: A novel decoding mechanism enables active camouflage by replacing HVTs with plausible background content when unauthorized access is detected; 3) Unified optimization: The framework integrates compression efficiency and HVT protection within an endto- end trainable network. Extensive experiments demonstrate that our approach achieves state-of-the-art compression performance while providing superior HVT protection, significantly outperforming traditional encrypt-then-compress methods. The code and weights are open-source athttps://github.com/eecoder-dyf/rgbir-compress. Yufan Deng, Xin Deng 0002, Shengxi Li, Xiaowan Hu, Mai Xu |
IEEE Signal Process. Lett. | 2 |
| 2026 | Blur-Resistant Hyperspectral Image Super-Resolution via Dual-Degradation Fusion ModelabstractThe deep unfolding network represents a promising research avenue in fusion-based hyperspectral image super-resolution (HSI-SR). However, most current deep unfolding methodologies are anchored in idealized observation models, which overlook the degradation of the multispectral image (MSI), hindering their SR performance and practical applicability. To address this problem, this paper establishes a novel Dual-Degradation Fusion (D2-Fusion) model, which incorporates both HSI degradation and MSI blurring into the HSI-SR modeling process. Subsequently, we apply the second-order semismooth Newton algorithm to solve the optimization problem in D2-Fusion model. The solution steps are then mapped into an end-to-end trainable network, termed Blur-resistant Hyperspectral image Super-Resolution Network (BHSR-Net). To the best of our knowledge, the proposed network is the first successful attempt to consider MSI blurring artifacts in the HSI-SR task. It offers several distinct advantages: 1) The network structure maintains a strict mathematical correspondence with the optimization algorithm, ensuring each module retains strong physical interpretability; 2) The network exhibits superior SR performance and strong generalization ability on both standard and real-world scenarios across five datasets; 3) The network demonstrates excellent learning efficiency with a compact architecture, and its lightweight variant achieves comparable results with only 38K parameters. The code is available at https://github.com/Dou0405/BHSR-Net. Mai Xu, Yongxuan Dou, Xin Deng 0002, Zhenwei Shi 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Spherical Manifold Guided Diffusion Model for Panoramic Image GenerationabstractPanoramic image essentially acts as a pivotal role in emerging virtual reality and augmented reality scenarios; however, the generation of panoramic images are essentially challenging due to the intrinsic spherical geometry and spherical distortions caused by equirectangular projection (ERP). To address this, we start from the very basics of S2manifold inherent to panoramic images, and propose a novel spherical manifold convolution (SMConv) on S2manifold. Based on the SMConv operation, we propose a spherical manifold guided diffusion (SMGD) model for text-conditioned panoramic image generation, which can well accommodate the spherical geometry during generation. We further develop a novel evaluation method by calculating grouped Fréchet inception distance (FID) on cube-map projections, which can well reflect the quality of generated panoramic images, compared to existing methods that randomly crop ERP-distorted content. Experiment results demonstrate that our SMGD model achieves the state-of-the-art generation quality and accuracy, whilst retaining the shortest sampling time in the text-conditioned panoramic image generation task. Codes are publicly available at https://github.com/chronos123/SMGD. Xiancheng Sun, Mai Xu, Shengxi Li, Senmao Ma, Xin Deng 0002, Lai Jiang 0004 |
CVPR | 5 |
| 2025 | Quality Control For HEVC: A Deep Reinforcement Learning ApproachabstractIn video coding, large quality fluctuations exist in compressed videos, significantly degrading their quality of experience (QoE). Most works in literature focus on controlling bit-rates, however, paying few attention on reducing the quality fluctuations. In this paper, we propose a novel deep reinforcement learning (DRL) method for quality control in video coding. Specifically, we first propose the formulation of quality control, which targets at both controlling the target quality and reducing fluctuations. Then, we solve the quality control formulation by proposing a DRL method, in which the DRL elements are modeled by considering the features of both current frame and previous encoded frames. Specifically, for the DRL elements, we take the encoding information, content complexity and hidden features of long short-term memory (LSTM) as the state of DRL, and the selection of quantization parameters (QP) as the action of DRL. Subsequently, an algorithm, based on proximal policy optimization, is utilized to update our DRL model for decision-making on the actions of QP selection. In this way, the videos can be compressed under given and constant quality. We implement our DRL-based quality control method on the standard of high efficiency video coding (HEVC) with the HM 16.15 platform, and experimental results show that our method achieves the state-of-the-art performance on both quality control accuracy and fluctuations, in comparison with other quality control baselines. Mai Xu, Lai Jiang 0004, Shengxi Li, Xin Deng 0002 |
ICME | 6 |
| 2025 | Spherical-Nested Diffusion Model for Panoramic Image OutpaintingabstractPanoramic image outpainting acts as a pivotal role in immersive content generation, allowing for seamless restoration and completion of panoramic content. Given the fact that the majority of generative outpainting solutions operates on planar images, existing methods for panoramic images address the sphere nature by soft regularisation during the end-to-end learning, which still fails to fully exploit the spherical content. In this paper, we set out the first attempt to impose the sphere nature in the design of diffusion model, such that the panoramic format is intrinsically ensured during the learning procedure, named as spherical-nested diffusion (SpND) model. This is achieved by employing spherical noise in the diffusion process to address the structural prior, together with a newly proposed spherical deformable convolution (SDC) module to intrinsically learn the panoramic knowledge. Upon this, the proposed method is effectively integrated into a pre-trained diffusion model, outperforming existing state-of-the-art methods for panoramic image outpainting. In particular, our SpND method reduces the FID values by more than 50\% against the state-of-the-art PanoDiffusion method. Codes are publicly available at \url{https://github.com/chronos123/SpND}. Xiancheng Sun, Senmao Ma, Shengxi Li, Mai Xu, Jingyuan Xia, Lai Jiang 0004, Xin Deng 0002 |
ICML | 7 |
| 2025 | DeepSN-Net: Deep Semi-Smooth Newton Driven Network for Blind Image RestorationabstractThe deep unfolding network represents a promising research avenue in image restoration. However, most current deep unfolding methodologies are anchored in first-order optimization algorithms, which suffer from sluggish convergence speed and unsatisfactory learning efficiency. In this paper, to address this issue, we first formulate an improved second-order semi-smooth Newton (ISN) algorithm, transforming the original nonlinear equations into an optimization problem amenable to network implementation. After that, we propose an innovative network architecture based on the ISN algorithm for blind image restoration, namely DeepSN-Net. To the best of our knowledge, DeepSN-Net is the first successful endeavor to design a second-order deep unfolding network for image restoration, which fills the blank of this area. Furthermore, it offers several distinct advantages: 1) DeepSN-Net provides a unified framework to a variety of image restoration tasks in both synthetic and real-world contexts, without imposing constraints on the degradation conditions. 2) The network architecture is meticulously aligned with the ISN algorithm, ensuring that each module possesses robust physical interpretability. 3) The network exhibits high learning efficiency, superior restoration accuracy and good generalization ability across 11 datasets on three typical restoration tasks. The success of DeepSN-Net on image restoration may ignite many subsequent works centered around the second-order optimization algorithms, which is good for the community. Xin Deng 0002, Lai Jiang 0004, Jingyuan Xia, Mai Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Spherical Patch Generative Adversarial Net for Unconditional Panoramic Image GenerationabstractRecent advancements in virtual reality (VR) and augmented reality (AR) have popularised the emerging panoramic content for the immersive visual experience. The difficulty in acquisition and display of 360° format further highlights the necessity of unconditional panoramic image generation. Existing methods essentially generate planar images mapped from panoramic images, and fail to address the deformation and closed-loop characteristics when inverted back to the panoramic images. Thus leading to the generation of pseudo-panoramic content. This paper aims to directly generate spherical content, in a patch-by-patch style; besides computation friendly, this promises the anywhere continuity on the panoramic image and proper accommodation of panoramic deformation. More specifically, we first propose a novel spherical patch convolution (SPConv) that operates on the local spherical patch, which naturally addresses the deformation of panoramic content. We then propose our spherical patch generative adversarial net (SP-GAN) that consists of spherical local embedding (SLE) and spherical content synthesiser (SCS) modules, which seamlessly incorporate our SPConv so as to generate continuous panoramic patches. To the best of our knowledge, the proposed SP-GAN is the first successful attempt to accommodate the spherical distortion for closed-loop panoramic image generation in a patch-by-patch manner. The experimental results, with human-rated evaluations, have verified the consistently superior performances for unconditional panoramic image generation, from the perspectives of generation quality, computational memory, and generalisation to various resolutions. Codes are publicly available at https://github.com/chronos123/SP-GAN. Mai Xu, Xiancheng Sun, Shengxi Li, Lai Jiang 0004, Jingyuan Xia, Xin Deng 0002 |
IEEE Trans. Image Process. | 6 |
| 2025 | Deep Semi-Smooth Newton-Driven Unfolding Network for Multi-Modal Image Super-ResolutionabstractDeep unfolding has emerged as a powerful solution for Multi-modal Image Super-Resolution (MISR) through strategic integration of cross-modal priors in network architecture. However, current deep unfolding approaches rely on first-order optimization, which exhibit limitations in learning efficiency and reconstruction accuracy. In this paper, to overcome these limitations, we propose a novel Semi-smooth Newton driven Unfolding network for MISR, namely SNUM-Net. Specifically, we first develop a Semi-smooth Newton-driven MISR (SNM) algorithm that establishes a theoretical foundation for our approach. Then, we unfold the iterative solution of SNM into a novel network. To the best of our knowledge, the SNUM-Net is the first successful attempt to design a deep unfolding MISR network based on second-order optimization algorithm. Compared to existing methods, the SNUM-Net demonstrates three main advantages. 1) Universal paradigm: the SNUM-Net provides a unified paradigm for diverse MISR tasks without requiring scenario-specific constraints; 2) Explainable framework: the network preserves a mathematical correspondence with the SNM algorithm, ensuring that the topological relationships between modules are well explainable; 3) Superior performance: comprehensive evaluations across 10 datasets spanning 3 MISR tasks demonstrate the network's exceptional reconstruction accuracy and generalization capability. The software codes are available at https://github.com/pandazcx/SNUM-Net. Xin Deng 0002, Yongxuan Dou, Mai Xu |
IEEE Trans. Image Process. | 2 |
| 2025 | Recruiting Teacher IF Modality for Nephropathy Diagnosis: A Customized Distillation Method With Attention-Based Diffusion NetworkabstractThe joint use of multiple modalities for medical image processing has been widely studied in recent years. The fusion of information from different modalities has demonstrated the performance improvement for a lot of medical tasks. For nephropathy diagnosis, immunofluorescence (IF) is one of the most widely-used multi-modality medical images due to its ease of acquisition and the effectiveness for certain nephropathy. However, the existing methods mainly assume different modalities have the equal effect on the diagnosis task, failing to exploit multi-modality knowledge in details. To avoid this disadvantage, this paper proposes a novel customized multi-teacher knowledge distillation framework to transfer knowledge from the trained single-modality teacher networks to a multi-modality student network. Specifically, a new attention-based diffusion network is developed for IF based diagnosis, considering global, local, and modality attention. Besides, a teacher recruitment module and diffusion-aware distillation loss are developed to learn to select the effective teacher networks based on the medical priors of the input IF sequence. The experimental results in the test and external datasets show that the proposed method has a better nephropathy diagnosis performance and generalizability, in comparison with the state-of-the-art methods. Mai Xu, Lai Jiang 0004, Yibing Fu, Xin Deng 0002, Shengxi Li |
IEEE Trans. Medical Imaging | 5 |
| 2025 | MDSC-Net: Multi-Modal Discriminative Sparse Coding Driven RGB-D Classification NetworkabstractIn this paper, we propose a novel sparsity-driven deep neural network to solve the RGB-D image classification problem. Different from existing classification networks, our network architecture is designed by drawing inspirations from a new proposed multi-modal discriminative sparse coding (MDSC) model. The key feature of this model is that it can gradually separate the discriminative and non-discriminative features in RGB-D images in a coarse-to-fine manner. Only the discriminative features are integrated and refined for classification, while the non-discriminative features are discarded, to improve the classification accuracy and efficiency. Derived from the MDSC model, the proposed network is composed of three modules, i.e., the shared feature extraction (SFE) module, discriminative feature refinement (DFR) module, and classification module. The architecture of each module is derived from the optimization solution in the MDSC model. To the best of our knowledge, this is the first time a fully sparsity-driven network has been proposed for RGB-D image classification. Extensive results verify the effectiveness of our method on different RGB-D image datasets. Xin Deng 0002, Yibing Fu, Mai Xu, Shengxi Li |
IEEE Trans. Multim. | 2 |
| 2024 | Enhancing Quality of Compressed Images by Mitigating Enhancement Bias Towards Compression DomainabstractExisting quality enhancement methods for compressed images focus on aligning the enhancement domain with the raw domain to yield realistic images. However, these methods exhibit a pervasive enhancement bias towards the compression domain, inadvertently regarding it as more realistic than the raw domain. This bias makes enhanced images closely resemble their compressed counterparts, thus degrading their perceptual quality. In this paper, we propose a simple yet effective method to mitigate this bias and enhance the quality of compressed images. Our method employs a conditional discriminator with the compressed image as a key condition, and then incorporates a domain-divergence regularization to actively distance the enhancement domain from the compression domain. Through this dual strategy, our method enables the discrimination against the compression domain, and brings the enhancement domain closer to the raw domain. Comprehensive quality evaluations confirm the superiority of our method over other state-of-the-art methods without incurring inference overheads. Qunliang Xing, Mai Xu, Shengxi Li, Xin Deng 0002, Meisong Zheng, Huaida Liu, Ying Chen 0011 |
CVPR | 4 |
| 2024 | SN-NET: Semismooth Newton Driven Lightweight Network for Real-World Image DenoisingabstractSemismooth Newton is a powerful tool to tackle the regularization problems in image restoration. Compared to other optimization methods such as alternating direction method of multipliers (ADMM), the semismooth Newton method exhibits faster convergence and greater efficiency, since the non-linear and coupling system is solved simultaneously. However, its performance relies heavily on the handcrafted parameters and efficient solvers for calculating Newton steps. To tackle this issue, we first develop an improved semismooth Newton method, in which we turn the original nonlinear system solving problem into a network-friendly convex optimization problem. After that, we unfold it into a novel network namely SN-Net. We apply SN-Net on the most fundamental image denoising task, which shows great advantages in the following two aspects. (1) The network is quite lightweight, i.e., the number of network parameters is only 86 KB. (2) The network structure exhibits strong interpretability. To the best of our knowledge, the $\mathrm{SN}-\mathrm{Net}$ is the first attempt to successfully map the semismooth Newton method to a learnable network. The great success of it on image denoising may inspire many potential works on other image restoration tasks. The code and the pre-trained models are released at https://github.com/pandazcx/SN-Net. Xin Deng 0002, Hongpeng Sun, Mai Xu |
ICIP | 2 |
| 2024 | Causal Context Adjustment Loss for Learned Image CompressionabstractIn recent years, learned image compression (LIC) technologies have surpassed conventional methods notably in terms of rate-distortion (RD) performance. Most present learned techniques are VAE-based with an autoregressive entropy model, which obviously promotes the RD performance by utilizing the decoded causal context. However, extant methods are highly dependent on the fixed hand-crafted causal context. The question of how to guide the auto-encoder to generate a more effective causal context benefit for the autoregressive entropy models is worth exploring. In this paper, we make the first attempt in investigating the way to explicitly adjust the causal context with our proposed Causal Context Adjustment loss (CCA-loss). By imposing the CCA-loss, we enable the neural network to spontaneously adjust important information into the early stage of the autoregressive entropy model. Furthermore, as transformer technology develops remarkably, variants of which have been adopted by many state-of-the-art (SOTA) LIC techniques. The existing computing devices have not adapted the calculation of the attention mechanism well, which leads to a burden on computation quantity and inference latency. To overcome it, we establish a convolutional neural network (CNN) image compression model and adopt the unevenly channel-wise grouped strategy for high efficiency. Ultimately, the proposed CNN-based LIC network trained with our Causal Context Adjustment loss attains a great trade-off between inference latency and rate-distortion performance. Minghao Han, Shiyin Jiang, Shengxi Li, Xin Deng 0002, Mai Xu, Ce Zhu, Shuhang Gu |
NeurIPS | 4 |
| 2024 | Joint Learning of Audio-Visual Saliency Prediction and Sound Source Localization on Multi-face Videos
Minglang Qiao, Yufan Liu 0001, Mai Xu, Xin Deng 0002, Bing Li 0001, Weiming Hu 0004, Ali Borji |
Int. J. Comput. Vis. | 4 |
| 2024 | CrossHomo: Cross-Modality and Cross-Resolution Homography EstimationabstractMulti-modal homography estimation aims to spatially align the images from different modalities, which is quite challenging since both the image content and resolution are variant across modalities. In this paper, we introduce a novel framework namely CrossHomo to tackle this challenging problem. Our framework is motivated by two interesting findings which demonstrate the mutual benefits between image super-resolution and homography estimation. Based on these findings, we design a flexible multi-level homography estimation network to align the multi-modal images in a coarse-to-fine manner. Each level is composed of a multi-modal image super-resolution (MISR) module to shrink the resolution gap between different modalities, followed by a multi-modal homography estimation (MHE) module to predict the homography matrix. To the best of our knowledge, CrossHomo is the first attempt to address the homography estimation problem with both modality and resolution discrepancy. Extensive experimental results show that our CrossHomo can achieve high registration accuracy on various multi-modal datasets with different resolution gaps. In addition, the network has high efficiency in terms of both model complexity and running speed. Xin Deng 0002, Enpeng Liu, Shengxi Li, Shuhang Gu, Mai Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Deep$\mathrm {M^{2}}$M2CDL: Deep Multi-Scale Multi-Modal Convolutional Dictionary Learning NetworkabstractFor multi-modal image processing, network interpretability is essential due to the complicated dependency across modalities. Recently, a promising research direction for interpretable network is to incorporate dictionary learning into deep learning through unfolding strategy. However, the existing multi-modal dictionary learning models are both single-layer and single-scale, which restricts the representation ability. In this paper, we first introduce a multi-scale multi-modal convolutional dictionary learning (M2CDL) model, which is performed in a multi-layer strategy, to associate different image modalities in a coarse-to-fine manner. Then, we propose a unified framework namely DeepM2CDL derived from the M2CDL model for both multi-modal image restoration (MIR) and multi-modal image fusion (MIF) tasks. The network architecture of DeepM2CDL fully matches the optimization steps of the M2CDL model, which makes each network module with good interpretability. Different from handcrafted priors, both the dictionary and sparse feature priors are learned through the network. The performance of the proposed DeepM2CDL is evaluated on a wide variety of MIR and MIF tasks, which shows the superiority of it over many state-of-the-art methods both quantitatively and qualitatively. In addition, we also visualize the multi-modal sparse features and dictionary filters learned from the network, which demonstrates the good interpretability of the DeepM2CDL network. Xin Deng 0002, Fangyuan Gao, Xiancheng Sun, Mai Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Extremely Low Bit-Rate Image Compression via Invertible Image GenerationabstractImage compression at extremely low bit-rates has always been a challenging task in bandwidth limited scenarios, such as aerospace and deep-sea explorations. Recent years have seen great success of deep learning in image compression, however, few of them are specially designed for extremely low bit-rate conditions. To solve this issue, in this paper, we propose a novel invertible image generation based framework for extremely low bit-rate image compression. The proposed framework is composed of three modules, including an invertible image generation (IIG) module, a generated image compression (GIC) module and a compressed image adjustment (CIA) module. The role of IIG module is to generate a compression-friendly image from the original image. In the IIG module, image generation and restoration are modelled as two mutually reversible processes to avoid the information loss. After the IIG module, the GIC module is employed to compress the generated images to save the coding bit-rates. After that, the CIA module is used to shrink the quality gap between the compressed generated image and the un-compressed image. Finally, the image from the CIA module is sent back to the IIG module to restore the original image. The experimental results on three different datasets show that the proposed framework achieves state-of-the-art performance in image compression with extremely low bit-rates. We also extend the proposed framework to feature compression towards object detection, which saves 90% bit-rates than the VVC standard with the same detection accuracy. Fangyuan Gao, Xin Deng 0002, Junpeng Jing, Mai Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Proposal With Alignment: A Bi-Directional Transformer for 360° Video Viewport ProposalabstractPeople normally watch 360 ° videos through a head-mounted display, inside which only the content of viewports can be seen. Therefore, viewport proposal, referring to detecting potential viewport candidates, plays an important role in many 360 ° video processing tasks. In this paper, we advance the viewport proposal by further aligning the predicted viewports across frames for individual subject. This provides a better methodology and a deeper perspective to learn the human perceptual behaviours on 360 ° videos. Specifically, we first analyze three 360 ° video datasets and obtain several findings on human consistency, objectness and motion of viewports. Inspired by these findings, we propose a bi-directional transformer approach, named BiT, for 360 ° video viewport proposal and alignment. Specifically, BiT is composed of a multi-level residual module, a bi-directional encoder-decoder module and a spherical matching module. This way, the viewports can be well proposed and aligned via considering multi-level, bi-directional and non-local information. Moreover, the aligned viewports by BiT are used to refine the viewports and improve viewport proposal accuracy in return. Finally, we validate that our BiT approach is superior on viewport proposal, compared with the state-of-the-art approaches. Besides, the aligned viewports from BiT is verified to be effective in multiple applications, such as saliency prediction, trajectory prediction and perceptual video compression. Mai Xu, Lai Jiang 0004, Xin Deng 0002, Gaoxing Chen, Leonid Sigal |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Laplacian Gradient Consistency Prior for Flash Guided Non-Flash Image DenoisingabstractFor flash guided non-flash image denoising, the main challenge is to explore the consistency prior between the two modalities. Most existing methods attempt to model the flash/non-flash consistency in pixel level, which may easily lead to blurred edges. Different from these methods, we have an important finding in this paper, which reveals that the modality gap between flash and non-flash images conforms to the Laplacian distribution in gradient domain. Based on this finding, we establish a Laplacian gradient consistency (LGC) model for flash guided non-flash image denoising. This model is demonstrated to have faster convergence speed and denoising accuracy than the traditional pixel consistency model. Through solving the LGC model, we further design a deep network namely LGCNet. Different from existing image denoising networks, each component of the LGCNet strictly matches the solution of LGC model, giving the network good interpretability. The performance of the proposed LGCNet is evaluated on three different flash/non-flash image datasets, which demonstrates its superior denoising performance over many state-of-the-art methods both quantitatively and qualitatively. The intermediate features are also visualized to verify the effectiveness of the Laplacian gradient consistency prior. The source codes are available at https://github.com/JingyiXu404/LGCNet. Xin Deng 0002, Shengxi Li, Mai Xu |
IEEE Trans. Image Process. | 2 |
| 2023 | PIRNet: Privacy-Preserving Image Restoration Network via Wavelet LiftingabstractThe cloud-based multimedia service becomes increasingly popular in the last decade, however, it poses a serious threat to the client’s privacy. To address this issue, many methods utilized image encryption as a defense mechanism. However, the encrypted images look quite different from the natural images, making them vulnerable to attackers. In this paper, we propose a novel method namely PIRNet, which operates privacy-preserving image restoration in the steganographic domain. Compared to existing methods, our method offers significant advantages in terms of invisibility and security. Specifically, we first propose a wavelet Lifting-based Invertible Hiding (LIH) network to conceal the secret image into the stego image. Then, a Lifting-based Secure Restoration (LSR) network is utilized to perform image restoration in the steganographic domain. Since the secret image remains hidden throughout the whole image restoration process, the privacy of clients can be largely ensured. In addition, since the stego image looks visually the same as the cover image, the attackers can hardly discover it, which significantly improves the security. The experimental results on different datasets show the superiority of our PIRNet over the existing methods on various privacy-preserving image restoration tasks, including image denoising, deblurring and super-resolution. Xin Deng 0002, Mai Xu |
ICCV | 1 |
| 2023 | Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo MatchingabstractCorrelation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed model do not work uniformly well across various datasets, greatly limiting their real-world applicability. To tackle this issue, this paper proposes a new perspective to dynamically calculate correlation for robust stereo matching. A novel Uncertainty Guided Adaptive Correlation (UGAC) module is introduced to robustly adapt the same model for different scenarios. Specifically, a variance-based uncertainty estimation is employed to adaptively adjust the sampling area during warping operation. Additionally, we improve the traditional non-parametric warping with learnable parameters, such that the position-specific weights can be learned. We show that by empowering the recurrent network with the UGAC module, stereo matching can be exploited more robustly and effectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance over the ETH3D, KITTI, and Middlebury datasets when employing the same fixed model over these datasets without any retraining procedure. To target real-time applications, we further design a lightweight model based on UGAC, which also outperforms other methods over KITTI benchmarks with only 0.6 M parameters. Junpeng Jing, Jiankun Li, Pengfei Xiong, Jiangyu Liu, Shuaicheng Liu, Xin Deng 0002, Mai Xu, Lai Jiang 0004, Leonid Sigal |
ICCV | 7 |
| 2023 | Neural Characteristic Function Learning for Conditional Image GenerationabstractThe emergence of conditional generative adversarial networks (cGANs) has revolutionised the way we approach and control the generation, by means of adversarially learning joint distributions of data and auxiliary information. Despite the success, cGANs have been consistently put under scrutiny due to their ill-posed discrepancy measure between distributions, leading to mode collapse and instability problems in training. To address this issue, we propose a novel conditional characteristic function generative adversarial network (CCF-GAN) to reduce the discrepancy by the characteristic functions (CFs), which is able to learn accurate distance measure of joint distributions under theoretical soundness. More specifically, the difference between CFs is first proved to be complete and optimisation-friendly, for measuring the discrepancy of two joint distributions. To relieve the problem of curse of dimensionality in calculating CF difference, we propose to employ the neural network, namely neural CF (NCF), to efficiently minimise an upper bound of the difference. Based on the NCF, we establish the CCF-GAN framework to explicitly decompose CFs of joint distributions, which allows for learning the data distribution and auxiliary information with classified importance. The experimental results on synthetic and real-world datasets verify the superior performances of our CCF-GAN, on both the generation quality and stability. Shengxi Li, Mai Xu, Xin Deng 0002 |
ICCV | 5 |
| 2023 | ULcompress: A Unified low bit-rate image Compression Framework via Invertible Image RepresentationabstractIn this paper, we propose a unified low bit-rate image compression framework, namely ULCompress, via invertible image representation. The proposed framework is composed of two important modules, including an invertible image rescaling (IIR) module and a compressed quality enhancement (CQE) module. The role of IIR module is to learn a compression-friendly low-resolution (LR) image from the high-resolution (HR) image. Instead of the HR image, we compress the LR image to save the bit-rates. The compression codecs can be any existing codecs. After compression, we propose a CQE module to enhance the quality of the compressed LR image, which is then sent back to the IIR module to restore the original HR image. The network architecture of IIR module is specially designed to ensure the invertibility of LR and HR images, i.e., the downsampling and upsampling processes are invertible. The CQE module works as a buffer between IIR module and the codec, which plays an important role in improving the compatibility of our framework. Experimental results show that our ULCompress is compatible with both standard and learning-based codecs, and is able to significantly improve their performance at low bit-rates. Fangyuan Gao, Xin Deng 0002, Mai Xu |
ICIP | 2 |
| 2023 | A Two-stage hybrid CNN-Transformer Network for RGB Guided Indoor Depth CompletionabstractThe indoor captured raw depth images usually contain large in-homogeneous missing regions. Most existing methods are designed for the outdoor sparse depth completion, which struggle in completing the indoor depth with large holes. In this paper, to solve this problem, we propose a hybrid CNN-Transformer network for RGB guided indoor depth completion. The proposed network is composed of two stages to achieve depth completion in a coarse-to-fine manner. In the first stage, we propose a CNN based self-completion module (SCM) with cross scale attention to restore a coarse depth image. In the second stage, we further refine the completed depth image with the guidance of RGB image by proposing a guided completion module (GCM). To fully explore the guidance from the RGB image, we design a cross-modal Transformer (CMT) block to fuse the features from the depth and RGB modalities at different scales. Extensive experiments on NYUv2 and SUN RGB-D datasets demonstrate the superior performance of the proposed method over other state-of-the-art methods both quantitatively and qualitatively. The code is available at https://github.com/eecoder-dyf/ICME-2023-depth-completion. Yufan Deng, Xin Deng 0002, Mai Xu |
ICME | 2 |
| 2023 | Recruiting the Best Teacher Modality: A Customized Knowledge Distillation Method for if Based Nephropathy Diagnosis
Lai Jiang 0004, Yibing Fu, Sai Pan, Mai Xu, Xin Deng 0002, Xiangmei Chen |
MICCAI (5) | 6 |
| 2023 | DeepMIH: Deep Invertible Network for Multiple Image HidingabstractMultiple image hiding aims to hide multiple secret images into a single cover image, and then recover all secret images perfectly. Such high-capacity hiding may easily lead to contour shadows or color distortion, which makes multiple image hiding a very challenging task. In this paper, we propose a novel multiple image hiding framework based on invertible neural network, namely DeepMIH. Specifically, we develop an invertible hiding neural network (IHNN) to innovatively model the image concealing and revealing as its forward and backward processes, making them fully coupled and reversible. The IHNN is highly flexible, which can be cascaded as many times as required to achieve the hiding of multiple images. To enhance the invisibility, we design an importance map (IM) module to guide the current image hiding based on the previous image hiding results. In addition, we find that the image hidden in the high-frequency sub-bands tends to achieve better hiding performance, and thus propose a low-frequency wavelet loss to constrain that no secret information is hidden in the low-frequency sub-bands. Experimental results show that our DeepMIH significantly outperforms other state-of-the-art methods, in terms of hiding invisibility, security and recovery accuracy on a variety of datasets. Zhenyu Guan 0002, Junpeng Jing, Xin Deng 0002, Mai Xu, Lai Jiang 0004, Zhou Zhang 0016 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | DAQE: Enhancing the Quality of Compressed Images by Exploiting the Inherent Characteristic of DefocusabstractImage defocus is inherent in the physics of image formation caused by the optical aberration of lenses, providing plentiful information on image quality. Unfortunately, existing quality enhancement approaches for compressed images neglect the inherent characteristic of defocus, resulting in inferior performance. This paper finds that in compressed images, significantly defocused regions have better compression quality, and two regions with different defocus values possess diverse texture patterns. These observations motivate our defocus-aware quality enhancement (DAQE) approach. Specifically, we propose a novel dynamic region-based deep learning architecture of the DAQE approach, which considers the regionwise defocus difference of compressed images in two aspects. (1) The DAQE approach employs fewer computational resources to enhance the quality of significantly defocused regions and more resources to enhance the quality of other regions; (2) The DAQE approach learns to separately enhance diverse texture patterns for regions with different defocus values, such that texture-specific enhancement can be achieved. Extensive experiments validate the superiority of our DAQE approach over state-of-the-art approaches in terms of quality enhancement and resource savings. Qunliang Xing, Mai Xu, Xin Deng 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | New Finding and Unified Framework for Fake Image DetectionabstractRecently, fake face images generated by generative adversarial network (GAN) have been widely spread in social networks, raising serious social concerns and security risks. To identify the fake images, the top priority is to find what properties make the fake images different from the real images. In this letter, we reveal an important observation about real/fake images, i.e., the GAN generated fake images contain stronger non-local self-similarity than the real images. Motivated by this observation, we propose a simple yet effective non-local attention based fake image detection network, namely NAFID, to distinguish GAN generated fake images from real images. Specifically, we develop a non-local feature extraction (NFE) module to extract the non-local features of the real/fake images, followed by a multi-stage classification module to distinguish the images with the extracted non-local features. Experimental results on various datasets demonstrate the superiority of our NAFID over state-of-the-art (SOTA) face forgery detection methods. More importantly, since the NFE module is independent from classification, we can plug it into any other forgery detection models. The results show that the NFE module can consistently improve the detection accuracy of other models, which verifies the universality of the proposed method. Xin Deng 0002, Bihe Zhao, Zhenyu Guan 0002, Mai Xu |
IEEE Signal Process. Lett. | 1 |
| 2023 | MASIC: Deep Mask Stereo Image CompressionabstractStereo image compression (SIC) aims to simultaneously compress a pair of left and right stereoscopic images, which can achieve higher compression efficiency than single image compression. In this paper, to benefit the SIC tasks, we collect a large real-world stereo image dataset, namely Palace, which is composed of hundreds of stereo image pairs at high-resolution. More importantly, we propose a novel mask stereo image compression network, namely MASIC, which can jointly compress the stereo images with high compression efficiency. Specifically, we first estimate the homography matrix between the stereo images through a regression model. Then, the left image is spatially transformed by the homography matrix, so that only the residual information needs to be encoded for the right image. To avoid the wrong guidance between stereo image pair, we propose a mask prediction module (MPM) to generate a multi-channel guided mask to navigate both the encoding and decoding processes. Based on the guided mask, we introduce a new mask conditional stereo entropy (MCSE) model, to fully explore the correlation between the stereo images in entropy coding. In the decoder, we develop a stereo decoding module to simultaneously decode the stereo images and enhance their compression quality. Experimental results show that our MASIC significantly advances the performance of SIC both quantitatively and qualitatively on a variety of datasets, and is robust to the change of parallax level between stereo images. The software codes are available athttps://github.com/eecoder-dyf/MASIC. Xin Deng 0002, Yufan Deng, Radu Timofte, Mai Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Interpretable Multi-Modal Image Registration Network Based on Disentangled Convolutional Sparse CodingabstractMulti-modal image registration aims to spatially align two images from different modalities to make their feature points match with each other. Captured by different sensors, the images from different modalities often contain many distinct features, which makes it challenging to find their accurate correspondences. With the success of deep learning, many deep networks have been proposed to align multi-modal images, however, they are mostly lack of interpretability. In this paper, we first model the multi-modal image registration problem as a disentangled convolutional sparse coding (DCSC) model. In this model, the multi-modal features that are responsible for alignment (RA features) are well separated from the features that are not responsible for alignment (nRA features). By only allowing the RA features to participate in the deformation field prediction, we can eliminate the interference of the nRA features to improve the registration accuracy and efficiency. The optimization process of the DCSC model to separate the RA and nRA features is then turned into a deep network, namely Interpretable Multi-modal Image Registration Network (InMIR-Net). To ensure the accurate separation of RA and nRA features, we further design an accompanying guidance network (AG-Net) to supervise the extraction of RA features in InMIR-Net. The advantage of InMIR-Net is that it provides a universal framework to tackle both rigid and non-rigid multi-modal image registration tasks. Extensive experimental results verify the effectiveness of our method on both rigid and non-rigid registrations on various multi-modal image datasets, including RGB/depth images, RGB/near-infrared (NIR) images, RGB/multi-spectral images, T1/T2 weighted magnetic resonance (MR) images and computed tomography (CT)/MR images. The codes are available at https://github.com/lep990816/Interpretable-Multi-modal-Image-Registration. Xin Deng 0002, Enpeng Liu, Shengxi Li, Yiping Duan, Mai Xu |
IEEE Trans. Image Process. | 1 |
| 2023 | Omnidirectional Image Super-Resolution via Latitude Adaptive NetworkabstractOmnidirectional images (ODI), also known as 360 images, have recently attracted extensive attention from both academia and industry. However, due to storage and transmission limitations, ODIs are usually at extremely low resolution. Thus, it is necessary to restore a high-resolution ODI from a low-resolution ODI, i.e., omnidirectional image super-resolution (ODI-SR). Different from traditional two-dimensional (2D) image SR, the challenge of ODI-SR is the nonuniformly distributed pixel density and geometric distortion across latitudes, which makes traditional SR methods difficult to be applied in ODI-SR. Towards ODI-SR, we propose in this paper a novel latitude-aware upscaling network, namely LAU-Net+, which fully considers the above characteristics of ODIs. In our network, different latitude bands can learn to adopt distinct upscaling factors, which significantly saves the computational resources and improves the SR efficiency. Specifically, a Laplacian multilevel pyramid network is introduced in which the upscaling factor is gradually increased with the number of levels. Each level is composed of a feature enhancement module (FEM), a drop-band decision module (DDM) and a high-latitude enhancement module (HEM). The FEM module serves to enhance the high-level features extracted from the input ODI, while the role of DDM is to dynamically drop the unnecessary high latitude bands and send the remained bands to the next level. The HEM is adopted to further enhance high-level features of dropped latitude bands with a lightweight architecture. In DDM, we develop a reinforcement learning scheme with a latitude adaptive reward to determine which band should be dropped. To the best of our knowledge, our method is the first work which considers the latitude characteristics for ODI-SR task. Extensive experimental results demonstrate that our LAU-Net+ achieves state-of-the-art results on ODI-SR both quantitatively and qualitatively on various ODI datasets. Xin Deng 0002, Hao Wang 0049, Mai Xu, Zulin Wang |
IEEE Trans. Multim. | 1 |
| 2022 | SFIC: Sparsity-Driven Facial Image Compression NetworkabstractFacial image compression is crucial in many areas like social media and video surveillance. Considering the sparsity of facial features, sparse representation (SR) has been applied to compress facial images, in which each image patch is sparsely represented by a small number of dictionary atoms to save bit-rates. Along this line, we propose the first end-to-end sparsity-driven facial image compression network namely SFIC. In the proposed network, the traditional convolutional sparse coding (CSC) is turned into a learnable CSC block, which is combined with discrete wavelet transform (DWT) to form the sparsity encoding module (SEM). This is the first time that CSC has been explored in facial image compression. In the decoding side, a corresponding sparsity decoding module (SDM) is used to decode the image, and we further propose a quality enhancement module (QEM) to enhance the quality of decoded image. The experimental results verify that the proposed SFIC network achieves 74%, 55%, and 33% bit-rate savings over JPEG, JPEG-2000, and HEVC. Fangyuan Gao, Xin Deng 0002, Mai Xu |
ICIP | 2 |
| 2022 | A Learning-based Approach for Martian Image CompressionabstractFor the scientific exploration and research on Mars, it is an indispensable step to transmit high-quality Martian images from distant Mars to Earth. Image compression is the key technique given the extremely limited Mars-Earth bandwidth. Recently, deep learning has demonstrated remarkable performance in natural image compression, which provides a possibility for efficient Martian image compression. However, deep learning usually requires large training data. In this paper, we establish the first large-scale high-resolution Martian image compression (MIC) dataset. Through analyzing this dataset, we observe an important non-local self-similarity prior for Marian images. Benefiting from this prior, we propose a deep Martian image compression network with the non-local block to explore both local and non-local dependencies among Martian image patches. Experimental results verify the effectiveness of the proposed network in Martian image compression, which outperforms both the deep learning based compression methods and HEVC codec. Mai Xu, Shengxi Li, Xin Deng 0002, Qiu Shen |
VCIP | 4 |
| 2022 | Hierarchical Bayesian LSTM for Head Trajectory Prediction on Omnidirectional ImagesabstractWhen viewing omnidirectional images (ODIs), viewers can access different viewports via head movement (HM), which sequentially forms head trajectories in spatial-temporal domain. Thus, head trajectories play a key role in modeling human attention on ODIs. In this paper, we establish a large-scale dataset collecting 21,600 head trajectories on 1,080 ODIs. By mining our dataset, we find two important factors influencing head trajectories, i.e., temporal dependency and subject-specific variance. Accordingly, we propose a novel approach integrating hierarchical Bayesian inference into long short-term memory (LSTM) network for head trajectory prediction on ODIs, which is called HiBayes-LSTM. In HiBayes-LSTM, we develop a mechanism of Future Intention Estimation (FIE), which captures the temporal correlations from previous, current and estimated future information, for predicting viewport transition. Additionally, a training scheme called Hierarchical Bayesian inference (HBI) is developed for modeling inter-subject uncertainty in HiBayes-LSTM. For HBI, we introduce a joint Gaussian distribution in a hierarchy, to approximate the posterior distribution over network weights. By sampling subject-specific weights from the approximated posterior distribution, our HiBayes-LSTM approach can yield diverse viewport transition among different subjects and obtain multiple head trajectories. Extensive experiments validate that our HiBayes-LSTM approach significantly outperforms 9 state-of-the-art approaches for trajectory prediction on ODIs, and then it is successfully applied to predict saliency on ODIs. Li Yang 0014, Mai Xu, Xin Deng 0002, Fangyuan Gao, Zhenyu Guan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Revisiting Convolutional Sparse Coding for Image Denoising: From a Multi-Scale PerspectiveabstractRecently, convolutional sparse coding (CSC) has shown great success in many image processing tasks, such as image super-resolution and image separation. However, it performs poorly in image denoising task. In this letter, we provide a new insight for CSC denoising by revisiting the CSC from a multi-scale perspective. We propose a multi-scale CSC model for image denoising. By unrolling the multi-scale solution into a learnable network, we obtain an interpretable lightweight multiscale network, namely MCSCNet. Experimental results show that the proposed MCSCNet significantly advances the denoising performance, with an average PSNR improvement of 0.32 dB over the state-of-the-art (SOTA) CSC based method. In addition, our MCSCNet is on par with many SOTA deep learning based methods, with less network parameters and lower FLOPs. The ablation study also validates the effectiveness of the multi-scale CSC mechanism. Xin Deng 0002, Mai Xu |
IEEE Signal Process. Lett. | 2 |
| 2022 | MW-GAN+ for Perceptual Quality Enhancement on Compressed VideoabstractThe great success of deep learning has boosted the fast development of video quality enhancement. However, existing methods mainly focus on enhancing the objective quality of compressed video, and ignore their perceptual quality that plays a key role in determining quality of experience (QoE) of videos. In this paper, we aim at enhancing the perceptual quality of compressed video. Our main observation is that perceptual quality enhancement mostly relies on recovering the high-frequency details with fine textures. Accordingly, we propose a novel generative adversarial network (GAN) based on multi-level wavelet packet transform (WPT), which is called multi-level wavelet-based GAN+ (MW-GAN+), to exploit high-frequency details for enhancing the perceptual quality of compressed video. In MW-GAN+, we first propose a multi-level wavelet pixel-adaptive (MWP) module to extract temporal information across video frames, such that frame similarity can be utilized in recovering high-frequency details. Then, a wavelet reconstruction network, consisting of wavelet-dense residual blocks (WDRB), is developed to recover high-frequency details in a multi-level manner for enhanced frame reconstruction. Finally, we develop a 3D discriminator to encourage temporal coherence with a 3D-CNN based architecture. Experimental results demonstrate the superiority of our method over state-of-the-art methods in enhancing the perceptual quality of compressed video. Our code is available athttps://github.com/IceClear/MW-GAN. Jianyi Wang, Mai Xu, Xin Deng 0002, Liquan Shen, Yuhang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Multi-Modal Convolutional Dictionary LearningabstractConvolutional dictionary learning has become increasingly popular in signal and image processing for its ability to overcome the limitations of traditional patch-based dictionary learning. Although most studies on convolutional dictionary learning mainly focus on the unimodal case, real-world image processing tasks usually involve images from multiple modalities, e.g., visible and near-infrared (NIR) images. Thus, it is necessary to explore convolutional dictionary learning across different modalities. In this paper, we propose a novel multi-modal convolutional dictionary learning algorithm, which efficiently correlates different image modalities and fully considers neighborhood information at the image level. In this model, each modality is represented by two convolutional dictionaries, in which one dictionary is for common feature representation and the other is for unique feature representation. The model is constrained by the requirement that the convolutional sparse representations (CSRs) for the common features should be the same across different modalities, considering that these images are captured from the same scene. We propose a new training method based on the alternating direction method of multipliers (ADMM) to alternatively learn the common and unique dictionaries in the discrete Fourier transform (DFT) domain. We show that our model converges in less than 20 iterations between the convolutional dictionary updating and the CSRs calculation. The effectiveness of the proposed dictionary learning algorithm is demonstrated on various multimodal image processing tasks, achieves better performance than both dictionary learning methods and deep learning based methods with limited training data. Fangyuan Gao, Xin Deng 0002, Mai Xu, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 2 |
| 2021 | LAU-Net: Latitude Adaptive Upscaling Network for Omnidirectional Image Super-ResolutionabstractThe omnidirectional images (ODIs) are usually at low-resolution, due to the constraints of collection, storage and transmission. The traditional two-dimensional (2D) image super-resolution methods are not effective for spherical ODIs, because ODIs tend to have non-uniformly distributed pixel density and varying texture complexity across latitudes. In this work, we propose a novel latitude adaptive upscaling network (LAU-Net) for ODI super-resolution, which allows pixels at different latitudes to adopt distinct upscaling factors. Specifically, we introduce a Laplacian multi-level separation architecture to split an ODI into different latitude bands, and hierarchically upscale them with different factors. In addition, we propose a deep reinforcement learning scheme with a latitude adaptive reward, in order to automatically select optimal upscaling factors for different latitude bands. To the best of our knowledge, LAU-Net is the first attempt to consider the latitude difference for ODI super-resolution. Extensive results demonstrate that our LAU-Net significantly advances the super-resolution performance for ODIs. Codes are available at https://github.com/wangh-allen/LAU-Net. Xin Deng 0002, Hao Wang 0049, Mai Xu, Yuhang Song 0001, Li Yang 0014 |
CVPR | 1 |
| 2021 | Deep Homography for Efficient Stereo Image CompressionabstractIn this paper, we propose HESIC, an end-to-end trainable deep network for stereo image compression (SIC). To fully explore the mutual information across two stereo images, we use a deep regression model to estimate the homography matrix, i.e., H matrix. Then, the left image is spatially transformed by the H matrix, and only the residual information between the left and right images is encoded to save bitrates. A two-branch auto-encoder architecture is adopted in HESIC, corresponding to the left and right images, respectively. For entropy coding, we use two conditional stereo entropy models, i.e., Gaussian mixture model (GMM) based and context based entropy models, to fully explore the correlation between the two images to reduce the coding bit-rates. In decoding, a cross quality enhancement module is proposed to enhance the image quality based on inverse H matrix. Experimental results show that our HESIC outperforms state-of-the-art SIC methods on InStereo2K and KITTI datasets both quantitatively and qualitatively. Code is available at https://github.com/ywz978020607/HESIC. Xin Deng 0002, Mai Xu, Enpeng Liu, Qianhan Feng, Radu Timofte |
CVPR | 1 |
| 2021 | HiNet: Deep Image Hiding by Invertible NetworkabstractImage hiding aims to hide a secret image into a cover image in an imperceptible way, and then recover the secret image perfectly at the receiver end. Capacity, invisibility and security are three primary challenges in image hiding task. This paper proposes a novel invertible neural network (INN) based framework, HiNet, to simultaneously overcome the three challenges in image hiding. For large capacity, we propose an inverse learning mechanism by simultaneously learning the image concealing and revealing processes. Our method is able to achieve the concealing of a full-size secret image into a cover image with the same size. For high invisibility, instead of pixel domain hiding, we propose to hide the secret information in wavelet domain. Furthermore, we propose a new low-frequency wavelet loss to constrain that secret information is hidden in high-frequency wavelet subbands, which significantly improves the hiding security. Experimental results show that our HiNet significantly outperforms other state-of-the-art image hiding methods, with more than 10 dB PSNR improvement in secret image recovery on ImageNet, COCO and DIV2K datasets. Codes are available at https://github.com/TomTomTommi/HiNet. Junpeng Jing, Xin Deng 0002, Mai Xu, Jianyi Wang, Zhenyu Guan 0002 |
ICCV | 2 |
| 2021 | CU-Net+: Deep Fully Interpretable Network for Multi-Modal Image RestorationabstractThe network interpretability is critical in computer vision related tasks, especially for tasks involving multiple modalities. For multi-modal image restoration, one recent method, CU-Net, introduces an interpretable network based on a multi-modal convolutional sparse coding model. However, its network architecture does not mimic in full the proposed sparse model. In this paper, we overcome the limitation of CU-Net by using recurrent scheme, and this leads to a fully interpretable network which we call CU-Net+. In addition, we relax the constraint on the number of common and unique features in CU-Net, for making it more consistent with real condition. The effectiveness of the proposed CU-Net+ is evaluated on RGB guided depth image super-resolution and flash guided non-flash image denoising tasks. The numerical results show that CU-Net+ outperforms other interpretable or non-interpretable methods, with 0.16 RMSE and 0.66 dB PSNR improvement over CU-Net for the two mentioned tasks, respectively. Code is available at https://git;hub.com/JingyiXu404/CU-Net;-plus. Xin Deng 0002, Mai Xu, Pier Luigi Dragotti |
ICIP | 2 |
| 2021 | Spatial Attention-Based Non-Reference Perceptual Quality Prediction Network for Omnidirectional ImagesabstractDue to the strong correlation between visual attention and perceptual quality, many methods attempt to use human saliency information for image quality assessment. Although this mechanism can get good performance, the networks require human saliency labels, which is not easily accessible for omnidirectional images (ODI). To alleviate this issue, we propose a spatial attention-based perceptual quality prediction network for non-reference quality assessment on ODIs (SAP-net). Without any human saliency labels, our network can adaptively estimate human perceptual quality on impaired ODIs through a self-attention manner, which significantly promotes the prediction performance of quality scores. Moreover, our method greatly reduces the computational complexity in quality assessment task on ODIs. Extensive experiments validate that our network outperforms 9 state-of-the-art methods for quality assessment on ODIs. The dataset and code have been available on https://github.com/yanglixiaoshen/SAP-Net. Li Yang 0014, Mai Xu, Xin Deng 0002 |
ICME | 3 |
| 2021 | Deep Convolutional Neural Network for Multi-Modal Image Restoration and FusionabstractIn this paper, we propose a novel deep convolutional neural network to solve the general multi-modal image restoration (MIR) and multi-modal image fusion (MIF) problems. Different from other methods based on deep learning, our network architecture is designed by drawing inspirations from a new proposed multi-modal convolutional sparse coding (MCSC) model. The key feature of the proposed network is that it can automatically split the common information shared among different modalities, from the unique information that belongs to each single modality, and is therefore denoted with CU-Net, i.e., common and unique information splitting network. Specifically, the CU-Net is composed of three modules, i.e., the unique feature extraction module (UFEM), common feature preservation module (CFPM), and image reconstruction module (IRM). The architecture of each module is derived from the corresponding part in the MCSC model, which consists of several learned convolutional sparse coding (LCSC) blocks. Extensive numerical results verify the effectiveness of our method on a variety of MIR and MIF tasks, including RGB guided depth image super-resolution, flash guided non-flash image denoising, multi-focus and multi-exposure image fusion. Xin Deng 0002, Pier Luigi Dragotti |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Deep Coupled Feedback Network for Joint Exposure Fusion and Image Super-ResolutionabstractNowadays, people are getting used to taking photos to record their daily life, however, the photos are actually not consistent with the real natural scenes. The two main differences are that the photos tend to have low dynamic range (LDR) and low resolution (LR), due to the inherent imaging limitations of cameras. The multi-exposure image fusion (MEF) and image super-resolution (SR) are two widely-used techniques to address these two issues. However, they are usually treated as independent researches. In this paper, we propose a deep Coupled Feedback Network (CF-Net) to achieve MEF and SR simultaneously. Given a pair of extremely over-exposed and under-exposed LDR images with low-resolution, our CF-Net is able to generate an image with both high dynamic range (HDR) and high-resolution. Specifically, the CF-Net is composed of two coupled recursive sub-networks, with LR over-exposed and under-exposed images as inputs, respectively. Each sub-network consists of one feature extraction block (FEB), one super-resolution block (SRB) and several coupled feedback blocks (CFB). The FEB and SRB are to extract high-level features from the input LDR image, which are required to be helpful for resolution enhancement. The CFB is arranged after SRB, and its role is to absorb the learned features from the SRBs of the two sub-networks, so that it can produce a high-resolution HDR image. We have a series of CFBs in order to progressively refine the fused high-resolution HDR image. Extensive experimental results show that our CF-Net drastically outperforms other state-of-the-art methods in terms of both SR accuracy and fusion performance. The software code is available here https://github.com/ytZhang99/CF-Net. Xin Deng 0002, Mai Xu, Shuhang Gu, Yiping Duan |
IEEE Trans. Image Process. | 1 |
| 2021 | Joint Learning of 3D Lesion Segmentation and Classification for Explainable COVID-19 DiagnosisabstractGiven the outbreak of COVID-19 pandemic and the shortage of medical resource, extensive deep learning models have been proposed for automatic COVID-19 diagnosis, based on 3D computed tomography (CT) scans. However, the existing models independently process the 3D lesion segmentation and disease classification, ignoring the inherent correlation between these two tasks. In this paper, we propose a joint deep learning model of 3D lesion segmentation and classification for diagnosing COVID-19, called DeepSC-COVID, as the first attempt in this direction. Specifically, we establish a large-scale CT database containing 1,805 3D CT scans with fine-grained lesion annotations, and reveal 4 findings about lesion difference between COVID-19 and community acquired pneumonia (CAP). Inspired by our findings, DeepSC-COVID is designed with 3 subnets: a cross-task feature subnet for feature extraction, a 3D lesion subnet for lesion segmentation, and a classification subnet for disease diagnosis. Besides, the task-aware loss is proposed for learning the task interaction across the 3D lesion and classification subnets. Different from all existing models for COVID-19 diagnosis, our model is interpretable with fine-grained 3D lesion distribution. Finally, extensive experimental results show that the joint learning framework in our model significantly improves the performance of 3D lesion segmentation and disease classification in both efficiency and efficacy. Xiaofei Wang 0004, Lai Jiang 0004, Liu Li 0001, Mai Xu, Xin Deng 0002, Lisong Dai, Tianyi Li 0004, Zulin Wang, Pier Luigi Dragotti |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Multi-level Wavelet-Based Generative Adversarial Network for Perceptual Quality Enhancement of Compressed Video
Jianyi Wang, Xin Deng 0002, Mai Xu, Congyong Chen, Yuhang Song 0001 |
ECCV (14) | 2 |
| 2020 | RADAR: Robust Algorithm for Depth Image Super Resolution Based on FRI Theory and Multimodal Dictionary LearningabstractDepth image super-resolution is a challenging problem, since normally high upscaling factors are required (e.g., 16×), and depth images are often noisy. In order to achieve large upscaling factors and resilience to noise, we propose a Robust Algorithm for Depth imAge super Resolution (RADAR) that combines the power of finite rate of innovation (FRI) theory with multimodal dictionary learning. Given a low-resolution (LR) depth image, we first model its rows and columns as piece-wise polynomials and propose an FRI-based depth upscaling (FDU) algorithm to super-resolve the image. Then, the upscaled moderate quality (MQ) depth image is further enhanced with the guidance of a registered high-resolution (HR) intensity image. This is achieved by learning multimodal mappings from the joint MQ depth and HR intensity pairs to the HR depth, through a recently proposed triple dictionary learning (TDL) algorithm. Moreover, to speed up the super-resolution process, we introduce a new projection-based rapid upscaling (PRU) technique that pre-calculates the projections from the joint MQ depth and HR intensity pairs to the HR depth. Compared with the state-of-the-art deep learning-based methods, our approach has two distinct advantages: we need a fraction of training data but can achieve the best performance, and we are resilient to mismatches between training and testing datasets. The extensive numerical results show that the proposed method outperforms other state-of-the-art methods on either noise-free or noisy datasets with large upscaling factors up to 16× and can handle unknown blurring kernels well. Xin Deng 0002, Pingfan Song, Miguel R. D. Rodrigues, Pier Luigi Dragotti |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Deep Coupled ISTA Network for Multi-Modal Image Super-ResolutionabstractGiven a low-resolution (LR) image, multi-modal image super-resolution (MISR) aims to find the high-resolution (HR) version of this image with the guidance of an HR image from another modality. In this paper, we use a model-based approach to design a new deep network architecture for MISR. We first introduce a novel joint multi-modal dictionary learning (JMDL) algorithm to model cross-modality dependency. In JMDL, we simultaneously learn three dictionaries and two transform matrices to combine the modalities. Then, by unfolding the iterative shrinkage and thresholding algorithm (ISTA), we turn the JMDL model into a deep neural network, called deep coupled ISTA network. Since the network initialization plays an important role in deep network training, we further propose a layer-wise optimization algorithm (LOA) to initialize the parameters of the network before running back-propagation strategy. Specifically, we model the network initialization as a multi-layer dictionary learning problem, and solve it through convex optimization. The proposed LOA is demonstrated to effectively decrease the training loss and increase the reconstruction accuracy. Finally, we compare our method with other state-of-the-art methods in the MISR task. The numerical results show that our method consistently outperforms others both quantitatively and qualitatively at different upscaling factors for various multi-modal scenarios. Xin Deng 0002, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 1 |
| 2020 | Accelerate CTU Partition to Real Time for HEVC Encoding With Complexity ControlabstractRecently, extensive approaches have been proposed for reducing the encoding complexity of high efficiency video coding, by predicting the coding tree unit partition using deep neural networks. However, these approaches cannot work in real time due to the complexity of the network architectures. In this paper, we propose a network pruning approach to accelerate a state-of-the-art deep neural network model, for real-time coding tree unit partition. Specifically, we first investigate the computational complexity throughout the network, and find that most calculations can be simplified by pruning the weight parameters. Considering that the number of weight parameters drastically differs by network layer and partition level, we design an adaptive pruning scheme by applying a well-suitable retention ratio of weight parameters to each layer at a level. The retention ratio indicates the ratio of weight parameters after and before pruning. By varying the retention ratios, we can obtain several accelerated network models with different levels of complexity. We further propose a complexity control algorithm by applying different accelerated models to different coding tree units, to ensure that the actual encoding complexity is close to a given target. To guarantee the rate-distortion performance, we model the complexity control algorithm as a convex optimization problem, and we can obtain a closed-form solution. Experimental results show that our approach can accelerate the original deep neural network model by 17-20 times, with little expense on the Bjøntegaard delta bit-rate. For complexity control, we achieve high control accuracy with a control error of less than 2% for most video sequences. Tianyi Li 0004, Mai Xu, Xin Deng 0002, Liquan Shen |
IEEE Trans. Image Process. | 3 |
| 2019 | Coupled Ista Network for Multi-modal Image Super-resolutionabstractIn this paper, we propose a novel deep neural network architecture for multi-modal image super-resolution (MISR). The architecture is based on a new joint multi-modal dictionary learning (JMDL) algorithm to model cross-modality dependency and to map them to a high-resolution version of one modality. In JMDL, we learn three dictionaries and two transform matrices to combine the modalities. By using the learned model, we then design the network architecture by a coupled unfolding of the iterative shrinkage and thresholding algorithm (ISTA). We finally initialize the parameters of our network with a new optimization strategy. The initialized parameters are demonstrated to effectively decrease the training loss and increase the reconstruction accuracy. The numerical results show that our method outperforms other state-of-the-art methods quantitatively and qualitatively for MISR. Xin Deng 0002, Pier Luigi Dragotti |
ICASSP | 1 |
| 2019 | Wavelet Domain Style Transfer for an Effective Perception-Distortion Tradeoff in Single Image Super-ResolutionabstractIn single image super-resolution (SISR), given a low-resolution (LR) image, one wishes to find a high-resolution (HR) version of it which is both accurate and photorealistic. Recently, it has been shown that there exists a fundamental tradeoff between low distortion and high perceptual quality, and the generative adversarial network (GAN) is demonstrated to approach the perception-distortion (PD) bound effectively. In this paper, we propose a novel method based on wavelet domain style transfer (WDST), which achieves a better PD tradeoff than the GAN based methods. Specifically, we propose to use 2D stationary wavelet transform (SWT) to decompose one image into low-frequency and high-frequency sub-bands. For the low-frequency sub-band, we improve its objective quality through an enhancement network. For the high-frequency sub-band, we propose to use WDST to effectively improve its perceptual quality. By feat of the perfect reconstruction property of wavelets, these sub-bands can be re-combined to obtain an image which has simultaneously high objective and perceptual quality. The numerical results on various datasets show that our method achieves the best trade-off between the distortion and perceptual quality among the existing state-of-the-art SISR methods. Xin Deng 0002, Mai Xu, Pier Luigi Dragotti |
ICCV | 1 |
| 2018 | U-Fresh: An Fri-Based Single Image Super Resolution Algorithm and An Application in Image CompressionabstractLearning based single image super resolution (SISR) methods have achieved notable results, however, they require large datasets for training, and may struggle when there is a mismatch between the testing and training data. To overcome these drawbacks, we propose an approach, named U - FRESH, which only requires a small dataset but can achieve state-of-the-art performance also in the presence of training and testing mismatches. We accomplish this by leveraging a method called FRESH, which enhances the image resolution using FRI theory. We start upscaling from the FRESH generated low resolution image. To minimize the reconstruction error, we propose a new regression selection technique to make the mapping more reliable and robust, and a wavelet based back projection technique to improve the quality of the reconstructed image. Based on U - FRESH, we also propose a new framework based on JPEG 2000 for image compression. Numerical results show that our U-FRESH method achieves state-of-the-art performance in SISR and provides better compression results than JPEG 2000. Xin Deng 0002, Junjie Huang 0001, Mengying Liu, Pier Luigi Dragotti |
ICASSP | 1 |
| 2018 | Enhancing Image Quality via Style Transfer for Single Image Super-ResolutionabstractRecently, by feat of the Generative Adversarial Network (GAN), single image super-resolution (SISR) has achieved great breakthroughs in enhancing the perceptual image quality. However, since the network is trained by minimizing the perceptual loss, the GAN based SISR method (SRGAN) [1] results in images with very low objective quality, i.e., peak signal-to-noise ratio (PSNR). In this letter, we aim to solve this problem in an image style transfer way, to generate an image with similar perceptual quality as SRGAN, but with much higher objective quality. Moreover, we propose a threshold-based method to automatically alter the objective and perceptual quality of the reconstructed image through adjusting only one parameter. Experimental results show that our method can achieve more than 1.6 dB PSNR improvement over SRGAN with similar Mean Opinion Score value. Also, with the same objective quality, our method can provide significantly better perceptual results than other state-of-the-art SISR methods. Xin Deng 0002 |
IEEE Signal Process. Lett. | 1 |
| 2018 | Reducing Complexity of HEVC: A Deep Learning ApproachabstractHigh Efficiency Video Coding (HEVC) significantly reduces bit-rates over the preceding H.264 standard but at the expense of extremely high encoding complexity. In HEVC, the quad-tree partition of coding unit (CU) consumes a large proportion of the HEVC encoding complexity, due to the brute-force search for rate-distortion optimization (RDO). Therefore, this paper proposes a deep learning approach to predict the CU partition for reducing the HEVC complexity at both intra-and inter-modes, which is based on convolutional neural network (CNN) and long-and short-term memory (LSTM) network. First, we establish a large-scale database including substantial CU partition data for HEVC intra-and inter-modes. This enables deep learning on the CU partition. Second, we represent the CU partition of an entire coding tree unit (CTU) in the form of a hierarchical CU partition map (HCPM). Then, we propose an early-terminated hierarchical CNN (ETH-CNN) for learning to predict the HCPM. Consequently, the encoding complexity of intra-mode HEVC can be drastically reduced by replacing the brute-force search with ETH-CNN to decide the CU partition. Third, an early-terminated hierarchical LSTM (ETH-LSTM) is proposed to learn the temporal correlation of the CU partition. Then, we combine ETH-LSTM and ETH-CNN to predict the CU partition for reducing the HEVC complexity at inter-mode. Finally, experimental results show that our approach outperforms other state-of-the-art approaches in reducing the HEVC complexity at both intra-and inter-modes. Mai Xu, Tianyi Li 0004, Zulin Wang, Xin Deng 0002, Zhenyu Guan 0002 |
IEEE Trans. Image Process. | 4 |
| 2017 | Complexity control of HEVC for video conferencingabstractIn this paper, we propose an effective complexity control approach for video conferencing scenarios on HEVC platform. A complexity control formulation is established to determine the number of depth-constrained largest coding units (LCUs) according to the target complexity. By limiting the maximum depths of different LCUs to different levels, the encoding complexity can be controlled with high accuracy. Different from other approaches, both the objective and perceptual-driven video quality are kindly preserved through taking both the objective and subjective weight maps into consideration when controlling the complexity. The experimental results demonstrate that our approach outperforms the state-of-the art approach with higher control accuracy. Also, despite of complexity reduction, our approach keeps the objective and perceptual-driven quality well. Xin Deng 0002, Mai Xu |
ICASSP | 1 |
| 2017 | A deep convolutional neural network approach for complexity reduction on intra-mode HEVCabstractThe High Efficiency Video Coding (HEVC) standard significantly saves coding bit-rate over the proceeding H.264 standard, but at the expense of extremely high encoding complexity. In fact, the coding tree unit (CTU) partition consumes a large proportion of HEVC encoding complexity, due to the brute-force search for rate-distortion optimization (RDO). Therefore, we propose in this paper a complexity reduction approach for intra-mode HEVC, which learns a deep convolutional neural network (CNN) model to predict CTU partition instead of RDO. Firstly, we establish a large-scale database with diversiform patterns of CTU partition. Secondly, we model the partition as a three-level classification problem. Then, for solving the classification problem, we develop a deep CNN structure with various sizes of convolutional kernels and extensive trainable parameters, which can be learnt from the established database. Finally, experimental results show that our approach reduces intramode encoding time by 62.25% and 69.06% with negligible Bjontegaard delta bit-rate of 2.12% and 1.38%, over the test sequences and images respectively, superior to other state-of-the-art approaches. Tianyi Li 0004, Mai Xu, Xin Deng 0002 |
ICME | 3 |
| 2017 | A subjective visual quality assessment method of panoramic videosabstractDifferent from 2-dimensional (2D) videos, panoramic videos contain spherical viewing direction with the support of head-mounted displays, thus improving immersive and interactive visual experience. Unfortunately, to our best knowledge, there are few subjective visual quality assessment (VQA) methods for panoramic videos. In this paper, we therefore propose a subjective VQA method for assessing quality loss of impaired panoramic videos. Specifically, we first establish a database containing viewing direction data of several subjects on watching panoramic videos. Then, we find out that there exists high consistency of viewing direction on panoramic videos across different subjects. Upon this finding, we present a procedure of subjective test in measuring quality of panoramic videos by different subjects, yielding different mean opinion score (DMOS). To couple with inconsistency of viewing directions on panoramic videos, we further propose a vectorized DMOS metric. Finally, experimental results verify that our subjective VQA method, in the forms of both overall and vectorized DMOS metrics, is effective in measuring subjective quality of panoramic videos. Mai Xu, Chen Li 0049, Yufan Liu 0001, Xin Deng 0002, Jiaxin Lu 0003 |
ICME | 4 |
| 2016 | Subjective-Driven Complexity Control Approach for HEVCabstractThe latest High Efficiency Video Coding (HEVC) standard significantly increases the encoding complexity for improving its coding efficiency, compared with the preceding H.264/Advanced Video Coding (AVC) standard. In this paper, we present a novel subjective-driven complexity control (SCC) approach to reduce and control the encoding complexity of HEVC. Through reasonably adjusting the maximum depth of each largest coding unit (LCU), the encoding complexity can be reduced to a target level with minimal visual distortion. Specifically, the maximum depths of different LCUs can be varied through solving the proposed optimization formulation of complexity control, based on two explored relationships: 1) the relationship between the maximum depth and encoding complexity and 2) the relationship between the maximum depth and visual distortion. Besides, the subjective visual quality is favored with a novel subjective-driven constraint imposed in the formulation, on the basis of a visual attention model. Finally, the experimental results show that our approach can achieve a wide range of encoding complexity control (as low as 20%) for HEVC, with the smallest complexity bias being 0.2%. Meanwhile, our SCC approach outperforms other two state-of-the-art complexity control approaches, in terms of both control accuracy and visual quality. Xin Deng 0002, Mai Xu, Lai Jiang 0004, Xiaoyan Sun 0001, Zulin Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Weight-based R-λ rate control for perceptual HEVC coding on conversational videos
Shengxi Li, Mai Xu, Xin Deng 0002, Zulin Wang |
Signal Process. Image Commun. | 3 |
| 2014 | A novel weight-based URQ scheme for perceptual video coding of conversational video in HEVCabstractIn this paper, we propose a novel weight-based unified rate-quantization (URQ) scheme for rate control in state-of-the-art HEVC standard, to improve its perceived visual quality for conversational videos. In conventional rate control of HEVC, a pixel-wise URQ scheme is proposed by introducing the concept of bit per pixel (bpp). This scheme is able to assign different amounts of bits to the blocks with various sizes, thus well suitable for flexible picture partition of HEVC. However, bpp does not reflect the visual importance of each pixel. Therefore, we propose a novel weight-based URQ scheme to take into account the visual importance for rate control in HEVC. In combination with the weight map acquired from a novel hierarchical perceptual model of face, such a scheme is capable of allocating more bits to the face and much more bits to the facial features, by using bit per weight (bpw) instead of bpp. As a result, the visual quality of face, especially facial features, can be improved such that perceptual video coding is achieved for HEVC. Finally, the experimental results validate such improvement. Shengxi Li, Mai Xu, Xin Deng 0002, Zulin Wang |
ICME | 3 |
| 2014 | Complexity control of HEVC based on region-of-interest attention modelabstractIn this paper, we present a novel complexity control method of HEVC to adjust its encoding complexity. First, a region-of-interest (ROI) attention model is established, which defines different weights for various regions according to their importance. Then, the complexity control algorithm is proposed with a distortion-complexity optimization model, to determine the maximum depth of the largest coding units (LCUs) according to their weights. We can reduce the encoding complexity to a given target level at the cost of little distortion loss. Finally, the experimental results show that the encoding complexity can drop to a pre-defined target complexity as low as 20% with bias less than 7%. Meanwhile, our method is verified to preserve the quality of ROI better than another state-of-the-art approach. Xin Deng 0002, Mai Xu, Shengxi Li, Zulin Wang |
VCIP | 1 |