Mai Xu

dblp:20/5353 · DBLP profile ↗
← Back
197ranked-venue papers
28as first author
107since 2021 · last 2026
0000-0002-0277-3301ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 142 · 18 first-author · 72 since 2021Artificial intelligence and machine learning · 64 · 13 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 11 since 2021Computer networks · 6 · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream Tasks
abstract
In recent years, the development of burst imaging technology has improved the capture and processing capabilities of visual data, enabling a wide range of applications. However, the redundancy in burst images leads to the increased storage and transmission demands, as well as reduced efficiency of downstream tasks. To address this, we propose a new task of Burst Image Quality Assessment (BuIQA), to evaluate the task-driven quality of each frame within a burst sequence, providing reasonable cues for burst image selection. Specifically, we establish the first benchmark dataset for BuIQA, consisting of 7,346 burst sequences with 45,827 images and 191,572 annotated quality scores for multiple downstream scenarios. Inspired by the data analysis, a unified BuIQA framework is proposed to achieve an efficient adaption for BuIQA under diverse downstream scenarios. Specifically, a task-driven prompt generation network is developed with heterogeneous knowledge distillation, to learn the priors of the downstream task. Then, the task-aware quality assessment network is introduced to assess the burst image quality based on the task prompt. Extensive experiments across 10 downstream scenarios demonstrate the impressive BuIQA performance of the proposed approach, outperforming the state-of-the-art. Furthermore, it can achieve 0.33 dB PSNR improvement in the downstream tasks of denoising and super-resolution, by applying our approach to select the high-quality burst frames.
Xiaoye Liang, Lai Jiang 0004, Minglang Qiao, Yue Zhang 0082, Xin Deng 0002, Shengxi Li, Yufan Liu 0001, Mai Xu
AAAI9
2026 FMF-DETR: A Frequency-Aware Multi-Scale Fusion DETR for Small Object Detection
abstract
ABSTRACT Small object detection, a critical technique for recognizing and localizing diminutive targets in visual data, plays a vital role in applications ranging from remote sensing and unmanned aerial vehicle (UAV) vision to autonomous driving. Current methodologies, however, face substantial challenges, including detection accuracy limitations, low‐resolution image processing difficulties, background noise interference, and target occlusion issues. To address these challenges, we propose FMF‐DETR, an innovative small object detection framework featuring frequency‐domain feature optimization through three key components: (1) High‐Low Frequency Fusion Model (HLFM), (2) Focused Diffusion Feature Pyramid Network (FDFPN), and (3) BiPathNet (BPNet). Specifically, the HLFM module enhances multiscale feature representation by emphasizing high‐frequency details while suppressing low‐frequency background noise. The FDFPN architecture improves detection performance in complex scenarios through multiscale feature fusion and saliency‐aware diffusion. BPNet introduces a dual‐path feature extraction mechanism that simultaneously enhances feature discriminability and reduces computational overhead. Through the synergistic integration of these components, the proposed framework enhances both detection accuracy and operational efficiency. Comprehensive evaluations on the VisDrone dataset demonstrate FMF‐DETR's superior performance, achieving a 2.2% accuracy improvement while reducing model parameters by 13.03M and computational complexity by 97.2G FLOPs compared to baseline methods. These results validate both the effectiveness and efficiency of our proposed framework.
Lingling Li 0004, Yang Mei, Xuezhuan Zhao, Xiaoyan Shao, Zonghao Zhu, Shiqin Diao, Mai Xu
Concurr. Comput. Pract. Exp.7
2026 Say the image: Auditory masking effect-driven invertible network for progressive image-in-audio steganography
Jinghang Song, Fangyuan Gao, Xin Deng 0002, Shengxi Li, Mai Xu
J. Inf. Secur. Appl.5
2026 AIRPNet: Adaptive Image Restoration With Privacy Protection in Steganographic Domain
abstract
Cloud-based third-party multimedia services have become increasingly popular in last decade, however, they pose serious threats to users' privacy. To address this issue, in this paper, we propose a novel Adaptive Image Restoration network with Privacy protection, namely AIRPNet, which first attempts to perform image restoration in steganographic domain. Compared with existing methods, our method has significant advantages in invisibility, security and flexibility. Specifically, we first propose a wavelet lifting-based Adaptive Invertible Hiding (AIH) module to conceal the low-quality (LQ) secret image into a stego image. Then, instead of performing single type of restoration on the secret image, an adaptive secure restoration (ASR) module is developed to deal with multiple image degradations on the stego image. Finally, a high-quality (HQ) secret image can be extracted from the restored stego image. Here, since the secret image remains hidden throughout the whole image restoration process, the privacy of users can be greatly protected. The framework can be flexibly extended to multiple image restoration, which can restore multiple secret images from the same stego image. Experimental results on various datasets demonstrate that our AIRPNet outperforms existing methods in terms of restoration accuracy, invisibility and security on different image restoration tasks.
Fangyuan Gao, Xin Deng 0002, Junjie Huang 0001, Mai Xu
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Learning Continuous Spatiotemporal Implicit Neural Fields for Unsupervised Video Denoising
abstract
Video denoising is fundamental to low-level vision and real-world imaging, yet existing self-supervised methods remain fragile under severe noise and complex motion. Most approaches still rely on spatially and temporally discrete grid-based representations: blind-spot networks enforce J-invariance by masking center pixels with a limited receptive field, while recurrent models build temporal dependencies on discretized frame sequences and noise-sensitive optical flow, leading to error accumulation and motion artifacts. We address this model bottleneck by reformulating self-supervised video denoising as learning a continuous spatiotemporal implicit field. Building on coordinate-based implicit neural representations, we propose a unified video denoising model with a spatiotemporal implicit neural field (SINF). In the spatial domain, a blind-spot implicit spatial field maps coordinates directly to pixel-level representations, enabling globally informed texture recovery beyond receptive-field limits. In the temporal domain, an implicit temporal embedding with periodic activations encodes motion continuously over time, while a time-aware spatial graph module refines cross-frame alignment. Together, SINF remodels discretized video signals into a continuous spatiotemporal intensity field, enabling more robust pixel-wise associations than coarse optical flow. Extensive experiments on synthetic and real noisy video benchmarks demonstrate that our SINF achieves state-of-the-art performance on synthetic and real noisy video benchmarks.
Xiaowan Hu, Henan Liu, Mai Xu
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 UniFES: A Unified Recurrent Network for Quality Enhancement and Stabilization in Face Videos
abstract
Recent years have witnessed an explosive increase of face content, which drives a distinct shift from static images to dynamic video formats. The shift of formats inherently alters the characteristics within face videos, whereby pixel-wise artifacts are intertwined with motion-related impairments. Addressing the emerging distortions that now always appear by twins in practice, however, is challenging and non-trivial, due to the distinct characteristics in addressing spatial-temporal frequencies in videos. In this paper, we propose a novel Unified recurrent network for joint Face video quality Enhancement and Stabilization (UniFES), as the first successful attempt for both quality enhancement and motion stabilization. Correspondingly, our UniFES method proposes to effectively aggregate the mutual information in the pixel and motion domains. For the quality enhancement, our UniFES method decomposes the shaking temporal alignment problem into progressive feature alignment with explicit physical information, which includes the global dynamics from the motion domain, i.e., from the stabilization task. Regarding the video stabilization, we integrate the mixed dynamics from the enhancement task (i.e., from pixel domain) to take into account both pixel-wise and motion-related characteristics, for ensuring robust trajectory estimation and motion stabilization. Subsequently, we refine the warping masks to achieve high-quality full frame rendering. We further establish a synthetic dataset for training and evaluation regarding this emerging task. Comprehensive experiments have illustrated the superior performances of our UniFES method over 32 comparing baselines on both newly established synthetic and real-world datasets.
Mai Xu, Shengxi Li, Lai Jiang 0004
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Breaking the Multi-Enhancement Bottleneck: Domain-Consistent Quality Enhancement for Compressed Images
abstract
Quality enhancement methods have been widely integrated into visual communication pipelines to mitigate artifacts in compressed images. Ideally, these quality enhancement methods should perform robustly when applied to images that have already undergone prior enhancement during transmission. We refer to this scenario as multi-enhancement, which generalizes the well-known multi-generation scenario of image compression. Unfortunately, current quality enhancement methods suffer from severe degradation when applied in multi-enhancement.To address this challenge, we propose a novel adaptation method that transforms existing quality enhancement models into domain-consistent ones. Specifically, our method enhances a low-quality compressed image into a high-quality image within the natural domain during the first enhancement, and ensures that subsequent enhancements preserve this quality without further degradation. Extensive experiments validate the effectiveness of our method and show that various existing models can be successfully adapted to maintain both fidelity and perceptual quality in multi-enhancement scenarios.
Qunliang Xing, Mai Xu, Shengxi Li
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 DeepELIC: Deep encrypted lossy image compression network via compressive sensing unfolding
Fangyuan Gao, Yufan Deng, Xin Deng 0002, Zhenyu Guan 0002, Mai Xu
Pattern Recognit.5
2026 Compressed image super-resolution based on invertible degradation and restoration
Mai Xu, Lai Jiang 0004, Xin Deng 0002, Yue Zhang 0082, Yufan Liu 0001
Pattern Recognit.2
2026 A novel geometry-aware spatio-temporal network for multi-view video feature learning
Yue Zhang 0082, Mai Xu, Lai Jiang 0004, Xin Deng 0002, Si Liu 0001
Pattern Recognit.2
2026 A Novel Visible-Infrared Image Compression Framework for High-Value Target Protection
abstract
The joint compression of visible-infrared images is crucial for military and surveillance applications. The challenge lies in the protection of high-value targets (HVT) while maintaining high compression efficiency. This paper proposes a novel dual-stream compression framework that effectively addresses this challenge. In our framework, the sensitive HVT infrared signatures are concealed within the visible image stream, while residual infrared image is encoded separately. This dual-stream compression framework introduces three key innovations. 1) HVT protection: The HVT information is physically isolated and hidden within public visible images through a dedicated concealment stream; 2) Key-conditioned reconstruction: A novel decoding mechanism enables active camouflage by replacing HVTs with plausible background content when unauthorized access is detected; 3) Unified optimization: The framework integrates compression efficiency and HVT protection within an endto- end trainable network. Extensive experiments demonstrate that our approach achieves state-of-the-art compression performance while providing superior HVT protection, significantly outperforming traditional encrypt-then-compress methods. The code and weights are open-source athttps://github.com/eecoder-dyf/rgbir-compress.
Yufan Deng, Xin Deng 0002, Shengxi Li, Xiaowan Hu, Mai Xu
IEEE Signal Process. Lett.5
2026 MarsQE: Semantic-Informed Quality Enhancement for Compressed Martian Image
abstract
Lossy image compression is essential for Mars exploration missions, due to the limited bandwidth between Earth and Mars. However, the compression may introduce visual artifacts that complicate the geological analysis of the Martian surface. Existing quality enhancement approaches, primarily designed for Earth images, fall short for Martian images due to a lack of consideration for the unique Martian semantics. In response to this challenge, we conduct an in-depth analysis of Martian images, yielding two key insights based on semantics: the presence of texture similarities and the compact nature of texture representations in Martian images. Inspired by these findings, we introduce MarsQE, an innovative, semantic-informed, two-phase quality enhancement approach specifically designed for Martian images. The first phase involves the semantic-based matching of texture-similar reference images, and the second phase enhances image quality by transferring texture patterns from these reference images to the compressed image. We also develop a post-enhancement network to further reduce compression artifacts and achieve superior compression quality. Our extensive experiments demonstrate that MarsQE significantly outperforms existing approaches for Earth images, establishing a new benchmark for the quality enhancement on Martian images. The code is available at https://github.com/keriphLiu/MarsQE.
Chengfeng Liu, Mai Xu, Qunliang Xing
IEEE Trans. Circuits Syst. Video Technol.2
2026 SportSal: Hypernetwork-Based Saliency Prediction for Sports Videos
abstract
Saliency prediction is crucial for improving sports video processing efficiency, thereby providing an enriched viewing experience for a wide-ranging audience. However, there is a long-term absence of well-established eye-tracking dataset and learning-based approach, particularly tailored for sports videos. In this paper, we establish a large-scale eye-tracking dataset dubbed audio-visual sports (AVS). AVS consists of 1,000 high-quality sports videos with eye fixations from 60 participants. Through data analysis on AVS, we observe that human attention patterns exhibit significant variations based on the specific scene context of the sports. Motivated by our observations, we propose a sports-aware saliency prediction approach, named SportSal, which can adaptively predict saliency maps in a hyper manner. Specifically, a hypernetwork is introduced to learn sports-aware priors. Meanwhile, an audio-visual fusion (AVF) block is developed to effectively fuse features from the visual and audio backbones. Given the learned priors and fused audio-visual features, we propose the hyper deformable convolutional (HDC) block and the hyper upsampling (HU) block for dynamic feature extraction and upsampling, respectively. The two blocks are alternatingly connected to adaptively predict saliency maps. Experimental results show that our approach outperforms 21 state-of-the-art saliency prediction approaches over three sports video eye-tracking datasets. Finally, we demonstrate the application of our SportSal approach in perceptual video compression. The dataset and code will be available at https://github.com/WeNsHiJIe-19950103/SportSal.
Mai Xu, Shijie Wen, Lai Jiang 0004, Minglang Qiao, Shengxi Li
IEEE Trans. Circuits Syst. Video Technol.1
2026 Blur-Resistant Hyperspectral Image Super-Resolution via Dual-Degradation Fusion Model
abstract
The deep unfolding network represents a promising research avenue in fusion-based hyperspectral image super-resolution (HSI-SR). However, most current deep unfolding methodologies are anchored in idealized observation models, which overlook the degradation of the multispectral image (MSI), hindering their SR performance and practical applicability. To address this problem, this paper establishes a novel Dual-Degradation Fusion (D2-Fusion) model, which incorporates both HSI degradation and MSI blurring into the HSI-SR modeling process. Subsequently, we apply the second-order semismooth Newton algorithm to solve the optimization problem in D2-Fusion model. The solution steps are then mapped into an end-to-end trainable network, termed Blur-resistant Hyperspectral image Super-Resolution Network (BHSR-Net). To the best of our knowledge, the proposed network is the first successful attempt to consider MSI blurring artifacts in the HSI-SR task. It offers several distinct advantages: 1) The network structure maintains a strict mathematical correspondence with the optimization algorithm, ensuring each module retains strong physical interpretability; 2) The network exhibits superior SR performance and strong generalization ability on both standard and real-world scenarios across five datasets; 3) The network demonstrates excellent learning efficiency with a compact architecture, and its lightweight variant achieves comparable results with only 38K parameters. The code is available at https://github.com/Dou0405/BHSR-Net.
Mai Xu, Yongxuan Dou, Xin Deng 0002, Zhenwei Shi 0001
IEEE Trans. Image Process.1
2026 Machines Serve Human: A Novel Variable Human-Machine Collaborative Compression Framework
abstract
Human-machine collaborative compression has been receiving increasing research efforts for reducing image/video data, serving as the basis for both human perception and machine intelligence. Existing collaborative methods are dominantly built upon the de facto human-vision compression pipeline, witnessing deficiency on complexity and bit-rates when aggregating the machine-vision compression. Indeed, machine vision solely focuses on the core regions within the image/video, requiring much less information compared with the compressed information for human vision. In this paper, we thus set out the first successful attempt by a novel collaborative compression method based on the machine-vision-oriented compression, instead of human-vision pipeline. In other words, machine vision serves as the basis for human vision within collaborative compression. A plug-and-play variable bit-rate strategy is also developed for machine vision tasks. Then, we propose to progressively aggregate the semantics from the machine-vision compression, whilst seamlessly tailing the diffusion prior to restore high-fidelity details for human vision, thus named as diffusion-prior based feature compression for human and machine visions (Diff-FCHM). Experimental results verify the consistently superior performances of our Diff-FCHM, on both machine-vision and human-vision compression with remarkable margins. The source code is available at https://github.com/bblgbr/Diff-FCHM.
Zifu Zhang, Shengxi Li, Xiancheng Sun, Mai Xu, Zhengyuan Liu, Jingyuan Xia
IEEE Trans. Image Process.4
2026 Dual-Domain Visual Prompt Learning for Multi-Modal Medical Image Saliency Prediction
abstract
Medical image saliency prediction plays a pivotal role in emulating clinician visual attention to prioritize diagnostically critical regions. Current methods remain constrained by their spatial-domain dependency and limited cross-modality generalizability, neglecting frequency-domain patterns critical for subtle pathology detection while suffering from over-specialization in specific imaging modalities. Therefore, we propose a dual-domain visual prompt network (DVPNet) that integrates cross-modality generalization with spectral pattern awareness. On the one hand, DVPNet establishes a dataset prompt branch that dynamically modulates spatial feature encoding through modality-specific priors, allowing adaptive interpretation of heterogeneous medical imaging domains. On the other hand, a spatial-frequency hybrid prompt module employs learnable wavelet filters to decompose images into multi-scale spectral components, preserving low-frequency anatomical context while enhancing discriminative high-frequency biomarkers that are typically obscured in previous pixel-level analysis. By seamlessly integrating these complementary representations, DVPNet optimally synthesizes spatial and spectral evidence, enabling robust generalization across diverse medical imaging modalities while sustaining computational efficiency. Extensive experimental results on two distinct datasets demonstrate that the proposed method outperforms state-of-the-art approaches, showing superior saliency prediction performance and enhanced generalizability across medical contexts.
Mai Xu, Xiaowan Hu, Lai Jiang 0004
IEEE J. Biomed. Health Informatics2
2026 Few-Shot Pulmonary Vessel Segmentation Based on Tubular-Aware Prompt-Tuning
abstract
Segmentation of the pulmonary vessel from computed tomography (CT) images plays a crucial role in the diagnosis and treatment of various lung diseases. Although deep learning-based approaches have shown remarkable progress in recent years, their performance is often hindered by the lack of high-quality annotated datasets, in which the complex anatomy and morphology of pulmonary vessels make manual annotation challenging, time-consuming, and prone to errors. To address this, we propose PV25, the first dataset that features finely paired annotations of both pulmonary vessels and airways. Moreover, we propose TPNet, a novel tubular-aware prompt-tuning framework for pulmonary vessel segmentation under few-shot training with limited annotations. Specifically, based on an advanced and frozen segmentation backbone, TPNet proposes tunable encoding and decoding networks that learn tubular structures as transfer learning priors, bridging the gap between the source and target pulmonary vessel domains. Specifically, TPNet is built in an encoder-decoder manner, including the fixed segmentation backbone, tunable encoding and decoding networks. In encoding stage, the Morphology-Driven Region Growing (MDRG) module is developed to leverage the tubular connectivity of vessels to guide the network in capturing fine-grained features of pulmonary vessels. In decoding stage, the Cross-Correlation Guidance (CCG) module is introduced to integrate multi-scale correlations between airway and vessel structures in a coarse-to-fine manner. Extensive experiments conducted on multiple datasets demonstrate that TPNet achieves state-of-the-art performance in pulmonary vessel segmentation under limited training data. Besides, TPNet shows strong performance in related tasks such as airway segmentation and artery-vein classification, highlighting its robustness and versatility.
Zijian Gao, Lai Jiang 0004, Sukun Tian, Yuchun Sun, Mai Xu, Liyuan Tao
IEEE Trans. Medical Imaging6
2025 Spherical Manifold Guided Diffusion Model for Panoramic Image Generation
abstract
Panoramic image essentially acts as a pivotal role in emerging virtual reality and augmented reality scenarios; however, the generation of panoramic images are essentially challenging due to the intrinsic spherical geometry and spherical distortions caused by equirectangular projection (ERP). To address this, we start from the very basics of S2manifold inherent to panoramic images, and propose a novel spherical manifold convolution (SMConv) on S2manifold. Based on the SMConv operation, we propose a spherical manifold guided diffusion (SMGD) model for text-conditioned panoramic image generation, which can well accommodate the spherical geometry during generation. We further develop a novel evaluation method by calculating grouped Fréchet inception distance (FID) on cube-map projections, which can well reflect the quality of generated panoramic images, compared to existing methods that randomly crop ERP-distorted content. Experiment results demonstrate that our SMGD model achieves the state-of-the-art generation quality and accuracy, whilst retaining the shortest sampling time in the text-conditioned panoramic image generation task. Codes are publicly available at https://github.com/chronos123/SMGD.
Xiancheng Sun, Mai Xu, Shengxi Li, Senmao Ma, Xin Deng 0002, Lai Jiang 0004
CVPR2
2025 Uncover Treasures in DCT: Advancing JPEG Quality Enhancement by Exploiting Latent Correlations
abstract
Joint Photographic Experts Group (JPEG) achieves data compression by quantizing Discrete Cosine Transform (DCT) coefficients, which inevitably introduces compression artifacts. Most existing JPEG quality enhancement methods operate in the pixel domain, suffering from the high computational costs of decoding. Consequently, direct enhancement of JPEG images in the DCT domain has gained increasing attention. However, current DCT-domain methods often exhibit limited performance. To address this challenge, we identify two critical types of correlations within the DCT coefficients of JPEG images. Building on this insight, we propose an Advanced DCT-domain JPEG Quality Enhancement (AJQE) method that fully exploits these correlations. The AJQE method enables the adaptation of numerous well-established pixel-domain models to the DCT domain, achieving superior performance with reduced computational complexity. Compared to the pixel-domain counterparts, the DCT-domain models derived by our method demonstrate a 0.35 dB improvement in PSNR and a 60.5% increase in enhancement throughput on average.
Qunliang Xing, Mai Xu, Minglang Qiao
ICCV3
2025 Deformable Spherical Geometry Transformer For Panoramic Semantic Segmentation
abstract
The increasing availability of 360° images has created a demand for effective Panoramic Semantic Segmentation (PASS) to enable comprehensive scene understanding. However, the spherical nature of 360° image introduces significant spatial distortions due to Equirectangular Projection (ERP), making it challenging for traditional 2D methods, which are designed for Euclidean spaces. Existing PASS methods typically mitigate these distortions through developing spherical-to-tangent polyhedron transformations or special-ized convolutional structures. Nevertheless, these approaches still struggle to preserve the spherical geometry and fail to adequately capture the semantic context of 360° images. In this paper, we propose a Deformable Spherical Geometry Transformer (DSGT) network that adapts to spherical distortions through a local-global self-attention mechanism. The local self-attention module captures local semantic information to alleviate distortions, while the global self-attention module integrates spherical geometric priors to enhance predictions. Experimental results on the Stanford2D3D panoramic dataset demonstrate that DSGT outperforms state-of-the-art PASS methods.
Boyang Lan, Li Yang 0014, Mai Xu, Lai Jiang 0004, Yufeng Wang 0004
ICIP3
2025 Quality Control For HEVC: A Deep Reinforcement Learning Approach
abstract
In video coding, large quality fluctuations exist in compressed videos, significantly degrading their quality of experience (QoE). Most works in literature focus on controlling bit-rates, however, paying few attention on reducing the quality fluctuations. In this paper, we propose a novel deep reinforcement learning (DRL) method for quality control in video coding. Specifically, we first propose the formulation of quality control, which targets at both controlling the target quality and reducing fluctuations. Then, we solve the quality control formulation by proposing a DRL method, in which the DRL elements are modeled by considering the features of both current frame and previous encoded frames. Specifically, for the DRL elements, we take the encoding information, content complexity and hidden features of long short-term memory (LSTM) as the state of DRL, and the selection of quantization parameters (QP) as the action of DRL. Subsequently, an algorithm, based on proximal policy optimization, is utilized to update our DRL model for decision-making on the actions of QP selection. In this way, the videos can be compressed under given and constant quality. We implement our DRL-based quality control method on the standard of high efficiency video coding (HEVC) with the HM 16.15 platform, and experimental results show that our method achieves the state-of-the-art performance on both quality control accuracy and fluctuations, in comparison with other quality control baselines.
Mai Xu, Lai Jiang 0004, Shengxi Li, Xin Deng 0002
ICME3
2025 SANE: Enhancing Large-scale Scene Representation with Semantic-aware NeRF Experts
abstract
We propose the Semantic-aware NeRF Experts (SANE), which fully exploits the intrinsic characteristics of large-scale scenes, including semantics and material features, to achieve high-quality novel view synthesis results and provide accurate 3D semantic information. SANE begins by building a semantic Mixture of Experts (MoE), utilizing a learnable gating network to semantically partition the scene into blocks for corresponding NeRF experts. We then develop a semantic volume rendering scheme that integrates discrete semantics into the end-to-end differentiable process of NeRF, enabling refined semantic labeling of each scene point. Additionally, we implement a dual-implicit encoding strategy: intra-block encoding captures lighting variations across viewpoints, while inter-block one captures texture features among different semantic objects. Experiments on benchmark datasets show that SANE delivers higher-quality scene representations and effective semantic decomposition for downstream tasks, such as precise editing of large-scale scenes based on semantics.
Zesheng Wang 0002, Yufeng Wang 0004, Shuangkang Fang, Dacheng Qi, Shengxi Li, Mai Xu, Wenrui Ding
ICME7
2025 Spherical-Nested Diffusion Model for Panoramic Image Outpainting
abstract
Panoramic image outpainting acts as a pivotal role in immersive content generation, allowing for seamless restoration and completion of panoramic content. Given the fact that the majority of generative outpainting solutions operates on planar images, existing methods for panoramic images address the sphere nature by soft regularisation during the end-to-end learning, which still fails to fully exploit the spherical content. In this paper, we set out the first attempt to impose the sphere nature in the design of diffusion model, such that the panoramic format is intrinsically ensured during the learning procedure, named as spherical-nested diffusion (SpND) model. This is achieved by employing spherical noise in the diffusion process to address the structural prior, together with a newly proposed spherical deformable convolution (SDC) module to intrinsically learn the panoramic knowledge. Upon this, the proposed method is effectively integrated into a pre-trained diffusion model, outperforming existing state-of-the-art methods for panoramic image outpainting. In particular, our SpND method reduces the FID values by more than 50\% against the state-of-the-art PanoDiffusion method. Codes are publicly available at \url{https://github.com/chronos123/SpND}.
Xiancheng Sun, Senmao Ma, Shengxi Li, Mai Xu, Jingyuan Xia, Lai Jiang 0004, Xin Deng 0002
ICML4
2025 Collateral Circulation Guided Multi-Modality Fusion Network for Postoperative Infarct Prediction
Lisong Dai, Heming Dong, Lai Jiang 0004, Mai Xu, Shengxi Li
MICCAI (15)6
2025 Blind Multimodal Quality Assessment of Low-Light Images
Miaohui Wang, Zhuowei Xu, Mai Xu, Weisi Lin
Int. J. Comput. Vis.3
2025 DeepSN-Net: Deep Semi-Smooth Newton Driven Network for Blind Image Restoration
abstract
The deep unfolding network represents a promising research avenue in image restoration. However, most current deep unfolding methodologies are anchored in first-order optimization algorithms, which suffer from sluggish convergence speed and unsatisfactory learning efficiency. In this paper, to address this issue, we first formulate an improved second-order semi-smooth Newton (ISN) algorithm, transforming the original nonlinear equations into an optimization problem amenable to network implementation. After that, we propose an innovative network architecture based on the ISN algorithm for blind image restoration, namely DeepSN-Net. To the best of our knowledge, DeepSN-Net is the first successful endeavor to design a second-order deep unfolding network for image restoration, which fills the blank of this area. Furthermore, it offers several distinct advantages: 1) DeepSN-Net provides a unified framework to a variety of image restoration tasks in both synthetic and real-world contexts, without imposing constraints on the degradation conditions. 2) The network architecture is meticulously aligned with the ISN algorithm, ensuring that each module possesses robust physical interpretability. 3) The network exhibits high learning efficiency, superior restoration accuracy and good generalization ability across 11 datasets on three typical restoration tasks. The success of DeepSN-Net on image restoration may ignite many subsequent works centered around the second-order optimization algorithms, which is good for the community.
Xin Deng 0002, Lai Jiang 0004, Jingyuan Xia, Mai Xu
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Patch Inverter: A Novel Block-Wise GAN Inversion Method for Arbitrary Image Resolutions
abstract
Generative adversarial networks (GANs) have achieved remarkable progress in generating realistic images from merely small dimensions, which essentially establishes the latent generating space by rich semantics. GAN inversion thus aims at mapping real-world images back into the latent space, allowing for the access of semantics from images. However, existing GAN inversion methods can only invert images with fixed resolutions; this significantly restricts the representation capability in real-world scenarios. To address this issue, we propose to invert images by patches, thus named as patch inverter, which is the first attempt in terms of block-wise inversion for arbitrary resolutions. More specifically, we develop the padding-free operation to ensure the continuity across patches, and analyse the intrinsic mismatch within the inversion procedure. To relieve the mismatch, we propose a shifted convolution operation, which retains the continuity across image patches and simultaneously enlarges the receptive field for each convolution layer. We further propose the reciprocal loss to regularize the inverted latent codes to reside on the original latent generating space, such that the rich semantics can be maximally preserved. Experimental results have demonstrated that our patch inverter is able to accurately invert images with arbitrary resolutions, whilst representing precise and rich image semantics in real-world scenarios.
Mai Xu, Shengxi Li, Zhenyu Guan 0002
IEEE Signal Process. Lett.2
2025 Continuous Patch Stitching for Block-Wise Image Compression
abstract
Most recently, learned image compression methods have outpaced traditional hand-crafted standard codecs. However, their inference typically requires to input the whole image at the cost of heavy computing resources, especially for high-resolution image compression; otherwise, the block artefact can exist when compressed by blocks within existing learned image compression methods. To address this issue, we propose a novel continuous patch stitching (CPS) framework for block-wise image compression that is able to achieve seamlessly patch stitching and mathematically eliminate block artefact, thus capable of significantly reducing the required computing resources when compressing images. More specifically, the proposed CPS framework is achieved by padding-free operations throughout, with a newly established parallel overlapping stitching strategy to provide a general upper bound for ensuring the continuity. Upon this, we further propose functional residual blocks with even-sized kernels to achieve down-sampling and up-sampling, together with bottleneck residual blocks retaining feature size to increase network depth. Experimental results demonstrate that our CPS framework achieves the state-of-the-art performance against existing baselines, whilst requiring less than half of computing resources of existing models. The source code and trained models are available athttps://github.com/bblgbr/SPL-CPS.
Zifu Zhang, Shengxi Li, Henan Liu, Mai Xu, Ce Zhu
IEEE Signal Process. Lett.4
2025 A Deep Transformer-Based Fast CU Partition Approach for Inter-Mode VVC
abstract
The latest versatile video coding (VVC) standard proposed by the Joint Video Exploration Team (JVET) has significantly improved coding efficiency compared to that of its predecessor, while introducing an extremely higher computational complexity by $6\sim 26$ times. The quad-tree plus multi-type tree (QTMT)-based coding unit (CU) partition accounts for most of the encoding time in VVC encoding. This paper proposes a data-driven fast CU partition approach based on an efficient Transformer model to accelerate VVC inter-coding. First, we establish a large-scale database for inter-mode VVC, comprising diverse CU partition patterns from more than 800 raw video sequences across various resolutions and contents. Next, we propose a deep neural network model with a Transformer-based temporal topology for predicting the CU partition, named as TCP-Net, which is adaptive to the group of pictures (GOP) hierarchy in VVC. Then, we design a two-stage structured output for TCP-Net, reflecting both the locations of CU edges and the split modes of all possible CUs. Accordingly, we develop a dual-supervised optimization mechanism to train the TCP-Net model with improved accuracy. The experimental results have verified that our approach can reduce the encoding time by $46.89\sim 55.91$ % with negligible rate-distortion (RD) degradation, outperforming other state-of-the-art approaches.
Tianyi Li 0004, Mai Xu, Ying Chen 0011, Kai Li 0012
IEEE Trans. Image Process.2
2025 Hierarchical Semantic Compression for Consistent Image Semantic Restoration
abstract
The emerging semantic compression has been receiving increasing research efforts most recently, capable of achieving high fidelity restoration during compression, even at extremely low bitrates. However, existing semantic compression methods typically combine standard pipelines with either pre-defined or high-dimensional semantics, thus suffering from deficiency in compression. To address this issue, we propose a novel hierarchical semantic compression (HSC) framework that purely operates within intrinsic semantic spaces from generative models, which is able to achieve efficient compression for consistent semantic restoration. More specifically, we first analyse the entropy models for the semantic compression, which motivates us to employ a hierarchical architecture based on a newly developed general inversion encoder. Then, we propose the feature compression network (FCN) and semantic compression network (SCN), such that the middle-level semantic feature and core semantics are hierarchically compressed to restore both accuracy and consistency of image semantics, via an entropy model progressively shared by channel-wise context. Experimental results demonstrate that the proposed HSC framework achieves the state-of-the-art performance on subjective quality and consistency for human vision, together with superior performances on machine vision tasks given compressed bitstreams. This essentially coincides with human visual system in understanding images, thus providing a new framework for future image/video compression paradigms. The source code and trained models are available at https://github.com/bblgbr/HSC-TIP2025.
Shengxi Li, Zifu Zhang, Mai Xu, Lai Jiang 0004, Yufan Liu 0001, Ce Zhu
IEEE Trans. Image Process.3
2025 Spherical Patch Generative Adversarial Net for Unconditional Panoramic Image Generation
abstract
Recent advancements in virtual reality (VR) and augmented reality (AR) have popularised the emerging panoramic content for the immersive visual experience. The difficulty in acquisition and display of 360° format further highlights the necessity of unconditional panoramic image generation. Existing methods essentially generate planar images mapped from panoramic images, and fail to address the deformation and closed-loop characteristics when inverted back to the panoramic images. Thus leading to the generation of pseudo-panoramic content. This paper aims to directly generate spherical content, in a patch-by-patch style; besides computation friendly, this promises the anywhere continuity on the panoramic image and proper accommodation of panoramic deformation. More specifically, we first propose a novel spherical patch convolution (SPConv) that operates on the local spherical patch, which naturally addresses the deformation of panoramic content. We then propose our spherical patch generative adversarial net (SP-GAN) that consists of spherical local embedding (SLE) and spherical content synthesiser (SCS) modules, which seamlessly incorporate our SPConv so as to generate continuous panoramic patches. To the best of our knowledge, the proposed SP-GAN is the first successful attempt to accommodate the spherical distortion for closed-loop panoramic image generation in a patch-by-patch manner. The experimental results, with human-rated evaluations, have verified the consistently superior performances for unconditional panoramic image generation, from the perspectives of generation quality, computational memory, and generalisation to various resolutions. Codes are publicly available at https://github.com/chronos123/SP-GAN.
Mai Xu, Xiancheng Sun, Shengxi Li, Lai Jiang 0004, Jingyuan Xia, Xin Deng 0002
IEEE Trans. Image Process.1
2025 Deep Semi-Smooth Newton-Driven Unfolding Network for Multi-Modal Image Super-Resolution
abstract
Deep unfolding has emerged as a powerful solution for Multi-modal Image Super-Resolution (MISR) through strategic integration of cross-modal priors in network architecture. However, current deep unfolding approaches rely on first-order optimization, which exhibit limitations in learning efficiency and reconstruction accuracy. In this paper, to overcome these limitations, we propose a novel Semi-smooth Newton driven Unfolding network for MISR, namely SNUM-Net. Specifically, we first develop a Semi-smooth Newton-driven MISR (SNM) algorithm that establishes a theoretical foundation for our approach. Then, we unfold the iterative solution of SNM into a novel network. To the best of our knowledge, the SNUM-Net is the first successful attempt to design a deep unfolding MISR network based on second-order optimization algorithm. Compared to existing methods, the SNUM-Net demonstrates three main advantages. 1) Universal paradigm: the SNUM-Net provides a unified paradigm for diverse MISR tasks without requiring scenario-specific constraints; 2) Explainable framework: the network preserves a mathematical correspondence with the SNM algorithm, ensuring that the topological relationships between modules are well explainable; 3) Superior performance: comprehensive evaluations across 10 datasets spanning 3 MISR tasks demonstrate the network's exceptional reconstruction accuracy and generalization capability. The software codes are available at https://github.com/pandazcx/SNUM-Net.
Xin Deng 0002, Yongxuan Dou, Mai Xu
IEEE Trans. Image Process.5
2025 Recruiting Teacher IF Modality for Nephropathy Diagnosis: A Customized Distillation Method With Attention-Based Diffusion Network
abstract
The joint use of multiple modalities for medical image processing has been widely studied in recent years. The fusion of information from different modalities has demonstrated the performance improvement for a lot of medical tasks. For nephropathy diagnosis, immunofluorescence (IF) is one of the most widely-used multi-modality medical images due to its ease of acquisition and the effectiveness for certain nephropathy. However, the existing methods mainly assume different modalities have the equal effect on the diagnosis task, failing to exploit multi-modality knowledge in details. To avoid this disadvantage, this paper proposes a novel customized multi-teacher knowledge distillation framework to transfer knowledge from the trained single-modality teacher networks to a multi-modality student network. Specifically, a new attention-based diffusion network is developed for IF based diagnosis, considering global, local, and modality attention. Besides, a teacher recruitment module and diffusion-aware distillation loss are developed to learn to select the effective teacher networks based on the medical priors of the input IF sequence. The experimental results in the test and external datasets show that the proposed method has a better nephropathy diagnosis performance and generalizability, in comparison with the state-of-the-art methods.
Mai Xu, Lai Jiang 0004, Yibing Fu, Xin Deng 0002, Shengxi Li
IEEE Trans. Medical Imaging1
2025 MDSC-Net: Multi-Modal Discriminative Sparse Coding Driven RGB-D Classification Network
abstract
In this paper, we propose a novel sparsity-driven deep neural network to solve the RGB-D image classification problem. Different from existing classification networks, our network architecture is designed by drawing inspirations from a new proposed multi-modal discriminative sparse coding (MDSC) model. The key feature of this model is that it can gradually separate the discriminative and non-discriminative features in RGB-D images in a coarse-to-fine manner. Only the discriminative features are integrated and refined for classification, while the non-discriminative features are discarded, to improve the classification accuracy and efficiency. Derived from the MDSC model, the proposed network is composed of three modules, i.e., the shared feature extraction (SFE) module, discriminative feature refinement (DFR) module, and classification module. The architecture of each module is derived from the optimization solution in the MDSC model. To the best of our knowledge, this is the first time a fully sparsity-driven network has been proposed for RGB-D image classification. Extensive results verify the effectiveness of our method on different RGB-D image datasets.
Xin Deng 0002, Yibing Fu, Mai Xu, Shengxi Li
IEEE Trans. Multim.4
2025 MVL-Net: Pairwise Learning for Multi-View Multiple People Labelling
abstract
In the multi-view domain, it is challenging to correctly label multiple people across viewpoints because of occlusions, visual ambiguities, appearance variation, etc. Deep learning, although having witnessed remarkable success in computer vision tasks, still remains underexplored for the multi-view labelling task, due to the lack of labelled multi-view datasets. In this paper, we propose a novel end-to-end deep neural network named Multi-View Labelling network (MVL-net) that addresses this issue. To overcome the dataset shortage, a large-scale multi-view dataset is generated by combining 3D human models and panoramic backgrounds, along with human poses and realistic rendering. In the proposed MVL-net, we first incorporate Transformer blocks to capture the non-local information for multi-view feature extraction. A matching net is then introduced to achieve multiple people labelling, by predicting matching confidence scores for pairwise instances from two views, thus addressing the problem of the unknown number of people when labelling across views. An additional geometry feature obtained from the epipolar geometry is integrated to leverage multi-view cues during training. To the best of our knowledge, the MVL-net is the first work using deep learning to train a multi-view labelling network. Comprehensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed method, which outperforms the existing state-of-the-art approaches.
Yue Zhang 0082, Akin Caliskan, Mai Xu, Adrian Hilton 0001, Jean-Yves Guillemaut
IEEE Trans. Multim.3
2024 Enhancing Quality of Compressed Images by Mitigating Enhancement Bias Towards Compression Domain
abstract
Existing quality enhancement methods for compressed images focus on aligning the enhancement domain with the raw domain to yield realistic images. However, these methods exhibit a pervasive enhancement bias towards the compression domain, inadvertently regarding it as more realistic than the raw domain. This bias makes enhanced images closely resemble their compressed counterparts, thus degrading their perceptual quality. In this paper, we propose a simple yet effective method to mitigate this bias and enhance the quality of compressed images. Our method employs a conditional discriminator with the compressed image as a key condition, and then incorporates a domain-divergence regularization to actively distance the enhancement domain from the compression domain. Through this dual strategy, our method enables the discrimination against the compression domain, and brings the enhancement domain closer to the raw domain. Comprehensive quality evaluations confirm the superiority of our method over other state-of-the-art methods without incurring inference overheads.
Qunliang Xing, Mai Xu, Shengxi Li, Xin Deng 0002, Meisong Zheng, Huaida Liu, Ying Chen 0011
CVPR2
2024 Saliency Prediction of Sports Videos: A Large-Scale Database and a Self-Adaptive Approach
abstract
Predicting video saliency is crucial for improving sports video processing efficiency, thereby providing an enriched viewing experience for a wide-ranging audience. However, there is a long-term absence of well-established eye-tracking database and learning-based approach, particularly tailored for sports videos. In this paper, we establish a large-scale eye-tracking database dubbed audio-visual sports (AVS). AVS consists of 1,000 high-quality sports videos with eye fixations from 60 participants. Through the data analysis on AVS, we observe that human attention patterns exhibit significant variations based on the specific scene context of the sports. Motivated by this, we propose a sport-aware audiovisual saliency model, which can adaptively learn the scene context in a hyper manner. Specifically, a new audio-visual fusion (AVF) block is developed to effectively fuse features from the visual and audio backbone. After that, a hyper network is introduced to learn sport-aware priors, which are then adopted to guide the self-adaptive saliency predictor for predicting saliency map. Experimental results demonstrate that our approach outperforms other state-of-the-art saliency prediction models over the only two sports video eye-tracking databases.
Minglang Qiao, Mai Xu, Shijie Wen, Lai Jiang 0004, Shengxi Li, Yunjin Chen, Leonid Sigal
ICASSP2
2024 SN-NET: Semismooth Newton Driven Lightweight Network for Real-World Image Denoising
abstract
Semismooth Newton is a powerful tool to tackle the regularization problems in image restoration. Compared to other optimization methods such as alternating direction method of multipliers (ADMM), the semismooth Newton method exhibits faster convergence and greater efficiency, since the non-linear and coupling system is solved simultaneously. However, its performance relies heavily on the handcrafted parameters and efficient solvers for calculating Newton steps. To tackle this issue, we first develop an improved semismooth Newton method, in which we turn the original nonlinear system solving problem into a network-friendly convex optimization problem. After that, we unfold it into a novel network namely SN-Net. We apply SN-Net on the most fundamental image denoising task, which shows great advantages in the following two aspects. (1) The network is quite lightweight, i.e., the number of network parameters is only 86 KB. (2) The network structure exhibits strong interpretability. To the best of our knowledge, the $\mathrm{SN}-\mathrm{Net}$ is the first attempt to successfully map the semismooth Newton method to a learnable network. The great success of it on image denoising may inspire many potential works on other image restoration tasks. The code and the pre-trained models are released at https://github.com/pandazcx/SN-Net.
Xin Deng 0002, Hongpeng Sun, Mai Xu
ICIP5
2024 Hybrid Single Input and Multiple Output Method For Compressing Features Towards Machine Vision Tasks
abstract
With the advance of deep learning in the BigData era, image/video coding for machines (VCM) as called for proposals by the moving picture experts group (MPEG) now becomes the pivotal technique for extensive intelligent vision tasks. However, existing VCM methods typically focus on compressing features independently at each scale, ignoring the redundancy of features across multiple scales. This paper thus introduces a simple yet effective architecture called hybrid single input and multiple output (H-SIMO) for VCM, which can significantly reduce the redundancy across scales of features. More specifically, as the pyramid structure is commonly employed for localising multi-scale objects, our H-SIMO method proposes to compress all features by inputting a single-scale feature while retaining the ability to decompress all the features. Moreover, an entropy model is seamlessly integrated into the training process to efficiently reduce the statistical redundancy of features. During the testing phase, the hybrid coding method, in conjunction with the versatile video coding (VVC), is employed to compress the features from both images and videos. We comprehensively evaluate the performance of our H-SIMO method in two standard machine vision tasks: object detection and instance segmentation, in which the experimental results verify the superior performances of our H-SIMO method.
Zifu Zhang, Shengxi Li, Mai Xu, Zhenyu Guan 0002, Zhuoyi Lv
ICIP4
2024 Causal Context Adjustment Loss for Learned Image Compression
abstract
In recent years, learned image compression (LIC) technologies have surpassed conventional methods notably in terms of rate-distortion (RD) performance. Most present learned techniques are VAE-based with an autoregressive entropy model, which obviously promotes the RD performance by utilizing the decoded causal context. However, extant methods are highly dependent on the fixed hand-crafted causal context. The question of how to guide the auto-encoder to generate a more effective causal context benefit for the autoregressive entropy models is worth exploring. In this paper, we make the first attempt in investigating the way to explicitly adjust the causal context with our proposed Causal Context Adjustment loss (CCA-loss). By imposing the CCA-loss, we enable the neural network to spontaneously adjust important information into the early stage of the autoregressive entropy model. Furthermore, as transformer technology develops remarkably, variants of which have been adopted by many state-of-the-art (SOTA) LIC techniques. The existing computing devices have not adapted the calculation of the attention mechanism well, which leads to a burden on computation quantity and inference latency. To overcome it, we establish a convolutional neural network (CNN) image compression model and adopt the unevenly channel-wise grouped strategy for high efficiency. Ultimately, the proposed CNN-based LIC network trained with our Causal Context Adjustment loss attains a great trade-off between inference latency and rate-distortion performance.
Minghao Han, Shiyin Jiang, Shengxi Li, Xin Deng 0002, Mai Xu, Ce Zhu, Shuhang Gu
NeurIPS5
2024 Physical layer signal processing for XR communications and systems
Yongpeng Wu 0001, Mai Xu, Guangtao Zhai, Wenjun Zhang 0001
Sci. China Inf. Sci.2
2024 Joint Learning of Audio-Visual Saliency Prediction and Sound Source Localization on Multi-face Videos
Minglang Qiao, Yufan Liu 0001, Mai Xu, Xin Deng 0002, Bing Li 0001, Weiming Hu 0004, Ali Borji
Int. J. Comput. Vis.3
2024 CrossHomo: Cross-Modality and Cross-Resolution Homography Estimation
abstract
Multi-modal homography estimation aims to spatially align the images from different modalities, which is quite challenging since both the image content and resolution are variant across modalities. In this paper, we introduce a novel framework namely CrossHomo to tackle this challenging problem. Our framework is motivated by two interesting findings which demonstrate the mutual benefits between image super-resolution and homography estimation. Based on these findings, we design a flexible multi-level homography estimation network to align the multi-modal images in a coarse-to-fine manner. Each level is composed of a multi-modal image super-resolution (MISR) module to shrink the resolution gap between different modalities, followed by a multi-modal homography estimation (MHE) module to predict the homography matrix. To the best of our knowledge, CrossHomo is the first attempt to address the homography estimation problem with both modality and resolution discrepancy. Extensive experimental results show that our CrossHomo can achieve high registration accuracy on various multi-modal datasets with different resolution gaps. In addition, the network has high efficiency in terms of both model complexity and running speed.
Xin Deng 0002, Enpeng Liu, Shengxi Li, Shuhang Gu, Mai Xu
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Deep$\mathrm {M^{2}}$M2CDL: Deep Multi-Scale Multi-Modal Convolutional Dictionary Learning Network
abstract
For multi-modal image processing, network interpretability is essential due to the complicated dependency across modalities. Recently, a promising research direction for interpretable network is to incorporate dictionary learning into deep learning through unfolding strategy. However, the existing multi-modal dictionary learning models are both single-layer and single-scale, which restricts the representation ability. In this paper, we first introduce a multi-scale multi-modal convolutional dictionary learning (M2CDL) model, which is performed in a multi-layer strategy, to associate different image modalities in a coarse-to-fine manner. Then, we propose a unified framework namely DeepM2CDL derived from the M2CDL model for both multi-modal image restoration (MIR) and multi-modal image fusion (MIF) tasks. The network architecture of DeepM2CDL fully matches the optimization steps of the M2CDL model, which makes each network module with good interpretability. Different from handcrafted priors, both the dictionary and sparse feature priors are learned through the network. The performance of the proposed DeepM2CDL is evaluated on a wide variety of MIR and MIF tasks, which shows the superiority of it over many state-of-the-art methods both quantitatively and qualitatively. In addition, we also visualize the multi-modal sparse features and dictionary filters learned from the network, which demonstrates the good interpretability of the DeepM2CDL network.
Xin Deng 0002, Fangyuan Gao, Xiancheng Sun, Mai Xu
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Assessing Face Image Quality: A Large-Scale Database and a Transformer Method
abstract
The amount of face images has been witnessing an explosive increase in the last decade, where various distortions inevitably exist on transmitted or stored face images. The distortions lead to visible and undesirable degradation on face images, affecting their quality of experience (QoE). To address this issue, this paper proposes a novel Transformer-based method for quality assessment on face images (named as TransFQA). Specifically, we first establish a large-scale face image quality assessment (FIQA) database, which includes 42,125 face images with diversifying content at different distortion types. Through an extensive crowdsource study, we obtain 712,808 subjective scores, which to the best of our knowledge contribute to the largest database for assessing the quality of face images. Furthermore, by investigating the established database, we comprehensively analyze the impacts of distortion types and facial components (FCs) on the overall image quality. Accordingly, we propose the TransFQA method, in which the FC-guided Transformer network (FT-Net) is developed to integrate the global context, face region and FC detailed features via a new progressive attention mechanism. Then, a distortion-specific prediction network (DP-Net) is designed to weight different distortions and accurately predict final quality scores. Finally, the experiments comprehensively verify that our TransFQA method significantly outperforms other state-of-the-art methods for quality assessment on face images.
Shengxi Li, Mai Xu, Li Yang 0014, Xiaofei Wang 0004
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 HyperSOR: Context-Aware Graph Hypernetwork for Salient Object Ranking
abstract
Salient object ranking (SOR) aims to segment salient objects in an image and simultaneously predict their saliency rankings, according to the shifted human attention over different objects. The existing SOR approaches mainly focus on object-based attention, e.g., the semantic and appearance of object. However, we find that the scene context plays a vital role in SOR, in which the saliency ranking of the same object varies a lot at different scenes. In this paper, we thus make the first attempt towards explicitly learning scene context for SOR. Specifically, we establish a large-scale SOR dataset of 24,373 images with rich context annotations, i.e., scene graphs, segmentation, and saliency rankings. Inspired by the data analysis on our dataset, we propose a novel graph hypernetwork, named HyperSOR, for context-aware SOR. In HyperSOR, an initial graph module is developed to segment objects and construct an initial graph by considering both geometry and semantic information. Then, a scene graph generation module with multi-path graph attention mechanism is designed to learn semantic relationships among objects based on the initial graph. Finally, a saliency ranking prediction module dynamically adopts the learned scene context through a novel graph hypernetwork, for inferring the saliency rankings. Experimental results show that our HyperSOR can significantly improve the performance of SOR.
Minglang Qiao, Mai Xu, Lai Jiang 0004, Shijie Wen, Yunjin Chen, Leonid Sigal
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Extremely Low Bit-Rate Image Compression via Invertible Image Generation
abstract
Image compression at extremely low bit-rates has always been a challenging task in bandwidth limited scenarios, such as aerospace and deep-sea explorations. Recent years have seen great success of deep learning in image compression, however, few of them are specially designed for extremely low bit-rate conditions. To solve this issue, in this paper, we propose a novel invertible image generation based framework for extremely low bit-rate image compression. The proposed framework is composed of three modules, including an invertible image generation (IIG) module, a generated image compression (GIC) module and a compressed image adjustment (CIA) module. The role of IIG module is to generate a compression-friendly image from the original image. In the IIG module, image generation and restoration are modelled as two mutually reversible processes to avoid the information loss. After the IIG module, the GIC module is employed to compress the generated images to save the coding bit-rates. After that, the CIA module is used to shrink the quality gap between the compressed generated image and the un-compressed image. Finally, the image from the CIA module is sent back to the IIG module to restore the original image. The experimental results on three different datasets show that the proposed framework achieves state-of-the-art performance in image compression with extremely low bit-rates. We also extend the proposed framework to feature compression towards object detection, which saves 90% bit-rates than the VVC standard with the same detection accuracy.
Fangyuan Gao, Xin Deng 0002, Junpeng Jing, Mai Xu
IEEE Trans. Circuits Syst. Video Technol.5
2024 Proposal With Alignment: A Bi-Directional Transformer for 360° Video Viewport Proposal
abstract
People normally watch 360 ° videos through a head-mounted display, inside which only the content of viewports can be seen. Therefore, viewport proposal, referring to detecting potential viewport candidates, plays an important role in many 360 ° video processing tasks. In this paper, we advance the viewport proposal by further aligning the predicted viewports across frames for individual subject. This provides a better methodology and a deeper perspective to learn the human perceptual behaviours on 360 ° videos. Specifically, we first analyze three 360 ° video datasets and obtain several findings on human consistency, objectness and motion of viewports. Inspired by these findings, we propose a bi-directional transformer approach, named BiT, for 360 ° video viewport proposal and alignment. Specifically, BiT is composed of a multi-level residual module, a bi-directional encoder-decoder module and a spherical matching module. This way, the viewports can be well proposed and aligned via considering multi-level, bi-directional and non-local information. Moreover, the aligned viewports by BiT are used to refine the viewports and improve viewport proposal accuracy in return. Finally, we validate that our BiT approach is superior on viewport proposal, compared with the state-of-the-art approaches. Besides, the aligned viewports from BiT is verified to be effective in multiple applications, such as saliency prediction, trajectory prediction and perceptual video compression.
Mai Xu, Lai Jiang 0004, Xin Deng 0002, Gaoxing Chen, Leonid Sigal
IEEE Trans. Circuits Syst. Video Technol.2
2024 Saliency Prediction on Mobile Videos: A Fixation Mapping-Based Dataset and A Transformer Approach
abstract
With the booming development of smart devices, mobile videos have drawn broad interest when humans surf social media. Different from traditional long-form videos, mobile videos are featured with uncertain human attention behavior so far owing to the specific displaying mode, thus promoting the research on saliency prediction for mobile videos. Unfortunately, the current eye-tracking experiments are not applicable for mobile videos, since the stationary eye-tracker and eye fixation acquisition are dedicated to the videos presented on computers. To tackle this issue, we propose performing the wearable eye-tracker to record viewers’ egocentric fixations and then devising a fixation mapping technique to project the eye fixations from egocentric videos onto mobile videos. Resorting to this technique, the large-scale mobile video saliency (MVS) dataset is established, including 1,007 mobile videos and 5,935,927 fixations. Given this dataset, we exhaustively analyze the characteristics of subjects’ fixations and obtain two findings. Based on the MVS dataset and these findings, we propose a saliency prediction approach on mobile videos upon Video Swin Transformer (MVFormer), wherein long-range spatio-temporal dependency is captured to derive the human attention mechanism on mobile videos. In MVFormer, we develop the selective feature fusion module to balance multi-scale features, and the progressive saliency prediction module to generate saliency maps via progressive aggregation of multi-scale features. Extensive experiments show that our MVFormer approach significantly outperforms other state-of-the-art saliency prediction approaches. Finally, we demonstrate the potential application of our MVFormer approach in the H.265 video coding standard by embedding it into the rate control scheme, such that the perceptual quality of compressed mobile videos can be significantly improved. The dataset and code will be available at https://github.com/wenshijie110/MVFormer.
Shijie Wen, Li Yang 0014, Mai Xu, Minglang Qiao, Lin Bai 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 HDDet: A More Common Heading Direction Detector for Remote Sensing and Arbitrary Viewing Angle Images
abstract
Object heading detection (OHD) offers potential for research in the control and traffic analysis sectors. Contemporary methodologies in OHD grapple with a set of distinct limitations: a constrained range of detectable object types, a discernible drop in accuracy for oriented bounding box (OBB) predictions influenced by heading estimations, and the inherent limitations associated with the exclusive perspective of bird’s-eye view imagery. This paper introduces HDDet, an advanced method devised for detecting the OBB of objects with heading direction from various viewpoints. It sequentially delineates the Circular Annotation Method (CAM), Multi-dimensional Angle Encoding (MDAE), and the CosWeight strategy. CAM initially couples oriented bounding box data with heading details for accurate target delineation. MDAE follows, optimizing angle encoding to boost the model’s training process. The culmination of this approach is CosWeight, which integrates the Rotated-IoU and heading information into the horizontal box’s loss, thereby enhancing the precision of heading predictions. Rigorous testing across diverse datasets, including the SJTU-L dataset where the heading accuracy increased from 64.41% to 94.82% over OHDet, validates HDDet’s enhanced capabilities, surpassing existing methodologies in both OBB detection and heading accuracy, and marking a significant advancement over the state-of-the-art OHDet. Additionally, the paper presents an engineering vehicle dataset, which is conducive to multi-perspective object heading detection research.
Siran Ding, Jingxian Liu, Mai Xu
IEEE Trans. Geosci. Remote. Sens.4
2024 Laplacian Gradient Consistency Prior for Flash Guided Non-Flash Image Denoising
abstract
For flash guided non-flash image denoising, the main challenge is to explore the consistency prior between the two modalities. Most existing methods attempt to model the flash/non-flash consistency in pixel level, which may easily lead to blurred edges. Different from these methods, we have an important finding in this paper, which reveals that the modality gap between flash and non-flash images conforms to the Laplacian distribution in gradient domain. Based on this finding, we establish a Laplacian gradient consistency (LGC) model for flash guided non-flash image denoising. This model is demonstrated to have faster convergence speed and denoising accuracy than the traditional pixel consistency model. Through solving the LGC model, we further design a deep network namely LGCNet. Different from existing image denoising networks, each component of the LGCNet strictly matches the solution of LGC model, giving the network good interpretability. The performance of the proposed LGCNet is evaluated on three different flash/non-flash image datasets, which demonstrates its superior denoising performance over many state-of-the-art methods both quantitatively and qualitatively. The intermediate features are also visualized to verify the effectiveness of the Laplacian gradient consistency prior. The source codes are available at https://github.com/JingyiXu404/LGCNet.
Xin Deng 0002, Shengxi Li, Mai Xu
IEEE Trans. Image Process.5
2024 Blind Quality Enhancement for Compressed Video
abstract
Deep convolutional neural networks (CNNs) have achieved impressive success in enhancing the quality of compressed images/videos. These approaches mostly obtain the noise level in advance and train multiple architecture-identical models for enhancement on images/videos of known levels of noise. It largely hinders their practical applications where the noise level is unknown and resource is limited. To practically perform quality enhancement, we propose a novel blind quality enhancement framework for compressed video (BQEV), which utilizes a single network to conduct enhancement on videos compressed at various and unknown quality parameters (QPs). Since there exists feature similarity and difference among videos compressed at multiple QPs, BQEV utilizes this prior to efficiently handle enhancement on videos compressed at blind QPs, which consists of progressive feature extraction and QP-adaptive feature fusion subnets. They utilize temporal information and feature similarity to progressively extract valuable features and further employ the feature difference to conduct reasonable QP-adaptive feature fusion and quality enhancement, respectively. In the progressive feature extraction subnet, we first design a quality rank module to assign more attention to higher-quality frames for efficient utilization of temporal information, then propose a progressive extraction module to further extract features from different QPs. In the QP-adaptive feature fusion subnet, we develop a quality estimation module to guide reasonable feature fusion of these extracted progressive features for stable and promising enhancement results on multiple QPs. Experimental results demonstrate that BQEV achieves 0.31–0.69 dB PSNR improvement compared with videos compressed at various QPs, outperforming state-of-the-art approaches.
Liquan Shen, Liangwei Yu, Hao Yang 0008, Mai Xu
IEEE Trans. Multim.5
2023 Learning Noise-Induced Reward Functions for Surpassing Demonstrations in Imitation Learning
abstract
Imitation learning (IL) has recently shown impressive performance in training a reinforcement learning agent with human demonstrations, eliminating the difficulty of designing elaborate reward functions in complex environments. However, most IL methods work under the assumption of the optimality of the demonstrations and thus cannot learn policies to surpass the demonstrators. Some methods have been investigated to obtain better-than-demonstration (BD) performance with inner human feedback or preference labels. In this paper, we propose a method to learn rewards from suboptimal demonstrations via a weighted preference learning technique (LERP). Specifically, we first formulate the suboptimality of demonstrations as the inaccurate estimation of rewards. The inaccuracy is modeled with a reward noise random variable following the Gumbel distribution. Moreover, we derive an upper bound of the expected return with different noise coefficients and propose a theorem to surpass the demonstrations. Unlike existing literature, our analysis does not depend on the linear reward constraint. Consequently, we develop a BD model with a weighted preference learning technique. Experimental results on continuous control and high-dimensional discrete control tasks show the superiority of our LERP method over other state-of-the-art BD methods.
Liangyu Huo, Zulin Wang, Mai Xu
AAAI3
2023 DINN360: Deformable Invertible Neural Network for Latitude-aware 360° Image Rescaling
abstract
With the rapid development of virtual reality, 360° images have gained increasing popularity. Their wide field of view necessitates high resolution to ensure image quality. This, however, makes it harder to acquire, store and even process such 360° images. To alleviate this issue, we propose the first attempt at 360° image rescaling, which refers to downscaling a 360° image to a visually valid lowresolution (LR) counterpart and then upscaling to a highresolution (HR) 360° image given the LR variant. Specifically, we first analyze two 360° image datasets and observe several findings that characterize how 360° images typically change along their latitudes. Inspired by these findings, we propose a novel deformable invertible neural network (INN), named DINN360, for latitude-aware 360° image rescaling. In DINN360, a deformable INN is designed to downscale the LR image, and project the high-frequency (HF) component to the latent space by adaptively handling various deformations occurring at different latitude regions. Given the downscaled LR image, the high-quality HR image is then reconstructed in a conditional latitude-aware manner by recovering the structure-related HF component from the latent space. Extensive experiments over four public datasets show that our DINN360 method performs considerably better than other state-of-the-art methods for 2 x, 4 x and 8 x 360° image rescaling.
Mai Xu, Lai Jiang 0004, Leonid Sigal, Yunjin Chen
CVPR2
2023 Learnt Mutual Feature Compression for Machine Vision
abstract
Recently, image coding for machines (ICM) has been playing an important role in facilitating intelligent vision tasks. Unfortunately, the existing ICM methods separately compress features at each scale, neglecting the redundancy across multi-scale features. To address this issue, this paper proposes an end-to-end mutual compression framework for the ICM, such that the compression efficiency can be significantly improved by removing the cross-scale redundancy. Specifically, the proposed framework consists of a mutual feature compression network (MFCNet) and a basic feature compression network (BFCNet). The MFCNet predicts large-scale features from basic small-scale features, such that the large amount of bitrates assigned to compress large-scale features can be saved. Moreover, the BFCNet is proposed to compress small-scale features of high quality by removing spatial and channel-wise redundancy. This guarantees superior performances whilst consuming extremely small amount of bit-rates. The experimental results show that our method achieves 90.10% and 74.97% BD-rate saving against the VVC feature anchor and VVC image anchor that have been recently accepted by the moving picture experts group (MPEG).
Mai Xu, Shengxi Li, Li Yang 0014, Zhuoyi Lv
ICASSP2
2023 PIRNet: Privacy-Preserving Image Restoration Network via Wavelet Lifting
abstract
The cloud-based multimedia service becomes increasingly popular in the last decade, however, it poses a serious threat to the client’s privacy. To address this issue, many methods utilized image encryption as a defense mechanism. However, the encrypted images look quite different from the natural images, making them vulnerable to attackers. In this paper, we propose a novel method namely PIRNet, which operates privacy-preserving image restoration in the steganographic domain. Compared to existing methods, our method offers significant advantages in terms of invisibility and security. Specifically, we first propose a wavelet Lifting-based Invertible Hiding (LIH) network to conceal the secret image into the stego image. Then, a Lifting-based Secure Restoration (LSR) network is utilized to perform image restoration in the steganographic domain. Since the secret image remains hidden throughout the whole image restoration process, the privacy of clients can be largely ensured. In addition, since the stego image looks visually the same as the cover image, the attackers can hardly discover it, which significantly improves the security. The experimental results on different datasets show the superiority of our PIRNet over the existing methods on various privacy-preserving image restoration tasks, including image denoising, deblurring and super-resolution.
Xin Deng 0002, Mai Xu
ICCV3
2023 Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo Matching
abstract
Correlation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed model do not work uniformly well across various datasets, greatly limiting their real-world applicability. To tackle this issue, this paper proposes a new perspective to dynamically calculate correlation for robust stereo matching. A novel Uncertainty Guided Adaptive Correlation (UGAC) module is introduced to robustly adapt the same model for different scenarios. Specifically, a variance-based uncertainty estimation is employed to adaptively adjust the sampling area during warping operation. Additionally, we improve the traditional non-parametric warping with learnable parameters, such that the position-specific weights can be learned. We show that by empowering the recurrent network with the UGAC module, stereo matching can be exploited more robustly and effectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance over the ETH3D, KITTI, and Middlebury datasets when employing the same fixed model over these datasets without any retraining procedure. To target real-time applications, we further design a lightweight model based on UGAC, which also outperforms other methods over KITTI benchmarks with only 0.6 M parameters.
Junpeng Jing, Jiankun Li, Pengfei Xiong, Jiangyu Liu, Shuaicheng Liu, Xin Deng 0002, Mai Xu, Lai Jiang 0004, Leonid Sigal
ICCV8
2023 Neural Characteristic Function Learning for Conditional Image Generation
abstract
The emergence of conditional generative adversarial networks (cGANs) has revolutionised the way we approach and control the generation, by means of adversarially learning joint distributions of data and auxiliary information. Despite the success, cGANs have been consistently put under scrutiny due to their ill-posed discrepancy measure between distributions, leading to mode collapse and instability problems in training. To address this issue, we propose a novel conditional characteristic function generative adversarial network (CCF-GAN) to reduce the discrepancy by the characteristic functions (CFs), which is able to learn accurate distance measure of joint distributions under theoretical soundness. More specifically, the difference between CFs is first proved to be complete and optimisation-friendly, for measuring the discrepancy of two joint distributions. To relieve the problem of curse of dimensionality in calculating CF difference, we propose to employ the neural network, namely neural CF (NCF), to efficiently minimise an upper bound of the difference. Based on the NCF, we establish the CCF-GAN framework to explicitly decompose CFs of joint distributions, which allows for learning the data distribution and auxiliary information with classified importance. The experimental results on synthetic and real-world datasets verify the superior performances of our CCF-GAN, on both the generation quality and stability.
Shengxi Li, Mai Xu, Xin Deng 0002
ICCV4
2023 ULcompress: A Unified low bit-rate image Compression Framework via Invertible Image Representation
abstract
In this paper, we propose a unified low bit-rate image compression framework, namely ULCompress, via invertible image representation. The proposed framework is composed of two important modules, including an invertible image rescaling (IIR) module and a compressed quality enhancement (CQE) module. The role of IIR module is to learn a compression-friendly low-resolution (LR) image from the high-resolution (HR) image. Instead of the HR image, we compress the LR image to save the bit-rates. The compression codecs can be any existing codecs. After compression, we propose a CQE module to enhance the quality of the compressed LR image, which is then sent back to the IIR module to restore the original HR image. The network architecture of IIR module is specially designed to ensure the invertibility of LR and HR images, i.e., the downsampling and upsampling processes are invertible. The CQE module works as a buffer between IIR module and the codec, which plays an important role in improving the compatibility of our framework. Experimental results show that our ULCompress is compatible with both standard and learning-based codecs, and is able to significantly improve their performance at low bit-rates.
Fangyuan Gao, Xin Deng 0002, Mai Xu
ICIP4
2023 Residual based hierarchical feature compression for multi-task machine vision
abstract
With the remarkable success of deep learning, image/video coding for machines (VCM) has been playing an important role in facilitating intelligent vision tasks. However, the existing VCM methods suffer from either sub-optimality of using image compression standards, or generalisation issues of learning-based methods. To address these issues, this paper proposes a residual-based hierarchical feature compression (RHFC) method to achieve optimal and universal feature compression for object detection and segmentation. More specifically, we first analyse the redundancy that exists in features at multiple scales, by finding that large-scale features are surprisingly less important to the vision tasks. Thus, we propose a pair of compression and enhancement networks to extract the very basic cues from the large-scale features, which are then compressed by the VVC codec. To compensate the inevitable detail loss, we further propose the hierarchical framework to compress the residuals between the reconstructed and original features, such that the performances can be significantly improved at low bit-rate cost. Experimental results have verified our superior performances, against both the state-of-the-art learning-based and standard feature compression methods. Our RHFC method also generalises well to other scenarios without the need of any further fine-tuning.
Mai Xu, Shengxi Li, Minglang Qiao, Zhuoyi Lv
ICME2
2023 A Two-stage hybrid CNN-Transformer Network for RGB Guided Indoor Depth Completion
abstract
The indoor captured raw depth images usually contain large in-homogeneous missing regions. Most existing methods are designed for the outdoor sparse depth completion, which struggle in completing the indoor depth with large holes. In this paper, to solve this problem, we propose a hybrid CNN-Transformer network for RGB guided indoor depth completion. The proposed network is composed of two stages to achieve depth completion in a coarse-to-fine manner. In the first stage, we propose a CNN based self-completion module (SCM) with cross scale attention to restore a coarse depth image. In the second stage, we further refine the completed depth image with the guidance of RGB image by proposing a guided completion module (GCM). To fully explore the guidance from the RGB image, we design a cross-modal Transformer (CMT) block to fuse the features from the depth and RGB modalities at different scales. Extensive experiments on NYUv2 and SUN RGB-D datasets demonstrate the superior performance of the proposed method over other state-of-the-art methods both quantitatively and qualitatively. The code is available at https://github.com/eecoder-dyf/ICME-2023-depth-completion.
Yufan Deng, Xin Deng 0002, Mai Xu
ICME3
2023 Optimizing DNN based quality assessment metric for image compression: A novel rate control method
abstract
In the existing coding standards, rate control (RC) plays a critical role in optimally allocating bit-rates to each coding unit, for improving rate-distortion performance under the limited bandwidth. However, the existing RC methods are mainly based on traditional distortion metrics, which fail to take the advantage of the emerging DNN based image quality assessment (IQA) metrics. In this paper, we set up the first attempt to achieve IQA score based RC for image compression. Specifically, a novel visualization based score-distortion (VSD) model and ρ-slope model are proposed to explicitly establish the relationship between IQA score and bit-rates. Then, by solving optimal rate-distortion optimization based on the IQA score, we propose a novel RC method for the HEVC standard. The experimental results show that, given the target bit-rates, the proposed RC method can accurately control the bit-rates and generate the compressed images with higher IQA score and better perceptual quality. More importantly, the proposed RC method is evaluated to be effective over two DNN based IQA metrics and four image datasets, exhibiting the potential in practical use. The code is available at https://github.com/Ffangqy/IQA-RC.
Qiuyue Fang, Lai Jiang 0004, Shengxi Li, Mai Xu, Yunjin Chen, Leonid Sigal
ICME5
2023 Recruiting the Best Teacher Modality: A Customized Knowledge Distillation Method for if Based Nephropathy Diagnosis
Lai Jiang 0004, Yibing Fu, Sai Pan, Mai Xu, Xin Deng 0002, Xiangmei Chen
MICCAI (5)5
2023 Predicting the Invariance Behind Residuals: A Novel GAN Inversion Method for Image Editing and Detail Retaining
abstract
Generative adversarial network (GAN) inversion has been serving as the vehicle to enable the restoration of real-world images by GANs, rather than the realistic generation from random noise. Existing GAN inversion methods, however, suffer from the fidelity-editability trade-off, which mainly invert the semantics within images and fail to reconstruct the details. To address this issue, we propose a novel adaptive detail compensation method for GAN inversion (ADC-GInv), which automatically locates and restores the non-semantic details, whilst maintaining the editability on the semantics of real-world images. More specifically, we first develop an adversarial reciprocal learning GAN (ARL-GAN) so as to seamlessly optimise the reconstruction during the training of GANs, followed by a sophisticated fine-tuning technique for ARL-GAN inversion. This ensures superior restoration and editability on the semantic cues of images. Then, regarding the non-semantic details, our ADC-GInv method adaptively locates the details by predicting the invariance given edited and non-edited residuals of restoration, which are then compensated at the pixel-level for high fidelity. As a consequence, the experimental results have verified the superior performance of our ADC-GInv, on both fidelity and editability during inversion.
Zhimo Yan, Hengyang He, Shengxi Li, Mai Xu, Ce Zhu
MMSP6
2023 Learned Structure-Based Hybrid Framework for Martian Image Compression
abstract
Recent landing marches on Mars have enabled the access to Martian surface images, which act as an important vehicle to demystify the evolution and habitability of Mars, in terms of climate, geography, etc. Transmitting Martian images thus calls for efficient compression methods to ensure the high-quality reconstruction from distant communication, in which the research is yet to start. To address this issue, we propose in this letter a learned structure-based hybrid (LSH) framework to compress Martian images. More specifically, we first observe that the structural consistency exists across Martian images, which motivates us to propose a structural compression network (SCN). The aim of SCN is to compactly represent the structural information of Martian images, thus allowing for the compression at extremely low bit-rates. Then, we propose a detail compensation network (DCN) to reconstruct the missing details when we restore from the structural information, which benefits from improved compression efficiency by reduced bit-rates. The experimental results have verified the superior performances of our LSH method on compressing Martian images, against existing state-of-the-art methods.
Shengxi Li, Xiancheng Sun, Mai Xu, Lai Jiang 0004
IEEE Geosci. Remote. Sens. Lett.3
2023 DeepMIH: Deep Invertible Network for Multiple Image Hiding
abstract
Multiple image hiding aims to hide multiple secret images into a single cover image, and then recover all secret images perfectly. Such high-capacity hiding may easily lead to contour shadows or color distortion, which makes multiple image hiding a very challenging task. In this paper, we propose a novel multiple image hiding framework based on invertible neural network, namely DeepMIH. Specifically, we develop an invertible hiding neural network (IHNN) to innovatively model the image concealing and revealing as its forward and backward processes, making them fully coupled and reversible. The IHNN is highly flexible, which can be cascaded as many times as required to achieve the hiding of multiple images. To enhance the invisibility, we design an importance map (IM) module to guide the current image hiding based on the previous image hiding results. In addition, we find that the image hidden in the high-frequency sub-bands tends to achieve better hiding performance, and thus propose a low-frequency wavelet loss to constrain that no secret information is hidden in the low-frequency sub-bands. Experimental results show that our DeepMIH significantly outperforms other state-of-the-art methods, in terms of hiding invisibility, security and recovery accuracy on a variety of datasets.
Zhenyu Guan 0002, Junpeng Jing, Xin Deng 0002, Mai Xu, Lai Jiang 0004, Zhou Zhang 0016
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 DAQE: Enhancing the Quality of Compressed Images by Exploiting the Inherent Characteristic of Defocus
abstract
Image defocus is inherent in the physics of image formation caused by the optical aberration of lenses, providing plentiful information on image quality. Unfortunately, existing quality enhancement approaches for compressed images neglect the inherent characteristic of defocus, resulting in inferior performance. This paper finds that in compressed images, significantly defocused regions have better compression quality, and two regions with different defocus values possess diverse texture patterns. These observations motivate our defocus-aware quality enhancement (DAQE) approach. Specifically, we propose a novel dynamic region-based deep learning architecture of the DAQE approach, which considers the regionwise defocus difference of compressed images in two aspects. (1) The DAQE approach employs fewer computational resources to enhance the quality of significantly defocused regions and more resources to enhance the quality of other regions; (2) The DAQE approach learns to separately enhance diverse texture patterns for regions with different defocus values, such that texture-specific enhancement can be achieved. Extensive experiments validate the superiority of our DAQE approach over state-of-the-art approaches in terms of quality enhancement and resource savings.
Qunliang Xing, Mai Xu, Xin Deng 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 New Finding and Unified Framework for Fake Image Detection
abstract
Recently, fake face images generated by generative adversarial network (GAN) have been widely spread in social networks, raising serious social concerns and security risks. To identify the fake images, the top priority is to find what properties make the fake images different from the real images. In this letter, we reveal an important observation about real/fake images, i.e., the GAN generated fake images contain stronger non-local self-similarity than the real images. Motivated by this observation, we propose a simple yet effective non-local attention based fake image detection network, namely NAFID, to distinguish GAN generated fake images from real images. Specifically, we develop a non-local feature extraction (NFE) module to extract the non-local features of the real/fake images, followed by a multi-stage classification module to distinguish the images with the extracted non-local features. Experimental results on various datasets demonstrate the superiority of our NAFID over state-of-the-art (SOTA) face forgery detection methods. More importantly, since the NFE module is independent from classification, we can plug it into any other forgery detection models. The results show that the NFE module can consistently improve the detection accuracy of other models, which verifies the universality of the proposed method.
Xin Deng 0002, Bihe Zhao, Zhenyu Guan 0002, Mai Xu
IEEE Signal Process. Lett.4
2023 MASIC: Deep Mask Stereo Image Compression
abstract
Stereo image compression (SIC) aims to simultaneously compress a pair of left and right stereoscopic images, which can achieve higher compression efficiency than single image compression. In this paper, to benefit the SIC tasks, we collect a large real-world stereo image dataset, namely Palace, which is composed of hundreds of stereo image pairs at high-resolution. More importantly, we propose a novel mask stereo image compression network, namely MASIC, which can jointly compress the stereo images with high compression efficiency. Specifically, we first estimate the homography matrix between the stereo images through a regression model. Then, the left image is spatially transformed by the homography matrix, so that only the residual information needs to be encoded for the right image. To avoid the wrong guidance between stereo image pair, we propose a mask prediction module (MPM) to generate a multi-channel guided mask to navigate both the encoding and decoding processes. Based on the guided mask, we introduce a new mask conditional stereo entropy (MCSE) model, to fully explore the correlation between the stereo images in entropy coding. In the decoder, we develop a stereo decoding module to simultaneously decode the stereo images and enhance their compression quality. Experimental results show that our MASIC significantly advances the performance of SIC both quantitatively and qualitatively on a variety of datasets, and is robust to the change of parallax level between stereo images. The software codes are available athttps://github.com/eecoder-dyf/MASIC.
Xin Deng 0002, Yufan Deng, Radu Timofte, Mai Xu
IEEE Trans. Circuits Syst. Video Technol.6
2023 Deep Multi-Task Learning Based Fast Intra-Mode Decision for Versatile Video Coding
abstract
The latest Versatile Video Coding (VVC) standard has significantly coding efficiency improvement compared with its ancestor High Efficiency Video Coding (HEVC) standard, but at the expense of over-high complexity. As measured by the VVC test model (VTM), the intra-mode comparison and selection in the rate-distortion optimization (RDO) search consume most of the encoding time. In this paper, we propose a deep multi-task learning based fast intra-mode decision approach via adaptively pruning off most redundant modes. First, we create a large-scale intra-mode database for VVC, including both normal angular modes and the newly introduced tools, i.e., intra sub-partition (ISP) and matrix-based intra prediction (MIP). Next, we propose a multi-task intra-mode decision network (MID-Net) model to effectively predict the most probable angular modes and whether to skip ISP and MIP modes. Then, a fast intra-coding workflow is designed accordingly, involving rough mode decision (RMD) acceleration and candidate mode list (CML) pruning. For the workflow output, the learning-oriented probability and the statistics-oriented probability are synthesized together to further improve the prediction accuracy, ensuring that only unnecessary intra-modes are skipped. Finally, experimental results show that our approach can significantly reduce 40.48% of encoding time of VVC intra-coding with negligible rate-distortion degradation, outperforming other state-of-the-art approaches.
Tianyi Li 0004, Ying Chen 0011, Kaijin Wei, Mai Xu, Honggang Qi
IEEE Trans. Circuits Syst. Video Technol.5
2023 Interpretable Multi-Modal Image Registration Network Based on Disentangled Convolutional Sparse Coding
abstract
Multi-modal image registration aims to spatially align two images from different modalities to make their feature points match with each other. Captured by different sensors, the images from different modalities often contain many distinct features, which makes it challenging to find their accurate correspondences. With the success of deep learning, many deep networks have been proposed to align multi-modal images, however, they are mostly lack of interpretability. In this paper, we first model the multi-modal image registration problem as a disentangled convolutional sparse coding (DCSC) model. In this model, the multi-modal features that are responsible for alignment (RA features) are well separated from the features that are not responsible for alignment (nRA features). By only allowing the RA features to participate in the deformation field prediction, we can eliminate the interference of the nRA features to improve the registration accuracy and efficiency. The optimization process of the DCSC model to separate the RA and nRA features is then turned into a deep network, namely Interpretable Multi-modal Image Registration Network (InMIR-Net). To ensure the accurate separation of RA and nRA features, we further design an accompanying guidance network (AG-Net) to supervise the extraction of RA features in InMIR-Net. The advantage of InMIR-Net is that it provides a universal framework to tackle both rigid and non-rigid multi-modal image registration tasks. Extensive experimental results verify the effectiveness of our method on both rigid and non-rigid registrations on various multi-modal image datasets, including RGB/depth images, RGB/near-infrared (NIR) images, RGB/multi-spectral images, T1/T2 weighted magnetic resonance (MR) images and computed tomography (CT)/MR images. The codes are available at https://github.com/lep990816/Interpretable-Multi-modal-Image-Registration.
Xin Deng 0002, Enpeng Liu, Shengxi Li, Yiping Duan, Mai Xu
IEEE Trans. Image Process.5
2023 Domain Adaptation for Underwater Image Enhancement
abstract
Recently, learning-based algorithms have shown impressive performance in underwater image enhancement. Most of them resort to training on synthetic data and obtain outstanding performance. However, these deep methods ignore the significant domain gap between the synthetic and real data (i.e., inter-domain gap), and thus the models trained on synthetic data often fail to generalize well to real-world underwater scenarios. Moreover, the complex and changeable underwater environment also causes a great distribution gap among the real data itself (i.e., intra-domain gap). However, almost no research focuses on this problem and thus their techniques often produce visually unpleasing artifacts and color distortions on various real images. Motivated by these observations, we propose a novel Two-phase Underwater Domain Adaptation network (TUDA) to simultaneously minimize the inter-domain and intra-domain gap. Concretely, in the first phase, a new triple-alignment network is designed, including a translation part for enhancing realism of input images, followed by a task-oriented enhancement part. With performing image-level, feature-level and output-level adaptation in these two parts through jointly adversarial learning, the network can better build invariance across domains and thus bridging the inter-domain gap. In the second phase, an easy-hard classification of real data according to the assessed quality of enhanced images is performed, in which a new rank-based underwater quality assessment method is embedded. By leveraging implicit quality information learned from rankings, this method can more accurately assess the perceptual quality of enhanced images. Using pseudo labels from the easy part, an easy-hard adaptation technique is then conducted to effectively decrease the intra-domain gap between easy and hard samples. Extensive experimental results demonstrate that the proposed TUDA is significantly superior to existing works in terms of both visual quality and quantitative metrics.
Zhengyong Wang, Liquan Shen, Mai Xu, Mei Yu 0001, Kun Wang 0048
IEEE Trans. Image Process.3
2023 Blind VQA on 360° Video via Progressively Learning From Pixels, Frames, and Video
abstract
Blind visual quality assessment (BVQA) on 360° video plays a key role in optimizing immersive multimedia systems. When assessing the quality of 360° video, human tends to perceive its quality degradation from the viewport-based spatial distortion of each spherical frame to motion artifact across adjacent frames, ending with the video-level quality score, i.e., a progressive quality assessment paradigm. However, the existing BVQA approaches for 360° video neglect this paradigm. In this paper, we take into account the progressive paradigm of human perception towards spherical video quality, and thus propose a novel BVQA approach (namely ProVQA) for 360° video via progressively learning from pixels, frames and video. Corresponding to the progressive learning of pixels, frames and video, three sub-nets are designed in our ProVQA approach, i.e., the spherical perception aware quality prediction (SPAQ), motion perception aware quality prediction (MPAQ) and multi-frame temporal non-local (MFTN) sub-nets. The SPAQ sub-net first models the spatial quality degradation based on spherical perception mechanism of human. Then, by exploiting motion cues across adjacent frames, the MPAQ sub-net properly incorporates motion contextual information for quality assessment on 360° video. Finally, the MFTN sub-net aggregates multi-frame quality degradation to yield the final quality score, via exploring long-term quality correlation from multiple frames. The experiments validate that our approach significantly advances the state-of-the-art BVQA performance on 360° video over two datasets, the code of which has been public in https://github.com/yanglixiaoshen/ProVQA.
Li Yang 0014, Mai Xu, Shengxi Li, Zulin Wang
IEEE Trans. Image Process.2
2023 Omnidirectional Image Super-Resolution via Latitude Adaptive Network
abstract
Omnidirectional images (ODI), also known as 360 images, have recently attracted extensive attention from both academia and industry. However, due to storage and transmission limitations, ODIs are usually at extremely low resolution. Thus, it is necessary to restore a high-resolution ODI from a low-resolution ODI, i.e., omnidirectional image super-resolution (ODI-SR). Different from traditional two-dimensional (2D) image SR, the challenge of ODI-SR is the nonuniformly distributed pixel density and geometric distortion across latitudes, which makes traditional SR methods difficult to be applied in ODI-SR. Towards ODI-SR, we propose in this paper a novel latitude-aware upscaling network, namely LAU-Net+, which fully considers the above characteristics of ODIs. In our network, different latitude bands can learn to adopt distinct upscaling factors, which significantly saves the computational resources and improves the SR efficiency. Specifically, a Laplacian multilevel pyramid network is introduced in which the upscaling factor is gradually increased with the number of levels. Each level is composed of a feature enhancement module (FEM), a drop-band decision module (DDM) and a high-latitude enhancement module (HEM). The FEM module serves to enhance the high-level features extracted from the input ODI, while the role of DDM is to dynamically drop the unnecessary high latitude bands and send the remained bands to the next level. The HEM is adopted to further enhance high-level features of dropped latitude bands with a lightweight architecture. In DDM, we develop a reinforcement learning scheme with a latitude adaptive reward to determine which band should be dropped. To the best of our knowledge, our method is the first work which considers the latitude characteristics for ODI-SR task. Extensive experimental results demonstrate that our LAU-Net+ achieves state-of-the-art results on ODI-SR both quantitatively and qualitatively on various ODI datasets.
Xin Deng 0002, Hao Wang 0049, Mai Xu, Zulin Wang
IEEE Trans. Multim.3
2023 A Task-Agnostic Regularizer for Diverse Subpolicy Discovery in Hierarchical Reinforcement Learning
abstract
The automatic subpolicy discovery approach in hierarchical reinforcement learning (HRL) has recently achieved promising performance on sparse reward tasks. This accelerates transfer learning and unsupervised intelligent creatures while eliminating the domain-specific knowledge constraint. Most previously developed approaches are demonstrated to suffer from collapsing into the situation where one subpolicy dominates the whole task, since they cannot ensure the diversity of different subpolicies. In contrast, this article proposes a task-agnostic regularizer (TAR) for learning diverse subpolicies in HRL. Specifically, we first formulate the discovery of diverse subpolicies as a trajectory inference problem and then propose a corresponding information-theoretic objective to encourage diversity. Subsequently, considering computability, we instantiate the objective as two simplifications for discrete and continuous action spaces. We extensively evaluate the proposed diversity-driven regularizer on three HRL task domains: 1) meta reinforcement learning; 2) hierarchical policy learning in the option framework; and 3) unsupervised subpolicy discovery. The extensive results obtained show that our TAR approach can improve upon the state-of-the-art performance on all three HRL domains without modifying any existing hyperparameters, indicating the wide applicability and robustness of our approach.
Liangyu Huo, Zulin Wang, Mai Xu, Yuhang Song 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2022 Does text attract attention on e-commerce images: A novel saliency prediction dataset and method
abstract
E-commerce images are playing a central role in attracting people's attention when retailing and shopping online, and an accurate attention prediction is of significant importance for both customers and retailers, where its research is yet to start. In this paper, we establish the first dataset of saliency e-commerce images (SalECI), which allows for learning to predict saliency on the e-commerce images. We then provide specialized and thorough analysis by high-lighting the distinct features of e-commerce images, e.g., non-locality and correlation to text regions. Correspondingly, taking advantages of the non-local and self-attention mechanisms, we propose a salient SWin-Transformer back-bone, followed by a multi-task learning with saliency and text detection heads, where an information flow mechanism is proposed to further benefit both tasks. Experimental results have verified the state-of-the-art performances of our work in the e-commerce scenario.
Lai Jiang 0004, Shengxi Li, Mai Xu, Se Lei
CVPR4
2022 SFIC: Sparsity-Driven Facial Image Compression Network
abstract
Facial image compression is crucial in many areas like social media and video surveillance. Considering the sparsity of facial features, sparse representation (SR) has been applied to compress facial images, in which each image patch is sparsely represented by a small number of dictionary atoms to save bit-rates. Along this line, we propose the first end-to-end sparsity-driven facial image compression network namely SFIC. In the proposed network, the traditional convolutional sparse coding (CSC) is turned into a learnable CSC block, which is combined with discrete wavelet transform (DWT) to form the sparsity encoding module (SEM). This is the first time that CSC has been explored in facial image compression. In the decoding side, a corresponding sparsity decoding module (SDM) is used to decode the image, and we further propose a quality enhancement module (QEM) to enhance the quality of decoded image. The experimental results verify that the proposed SFIC network achieves 74%, 55%, and 33% bit-rate savings over JPEG, JPEG-2000, and HEVC.
Fangyuan Gao, Xin Deng 0002, Mai Xu
ICIP4
2022 TVFormer: Trajectory-guided Visual Quality Assessment on 360° Images with Transformers
abstract
Visual quality assessment (VQA) on 360° images plays an important role in optimizing immersive multimedia systems. Due to the absence of pristine 360° images in real world, blind VQA (BVQA) on 360° images has drawn much research attention. In subjective VQA on 360^ images, human intuitively make the quality-scoring decisions through the quality degradation of each observed viewport on the head trajectories. Unfortunately, the existing BVQA works for 360° images neglect the dynamic property of head trajectories with viewport interactions, thus failing to obtain human-like quality scores. In this paper, we propose a novel Transformer-based approach for trajectory-guided VQA on 360° images (named TVFormer), in which both the tasks of head trajectory prediction and BVQA can be accomplished for 360° images. In the first task, we develop a trajectory-aware memory updater (TMU) module, for maintaining the coherence and accuracy of predicted head trajectories. To capture the long-range quality dependency across time-ordered viewports, we propose a spatio-temporal factorized self-attention (STF) module in the encoder of TVFormer for the BVQA task. By implanting the predicted head trajectories into the BVQA task, we can obtain the human-like quality scores. Extensive experiments demonstrate the superior BVQA performance of TVFormer over state-of-the-art approaches on three benchmark datasets.
Li Yang 0014, Mai Xu, Liangyu Huo, Xinbo Gao 0001
ACM Multimedia2
2022 A Learning-based Approach for Martian Image Compression
abstract
For the scientific exploration and research on Mars, it is an indispensable step to transmit high-quality Martian images from distant Mars to Earth. Image compression is the key technique given the extremely limited Mars-Earth bandwidth. Recently, deep learning has demonstrated remarkable performance in natural image compression, which provides a possibility for efficient Martian image compression. However, deep learning usually requires large training data. In this paper, we establish the first large-scale high-resolution Martian image compression (MIC) dataset. Through analyzing this dataset, we observe an important non-local self-similarity prior for Marian images. Benefiting from this prior, we propose a deep Martian image compression network with the non-local block to explore both local and non-local dependencies among Martian image patches. Experimental results verify the effectiveness of the proposed network in Martian image compression, which outperforms both the deep learning based compression methods and HEVC codec.
Mai Xu, Shengxi Li, Xin Deng 0002, Qiu Shen
VCIP2
2022 Guest Editorial Special Issue on Space-Air-Ground-Integrated Networks for Internet of Vehicles
abstract
Internet of Vehicles (IoV) is one of the most promising applications of Internet of Things (IoT) in the automotive industry, which can empower moving vehicles to exchange information with neighboring cars, roadside infrastructure, remote servers, traffic control centers, and so on. IoV expects to support a wide range of vehicular services, such as road safety, path planning, infotainment, and smart parking, which will play a vital role in intelligent transportation systems (ITSs)[1]–[3]. The main enabling platforms for IoV consist of dedicated short-range communications (DSRCs)-based networks and cellular networks (C-V2X). However, these terrestrial networks alone might not be able to support the vehicular applications well in all the cases and scenarios, due to the issues of limited coverage and capacity, as well as costly deployment.
Tingting Yang 0001, Ning Zhang 0007, Mai Xu, Mehrdad Dianati, F. Richard Yu
IEEE Internet Things J.3
2022 Viewport-Based CNN: A Multi-Task Approach for Assessing 360° Video Quality
abstract
For 360° video, the existing visual quality assessment (VQA) approaches are designed based on either the whole frames or the cropped patches, ignoring the fact that subjects can only access viewports. When watching 360° video, subjects select viewports through head movement (HM) and then fixate on attractive regions within the viewports through eye movement (EM). Therefore, this paper proposes a two-staged multi-task approach for viewport-based VQA on 360° video. Specifically, we first establish a large-scale VQA dataset of 360° video, called VQA-ODV, which collects the subjective quality scores and the HM and EM data on 600 video sequences. By mining our dataset, we find that the subjective quality of 360° video is related to camera motion, viewport positions and saliency within viewports. Accordingly, we propose a viewport-based convolutional neural network (V-CNN) approach for VQA on 360° video, which has a novel multi-task architecture composed of a viewport proposal network (VP-net) and viewport quality network (VQ-net). The VP-net handles the auxiliary tasks of camera motion detection and viewport proposal, while the VQ-net accomplishes the auxiliary task of viewport saliency prediction and the main task of VQA. The experiments validate that our V-CNN approach significantly advances state-of-the-art VQA performance on 360° video and it is also effective in the three auxiliary tasks.
Mai Xu, Lai Jiang 0004, Chen Li 0049, Zulin Wang, Xiaoming Tao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Hierarchical Bayesian LSTM for Head Trajectory Prediction on Omnidirectional Images
abstract
When viewing omnidirectional images (ODIs), viewers can access different viewports via head movement (HM), which sequentially forms head trajectories in spatial-temporal domain. Thus, head trajectories play a key role in modeling human attention on ODIs. In this paper, we establish a large-scale dataset collecting 21,600 head trajectories on 1,080 ODIs. By mining our dataset, we find two important factors influencing head trajectories, i.e., temporal dependency and subject-specific variance. Accordingly, we propose a novel approach integrating hierarchical Bayesian inference into long short-term memory (LSTM) network for head trajectory prediction on ODIs, which is called HiBayes-LSTM. In HiBayes-LSTM, we develop a mechanism of Future Intention Estimation (FIE), which captures the temporal correlations from previous, current and estimated future information, for predicting viewport transition. Additionally, a training scheme called Hierarchical Bayesian inference (HBI) is developed for modeling inter-subject uncertainty in HiBayes-LSTM. For HBI, we introduce a joint Gaussian distribution in a hierarchy, to approximate the posterior distribution over network weights. By sampling subject-specific weights from the approximated posterior distribution, our HiBayes-LSTM approach can yield diverse viewport transition among different subjects and obtain multiple head trajectories. Extensive experiments validate that our HiBayes-LSTM approach significantly outperforms 9 state-of-the-art approaches for trajectory prediction on ODIs, and then it is successfully applied to predict saliency on ODIs.
Li Yang 0014, Mai Xu, Xin Deng 0002, Fangyuan Gao, Zhenyu Guan 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Revisiting Convolutional Sparse Coding for Image Denoising: From a Multi-Scale Perspective
abstract
Recently, convolutional sparse coding (CSC) has shown great success in many image processing tasks, such as image super-resolution and image separation. However, it performs poorly in image denoising task. In this letter, we provide a new insight for CSC denoising by revisiting the CSC from a multi-scale perspective. We propose a multi-scale CSC model for image denoising. By unrolling the multi-scale solution into a learnable network, we obtain an interpretable lightweight multiscale network, namely MCSCNet. Experimental results show that the proposed MCSCNet significantly advances the denoising performance, with an average PSNR improvement of 0.32 dB over the state-of-the-art (SOTA) CSC based method. In addition, our MCSCNet is on par with many SOTA deep learning based methods, with less network parameters and lower FLOPs. The ablation study also validates the effectiveness of the multi-scale CSC mechanism.
Xin Deng 0002, Mai Xu
IEEE Signal Process. Lett.3
2022 MRS-Net+ for Enhancing Face Quality of Compressed Videos
abstract
During the past few years, face videos, e.g., video conference, interviews and variety shows, have grown explosively with millions of users over social media networks. Unfortunately, the existing compression algorithms are applied to these videos for reducing bandwidth, which also bring annoying artifacts to face regions. This paper addresses the problem of face quality enhancement in compressed videos by reducing the artifacts of face regions. Specifically, we establish a compressed face video (CFV) database, which includes 196,337 faces in 214 high-quality video sequences and their corresponding 1,712 compressed sequences. We find that the faces of compressed videos exhibit tremendous scale variation and quality fluctuation. Motivated by scalable video coding, we propose a multi-scale recurrent scalable network (MRS-Net+) to enhance the quality of multi-scale faces in compressed videos. The MRS-Net+ is comprised by one base and two refined enhancement levels, corresponding to the quality enhancement of small-, medium- and large-scale faces, respectively. In the multi-level architecture of our MRS-Net+, small-/medium-scale face quality enhancement serves as the basis for facilitating the quality enhancement of medium-/large-scale faces. We further develop a landmark-assisted pyramid alignment (LPA) subnet to align faces across consecutive frames, and then apply the mask-guided quality enhancement (QE) subnet for enhancing multi-scale faces. Finally, experimental results show that our MRS-Net+ method achieves averagely 1.196 dB improvement of peak signal-to-noise ratio (PSNR) and 23.54% saving of Bjøntegaard distortion-rate (BD-rate), significantly outperforming other state-of-the-art methods.
Mai Xu, Shengxi Li, Huaida Liu
IEEE Trans. Circuits Syst. Video Technol.2
2022 MW-GAN+ for Perceptual Quality Enhancement on Compressed Video
abstract
The great success of deep learning has boosted the fast development of video quality enhancement. However, existing methods mainly focus on enhancing the objective quality of compressed video, and ignore their perceptual quality that plays a key role in determining quality of experience (QoE) of videos. In this paper, we aim at enhancing the perceptual quality of compressed video. Our main observation is that perceptual quality enhancement mostly relies on recovering the high-frequency details with fine textures. Accordingly, we propose a novel generative adversarial network (GAN) based on multi-level wavelet packet transform (WPT), which is called multi-level wavelet-based GAN+ (MW-GAN+), to exploit high-frequency details for enhancing the perceptual quality of compressed video. In MW-GAN+, we first propose a multi-level wavelet pixel-adaptive (MWP) module to extract temporal information across video frames, such that frame similarity can be utilized in recovering high-frequency details. Then, a wavelet reconstruction network, consisting of wavelet-dense residual blocks (WDRB), is developed to recover high-frequency details in a multi-level manner for enhanced frame reconstruction. Finally, we develop a 3D discriminator to encourage temporal coherence with a 3D-CNN based architecture. Experimental results demonstrate the superiority of our method over state-of-the-art methods in enhancing the perceptual quality of compressed video. Our code is available athttps://github.com/IceClear/MW-GAN.
Jianyi Wang, Mai Xu, Xin Deng 0002, Liquan Shen, Yuhang Song 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Multi-Modal Convolutional Dictionary Learning
abstract
Convolutional dictionary learning has become increasingly popular in signal and image processing for its ability to overcome the limitations of traditional patch-based dictionary learning. Although most studies on convolutional dictionary learning mainly focus on the unimodal case, real-world image processing tasks usually involve images from multiple modalities, e.g., visible and near-infrared (NIR) images. Thus, it is necessary to explore convolutional dictionary learning across different modalities. In this paper, we propose a novel multi-modal convolutional dictionary learning algorithm, which efficiently correlates different image modalities and fully considers neighborhood information at the image level. In this model, each modality is represented by two convolutional dictionaries, in which one dictionary is for common feature representation and the other is for unique feature representation. The model is constrained by the requirement that the convolutional sparse representations (CSRs) for the common features should be the same across different modalities, considering that these images are captured from the same scene. We propose a new training method based on the alternating direction method of multipliers (ADMM) to alternatively learn the common and unique dictionaries in the discrete Fourier transform (DFT) domain. We show that our model converges in less than 20 iterations between the convolutional dictionary updating and the CSRs calculation. The effectiveness of the proposed dictionary learning algorithm is demonstrated on various multimodal image processing tasks, achieves better performance than both dictionary learning methods and deep learning based methods with limited training data.
Fangyuan Gao, Xin Deng 0002, Mai Xu, Pier Luigi Dragotti
IEEE Trans. Image Process.3
2022 From Whole Video to Frames: Weakly-Supervised Domain Adaptive Continuous-Time QoE Evaluation
abstract
Due to the rapid increase in video traffic and relatively limited delivery infrastructure, end users often experience dynamically varying quality over time when viewing streaming videos. The user quality-of-experience (QoE) must be continuously monitored to deliver an optimized service. However, modern approaches for continuous-time video QoE estimation require densely annotating the continuous-time QoE labels, which is labor-intensive and time-consuming. To cope with such limitations, we propose a novel weakly-supervised domain adaptation approach for continuous-time QoE evaluation, by making use of a small amount of continuously labeled data in the source domain and abundant weakly-labeled data (only containing the retrospective QoE labels) in the target domain. Specifically, given a pair of videos from source and target domains, effective spatiotemporal segment-level feature representation is first learned by a combination of 2D and 3D convolutional networks. Then, a multi-task prediction framework is developed to simultaneously achieve continuous-time and retrospective QoE predictions, where a quality attentive adaptation approach is investigated to effectively alleviate the domain discrepancy without hampering the prediction performance. This approach is enabled by explicitly attending to the video-level discrimination and segment-level transferability in terms of the domain discrepancy. Experiments on benchmark databases demonstrate that the proposed method significantly improves the prediction performance under the cross-domain setting.
Leida Li, Pengfei Chen 0003, Weisi Lin, Mai Xu, Guangming Shi
IEEE Trans. Image Process.4
2022 Joint Learning of Multi-Level Tasks for Diabetic Retinopathy Grading on Low-Resolution Fundus Images
abstract
Diabetic retinopathy (DR) is a leading cause of permanent blindness among the working-age people. Automatic DR grading can help ophthalmologists make timely treatment for patients. However, the existing grading methods are usually trained with high resolution (HR) fundus images, such that the grading performance decreases a lot given low resolution (LR) images, which are common in clinic. In this paper, we mainly focus on DR grading with LR fundus images. According to our analysis on the DR task, we find that: 1) image super-resolution (ISR) can boost the performance of both DR grading and lesion segmentation; 2) the lesion segmentation regions of fundus images are highly consistent with pathological regions for DR grading. Based on our findings, we propose a convolutional neural network (CNN)-based method for joint learning of multi-level tasks for DR grading, called DeepMT-DR, which can simultaneously handle the low-level task of ISR, the mid-level task of lesion segmentation and the high-level task of disease severity classification on LR fundus images. Moreover, a novel task-aware loss is developed to encourage ISR to focus on the pathological regions for its subsequent tasks: lesion segmentation and DR grading. Extensive experimental results show that our DeepMT-DR method significantly outperforms other state-of-the-art methods for DR grading over three datasets. In addition, our method achieves comparable performance in two auxiliary tasks of ISR and lesion segmentation.
Xiaofei Wang 0004, Mai Xu, Jicong Zhang, Lai Jiang 0004, Liu Li 0001, Mengxian He, Ningli Wang, Hanruo Liu, Zulin Wang
IEEE J. Biomed. Health Informatics2
2021 Deep Multi-Task Learning for Diabetic Retinopathy Grading in Fundus Images
abstract
Recent years have witnessed the growing interest in disease severity grading, especially for ocular diseases based on fundus images. The existing grading methods are usually trained with high resolution (HR) images. However, the grading performance decreases a lot given low resolution (LR) images, which are common in practice. In this paper, we mainly focus on diabetic retinopathy (DR) grading with LR fundus images. According to our analysis on the DR task, we find that: 1) image super-resolution (ISR) can boost the performance of DR grading and lesion segmentation; 2) the lesion segmentation regions of fundus images are highly consistent with pathological regions for DR grading. Thus, we propose a deep multi-task learning based DR grading (DeepMT-DR) method for LR fundus images, which simultaneously handles the auxiliary tasks of ISR and lesion segmentation. Specifically, based on our findings, we propose a hierarchical deep learning structure that simultaneously processes the low-level task of ISR, the mid-level task of lesion segmentation and the high-level task of DR grading. Moreover, a novel task-aware loss is developed to encourage ISR to focus on the pathological regions for its subsequent tasks: lesion segmentation and DR grading. Extensive experimental results show that our DeepMT-DR method significantly outperforms other state-of-the-art methods for DR grading over two public datasets. In addition, our method achieves comparable performance in two auxiliary tasks of ISR and lesion segmentation.
Xiaofei Wang 0004, Mai Xu, Jicong Zhang, Lai Jiang 0004, Liu Li 0001
AAAI2
2021 LAU-Net: Latitude Adaptive Upscaling Network for Omnidirectional Image Super-Resolution
abstract
The omnidirectional images (ODIs) are usually at low-resolution, due to the constraints of collection, storage and transmission. The traditional two-dimensional (2D) image super-resolution methods are not effective for spherical ODIs, because ODIs tend to have non-uniformly distributed pixel density and varying texture complexity across latitudes. In this work, we propose a novel latitude adaptive upscaling network (LAU-Net) for ODI super-resolution, which allows pixels at different latitudes to adopt distinct upscaling factors. Specifically, we introduce a Laplacian multi-level separation architecture to split an ODI into different latitude bands, and hierarchically upscale them with different factors. In addition, we propose a deep reinforcement learning scheme with a latitude adaptive reward, in order to automatically select optimal upscaling factors for different latitude bands. To the best of our knowledge, LAU-Net is the first attempt to consider the latitude difference for ODI super-resolution. Extensive results demonstrate that our LAU-Net significantly advances the super-resolution performance for ODIs. Codes are available at https://github.com/wangh-allen/LAU-Net.
Xin Deng 0002, Hao Wang 0049, Mai Xu, Yuhang Song 0001, Li Yang 0014
CVPR3
2021 Deep Homography for Efficient Stereo Image Compression
abstract
In this paper, we propose HESIC, an end-to-end trainable deep network for stereo image compression (SIC). To fully explore the mutual information across two stereo images, we use a deep regression model to estimate the homography matrix, i.e., H matrix. Then, the left image is spatially transformed by the H matrix, and only the residual information between the left and right images is encoded to save bitrates. A two-branch auto-encoder architecture is adopted in HESIC, corresponding to the left and right images, respectively. For entropy coding, we use two conditional stereo entropy models, i.e., Gaussian mixture model (GMM) based and context based entropy models, to fully explore the correlation between the two images to reduce the coding bit-rates. In decoding, a cross quality enhancement module is proposed to enhance the image quality based on inverse H matrix. Experimental results show that our HESIC outperforms state-of-the-art SIC methods on InStereo2K and KITTI datasets both quantitatively and qualitatively. Code is available at https://github.com/ywz978020607/HESIC.
Xin Deng 0002, Mai Xu, Enpeng Liu, Qianhan Feng, Radu Timofte
CVPR4
2021 Saliency-Guided Image Translation
abstract
In this paper, we propose a novel task for saliency-guided image translation, with the goal of image-to-image translation conditioned on the user specified saliency map. To address this problem, we develop a novel Generative Adversarial Network (GAN)-based model, called SalG-GAN. Given the original image and target saliency map, SalG-GAN can generate a translated image that satisfies the target saliency map. In SalG-GAN, a disentangled representation framework is proposed to encourage the model to learn diverse translations for the same target saliency condition. A saliency-based attention module is introduced as a special attention mechanism for facilitating the developed structures of saliency-guided generator, saliency cue encoder and saliency-guided global and local discriminators. Furthermore, we build a synthetic dataset and a real-world dataset with labeled visual attention for training and evaluating our SalG-GAN. The experimental results over both datasets verify the effectiveness of our model for saliency-guided image translation.
Lai Jiang 0004, Mai Xu, Xiaofei Wang 0004, Leonid Sigal
CVPR2
2021 A Viewport-Adaptive Rate Control Approach for Omnidirectional Video Coding
abstract
For omnidirectional videos (ODVs), the existing off-line coding approaches are designed based on the spatial or perceptual distortion in a whole ODV frame, ignoring the fact that subjects can only access viewports. To improve the subjective quality inside the viewports, this paper proposes an off-line viewport-adaptive rate control (RC) approach for ODVs in high efficiency video coding (HEVC) framework. Specifically, we predict the viewport candidates with importance weights and develop a viewport saliency detection model. Then, the predicted candidates and detected saliency are taken into account in our viewport-adaptive CTU traversal and bit allocation scheme. Finally, the experimental results validate that our approach is effective in saving bit-rates and improving subjective quality for encoding ODVs; meanwhile, our approach is also effective in the auxiliary task of saliency detection in viewports.
Mai Xu, Li Yang 0014
DCC2
2021 HiNet: Deep Image Hiding by Invertible Network
abstract
Image hiding aims to hide a secret image into a cover image in an imperceptible way, and then recover the secret image perfectly at the receiver end. Capacity, invisibility and security are three primary challenges in image hiding task. This paper proposes a novel invertible neural network (INN) based framework, HiNet, to simultaneously overcome the three challenges in image hiding. For large capacity, we propose an inverse learning mechanism by simultaneously learning the image concealing and revealing processes. Our method is able to achieve the concealing of a full-size secret image into a cover image with the same size. For high invisibility, instead of pixel domain hiding, we propose to hide the secret information in wavelet domain. Furthermore, we propose a new low-frequency wavelet loss to constrain that secret information is hidden in high-frequency wavelet subbands, which significantly improves the hiding security. Experimental results show that our HiNet significantly outperforms other state-of-the-art image hiding methods, with more than 10 dB PSNR improvement in secret image recovery on ImageNet, COCO and DIV2K datasets. Codes are available at https://github.com/TomTomTommi/HiNet.
Junpeng Jing, Xin Deng 0002, Mai Xu, Jianyi Wang, Zhenyu Guan 0002
ICCV3
2021 CU-Net+: Deep Fully Interpretable Network for Multi-Modal Image Restoration
abstract
The network interpretability is critical in computer vision related tasks, especially for tasks involving multiple modalities. For multi-modal image restoration, one recent method, CU-Net, introduces an interpretable network based on a multi-modal convolutional sparse coding model. However, its network architecture does not mimic in full the proposed sparse model. In this paper, we overcome the limitation of CU-Net by using recurrent scheme, and this leads to a fully interpretable network which we call CU-Net+. In addition, we relax the constraint on the number of common and unique features in CU-Net, for making it more consistent with real condition. The effectiveness of the proposed CU-Net+ is evaluated on RGB guided depth image super-resolution and flash guided non-flash image denoising tasks. The numerical results show that CU-Net+ outperforms other interpretable or non-interpretable methods, with 0.16 RMSE and 0.66 dB PSNR improvement over CU-Net for the two mentioned tasks, respectively. Code is available at https://git;hub.com/JingyiXu404/CU-Net;-plus.
Xin Deng 0002, Mai Xu, Pier Luigi Dragotti
ICIP3
2021 Spatial Attention-Based Non-Reference Perceptual Quality Prediction Network for Omnidirectional Images
abstract
Due to the strong correlation between visual attention and perceptual quality, many methods attempt to use human saliency information for image quality assessment. Although this mechanism can get good performance, the networks require human saliency labels, which is not easily accessible for omnidirectional images (ODI). To alleviate this issue, we propose a spatial attention-based perceptual quality prediction network for non-reference quality assessment on ODIs (SAP-net). Without any human saliency labels, our network can adaptively estimate human perceptual quality on impaired ODIs through a self-attention manner, which significantly promotes the prediction performance of quality scores. Moreover, our method greatly reduces the computational complexity in quality assessment task on ODIs. Extensive experiments validate that our network outperforms 9 state-of-the-art methods for quality assessment on ODIs. The dataset and code have been available on https://github.com/yanglixiaoshen/SAP-Net.
Li Yang 0014, Mai Xu, Xin Deng 0002
ICME2
2021 DeepVS2.0: A Saliency-Structured Deep Learning Method for Predicting Dynamic Visual Attention
Lai Jiang 0004, Mai Xu, Zulin Wang, Leonid Sigal
Int. J. Comput. Vis.2
2021 Using Minimum Component and CNN for Satellite Remote Sensing Image Cloud Detection
abstract
Cloud detection is an important part of remote sensing (RS) image preprocessing. For earth observation tasks, the reliability of RS images will be judged based on the presence of clouds. A large number of cloud detection methods have been developed. There are two difficulties for cloud detection. First, it is hard to detect thin clouds and ragged clouds. Second, clouds are hard to distinguish from photometrically similar regions, such as snow. The rise of deep learning has brought new methods to address the above problems. In this letter, we propose a novel end-to-end neural network that detects clouds without additional manual work. Furthermore, we develop an RGB minimum component transformation mechanism that is useful for discriminating clouds from snow. Moreover, we have made our data set public to help others for further research. Our method increases the precision of cloud detection to 93.73%.
Mai Xu, Qinpeng Li
IEEE Geosci. Remote. Sens. Lett.3
2021 MFQE 2.0: A New Approach for Multi-Frame Quality Enhancement on Compressed Video
abstract
The past few years have witnessed great success in applying deep learning to enhance the quality of compressed image/video. The existing approaches mainly focus on enhancing the quality of a single frame, not considering the similarity between consecutive frames. Since heavy fluctuation exists across compressed video frames as investigated in this paper, frame similarity can be utilized for quality enhancement of low-quality frames given their neighboring high-quality frames. This task is Multi-Frame Quality Enhancement (MFQE). Accordingly, this paper proposes an MFQE approach for compressed video, as the first attempt in this direction. In our approach, we first develop a Bidirectional Long Short-Term Memory (BiLSTM) based detector to locate Peak Quality Frames (PQFs) in compressed video. Then, a novel Multi-Frame Convolutional Neural Network (MF-CNN) is designed to enhance the quality of compressed video, in which the non-PQF and its nearest two PQFs are the input. In MF-CNN, motion between the non-PQF and PQFs is compensated by a motion compensation subnet. Subsequently, a quality enhancement subnet fuses the non-PQF and compensated PQFs, and then reduces the compression artifacts of the non-PQF. Also, PQF quality is enhanced in the same way. Finally, experiments validate the effectiveness and generalization ability of our MFQE approach in advancing the state-of-the-art quality enhancement of compressed video.
Zhenyu Guan 0002, Qunliang Xing, Mai Xu, Zulin Wang
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Deep Coupled Feedback Network for Joint Exposure Fusion and Image Super-Resolution
abstract
Nowadays, people are getting used to taking photos to record their daily life, however, the photos are actually not consistent with the real natural scenes. The two main differences are that the photos tend to have low dynamic range (LDR) and low resolution (LR), due to the inherent imaging limitations of cameras. The multi-exposure image fusion (MEF) and image super-resolution (SR) are two widely-used techniques to address these two issues. However, they are usually treated as independent researches. In this paper, we propose a deep Coupled Feedback Network (CF-Net) to achieve MEF and SR simultaneously. Given a pair of extremely over-exposed and under-exposed LDR images with low-resolution, our CF-Net is able to generate an image with both high dynamic range (HDR) and high-resolution. Specifically, the CF-Net is composed of two coupled recursive sub-networks, with LR over-exposed and under-exposed images as inputs, respectively. Each sub-network consists of one feature extraction block (FEB), one super-resolution block (SRB) and several coupled feedback blocks (CFB). The FEB and SRB are to extract high-level features from the input LDR image, which are required to be helpful for resolution enhancement. The CFB is arranged after SRB, and its role is to absorb the learned features from the SRBs of the two sub-networks, so that it can produce a high-resolution HDR image. We have a series of CFBs in order to progressively refine the fused high-resolution HDR image. Extensive experimental results show that our CF-Net drastically outperforms other state-of-the-art methods in terms of both SR accuracy and fusion performance. The software code is available here https://github.com/ytZhang99/CF-Net.
Xin Deng 0002, Mai Xu, Shuhang Gu, Yiping Duan
IEEE Trans. Image Process.3
2021 Patch-Wise Spatial-Temporal Quality Enhancement for HEVC Compressed Video
abstract
Recently, many deep learning based researches are conducted to explore the potential quality improvement of compressed videos. These methods mostly utilize either the spatial or temporal information to perform frame-level video enhancement. However, they fail in combining different spatial-temporal information to adaptively utilize adjacent patches to enhance the current patch and achieve limited enhancement performance especially on scene-changing and strong-motion videos. To overcome these limitations, we propose a patch-wise spatial-temporal quality enhancement network which firstly extracts spatial and temporal features, then recalibrates and fuses the obtained spatial and temporal features. Specifically, we design a temporal and spatial-wise attention-based feature distillation structure to adaptively utilize the adjacent patches for distilling patch-wise temporal features. For adaptively enhancing different patch with spatial and temporal information, a channel and spatial-wise attention fusion block is proposed to achieve patch-wise recalibration and fusion of spatial and temporal features. Experimental results demonstrate our network achieves peak signal-to-noise ratio improvement, 0.55 - 0.69 dB compared with the compressed videos at different quantization parameters, outperforming state-of-the-art approach.
Liquan Shen, Liangwei Yu, Hao Yang 0008, Mai Xu
IEEE Trans. Image Process.5
2021 DeepQTMT: A Deep Learning Approach for Fast QTMT-Based CU Partition of Intra-Mode VVC
abstract
Versatile Video Coding (VVC), as the latest standard, significantly improves the coding efficiency over its predecessor standard High Efficiency Video Coding (HEVC), but at the expense of sharply increased complexity. In VVC, the quad-tree plus multi-type tree (QTMT) structure of the coding unit (CU) partition accounts for over 97% of the encoding time, due to the brute-force search for recursive rate-distortion (RD) optimization. Instead of the brute-force QTMT search, this paper proposes a deep learning approach to predict the QTMT-based CU partition, for drastically accelerating the encoding process of intra-mode VVC. First, we establish a large-scale database containing sufficient CU partition patterns with diverse video content, which can facilitate the data-driven VVC complexity reduction. Next, we propose a multi-stage exit CNN (MSE-CNN) model with an early-exit mechanism to determine the CU partition, in accord with the flexible QTMT structure at multiple stages. Then, we design an adaptive loss function for training the MSE-CNN model, synthesizing both the uncertain number of split modes and the target on minimized RD cost. Finally, a multi-threshold decision scheme is developed, achieving a desirable trade-off between complexity and RD performance. The experimental results demonstrate that our approach can reduce the encoding time of VVC by 44.65%~66.88% with a negligible Bjøntegaard delta bit-rate (BD-BR) of 1.322%~3.188%, significantly outperforming other state-of-the-art approaches.
Tianyi Li 0004, Mai Xu, Runzhi Tang, Ying Chen 0011, Qunliang Xing
IEEE Trans. Image Process.2
2021 Semantic Perceptual Image Compression With a Laplacian Pyramid of Convolutional Networks
abstract
The existing image compression methods usually choose or optimize low-level representation manually. Actually, these methods struggle for the texture restoration at low bit rates. Recently, deep neural network (DNN)-based image compression methods have achieved impressive results. To achieve better perceptual quality, generative models are widely used, especially generative adversarial networks (GAN). However, training GAN is intractable, especially for high-resolution images, with the challenges of unconvincing reconstructions and unstable training. To overcome these problems, we propose a novel DNN-based image compression framework in this paper. The key point is decomposing an image into multi-scale sub-images using the proposed Laplacian pyramid based multi-scale networks. For each pyramid scale, we train a specific DNN to exploit the compressive representation. Meanwhile, each scale is optimized with different aspects, including pixel, semantics, distribution and entropy, for a good "rate-distortion-perception" trade-off. By independently optimizing each pyramid scale, we make each stage manageable and make each sub-image plausible. Experimental results demonstrate that our method achieves state-of-the-art performance, with advantages over existing methods in providing improved visual quality. Additionally, a better performance in the down-stream visual analysis tasks which are conducted on the reconstructed images, validates the excellent semantics-preserving ability of the proposed method.
Juan Wang 0012, Yiping Duan, Xiaoming Tao 0001, Mai Xu, Jianhua Lu
IEEE Trans. Image Process.4
2021 Saliency Prediction on Omnidirectional Image With Generative Adversarial Imitation Learning
abstract
When watching omnidirectional images (ODIs), subjects can access different viewports by moving their heads. Therefore, it is necessary to predict subjects' head fixations on ODIs. Inspired by generative adversarial imitation learning (GAIL), this paper proposes a novel approach to predict saliency of head fixations on ODIs, named SalGAIL. First, we establish a dataset for attention on ODIs (AOI). In contrast to traditional datasets, our AOI dataset is large-scale, which contains the head fixations of 30 subjects viewing 600 ODIs. Next, we mine our AOI dataset and discover three findings: (1) the consistency of head fixations are consistent among subjects, and it grows alongside the increased subject number; (2) the head fixations exist with a front center bias (FCB); and (3) the magnitude of head movement is similar across the subjects. According to these findings, our SalGAIL approach applies deep reinforcement learning (DRL) to predict the head fixations of one subject, in which GAIL learns the reward of DRL, rather than the traditional human-designed reward. Then, multi-stream DRL is developed to yield the head fixations of different subjects, and the saliency map of an ODI is generated via convoluting predicted head fixations. Finally, experiments validate the effectiveness of our approach in predicting saliency maps of ODIs, significantly better than 11 state-of-the-art approaches. Our AOI dataset and code of SalGAIL are available online at https://github.com/yanglixiaoshen/SalGAIL.
Mai Xu, Li Yang 0014, Xiaoming Tao 0001, Yiping Duan, Zulin Wang
IEEE Trans. Image Process.1
2021 Joint Learning of 3D Lesion Segmentation and Classification for Explainable COVID-19 Diagnosis
abstract
Given the outbreak of COVID-19 pandemic and the shortage of medical resource, extensive deep learning models have been proposed for automatic COVID-19 diagnosis, based on 3D computed tomography (CT) scans. However, the existing models independently process the 3D lesion segmentation and disease classification, ignoring the inherent correlation between these two tasks. In this paper, we propose a joint deep learning model of 3D lesion segmentation and classification for diagnosing COVID-19, called DeepSC-COVID, as the first attempt in this direction. Specifically, we establish a large-scale CT database containing 1,805 3D CT scans with fine-grained lesion annotations, and reveal 4 findings about lesion difference between COVID-19 and community acquired pneumonia (CAP). Inspired by our findings, DeepSC-COVID is designed with 3 subnets: a cross-task feature subnet for feature extraction, a 3D lesion subnet for lesion segmentation, and a classification subnet for disease diagnosis. Besides, the task-aware loss is proposed for learning the task interaction across the 3D lesion and classification subnets. Different from all existing models for COVID-19 diagnosis, our model is interpretable with fine-grained 3D lesion distribution. Finally, extensive experimental results show that the joint learning framework in our model significantly improves the performance of 3D lesion segmentation and disease classification in both efficiency and efficacy.
Xiaofei Wang 0004, Lai Jiang 0004, Liu Li 0001, Mai Xu, Xin Deng 0002, Lisong Dai, Tianyi Li 0004, Zulin Wang, Pier Luigi Dragotti
IEEE Trans. Medical Imaging4
2021 Viewport-Dependent Saliency Prediction in 360° Video
abstract
Saliency prediction in traditional images and videos has drawn extensive research interests in recent years. Few works have been proposed for saliency prediction over 360° videos. They focus on directly predicting fixations over the whole panorama. When viewing 360° videos, a person can only observe the content in her viewport, which means that only a fraction of the 360° scene can be seen at any given time. In this paper, we study human attention over viewport of 360° videos and propose a novel visual saliency model, dubbed viewport saliency, to predict fixations over 360° videos. Two contributions are introduced. First, we find that where people look is affected by the content and location of the viewport in 360° video. We study this over 200+ 360° videos viewed by 30+ subjects over two recent benchmark databases. Second, we propose a Multi-Task Deep Neural Network (MT-DNN) method for Viewport Saliency (VS) prediction in 360° video, which considers the input content and location of the viewport. Extensive experiments and analyses show that our method outperforms other state-of-the-art methods in this task. In particular, over the two recent 360° video databases, our MT-DNN raises the average CC score by 0.149 and 0.205, compared to SalGAN and DeepVS methods, respectively.
Minglang Qiao, Mai Xu, Zulin Wang, Ali Borji
IEEE Trans. Multim.2
2021 Attention-Based Deep Reinforcement Learning for Virtual Cinematography of 360$^{\circ}$ Videos
abstract
Virtual cinematography refers to automatically selecting a natural-looking normal field-of-view (NFOV) from an entire 360$^{\circ}$video. In fact, virtual cinematography can be modeled as a deep reinforcement learning (DRL) problem, in which an agent makes actions related to NFOV selection according to the environment of 360$^{\circ}$video frames. More importantly, we find from our data analysis that the selected NFOVs attract significantly more attention than other regions, i.e., the NFOVs have high saliency. Therefore, in this paper, we propose an attention-based DRL (A-DRL) approach for virtual cinematography in 360$^{\circ}$video. Specifically, we develop a new DRL framework for automatic NFOV selection with the input of both the content, and saliency map of each 360$^{\circ}$frame. Then, we propose a new reward function for the DRL framework in our approach, which considers the saliency values, ground-truth, and smooth transition for NFOV selection. Subsequently, a simplified DenseNet (called Mini-DenseNet) is designed to learn the optimal policy via maximizing the reward. Based on the learned policy, the actions of NFOV can be made in our A-DRL approach for virtual cinematography of 360$^{\circ}$video. Extensive experiments show that our A-DRL approach outperforms other state-of-the-art virtual cinematography methods, over the datasets of Sports-360 video, and Pano2Vid.
Jianyi Wang, Mai Xu, Lai Jiang 0004, Yuhang Song 0001
IEEE Trans. Multim.2
2020 Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent Intelligence
abstract
Learning agents that are not only capable of taking tests, but also innovating is becoming a hot topic in AI. One of the most promising paths towards this vision is multi-agent learning, where agents act as the environment for each other, and improving each agent means proposing new problems for others. However, existing evaluation platforms are either not compatible with multi-agent settings, or limited to a specific game. That is, there is not yet a general evaluation platform for research on multi-agent intelligence. To this end, we introduce Arena, a general evaluation platform for multi-agent intelligence with 35 games of diverse logics and representations. Furthermore, multi-agent intelligence is still at the stage where many problems remain unexplored. Therefore, we provide a building toolkit for researchers to easily invent and build novel multi-agent problems from the provided game set based on a GUI-configurable social tree and five basic multi-agent reward schemes. Finally, we provide Python implementations of five state-of-the-art deep multi-agent reinforcement learning baselines. Along with the baseline implementations, we release a set of 100 best agents/teams that we can train with different training schemes for each game, as the base for evaluating agents with population performance. As such, the research community can perform comparisons under a stable and uniform standard. All the implementations and accompanied tutorials have been open-sourced for the community at https://sites.google.com/view/arena-unity/.
Yuhang Song 0001, Andrzej Wojcicki, Thomas Lukasiewicz, Jianyi Wang, Abi Aryan, Zhenghua Xu 0001, Mai Xu, Lianlong Wu
AAAI7
2020 Mega-Reward: Achieving Human-Level Play without Extrinsic Rewards
abstract
Intrinsic rewards were introduced to simulate how human intelligence works; they are usually evaluated by intrinsically-motivated play, i.e., playing games without extrinsic rewards but evaluated with extrinsic rewards. However, none of the existing intrinsic reward approaches can achieve human-level performance under this very challenging setting of intrinsically-motivated play. In this work, we propose a novel megalomania-driven intrinsic reward (called mega-reward), which, to our knowledge, is the first approach that achieves human-level performance in intrinsically-motivated play. Intuitively, mega-reward comes from the observation that infants' intelligence develops when they try to gain more control on entities in an environment; therefore, mega-reward aims to maximize the control capabilities of agents on given entities in a given environment. To formalize mega-reward, a relational transition model is proposed to bridge the gaps between direct and latent control. Experimental studies show that mega-reward (i) can greatly outperform all state-of-the-art intrinsic reward approaches, (ii) generally achieves the same level of performance as Ex-PPO and professional human-level scores, and (iii) has also a superior performance when it is incorporated with extrinsic rewards.
Yuhang Song 0001, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu 0001, Shangtong Zhang, Andrzej Wojcicki, Mai Xu
AAAI7
2020 Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model
Yufan Liu 0001, Minglang Qiao, Mai Xu, Bing Li 0001, Weiming Hu 0004, Ali Borji
ECCV (20)3
2020 Multi-level Wavelet-Based Generative Adversarial Network for Perceptual Quality Enhancement of Compressed Video
Jianyi Wang, Xin Deng 0002, Mai Xu, Congyong Chen, Yuhang Song 0001
ECCV (14)3
2020 Early Exit or Not: Resource-Efficient Blind Quality Enhancement for Compressed Images
Qunliang Xing, Mai Xu, Tianyi Li 0004, Zhenyu Guan 0002
ECCV (16)2
2020 Learning Diverse Sub-Policies via a Task-Agnostic Regularization on Action Distributions
abstract
Automatic sub-policy discovery has recently received much attention in hierarchical reinforcement learning (HRL). The conventional approaches to learning sub-policies suffer from collapsing into just one sub-policy dominating the whole task, lacking techniques to ensure the diversity of different subpolicies. In this paper, we formulate the discovery of diverse sub-policies as a trajectory inference. Then, we propose an information-theoretic objective based on action distributions to encourage diversity. Moreover, two simplifications are derived on discrete and continuous action space for reducing the computation. Finally, the experimental results show that the proposed approach can further improve the state-of-theart approaches without modifying existing hyperparameters on two different HRL domains, suggesting the wide applicability and robustness of our approach.
Liangyu Huo, Zulin Wang, Mai Xu, Yuhang Song 0001
ICASSP3
2020 DeepGF: Glaucoma Forecast Using the Sequential Fundus Images
Liu Li 0001, Xiaofei Wang 0004, Mai Xu, Hanruo Liu, Ximeng Chen
MICCAI (5)3
2020 MRS-Net: Multi-Scale Recurrent Scalable Network for Face Quality Enhancement of Compressed Videos
abstract
The past decade has witnessed the explosive growth of faces in video multimedia systems, e.g., videoconferencing and live shows. However, these videos are normally compressed at low bit-rates due to the bandwidth-hungry issue, leading to heavy quality degradation on face regions. This paper addresses the problem of face quality enhancement in compressed videos. Specifically, we establish a compressed face video (CFV) database, which includes 87,607 faces in 113 raw video sequences and their corresponding 904 compressed sequences. We find that the faces of compressed videos exhibit tremendous scale variation and quality fluctuation. Motivated by scalable video coding, we propose a multi-scale recurrent scalable network (MRS-Net) to enhance the quality of multi-scale faces in compressed videos. The MRS-Net is comprised by one base and two refined enhancement levels, corresponding to the quality enhancement of small-, medium- and large-scale faces, respectively. In the multi-level architecture of our MRS-Net, small-/medium-scale face quality enhancement serves as the basis for facilitating the quality enhancement of medium-/large-scale faces. Finally, experimental results show that our MRS-Net method is effective in enhancing the quality of multi-scale faces for compressed videos, significantly outperforming other state-of-the-art methods.
Mai Xu, Shengxi Li, Huaida Liu
ACM Multimedia2
2020 DeepCT: A novel deep complex-valued network with learnable transform for video saliency prediction
Lai Jiang 0004, Mai Xu, Shanyi Zhang, Leonid Sigal
Pattern Recognit.2
2020 VRFCNN: Virtual Reference Frame Generation Network for Quality SHVC
abstract
For more efficient inter prediction in quality scalable high efficiency video coding (SHVC), a learning-based framework for generating a virtual reference frame (VRF) is proposed in this letter. In our method, reconstructed base layer (BL) and enhancement layer (EL) frames are employed to make the generated VRF and current EL frame as same as possible. To this end, this letter proposes a novel VRF generation convolutional neural network (VRFCNN) to jointly handle enhancement of corresponding BL frame and compensation of previous EL frame. Specifically, the VRFCNN consists of BL enhancement, EL compensation and feature fusion subnets. Previous EL frames are firstly compensated with the learned coarse flow between two adjacent BL frames and then adopted to provide interlayer information for corresponding BL enhancement. The learned finer flow between the EL and enhanced BL features is adopted to provide temporal information for previous EL compensation. For efficiently handling slow- and fast-motion videos, the enhanced BL and compensated EL features are fused to generate a VRF. Experimental results show that VRFCNN averagely achieves 11.8% BD-rate reduction under low delay P configuration, which outperforms other methods. The code of our VRFCNN approach is available at https://github.com/dq0309/VRFCNN.
Liquan Shen, Hao Yang 0008, Xinchao Dong, Mai Xu
IEEE Signal Process. Lett.5
2020 Toward Variable-Rate Generative Compression by Reducing the Channel Redundancy
abstract
Compressing large images with a generative model goes beyond typical image encoding standards under a notably low bitrate. In this paper, we step toward practical generative compression systems based on recent advances. Specifically, we show that the channel redundancy of the latent representation produced by an autoencoder network can be effectively compressed via mask compression. The mask compression performs quantization on the channel variance of latent representation instead of original values. Instead of training multiple models, changing the mask leads to a simple and efficient variable rate compression scheme. Then, we estimate the relative bitrate by measuring the L1 norm of the channel variance and hence obtain the rate-distortion formulation. The L1 regularizer assumes a Laplacian prior on the channel variance, through which model we develop corresponding methods to produce approximate images at a target bitrate. This eliminates the need for manually searching hyperparameters for our variable-rate compression. We conduct exhaustive experiments to demonstrate the advanced performance of the proposed method in preserving image quality and semantics.
Chaoyi Han, Yiping Duan, Xiaoming Tao 0001, Mai Xu, Jianhua Lu
IEEE Trans. Circuits Syst. Video Technol.4
2020 A Meta-Learning Framework for Learning Multi-User Preferences in QoE Optimization of DASH
abstract
Dynamic adaptive video streaming over hypertext transfer protocol (DASH) plays a key role in video transmission over the Internet. The conventional DASH adaptation approaches mainly focus on optimizing the overall quality of experience (QoE) for all client sides, neglecting the QoE diversity of different users. In this paper, we propose a meta-learning framework for multi-user preferences (MLMP) as a new DASH adaptation approach, which is able to optimize the diverse QoE of different users. Specifically, we first design a subjective experiment to analyze the difference of QoE preferences across users, in which QoE refers to the metrics of visual quality, fluctuation, and rebuffering events. Based on our findings, we formulate the QoE optimization of multi-user preferences as a multi-task deep reinforcement learning (DRL) problem. In our formulation, the QoE preference of each user is modeled in the overall QoE calculation via assigning the weights to the three QoE metrics. Then, the MLMP framework is developed to solve the proposed multi-task DRL problem, such that the preferences regarding visual quality, fluctuation, and rebuffering events can be optimized for different users in DASH adaptation. Finally, the simulation results show that the proposed approach outperforms state-of-the-art DASH adaptation approaches in satisfying the different users' QoE preferences regarding visual quality, fluctuation, and rebuffering events.
Liangyu Huo, Zulin Wang, Mai Xu, Yong Li 0008, Zhiguo Ding 0001, Hao Wang 0049
IEEE Trans. Circuits Syst. Video Technol.3
2020 Accelerate CTU Partition to Real Time for HEVC Encoding With Complexity Control
abstract
Recently, extensive approaches have been proposed for reducing the encoding complexity of high efficiency video coding, by predicting the coding tree unit partition using deep neural networks. However, these approaches cannot work in real time due to the complexity of the network architectures. In this paper, we propose a network pruning approach to accelerate a state-of-the-art deep neural network model, for real-time coding tree unit partition. Specifically, we first investigate the computational complexity throughout the network, and find that most calculations can be simplified by pruning the weight parameters. Considering that the number of weight parameters drastically differs by network layer and partition level, we design an adaptive pruning scheme by applying a well-suitable retention ratio of weight parameters to each layer at a level. The retention ratio indicates the ratio of weight parameters after and before pruning. By varying the retention ratios, we can obtain several accelerated network models with different levels of complexity. We further propose a complexity control algorithm by applying different accelerated models to different coding tree units, to ensure that the actual encoding complexity is close to a given target. To guarantee the rate-distortion performance, we model the complexity control algorithm as a convex optimization problem, and we can obtain a closed-form solution. Experimental results show that our approach can accelerate the original deep neural network model by 17-20 times, with little expense on the Bjøntegaard delta bit-rate. For complexity control, we achieve high control accuracy with a control error of less than 2% for most video sequences.
Tianyi Li 0004, Mai Xu, Xin Deng 0002, Liquan Shen
IEEE Trans. Image Process.2
2020 Model-Free Distortion Rectification Framework Bridged by Distortion Distribution Map
abstract
Recently, learning-based distortion rectification schemes have shown high efficiency. However, most of these methods only focus on a specific camera model with fixed parameters, thus failing to be extended to other models. To avoid such a disadvantage, we propose a model-free distortion rectification framework for the single-shot case, bridged by the distortion distribution map (DDM). Our framework is based on an observation that the pixel-wise distortion information is mathematically regular in a distorted image, despite different models having different types and numbers of distortion parameters. Motivated by this observation, instead of estimating the heterogeneous distortion parameters, we construct a proposed distortion distribution map that intuitively indicates the global distortion features of a distorted image. In addition, we develop a dual-stream feature learning module, benefitting from both the advantages of traditional methods that leverage the local handcrafted feature and learning-based methods that focus on the global semantic feature perception. Due to the sparsity of handcrafted features, we discrete the features into a 2D point map and learn the structure inspired by PointNet. Finally, a multimodal attention fusion module is designed to attentively fuse the local structural and global semantic features, providing the hybrid features for the more reasonable scene recovery. The experimental results demonstrate the excellent generalization ability and more significant performance of our method in both quantitative and qualitative evaluations, compared with the stateof- the-art methods.
Kang Liao, Chunyu Lin, Yao Zhao 0001, Mai Xu
IEEE Trans. Image Process.4
2020 Understanding and Predicting the Memorability of Outdoor Natural Scenes
abstract
Memorability measures how easily an image is to be memorized after glancing, which may contribute to designing magazine covers, tourism publicity materials, and so forth. Recent works have shed light on the visual features that make generic images, object images or face photographs memorable. However, these methods are not able to effectively predict the memorability of outdoor natural scene images. To overcome this shortcoming of previous works, in this paper, we provide an attempt to answer: “what exactly makes outdoor natural scenes memorable”. To this end, we first establish a large-scale outdoor natural scene image memorability (LNSIM) database, containing 2,632 outdoor natural scene images with their ground truth memorability scores and the multi-label scene category annotations. Then, similar to previous works, we mine our database to investigate how low-, middle- and high-level handcrafted features affect the memorability of outdoor natural scenes. In particular, we find that the high-level feature of scene category is rather correlated with outdoor natural scene memorability, and the deep features learnt by deep neural network (DNN) are also effective in predicting the memorability scores. Moreover, combining the deep features with the category feature can further boost the performance of memorability prediction. Therefore, we propose an end-to-end DNN based outdoor natural scene memorability (DeepNSM) predictor, which takes advantage of the learned category-related features. Then, the experimental results validate the effectiveness of our DeepNSM model, exceeding the state-of-the-art methods. Finally, we try to understand the reason of the good performance for our DeepNSM model, and also study the cases that our DeepNSM model succeeds or fails to accurately predict the memorability of outdoor natural scenes.
Jiaxin Lu 0003, Mai Xu, Zulin Wang
IEEE Trans. Image Process.2
2020 A Large-Scale Database and a CNN Model for Attention-Based Glaucoma Detection
abstract
Glaucoma is one of the leading causes of irreversible vision loss. Many approaches have recently been proposed for automatic glaucoma detection based on fundus images. However, none of the existing approaches can efficiently remove high redundancy in fundus images for glaucoma detection, which may reduce the reliability and accuracy of glaucoma detection. To avoid this disadvantage, this paper proposes an attention-based convolutional neural network (CNN) for glaucoma detection, called AG-CNN. Specifically, we first establish a large-scale attention-based glaucoma (LAG) database, which includes 11 760 fundus images labeled as either positive glaucoma (4878) or negative glaucoma (6882). Among the 11 760 fundus images, the attention maps of 5824 images are further obtained from ophthalmologists through a simulated eye-tracking experiment. Then, a new structure of AG-CNN is designed, including an attention prediction subnet, a pathological area localization subnet, and a glaucoma classification subnet. The attention maps are predicted in the attention prediction subnet to highlight the salient regions for glaucoma detection, under a weakly supervised training manner. In contrast to other attention-based CNN methods, the features are also visualized as the localized pathological area, which are further added in our AG-CNN structure to enhance the glaucoma detection performance. Finally, the experiment results from testing over our LAG database and another public glaucoma database show that the proposed AG-CNN approach significantly advances the state-of-the-art in glaucoma detection.
Liu Li 0001, Mai Xu, Hanruo Liu, Yang Li 0010, Xiaofei Wang 0004, Lai Jiang 0004, Zulin Wang, Ningli Wang
IEEE Trans. Medical Imaging2
2020 An EEG-Based Study on Perception of Video Distortion Under Various Content Motion Conditions
abstract
Human perception sensitivity to video distortion is vital for visual quality assessment (VQA). Different from the perception mechanism of image distortion that has been thoroughly studied, the perception of video distortion is inevitably influenced by motion of dynamic content due to the characteristics of the human visual system (HVS). In this paper, electroencephalography (EEG) is used as a novel psychophysiological method to study the human perception sensitivity to quantification-aroused video distortion under various content motion conditions. For this purpose, we conduct experiments to record the EEG signals of the subjects when they are watching distorted videos. According to the feature analysis of EEG data, the P300 component aroused by human perception of video quality change is selected as the indicator of human perception of distortion. By the means of classification based on linear discriminant analysis (LDA), it is found that the separability of the P300 component, which is measured by the area under curve (AUC) of the receiver operating characteristic (ROC), is positively correlated with the perceptibility of distortion. The correlation provides a valid psychophysiological method, which is exempt from being influenced by subjective bias due to human high-level cognitive activities, for evaluating distortion perceptibility. In addition, the regression analysis results demonstrate a sigmoid-typed quantitative relation between the perceptibility of distortion and separability of the P300 component. Based on such relation, the perceptibility thresholds of distortion corresponding to various content motion speeds are calibrated by EEG signals and it is found that the content motion speed has a significant impact on distortion perceptibility.
Xiaoming Tao 0001, Mai Xu, Yafeng Zhan, Jianhua Lu
IEEE Trans. Multim.3
2019 Image Saliency Prediction in Transformed Domain: A Deep Complex Neural Network Method
abstract
The transformed domain fearures of images show effectiveness in distinguishing salient and non-salient regions. In this paper, we propose a novel deep complex neural network, named SalDCNN, to predict image saliency by learning features in both pixel and transformed domains. Before proposing Sal-DCNN, we analyze the saliency cues encoded in discrete Fourier transform (DFT) domain. Consequently, we have the following findings: 1) the phase spectrum encodes most saliency cues; 2) a certain pattern of the amplitude spectrum is important for saliency prediction; 3) the transformed domain spectrum is robust to noise and down-sampling for saliency prediction. According to these findings, we develop the structure of SalDCNN, including two main stages: the complex dense encoder and three-stream multi-domain decoder. Given the new SalDCNN structure, the saliency maps can be predicted under the supervision of ground-truth fixation maps in both pixel and transformed domains. Finally, the experimental results show that our Sal-DCNN method outperforms other 8 state-of-theart methods for image saliency prediction on 3 databases.
Lai Jiang 0004, Mai Xu, Zulin Wang
AAAI3
2019 Diversity-Driven Extensible Hierarchical Reinforcement Learning
abstract
Hierarchical reinforcement learning (HRL) has recently shown promising advances on speeding up learning, improving the exploration, and discovering intertask transferable skills. Most recent works focus on HRL with two levels, i.e., a master policy manipulates subpolicies, which in turn manipulate primitive actions. However, HRL with multiple levels is usually needed in many real-world scenarios, whose ultimate goals are highly abstract, while their actions are very primitive. Therefore, in this paper, we propose a diversitydriven extensible HRL (DEHRL), where an extensible and scalable framework is built and learned levelwise to realize HRL with multiple levels. DEHRL follows a popular assumption: diverse subpolicies are useful, i.e., subpolicies are believed to be more useful if they are more diverse. However, existing implementations of this diversity assumption usually have their own drawbacks, which makes them inapplicable to HRL with multiple levels. Consequently, we further propose a novel diversity-driven solution to achieve this assumption in DEHRL. Experimental studies evaluate DEHRL with nine baselines from four perspectives in two domains; the results show that DEHRL outperforms the state-of-the-art baselines in all four aspects.
Yuhang Song 0001, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu 0001, Mai Xu
AAAI5
2019 Viewport Proposal CNN for 360deg Video Quality Assessment
abstract
Recent years have witnessed the growing interest in visual quality assessment (VQA) for 360° video. Unfortunately, the existing VQA approaches do not consider the facts that: 1) Observers only see viewports of 360° video, rather than patches or whole 360° frames. 2) Within the viewport, only salient regions can be perceived by observers with high resolution. Thus, this paper proposes a viewport-based convolutional neural network (V-CNN) approach for VQA on 360° video, considering both auxiliary tasks of viewport proposal and viewport saliency prediction. Our V-CNN approach is composed of two stages, i.e., viewport proposal and VQA. In the first stage, the viewport proposal network (VP-net) is developed to yield several potential viewports, seen as the first auxiliary task. In the second stage, a viewport quality network (VQ-net) is designed to rate the VQA score for each proposed viewport, in which the saliency map of the viewport is predicted and then utilized in VQA score rating. Consequently, another auxiliary task of viewport saliency prediction can be achieved. More importantly, the main task of VQA on 360° video can be accomplished via integrating the VQA scores of all viewports. The experiments validate the effectiveness of our V-CNN approach in significantly advancing the state-of-the-art performance of VQA on 360° video. In addition, our approach achieves comparable performance in two auxiliary tasks. The code of our V-CNN approach is available at https://github.com/Archer-Tatsu/V-CNN.
Chen Li 0049, Mai Xu, Lai Jiang 0004, Shanyi Zhang, Xiaoming Tao 0001
CVPR2
2019 Attention Based Glaucoma Detection: A Large-Scale Database and CNN Model
abstract
Recently, the attention mechanism has been successfully applied in convolutional neural networks (CNNs), significantly boosting the performance of many computer vision tasks. Unfortunately, few medical image recognition approaches incorporate the attention mechanism in the CNNs. In particular, there exists high redundancy in fundus images for glaucoma detection, such that the attention mechanism has potential in improving the performance of CNN-based glaucoma detection. This paper proposes an attention-based CNN for glaucoma detection (AG-CNN). Specifically, we first establish a large-scale attention based glaucoma (LAG) database, which includes 5,824 fundus images labeled with either positive glaucoma (2,392) or negative glaucoma (3,432). The attention maps of the ophthalmologists are also collected in LAG database through a simulated eye-tracking experiment. Then, a new structure of AG-CNN is designed, including an attention prediction subnet, a pathological area localization subnet and a glaucoma classification subnet. Different from other attention-based CNN methods, the features are also visualized as the localized pathological area, which can advance the performance of glaucoma detection. Finally, the experiment results show that the proposed AG-CNN approach significantly advances state-of-the-art glaucoma detection.
Liu Li 0001, Mai Xu, Xiaofei Wang 0004, Lai Jiang 0004, Hanruo Liu
CVPR2
2019 A DenseNet Based Approach for Multi-frame In-loop Filter in HEVC
abstract
High efficiency video coding (HEVC) has brought outperforming efficiency for video compression. To reduce the compression artifacts of HEVC, we propose a DenseNet based approach as the in-loop filter of HEVC, which leverages multiple adjacent frames to enhance the quality of each encoded frame. Specifically, the higher-quality frames are found by a reference frame selector (RFS). Then, a deep neural network for multi-frame in-loop filter (named MIF-Net) is developed to enhance the quality of each encoded frame by utilizing the spatial information of this frame and the temporal information of its neighboring higher-quality frames. The MIF-Net is built on the recently developed DenseNet, benefiting from the improved generalization capacity and computational efficiency. Finally, experimental results verify the effectiveness of our multi-frame in-loop filter, outperforming the HM baseline and other state-of-the-art approaches.
Tianyi Li 0004, Mai Xu, Xiaoming Tao 0001
DCC2
2019 Texture-Classification Accelerated CNN Scheme for Fast Intra CU Partition in HEVC
abstract
High Efficiency Video Coding (HEVC) achieves significant coding performance over H.264. However, the performance gain is achieved at the cost of substantially higher encoding complexity, in which the coding tree unit (CTU) partition is one of the most time-consuming parts due to the rate-distortion optimization-based ergodic search of all possible quad-tree partitions. To address this problem, this paper proposes a texture-classification accelerated convolutional neural network (CNN)-based fast intra CU partition scheme to reduce the encoding complexity for intra-coding in HEVC, by taking into consideration of the heterogeneous texture characteristics into the CNN-based classification. First, a threshold-based texture classification model is developed to identify the heterogeneous and homogeneous CTUs, through jointly consideration of the CU depth, quantization parameter and texture complexity. Second, three different CNN structures are designed and trained to predict the CU partition mode for each CU layer in the heterogeneous CTUs. Finally, extensive experimental results show that the proposed scheme can reduce intra-mode encoding time by 62.13% with negligible BD-rate loss of 2.01%, consistently outperforming two state-of-the-art CNN-based schemes in terms of both coding performance and complexity reduction.
Yongfei Zhang, Gang Wang 0023, Mai Xu, C.-C. Jay Kuo
DCC4
2019 Optimizing QoE of Multiple Users over DASH: A Meta-learning Approach
abstract
Dynamic adaptive video streaming over HTTP (DASH) plays a key role in video transmission over the Internet. The conventional DASH adaptation approaches concentrate on optimizing the overall quality of experience (QoE) for all client sides, neglecting the QoE diversity of different users. In this paper, we formulate the QoE optimization of multi-user preferences as a multi-task deep reinforcement learning problem, in which QoE refers to the metrics of visual quality, fluctuation and rebuffing events. Then, we propose a meta-learning framework for multi-user preferences (MLMP) as a new DASH adaptation approach. Finally, the simulation results show that the proposed approach outperforms state-of-the-art DASH adaptation approaches in satisfying the different users' QoE preferences regarding the three metrics.
Liangyu Huo, Zulin Wang, Mai Xu, Zhiguo Ding 0001, Xiaoming Tao 0001
ICASSP3
2019 A Deep Neural Network Based Maneuvering-target Tracking Algorithm
abstract
In the field of maneuvering-target tracking (MTT), the targets with changeable and uncertain maneuvering movements cannot be tracked precisely because there always exist time delays of maneuvering model estimation with traditional MT-T algorithms. To solve this problem, we propose a deep MTT (DeepMTT) algorithm based on a deep neural network, which can quickly track maneuvering targets once it has been well trained by abundant off-line trajectory data from existent ma-neuvering targets. To this end, we first build a Large-scale trajectory database to offer abundant off-line trajectory data for network training. Second, the DeepMTT algorithm is developed based on a deep neural network, which consists of three bidirectional long short-term memory layers, a filtering layer, a maxout layer and a linear output layer. The simulation results verify that our DeepMTT algorithm outperforms other state-of-the-art MTT algorithms.
Jingxian Liu, Zulin Wang, Mai Xu, Jie Ren 0004
ICASSP3
2019 Unsupervised User Clustering in Non-orthogonal Multiple Access
abstract
Non-orthogonal multiple access (NOMA) is one of the most promising technologies in fifth-generation mobile communication system for its advantages in serving multiuser simultaneously and enhancing spectrum efficiency. In this paper, we investigate the optimization problem of sum-rate maximization for NOMA-based system, and mainly focus on user clustering. Inspired by the correlation features of users, we introduce machine learning in user clustering. We first develop an expectation maximization (EM) based algorithm for fixed user scenario. Then, the dynamic user scenario is considered and an online EM (OLEM) based clustering algorithm is proposed. Simulation results show that the proposed EM-based and OLEM-based algorithms outperform the state-of-the-art algorithms in fixed and dynamic user scenario, respectively.
Jie Ren 0004, Zulin Wang, Mai Xu, Fang Fang 0005, Zhiguo Ding 0001
ICASSP3
2019 Wavelet Domain Style Transfer for an Effective Perception-Distortion Tradeoff in Single Image Super-Resolution
abstract
In single image super-resolution (SISR), given a low-resolution (LR) image, one wishes to find a high-resolution (HR) version of it which is both accurate and photorealistic. Recently, it has been shown that there exists a fundamental tradeoff between low distortion and high perceptual quality, and the generative adversarial network (GAN) is demonstrated to approach the perception-distortion (PD) bound effectively. In this paper, we propose a novel method based on wavelet domain style transfer (WDST), which achieves a better PD tradeoff than the GAN based methods. Specifically, we propose to use 2D stationary wavelet transform (SWT) to decompose one image into low-frequency and high-frequency sub-bands. For the low-frequency sub-band, we improve its objective quality through an enhancement network. For the high-frequency sub-band, we propose to use WDST to effectively improve its perceptual quality. By feat of the perfect reconstruction property of wavelets, these sub-bands can be re-combined to obtain an image which has simultaneously high objective and perceptual quality. The numerical results on various datasets show that our method achieves the best trade-off between the distortion and perceptual quality among the existing state-of-the-art SISR methods.
Xin Deng 0002, Mai Xu, Pier Luigi Dragotti
ICCV3
2019 Semantic Perceptual Image Compression with a Laplacian Pyramid of Convolutional Networks
abstract
Recently, deep neural network (DNN)-based image compression methods have achieved impressive results. These methods generally use thumbnail images or crop small patches from high-resolution images to train their networks. Instead of using patch-based training mode, we propose a novel DNN-based image compression framework in this paper. We apply the Laplacian pyramid to construct a multi-scale image representation. By learning the increasingly detailed representations, the proposed method is able to progressively restore an image. Furthermore, we use the adversarial networks for training to encourage the perceptual quality of the reconstructed image. Particularly at low bitrates, our model can only store the global semantics of an image and automatically synthesize the texture to achieve high subjective quality. Experimental results on demonstrate that our method achieves state-of-the-art performance, with advantages over existing methods in terms of visual quality.
Juan Wang 0012, Xiaoming Tao 0001, Mai Xu, Jianhua Lu
ICIP3
2019 Removing Rain in Videos: A Large-Scale Database and a Two-Stream ConvLSTM Approach
abstract
Rain removal has recently attracted increasing research attention, as it is able to enhance the visibility of rain videos. However, the existing learning based rain removal approaches for videos suffer from insufficient training data, especially when applying deep learning to remove rain. In this paper, we establish a large-scale video database for rain removal (LasVR), which consists of 316 rain videos. Then, we observe from our database that there exist the temporal correlation of clean content and similar patterns of rain across video frames. According to these two observations, we propose a two-stream convolutional long-and short-term memory (ConvLSTM) approach for rain removal in videos. The first stream is composed of the subnet for rain detection, while the second stream is the subnet of rain removal that leverages the features from the rain detection subnet. Finally, the experimental results on both synthetic and real rain videos show the proposed approach performs better than other state-of-the-art approaches.
Mai Xu, Zulin Wang
ICME2
2019 Quality-Gated Convolutional Lstm for Enhancing Compressed Video
abstract
The past decade has witnessed great success in applying deep learning to enhance the quality of compressed video. However, the existing approaches aim at quality enhancement on a single frame, or only using fixed neighboring frames. Thus they fail to take full advantage of the inter-frame correlation in the video. This paper proposes the Quality-Gated Convolutional Long Short-Term Memory (QG-ConvLSTM) network with bi-directional recurrent structure to fully exploit the advantageous information in a large range of frames. More importantly, due to the obvious quality fluctuation among compressed frames, higher quality frames can provide more useful information for other frames to enhance quality. Therefore, we propose learning the "forget" and "'input" gates in the ConvLSTM cell from quality-related features. As such, the frames with various quality contribute to the memory in ConvLSTM with different importance, making the information of each frame reasonably and adequately used. Finally, the experiments validate the effectiveness of our QG-ConvLSTM approach in advancing the state-of-the-art quality enhancement of compressed video, and the ablation study shows that our QG-ConvLSTM approach is learnt to make a trade-off between quality and correlation when leveraging multi-frame information. The project page: https://github.com/ryangchn/QG-ConvLSTM.git.
Xiaoyan Sun 0001, Mai Xu, Wenjun Zeng 0001
ICME3
2019 Pathology-Aware Deep Network Visualization and Its Application in Glaucoma Image Synthesis
Xiaofei Wang 0004, Mai Xu, Liu Li 0001, Zulin Wang, Zhenyu Guan 0002
MICCAI (1)2
2019 Rebuffering Optimization for DASH via Pricing and EEG-Based QoE Modeling
abstract
Pricing is an effective mechanism for network resource allocation that can be used to achieve a desirable balance between efficiency and fairness. However, sophisticated utility models are needed to guarantee the performance of price-based resource allocation, especially with regard to video transmission. Among various performance indices of video transmission, rebuffering is an important one that influences user quality of experience (QoE). Therefore, a price-based bandwidth allocation scheme for a dynamic adaptive streaming over hypertext transfer protocol (DASH) system is proposed for mitigating the effect of rebuffering on QoE. The utility model of the proposed scheme considers the relationship between the allocated bandwidth and the rebuffering length, as well as the effect of rebuffering length on QoE. Specifically, electroencephalography (EEG) experiments are conducted, and the distribution of the subjects' time limits at which rebuffering arouse negative emotions is used for calibration. Based on this model, the DASH server collects the buffer state information of all DASH clients periodically to adjust its bandwidth allocation. Assuming the server protects itself from congestion by pricing the clients' requested bandwidth, a Stackelberg game is formulated to study the joint utility maximization on the revenue of the server and the utility of the clients. The Stackelberg equilibrium of the game is characterized, and an efficient searching algorithm is proposed for its solution. EEG experiments are repeated on another group of subjects to verify the generalization ability of the results, and simulation results are presented to validate the effectiveness of the proposed algorithm. The proposed algorithm is shown to have low complexity and outperforms both traditional price-based scheme and QoE maximized scheme in terms of rebuffering.
Xiaoming Tao 0001, Zhao Chen 0002, Mai Xu, Jianhua Lu
IEEE J. Sel. Areas Commun.3
2019 Learning QoE of Mobile Video Transmission With Deep Neural Network: A Data-Driven Approach
abstract
Quality of experience (QoE) serves as a direct evaluation of users' experiences in mobile video transmission and thus essential for network management, such as network optimization. In this paper, we propose a deep learning-based QoE prediction approach with a large-scale QoE dataset for mobile video transmission. Specifically, we develop a mobile phone application for collecting user QoE data when viewing videos transmitted over the mobile internet in a practical environment. Then, we construct a large-scale dataset by collecting over 80000 piece of data with four kinds of subjective scores and 89 network parameters. Each QoE metric is related to only some of the 89 network parameters. Therefore, we apply the feature selection method to find the feature parameters related to user scores. Additionally, the boxplot method is used to clean the raw data by removing outliers. Finally, a deep neural network (DNN) is developed to learn the relationships between the network parameters and the subjective QoE scores. The proposed DNN can also be seen as a data-driven objective QoE prediction approach for mobile video transmission, which can be used to predict the user QoE scores. The experimental results show that the proposed approach can effectively remove most features irrelevant to QoE prediction. Moreover, the performance of QoE prediction by the proposed model outperforms other state-of-the-art approaches.
Xiaoming Tao 0001, Yiping Duan, Mai Xu, Zhishen Meng, Jianhua Lu
IEEE J. Sel. Areas Commun.3
2019 Predicting Head Movement in Panoramic Video: A Deep Reinforcement Learning Approach
abstract
Panoramic video provides immersive and interactive experience by enabling humans to control the field of view (FoV) through head movement (HM). Thus, HM plays a key role in modeling human attention on panoramic video. This paper establishes a database collecting subjects' HM in panoramic video sequences. From this database, we find that the HM data are highly consistent across subjects. Furthermore, we find that deep reinforcement learning (DRL) can be applied to predict HM positions, via maximizing the reward of imitating human HM scanpaths through the agent's actions. Based on our findings, we propose a DRL-based HM prediction (DHP) approach with offline and online versions, called offline-DHP and online-DHP. In offline-DHP, multiple DRL workflows are run to determine potential HM positions at each panoramic frame. Then, a heat map of the potential HM positions, named the HM map, is generated as the output of offline-DHP. In online-DHP, the next HM position of one subject is estimated given the currently observed HM position, which is achieved by developing a DRL algorithm upon the learned offline-DHP model. Finally, the experiments validate that our approach is effective in both offline and online prediction of HM positions for panoramic video, and that the learned offline-DHP model can improve the performance of online-DHP.
Mai Xu, Yuhang Song 0001, Jianyi Wang, Minglang Qiao, Liangyu Huo, Zulin Wang
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 An EM-Based User Clustering Method in Non-Orthogonal Multiple Access
abstract
Power domain non-orthogonal multiple access (NOMA), with the ability to serve multiple users within one resource block, is one of the most promising technologies for the fifth generation. In this paper, we study the downlink millimeter wave (mmWave) NOMA-based system, where the base station sends messages to multiple clusters and serves multiple users simultaneously, and sum-rate maximization problem is investigated. Since users are multiplexed on one resource block, user clustering is important for NOMA and has a great influence on sum-rate optimization problem. Inspired by correlation features of users' spatial distributions in mmWave NOMA-based system, we introduce unsupervised learning method into user clustering. We first develop an Expectation Maximization (EM)-based algorithm in fixed user scenario. Then, the dynamic user scenario is introduced, which includes user reduction, increment and movement situations. After that, an online EM-based clustering algorithm is proposed to fast update user distribution parameters with lower computational complexity compared to the conventional complete re-clustering methods. Simulation results show that the proposed EM-based algorithm can improve the performance of NOMA-based system in fixed user scenario. In addition, the proposed online EM-based algorithm can achieve similar performance as the complete EM-based algorithm with less computational complexity in dynamic user scenario.
Jie Ren 0004, Zulin Wang, Mai Xu, Fang Fang 0005, Zhiguo Ding 0001
IEEE Trans. Commun.3
2019 Assessing Visual Quality of Omnidirectional Videos
abstract
In contrast with traditional videos, omnidirectional videos enable spherical viewing direction with support for head-mounted displays, providing an interactive and immersive experience. Unfortunately, to the best of our knowledge, there are only a few visual quality assessment (VQA) methods, either subjective or objective, for omnidirectional video coding. This paper proposes both subjective and objective methods for assessing the quality loss in encoding an omnidirectional video. Specifically, we first present a new database, which includes the viewing direction data from several subjects watching omnidirectional video sequences. Then, from our database, we find a high consistency in viewing directions across different subjects. The viewing directions are normally distributed in the center of the front regions, but they sometimes fall into other regions, related to the video content. Given this finding, we present a subjective VQA method for measuring the difference mean opinion score (DMOS) of the whole and regional omnidirectional video, in terms of overall DMOS and vectorized DMOS, respectively. Moreover, we propose two objective VQA methods for the encoded omnidirectional video, in light of the human perception characteristics of the omnidirectional video. One method weighs the distortion of pixels with regard to their distances to the center of front regions, which considers human preference in a panorama. The other method predicts viewing directions according to the video content, and then the predicted viewing directions are leveraged to allocate weights to the distortion of each pixel in our objective VQA method. Finally, our experimental results verify that both the subjective and objective methods proposed in this paper advance the state-of-the-art VQA for omnidirectional videos.
Mai Xu, Chen Li 0049, Zhenzhong Chen 0001, Zulin Wang, Zhenyu Guan 0002
IEEE Trans. Circuits Syst. Video Technol.1
2019 Enhancing Quality for HEVC Compressed Videos
abstract
The latest High Efficiency Video Coding (HEVC) standard has been increasingly applied to generate video streams over the Internet. However, HEVC compressed videos may incur severe quality degradation, particularly at low bit rates. Thus, it is necessary to enhance the visual quality of HEVC videos at the decoder side. To this end, this paper proposes a quality enhancement convolutional neural network (QE-CNN) method that does not require any modification of the encoder to achieve quality enhancement for HEVC. In particular, our QE-CNN method learns QE-CNN-I and QE-CNN-P models to reduce the distortion of HEVC I and P/B frames, respectively. The proposed method differs from the existing CNN-based quality enhancement approaches, which only handle intra-coding distortion and are thus not suitable for P/B frames. Our experimental results validate that our QE-CNN method is effective in enhancing quality for both I and P/B frames of HEVC videos. To apply our QE-CNN method in time-constrained scenarios, we further propose a time-constrained quality enhancement optimization (TQEO) scheme. Our TQEO scheme controls the computational time of QE-CNN to meet a target, meanwhile maximizing the quality enhancement. Next, the experimental results demonstrate the effectiveness of our TQEO scheme from the aspects of time control accuracy and quality enhancement under different time constraints. Finally, we design a prototype to implement our TQEO scheme in a real-time scenario.
Mai Xu, Zulin Wang, Zhenyu Guan 0002
IEEE Trans. Circuits Syst. Video Technol.2
2019 A Deep Learning Approach for Multi-Frame In-Loop Filter of HEVC
abstract
An extensive study on the in-loop filter has been proposed for a high efficiency video coding (HEVC) standard to reduce compression artifacts, thus improving coding efficiency. However, in the existing approaches, the in-loop filter is always applied to each single frame, without exploiting the content correlation among multiple frames. In this paper, we propose a multi-frame in-loop filter (MIF) for HEVC, which enhances the visual quality of each encoded frame by leveraging its adjacent frames. Specifically, we first construct a large-scale database containing encoded frames and their corresponding raw frames of a variety of content, which can be used to learn the in-loop filter in HEVC. Furthermore, we find that there usually exist a number of reference frames of higher quality and of similar content for an encoded frame. Accordingly, a reference frame selector (RFS) is designed to identify these frames. Then, a deep neural network for MIF (known as MIF-Net) is developed to enhance the quality of each encoded frame by utilizing the spatial information of this frame and the temporal information of its neighboring higher-quality frames. The MIF-Net is built on the recently developed DenseNet, benefiting from its improved generalization capacity and computational efficiency. In addition, a novel block-adaptive convolutional layer is designed and applied in the MIF-Net, for handling the artifacts influenced by coding tree unit (CTU) structure in HEVC. Extensive experiments show that our MIF approach achieves on average 11.621% saving of the Bjøntegaard delta bit-rate (BD-BR) on the standard test set, significantly outperforming the standard in-loop filter in HEVC and other state-of-the-art approaches.
Tianyi Li 0004, Mai Xu, Ce Zhu, Zulin Wang, Zhenyu Guan 0002
IEEE Trans. Image Process.2
2019 Fast H.264 to HEVC Transcoding: A Deep Learning Method
abstract
With the development of video coding technology, high-efficiency video coding (HEVC) has become a promising alternative, compared with the previous coding standards, for example, H.264. In general, H.264 to HEVC transcoding can be accomplished by fully H.264 decoding and fully HEVC encoding, which suffers from considerable time consumption on the brute-force search of the HEVC coding tree unit (CTU) partition for rate-distortion optimization (RDO). In this paper, we propose a deep learning method to predict the HEVC CTU partition, instead of the brute-force RDO search, for H.264 to HEVC transcoding. First, we build a large-scale H.264 to HEVC transcoding database. Second, we investigate the correlation between the HEVC CTU partition and H.264 features, and analyze both temporal and spatial-temporal similarities of the CTU partition across video frames. Third, we propose a deep learning architecture of a hierarchical long short-term memory (H-LSTM) network to predict the CTU partition of HEVC. Then, the brute-force RDO search of the CTU partition is replaced by the H-LSTM prediction such that the computational time can be significantly reduced for fast H.264 to HEVC transcoding. Finally, the experimental results verify that the proposed H-LSTM method can achieve a better tradeoff between coding efficiency and complexity, compared to the state-of-the-art H.264 to HEVC transcoding methods.
Jingyao Xu 0002, Mai Xu, Yanan Wei, Zulin Wang, Zhenyu Guan 0002
IEEE Trans. Multim.2
2018 Multi-Frame Quality Enhancement for Compressed Video
abstract
The past few years have witnessed great success in applying deep learning to enhance the quality of compressed image/video. The existing approaches mainly focus on enhancing the quality of a single frame, ignoring the similarity between consecutive frames. In this paper, we investigate that heavy quality fluctuation exists across compressed video frames, and thus low quality frames can be enhanced using the neighboring high quality frames, seen as Multi-Frame Quality Enhancement (MFQE). Accordingly, this paper proposes an MFQE approach for compressed video, as a first attempt in this direction. In our approach, we firstly develop a Support Vector Machine (SVM) based detector to locate Peak Quality Frames (PQFs) in compressed video. Then, a novel Multi-Frame Convolutional Neural Network (MF-CNN) is designed to enhance the quality of compressed video, in which the non-PQF and its nearest two PQFs are as the input. The MF-CNN compensates motion between the non-PQF and PQFs through the Motion Compensation subnet (MC-subnet). Subsequently, the Quality Enhancement subnet (QE-subnet) reduces compression artifacts of the non-PQF with the help of its nearest PQFs. Finally, the experiments validate the effectiveness and generality of our MFQE approach in advancing the state-of-the-art quality enhancement of compressed video. The code of our MFQE approach is available at https://github.com/ryangBUAA/MFQE.git.
Mai Xu, Zulin Wang, Tianyi Li 0004
CVPR2
2018 DeepVS: A Deep Learning Based Video Saliency Prediction Approach
Lai Jiang 0004, Mai Xu, Minglang Qiao, Zulin Wang
ECCV (14)2
2018 Boundary Objectness Network for Object Detection and Localization
abstract
In this paper, we present the boundary objectness network (BON), an effective convolutional neural network (CNN) for object detection. Its core contribution is to accurately localize the objects. Generally, the CNN-based localizers predict four bounding box coordinates by learning a regression function. This method shows a low Intersection-of-Union (IoU) with the ground truth box. In our work, the localization is formu-lated as a probabilistic problem. Specifically, the deep features inside the candidate proposal are mapped into a row and a column feature vector, which are called boundary object-ness. The boundary objectness indicates the existence of an object in the horizontal and vertical direction of the proposal, enabling us to elaborately localize the object. Moreover, the modules of object detection share the common convolution-al layers. Meanwhile, a multi-task loss function is designed for joint training strategy. Experimental results on the PAS-CAL VOC datasets demonstrate the competitive performance of our method. For the VGG16 model, we achieve 77.6 % mAP at a speed of 4 frame per second (FPS), thus having the potential for real-time processing.
Juan Wang 0012, Xiaoming Tao 0001, Mai Xu, Jianhua Lu
ICASSP3
2018 Bridge the Gap Between VQA and Human Behavior on Omnidirectional Video: A Large-Scale Dataset and a Deep Learning Model
abstract
Omnidirectional video enables spherical stimuli with the $360 \times 180^ \circ$ viewing range. Meanwhile, only the viewport region of omnidirectional video can be seen by the observer through head movement (HM), and an even smaller region within the viewport can be clearly perceived through eye movement (EM). Thus, the subjective quality of omnidirectional video may be correlated with HM and EM of human behavior. To fill in the gap between subjective quality and human behavior, this paper proposes a large-scale visual quality assessment (VQA) dataset of omnidirectional video, called VQA-OV, which collects 60 reference sequences and 540 impaired sequences. Our VQA-OV dataset provides not only the subjective quality scores of sequences but also the HM and EM data of subjects. By mining our dataset, we find that the subjective quality of omnidirectional video is indeed related to HM and EM. Hence, we develop a deep learning model, which embeds HM and EM, for objective VQA on omnidirectional video. Experimental results show that our model significantly improves the state-of-the-art performance of VQA on omnidirectional video.
Chen Li 0049, Mai Xu, Xinzhe Du, Zulin Wang
ACM Multimedia2
2018 Rate control schemes for panoramic video coding
Yufan Liu 0001, Li Yang 0014, Mai Xu, Zulin Wang
J. Vis. Commun. Image Represent.3
2018 Embracing non-orthogonalmultiple access in future wireless networks
abstract
This paper provides a comprehensive survey of the impact of the emerging communication technique, non-orthogonal multiple access (NOMA), on future wireless networks. Particularly, how the NOMA principle affects the design of the generation multiple access techniques is introduced first. Then the applications of NOMA to other advanced communication techniques, such as wireless caching, multiple-input multiple-output techniques, millimeter-wave communications, and cooperative relaying, are discussed. The impact of NOMA on communication systems beyond cellular networks is also illustrated, through the examples of digital TV, satellite communications, vehicular networks, and visible light communications. Finally, the study is concluded with a discussion of important research challenges and promising future directions in NOMA.
Zhiguo Ding 0001, Mai Xu, Yan Chen 0010, Mugen Peng, H. Vincent Poor
Frontiers Inf. Technol. Electron. Eng.2
2018 Hierarchical objectness network for region proposal generation and object detection
Juan Wang 0012, Xiaoming Tao 0001, Mai Xu, Yiping Duan, Jianhua Lu
Pattern Recognit.3
2018 Find Who to Look at: Turning From Action to Saliency
abstract
The past decade has witnessed the use of highlevel features in saliency prediction for both videos and images. Unfortunately, the existing saliency prediction methods only handle high-level static features, such as face. In fact, high-level dynamic features (also called actions), such as speaking or head turning, are also extremely attractive to visual attention in videos. Thus, in this paper, we propose a data-driven method for learning to predict the saliency of multiple-face videos, by leveraging both static and dynamic features at high-level. Specifically, we introduce an eye-tracking database, collecting the fixations of 39 subjects viewing 65 multiple-face videos. Through analysis on our database, we find a set of high-level features that cause a face to receive extensive visual attention. These high-level features include the static features of face size, center-bias and head pose, as well as the dynamic features of speaking and head turning. Then, we present the techniques for extracting these high-level features. Afterwards, a novel model, namely multiple hidden Markov model (M-HMM), is developed in our method to enable the transition of saliency among faces. In our MHMM, the saliency transition takes into account both the state of saliency at previous frames and the observed high-level features at the current frame. The experimental results show that the proposed method is superior to other state-of-the-art methods in predicting visual attention on multiple-face videos. Finally, we shed light on a promising implementation of our saliency prediction method in locating the region-of-interest (ROI), for video conference compression with high efficiency video coding (HEVC).
Mai Xu, Yufan Liu 0001, Haoji Hu, Feng He 0007
IEEE Trans. Image Process.1
2018 Reducing Complexity of HEVC: A Deep Learning Approach
abstract
High Efficiency Video Coding (HEVC) significantly reduces bit-rates over the preceding H.264 standard but at the expense of extremely high encoding complexity. In HEVC, the quad-tree partition of coding unit (CU) consumes a large proportion of the HEVC encoding complexity, due to the brute-force search for rate-distortion optimization (RDO). Therefore, this paper proposes a deep learning approach to predict the CU partition for reducing the HEVC complexity at both intra-and inter-modes, which is based on convolutional neural network (CNN) and long-and short-term memory (LSTM) network. First, we establish a large-scale database including substantial CU partition data for HEVC intra-and inter-modes. This enables deep learning on the CU partition. Second, we represent the CU partition of an entire coding tree unit (CTU) in the form of a hierarchical CU partition map (HCPM). Then, we propose an early-terminated hierarchical CNN (ETH-CNN) for learning to predict the HCPM. Consequently, the encoding complexity of intra-mode HEVC can be drastically reduced by replacing the brute-force search with ETH-CNN to decide the CU partition. Third, an early-terminated hierarchical LSTM (ETH-LSTM) is proposed to learn the temporal correlation of the CU partition. Then, we combine ETH-LSTM and ETH-CNN to predict the CU partition for reducing the HEVC complexity at inter-mode. Finally, experimental results show that our approach outperforms other state-of-the-art approaches in reducing the HEVC complexity at both intra-and inter-modes.
Mai Xu, Tianyi Li 0004, Zulin Wang, Xin Deng 0002, Zhenyu Guan 0002
IEEE Trans. Image Process.1
2018 Closed-Form Optimization on Saliency-Guided Image Compression for HEVC-MSP
abstract
High efficiency video coding (HEVC) is the latest video coding standard, and it has the best performance among all the existing standards. HEVC main still picture profile (HEVC-MSP) also achieves top performance in image compr-ession. In this paper, we propose a closed-form bit allocation approach to optimize the saliency-guided PSNR (viewed as perceptual distortion) such that the coding efficiency of HEVC-based image compression can be significantly improved from a subjective perspective. Specifically, a bit allocation formulation is established to minimize perceptual distortion with a constraint on bit-rates. Then, this formulation is solved using the proposed recursive Taylor expansion method with a closed-form solution. On the basis of our solution, a bit allocation and re-allocation process is developed in our approach to minimize perceptual distortion, meanwhile accurately controlling bit-rates. In addition, we provide both theoretical and numerical analyses of the computational complexity, verifying the little extra time cost of our approach. The experimental results demonstrate the superior performance of our approach over the state-of-the-art HEVC-MSP, and the BD-rate savings are approximately 40% and 24% for face and generic images, respectively.
Shengxi Li, Mai Xu, Yun Ren, Zulin Wang
IEEE Trans. Multim.2
2018 Saliency Detection in Face Videos: A Data-Driven Approach
abstract
Recently, videoconferencing has been popular in multimedia systems, such as FaceTime and Skype. In videoconferencing, almost every frame contains a human face. Therefore, it is important to predict human visual attention on face videos by saliency detection, as saliency may be used as a guide to the region of interest for the content-based applications of face videos. In this paper, we propose a data-driven approach for saliency detection in face videos. From the data-driven perspective, we first establish an eye-tracking database that contains fixations of 76 face videos viewed by 40 subjects. Upon the analysis of our database, we find that visual attention is significantly attracted by faces in videos. More important, the attention distribution within face regions varies with regard to mouth movement. Since previous works have investigated that it is efficient to model face saliency in still images using a Gaussian mixture model (GMM), the variation of visual attention in videos can be modeled by dynamic GMM (DGMM). Accordingly, we propose adopting the particle filter (PF) in modeling DGMM for saliency detection of face videos, which is called PF-DGMM. Finally, the experimental results show that our PF-DGMM approach significantly outperforms other state-of-the-art approaches in saliency detection of face videos.
Mai Xu, Yun Ren, Zulin Wang, Jingxian Liu, Xiaoming Tao 0001
IEEE Trans. Multim.1
2017 Predicting Salient Face in Multiple-Face Videos
abstract
Although the recent success of convolutional neural network (CNN) advances state-of-the-art saliency prediction in static images, few work has addressed the problem of predicting attention in videos. On the other hand, we find that the attention of different subjects consistently focuses on a single face in each frame of videos involving multiple faces. Therefore, we propose in this paper a novel deep learning (DL) based method to predict salient face in multiple-face videos, which is capable of learning features and transition of salient faces across video frames. In particular, we first learn a CNN for each frame to locate salient face. Taking CNN features as input, we develop a multiple-stream long short-term memory (M-LSTM) network to predict the temporal transition of salient faces in video sequences. To evaluate our DL-based method, we build a new eye-tracking database of multiple-face videos. The experimental results show that our method outperforms the prior state-of-the-art methods in predicting visual attention on faces in multiple-face videos.
Yufan Liu 0001, Songyang Zhang 0001, Mai Xu, Xuming He 0001
CVPR3
2017 Watching Videos with Certain and Constant Quality: PID-Based Quality Control Method
Yuhang Song 0001, Mai Xu, Shengxi Li
DCC2
2017 Complexity control of HEVC for video conferencing
abstract
In this paper, we propose an effective complexity control approach for video conferencing scenarios on HEVC platform. A complexity control formulation is established to determine the number of depth-constrained largest coding units (LCUs) according to the target complexity. By limiting the maximum depths of different LCUs to different levels, the encoding complexity can be controlled with high accuracy. Different from other approaches, both the objective and perceptual-driven video quality are kindly preserved through taking both the objective and subjective weight maps into consideration when controlling the complexity. The experimental results demonstrate that our approach outperforms the state-of-the art approach with higher control accuracy. Also, despite of complexity reduction, our approach keeps the objective and perceptual-driven quality well.
Xin Deng 0002, Mai Xu
ICASSP2
2017 A deep convolutional neural network approach for complexity reduction on intra-mode HEVC
abstract
The High Efficiency Video Coding (HEVC) standard significantly saves coding bit-rate over the proceeding H.264 standard, but at the expense of extremely high encoding complexity. In fact, the coding tree unit (CTU) partition consumes a large proportion of HEVC encoding complexity, due to the brute-force search for rate-distortion optimization (RDO). Therefore, we propose in this paper a complexity reduction approach for intra-mode HEVC, which learns a deep convolutional neural network (CNN) model to predict CTU partition instead of RDO. Firstly, we establish a large-scale database with diversiform patterns of CTU partition. Secondly, we model the partition as a three-level classification problem. Then, for solving the classification problem, we develop a deep CNN structure with various sizes of convolutional kernels and extensive trainable parameters, which can be learnt from the established database. Finally, experimental results show that our approach reduces intramode encoding time by 62.25% and 69.06% with negligible Bjontegaard delta bit-rate of 2.12% and 1.38%, over the test sequences and images respectively, superior to other state-of-the-art approaches.
Tianyi Li 0004, Mai Xu, Xin Deng 0002
ICME2
2017 A novel rate control scheme for panoramic video coding
abstract
The popularity of multi-view panoramic videos has been considerably increased for producing Virtual Reality (VR) content, due to its immersive visual experience. We argue in this paper that PSNR is less effective in assessing visual quality of compressed panoramic videos than Sphere-based PSNR (S-PNSR), in which sphere-to-plain mapping of panoramic videos is considered. Thus, the conventional rate control (R-C) schemes of 2-Dimensional (2D) video coding, which optimize on PSNR, are not suitable for panoramic video coding. To optimize S-PSNR, we propose in this paper a novel RC scheme for panoramic video coding. Specifically, we develop an S-PSNR optimization formulation with constraint on bit-rate. Then, a solution is provided to the developed formulation, such that bits can be allocated to each coding block for achieving optimal S-PSNR in panoramic video coding. Finally, the experiment results validate the effectiveness of the proposed RC scheme in improving S-PSNR of panoramic video coding.
Yufan Liu 0001, Mai Xu, Chen Li 0049, Shengxi Li, Zulin Wang
ICME2
2017 A subjective visual quality assessment method of panoramic videos
abstract
Different from 2-dimensional (2D) videos, panoramic videos contain spherical viewing direction with the support of head-mounted displays, thus improving immersive and interactive visual experience. Unfortunately, to our best knowledge, there are few subjective visual quality assessment (VQA) methods for panoramic videos. In this paper, we therefore propose a subjective VQA method for assessing quality loss of impaired panoramic videos. Specifically, we first establish a database containing viewing direction data of several subjects on watching panoramic videos. Then, we find out that there exists high consistency of viewing direction on panoramic videos across different subjects. Upon this finding, we present a procedure of subjective test in measuring quality of panoramic videos by different subjects, yielding different mean opinion score (DMOS). To couple with inconsistency of viewing directions on panoramic videos, we further propose a vectorized DMOS metric. Finally, experimental results verify that our subjective VQA method, in the forms of both overall and vectorized DMOS metrics, is effective in measuring subjective quality of panoramic videos.
Mai Xu, Chen Li 0049, Yufan Liu 0001, Xin Deng 0002, Jiaxin Lu 0003
ICME1
2017 Decoder-side HEVC quality enhancement with scalable convolutional neural network
abstract
The latest High Efficiency Video Coding (HEVC) has been increasingly used to generate video streams over Internet. However, the decoded HEVC video streams may incur severe quality degradation, especially at low bit-rates. Thus, it is necessary to enhance visual quality of HEVC videos at the decoder side. To this end, we propose in this paper a Decoder-side Scalable Convolutional Neural Network (DS-CNN) approach to achieve quality enhancement for HEVC, which does not require any modification of the encoder. In particular, our DS-CNN approach learns a model of Convo-lutional Neural Network (CNN) to reduce distortion of both I and B/P frames in HEVC. It is different from the existing CNN-based quality enhancement approaches, which only handle intra coding distortion, thus not suitable for B/P frames. Furthermore, a scalable structure is included in our DS-CNN, suchthat the computational complexity of our DS-CNN approach is adjustable to the changing computational resources. Finally, the experimental results show the effectiveness of our DS-CNN approach in enhancing quality for both I and B/P frames of HEVC.
Mai Xu, Zulin Wang
ICME2
2017 An LSTM method for predicting CU splitting in H.264 to HEVC transcoding
abstract
For H.264 to high efficiency video coding (HEVC) transcoding, this paper proposes a hierarchical Long Short-Term Memory (LSTM) method to predict coding unit (CU) splitting. Specifically, we first analyze the correlation between CU splitting patterns and H.264 features. Upon our analysis, we further propose a hierarchical LSTM architecture for predicting CU splitting of HEVC, with regard to the explored H.264 features. The features of H.264, including residual, macroblock (MB) partition and bit allocation, are employed as the input to our LSTM method. Experimental results demonstrate that the proposed method outperforms the state-of-the-art H.264 to HEVC transcoding methods, in terms of both complexity reduction and PSNR performance.
Yanan Wei, Zulin Wang, Mai Xu, Shuhao Qiao
VCIP3
2017 Road Structure Refined CNN for Road Extraction in Aerial Image
abstract
In this letter, we propose a road structure refined convolutional neural network (RSRCNN) approach for road extraction in aerial images. In order to obtain structured output of road extraction, both deconvolutional and fusion layers are designed in the architecture of RSRCNN. For training RSRCNN, a new loss function is proposed to incorporate the geometric information of road structure in cross-entropy loss, thus called road-structure-based loss function. Experimental results demonstrate that the trained RSRCNN model is able to advance the state-of-the-art road extraction for aerial images, in terms of precision, recall, F-score, and accuracy.
Yanan Wei, Zulin Wang, Mai Xu
IEEE Geosci. Remote. Sens. Lett.3
2017 Optimal Bit Allocation for CTU Level Rate Control in HEVC
abstract
For High Efficiency Video Coding (HEVC), the R–$\lambda $scheme is the latest rate control (RC) scheme, which investigates the relationships among allocated bits, the slope of rate-distortion (R-D) curve$\lambda $, and quantization parameter. However, we argue that bit allocation in the existing R–$\lambda $scheme is not optimal. In this paper, we therefore propose an optimal bit allocation (OBA) scheme for coding tree unit level RC in HEVC. Specifically, to achieve the OBA, we first develop an optimization formulation with a novel R-D estimation, instead of the existing R–$\lambda $estimation. Unfortunately, it is intractable to obtain a closed-form solution to the optimization formulation. We thus propose a recursive Taylor expansion (RTE) method to iteratively solve the formulation. As a result, an approximate closed-form solution can be obtained, thus achieving OBA and bit reallocation. Both theoretical and numerical analyses show the fast convergence speed and little computational time of the proposed RTE method. Therefore, our OBA scheme can be achieved at little encoding complexity cost. Finally, the experimental results validate the effectiveness of our scheme in three aspects: R-D performance, RC accuracy, and robustness over dynamic scene changes.
Shengxi Li, Mai Xu, Zulin Wang, Xiaoyan Sun 0001
IEEE Trans. Circuits Syst. Video Technol.2
2017 Bayesian Hyperspectral and Multispectral Image Fusions via Double Matrix Factorization
abstract
This paper focuses on fusing hyperspectral and multispectral images with an unknown arbitrary point spread function (PSF). Instead of obtaining the fused image based on the estimation of the PSF, a novel model is proposed without intervention of the PSF under Bayesian framework, in which the fused image is decomposed into double subspace-constrained matrix-factorization-based components and residuals. On the basis of the model, the fusion problem is cast as a minimum mean square error estimator of three factor matrices. Then, to approximate the posterior distribution of the unknowns efficiently, an estimation approach is developed based on variational Bayesian inference. Different from most previous works, the PSF is not required in the proposed model and is not pre-assumed to be spatially invariant. Hence, the proposed approach is not related to the estimation errors of the PSF and has potential computational benefits when extended to spatially variant imaging system. Moreover, model parameters in our approach are less dependent on the input data sets and most of them can be learned automatically without manual intervention. Exhaustive experiments on three data sets verify that our approach shows excellent performance and more robustness to the noise with acceptable computational complexity, compared with other state-of-the-art methods.
Baihong Lin, Xiaoming Tao 0001, Mai Xu, Linhao Dong, Jianhua Lu
IEEE Trans. Geosci. Remote. Sens.3
2017 Learning to Detect Video Saliency With HEVC Features
abstract
Saliency detection has been widely studied to predict human fixations, with various applications in computer vision and image processing. For saliency detection, we argue in this paper that the state-of-the-art High Efficiency Video Coding (HEVC) standard can be used to generate the useful features in compressed domain. Therefore, this paper proposes to learn the video saliency model, with regard to HEVC features. First, we establish an eye tracking database for video saliency detection, which can be downloaded from https://github.com/remega/video_database. Through the statistical analysis on our eye tracking database, we find out that human fixations tend to fall into the regions with large-valued HEVC features on splitting depth, bit allocation, and motion vector (MV). In addition, three observations are obtained with the further analysis on our eye tracking database. Accordingly, several features in HEVC domain are proposed on the basis of splitting depth, bit allocation, and MV. Next, a kind of support vector machine is learned to integrate those HEVC features together, for video saliency detection. Since almost all video data are stored in the compressed form, our method is able to avoid both the computational cost on decoding and the storage cost on raw data. More importantly, experimental results show that the proposed method is superior to other state-of-the-art saliency detection methods, either in compressed or uncompressed domain.
Mai Xu, Lai Jiang 0004, Xiaoyan Sun 0001, Zhaoting Ye, Zulin Wang
IEEE Trans. Image Process.1
2016 Optimizing Subjective Quality in HEVC-MSP: An Approximate Closed-form Image Compression Approach
abstract
HEVC, as the latest video coding standard, achieves top performance on image compression. On the basis of this, we propose a novel approach to optimize subjective quality for HEVC-based image compression. Specifically, a bit allocation formulation is established to optimize subjective quality with constraint on bit-rates. Then, we propose a recursive Taylor expansion method to quickly solve such a formulation with an approximate closed-form solution. The experimental results show the superior performance of our approach, with ~40% BD-rate saving over the state-of-the-art HEVC-MSP for face image compression.
Shengxi Li, Mai Xu, Yun Ren, Chengzhang Ma, Zulin Wang
DCC2
2016 Subjective-quality-optimized complexity control for HEVC decoding
abstract
The latest High Efficiency Video Coding (HEVC) standard significantly improves coding efficiency over H.264/AVC, at the cost of heavy encoding and decoding complexity. For reducing HEVC decoding complexity to a target, we propose in this paper a Subjective-Quality-Optimized Complexity Control (SQOCC) approach, which optimizes subjective quality loss caused by the decoding complexity reduction. First, a saliency detection method in HEVC domain is developed as the preliminary of subjective quality metric. Based on detected saliency, we establish a formulation to minimize subjective quality loss at the constraint of specific decoding complexity reduction, via disabling the deblocking filters of some Largest Coding Units (LCUs). Next, we utilize least square fitting to model functions in our formulation. We then provide a solution to our formulation, achieving subjective-quality-optimized complexity control for HEVC decoding. Finally, the experimental results show the effectiveness of our SQOC-C approach in terms of both control accuracy and subjective quality.
Mai Xu, Lai Jiang 0004, Zulin Wang
ICME2
2016 Predicting the memorability of natural-scene images
abstract
Recent work has shown that image memorability, in general, can be reliably predicted using some state-of-the-art features. However, all existing methods are not effective in predicting memorability of natural-scene images, far from human. In this paper, we propose a novel method to improve the effectiveness of memorability prediction for natural-scene images. Specifically, we argue that some of HSV colors have either positive or negative impact on memorability of natural-scene images in our Natural-Scene Image Memorability (NSIM) dataset. Then, we develop an HSV-based feature for memorability prediction. Finally, the HSV-based feature is combined with other efficient state-of-the-art features in our approach to predict memorability on natural-scene images. Experimental results validate the effectiveness of our method.
Jiaxin Lu 0003, Mai Xu, Zulin Wang
VCIP2
2016 A novel double-layer sparse representation approach for unsupervised dictionary learning
Mai Xu, Zulin Wang
Comput. Vis. Image Underst.1
2016 Unsupervised dictionary learning with Fisher discriminant for clustering
Mai Xu, Ling Li 0010
Neurocomputing1
2016 Bottom-up saliency detection with sparse representation of learnt texture atoms
Mai Xu, Lai Jiang 0004, Zhaoting Ye, Zulin Wang
Pattern Recognit.1
2016 Subjective-Driven Complexity Control Approach for HEVC
abstract
The latest High Efficiency Video Coding (HEVC) standard significantly increases the encoding complexity for improving its coding efficiency, compared with the preceding H.264/Advanced Video Coding (AVC) standard. In this paper, we present a novel subjective-driven complexity control (SCC) approach to reduce and control the encoding complexity of HEVC. Through reasonably adjusting the maximum depth of each largest coding unit (LCU), the encoding complexity can be reduced to a target level with minimal visual distortion. Specifically, the maximum depths of different LCUs can be varied through solving the proposed optimization formulation of complexity control, based on two explored relationships: 1) the relationship between the maximum depth and encoding complexity and 2) the relationship between the maximum depth and visual distortion. Besides, the subjective visual quality is favored with a novel subjective-driven constraint imposed in the formulation, on the basis of a visual attention model. Finally, the experimental results show that our approach can achieve a wide range of encoding complexity control (as low as 20%) for HEVC, with the smallest complexity bias being 0.2%. Meanwhile, our SCC approach outperforms other two state-of-the-art complexity control approaches, in terms of both control accuracy and visual quality.
Xin Deng 0002, Mai Xu, Lai Jiang 0004, Xiaoyan Sun 0001, Zulin Wang
IEEE Trans. Circuits Syst. Video Technol.2
2015 Learning to Predict Saliency on Face Images
abstract
This paper proposes a novel method, which learns to detect saliency of face images. To be more specific, we obtain a database of eye tracking over extensive face images, via conducting an eye tracking experiment. With analysis on eye tracking database, we verify that the fixations tend to cluster around facial features, when viewing images with large faces. For modeling attention on faces and facial features, the proposed method learns the Gaussian mixture model (GMM) distribution from the fixations of eye tracking data as the top-down features for saliency detection of face images. Then, in our method, the top-down features (i.e., face and facial features) upon the the learnt GMM are linearly combined with the conventional bottom-up features (i.e., color, intensity, and orientation), for saliency detection. In the linear combination, we argue that the weights corresponding to top-down feature channels depend on the face size in images, and the relationship between the weights and face size is thus investigated via learning from the training eye tracking data. Finally, experimental results show that our learning-based method is able to advance state-of-the-art saliency prediction for face images. The corresponding database and code are available online: www.ee.buaa.edu.cn/xumfiles/saliency_detection.html.
Mai Xu, Yun Ren, Zulin Wang
ICCV1
2015 A novel method on optimal bit allocation at LCU level for rate control in HEVC
abstract
In this paper, we propose a new method, namely recursive Taylor expansion (RTE) method, for optimally allocating bits to each LCU in the R-λ rate control scheme for HEVC. Specifically, we first set up an optimization formulation on optimal bit allocation. Unfortunately, it is intractable to achieve a closed-form solution for this formulation. We therefore propose a RTE solution to iteratively solve the formulation with a fast convergence speed. Then, an approximate closed-form solution can be obtained. This way, the optimal bit allocation can be achieved at little encoding complexity cost. Finally, the experimental results validate the effectiveness of our method in three aspects: compressed distortion, bit-rate control error, and bit fluctuation.
Shengxi Li, Mai Xu, Zulin Wang
ICME2
2015 Learning Gaussian mixture model for saliency detection on face images
abstract
The previous work has demonstrated that integrating top-down features in bottom-up saliency methods can improve the saliency prediction accuracy. Therefore, for face images, this paper proposes a saliency detection method based on Gaussian mixture model (GMM), which learns the distribution of saliency over face regions as the top-down feature. Specifically, we verify that fixations tend to cluster around facial features, when viewing images with large faces. Thus, the GMM is learnt from fixations of eye tracking data, for establishing the distribution of saliency in faces. Then, in our method, the top-down feature upon the the learnt GMM is combined with the conventional bottom-up features (i.e., color, intensity, and orientation), for saliency detection. Finally, experimental results validate that our method is capable of improving the accuracy of saliency prediction for face images.
Yun Ren, Mai Xu, Ruihan Pan, Zulin Wang
ICME2
2015 Subjective rate-distortion optimization in HEVC with perceptual model of multiple faces
abstract
This paper proposes a novel perceptual video coding approach with a perceptual model of multiple faces, to improve the coding efficiency of HEVC in video conferencing scenarios. For the perceptual model, a latest active appearance model (AAM) is used to detect multiple faces in a video frame. Then, the perceptual model of multiple faces can be established on the basis of the detected multiple faces. With the established perceptual model of multiple faces, all faces in a video frame can be taken into account for subjective rate-distortion optimization, which is based on the state-of-the-art r-λ rate control scheme of HEVC. As such, the perceptual video coding can be achieved for HEVC of video conferencing scenarios. Finally, the experimental results validate the effectiveness of the proposed perceptual video coding approach, in terms of subjective quality.
Yufan Liu 0001, Haoji Hu, Mai Xu
VCIP3
2015 Robust secrecy rate optimisations for multiuser multiple-input-single-output channel with device-to-device communications
abstract
In the present study, the authors investigate robust secrecy rate optimisation problems for a multiple‐input‐single‐output secrecy channel with multiple device‐to‐device (D2D) communications. The D2D communication nodes access this secrecy channel by sharing the same spectrum, and help to improve the secrecy communications by confusing the eavesdroppers. In return, the legitimate transmitter ensures that the D2D communication nodes achieve their required rates. In addition, it is assumed that the legitimate transmitter has imperfect channel state information of different nodes. For this secrecy network, the authors solve two robust secrecy rate optimisation problems: (a) robust power minimisation problem, subject to the probability based secrecy rate and the D2D transmission rate constraints; (b) robust secrecy rate maximisation problem with the transmit power, the probabilistic based secrecy rate and the D2D transmission rate constraints. Owing to the non‐convexity of robust beamforming design based on two statistical channel uncertainty models, the authors present two conservative approximation approaches based on ‘Bernstein‐type’ inequality and ‘S‐Procedure’ to solve these robust optimisation problems. Simulation results are provided to validate the performance of these two conservative approximation methods, where it is shown that ‘Bernstein‐type’ inequality based approach outperforms the ‘S‐Procedure’ approach in terms of achievable secrecy rates.
Zheng Chu 0001, K. Cumanan, Mai Xu, Zhiguo Ding 0001
IET Commun.3
2015 Weight-based R-λ rate control for perceptual HEVC coding on conversational videos
Shengxi Li, Mai Xu, Xin Deng 0002, Zulin Wang
Signal Process. Image Commun.2
2014 A novel weight-based URQ scheme for perceptual video coding of conversational video in HEVC
abstract
In this paper, we propose a novel weight-based unified rate-quantization (URQ) scheme for rate control in state-of-the-art HEVC standard, to improve its perceived visual quality for conversational videos. In conventional rate control of HEVC, a pixel-wise URQ scheme is proposed by introducing the concept of bit per pixel (bpp). This scheme is able to assign different amounts of bits to the blocks with various sizes, thus well suitable for flexible picture partition of HEVC. However, bpp does not reflect the visual importance of each pixel. Therefore, we propose a novel weight-based URQ scheme to take into account the visual importance for rate control in HEVC. In combination with the weight map acquired from a novel hierarchical perceptual model of face, such a scheme is capable of allocating more bits to the face and much more bits to the facial features, by using bit per weight (bpw) instead of bpp. As a result, the visual quality of face, especially facial features, can be improved such that perceptual video coding is achieved for HEVC. Finally, the experimental results validate such improvement.
Shengxi Li, Mai Xu, Xin Deng 0002, Zulin Wang
ICME2
2014 Complexity control of HEVC based on region-of-interest attention model
abstract
In this paper, we present a novel complexity control method of HEVC to adjust its encoding complexity. First, a region-of-interest (ROI) attention model is established, which defines different weights for various regions according to their importance. Then, the complexity control algorithm is proposed with a distortion-complexity optimization model, to determine the maximum depth of the largest coding units (LCUs) according to their weights. We can reduce the encoding complexity to a given target level at the cost of little distortion loss. Finally, the experimental results show that the encoding complexity can drop to a pre-defined target complexity as low as 20% with bias less than 7%. Meanwhile, our method is verified to preserve the quality of ROI better than another state-of-the-art approach.
Xin Deng 0002, Mai Xu, Shengxi Li, Zulin Wang
VCIP2
2014 A novel objective quality assessment method for perceptual video coding in conversational scenarios
abstract
Recently, numerous perceptual video coding approaches have been proposed to use face as ROI regions, for improving perceived visual quality of compressed conversational videos. However, there exists no objective metric, specialized for efficiently evaluating the perceived visual quality of compressed conversational videos. This paper thus proposes an efficient objective quality assessment method, namely Gaussian mixture model based PSNR (GMM-PSNR), for conversational videos. First, eye tracking experiments, together with a face extraction technique, were carried out to identify importance of the regions of background, face, and facial features, through eye fixation points. Next, assuming that the distribution of some eye fixation points obeys Gaussian mixture model, an importance weight map is generated by introducing a new term, eye fixation points/pixel(efp/p). Finally, GMM-PSNR is computed by assigning different penalties to the distortion of each pixel in a video frame, according to the generated weight map. The experimental results show the effectiveness of our GMM-PSNR by investigating its correlation with subjective quality on several test video sequences.
Mai Xu, Jingze Zhang, Zulin Wang
VCIP1
2014 Unsupervised dictionary learning with double-layer sparse representation
abstract
This paper presents a novel double-layer sparse representation (DLSR) approach for unsupervised dictionary learning. In supervised/unsupervised discriminative dictionary learning, classical approaches usually develop a discriminative term for learning multiple sub-dictionaries, each of which corresponds to one-class training image patches. However, in unsupervised scenario, some of the training patches for learning sub-dictionaries of each class are related to more than one class. Thus, we propose a DLSR formulation, in this paper, to impose the first-layer sparsity on the coefficients and the second-layer sparsity on the classes for each training patch, embedding both the reconstructive (via the first-layer) and discriminative (via the second-layer) abilities in the dictionary. To address the proposed DLSR formulation, a simple yet effective algorithm, called DLSR-OMP, is developed in light of the conventional OMP. Finally, the experimental results show the effectiveness of our approach in the reconstruction task of image denoising and the clustering task of texture segmentation.
Mai Xu, Zulin Wang
WACV1
2014 Tower of Knowledge for scene interpretation: A survey
Mai Xu, Zulin Wang, Maria Petrou
Pattern Recognit. Lett.1
2014 Compressibility Constrained Sparse Representation With Learnt Dictionary for Low Bit-Rate Image Compression
abstract
This paper proposes a compressibility constrained sparse representation (CCSR) approach to low bit-rate image compression using a learnt over-complete dictionary of texture patches. Conventional sparse representation approaches for image compression are based on matching pursuit (MP) algorithms. Actually, the weakness of these approaches is that they are not stable in terms of sparsity of the estimated coefficients, thereby resulting in the inferior performance in low bit-rate image compression. In comparison with MP, convex relaxation approaches are more stable for sparse representation. However, it is intractable to directly apply convex relaxation approaches to image compression, as their coefficients are not always compressible. To utilize convex relaxation in image compression, we first propose in this paper a CCSR formulation, imposing the compressibility constraint on the coefficients of sparse representation for each image patch. In addition, we work out the CCSR formulation to obtain sparse and compressible coefficients, through recursively solving the \(\ell _{1}\) -norm optimization problem of sparse representation. Given these coefficients, each image patch can be represented by the linear combination of texture elements encoded in an over-complete dictionary, learnt from other training images. Finally, low bit-rate image compression can be achieved, owing to the sparsity and compressibility of coefficients by our CCSR approach. The experimental results demonstrate the effectiveness and superiority of the CCSR approach on compressing the natural and remote sensing images at low bit-rates.
Mai Xu, Shengxi Li, Jianhua Lu, Wenwu Zhu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2013 Modeling daily patient arrivals at Emergency Department and quantifying the relative importance of contributing variables using artificial neural network
Mai Xu, T. C. Wong 0001, Kwai-Sang Chin
Decis. Support Syst.1
2012 Sparse representation of texture patches for low bit-rate image compression
abstract
This paper proposes a sparse representation based approach for low bit-rate image compression using the learnt over-complete dictionary of texture patches. We first propose to compress each patch of the image with sparse and compressible linear combinations (via nonzero coefficients) of texture patterns encoded in a dictionary for image patches. Then, we find out that the compressibility and sparsity of coefficients can be achieved by the proposed recursive procedure of solving ℓ1optimization problem of sparse representation. Moreover, rather than transform-based patterns (e.g. DCT), we explore the basic texture patterns from other training images with a learning algorithm based on the gradient descent, to form the over-complete dictionary. The experimental results demonstrate the effectiveness of the proposed approach.
Mai Xu, Jianhua Lu, Wenwu Zhu 0001
VCIP1
2012 Rateless Codes with Progressive Recovery for Layered Multimedia Delivery
abstract
This paper proposes a novel approach, based on unequal error protection, to enhance rateless codes with progressive recovery for layered multimedia delivery. With a parallel encoding structure, the proposed Progressive Rateless codes (PRC) assign unequal redundancy to each layer in accordance with their importance. Each output symbol contains information from all layers, and thus the stream layers can be recovered progressively at the expected received ratios of output symbols. Furthermore, the dependency between layers is naturally considered. The performance of the PRC is evaluated and compared with some related UEP approaches. Results show that our PRC approach provides better recovery performance with lower overhead both theoretically and numerically.
Zhao Chen 0002, Liuguo Yin, Mai Xu, Jianhua Lu
VTC Fall3
2012 A General Transmission Scheme for Bi-Directional Communication by Using Eigenmode Sharing
abstract
In this paper, we develop a general transmission scheme for bi-directional communications by using eigenmode sharing, where existing two-way relaying protocols can be viewed as a special case of the proposed transmission scheme. In addition, the proposed bi-directional scheme can also be applied to more challenging scenarios with more than one source pair. Asymptotical behavior of the outage probability achieved by the proposed transmission protocol is studied in order to obtain insightful understandings for the fundamental limits of the proposed scheme. Our developed results show that the proposed bi-directional transmission scheme can realize larger system throughput than time sharing based approaches, and serving more than one pair at the same time is more beneficial than simple two-way relaying in terms of multiplexing gains.
Zhiguo Ding 0001, Mai Xu, Bayan S. Sharif, Jianhua Lu
IEEE J. Sel. Areas Commun.2
2011 3D Scene interpretation by combining probability theory and logic: The tower of knowledge
Mai Xu, Maria Petrou
Comput. Vis. Image Underst.1
2011 Learning Logic Rules for the Tower of Knowledge Using Markov Logic Networks
abstract
In this paper, we propose a novel logic-rule learning approach for the Tower of Knowledge (ToK) architecture, based on Markov logic networks, for scene interpretation. This approach is in the spirit of the recently proposed Markov logic networks for machine learning. Its purpose is to learn the soft-constraint logic rules for labeling the components of a scene. In our approach, FOIL (First Order Inductive Learner) is applied to learn the logic rules for MLN and then gradient ascent search is utilized to compute weights attached to each rule for softening the rules. This approach also benefits from the architecture of ToK, in reasoning whether a component in a scene has the right characteristics in order to fulfil the functions a label implies, from the logic point of view. One significant advantage of the proposed approach, rather than the previous versions of ToK, is its automatic logic learning capability such that the manual insertion of logic rules is not necessary. Experiments of labeling the identified components in buildings, for building scene interpretation, illustrate the promise of this approach.
Mai Xu, Maria Petrou, Jianhua Lu
Int. J. Pattern Recognit. Artif. Intell.1
2010 Component Identification in the 3D Model of a Building
abstract
This paper addresses the problem of identifying the components (such as balconies and windows) of the 3D model of a building. A novel method, based on a voting scheme, is presented for solving such a problem. It is intuitive that interference (such as shadows and occlusions) rarely happen at the same place or at different times when looking at a scene from different directions. In the spirit of this intuition, the voting-based method combines the information from various images to identify and segment the components of a building.
Mai Xu, Maria Petrou, Mohammad Jahangiri
ICPR1
2009 Learning Logic Rules for Scene Interpretation Based on Markov Logic Networks
Mai Xu, Maria Petrou
ACCV (3)1
2008 Recursive Tower of Knowledge for Learning to Interpret Scenes
abstract
The Tower of Knowledge architecture integrates probability theory and logic for making decisions. The scheme models the causal dependencies between the functionalities of objects and their descriptions, and then employs the maximum expected utility principle, which combines probability theory and logic, to select the most appropriate label for the object. Since most existing scene interpretation methods rely heavily on training data, we develop in this paper a recursive version of ToK to avoid such dependency. Recursive ToK learns the prior distributions iteratively from the decisions of labelling components made in the last iteration, partly by functionalities of components, and partly by the already learnt prior distributions in previous iterations. To validate our method in the domain of 3D outdoor scene interpretation, we compare ToK against a state-of-the-art method, Expandable Bayesian Networks (EBN), for labelling components of buildings. Experimental results then show that the labelling accuracy of ToK is superior to that of EBN. Also, these results reveal that recursive ToK improves the accuracy of ToK for labelling 3D components in the worst case when lacking any training data. 1
Mai Xu, Maria Petrou
BMVC1