VLDB 2026 Research / reviewers in the wild / expert
Shengxi Li
dblp:147/5453
· DBLP profile ↗
69ranked-venue papers
15as first author
53since 2021 · last 2026
0000-0003-4979-9290ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 48 · 10 first-author · 34 since 2021Artificial intelligence and machine learning · 27 · 6 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream TasksabstractIn recent years, the development of burst imaging technology has improved the capture and processing capabilities of visual data, enabling a wide range of applications. However, the redundancy in burst images leads to the increased storage and transmission demands, as well as reduced efficiency of downstream tasks. To address this, we propose a new task of Burst Image Quality Assessment (BuIQA), to evaluate the task-driven quality of each frame within a burst sequence, providing reasonable cues for burst image selection. Specifically, we establish the first benchmark dataset for BuIQA, consisting of 7,346 burst sequences with 45,827 images and 191,572 annotated quality scores for multiple downstream scenarios. Inspired by the data analysis, a unified BuIQA framework is proposed to achieve an efficient adaption for BuIQA under diverse downstream scenarios. Specifically, a task-driven prompt generation network is developed with heterogeneous knowledge distillation, to learn the priors of the downstream task. Then, the task-aware quality assessment network is introduced to assess the burst image quality based on the task prompt. Extensive experiments across 10 downstream scenarios demonstrate the impressive BuIQA performance of the proposed approach, outperforming the state-of-the-art. Furthermore, it can achieve 0.33 dB PSNR improvement in the downstream tasks of denoising and super-resolution, by applying our approach to select the high-quality burst frames. Xiaoye Liang, Lai Jiang 0004, Minglang Qiao, Yue Zhang 0082, Xin Deng 0002, Shengxi Li, Yufan Liu 0001, Mai Xu |
AAAI | 7 |
| 2026 | Dynamic Semantic Tokenization for Time Series via Elastic Sampling on Physics-aware PerceptionabstractDespite the remarkable success of semantic token learning in NLP and vision domains, token-level representation mechanisms face fundamental challenges when extended to continuous time series analysis. We identify a core limitation lies in the intrinsic absence of semantically meaningful tokenization boundaries within time-series, which differs substantially from discrete text tokens and presents unique complexities compared to spatially coherent image patches. While existing works mechanically apply fixed-length partitioning, recent evidence from time series foundation models reveals performance ceilings in prediction tasks under such paradigms. This paper introduces a novel tokenization framework known as physics-aware tokenization (PATK), designed to implement adaptive time-frequency tokenization via distribution-sensitive sampling strategies. Key innovations include: 1) A Rate-of-Variation (RoV) distribution is meticulously structured to encompass multi-scale temporal dynamics in the time domain, alongside a Spectral Energy Intensity (SEI) distribution devised to reveal global seasonal patterns within the frequency domain; 2) A physics-aware hidden Markov modeling (PA-HMM) is then established to adaptively breaks down continuous time-series into distinct tokens with elastic lengths, responding to physics-aware probabilities sampled from RoV and SEI distributions. The proposed PATK allows steady integration with both conventional Transformers and advanced large-scale time series models (including LLM-transferred methods and pretrained time series foundation models). Simulations across various datasets demonstrate that PATK excels in classification and forecasting tasks, showing notable adaptability to model long-term dependencies, strengthening resilience against disturbances, and robustness to missing data events. Huaizhang Liao, Zhixiong Yang 0001, Jingyuan Xia, Yuheng Sun, Yue Zhang 0082, Shengxi Li, Yongxiang Liu |
AAAI | 6 |
| 2026 | Say the image: Auditory masking effect-driven invertible network for progressive image-in-audio steganography
Jinghang Song, Fangyuan Gao, Xin Deng 0002, Shengxi Li, Mai Xu |
J. Inf. Secur. Appl. | 4 |
| 2026 | UniFES: A Unified Recurrent Network for Quality Enhancement and Stabilization in Face VideosabstractRecent years have witnessed an explosive increase of face content, which drives a distinct shift from static images to dynamic video formats. The shift of formats inherently alters the characteristics within face videos, whereby pixel-wise artifacts are intertwined with motion-related impairments. Addressing the emerging distortions that now always appear by twins in practice, however, is challenging and non-trivial, due to the distinct characteristics in addressing spatial-temporal frequencies in videos. In this paper, we propose a novel Unified recurrent network for joint Face video quality Enhancement and Stabilization (UniFES), as the first successful attempt for both quality enhancement and motion stabilization. Correspondingly, our UniFES method proposes to effectively aggregate the mutual information in the pixel and motion domains. For the quality enhancement, our UniFES method decomposes the shaking temporal alignment problem into progressive feature alignment with explicit physical information, which includes the global dynamics from the motion domain, i.e., from the stabilization task. Regarding the video stabilization, we integrate the mixed dynamics from the enhancement task (i.e., from pixel domain) to take into account both pixel-wise and motion-related characteristics, for ensuring robust trajectory estimation and motion stabilization. Subsequently, we refine the warping masks to achieve high-quality full frame rendering. We further establish a synthetic dataset for training and evaluation regarding this emerging task. Comprehensive experiments have illustrated the superior performances of our UniFES method over 32 comparing baselines on both newly established synthetic and real-world datasets. Mai Xu, Shengxi Li, Lai Jiang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Breaking the Multi-Enhancement Bottleneck: Domain-Consistent Quality Enhancement for Compressed ImagesabstractQuality enhancement methods have been widely integrated into visual communication pipelines to mitigate artifacts in compressed images. Ideally, these quality enhancement methods should perform robustly when applied to images that have already undergone prior enhancement during transmission. We refer to this scenario as multi-enhancement, which generalizes the well-known multi-generation scenario of image compression. Unfortunately, current quality enhancement methods suffer from severe degradation when applied in multi-enhancement.To address this challenge, we propose a novel adaptation method that transforms existing quality enhancement models into domain-consistent ones. Specifically, our method enhances a low-quality compressed image into a high-quality image within the natural domain during the first enhancement, and ensures that subsequent enhancements preserve this quality without further degradation. Extensive experiments validate the effectiveness of our method and show that various existing models can be successfully adapted to maintain both fidelity and perceptual quality in multi-enhancement scenarios. Qunliang Xing, Mai Xu, Shengxi Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Band-Kernel Stochastic Learning for Unsupervised Blind Hyperspectral Image Super-ResolutionabstractHyperspectral image super-resolution (HSI-SR) is fundamentally more difficult than RGB image SR, since its ultrahigh spectral dimensionality. Existing supervised methods rely on labeled training data to obtain data prior, which incurs prohibitive collection costs and limits generalization. Unsupervised methods individually preset the band and kernel with handcrafted priors, whereas this decoupling modeling artificially creates a complexity-performance trade-off in the selected band number. To address these issues, we propose BKX-HMM, a unified statistical framework for blind HSI-SR, which uniformly models the band selection, kernel estimation, and HSI restoration through the state transition of a hidden Markov model (HMM). BKX-HMM redefines the trade-off as a distributional fitting problem: each Markov transition progressively learns optimal parameters of full-band distribution via limited spectral observations. Based on BKX-HMM, we propose BKSR, the first unsupervised blind HSI-SR method, which consists of three synergistic modules: Gibbs sampling-based band selection (GBS), test-time-training kernel estimation (TKE), and robust HSI restoration (RHR). These modules form a closed-loop optimization cycle: i) In GBS, the dynamic ergodicity of Gibbs sampling provides a global spectral view for kernel estimation and HSI restoration while maintaining local spectral computations; ii) In TKE, the GBS-sampled bands guide the kernel estimator update, achieving a learnable sampling-based mechanism, which refines kernel estimation to regularize RHR's diffusion trajectory; iii) In RHR, a spectral hyper-Laplacian prior is integrated into the reverse process of an off-the-shelf diffusion model, which achieves non-i.i.d. noise robust HSI restoration, feedback reweights band and kernel importance for subsequent GBS and TKE iterations. Extensive experiments on both synthetic and real HSI datasets demonstrate our BKSR's superiority over baseline methods across diverse scenarios (e.g., unknown Gaussian/motion kernel, non-i.i.d. noise) while maintaining comparable computational costs to the classic band selection methods. Zhixiong Yang 0001, Jingyuan Xia, Shengxi Li, Lingyu Zheng, Shuanghui Zhang, Li Liu 0002, Yaowen Fu, Yongxiang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Multi-modal face anti-spoofing via self-supervised learning
Yufan Liu 0001, Lai Jiang 0004, Shengxi Li, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jinlong Lin |
Pattern Recognit. Lett. | 4 |
| 2026 | A Novel Visible-Infrared Image Compression Framework for High-Value Target ProtectionabstractThe joint compression of visible-infrared images is crucial for military and surveillance applications. The challenge lies in the protection of high-value targets (HVT) while maintaining high compression efficiency. This paper proposes a novel dual-stream compression framework that effectively addresses this challenge. In our framework, the sensitive HVT infrared signatures are concealed within the visible image stream, while residual infrared image is encoded separately. This dual-stream compression framework introduces three key innovations. 1) HVT protection: The HVT information is physically isolated and hidden within public visible images through a dedicated concealment stream; 2) Key-conditioned reconstruction: A novel decoding mechanism enables active camouflage by replacing HVTs with plausible background content when unauthorized access is detected; 3) Unified optimization: The framework integrates compression efficiency and HVT protection within an endto- end trainable network. Extensive experiments demonstrate that our approach achieves state-of-the-art compression performance while providing superior HVT protection, significantly outperforming traditional encrypt-then-compress methods. The code and weights are open-source athttps://github.com/eecoder-dyf/rgbir-compress. Yufan Deng, Xin Deng 0002, Shengxi Li, Xiaowan Hu, Mai Xu |
IEEE Signal Process. Lett. | 3 |
| 2026 | A New Compliant Bonder With Fast Constant-Force and Precise Self-Leveling Functions Toward Advanced Chip Packaging
Zhishen Liao, Yingjie Jia, Chengsi Huang, Songchao Tang, Shengxi Li, Yuzhang Wei, Hui Tang 0003 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | SportSal: Hypernetwork-Based Saliency Prediction for Sports VideosabstractSaliency prediction is crucial for improving sports video processing efficiency, thereby providing an enriched viewing experience for a wide-ranging audience. However, there is a long-term absence of well-established eye-tracking dataset and learning-based approach, particularly tailored for sports videos. In this paper, we establish a large-scale eye-tracking dataset dubbed audio-visual sports (AVS). AVS consists of 1,000 high-quality sports videos with eye fixations from 60 participants. Through data analysis on AVS, we observe that human attention patterns exhibit significant variations based on the specific scene context of the sports. Motivated by our observations, we propose a sports-aware saliency prediction approach, named SportSal, which can adaptively predict saliency maps in a hyper manner. Specifically, a hypernetwork is introduced to learn sports-aware priors. Meanwhile, an audio-visual fusion (AVF) block is developed to effectively fuse features from the visual and audio backbones. Given the learned priors and fused audio-visual features, we propose the hyper deformable convolutional (HDC) block and the hyper upsampling (HU) block for dynamic feature extraction and upsampling, respectively. The two blocks are alternatingly connected to adaptively predict saliency maps. Experimental results show that our approach outperforms 21 state-of-the-art saliency prediction approaches over three sports video eye-tracking datasets. Finally, we demonstrate the application of our SportSal approach in perceptual video compression. The dataset and code will be available at https://github.com/WeNsHiJIe-19950103/SportSal. Mai Xu, Shijie Wen, Lai Jiang 0004, Minglang Qiao, Shengxi Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Machines Serve Human: A Novel Variable Human-Machine Collaborative Compression FrameworkabstractHuman-machine collaborative compression has been receiving increasing research efforts for reducing image/video data, serving as the basis for both human perception and machine intelligence. Existing collaborative methods are dominantly built upon the de facto human-vision compression pipeline, witnessing deficiency on complexity and bit-rates when aggregating the machine-vision compression. Indeed, machine vision solely focuses on the core regions within the image/video, requiring much less information compared with the compressed information for human vision. In this paper, we thus set out the first successful attempt by a novel collaborative compression method based on the machine-vision-oriented compression, instead of human-vision pipeline. In other words, machine vision serves as the basis for human vision within collaborative compression. A plug-and-play variable bit-rate strategy is also developed for machine vision tasks. Then, we propose to progressively aggregate the semantics from the machine-vision compression, whilst seamlessly tailing the diffusion prior to restore high-fidelity details for human vision, thus named as diffusion-prior based feature compression for human and machine visions (Diff-FCHM). Experimental results verify the consistently superior performances of our Diff-FCHM, on both machine-vision and human-vision compression with remarkable margins. The source code is available at https://github.com/bblgbr/Diff-FCHM. Zifu Zhang, Shengxi Li, Xiancheng Sun, Mai Xu, Zhengyuan Liu, Jingyuan Xia |
IEEE Trans. Image Process. | 2 |
| 2025 | InstructOCR: Instruction Boosting Scene Text SpottingabstractIn the field of scene text spotting, previous OCR methods primarily relied on image encoders and pre-trained text information, but they often overlooked the advantages of incorporating human language instructions. To address this gap, we propose InstructOCR, an innovative instruction-based scene text spotting model that leverages human language instructions to enhance the understanding of text within images. Our framework employs both text and image encoders during training and inference, along with instructions meticulously designed based on text attributes. This approach enables the model to interpret text more accurately and flexibly. Extensive experiments demonstrate the effectiveness of our model and we achieve state-of-the-art results on widely used benchmarks. Furthermore, the proposed framework can be seamlessly applied to scene text VQA tasks. By leveraging instruction strategies during pre-training, the performance on downstream VQA tasks can be significantly improved, with a 2.6% increase on the TextVQA dataset and a 2.1% increase on the ST-VQA dataset. These experimental results provide insights into the benefits of incorporating human language instructions for OCR-related tasks. Chen Duan, Qianyi Jiang, Pei Fu, Shengxi Li, Shan Guo, Junfeng Luo |
AAAI | 5 |
| 2025 | Spherical Manifold Guided Diffusion Model for Panoramic Image GenerationabstractPanoramic image essentially acts as a pivotal role in emerging virtual reality and augmented reality scenarios; however, the generation of panoramic images are essentially challenging due to the intrinsic spherical geometry and spherical distortions caused by equirectangular projection (ERP). To address this, we start from the very basics of S2manifold inherent to panoramic images, and propose a novel spherical manifold convolution (SMConv) on S2manifold. Based on the SMConv operation, we propose a spherical manifold guided diffusion (SMGD) model for text-conditioned panoramic image generation, which can well accommodate the spherical geometry during generation. We further develop a novel evaluation method by calculating grouped Fréchet inception distance (FID) on cube-map projections, which can well reflect the quality of generated panoramic images, compared to existing methods that randomly crop ERP-distorted content. Experiment results demonstrate that our SMGD model achieves the state-of-the-art generation quality and accuracy, whilst retaining the shortest sampling time in the text-conditioned panoramic image generation task. Codes are publicly available at https://github.com/chronos123/SMGD. Xiancheng Sun, Mai Xu, Shengxi Li, Senmao Ma, Xin Deng 0002, Lai Jiang 0004 |
CVPR | 3 |
| 2025 | Quality Control For HEVC: A Deep Reinforcement Learning ApproachabstractIn video coding, large quality fluctuations exist in compressed videos, significantly degrading their quality of experience (QoE). Most works in literature focus on controlling bit-rates, however, paying few attention on reducing the quality fluctuations. In this paper, we propose a novel deep reinforcement learning (DRL) method for quality control in video coding. Specifically, we first propose the formulation of quality control, which targets at both controlling the target quality and reducing fluctuations. Then, we solve the quality control formulation by proposing a DRL method, in which the DRL elements are modeled by considering the features of both current frame and previous encoded frames. Specifically, for the DRL elements, we take the encoding information, content complexity and hidden features of long short-term memory (LSTM) as the state of DRL, and the selection of quantization parameters (QP) as the action of DRL. Subsequently, an algorithm, based on proximal policy optimization, is utilized to update our DRL model for decision-making on the actions of QP selection. In this way, the videos can be compressed under given and constant quality. We implement our DRL-based quality control method on the standard of high efficiency video coding (HEVC) with the HM 16.15 platform, and experimental results show that our method achieves the state-of-the-art performance on both quality control accuracy and fluctuations, in comparison with other quality control baselines. Mai Xu, Lai Jiang 0004, Shengxi Li, Xin Deng 0002 |
ICME | 5 |
| 2025 | SANE: Enhancing Large-scale Scene Representation with Semantic-aware NeRF ExpertsabstractWe propose the Semantic-aware NeRF Experts (SANE), which fully exploits the intrinsic characteristics of large-scale scenes, including semantics and material features, to achieve high-quality novel view synthesis results and provide accurate 3D semantic information. SANE begins by building a semantic Mixture of Experts (MoE), utilizing a learnable gating network to semantically partition the scene into blocks for corresponding NeRF experts. We then develop a semantic volume rendering scheme that integrates discrete semantics into the end-to-end differentiable process of NeRF, enabling refined semantic labeling of each scene point. Additionally, we implement a dual-implicit encoding strategy: intra-block encoding captures lighting variations across viewpoints, while inter-block one captures texture features among different semantic objects. Experiments on benchmark datasets show that SANE delivers higher-quality scene representations and effective semantic decomposition for downstream tasks, such as precise editing of large-scale scenes based on semantics. Zesheng Wang 0002, Yufeng Wang 0004, Shuangkang Fang, Dacheng Qi, Shengxi Li, Mai Xu, Wenrui Ding |
ICME | 6 |
| 2025 | Spherical-Nested Diffusion Model for Panoramic Image OutpaintingabstractPanoramic image outpainting acts as a pivotal role in immersive content generation, allowing for seamless restoration and completion of panoramic content. Given the fact that the majority of generative outpainting solutions operates on planar images, existing methods for panoramic images address the sphere nature by soft regularisation during the end-to-end learning, which still fails to fully exploit the spherical content. In this paper, we set out the first attempt to impose the sphere nature in the design of diffusion model, such that the panoramic format is intrinsically ensured during the learning procedure, named as spherical-nested diffusion (SpND) model. This is achieved by employing spherical noise in the diffusion process to address the structural prior, together with a newly proposed spherical deformable convolution (SDC) module to intrinsically learn the panoramic knowledge. Upon this, the proposed method is effectively integrated into a pre-trained diffusion model, outperforming existing state-of-the-art methods for panoramic image outpainting. In particular, our SpND method reduces the FID values by more than 50\% against the state-of-the-art PanoDiffusion method. Codes are publicly available at \url{https://github.com/chronos123/SpND}. Xiancheng Sun, Senmao Ma, Shengxi Li, Mai Xu, Jingyuan Xia, Lai Jiang 0004, Xin Deng 0002 |
ICML | 3 |
| 2025 | Collateral Circulation Guided Multi-Modality Fusion Network for Postoperative Infarct Prediction
Lisong Dai, Heming Dong, Lai Jiang 0004, Mai Xu, Shengxi Li |
MICCAI (15) | 7 |
| 2025 | Luminance-Aware Statistical Quantization: Unsupervised Hierarchical Learning for Illumination EnhancementabstractLow-light image enhancement (LLIE) faces persistent challenges in balancing reconstruction fidelity with cross-scenario generalization. While existing methods predominantly focus on deterministic pixel-level mappings between paired low/normal-light images, they often neglect the continuous physical process of luminance transitions in real-world environments, leading to performance drop when normal-light references are unavailable. Inspired by empirical analysis of natural luminance dynamics revealing power-law distributed intensity transitions, this paper introduces Luminance-Aware Statistical Quantification (LASQ), a novel framework that reformulates LLIE as a statistical sampling process over hierarchical luminance distributions. Our LASQ re-conceptualizes luminance transition as a power-law distribution in intensity coordinate space that can be approximated by stratified power functions, therefore, replacing deterministic mappings with probabilistic sampling over continuous luminance layers. A diffusion forward process is designed to autonomously discover optimal transition paths between luminance layers, achieving unsupervised distribution emulation without normal-light references.
In this way, it considerably improves the performance in practical situations, enabling more adaptable and versatile light restoration. This framework is also readily applicable to cases with normal-light references, where it achieves superior performance on domain-specific datasets alongside better generalization-ability across non-reference datasets. The code is available at: https://github.com/XYLGroup/LASQ. Derong Kong, Zhixiong Yang 0001, Shengxi Li, Shuaifeng Zhi, Li Liu 0002, Zhen Liu 0004, Jingyuan Xia |
NeurIPS | 3 |
| 2025 | Patch Inverter: A Novel Block-Wise GAN Inversion Method for Arbitrary Image ResolutionsabstractGenerative adversarial networks (GANs) have achieved remarkable progress in generating realistic images from merely small dimensions, which essentially establishes the latent generating space by rich semantics. GAN inversion thus aims at mapping real-world images back into the latent space, allowing for the access of semantics from images. However, existing GAN inversion methods can only invert images with fixed resolutions; this significantly restricts the representation capability in real-world scenarios. To address this issue, we propose to invert images by patches, thus named as patch inverter, which is the first attempt in terms of block-wise inversion for arbitrary resolutions. More specifically, we develop the padding-free operation to ensure the continuity across patches, and analyse the intrinsic mismatch within the inversion procedure. To relieve the mismatch, we propose a shifted convolution operation, which retains the continuity across image patches and simultaneously enlarges the receptive field for each convolution layer. We further propose the reciprocal loss to regularize the inverted latent codes to reside on the original latent generating space, such that the rich semantics can be maximally preserved. Experimental results have demonstrated that our patch inverter is able to accurately invert images with arbitrary resolutions, whilst representing precise and rich image semantics in real-world scenarios. Mai Xu, Shengxi Li, Zhenyu Guan 0002 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Continuous Patch Stitching for Block-Wise Image CompressionabstractMost recently, learned image compression methods have outpaced traditional hand-crafted standard codecs. However, their inference typically requires to input the whole image at the cost of heavy computing resources, especially for high-resolution image compression; otherwise, the block artefact can exist when compressed by blocks within existing learned image compression methods. To address this issue, we propose a novel continuous patch stitching (CPS) framework for block-wise image compression that is able to achieve seamlessly patch stitching and mathematically eliminate block artefact, thus capable of significantly reducing the required computing resources when compressing images. More specifically, the proposed CPS framework is achieved by padding-free operations throughout, with a newly established parallel overlapping stitching strategy to provide a general upper bound for ensuring the continuity. Upon this, we further propose functional residual blocks with even-sized kernels to achieve down-sampling and up-sampling, together with bottleneck residual blocks retaining feature size to increase network depth. Experimental results demonstrate that our CPS framework achieves the state-of-the-art performance against existing baselines, whilst requiring less than half of computing resources of existing models. The source code and trained models are available athttps://github.com/bblgbr/SPL-CPS. Zifu Zhang, Shengxi Li, Henan Liu, Mai Xu, Ce Zhu |
IEEE Signal Process. Lett. | 2 |
| 2025 | Hierarchical Semantic Compression for Consistent Image Semantic RestorationabstractThe emerging semantic compression has been receiving increasing research efforts most recently, capable of achieving high fidelity restoration during compression, even at extremely low bitrates. However, existing semantic compression methods typically combine standard pipelines with either pre-defined or high-dimensional semantics, thus suffering from deficiency in compression. To address this issue, we propose a novel hierarchical semantic compression (HSC) framework that purely operates within intrinsic semantic spaces from generative models, which is able to achieve efficient compression for consistent semantic restoration. More specifically, we first analyse the entropy models for the semantic compression, which motivates us to employ a hierarchical architecture based on a newly developed general inversion encoder. Then, we propose the feature compression network (FCN) and semantic compression network (SCN), such that the middle-level semantic feature and core semantics are hierarchically compressed to restore both accuracy and consistency of image semantics, via an entropy model progressively shared by channel-wise context. Experimental results demonstrate that the proposed HSC framework achieves the state-of-the-art performance on subjective quality and consistency for human vision, together with superior performances on machine vision tasks given compressed bitstreams. This essentially coincides with human visual system in understanding images, thus providing a new framework for future image/video compression paradigms. The source code and trained models are available at https://github.com/bblgbr/HSC-TIP2025. Shengxi Li, Zifu Zhang, Mai Xu, Lai Jiang 0004, Yufan Liu 0001, Ce Zhu |
IEEE Trans. Image Process. | 1 |
| 2025 | Spherical Patch Generative Adversarial Net for Unconditional Panoramic Image GenerationabstractRecent advancements in virtual reality (VR) and augmented reality (AR) have popularised the emerging panoramic content for the immersive visual experience. The difficulty in acquisition and display of 360° format further highlights the necessity of unconditional panoramic image generation. Existing methods essentially generate planar images mapped from panoramic images, and fail to address the deformation and closed-loop characteristics when inverted back to the panoramic images. Thus leading to the generation of pseudo-panoramic content. This paper aims to directly generate spherical content, in a patch-by-patch style; besides computation friendly, this promises the anywhere continuity on the panoramic image and proper accommodation of panoramic deformation. More specifically, we first propose a novel spherical patch convolution (SPConv) that operates on the local spherical patch, which naturally addresses the deformation of panoramic content. We then propose our spherical patch generative adversarial net (SP-GAN) that consists of spherical local embedding (SLE) and spherical content synthesiser (SCS) modules, which seamlessly incorporate our SPConv so as to generate continuous panoramic patches. To the best of our knowledge, the proposed SP-GAN is the first successful attempt to accommodate the spherical distortion for closed-loop panoramic image generation in a patch-by-patch manner. The experimental results, with human-rated evaluations, have verified the consistently superior performances for unconditional panoramic image generation, from the perspectives of generation quality, computational memory, and generalisation to various resolutions. Codes are publicly available at https://github.com/chronos123/SP-GAN. Mai Xu, Xiancheng Sun, Shengxi Li, Lai Jiang 0004, Jingyuan Xia, Xin Deng 0002 |
IEEE Trans. Image Process. | 3 |
| 2025 | Recruiting Teacher IF Modality for Nephropathy Diagnosis: A Customized Distillation Method With Attention-Based Diffusion NetworkabstractThe joint use of multiple modalities for medical image processing has been widely studied in recent years. The fusion of information from different modalities has demonstrated the performance improvement for a lot of medical tasks. For nephropathy diagnosis, immunofluorescence (IF) is one of the most widely-used multi-modality medical images due to its ease of acquisition and the effectiveness for certain nephropathy. However, the existing methods mainly assume different modalities have the equal effect on the diagnosis task, failing to exploit multi-modality knowledge in details. To avoid this disadvantage, this paper proposes a novel customized multi-teacher knowledge distillation framework to transfer knowledge from the trained single-modality teacher networks to a multi-modality student network. Specifically, a new attention-based diffusion network is developed for IF based diagnosis, considering global, local, and modality attention. Besides, a teacher recruitment module and diffusion-aware distillation loss are developed to learn to select the effective teacher networks based on the medical priors of the input IF sequence. The experimental results in the test and external datasets show that the proposed method has a better nephropathy diagnosis performance and generalizability, in comparison with the state-of-the-art methods. Mai Xu, Lai Jiang 0004, Yibing Fu, Xin Deng 0002, Shengxi Li |
IEEE Trans. Medical Imaging | 6 |
| 2025 | MDSC-Net: Multi-Modal Discriminative Sparse Coding Driven RGB-D Classification NetworkabstractIn this paper, we propose a novel sparsity-driven deep neural network to solve the RGB-D image classification problem. Different from existing classification networks, our network architecture is designed by drawing inspirations from a new proposed multi-modal discriminative sparse coding (MDSC) model. The key feature of this model is that it can gradually separate the discriminative and non-discriminative features in RGB-D images in a coarse-to-fine manner. Only the discriminative features are integrated and refined for classification, while the non-discriminative features are discarded, to improve the classification accuracy and efficiency. Derived from the MDSC model, the proposed network is composed of three modules, i.e., the shared feature extraction (SFE) module, discriminative feature refinement (DFR) module, and classification module. The architecture of each module is derived from the optimization solution in the MDSC model. To the best of our knowledge, this is the first time a fully sparsity-driven network has been proposed for RGB-D image classification. Extensive results verify the effectiveness of our method on different RGB-D image datasets. Xin Deng 0002, Yibing Fu, Mai Xu, Shengxi Li |
IEEE Trans. Multim. | 5 |
| 2024 | Enhancing Quality of Compressed Images by Mitigating Enhancement Bias Towards Compression DomainabstractExisting quality enhancement methods for compressed images focus on aligning the enhancement domain with the raw domain to yield realistic images. However, these methods exhibit a pervasive enhancement bias towards the compression domain, inadvertently regarding it as more realistic than the raw domain. This bias makes enhanced images closely resemble their compressed counterparts, thus degrading their perceptual quality. In this paper, we propose a simple yet effective method to mitigate this bias and enhance the quality of compressed images. Our method employs a conditional discriminator with the compressed image as a key condition, and then incorporates a domain-divergence regularization to actively distance the enhancement domain from the compression domain. Through this dual strategy, our method enables the discrimination against the compression domain, and brings the enhancement domain closer to the raw domain. Comprehensive quality evaluations confirm the superiority of our method over other state-of-the-art methods without incurring inference overheads. Qunliang Xing, Mai Xu, Shengxi Li, Xin Deng 0002, Meisong Zheng, Huaida Liu, Ying Chen 0011 |
CVPR | 3 |
| 2024 | A Dynamic Kernel Prior Model for Unsupervised Blind Image Super-ResolutionabstractDeep learning-based methods have achieved significant successes on solving the blind super-resolution (BSR) problem. However, most of them request supervised pretraining on labelled datasets. This paper proposes an unsupervised kernel estimation model, named dynamic kernel prior (DKP), to realize an unsupervised and pretraining-free learning-based algorithm for solving the BSR problem. DKP can adaptively learn dynamic kernel priors to realize real-time kernel estimation, and thereby enables superior HR image restoration performances. This is achieved by a Markov chain Monte Carlo sampling process on random kernel distributions. The learned kernel prior is then assigned to optimize a blur kernel estimation network, which entails a network-based Langevin dynamic optimization strategy. These two techniques ensure the accuracy of the kernel estimation. DKP can be easily used to replace the kernel estimation models in the existing methods, such as Double-DIP and FKP-DIP, or be added to the off-the-shelf image restoration model, such as diffusion model. In this paper, we incorporate our DKP model with DIP and diffusion model, referring to DIP-DKP and Diff-DKP, for validations. Extensive simulations on Gaussian and motion kernel scenarios demonstrate that the proposed DKP model can significantly improve the kernel estimation with comparable runtime and memory usage, leading to state-of-the-art BSR results. The code is available at https://github.com/XYLGroup/DKP. Zhixiong Yang 0001, Jingyuan Xia, Shengxi Li, Xinghua Huang, Shuanghui Zhang, Zhen Liu 0004, Yaowen Fu, Yongxiang Liu |
CVPR | 3 |
| 2024 | Saliency Prediction of Sports Videos: A Large-Scale Database and a Self-Adaptive ApproachabstractPredicting video saliency is crucial for improving sports video processing efficiency, thereby providing an enriched viewing experience for a wide-ranging audience. However, there is a long-term absence of well-established eye-tracking database and learning-based approach, particularly tailored for sports videos. In this paper, we establish a large-scale eye-tracking database dubbed audio-visual sports (AVS). AVS consists of 1,000 high-quality sports videos with eye fixations from 60 participants. Through the data analysis on AVS, we observe that human attention patterns exhibit significant variations based on the specific scene context of the sports. Motivated by this, we propose a sport-aware audiovisual saliency model, which can adaptively learn the scene context in a hyper manner. Specifically, a new audio-visual fusion (AVF) block is developed to effectively fuse features from the visual and audio backbone. After that, a hyper network is introduced to learn sport-aware priors, which are then adopted to guide the self-adaptive saliency predictor for predicting saliency map. Experimental results demonstrate that our approach outperforms other state-of-the-art saliency prediction models over the only two sports video eye-tracking databases. Minglang Qiao, Mai Xu, Shijie Wen, Lai Jiang 0004, Shengxi Li, Yunjin Chen, Leonid Sigal |
ICASSP | 5 |
| 2024 | Hybrid Single Input and Multiple Output Method For Compressing Features Towards Machine Vision TasksabstractWith the advance of deep learning in the BigData era, image/video coding for machines (VCM) as called for proposals by the moving picture experts group (MPEG) now becomes the pivotal technique for extensive intelligent vision tasks. However, existing VCM methods typically focus on compressing features independently at each scale, ignoring the redundancy of features across multiple scales. This paper thus introduces a simple yet effective architecture called hybrid single input and multiple output (H-SIMO) for VCM, which can significantly reduce the redundancy across scales of features. More specifically, as the pyramid structure is commonly employed for localising multi-scale objects, our H-SIMO method proposes to compress all features by inputting a single-scale feature while retaining the ability to decompress all the features. Moreover, an entropy model is seamlessly integrated into the training process to efficiently reduce the statistical redundancy of features. During the testing phase, the hybrid coding method, in conjunction with the versatile video coding (VVC), is employed to compress the features from both images and videos. We comprehensively evaluate the performance of our H-SIMO method in two standard machine vision tasks: object detection and instance segmentation, in which the experimental results verify the superior performances of our H-SIMO method. Zifu Zhang, Shengxi Li, Mai Xu, Zhenyu Guan 0002, Zhuoyi Lv |
ICIP | 2 |
| 2024 | QVD: Post-training Quantization for Video Diffusion ModelsabstractRecently, video diffusion models (VDMs) have garnered significant attention due to their notable advancements in generating coherent and realistic video content. However, processing multiple frame features concurrently, coupled with the considerable model size, results in high latency and extensive memory consumption, hindering their broader application. Post-training quantization (PTQ) is an effective technique to reduce memory footprint and improve computational efficiency. Unlike image diffusion, we observe that the temporal features, which are integrated into all frame features, exhibit pronounced skewness. Furthermore, we investigate significant inter-channel disparities and asymmetries in the activation of video diffusion models, resulting in low coverage of quantization levels by individual channels and increasing the challenge of quantization. To address these issues, we introduce the first PTQ strategy tailored for video diffusion models, dubbed QVD. Specifically, we propose the High Temporal Discriminability Quantization (HTDQ) method, designed for temporal features, which retains the high discriminability of quantized features, providing precise temporal guidance for all video frames. In addition, we present the Scattered Channel Range Integration (SCRI) method which aims to improve the coverage of quantization levels across individual channels. Experimental validations across various models, datasets, and bit-width settings demonstrate the effectiveness of our QVD in terms of diverse metrics. In particular, we achieve near-lossless performance degradation on W8A8, outperforming the current methods by 205.12 in FVD. Shilong Tian, Hong Chen 0014, Chengtao Lv, Yu Liu 0031, Jinyang Guo 0002, Xianglong Liu 0001, Shengxi Li, Hao Yang 0008 |
ACM Multimedia | 7 |
| 2024 | Causal Context Adjustment Loss for Learned Image CompressionabstractIn recent years, learned image compression (LIC) technologies have surpassed conventional methods notably in terms of rate-distortion (RD) performance. Most present learned techniques are VAE-based with an autoregressive entropy model, which obviously promotes the RD performance by utilizing the decoded causal context. However, extant methods are highly dependent on the fixed hand-crafted causal context. The question of how to guide the auto-encoder to generate a more effective causal context benefit for the autoregressive entropy models is worth exploring. In this paper, we make the first attempt in investigating the way to explicitly adjust the causal context with our proposed Causal Context Adjustment loss (CCA-loss). By imposing the CCA-loss, we enable the neural network to spontaneously adjust important information into the early stage of the autoregressive entropy model. Furthermore, as transformer technology develops remarkably, variants of which have been adopted by many state-of-the-art (SOTA) LIC techniques. The existing computing devices have not adapted the calculation of the attention mechanism well, which leads to a burden on computation quantity and inference latency. To overcome it, we establish a convolutional neural network (CNN) image compression model and adopt the unevenly channel-wise grouped strategy for high efficiency. Ultimately, the proposed CNN-based LIC network trained with our Causal Context Adjustment loss attains a great trade-off between inference latency and rate-distortion performance. Minghao Han, Shiyin Jiang, Shengxi Li, Xin Deng 0002, Mai Xu, Ce Zhu, Shuhang Gu |
NeurIPS | 3 |
| 2024 | Meta-learning based blind image super-resolution approach to different degradations
Zhixiong Yang 0001, Jingyuan Xia, Shengxi Li, Wende Liu, Shuaifeng Zhi, Shuanghui Zhang, Li Liu 0002, Yaowen Fu, Deniz Gündüz |
Neural Networks | 3 |
| 2024 | CrossHomo: Cross-Modality and Cross-Resolution Homography EstimationabstractMulti-modal homography estimation aims to spatially align the images from different modalities, which is quite challenging since both the image content and resolution are variant across modalities. In this paper, we introduce a novel framework namely CrossHomo to tackle this challenging problem. Our framework is motivated by two interesting findings which demonstrate the mutual benefits between image super-resolution and homography estimation. Based on these findings, we design a flexible multi-level homography estimation network to align the multi-modal images in a coarse-to-fine manner. Each level is composed of a multi-modal image super-resolution (MISR) module to shrink the resolution gap between different modalities, followed by a multi-modal homography estimation (MHE) module to predict the homography matrix. To the best of our knowledge, CrossHomo is the first attempt to address the homography estimation problem with both modality and resolution discrepancy. Extensive experimental results show that our CrossHomo can achieve high registration accuracy on various multi-modal datasets with different resolution gaps. In addition, the network has high efficiency in terms of both model complexity and running speed. Xin Deng 0002, Enpeng Liu, Shengxi Li, Shuhang Gu, Mai Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Assessing Face Image Quality: A Large-Scale Database and a Transformer MethodabstractThe amount of face images has been witnessing an explosive increase in the last decade, where various distortions inevitably exist on transmitted or stored face images. The distortions lead to visible and undesirable degradation on face images, affecting their quality of experience (QoE). To address this issue, this paper proposes a novel Transformer-based method for quality assessment on face images (named as TransFQA). Specifically, we first establish a large-scale face image quality assessment (FIQA) database, which includes 42,125 face images with diversifying content at different distortion types. Through an extensive crowdsource study, we obtain 712,808 subjective scores, which to the best of our knowledge contribute to the largest database for assessing the quality of face images. Furthermore, by investigating the established database, we comprehensively analyze the impacts of distortion types and facial components (FCs) on the overall image quality. Accordingly, we propose the TransFQA method, in which the FC-guided Transformer network (FT-Net) is developed to integrate the global context, face region and FC detailed features via a new progressive attention mechanism. Then, a distortion-specific prediction network (DP-Net) is designed to weight different distortions and accurately predict final quality scores. Finally, the experiments comprehensively verify that our TransFQA method significantly outperforms other state-of-the-art methods for quality assessment on face images. Shengxi Li, Mai Xu, Li Yang 0014, Xiaofei Wang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Blind Super-Resolution via Meta-Learning and Markov Chain Monte Carlo SimulationabstractLearning based approaches have witnessed great successes in blind single image super-resolution (SISR) tasks, however, handcrafted kernel priors and learning based kernel priors are typically required. In this paper, we propose a meta-learning and Markov Chain Monte Carlo (MCMC) based SISR approach to learn kernel priors from organized randomness. In concrete, a lightweight network is adopted as kernel generator, and is optimized via learning from the MCMC simulation on random Gaussian distributions. This procedure provides an approximation for the rational blur kernel, and introduces a network-level Langevin dynamics into SISR optimization processes, which contributes to preventing bad local optimal solutions for kernel estimation. Meanwhile, a meta-learning based alternating optimization procedure is proposed to optimize the kernel generator and image restorer, respectively. In contrast to the conventional alternating minimization strategy, a meta-learning based framework is applied to learn an adaptive optimization strategy, which is less-greedy and results in better convergence performance. These two procedures are iteratively processed in a plug-and-play fashion, for the first time, realizing a learning-based but plug-and-play blind SISR solution in unsupervised inference. Extensive simulations demonstrate the superior performance and generalization ability of the proposed approach when compared with the Start-of-the-Art solutions on synthesis and real-world datasets. Jingyuan Xia, Zhixiong Yang 0001, Shengxi Li, Shuanghui Zhang, Yaowen Fu, Deniz Gündüz, Xiang Li 0014 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Introduction to the Special Issue on AI-Generated Content for MultimediaabstractOur world is becoming rapidly dependent on data of increasing complexity, diversity, and volume which calls for robust and powerful tools to process such big data. Probabilistic generative models fulfill this goal by learning latent characteristic data relations, especially for the recent emergence of large-scale deep generative models that are able to create realistic content, namely, artificial intelligence-generated content (AIGC). The applications of AIGC span across various domains, and witness rich potential in multimedia content creation, including dialog generation, text-to-speech conversion, image/video generation, and cross-modal content generation. Shengxi Li, Xuelong Li 0001, Leonardo Chiariglione, Jiebo Luo 0001, Wenwu Wang 0001, Zhengyuan Yang, Danilo P. Mandic, Hamido Fujita |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Laplacian Gradient Consistency Prior for Flash Guided Non-Flash Image DenoisingabstractFor flash guided non-flash image denoising, the main challenge is to explore the consistency prior between the two modalities. Most existing methods attempt to model the flash/non-flash consistency in pixel level, which may easily lead to blurred edges. Different from these methods, we have an important finding in this paper, which reveals that the modality gap between flash and non-flash images conforms to the Laplacian distribution in gradient domain. Based on this finding, we establish a Laplacian gradient consistency (LGC) model for flash guided non-flash image denoising. This model is demonstrated to have faster convergence speed and denoising accuracy than the traditional pixel consistency model. Through solving the LGC model, we further design a deep network namely LGCNet. Different from existing image denoising networks, each component of the LGCNet strictly matches the solution of LGC model, giving the network good interpretability. The performance of the proposed LGCNet is evaluated on three different flash/non-flash image datasets, which demonstrates its superior denoising performance over many state-of-the-art methods both quantitatively and qualitatively. The intermediate features are also visualized to verify the effectiveness of the Laplacian gradient consistency prior. The source codes are available at https://github.com/JingyiXu404/LGCNet. Xin Deng 0002, Shengxi Li, Mai Xu |
IEEE Trans. Image Process. | 4 |
| 2023 | Learnt Mutual Feature Compression for Machine VisionabstractRecently, image coding for machines (ICM) has been playing an important role in facilitating intelligent vision tasks. Unfortunately, the existing ICM methods separately compress features at each scale, neglecting the redundancy across multi-scale features. To address this issue, this paper proposes an end-to-end mutual compression framework for the ICM, such that the compression efficiency can be significantly improved by removing the cross-scale redundancy. Specifically, the proposed framework consists of a mutual feature compression network (MFCNet) and a basic feature compression network (BFCNet). The MFCNet predicts large-scale features from basic small-scale features, such that the large amount of bitrates assigned to compress large-scale features can be saved. Moreover, the BFCNet is proposed to compress small-scale features of high quality by removing spatial and channel-wise redundancy. This guarantees superior performances whilst consuming extremely small amount of bit-rates. The experimental results show that our method achieves 90.10% and 74.97% BD-rate saving against the VVC feature anchor and VVC image anchor that have been recently accepted by the moving picture experts group (MPEG). Mai Xu, Shengxi Li, Li Yang 0014, Zhuoyi Lv |
ICASSP | 3 |
| 2023 | Neural Characteristic Function Learning for Conditional Image GenerationabstractThe emergence of conditional generative adversarial networks (cGANs) has revolutionised the way we approach and control the generation, by means of adversarially learning joint distributions of data and auxiliary information. Despite the success, cGANs have been consistently put under scrutiny due to their ill-posed discrepancy measure between distributions, leading to mode collapse and instability problems in training. To address this issue, we propose a novel conditional characteristic function generative adversarial network (CCF-GAN) to reduce the discrepancy by the characteristic functions (CFs), which is able to learn accurate distance measure of joint distributions under theoretical soundness. More specifically, the difference between CFs is first proved to be complete and optimisation-friendly, for measuring the discrepancy of two joint distributions. To relieve the problem of curse of dimensionality in calculating CF difference, we propose to employ the neural network, namely neural CF (NCF), to efficiently minimise an upper bound of the difference. Based on the NCF, we establish the CCF-GAN framework to explicitly decompose CFs of joint distributions, which allows for learning the data distribution and auxiliary information with classified importance. The experimental results on synthetic and real-world datasets verify the superior performances of our CCF-GAN, on both the generation quality and stability. Shengxi Li, Mai Xu, Xin Deng 0002 |
ICCV | 1 |
| 2023 | Residual based hierarchical feature compression for multi-task machine visionabstractWith the remarkable success of deep learning, image/video coding for machines (VCM) has been playing an important role in facilitating intelligent vision tasks. However, the existing VCM methods suffer from either sub-optimality of using image compression standards, or generalisation issues of learning-based methods. To address these issues, this paper proposes a residual-based hierarchical feature compression (RHFC) method to achieve optimal and universal feature compression for object detection and segmentation. More specifically, we first analyse the redundancy that exists in features at multiple scales, by finding that large-scale features are surprisingly less important to the vision tasks. Thus, we propose a pair of compression and enhancement networks to extract the very basic cues from the large-scale features, which are then compressed by the VVC codec. To compensate the inevitable detail loss, we further propose the hierarchical framework to compress the residuals between the reconstructed and original features, such that the performances can be significantly improved at low bit-rate cost. Experimental results have verified our superior performances, against both the state-of-the-art learning-based and standard feature compression methods. Our RHFC method also generalises well to other scenarios without the need of any further fine-tuning. Mai Xu, Shengxi Li, Minglang Qiao, Zhuoyi Lv |
ICME | 3 |
| 2023 | Optimizing DNN based quality assessment metric for image compression: A novel rate control methodabstractIn the existing coding standards, rate control (RC) plays a critical role in optimally allocating bit-rates to each coding unit, for improving rate-distortion performance under the limited bandwidth. However, the existing RC methods are mainly based on traditional distortion metrics, which fail to take the advantage of the emerging DNN based image quality assessment (IQA) metrics. In this paper, we set up the first attempt to achieve IQA score based RC for image compression. Specifically, a novel visualization based score-distortion (VSD) model and ρ-slope model are proposed to explicitly establish the relationship between IQA score and bit-rates. Then, by solving optimal rate-distortion optimization based on the IQA score, we propose a novel RC method for the HEVC standard. The experimental results show that, given the target bit-rates, the proposed RC method can accurately control the bit-rates and generate the compressed images with higher IQA score and better perceptual quality. More importantly, the proposed RC method is evaluated to be effective over two DNN based IQA metrics and four image datasets, exhibiting the potential in practical use. The code is available at https://github.com/Ffangqy/IQA-RC. Qiuyue Fang, Lai Jiang 0004, Shengxi Li, Mai Xu, Yunjin Chen, Leonid Sigal |
ICME | 4 |
| 2023 | Predicting the Invariance Behind Residuals: A Novel GAN Inversion Method for Image Editing and Detail RetainingabstractGenerative adversarial network (GAN) inversion has been serving as the vehicle to enable the restoration of real-world images by GANs, rather than the realistic generation from random noise. Existing GAN inversion methods, however, suffer from the fidelity-editability trade-off, which mainly invert the semantics within images and fail to reconstruct the details. To address this issue, we propose a novel adaptive detail compensation method for GAN inversion (ADC-GInv), which automatically locates and restores the non-semantic details, whilst maintaining the editability on the semantics of real-world images. More specifically, we first develop an adversarial reciprocal learning GAN (ARL-GAN) so as to seamlessly optimise the reconstruction during the training of GANs, followed by a sophisticated fine-tuning technique for ARL-GAN inversion. This ensures superior restoration and editability on the semantic cues of images. Then, regarding the non-semantic details, our ADC-GInv method adaptively locates the details by predicting the invariance given edited and non-edited residuals of restoration, which are then compensated at the pixel-level for high fidelity. As a consequence, the experimental results have verified the superior performance of our ADC-GInv, on both fidelity and editability during inversion. Zhimo Yan, Hengyang He, Shengxi Li, Mai Xu, Ce Zhu |
MMSP | 4 |
| 2023 | Learned Structure-Based Hybrid Framework for Martian Image CompressionabstractRecent landing marches on Mars have enabled the access to Martian surface images, which act as an important vehicle to demystify the evolution and habitability of Mars, in terms of climate, geography, etc. Transmitting Martian images thus calls for efficient compression methods to ensure the high-quality reconstruction from distant communication, in which the research is yet to start. To address this issue, we propose in this letter a learned structure-based hybrid (LSH) framework to compress Martian images. More specifically, we first observe that the structural consistency exists across Martian images, which motivates us to propose a structural compression network (SCN). The aim of SCN is to compactly represent the structural information of Martian images, thus allowing for the compression at extremely low bit-rates. Then, we propose a detail compensation network (DCN) to reconstruct the missing details when we restore from the structural information, which benefits from improved compression efficiency by reduced bit-rates. The experimental results have verified the superior performances of our LSH method on compressing Martian images, against existing state-of-the-art methods. Shengxi Li, Xiancheng Sun, Mai Xu, Lai Jiang 0004 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Reciprocal GAN Through Characteristic Functions (RCF-GAN)abstractThe integral probability metric (IPM) equips generative adversarial nets (GANs) with the necessary theoretical support for comparing statistical moments in an embedded domain of the critic, while stabilising their training and mitigating the mode collapse issues. For enhanced intuition and physical insight, we introduce a generalisation of IPM-GANs which operates by directly comparing probability distributions rather than their moments. This is achieved through characteristic functions (CFs), a powerful tool that uniquely comprises all information about any general distribution. For rigour, we first theoretically prove the ability of the CF loss to compare probability distributions, and proceed to establish the physical meaning of the phase and amplitude of CFs. An optimal sampling strategy is then developed to calculate the CFs, and an equivalence between the embedded and data domains is proved under the reciprocal theory. This makes it possible to seamlessly combine IPM-GAN with an auto-encoder structure by an advanced anchor architecture, which adversarially learns a semantic low-dimensional manifold for both generation and reconstruction. This efficient reciprocal CF GAN (RCF-GAN) structure, uses only two modules and a simple training strategy to achieve the state-of-the-art bi-directional generation. Experiments demonstrate the superior performance of RCF-GAN on both regular (images) and irregular (graph) domains. Shengxi Li, Zeyang Yu, Min Xiang, Danilo P. Mandic |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Interpretable Multi-Modal Image Registration Network Based on Disentangled Convolutional Sparse CodingabstractMulti-modal image registration aims to spatially align two images from different modalities to make their feature points match with each other. Captured by different sensors, the images from different modalities often contain many distinct features, which makes it challenging to find their accurate correspondences. With the success of deep learning, many deep networks have been proposed to align multi-modal images, however, they are mostly lack of interpretability. In this paper, we first model the multi-modal image registration problem as a disentangled convolutional sparse coding (DCSC) model. In this model, the multi-modal features that are responsible for alignment (RA features) are well separated from the features that are not responsible for alignment (nRA features). By only allowing the RA features to participate in the deformation field prediction, we can eliminate the interference of the nRA features to improve the registration accuracy and efficiency. The optimization process of the DCSC model to separate the RA and nRA features is then turned into a deep network, namely Interpretable Multi-modal Image Registration Network (InMIR-Net). To ensure the accurate separation of RA and nRA features, we further design an accompanying guidance network (AG-Net) to supervise the extraction of RA features in InMIR-Net. The advantage of InMIR-Net is that it provides a universal framework to tackle both rigid and non-rigid multi-modal image registration tasks. Extensive experimental results verify the effectiveness of our method on both rigid and non-rigid registrations on various multi-modal image datasets, including RGB/depth images, RGB/near-infrared (NIR) images, RGB/multi-spectral images, T1/T2 weighted magnetic resonance (MR) images and computed tomography (CT)/MR images. The codes are available at https://github.com/lep990816/Interpretable-Multi-modal-Image-Registration. Xin Deng 0002, Enpeng Liu, Shengxi Li, Yiping Duan, Mai Xu |
IEEE Trans. Image Process. | 3 |
| 2023 | Blind VQA on 360° Video via Progressively Learning From Pixels, Frames, and VideoabstractBlind visual quality assessment (BVQA) on 360° video plays a key role in optimizing immersive multimedia systems. When assessing the quality of 360° video, human tends to perceive its quality degradation from the viewport-based spatial distortion of each spherical frame to motion artifact across adjacent frames, ending with the video-level quality score, i.e., a progressive quality assessment paradigm. However, the existing BVQA approaches for 360° video neglect this paradigm. In this paper, we take into account the progressive paradigm of human perception towards spherical video quality, and thus propose a novel BVQA approach (namely ProVQA) for 360° video via progressively learning from pixels, frames and video. Corresponding to the progressive learning of pixels, frames and video, three sub-nets are designed in our ProVQA approach, i.e., the spherical perception aware quality prediction (SPAQ), motion perception aware quality prediction (MPAQ) and multi-frame temporal non-local (MFTN) sub-nets. The SPAQ sub-net first models the spatial quality degradation based on spherical perception mechanism of human. Then, by exploiting motion cues across adjacent frames, the MPAQ sub-net properly incorporates motion contextual information for quality assessment on 360° video. Finally, the MFTN sub-net aggregates multi-frame quality degradation to yield the final quality score, via exploring long-term quality correlation from multiple frames. The experiments validate that our approach significantly advances the state-of-the-art BVQA performance on 360° video over two datasets, the code of which has been public in https://github.com/yanglixiaoshen/ProVQA. Li Yang 0014, Mai Xu, Shengxi Li, Zulin Wang |
IEEE Trans. Image Process. | 3 |
| 2023 | Von Mises-Fisher Elliptical DistributionabstractModern probabilistic learning systems mainly assume symmetric distributions, however, real-world data typically obey skewed distributions and are thus not adequately modeled through symmetric distributions. To address this issue, a generalization of symmetric distributions called elliptical distributions are increasingly used, together with further improvements based on skewed elliptical distributions. However, existing approaches are either hard to estimate or have complicated and abstract representations. To this end, we propose a novel approach based on the von-Mises-Fisher (vMF) distribution to obtain an explicit and simple probability representation of skewed elliptical distributions. The analysis shows that this not only allows us to design and implement nonsymmetric learning systems but also provides a physically meaningful and intuitive way of generalizing skewed distributions. For rigor, the proposed framework is proven to share important and desirable properties with its symmetric counterpart. The proposed vMF distribution is demonstrated to be easy to generate and stable to estimate, both theoretically and through examples. Shengxi Li, Danilo P. Mandic |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Metalearning-Based Alternating Minimization Algorithm for Nonconvex OptimizationabstractIn this article, we propose a novel solution for nonconvex problems of multiple variables, especially for those typically solved by an alternating minimization (AM) strategy that splits the original optimization problem into a set of subproblems corresponding to each variable and then iteratively optimizes each subproblem using a fixed updating rule. However, due to the intrinsic nonconvexity of the original optimization problem, the optimization can be trapped into a spurious local minimum even when each subproblem can be optimally solved at each iteration. Meanwhile, learning-based approaches, such as deep unfolding algorithms, have gained popularity for nonconvex optimization; however, they are highly limited by the availability of labeled data and insufficient explainability. To tackle these issues, we propose a meta-learning based alternating minimization (MLAM) method that aims to minimize a part of the global losses over iterations instead of carrying minimization on each subproblem, and it tends to learn an adaptive strategy to replace the handcrafted counterpart resulting in advance on superior performance. The proposed MLAM maintains the original algorithmic principle, providing certain interpretability. We evaluate the proposed method on two representative problems, namely, bilinear inverse problem: matrix completion and nonlinear problem: Gaussian mixture models. The experimental results validate the proposed approach outperforms AM-based methods. Jingyuan Xia, Shengxi Li, Junjie Huang 0001, Zhixiong Yang 0001, Imad Jaimoukha, Deniz Gündüz |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Does text attract attention on e-commerce images: A novel saliency prediction dataset and methodabstractE-commerce images are playing a central role in attracting people's attention when retailing and shopping online, and an accurate attention prediction is of significant importance for both customers and retailers, where its research is yet to start. In this paper, we establish the first dataset of saliency e-commerce images (SalECI), which allows for learning to predict saliency on the e-commerce images. We then provide specialized and thorough analysis by high-lighting the distinct features of e-commerce images, e.g., non-locality and correlation to text regions. Correspondingly, taking advantages of the non-local and self-attention mechanisms, we propose a salient SWin-Transformer back-bone, followed by a multi-task learning with saliency and text detection heads, where an information flow mechanism is proposed to further benefit both tasks. Experimental results have verified the state-of-the-art performances of our work in the e-commerce scenario. Lai Jiang 0004, Shengxi Li, Mai Xu, Se Lei |
CVPR | 3 |
| 2022 | Low-Complexity Attention Modelling via Graph Tensor NetworksabstractThe attention mechanism is at the core of modern Natural Language Processing (NLP) models, owing to its ability to focus on the most contextually relevant part of a sequence. However, current attention models rely on "flat-view" matrix methods to process tokens embedded in vector spaces; this results in exceedingly high parameter complexity which is prohibitive for practical applications. To this end, we introduce a novel Tensorized Graph Attention (TGA) mechanism, which leverages on the recent Graph Tensor Network (GTN) framework to efficiently process tensorized token embeddings via attention based graph filters. Such tensorized token embeddings are shown to effectively bypass the Curse of Dimensionality, reducing the parameter complexity of the attention mechanism from an exponential to a linear one in the embedding dimensions. The expressive power of the TGA framework is further enhanced by virtue of domain-aware graph convolution filters. Simulations across benchmark NLP paradigms verify the advantages of the proposed framework over existing attention models, at drastically lower parameter complexity. Yao Lei Xu, Kriton Konstantinidis, Shengxi Li, Ljubisa Stankovic, Danilo P. Mandic |
ICASSP | 3 |
| 2022 | A Learning-based Approach for Martian Image CompressionabstractFor the scientific exploration and research on Mars, it is an indispensable step to transmit high-quality Martian images from distant Mars to Earth. Image compression is the key technique given the extremely limited Mars-Earth bandwidth. Recently, deep learning has demonstrated remarkable performance in natural image compression, which provides a possibility for efficient Martian image compression. However, deep learning usually requires large training data. In this paper, we establish the first large-scale high-resolution Martian image compression (MIC) dataset. Through analyzing this dataset, we observe an important non-local self-similarity prior for Marian images. Benefiting from this prior, we propose a deep Martian image compression network with the non-local block to explore both local and non-local dependencies among Martian image patches. Experimental results verify the effectiveness of the proposed network in Martian image compression, which outperforms both the deep learning based compression methods and HEVC codec. Mai Xu, Shengxi Li, Xin Deng 0002, Qiu Shen |
VCIP | 3 |
| 2022 | MRS-Net+ for Enhancing Face Quality of Compressed VideosabstractDuring the past few years, face videos, e.g., video conference, interviews and variety shows, have grown explosively with millions of users over social media networks. Unfortunately, the existing compression algorithms are applied to these videos for reducing bandwidth, which also bring annoying artifacts to face regions. This paper addresses the problem of face quality enhancement in compressed videos by reducing the artifacts of face regions. Specifically, we establish a compressed face video (CFV) database, which includes 196,337 faces in 214 high-quality video sequences and their corresponding 1,712 compressed sequences. We find that the faces of compressed videos exhibit tremendous scale variation and quality fluctuation. Motivated by scalable video coding, we propose a multi-scale recurrent scalable network (MRS-Net+) to enhance the quality of multi-scale faces in compressed videos. The MRS-Net+ is comprised by one base and two refined enhancement levels, corresponding to the quality enhancement of small-, medium- and large-scale faces, respectively. In the multi-level architecture of our MRS-Net+, small-/medium-scale face quality enhancement serves as the basis for facilitating the quality enhancement of medium-/large-scale faces. We further develop a landmark-assisted pyramid alignment (LPA) subnet to align faces across consecutive frames, and then apply the mask-guided quality enhancement (QE) subnet for enhancing multi-scale faces. Finally, experimental results show that our MRS-Net+ method achieves averagely 1.196 dB improvement of peak signal-to-noise ratio (PSNR) and 23.54% saving of Bjøntegaard distortion-rate (BD-rate), significantly outperforming other state-of-the-art methods. Mai Xu, Shengxi Li, Huaida Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Kernel Learning with Tensor NetworksabstractThe expressive power of Gaussian Processes (GPs) is largely attributed to their kernel function, which highlights the crucial role of kernel design. Efforts in this direction include modern Neural Network (NN) based kernel design, which despite success, suffers from the lack of interpretability and tendency to overfit. To this end, we introduce a Tensor Network (TN) approach to learning kernel embeddings, with a TN serving to map the input to a low dimensional manifold, where a suitable base kernel function can be applied. The proposed framework allows for joint learning of the TN and base kernel parameters using stochastic variational inference, while leveraging on the low-rank regularization and multi-linear nature of TNs to boost model performance and provide enhanced interpretability. Performance evaluation within the regression paradigm against TNs and Deep Kernels demonstrates the potential of the framework, providing conclusive evidence for promising future extensions to other learning paradigms. Kriton Konstantinidis, Shengxi Li, Danilo P. Mandic |
ICASSP | 2 |
| 2021 | A Universal Framework for Learning the Elliptical Mixture ModelabstractMixture modeling using elliptical distributions promises enhanced robustness, flexibility, and stability over the widely employed Gaussian mixture model (GMM). However, existing studies based on the elliptical mixture model (EMM) are restricted to several specific types of elliptical probability density functions, which are not supported by general solutions or systematic analysis frameworks; this significantly limits the rigor in the design and power of EMMs in applications. To this end, we propose a novel general framework for estimating and analyzing the EMMs, achieved through the Riemannian manifold optimization. First, we investigate the relationships between Riemannian manifolds and elliptical distributions, and the so established connection between the original manifold and a reformulated one indicates a mismatch between these manifolds, a major cause of failure of the existing optimization for solving general EMMs. We next propose a universal solver that is based on the optimization of a redesigned cost and prove the existence of the same optimum as in the original problem; this is achieved in a simple, fast and stable way. We further calculate the influence functions of the EMM as theoretical bounds to quantify robustness to outliers. Comprehensive numerical results demonstrate the ability of the proposed framework to accommodate EMMs with different properties of individual functions in a stable way and with fast convergence speed. Finally, the enhanced robustness and flexibility of the proposed framework over the standard GMM are demonstrated both analytically and through comprehensive simulations. Shengxi Li, Zeyang Yu, Danilo P. Mandic |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Solving General Elliptical Mixture Models through an Approximate Wasserstein ManifoldabstractWe address the estimation problem for general finite mixture models, with a particular focus on the elliptical mixture models (EMMs). Compared to the widely adopted Kullback–Leibler divergence, we show that the Wasserstein distance provides a more desirable optimisation space. We thus provide a stable solution to the EMMs that is both robust to initialisations and reaches a superior optimum by adaptively optimising along a manifold of an approximate Wasserstein distance. To this end, we first provide a unifying account of computable and identifiable EMMs, which serves as a basis to rigorously address the underpinning optimisation problem. Due to a probability constraint, solving this problem is extremely cumbersome and unstable, especially under the Wasserstein distance. To relieve this issue, we introduce an efficient optimisation method on a statistical manifold defined under an approximate Wasserstein distance, which allows for explicit metrics and computable operations, thus significantly stabilising and improving the EMM estimation. We further propose an adaptive method to accelerate the convergence. Experimental results demonstrate the excellent performance of the proposed EMM solver. Shengxi Li, Zeyang Yu, Min Xiang, Danilo P. Mandic |
AAAI | 1 |
| 2020 | MRS-Net: Multi-Scale Recurrent Scalable Network for Face Quality Enhancement of Compressed VideosabstractThe past decade has witnessed the explosive growth of faces in video multimedia systems, e.g., videoconferencing and live shows. However, these videos are normally compressed at low bit-rates due to the bandwidth-hungry issue, leading to heavy quality degradation on face regions. This paper addresses the problem of face quality enhancement in compressed videos. Specifically, we establish a compressed face video (CFV) database, which includes 87,607 faces in 113 raw video sequences and their corresponding 904 compressed sequences. We find that the faces of compressed videos exhibit tremendous scale variation and quality fluctuation. Motivated by scalable video coding, we propose a multi-scale recurrent scalable network (MRS-Net) to enhance the quality of multi-scale faces in compressed videos. The MRS-Net is comprised by one base and two refined enhancement levels, corresponding to the quality enhancement of small-, medium- and large-scale faces, respectively. In the multi-level architecture of our MRS-Net, small-/medium-scale face quality enhancement serves as the basis for facilitating the quality enhancement of medium-/large-scale faces. Finally, experimental results show that our MRS-Net method is effective in enhancing the quality of multi-scale faces for compressed videos, significantly outperforming other state-of-the-art methods. Mai Xu, Shengxi Li, Huaida Liu |
ACM Multimedia | 3 |
| 2020 | Reciprocal Adversarial Learning via Characteristic FunctionsabstractGenerative adversarial nets (GANs) have become a preferred tool for tasks involving complicated distributions. To stabilise the training and reduce the mode collapse of GANs, one of their main variants employs the integral probability metric (IPM) as the loss function. This provides extensive IPM-GANs with theoretical support for basically comparing moments in an embedded domain of the \textit{critic}. We generalise this by comparing the distributions rather than their moments via a powerful tool, i.e., the characteristic function (CF), which uniquely and universally comprising all the information about a distribution. For rigour, we first establish the physical meaning of the phase and amplitude in CF, and show that this provides a feasible way of balancing the accuracy and diversity of generation. We then develop an efficient sampling strategy to calculate the CFs. Within this framework, we further prove an equivalence between the embedded and data domains when a reciprocal exists, where we naturally develop the GAN in an auto-encoder structure, in a way of comparing everything in the embedded space (a semantically meaningful manifold). This efficient structure uses only two modules, together with a simple training strategy, to achieve bi-directionally generating clear images, which is referred to as the reciprocal CF GAN (RCF-GAN). Experimental results demonstrate the superior performances of the proposed RCF-GAN in terms of both generation and reconstruction. Shengxi Li, Zeyang Yu, Min Xiang, Danilo P. Mandic |
NeurIPS | 1 |
| 2019 | Enhanced Normalized Mean Error loss for Robust Facial Landmark detection
Shenqi Lai, Zhenhua Chai, Shengxi Li, Huanhuan Meng, Mengzhao Yang, Xiaoming Wei |
BMVC | 3 |
| 2019 | Widely Linear Complex-Valued Autoencoder: Dealing with Noncircularity in Generative-Discriminative Models
Zeyang Yu, Shengxi Li, Danilo P. Mandic |
ICANN (1) | 2 |
| 2019 | Tracking Dynamic Systems in α-Stable EnvironmentsabstractIn order to accommodate for modern adaptive filtering applications, the classic adaptive filtering paradigm is considered from a more general perspective. The new formulation allows for time dependent variations in the state of the system and more importantly it relaxes the Gaussian assumption to the generalized setting of α-stable distributions. In this work, based on the principles of gradient descent and fractional-order calculus, a cost-effective technique for tracking the state of such a system is derived. For rigour, performance of the derived filtering technique is analyzed and convergence conditions are established. Sayed Pouria Talebi, Stefan Werner 0001, Shengxi Li, Danilo P. Mandic |
ICASSP | 3 |
| 2018 | Closed-Form Optimization on Saliency-Guided Image Compression for HEVC-MSPabstractHigh efficiency video coding (HEVC) is the latest video coding standard, and it has the best performance among all the existing standards. HEVC main still picture profile (HEVC-MSP) also achieves top performance in image compr-ession. In this paper, we propose a closed-form bit allocation approach to optimize the saliency-guided PSNR (viewed as perceptual distortion) such that the coding efficiency of HEVC-based image compression can be significantly improved from a subjective perspective. Specifically, a bit allocation formulation is established to minimize perceptual distortion with a constraint on bit-rates. Then, this formulation is solved using the proposed recursive Taylor expansion method with a closed-form solution. On the basis of our solution, a bit allocation and re-allocation process is developed in our approach to minimize perceptual distortion, meanwhile accurately controlling bit-rates. In addition, we provide both theoretical and numerical analyses of the computational complexity, verifying the little extra time cost of our approach. The experimental results demonstrate the superior performance of our approach over the state-of-the-art HEVC-MSP, and the BD-rate savings are approximately 40% and 24% for face and generic images, respectively. Shengxi Li, Mai Xu, Yun Ren, Zulin Wang |
IEEE Trans. Multim. | 1 |
| 2017 | Watching Videos with Certain and Constant Quality: PID-Based Quality Control Method
Yuhang Song 0001, Mai Xu, Shengxi Li |
DCC | 3 |
| 2017 | A novel rate control scheme for panoramic video codingabstractThe popularity of multi-view panoramic videos has been considerably increased for producing Virtual Reality (VR) content, due to its immersive visual experience. We argue in this paper that PSNR is less effective in assessing visual quality of compressed panoramic videos than Sphere-based PSNR (S-PNSR), in which sphere-to-plain mapping of panoramic videos is considered. Thus, the conventional rate control (R-C) schemes of 2-Dimensional (2D) video coding, which optimize on PSNR, are not suitable for panoramic video coding. To optimize S-PSNR, we propose in this paper a novel RC scheme for panoramic video coding. Specifically, we develop an S-PSNR optimization formulation with constraint on bit-rate. Then, a solution is provided to the developed formulation, such that bits can be allocated to each coding block for achieving optimal S-PSNR in panoramic video coding. Finally, the experiment results validate the effectiveness of the proposed RC scheme in improving S-PSNR of panoramic video coding. Yufan Liu 0001, Mai Xu, Chen Li 0049, Shengxi Li, Zulin Wang |
ICME | 4 |
| 2017 | Optimal Bit Allocation for CTU Level Rate Control in HEVCabstractFor High Efficiency Video Coding (HEVC), the R–$\lambda $scheme is the latest rate control (RC) scheme, which investigates the relationships among allocated bits, the slope of rate-distortion (R-D) curve$\lambda $, and quantization parameter. However, we argue that bit allocation in the existing R–$\lambda $scheme is not optimal. In this paper, we therefore propose an optimal bit allocation (OBA) scheme for coding tree unit level RC in HEVC. Specifically, to achieve the OBA, we first develop an optimization formulation with a novel R-D estimation, instead of the existing R–$\lambda $estimation. Unfortunately, it is intractable to obtain a closed-form solution to the optimization formulation. We thus propose a recursive Taylor expansion (RTE) method to iteratively solve the formulation. As a result, an approximate closed-form solution can be obtained, thus achieving OBA and bit reallocation. Both theoretical and numerical analyses show the fast convergence speed and little computational time of the proposed RTE method. Therefore, our OBA scheme can be achieved at little encoding complexity cost. Finally, the experimental results validate the effectiveness of our scheme in three aspects: R-D performance, RC accuracy, and robustness over dynamic scene changes. Shengxi Li, Mai Xu, Zulin Wang, Xiaoyan Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Optimizing Subjective Quality in HEVC-MSP: An Approximate Closed-form Image Compression ApproachabstractHEVC, as the latest video coding standard, achieves top performance on image compression. On the basis of this, we propose a novel approach to optimize subjective quality for HEVC-based image compression. Specifically, a bit allocation formulation is established to optimize subjective quality with constraint on bit-rates. Then, we propose a recursive Taylor expansion method to quickly solve such a formulation with an approximate closed-form solution. The experimental results show the superior performance of our approach, with ~40% BD-rate saving over the state-of-the-art HEVC-MSP for face image compression. Shengxi Li, Mai Xu, Yun Ren, Chengzhang Ma, Zulin Wang |
DCC | 1 |
| 2015 | A novel method on optimal bit allocation at LCU level for rate control in HEVCabstractIn this paper, we propose a new method, namely recursive Taylor expansion (RTE) method, for optimally allocating bits to each LCU in the R-λ rate control scheme for HEVC. Specifically, we first set up an optimization formulation on optimal bit allocation. Unfortunately, it is intractable to achieve a closed-form solution for this formulation. We therefore propose a RTE solution to iteratively solve the formulation with a fast convergence speed. Then, an approximate closed-form solution can be obtained. This way, the optimal bit allocation can be achieved at little encoding complexity cost. Finally, the experimental results validate the effectiveness of our method in three aspects: compressed distortion, bit-rate control error, and bit fluctuation. Shengxi Li, Mai Xu, Zulin Wang |
ICME | 1 |
| 2015 | Weight-based R-λ rate control for perceptual HEVC coding on conversational videos
Shengxi Li, Mai Xu, Xin Deng 0002, Zulin Wang |
Signal Process. Image Commun. | 1 |
| 2014 | A novel weight-based URQ scheme for perceptual video coding of conversational video in HEVCabstractIn this paper, we propose a novel weight-based unified rate-quantization (URQ) scheme for rate control in state-of-the-art HEVC standard, to improve its perceived visual quality for conversational videos. In conventional rate control of HEVC, a pixel-wise URQ scheme is proposed by introducing the concept of bit per pixel (bpp). This scheme is able to assign different amounts of bits to the blocks with various sizes, thus well suitable for flexible picture partition of HEVC. However, bpp does not reflect the visual importance of each pixel. Therefore, we propose a novel weight-based URQ scheme to take into account the visual importance for rate control in HEVC. In combination with the weight map acquired from a novel hierarchical perceptual model of face, such a scheme is capable of allocating more bits to the face and much more bits to the facial features, by using bit per weight (bpw) instead of bpp. As a result, the visual quality of face, especially facial features, can be improved such that perceptual video coding is achieved for HEVC. Finally, the experimental results validate such improvement. Shengxi Li, Mai Xu, Xin Deng 0002, Zulin Wang |
ICME | 1 |
| 2014 | Complexity control of HEVC based on region-of-interest attention modelabstractIn this paper, we present a novel complexity control method of HEVC to adjust its encoding complexity. First, a region-of-interest (ROI) attention model is established, which defines different weights for various regions according to their importance. Then, the complexity control algorithm is proposed with a distortion-complexity optimization model, to determine the maximum depth of the largest coding units (LCUs) according to their weights. We can reduce the encoding complexity to a given target level at the cost of little distortion loss. Finally, the experimental results show that the encoding complexity can drop to a pre-defined target complexity as low as 20% with bias less than 7%. Meanwhile, our method is verified to preserve the quality of ROI better than another state-of-the-art approach. Xin Deng 0002, Mai Xu, Shengxi Li, Zulin Wang |
VCIP | 3 |
| 2014 | Compressibility Constrained Sparse Representation With Learnt Dictionary for Low Bit-Rate Image CompressionabstractThis paper proposes a compressibility constrained sparse representation (CCSR) approach to low bit-rate image compression using a learnt over-complete dictionary of texture patches. Conventional sparse representation approaches for image compression are based on matching pursuit (MP) algorithms. Actually, the weakness of these approaches is that they are not stable in terms of sparsity of the estimated coefficients, thereby resulting in the inferior performance in low bit-rate image compression. In comparison with MP, convex relaxation approaches are more stable for sparse representation. However, it is intractable to directly apply convex relaxation approaches to image compression, as their coefficients are not always compressible. To utilize convex relaxation in image compression, we first propose in this paper a CCSR formulation, imposing the compressibility constraint on the coefficients of sparse representation for each image patch. In addition, we work out the CCSR formulation to obtain sparse and compressible coefficients, through recursively solving the \(\ell _{1}\) -norm optimization problem of sparse representation. Given these coefficients, each image patch can be represented by the linear combination of texture elements encoded in an over-complete dictionary, learnt from other training images. Finally, low bit-rate image compression can be achieved, owing to the sparsity and compressibility of coefficients by our CCSR approach. The experimental results demonstrate the effectiveness and superiority of the CCSR approach on compressing the natural and remote sensing images at low bit-rates. Mai Xu, Shengxi Li, Jianhua Lu, Wenwu Zhu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |