Lai Jiang 0004

dblp:49/9147-4 · DBLP profile ↗
← Back
44ranked-venue papers
6as first author
34since 2021 · last 2026
0000-0002-4639-8136ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 23 · 6 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021
YearPublicationVenuePosition
2026 Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream Tasks
abstract
In recent years, the development of burst imaging technology has improved the capture and processing capabilities of visual data, enabling a wide range of applications. However, the redundancy in burst images leads to the increased storage and transmission demands, as well as reduced efficiency of downstream tasks. To address this, we propose a new task of Burst Image Quality Assessment (BuIQA), to evaluate the task-driven quality of each frame within a burst sequence, providing reasonable cues for burst image selection. Specifically, we establish the first benchmark dataset for BuIQA, consisting of 7,346 burst sequences with 45,827 images and 191,572 annotated quality scores for multiple downstream scenarios. Inspired by the data analysis, a unified BuIQA framework is proposed to achieve an efficient adaption for BuIQA under diverse downstream scenarios. Specifically, a task-driven prompt generation network is developed with heterogeneous knowledge distillation, to learn the priors of the downstream task. Then, the task-aware quality assessment network is introduced to assess the burst image quality based on the task prompt. Extensive experiments across 10 downstream scenarios demonstrate the impressive BuIQA performance of the proposed approach, outperforming the state-of-the-art. Furthermore, it can achieve 0.33 dB PSNR improvement in the downstream tasks of denoising and super-resolution, by applying our approach to select the high-quality burst frames.
Xiaoye Liang, Lai Jiang 0004, Minglang Qiao, Yue Zhang 0082, Xin Deng 0002, Shengxi Li, Yufan Liu 0001, Mai Xu
AAAI2
2026 UniFES: A Unified Recurrent Network for Quality Enhancement and Stabilization in Face Videos
abstract
Recent years have witnessed an explosive increase of face content, which drives a distinct shift from static images to dynamic video formats. The shift of formats inherently alters the characteristics within face videos, whereby pixel-wise artifacts are intertwined with motion-related impairments. Addressing the emerging distortions that now always appear by twins in practice, however, is challenging and non-trivial, due to the distinct characteristics in addressing spatial-temporal frequencies in videos. In this paper, we propose a novel Unified recurrent network for joint Face video quality Enhancement and Stabilization (UniFES), as the first successful attempt for both quality enhancement and motion stabilization. Correspondingly, our UniFES method proposes to effectively aggregate the mutual information in the pixel and motion domains. For the quality enhancement, our UniFES method decomposes the shaking temporal alignment problem into progressive feature alignment with explicit physical information, which includes the global dynamics from the motion domain, i.e., from the stabilization task. Regarding the video stabilization, we integrate the mixed dynamics from the enhancement task (i.e., from pixel domain) to take into account both pixel-wise and motion-related characteristics, for ensuring robust trajectory estimation and motion stabilization. Subsequently, we refine the warping masks to achieve high-quality full frame rendering. We further establish a synthetic dataset for training and evaluation regarding this emerging task. Comprehensive experiments have illustrated the superior performances of our UniFES method over 32 comparing baselines on both newly established synthetic and real-world datasets.
Mai Xu, Shengxi Li, Lai Jiang 0004
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Compressed image super-resolution based on invertible degradation and restoration
Mai Xu, Lai Jiang 0004, Xin Deng 0002, Yue Zhang 0082, Yufan Liu 0001
Pattern Recognit.3
2026 A novel geometry-aware spatio-temporal network for multi-view video feature learning
Yue Zhang 0082, Mai Xu, Lai Jiang 0004, Xin Deng 0002, Si Liu 0001
Pattern Recognit.3
2026 Multi-modal face anti-spoofing via self-supervised learning
Yufan Liu 0001, Lai Jiang 0004, Shengxi Li, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jinlong Lin
Pattern Recognit. Lett.3
2026 SportSal: Hypernetwork-Based Saliency Prediction for Sports Videos
abstract
Saliency prediction is crucial for improving sports video processing efficiency, thereby providing an enriched viewing experience for a wide-ranging audience. However, there is a long-term absence of well-established eye-tracking dataset and learning-based approach, particularly tailored for sports videos. In this paper, we establish a large-scale eye-tracking dataset dubbed audio-visual sports (AVS). AVS consists of 1,000 high-quality sports videos with eye fixations from 60 participants. Through data analysis on AVS, we observe that human attention patterns exhibit significant variations based on the specific scene context of the sports. Motivated by our observations, we propose a sports-aware saliency prediction approach, named SportSal, which can adaptively predict saliency maps in a hyper manner. Specifically, a hypernetwork is introduced to learn sports-aware priors. Meanwhile, an audio-visual fusion (AVF) block is developed to effectively fuse features from the visual and audio backbones. Given the learned priors and fused audio-visual features, we propose the hyper deformable convolutional (HDC) block and the hyper upsampling (HU) block for dynamic feature extraction and upsampling, respectively. The two blocks are alternatingly connected to adaptively predict saliency maps. Experimental results show that our approach outperforms 21 state-of-the-art saliency prediction approaches over three sports video eye-tracking datasets. Finally, we demonstrate the application of our SportSal approach in perceptual video compression. The dataset and code will be available at https://github.com/WeNsHiJIe-19950103/SportSal.
Mai Xu, Shijie Wen, Lai Jiang 0004, Minglang Qiao, Shengxi Li
IEEE Trans. Circuits Syst. Video Technol.3
2026 Dual-Domain Visual Prompt Learning for Multi-Modal Medical Image Saliency Prediction
abstract
Medical image saliency prediction plays a pivotal role in emulating clinician visual attention to prioritize diagnostically critical regions. Current methods remain constrained by their spatial-domain dependency and limited cross-modality generalizability, neglecting frequency-domain patterns critical for subtle pathology detection while suffering from over-specialization in specific imaging modalities. Therefore, we propose a dual-domain visual prompt network (DVPNet) that integrates cross-modality generalization with spectral pattern awareness. On the one hand, DVPNet establishes a dataset prompt branch that dynamically modulates spatial feature encoding through modality-specific priors, allowing adaptive interpretation of heterogeneous medical imaging domains. On the other hand, a spatial-frequency hybrid prompt module employs learnable wavelet filters to decompose images into multi-scale spectral components, preserving low-frequency anatomical context while enhancing discriminative high-frequency biomarkers that are typically obscured in previous pixel-level analysis. By seamlessly integrating these complementary representations, DVPNet optimally synthesizes spatial and spectral evidence, enabling robust generalization across diverse medical imaging modalities while sustaining computational efficiency. Extensive experimental results on two distinct datasets demonstrate that the proposed method outperforms state-of-the-art approaches, showing superior saliency prediction performance and enhanced generalizability across medical contexts.
Mai Xu, Xiaowan Hu, Lai Jiang 0004
IEEE J. Biomed. Health Informatics5
2026 Few-Shot Pulmonary Vessel Segmentation Based on Tubular-Aware Prompt-Tuning
abstract
Segmentation of the pulmonary vessel from computed tomography (CT) images plays a crucial role in the diagnosis and treatment of various lung diseases. Although deep learning-based approaches have shown remarkable progress in recent years, their performance is often hindered by the lack of high-quality annotated datasets, in which the complex anatomy and morphology of pulmonary vessels make manual annotation challenging, time-consuming, and prone to errors. To address this, we propose PV25, the first dataset that features finely paired annotations of both pulmonary vessels and airways. Moreover, we propose TPNet, a novel tubular-aware prompt-tuning framework for pulmonary vessel segmentation under few-shot training with limited annotations. Specifically, based on an advanced and frozen segmentation backbone, TPNet proposes tunable encoding and decoding networks that learn tubular structures as transfer learning priors, bridging the gap between the source and target pulmonary vessel domains. Specifically, TPNet is built in an encoder-decoder manner, including the fixed segmentation backbone, tunable encoding and decoding networks. In encoding stage, the Morphology-Driven Region Growing (MDRG) module is developed to leverage the tubular connectivity of vessels to guide the network in capturing fine-grained features of pulmonary vessels. In decoding stage, the Cross-Correlation Guidance (CCG) module is introduced to integrate multi-scale correlations between airway and vessel structures in a coarse-to-fine manner. Extensive experiments conducted on multiple datasets demonstrate that TPNet achieves state-of-the-art performance in pulmonary vessel segmentation under limited training data. Besides, TPNet shows strong performance in related tasks such as airway segmentation and artery-vein classification, highlighting its robustness and versatility.
Zijian Gao, Lai Jiang 0004, Sukun Tian, Yuchun Sun, Mai Xu, Liyuan Tao
IEEE Trans. Medical Imaging2
2025 Spherical Manifold Guided Diffusion Model for Panoramic Image Generation
abstract
Panoramic image essentially acts as a pivotal role in emerging virtual reality and augmented reality scenarios; however, the generation of panoramic images are essentially challenging due to the intrinsic spherical geometry and spherical distortions caused by equirectangular projection (ERP). To address this, we start from the very basics of S2manifold inherent to panoramic images, and propose a novel spherical manifold convolution (SMConv) on S2manifold. Based on the SMConv operation, we propose a spherical manifold guided diffusion (SMGD) model for text-conditioned panoramic image generation, which can well accommodate the spherical geometry during generation. We further develop a novel evaluation method by calculating grouped Fréchet inception distance (FID) on cube-map projections, which can well reflect the quality of generated panoramic images, compared to existing methods that randomly crop ERP-distorted content. Experiment results demonstrate that our SMGD model achieves the state-of-the-art generation quality and accuracy, whilst retaining the shortest sampling time in the text-conditioned panoramic image generation task. Codes are publicly available at https://github.com/chronos123/SMGD.
Xiancheng Sun, Mai Xu, Shengxi Li, Senmao Ma, Xin Deng 0002, Lai Jiang 0004
CVPR6
2025 Deformable Spherical Geometry Transformer For Panoramic Semantic Segmentation
abstract
The increasing availability of 360° images has created a demand for effective Panoramic Semantic Segmentation (PASS) to enable comprehensive scene understanding. However, the spherical nature of 360° image introduces significant spatial distortions due to Equirectangular Projection (ERP), making it challenging for traditional 2D methods, which are designed for Euclidean spaces. Existing PASS methods typically mitigate these distortions through developing spherical-to-tangent polyhedron transformations or special-ized convolutional structures. Nevertheless, these approaches still struggle to preserve the spherical geometry and fail to adequately capture the semantic context of 360° images. In this paper, we propose a Deformable Spherical Geometry Transformer (DSGT) network that adapts to spherical distortions through a local-global self-attention mechanism. The local self-attention module captures local semantic information to alleviate distortions, while the global self-attention module integrates spherical geometric priors to enhance predictions. Experimental results on the Stanford2D3D panoramic dataset demonstrate that DSGT outperforms state-of-the-art PASS methods.
Boyang Lan, Li Yang 0014, Mai Xu, Lai Jiang 0004, Yufeng Wang 0004
ICIP4
2025 Quality Control For HEVC: A Deep Reinforcement Learning Approach
abstract
In video coding, large quality fluctuations exist in compressed videos, significantly degrading their quality of experience (QoE). Most works in literature focus on controlling bit-rates, however, paying few attention on reducing the quality fluctuations. In this paper, we propose a novel deep reinforcement learning (DRL) method for quality control in video coding. Specifically, we first propose the formulation of quality control, which targets at both controlling the target quality and reducing fluctuations. Then, we solve the quality control formulation by proposing a DRL method, in which the DRL elements are modeled by considering the features of both current frame and previous encoded frames. Specifically, for the DRL elements, we take the encoding information, content complexity and hidden features of long short-term memory (LSTM) as the state of DRL, and the selection of quantization parameters (QP) as the action of DRL. Subsequently, an algorithm, based on proximal policy optimization, is utilized to update our DRL model for decision-making on the actions of QP selection. In this way, the videos can be compressed under given and constant quality. We implement our DRL-based quality control method on the standard of high efficiency video coding (HEVC) with the HM 16.15 platform, and experimental results show that our method achieves the state-of-the-art performance on both quality control accuracy and fluctuations, in comparison with other quality control baselines.
Mai Xu, Lai Jiang 0004, Shengxi Li, Xin Deng 0002
ICME4
2025 Spherical-Nested Diffusion Model for Panoramic Image Outpainting
abstract
Panoramic image outpainting acts as a pivotal role in immersive content generation, allowing for seamless restoration and completion of panoramic content. Given the fact that the majority of generative outpainting solutions operates on planar images, existing methods for panoramic images address the sphere nature by soft regularisation during the end-to-end learning, which still fails to fully exploit the spherical content. In this paper, we set out the first attempt to impose the sphere nature in the design of diffusion model, such that the panoramic format is intrinsically ensured during the learning procedure, named as spherical-nested diffusion (SpND) model. This is achieved by employing spherical noise in the diffusion process to address the structural prior, together with a newly proposed spherical deformable convolution (SDC) module to intrinsically learn the panoramic knowledge. Upon this, the proposed method is effectively integrated into a pre-trained diffusion model, outperforming existing state-of-the-art methods for panoramic image outpainting. In particular, our SpND method reduces the FID values by more than 50\% against the state-of-the-art PanoDiffusion method. Codes are publicly available at \url{https://github.com/chronos123/SpND}.
Xiancheng Sun, Senmao Ma, Shengxi Li, Mai Xu, Jingyuan Xia, Lai Jiang 0004, Xin Deng 0002
ICML6
2025 Collateral Circulation Guided Multi-Modality Fusion Network for Postoperative Infarct Prediction
Lisong Dai, Heming Dong, Lai Jiang 0004, Mai Xu, Shengxi Li
MICCAI (15)5
2025 DeepSN-Net: Deep Semi-Smooth Newton Driven Network for Blind Image Restoration
abstract
The deep unfolding network represents a promising research avenue in image restoration. However, most current deep unfolding methodologies are anchored in first-order optimization algorithms, which suffer from sluggish convergence speed and unsatisfactory learning efficiency. In this paper, to address this issue, we first formulate an improved second-order semi-smooth Newton (ISN) algorithm, transforming the original nonlinear equations into an optimization problem amenable to network implementation. After that, we propose an innovative network architecture based on the ISN algorithm for blind image restoration, namely DeepSN-Net. To the best of our knowledge, DeepSN-Net is the first successful endeavor to design a second-order deep unfolding network for image restoration, which fills the blank of this area. Furthermore, it offers several distinct advantages: 1) DeepSN-Net provides a unified framework to a variety of image restoration tasks in both synthetic and real-world contexts, without imposing constraints on the degradation conditions. 2) The network architecture is meticulously aligned with the ISN algorithm, ensuring that each module possesses robust physical interpretability. 3) The network exhibits high learning efficiency, superior restoration accuracy and good generalization ability across 11 datasets on three typical restoration tasks. The success of DeepSN-Net on image restoration may ignite many subsequent works centered around the second-order optimization algorithms, which is good for the community.
Xin Deng 0002, Lai Jiang 0004, Jingyuan Xia, Mai Xu
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Hierarchical Semantic Compression for Consistent Image Semantic Restoration
abstract
The emerging semantic compression has been receiving increasing research efforts most recently, capable of achieving high fidelity restoration during compression, even at extremely low bitrates. However, existing semantic compression methods typically combine standard pipelines with either pre-defined or high-dimensional semantics, thus suffering from deficiency in compression. To address this issue, we propose a novel hierarchical semantic compression (HSC) framework that purely operates within intrinsic semantic spaces from generative models, which is able to achieve efficient compression for consistent semantic restoration. More specifically, we first analyse the entropy models for the semantic compression, which motivates us to employ a hierarchical architecture based on a newly developed general inversion encoder. Then, we propose the feature compression network (FCN) and semantic compression network (SCN), such that the middle-level semantic feature and core semantics are hierarchically compressed to restore both accuracy and consistency of image semantics, via an entropy model progressively shared by channel-wise context. Experimental results demonstrate that the proposed HSC framework achieves the state-of-the-art performance on subjective quality and consistency for human vision, together with superior performances on machine vision tasks given compressed bitstreams. This essentially coincides with human visual system in understanding images, thus providing a new framework for future image/video compression paradigms. The source code and trained models are available at https://github.com/bblgbr/HSC-TIP2025.
Shengxi Li, Zifu Zhang, Mai Xu, Lai Jiang 0004, Yufan Liu 0001, Ce Zhu
IEEE Trans. Image Process.4
2025 Spherical Patch Generative Adversarial Net for Unconditional Panoramic Image Generation
abstract
Recent advancements in virtual reality (VR) and augmented reality (AR) have popularised the emerging panoramic content for the immersive visual experience. The difficulty in acquisition and display of 360° format further highlights the necessity of unconditional panoramic image generation. Existing methods essentially generate planar images mapped from panoramic images, and fail to address the deformation and closed-loop characteristics when inverted back to the panoramic images. Thus leading to the generation of pseudo-panoramic content. This paper aims to directly generate spherical content, in a patch-by-patch style; besides computation friendly, this promises the anywhere continuity on the panoramic image and proper accommodation of panoramic deformation. More specifically, we first propose a novel spherical patch convolution (SPConv) that operates on the local spherical patch, which naturally addresses the deformation of panoramic content. We then propose our spherical patch generative adversarial net (SP-GAN) that consists of spherical local embedding (SLE) and spherical content synthesiser (SCS) modules, which seamlessly incorporate our SPConv so as to generate continuous panoramic patches. To the best of our knowledge, the proposed SP-GAN is the first successful attempt to accommodate the spherical distortion for closed-loop panoramic image generation in a patch-by-patch manner. The experimental results, with human-rated evaluations, have verified the consistently superior performances for unconditional panoramic image generation, from the perspectives of generation quality, computational memory, and generalisation to various resolutions. Codes are publicly available at https://github.com/chronos123/SP-GAN.
Mai Xu, Xiancheng Sun, Shengxi Li, Lai Jiang 0004, Jingyuan Xia, Xin Deng 0002
IEEE Trans. Image Process.4
2025 Recruiting Teacher IF Modality for Nephropathy Diagnosis: A Customized Distillation Method With Attention-Based Diffusion Network
abstract
The joint use of multiple modalities for medical image processing has been widely studied in recent years. The fusion of information from different modalities has demonstrated the performance improvement for a lot of medical tasks. For nephropathy diagnosis, immunofluorescence (IF) is one of the most widely-used multi-modality medical images due to its ease of acquisition and the effectiveness for certain nephropathy. However, the existing methods mainly assume different modalities have the equal effect on the diagnosis task, failing to exploit multi-modality knowledge in details. To avoid this disadvantage, this paper proposes a novel customized multi-teacher knowledge distillation framework to transfer knowledge from the trained single-modality teacher networks to a multi-modality student network. Specifically, a new attention-based diffusion network is developed for IF based diagnosis, considering global, local, and modality attention. Besides, a teacher recruitment module and diffusion-aware distillation loss are developed to learn to select the effective teacher networks based on the medical priors of the input IF sequence. The experimental results in the test and external datasets show that the proposed method has a better nephropathy diagnosis performance and generalizability, in comparison with the state-of-the-art methods.
Mai Xu, Lai Jiang 0004, Yibing Fu, Xin Deng 0002, Shengxi Li
IEEE Trans. Medical Imaging3
2024 Saliency Prediction of Sports Videos: A Large-Scale Database and a Self-Adaptive Approach
abstract
Predicting video saliency is crucial for improving sports video processing efficiency, thereby providing an enriched viewing experience for a wide-ranging audience. However, there is a long-term absence of well-established eye-tracking database and learning-based approach, particularly tailored for sports videos. In this paper, we establish a large-scale eye-tracking database dubbed audio-visual sports (AVS). AVS consists of 1,000 high-quality sports videos with eye fixations from 60 participants. Through the data analysis on AVS, we observe that human attention patterns exhibit significant variations based on the specific scene context of the sports. Motivated by this, we propose a sport-aware audiovisual saliency model, which can adaptively learn the scene context in a hyper manner. Specifically, a new audio-visual fusion (AVF) block is developed to effectively fuse features from the visual and audio backbone. After that, a hyper network is introduced to learn sport-aware priors, which are then adopted to guide the self-adaptive saliency predictor for predicting saliency map. Experimental results demonstrate that our approach outperforms other state-of-the-art saliency prediction models over the only two sports video eye-tracking databases.
Minglang Qiao, Mai Xu, Shijie Wen, Lai Jiang 0004, Shengxi Li, Yunjin Chen, Leonid Sigal
ICASSP4
2024 HyperSOR: Context-Aware Graph Hypernetwork for Salient Object Ranking
abstract
Salient object ranking (SOR) aims to segment salient objects in an image and simultaneously predict their saliency rankings, according to the shifted human attention over different objects. The existing SOR approaches mainly focus on object-based attention, e.g., the semantic and appearance of object. However, we find that the scene context plays a vital role in SOR, in which the saliency ranking of the same object varies a lot at different scenes. In this paper, we thus make the first attempt towards explicitly learning scene context for SOR. Specifically, we establish a large-scale SOR dataset of 24,373 images with rich context annotations, i.e., scene graphs, segmentation, and saliency rankings. Inspired by the data analysis on our dataset, we propose a novel graph hypernetwork, named HyperSOR, for context-aware SOR. In HyperSOR, an initial graph module is developed to segment objects and construct an initial graph by considering both geometry and semantic information. Then, a scene graph generation module with multi-path graph attention mechanism is designed to learn semantic relationships among objects based on the initial graph. Finally, a saliency ranking prediction module dynamically adopts the learned scene context through a novel graph hypernetwork, for inferring the saliency rankings. Experimental results show that our HyperSOR can significantly improve the performance of SOR.
Minglang Qiao, Mai Xu, Lai Jiang 0004, Shijie Wen, Yunjin Chen, Leonid Sigal
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Proposal With Alignment: A Bi-Directional Transformer for 360° Video Viewport Proposal
abstract
People normally watch 360 ° videos through a head-mounted display, inside which only the content of viewports can be seen. Therefore, viewport proposal, referring to detecting potential viewport candidates, plays an important role in many 360 ° video processing tasks. In this paper, we advance the viewport proposal by further aligning the predicted viewports across frames for individual subject. This provides a better methodology and a deeper perspective to learn the human perceptual behaviours on 360 ° videos. Specifically, we first analyze three 360 ° video datasets and obtain several findings on human consistency, objectness and motion of viewports. Inspired by these findings, we propose a bi-directional transformer approach, named BiT, for 360 ° video viewport proposal and alignment. Specifically, BiT is composed of a multi-level residual module, a bi-directional encoder-decoder module and a spherical matching module. This way, the viewports can be well proposed and aligned via considering multi-level, bi-directional and non-local information. Moreover, the aligned viewports by BiT are used to refine the viewports and improve viewport proposal accuracy in return. Finally, we validate that our BiT approach is superior on viewport proposal, compared with the state-of-the-art approaches. Besides, the aligned viewports from BiT is verified to be effective in multiple applications, such as saliency prediction, trajectory prediction and perceptual video compression.
Mai Xu, Lai Jiang 0004, Xin Deng 0002, Gaoxing Chen, Leonid Sigal
IEEE Trans. Circuits Syst. Video Technol.3
2023 DINN360: Deformable Invertible Neural Network for Latitude-aware 360° Image Rescaling
abstract
With the rapid development of virtual reality, 360° images have gained increasing popularity. Their wide field of view necessitates high resolution to ensure image quality. This, however, makes it harder to acquire, store and even process such 360° images. To alleviate this issue, we propose the first attempt at 360° image rescaling, which refers to downscaling a 360° image to a visually valid lowresolution (LR) counterpart and then upscaling to a highresolution (HR) 360° image given the LR variant. Specifically, we first analyze two 360° image datasets and observe several findings that characterize how 360° images typically change along their latitudes. Inspired by these findings, we propose a novel deformable invertible neural network (INN), named DINN360, for latitude-aware 360° image rescaling. In DINN360, a deformable INN is designed to downscale the LR image, and project the high-frequency (HF) component to the latent space by adaptively handling various deformations occurring at different latitude regions. Given the downscaled LR image, the high-quality HR image is then reconstructed in a conditional latitude-aware manner by recovering the structure-related HF component from the latent space. Extensive experiments over four public datasets show that our DINN360 method performs considerably better than other state-of-the-art methods for 2 x, 4 x and 8 x 360° image rescaling.
Mai Xu, Lai Jiang 0004, Leonid Sigal, Yunjin Chen
CVPR3
2023 Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo Matching
abstract
Correlation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed model do not work uniformly well across various datasets, greatly limiting their real-world applicability. To tackle this issue, this paper proposes a new perspective to dynamically calculate correlation for robust stereo matching. A novel Uncertainty Guided Adaptive Correlation (UGAC) module is introduced to robustly adapt the same model for different scenarios. Specifically, a variance-based uncertainty estimation is employed to adaptively adjust the sampling area during warping operation. Additionally, we improve the traditional non-parametric warping with learnable parameters, such that the position-specific weights can be learned. We show that by empowering the recurrent network with the UGAC module, stereo matching can be exploited more robustly and effectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance over the ETH3D, KITTI, and Middlebury datasets when employing the same fixed model over these datasets without any retraining procedure. To target real-time applications, we further design a lightweight model based on UGAC, which also outperforms other methods over KITTI benchmarks with only 0.6 M parameters.
Junpeng Jing, Jiankun Li, Pengfei Xiong, Jiangyu Liu, Shuaicheng Liu, Xin Deng 0002, Mai Xu, Lai Jiang 0004, Leonid Sigal
ICCV9
2023 Optimizing DNN based quality assessment metric for image compression: A novel rate control method
abstract
In the existing coding standards, rate control (RC) plays a critical role in optimally allocating bit-rates to each coding unit, for improving rate-distortion performance under the limited bandwidth. However, the existing RC methods are mainly based on traditional distortion metrics, which fail to take the advantage of the emerging DNN based image quality assessment (IQA) metrics. In this paper, we set up the first attempt to achieve IQA score based RC for image compression. Specifically, a novel visualization based score-distortion (VSD) model and ρ-slope model are proposed to explicitly establish the relationship between IQA score and bit-rates. Then, by solving optimal rate-distortion optimization based on the IQA score, we propose a novel RC method for the HEVC standard. The experimental results show that, given the target bit-rates, the proposed RC method can accurately control the bit-rates and generate the compressed images with higher IQA score and better perceptual quality. More importantly, the proposed RC method is evaluated to be effective over two DNN based IQA metrics and four image datasets, exhibiting the potential in practical use. The code is available at https://github.com/Ffangqy/IQA-RC.
Qiuyue Fang, Lai Jiang 0004, Shengxi Li, Mai Xu, Yunjin Chen, Leonid Sigal
ICME3
2023 Recruiting the Best Teacher Modality: A Customized Knowledge Distillation Method for if Based Nephropathy Diagnosis
Lai Jiang 0004, Yibing Fu, Sai Pan, Mai Xu, Xin Deng 0002, Xiangmei Chen
MICCAI (5)2
2023 Learned Structure-Based Hybrid Framework for Martian Image Compression
abstract
Recent landing marches on Mars have enabled the access to Martian surface images, which act as an important vehicle to demystify the evolution and habitability of Mars, in terms of climate, geography, etc. Transmitting Martian images thus calls for efficient compression methods to ensure the high-quality reconstruction from distant communication, in which the research is yet to start. To address this issue, we propose in this letter a learned structure-based hybrid (LSH) framework to compress Martian images. More specifically, we first observe that the structural consistency exists across Martian images, which motivates us to propose a structural compression network (SCN). The aim of SCN is to compactly represent the structural information of Martian images, thus allowing for the compression at extremely low bit-rates. Then, we propose a detail compensation network (DCN) to reconstruct the missing details when we restore from the structural information, which benefits from improved compression efficiency by reduced bit-rates. The experimental results have verified the superior performances of our LSH method on compressing Martian images, against existing state-of-the-art methods.
Shengxi Li, Xiancheng Sun, Mai Xu, Lai Jiang 0004
IEEE Geosci. Remote. Sens. Lett.4
2023 DeepMIH: Deep Invertible Network for Multiple Image Hiding
abstract
Multiple image hiding aims to hide multiple secret images into a single cover image, and then recover all secret images perfectly. Such high-capacity hiding may easily lead to contour shadows or color distortion, which makes multiple image hiding a very challenging task. In this paper, we propose a novel multiple image hiding framework based on invertible neural network, namely DeepMIH. Specifically, we develop an invertible hiding neural network (IHNN) to innovatively model the image concealing and revealing as its forward and backward processes, making them fully coupled and reversible. The IHNN is highly flexible, which can be cascaded as many times as required to achieve the hiding of multiple images. To enhance the invisibility, we design an importance map (IM) module to guide the current image hiding based on the previous image hiding results. In addition, we find that the image hidden in the high-frequency sub-bands tends to achieve better hiding performance, and thus propose a low-frequency wavelet loss to constrain that no secret information is hidden in the low-frequency sub-bands. Experimental results show that our DeepMIH significantly outperforms other state-of-the-art methods, in terms of hiding invisibility, security and recovery accuracy on a variety of datasets.
Zhenyu Guan 0002, Junpeng Jing, Xin Deng 0002, Mai Xu, Lai Jiang 0004, Zhou Zhang 0016
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Does text attract attention on e-commerce images: A novel saliency prediction dataset and method
abstract
E-commerce images are playing a central role in attracting people's attention when retailing and shopping online, and an accurate attention prediction is of significant importance for both customers and retailers, where its research is yet to start. In this paper, we establish the first dataset of saliency e-commerce images (SalECI), which allows for learning to predict saliency on the e-commerce images. We then provide specialized and thorough analysis by high-lighting the distinct features of e-commerce images, e.g., non-locality and correlation to text regions. Correspondingly, taking advantages of the non-local and self-attention mechanisms, we propose a salient SWin-Transformer back-bone, followed by a multi-task learning with saliency and text detection heads, where an information flow mechanism is proposed to further benefit both tasks. Experimental results have verified the state-of-the-art performances of our work in the e-commerce scenario.
Lai Jiang 0004, Shengxi Li, Mai Xu, Se Lei
CVPR1
2022 Viewport-Based CNN: A Multi-Task Approach for Assessing 360° Video Quality
abstract
For 360° video, the existing visual quality assessment (VQA) approaches are designed based on either the whole frames or the cropped patches, ignoring the fact that subjects can only access viewports. When watching 360° video, subjects select viewports through head movement (HM) and then fixate on attractive regions within the viewports through eye movement (EM). Therefore, this paper proposes a two-staged multi-task approach for viewport-based VQA on 360° video. Specifically, we first establish a large-scale VQA dataset of 360° video, called VQA-ODV, which collects the subjective quality scores and the HM and EM data on 600 video sequences. By mining our dataset, we find that the subjective quality of 360° video is related to camera motion, viewport positions and saliency within viewports. Accordingly, we propose a viewport-based convolutional neural network (V-CNN) approach for VQA on 360° video, which has a novel multi-task architecture composed of a viewport proposal network (VP-net) and viewport quality network (VQ-net). The VP-net handles the auxiliary tasks of camera motion detection and viewport proposal, while the VQ-net accomplishes the auxiliary task of viewport saliency prediction and the main task of VQA. The experiments validate that our V-CNN approach significantly advances state-of-the-art VQA performance on 360° video and it is also effective in the three auxiliary tasks.
Mai Xu, Lai Jiang 0004, Chen Li 0049, Zulin Wang, Xiaoming Tao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Joint Learning of Multi-Level Tasks for Diabetic Retinopathy Grading on Low-Resolution Fundus Images
abstract
Diabetic retinopathy (DR) is a leading cause of permanent blindness among the working-age people. Automatic DR grading can help ophthalmologists make timely treatment for patients. However, the existing grading methods are usually trained with high resolution (HR) fundus images, such that the grading performance decreases a lot given low resolution (LR) images, which are common in clinic. In this paper, we mainly focus on DR grading with LR fundus images. According to our analysis on the DR task, we find that: 1) image super-resolution (ISR) can boost the performance of both DR grading and lesion segmentation; 2) the lesion segmentation regions of fundus images are highly consistent with pathological regions for DR grading. Based on our findings, we propose a convolutional neural network (CNN)-based method for joint learning of multi-level tasks for DR grading, called DeepMT-DR, which can simultaneously handle the low-level task of ISR, the mid-level task of lesion segmentation and the high-level task of disease severity classification on LR fundus images. Moreover, a novel task-aware loss is developed to encourage ISR to focus on the pathological regions for its subsequent tasks: lesion segmentation and DR grading. Extensive experimental results show that our DeepMT-DR method significantly outperforms other state-of-the-art methods for DR grading over three datasets. In addition, our method achieves comparable performance in two auxiliary tasks of ISR and lesion segmentation.
Xiaofei Wang 0004, Mai Xu, Jicong Zhang, Lai Jiang 0004, Liu Li 0001, Mengxian He, Ningli Wang, Hanruo Liu, Zulin Wang
IEEE J. Biomed. Health Informatics4
2021 Deep Multi-Task Learning for Diabetic Retinopathy Grading in Fundus Images
abstract
Recent years have witnessed the growing interest in disease severity grading, especially for ocular diseases based on fundus images. The existing grading methods are usually trained with high resolution (HR) images. However, the grading performance decreases a lot given low resolution (LR) images, which are common in practice. In this paper, we mainly focus on diabetic retinopathy (DR) grading with LR fundus images. According to our analysis on the DR task, we find that: 1) image super-resolution (ISR) can boost the performance of DR grading and lesion segmentation; 2) the lesion segmentation regions of fundus images are highly consistent with pathological regions for DR grading. Thus, we propose a deep multi-task learning based DR grading (DeepMT-DR) method for LR fundus images, which simultaneously handles the auxiliary tasks of ISR and lesion segmentation. Specifically, based on our findings, we propose a hierarchical deep learning structure that simultaneously processes the low-level task of ISR, the mid-level task of lesion segmentation and the high-level task of DR grading. Moreover, a novel task-aware loss is developed to encourage ISR to focus on the pathological regions for its subsequent tasks: lesion segmentation and DR grading. Extensive experimental results show that our DeepMT-DR method significantly outperforms other state-of-the-art methods for DR grading over two public datasets. In addition, our method achieves comparable performance in two auxiliary tasks of ISR and lesion segmentation.
Xiaofei Wang 0004, Mai Xu, Jicong Zhang, Lai Jiang 0004, Liu Li 0001
AAAI4
2021 Saliency-Guided Image Translation
abstract
In this paper, we propose a novel task for saliency-guided image translation, with the goal of image-to-image translation conditioned on the user specified saliency map. To address this problem, we develop a novel Generative Adversarial Network (GAN)-based model, called SalG-GAN. Given the original image and target saliency map, SalG-GAN can generate a translated image that satisfies the target saliency map. In SalG-GAN, a disentangled representation framework is proposed to encourage the model to learn diverse translations for the same target saliency condition. A saliency-based attention module is introduced as a special attention mechanism for facilitating the developed structures of saliency-guided generator, saliency cue encoder and saliency-guided global and local discriminators. Furthermore, we build a synthetic dataset and a real-world dataset with labeled visual attention for training and evaluating our SalG-GAN. The experimental results over both datasets verify the effectiveness of our model for saliency-guided image translation.
Lai Jiang 0004, Mai Xu, Xiaofei Wang 0004, Leonid Sigal
CVPR1
2021 DeepVS2.0: A Saliency-Structured Deep Learning Method for Predicting Dynamic Visual Attention
Lai Jiang 0004, Mai Xu, Zulin Wang, Leonid Sigal
Int. J. Comput. Vis.1
2021 Joint Learning of 3D Lesion Segmentation and Classification for Explainable COVID-19 Diagnosis
abstract
Given the outbreak of COVID-19 pandemic and the shortage of medical resource, extensive deep learning models have been proposed for automatic COVID-19 diagnosis, based on 3D computed tomography (CT) scans. However, the existing models independently process the 3D lesion segmentation and disease classification, ignoring the inherent correlation between these two tasks. In this paper, we propose a joint deep learning model of 3D lesion segmentation and classification for diagnosing COVID-19, called DeepSC-COVID, as the first attempt in this direction. Specifically, we establish a large-scale CT database containing 1,805 3D CT scans with fine-grained lesion annotations, and reveal 4 findings about lesion difference between COVID-19 and community acquired pneumonia (CAP). Inspired by our findings, DeepSC-COVID is designed with 3 subnets: a cross-task feature subnet for feature extraction, a 3D lesion subnet for lesion segmentation, and a classification subnet for disease diagnosis. Besides, the task-aware loss is proposed for learning the task interaction across the 3D lesion and classification subnets. Different from all existing models for COVID-19 diagnosis, our model is interpretable with fine-grained 3D lesion distribution. Finally, extensive experimental results show that the joint learning framework in our model significantly improves the performance of 3D lesion segmentation and disease classification in both efficiency and efficacy.
Xiaofei Wang 0004, Lai Jiang 0004, Liu Li 0001, Mai Xu, Xin Deng 0002, Lisong Dai, Tianyi Li 0004, Zulin Wang, Pier Luigi Dragotti
IEEE Trans. Medical Imaging2
2021 Attention-Based Deep Reinforcement Learning for Virtual Cinematography of 360$^{\circ}$ Videos
abstract
Virtual cinematography refers to automatically selecting a natural-looking normal field-of-view (NFOV) from an entire 360$^{\circ}$video. In fact, virtual cinematography can be modeled as a deep reinforcement learning (DRL) problem, in which an agent makes actions related to NFOV selection according to the environment of 360$^{\circ}$video frames. More importantly, we find from our data analysis that the selected NFOVs attract significantly more attention than other regions, i.e., the NFOVs have high saliency. Therefore, in this paper, we propose an attention-based DRL (A-DRL) approach for virtual cinematography in 360$^{\circ}$video. Specifically, we develop a new DRL framework for automatic NFOV selection with the input of both the content, and saliency map of each 360$^{\circ}$frame. Then, we propose a new reward function for the DRL framework in our approach, which considers the saliency values, ground-truth, and smooth transition for NFOV selection. Subsequently, a simplified DenseNet (called Mini-DenseNet) is designed to learn the optimal policy via maximizing the reward. Based on the learned policy, the actions of NFOV can be made in our A-DRL approach for virtual cinematography of 360$^{\circ}$video. Extensive experiments show that our A-DRL approach outperforms other state-of-the-art virtual cinematography methods, over the datasets of Sports-360 video, and Pano2Vid.
Jianyi Wang, Mai Xu, Lai Jiang 0004, Yuhang Song 0001
IEEE Trans. Multim.3
2020 DeepCT: A novel deep complex-valued network with learnable transform for video saliency prediction
Lai Jiang 0004, Mai Xu, Shanyi Zhang, Leonid Sigal
Pattern Recognit.1
2020 A Large-Scale Database and a CNN Model for Attention-Based Glaucoma Detection
abstract
Glaucoma is one of the leading causes of irreversible vision loss. Many approaches have recently been proposed for automatic glaucoma detection based on fundus images. However, none of the existing approaches can efficiently remove high redundancy in fundus images for glaucoma detection, which may reduce the reliability and accuracy of glaucoma detection. To avoid this disadvantage, this paper proposes an attention-based convolutional neural network (CNN) for glaucoma detection, called AG-CNN. Specifically, we first establish a large-scale attention-based glaucoma (LAG) database, which includes 11 760 fundus images labeled as either positive glaucoma (4878) or negative glaucoma (6882). Among the 11 760 fundus images, the attention maps of 5824 images are further obtained from ophthalmologists through a simulated eye-tracking experiment. Then, a new structure of AG-CNN is designed, including an attention prediction subnet, a pathological area localization subnet, and a glaucoma classification subnet. The attention maps are predicted in the attention prediction subnet to highlight the salient regions for glaucoma detection, under a weakly supervised training manner. In contrast to other attention-based CNN methods, the features are also visualized as the localized pathological area, which are further added in our AG-CNN structure to enhance the glaucoma detection performance. Finally, the experiment results from testing over our LAG database and another public glaucoma database show that the proposed AG-CNN approach significantly advances the state-of-the-art in glaucoma detection.
Liu Li 0001, Mai Xu, Hanruo Liu, Yang Li 0010, Xiaofei Wang 0004, Lai Jiang 0004, Zulin Wang, Ningli Wang
IEEE Trans. Medical Imaging6
2019 Image Saliency Prediction in Transformed Domain: A Deep Complex Neural Network Method
abstract
The transformed domain fearures of images show effectiveness in distinguishing salient and non-salient regions. In this paper, we propose a novel deep complex neural network, named SalDCNN, to predict image saliency by learning features in both pixel and transformed domains. Before proposing Sal-DCNN, we analyze the saliency cues encoded in discrete Fourier transform (DFT) domain. Consequently, we have the following findings: 1) the phase spectrum encodes most saliency cues; 2) a certain pattern of the amplitude spectrum is important for saliency prediction; 3) the transformed domain spectrum is robust to noise and down-sampling for saliency prediction. According to these findings, we develop the structure of SalDCNN, including two main stages: the complex dense encoder and three-stream multi-domain decoder. Given the new SalDCNN structure, the saliency maps can be predicted under the supervision of ground-truth fixation maps in both pixel and transformed domains. Finally, the experimental results show that our Sal-DCNN method outperforms other 8 state-of-theart methods for image saliency prediction on 3 databases.
Lai Jiang 0004, Mai Xu, Zulin Wang
AAAI1
2019 Viewport Proposal CNN for 360deg Video Quality Assessment
abstract
Recent years have witnessed the growing interest in visual quality assessment (VQA) for 360° video. Unfortunately, the existing VQA approaches do not consider the facts that: 1) Observers only see viewports of 360° video, rather than patches or whole 360° frames. 2) Within the viewport, only salient regions can be perceived by observers with high resolution. Thus, this paper proposes a viewport-based convolutional neural network (V-CNN) approach for VQA on 360° video, considering both auxiliary tasks of viewport proposal and viewport saliency prediction. Our V-CNN approach is composed of two stages, i.e., viewport proposal and VQA. In the first stage, the viewport proposal network (VP-net) is developed to yield several potential viewports, seen as the first auxiliary task. In the second stage, a viewport quality network (VQ-net) is designed to rate the VQA score for each proposed viewport, in which the saliency map of the viewport is predicted and then utilized in VQA score rating. Consequently, another auxiliary task of viewport saliency prediction can be achieved. More importantly, the main task of VQA on 360° video can be accomplished via integrating the VQA scores of all viewports. The experiments validate the effectiveness of our V-CNN approach in significantly advancing the state-of-the-art performance of VQA on 360° video. In addition, our approach achieves comparable performance in two auxiliary tasks. The code of our V-CNN approach is available at https://github.com/Archer-Tatsu/V-CNN.
Chen Li 0049, Mai Xu, Lai Jiang 0004, Shanyi Zhang, Xiaoming Tao 0001
CVPR3
2019 Attention Based Glaucoma Detection: A Large-Scale Database and CNN Model
abstract
Recently, the attention mechanism has been successfully applied in convolutional neural networks (CNNs), significantly boosting the performance of many computer vision tasks. Unfortunately, few medical image recognition approaches incorporate the attention mechanism in the CNNs. In particular, there exists high redundancy in fundus images for glaucoma detection, such that the attention mechanism has potential in improving the performance of CNN-based glaucoma detection. This paper proposes an attention-based CNN for glaucoma detection (AG-CNN). Specifically, we first establish a large-scale attention based glaucoma (LAG) database, which includes 5,824 fundus images labeled with either positive glaucoma (2,392) or negative glaucoma (3,432). The attention maps of the ophthalmologists are also collected in LAG database through a simulated eye-tracking experiment. Then, a new structure of AG-CNN is designed, including an attention prediction subnet, a pathological area localization subnet and a glaucoma classification subnet. Different from other attention-based CNN methods, the features are also visualized as the localized pathological area, which can advance the performance of glaucoma detection. Finally, the experiment results show that the proposed AG-CNN approach significantly advances state-of-the-art glaucoma detection.
Liu Li 0001, Mai Xu, Xiaofei Wang 0004, Lai Jiang 0004, Hanruo Liu
CVPR4
2018 DeepVS: A Deep Learning Based Video Saliency Prediction Approach
Lai Jiang 0004, Mai Xu, Minglang Qiao, Zulin Wang
ECCV (14)1
2017 Learning to Detect Video Saliency With HEVC Features
abstract
Saliency detection has been widely studied to predict human fixations, with various applications in computer vision and image processing. For saliency detection, we argue in this paper that the state-of-the-art High Efficiency Video Coding (HEVC) standard can be used to generate the useful features in compressed domain. Therefore, this paper proposes to learn the video saliency model, with regard to HEVC features. First, we establish an eye tracking database for video saliency detection, which can be downloaded from https://github.com/remega/video_database. Through the statistical analysis on our eye tracking database, we find out that human fixations tend to fall into the regions with large-valued HEVC features on splitting depth, bit allocation, and motion vector (MV). In addition, three observations are obtained with the further analysis on our eye tracking database. Accordingly, several features in HEVC domain are proposed on the basis of splitting depth, bit allocation, and MV. Next, a kind of support vector machine is learned to integrate those HEVC features together, for video saliency detection. Since almost all video data are stored in the compressed form, our method is able to avoid both the computational cost on decoding and the storage cost on raw data. More importantly, experimental results show that the proposed method is superior to other state-of-the-art saliency detection methods, either in compressed or uncompressed domain.
Mai Xu, Lai Jiang 0004, Xiaoyan Sun 0001, Zhaoting Ye, Zulin Wang
IEEE Trans. Image Process.2
2016 Subjective-quality-optimized complexity control for HEVC decoding
abstract
The latest High Efficiency Video Coding (HEVC) standard significantly improves coding efficiency over H.264/AVC, at the cost of heavy encoding and decoding complexity. For reducing HEVC decoding complexity to a target, we propose in this paper a Subjective-Quality-Optimized Complexity Control (SQOCC) approach, which optimizes subjective quality loss caused by the decoding complexity reduction. First, a saliency detection method in HEVC domain is developed as the preliminary of subjective quality metric. Based on detected saliency, we establish a formulation to minimize subjective quality loss at the constraint of specific decoding complexity reduction, via disabling the deblocking filters of some Largest Coding Units (LCUs). Next, we utilize least square fitting to model functions in our formulation. We then provide a solution to our formulation, achieving subjective-quality-optimized complexity control for HEVC decoding. Finally, the experimental results show the effectiveness of our SQOC-C approach in terms of both control accuracy and subjective quality.
Mai Xu, Lai Jiang 0004, Zulin Wang
ICME3
2016 Bottom-up saliency detection with sparse representation of learnt texture atoms
Mai Xu, Lai Jiang 0004, Zhaoting Ye, Zulin Wang
Pattern Recognit.2
2016 Subjective-Driven Complexity Control Approach for HEVC
abstract
The latest High Efficiency Video Coding (HEVC) standard significantly increases the encoding complexity for improving its coding efficiency, compared with the preceding H.264/Advanced Video Coding (AVC) standard. In this paper, we present a novel subjective-driven complexity control (SCC) approach to reduce and control the encoding complexity of HEVC. Through reasonably adjusting the maximum depth of each largest coding unit (LCU), the encoding complexity can be reduced to a target level with minimal visual distortion. Specifically, the maximum depths of different LCUs can be varied through solving the proposed optimization formulation of complexity control, based on two explored relationships: 1) the relationship between the maximum depth and encoding complexity and 2) the relationship between the maximum depth and visual distortion. Besides, the subjective visual quality is favored with a novel subjective-driven constraint imposed in the formulation, on the basis of a visual attention model. Finally, the experimental results show that our approach can achieve a wide range of encoding complexity control (as low as 20%) for HEVC, with the smallest complexity bias being 0.2%. Meanwhile, our SCC approach outperforms other two state-of-the-art complexity control approaches, in terms of both control accuracy and visual quality.
Xin Deng 0002, Mai Xu, Lai Jiang 0004, Xiaoyan Sun 0001, Zulin Wang
IEEE Trans. Circuits Syst. Video Technol.3