Jianjun Xiang

dblp:272/6124 · DBLP profile ↗
← Back
10ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0002-9025-9520ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning Decoupled Features With Perceptual Distillation for Blind Image Quality Assessment
abstract
Existing Blind Image Quality Assessment (BIQA) approaches typically employ subjective scores as optimization targets to train the model, aiming for results consistent with human judgments. Such judgments are derived from a comprehensive analysis of complex distortions and diverse semantics from images, whereas subjective scores represent the overall quality. This poses a significant challenge for a single model to learn diverse perceptual cues under weak supervision. To address this, we propose a Decoupled Feature Learning (DFL) framework that learns compact global content-aware and local distortion-aware features in a disentangled modeling for BIQA. Our key insight is to leverage global-local input pairs to decompose content-aware and distortion-aware cues entangled in distorted images, and aggregate decoupled perceptual features into a single network. We design a perceptual knowledge distillation strategy that progressively guides the student from fragmented representations to build local-to-global correspondences by distilling self-supervised semantic knowledge, while incorporating the Just-Noticeable-Difference (JND) model to highlight the transfer of perceptually sensitive content features. Finally, we introduce a local distortion-guided attention module to model synergistic effects of different perceptual features from the student for quality evaluation. Extensive experiments on eight benchmark datasets demonstrate the superior performance of the proposed model over the state-of-the-arts. In addition, the DFL framework is flexibly used to improve the perception ability of other Transformer variants. The code is released at https://github.com/JianjunXiang/DFT.
Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Weisi Lin
IEEE Trans. Image Process.1
2024 Semantic-Aware and Quality-Aware Interaction Network for Blind Video Quality Assessment
abstract
Current state-of-the-art video quality assessment (VQA) models typically integrate various perceptual features to comprehensively represent video quality degradation. These models either directly concatenate features or fuse different perceptual scores while ignoring the domain gaps between cross-aware features, thus failing to adequately learn the correlations and interactions between different perceptual features. To this end, we analyze the independent effects and information gaps of quality-and semantic-aware features on video quality. Based on an analysis of the spatial and temporal differences between two aware features, we propose a semantic-Aware and quality-Aware Interaction Network (A2INet) for blind VQA. For spatial gaps, we introduce a cross-aware guided interaction module to enhance the interaction between semantic-and quality-aware features in a local-to-global manner. Considering temporal discrepancies, we design a cross-aware temporal modeling module to further perceive temporal content variation and quality saliency information, and perceptual features are regressed into quality score by a temporal network and a temporal pooling. Extensive experiments on six benchmark VQA datasets show that our model achieves state-of-the-art performance, and ablation studies further validate the effectiveness of each module. We also present a simple video sampling strategy to balance the effectiveness and efficiency of the model. The code for the proposed method will be released at https://github.com/JianjunXiang/A2INet.
Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Ruohong Huan, Nan Gao 0001
ACM Multimedia1
2024 Pseudo Light Field Image and 4D Wavelet-Transform-Based Reduced-Reference Light Field Image Quality Assessment
abstract
Reduced-reference light field image (LFI) quality assessment (RR LFIQA) automatically assesses image quality with only partial information about the reference LFI is available. Existing RR LFIQA has difficulty extracting effective RR information and perceptual features to represent the LFI quality. In this article, we propose an RR LFIQA model based on pseudo LFI (PLFI) and four-dimensional (4D) wavelet transform. To extract RR information related to LFI perceptual quality, a PLFI is created as the RR information of the LFI using a view synthesis algorithm. Considering that the high-dimensional characteristics of the PLFI, 4D wavelet transform is used to decompose the original and distorted PLFIs. The 4D wavelet transform essentially performs a continuous 1D wavelet transform for the 4D signal to enable the local 4D structure of the PLFIs to be characterized effectively in the 4D wavelet domain. A novel spatial-angular weighting strategy is proposed to describe the importance of each location for quality evaluation, to further improve the performance of the proposed method. Experimental results on four benchmark datasets show that the proposed model performs better than the representative 2DIQA and LFIQA models.
Jianjun Xiang, Peng Chen 0008, Yuanjie Dang, Ronghua Liang, Gangyi Jiang
IEEE Trans. Multim.1
2023 STAN: Spatio-Temporal Alignment Network for No-Reference Video Quality Assessment
Zhengyi Yang 0008, Yuanjie Dang, Jianjun Xiang, Peng Chen 0008
ICANN (3)3
2023 Spatial-angular Quality-aware Representation Learning for Blind Light Field Image Quality Assessment
abstract
Blind light field image quality assessment (BLFIQA) remains a challenging task in deep learning due to the unique spatial-angular structure of light field images (LFIs) and the lack of large-scale labeled data for training. In this work, we propose a novel BLFIQA method using spatial-angular quality-aware representation learning in a self-supervised learning manner. Visual content and distortion type are important factors affecting the perceived quality of LFIs. In our observation, the band-pass transform maps of LFIs with the same distortion type exhibit similar Gaussian distributions. Thus, we learn spatial-angular quality-aware representations by minimizing the distance in the embedding space between the luminance map and the band-pass transform map of the same LFI. To implement spatial-angular quality-aware representations of LFI, we also build a large-scale unlabeled dataset containing 40k distorted LFIs with different distortion types and visual content. Further, we propose a fusion-separation-fusion network (FSFNet) to extract features for representing the intrinsic spatial-angular structure of the LFI. After pre-training on the unlabeled dataset using the proposed self-supervised learning, the FSFNet is employed for downstream BLFIQA tasks and achieves good performance. Experimental results show that our proposed method outperforms seventeen state-of-the-art models on the Win5-LID, NBU-LF1.0 and LFDD datasets, and achieves 3.78%, 6.61% and 4.06% SRCC improvements, respectively. The code and dataset will be publicly available in https://github.com/JianjunXiang/SSL_and_FSFNet.
Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Ruohong Huan
ACM Multimedia1
2023 Blind light field image quality assessment with tensor color domain and 3D shearlet transform
Jianjun Xiang, Mei Yu 0001, Gangyi Jiang, Haiyong Xu
Signal Process.1
2023 No-Reference Light Field Image Quality Assessment Using Four-Dimensional Sparse Transform
abstract
Light field imaging can simultaneously capture the intensity and direction information of light rays in the real world. Light field image (LFI) with four-dimensional (4D) data suffers from quality degradation in the process of compression, reconstruction and processing. How to evaluate the visual quality of LFI is thought-provoking. This paper proposes a no-reference LFI quality assessment metric based on high-dimensional sparse transform. Firstly, LFI's sub-aperture gradient image array (SAGIA), which is still a 4D signal, is generated by high-pass filtering between adjacent SAIs. Then, SAGIA is transformed with 4D discrete cosine transform (4D-DCT). 4D-DCT coefficients of SAGIA can characterize the angular and spatial information of LFI. And the logarithmic amplitudes of the coefficients at the same position of SAGIA?s transformed 4D blocks are averaged as the coefficient energy. Subsequently, the 4D-DCT coefficients of SAGIA are divided into the spatial-angular frequency bands and spatial-angular orientation bands, and the corresponding energy features are extracted by converging the coefficient energy of the same band. In addition, the coefficients' amplitudes at the same position of blocks are fitted by the Weibull distribution. Then, the fitted parameters of each position are concatenated, and cropped with principal component analysis to obtain the compact features. Finally, the extracted features are pooled to predict the visual quality of the distorted LFIs. The experimental results demonstrate that the proposed method is more consistent with the subjective evaluation on three LFI databases, compared with the state-of-the-art image quality assessment methods and LFI quality assessment methods.
Jianjun Xiang, Gangyi Jiang, Mei Yu 0001, Zhidi Jiang, Yo-Sung Ho
IEEE Trans. Multim.1
2021 No-reference light field image quality assessment based on depth, structural and angular information
Jianjun Xiang, Gangyi Jiang, Mei Yu 0001, Yongqiang Bai, Zhongjie Zhu
Signal Process.1
2021 Pseudo Video and Refocused Images-Based Blind Light Field Image Quality Assessment
abstract
The commercial light field camera is able to capture four-dimensional Light Field Image (LFI), which can be visualized to LFI contents on 2D displays by means of the Pseudo Video (PV) or the Refocused Images (RIs) generated with the refocusing function of LFI. However, the quality degradation of LFI will affect user’s visual experience of LFI contents. Hence, it is crucial to develop an effective LFI quality assessment method to monitor the LFI quality. Most existing subjective databases of LFI use PV and RIs visualization techniques to assess the quality of LFI. Therefore, as the way of presenting LFI on 2D display, PV and RIs are closely related to the subjective perception of LFI by human eyes. Based on these two visualization techniques, this article proposes a novel PV and RIs based blind LFI quality assessment method, in which the feature extraction is divided into two parts. In the first part, the PV’s structure, motion and disparity information are extracted with multi-scale and multi-directional Shearlet transform. In the other part, the spatial structure, depth and semantic information of the RIs are obtained. Finally, support vector regression is used to nonlinear map the perceptual features to quality score of LFI. The experimental results on four LFI databases show that the proposed method has better correlation with human visual perception, compared with the classical 2D image quality assessment methods as well as the state-of-the-art LFI quality assessment methods.
Jianjun Xiang, Mei Yu 0001, Gangyi Jiang, Haiyong Xu, Yang Song 0015, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.1
2020 VBLFI: Visualization-Based Blind Light Field Image Quality Assessment
abstract
Light field image (LFI) contains the intensity and direction information of the scene. The huge amount of data and different visualization methods of LFI brings great challenges to LFI processing and its blind LFI quality assessment. This paper analyzes the human visual perception from the LFI's visualization, and proposes a novel Visualization-based Blind Light Field Image quality assessment (VBLFI) model. With LFI's visualization and its depth cues, we compute mean difference image from LFI to reduce redundant information of LFI and to describe depth and structural information of LFI. LFI's multi-scale expression with curvelet transform is used to reflect the multi-channel characteristics of human visual system. So, the corresponding natural scene statistical features and energy features are extracted from the mean difference image and sub-aperture images of LFI in curvelet transform domain to form the feature vector, further used to predict the LFI quality. Compared to the representative 2D image quality assessment models and the state-of-the-art LFIQA models, the proposed VBLFI model has better prediction accuracy and stability in the public LFI databases.
Jianjun Xiang, Mei Yu 0001, Hua Chen 0004, Haiyong Xu, Yang Song 0015, Gangyi Jiang
ICME1