Jiebin Yan

dblp:190/4341 · DBLP profile ↗
← Back
55ranked-venue papers
15as first author
49since 2021 · last 2026
0000-0002-0337-6877ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 11 first-author · 40 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-Cancer
abstract
Single-cell RNA sequencing (scRNA-seq) is essential for decoding tumor heterogeneity. However, pan-cancer research still faces two key challenges: learning discriminative and efficient single-cell representations, and establishing a comprehensive evaluation benchmark. In this paper, we introduce \algoname, a lightweight hybrid neural network that combines the strengths of Transformers and state-space models to achieve a balance between performance and efficiency. \algoname consists of a front-end local-context encoder with shared self-attention layers to capture complex, order-independent gene interactions; and a back-end global sequential feature decoder that efficiently integrates global context using a linear-time state-space model. This modular design preserves the expressive power of Transformers while leveraging the scalability of Mamba to enable transcriptome modeling, effectively capturing both local and global regulatory signals. To enable robust evaluation, we also construct a large-scale pan-cancer single-cell benchmark, \algoname Bench, containing over 3.5 million high-quality cells across 33 cancer subtypes, curated through a rigorous preprocessing pipeline. Experimental results show that \algoname outperforms state-of-the-art models on our pan-cancer benchmark (+4.0\%) and across multiple public tasks, including cell type annotation (+7.4\%), batch integration (+4.0\%) and multi-omics integration (+3.1\%).
Xiaoshui Huang, Tianlin Zhu, Yifan Zuo 0001, Xue Xia 0005, Zonghan Wu, Jiebin Yan, Dingli Hua, Zongyi Xu, Yuming Fang 0001, Jian Zhang 0002
AAAI6
2026 Blind Omnidirectional Image Quality Assessment: Embracing the Magic Power of Multimodal Large Language Models
Jiebin Yan, Junjie Chen 0008, Pengfei Chen 0003, Xuelin Liu, Ziwen Tan, Yuming Fang 0001
Int. J. Comput. Vis.1
2026 IC-Bench: Benchmarking robustness of large multimodal models to common corruptions on image captioning
Xuelin Liu, Xinpeng Fang, Jiebin Yan, Chengyang Fang, Yuming Fang 0001
Pattern Recognit.3
2026 Audio-visual saliency prediction based on joint adversarial learning and Co-Attention mechanism
Amin Mao, Jiebin Yan, Yuming Fang 0001
Pattern Recognit.2
2026 RGB-D salient object detection via cross-modal adaptive correlation learning network
Guanqun Ding, Yuming Fang 0001, Jiebin Yan
Signal Process. Image Commun.3
2026 Objective Quality Assessment of AI-Generated Content Videos With Transformation Consistency Focus
abstract
Unnatural motion artifacts—such as implausible object dynamics or discontinuous scene transitions—characterize a critical challenge in AI-generated content (AIGC) videos. Assessing these temporal inconsistencies is essential for benchmarking the performance of text-to-video (T2V) models. Current assessment methods derive motion features from either action recognition models or optical flow estimators. However, these motion features cannot faithfully reflect the human-aligned interpretability of how objects or scenes transform between frames. For instance, a model might detect a “running” action but fail to penalize implausible leg movements. To address this gap, we propose Transformation Consistency-based Video Quality Assessment (TCVQA), a novel framework that quantifies transformation consistency by measuring the recognizability of semantic transformations across frames. The core module of TCVQA is the TC-branch, which includes three core components: A Feature Extractor to capture high-level, fine-grained, and low-level motion features. A Flow-Driven Transformation module Warps extracted features from the source frame to the target frame using predicted optical flow. A Differential Perceiver computes discrepancies between warped source features and actual target features, yielding a consistency score that reflects deviations from natural motion patterns. Besides, the TCVQA also integrates three other branches, the TV-branch, the V-branch, and the F-branch, to perceive multiple aspects of distortions. Extensive experiments on AIGC-VQA benchmarks—including T2VQA-DB, LGVQ, FETV, and MQT demonstrate TCVQA’s superiority, achieving consistent improvement in correlation with human judgments over state-of-the-art methods. Our work establishes transformation consistency as a pivotal axis, enabling more reliable evaluation of AIGC video quality assessment.
Jiebin Yan, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Viewport-Unaware Full-Reference Omnidirectional Image Quality Assessment With Inter-Patch and Sequence Similarity
abstract
Full-reference (FR) image quality assessment (IQA) (FR-IQA) has been extensively explored in the past two decades and is one of the most basic and hot topics in the image processing community, due to its indispensable role in quantitatively describing image quality degradation and guiding algorithm and system optimization. However, FR omnidirectional image quality assessment (OIQA) (FR-OIQA) has achieved less success, due to the natural gap between 2D images and omnidirectional images (OIs). To this end, we present a novel FR-OIQA model with Inter-Patch and Sequence Similarity (IPSS). Specifically, to avoid the extra computational load of viewport generation/prediction methods, IPSS processes OIs in aviewport-unawaremanner,i.e., directly extracting a patch sequence from an OI in the format of Equirectangular Projection (ERP) with retaining regions of interest. Furthermore, since the patches from ERP image contain inborn geometry deformation, thedeformation-awareconvolution is plugged into feature extraction and used to distill quality-aware features from theintrinsic pseudo-degradation, which are then utilized to measure inter-patch similarity. Finally, a distortion-aware interaction module is used to aggregate patch-wise quality-aware features, whose output is used to calculate patch-sequence similarity,i.e., the global quality of OI. Through comprehensive experiments on a large-scale OIQA database, we demonstrate the superiority of the proposed IPSS and the effectiveness of each module.
Jiebin Yan, Junjie Chen 0008, Pengfei Chen 0003, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 RAM-VQA: Restoration Assisted Multi-Modality Video Quality Assessment
abstract
Video Quality Assessment (VQA) strives to computationally emulate human perceptual judgments and has garnered significant attention given its widespread applicability. However, existing methodologies face two primary impediments: (1) limited proficiency in evaluating samples at quality extremes (e.g., severely degraded or near-perfect videos), and (2) insufficient sensitivity to nuanced quality variations arising from a misalignment with human perceptual mechanisms. Although vision-language models offer promising semantic understanding, their reliance on visual encoders pre-trained for high-level tasks often compromises their sensitivity to low-level distortions. To surmount these challenges, we propose the Restoration-Assisted Multi-modality VQA (RAM-VQA) framework. Uniquely, our approach leverages video restoration as a proxy to explicitly model distortion-sensitive features. The framework operates through two synergistic stages: a prompt learning stage that constructs a quality-aware textual space using triple-level references (degraded, restored, and pristine) derived from the restoration process, and a dual-branch evaluation stage that integrates semantic cues with technical quality indicators via spatio-temporal differential analysis. Extensive experiments demonstrate that RAM-VQA achieves state-of-the-art performance across diverse benchmarks, exhibiting superior capability in handling extreme-quality content while ensuring robust generalization.
Pengfei Chen 0003, Jiebin Yan, Rajiv Soundararajan, Giuseppe Valenzise, Leida Li
IEEE Trans. Image Process.2
2026 Exploring Cross-Modal Mutual Prompt Learning for Video Quality Assessment
abstract
Enhancing video quality assessment (VQA) through semantic information integration is a critical research focus. Recent research has employed the Contrastive Language-Image Pre-training (CLIP) model as a foundation to improve semantic perception. However, the image-text alignment inherent in these pre-trained Vision-Language (VL) models frequently results in suboptimal VQA performance. While prompt engineering has recently targeted the language component to address this alignment issue, the unique insights resided in visual analysis is still overlooked for further advancing VQA tasks. Additionally, seeking a trade-off between quality separability and domain invariance in VQA remains largely unresolved within the VL paradigm. In this paper, we introduce a novel cross-modal prompt-based approach to tackle these challenges. Specifically, we propose learnable prompts within the vision branch to foster synergy between visual and language modalities through a language-to-vision coupling function. The multi-view backbone is then carefully crafted with content enhancement and distortion-aware temporal modulation to ensure quality separability. The language prompts, derived from visual representations, are further supported by adaptive weighting mechanisms to optimize the balance between quality separability and domain invariance. Experimental results demonstrate the effectiveness of our proposed method over leading VQA models, showing significant improvements in generalization across diverse datasets. The source code for this work is publicly available athttps://github.com/cpf0079/CM2PL.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Jiebin Yan, Vinit Jakhetiya, Aladine Chetouani
IEEE Trans. Multim.4
2025 Cross-Structure and Semantic Enhancement for Diabetic Retinopathy Grading
abstract
Challenges such as highly variable lesion appearances and complex structural distributions hinder model performance in diabetic retinopathy (DR) grading tasks. To address these issues, we focus on guiding the network toward discriminative feature representation by prioritizing diagnostically relevant information and modeling intricate dependencies between features and DR grades. Specifically, we propose a DR grading network that integrates the Kolmogorov-Arnold Network (KAN) into a Convolution-Vision Transformer (CNN-ViT) cooperative framework, where convolutions and Transformer encoders capture spatial patterns, hierarchical structures, and context, while KAN enhances non-linear semantic dependency modeling. Additionally, we introduce a Cross-Structure (CS) attention module to emphasize relevant features. The proposed modules form the Convolution-Cross-Structure-KAN (CCSK) block, which serves as the backbone of our network, CCSKFormer, enabling more accurate DR grading. The proposed model achieves outstanding performance on two public datasets, with comparisons and ablation studies further validating the effectiveness of the individual modules (https://github.com/xia-xx-cv/CCSKformer).
Xue Xia 0005, Zipeng Lin, Jingying Zhu, Jiebin Yan, Yuming Fang 0001
ICME4
2025 Deep Opinion-Unaware Blind Image Quality Assessment by Learning and Adapting from Multiple Annotators
abstract
Existing deep neural network (DNN)-based blind image quality assessment (BIQA) methods primarily rely on human-rated datasets for training. However, collecting human labels is extremely time-consuming and labor-intensive, posing a significant bottleneck for practical applications. To address this challenge, we propose a Deep opinion-Unaware BIQA model by learning and adapting from Multiple Annotators, termed DUBMA, thereby eliminating the need for human annotations. Specifically, we first generate a large-scale set of distorted image pairs and then assign relative quality rankings using existing full-reference IQA models. The resulting dataset is subsequently employed for training our DUBMA. Due to the inherent discrepancies between synthetic and real-world distortions, a domain shift may occur. To address this, we propose an outlier-robust unsupervised domain adaptation approach leveraging optimal transport. This strategy effectively reduces the gap between synthetic and real-world distortion domains, thereby boosting the model’s adaptability and overall performance. Extensive experiments show that DUBMA outperforms existing opinion-unaware BIQA methods in terms of prediction accuracy across multiple datasets.
Zhihua Wang 0002, Xuelin Liu, Jiebin Yan, Jie Wen 0001, Wei Wang 0169, Chao Huang 0008
IJCAI3
2025 Frequency-Aware Native Resolution Assessment of 8K Omnidirectional Images
abstract
Omnidirectional images (ODIs) serve as fundamental visual medium for presenting virtual reality (VR) contents, supporting fully immersive experiences through 360-degree scene representation. Typically, a high pixel density is essential for visual quality in VR environments, which in turn requires sufficiently high-resolution imagery to achieve. However, capturing native high-resolution ODIs requires expensive omnidirectional cameras with large sensors (e.g., Insta360 TITAN). An alternative approach is to use low-resolution cameras to acquire original images and then enhance their resolution via super-resolution algorithms. In this work, we explore whether super-resolution ODIs can be easily distinguished from native high-resolution ODIs at 8K scale. To this end, we firstly construct the Native Resolution Assessment of 8K Omnidirectional Images (NRA- 8KODI) dataset, whose native 8K ODIs are collected with an Insta360 TITAN camera and 8K super-resolution images are generated from SOTA open-sourced algorithms. Recognizing high-frequency signals are essential for differentiating non-native 8K ODIs, a frequency-aware model is designed to capture high-frequency details. Specially, to maintain high-frequency details kept in high-resolutions while reduce computational costs brought by high-resolutions, we propose a frequency-aware compressor module to suppress feature channels dominated by low-frequency details. Finally, our model achieves 97.2% accuracy in detecting non-native 8K ODIs, implying that super-resolution for ODIs can still be improved for visual experience in VR applications.
Jingwen Hou, Zengliang Li, Jiebin Yan, Weide Liu, Yuming Fang 0001, Wei Zhou 0021
VCIP3
2025 A survey of super-resolution image quality assessment
Qinru Zhu, Jiebin Yan
Neurocomputing5
2025 Opinion-unaware blind quality assessment of AI-generated omnidirectional images based on deep feature statistics
Xuelin Liu, Jiebin Yan, Yuming Fang 0001, Jingwen Hou
J. Vis. Commun. Image Represent.2
2025 Hierarchical boundary feature alignment network for video salient object detection
abstract
The deep learning based video salient object detection (VSOD) models have achieved great success in the past few years, however, these VSOD models still suffer from the following two problems: i) struggle in accurately predicting those pixels surrounding salient objects; ii) unaligned features of different scales lead to deviations in feature fusion . To tackle these problems, we propose a hierarchical boundary feature alignment network (HBFA). Specifically, the proposed HBFA consists of a temporal–spatial fusion module (TSM) and three decoding branches. TSM captures multi-scale spatiotemporal information. The two boundary feature branches are used to guide the whole network to pay more attention to the boundary of salient objects, while the feature alignment branch is capable of fusing the features from the internal and external branches while aligning features across different scales. Our extensive experiments show that the proposed method reaches a new state-of-the-art performance.
Amin Mao, Jiebin Yan, Yuming Fang 0001, Hantao Liu
J. Vis. Commun. Image Represent.2
2025 Opinion-unaware blind stereoscopic image quality assessment: A comprehensive study
Jiebin Yan, Yuming Fang 0001, Xuelin Liu, Wenhui Jiang 0001, Yang Liu 0293
Pattern Recognit.1
2025 Max360IQ: Blind omnidirectional image quality assessment with multi-axis attention
Jiebin Yan, Ziwen Tan, Yuming Fang 0001, Jiale Rao, Yifan Zuo 0001
Pattern Recognit.1
2025 Viewport-Independent Blind Quality Assessment of AI-Generated Omnidirectional Images via Vision-Language Correspondence
abstract
The advancement of deep generation technology has significantly enhanced the growth of artificial intelligencegenerated content (AIGC). Among these, AI-generated omnidirectional images (AGOIs), hold considerable promise for applications in virtual reality (VR). However, the quality of AGOIs varies widely, and there has been limited research focused on their quality assessment. In this letter, inspired by the characteristics of the human visual system, we propose a novel viewportindependent blind quality assessment method for AGOIs, termed VI-AGOIQA, which leverages vision-language correspondence. Specifically, to minimize the computational burden associated with viewport-based prediction methods for omnidirectional image quality assessment, a set of image patches are first extracted from AGOIs in Equirectangular Projection (ERP) format. Then, the correspondence between visual and textual inputs is effectively learned by utilizing the pre-trained image and text encoders of the Contrastive Language-Image Pre-training (CLIP) model. Finally, a multimodal feature fusion module is applied to predict human visual preferences based on the learned knowledge of visual-language consistency. Extensive experiments conducted on publicly available database demonstrate the promising performance of the proposed method. The source code will be made available at https://github.com/LXLHXL123/VI-AGOIQA.
Xuelin Liu, Jiebin Yan, Chenyi Lai, Yuming Fang 0001
IEEE Signal Process. Lett.2
2025 Towards Scalable and Efficient Full-Reference Omnidirectional Image Quality Assessment
abstract
Full-Reference (FR) image quality assessment (IQA) (FR-IQA) has achieved notable success due to its irreplaceable role in algorithm and system optimization; however, it has less been investigated in omnidirectional image quality assessment (OIQA). In this paper, we make an attempt to FR-OIQA considering the constraint of the computation budget, in which this issue is formulated as “quality perception from patch to sequence”,i.e.,Intra-PatchSequence degradation modeling andInter-PatchSequence similarity calculation (denoted by IPS$^{2}$). Specifically, IPS$^{2}$directly accepts local patches from the omnidirectional image (OI) in the format of Equirectangular Projection as input, avoiding other preprocessing operations, such as scan-path prediction and projection transformation. Subsequently, IPS$^{2}$uses a deep feature extractor to capture patch quality and then sends the patch- wise quality maps to the cross-patch similarity (CPS) module, which explicitly models intra-patch sequence degradation and inter-patch sequence similarity via self-attention. Finally, a quality regressor is used to aggregate these features of the CPS module and predict the global quality of the OI. The experimental results on a large-scale OIQA database show that the proposed IPS$^{2}$outperforms most state-of-the-art methods in quality prediction accuracy while offering substantial reductions in computational cost and model size.
Jiebin Yan, Zhihua Wang 0002, Yuming Fang 0001, Hantao Liu
IEEE Signal Process. Lett.1
2025 Multitask Auxiliary Network for Perceptual Quality Assessment of Non-Uniformly Distorted Omnidirectional Images
abstract
Omnidirectional image quality assessment (OIQA) has been widely investigated in the past few years and achieved much success. However, most of existing studies are dedicated to solve the uniform distortion problem in OIQA, which has a natural gap with the non-uniform distortion problem, and their ability in capturing non-uniform distortion is far from satisfactory. To narrow this gap, in this paper, we propose a multitask auxiliary network for non-uniformly distorted omnidirectional images, where the parameters are optimized by jointly training the main task and other auxiliary tasks. The proposed network mainly consists of three parts: a backbone for extracting multiscale features from the viewport sequence, a multitask feature selection module for dynamically allocating specific features to different tasks, and auxiliary sub-networks for guiding the proposed model to capture local distortion and global quality change. Extensive experiments conducted on two large-scale OIQA databases demonstrate that the proposed model outperforms other state-of-the-art OIQA metrics, and these auxiliary sub-networks contribute to improve the performance of the proposed model. The source code is available athttps://github.com/RJL2000/MTAOIQA.
Jiebin Yan, Jiale Rao, Junjie Chen 0008, Ziwen Tan, Weide Liu, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Webly Supervised Fine-Grained Classification by Integrally Tackling Noises and Subtle Differences
abstract
Webly-supervised fine-grained visual classification (WSL-FGVC) aims to learn similar sub-classes from cheap web images, which suffers from two major issues: label noises in web images and subtle differences among fine-grained classes. However, existing methods for WSL-FGVC only focus on suppressing noise at image-level, but neglect to mine cues at pixel-level to distinguish the subtle differences among fine-grained classes. In this paper, we propose a bag-level top-down attention framework, which could tackle label noises and mine subtle cues simultaneously and integrally. Specifically, our method first extracts high-level semantic information from a bag of images belonging to the same class, and then uses the bag-level information to mine discriminative regions in various scales of each image. Besides, we propose to derive attention weights from attention maps to weight the bag-level fusion for a robust supervision. We also propose an attention loss on self-bag attention and cross-bag attention to facilitate the learning of valid attention. Extensive experiments on four WSL-FGVC datasets, i.e., Web-Aircraft, Web-Bird, Web-Car, and WebiNat-5089, demonstrate the effectiveness of our method against the state-of-the-art methods.
Junjie Chen 0008, Jiebin Yan, Yuming Fang 0001, Li Niu 0002
IEEE Trans. Image Process.2
2025 Diffusion-Based Facial Aesthetics Enhancement With 3D Structure Guidance
abstract
Facial Aesthetics Enhancement (FAE) aims to improve facial attractiveness by adjusting the structure and appearance of a facial image while preserving its identity as much as possible. Most existing methods adopted deep feature-based or score-based guidance for generation models to conduct FAE. Although these methods achieved promising results, they potentially produced excessively beautified results with lower identity consistency or insufficiently improved facial attractiveness. To enhance facial aesthetics with less loss of identity, we propose the Nearest Neighbor Structure Guidance based on Diffusion (NNSG-Diffusion), a diffusion-based FAE method that beautifies a 2D facial image with 3D structure guidance. Specifically, we propose to extract FAE guidance from a nearest neighbor reference face. To allow for less change of facial structures in the FAE process, a 3D face model is recovered by referring to both the matched 2D reference face and the 2D input face, so that the depth and contour guidance can be extracted from the 3D face model. Then the depth and contour clues can provide effective guidance to Stable Diffusion with ControlNet for FAE. Extensive experiments demonstrate that our method is superior to previous relevant methods in enhancing facial aesthetics while preserving facial identity.
Lisha Li, Jingwen Hou, Weide Liu, Yuming Fang 0001, Jiebin Yan
IEEE Trans. Image Process.5
2025 Omnidirectional Image Quality Captioning: A Large-Scale Database and a New Model
abstract
The fast growing application of omnidirectional images calls for effective approaches for omnidirectional image quality assessment (OIQA). Existing OIQA methods have been developed and tested on homogeneously distorted omnidirectional images, but it is hard to transfer their success directly to the heterogeneously distorted omnidirectional images. In this paper, we conduct the largest study so far on OIQA, where we establish a large-scale database called OIQ-10K containing 10,000 omnidirectional images with both homogeneous and heterogeneous distortions. A comprehensive psychophysical study is elaborated to collect human opinions for each omnidirectional image, together with the spatial distributions (within local regions or globally) of distortions, and the head and eye movements of the subjects. Furthermore, we propose a novel multitask-derived adaptive feature-tailoring OIQA model named IQCaption360, which is capable of generating a quality caption for an omnidirectional image in a manner of textual template. Extensive experiments demonstrate the effectiveness of IQCaption360, which outperforms state-of-the-art methods by a significant margin on the proposed OIQ-10K database. The OIQ-10K database and the related source codes are available at https://github.com/WenJuing/IQCaption360.
Jiebin Yan, Ziwen Tan, Yuming Fang 0001, Junjie Chen 0008, Wenhui Jiang 0001, Zhou Wang 0001
IEEE Trans. Image Process.1
2025 Learning Guided Implicit Depth Function With Scale-Aware Feature Fusion
abstract
Recently, the single image super-resolution based on implicit image function is a hot topic, which learns a universal model for arbitrary upsampling scales. By contrast, color-guided depth map super-resolution is less explored based on implicit function learning. The related research faces three questions. First, is it also necessary and applicable to fuse the depth feature and the color feature in the encoder with continuous upsampling scales? Second, is the scale information in the encoder as important as that in the decoder? Third, how to efficiently and effectively model the affinity of location distance and content similarity within cross domains in the decoder? This paper proposes a transformer-based network to answer the above questions, which includes a depth super-resolution branch and a guidance extraction branch. Specifically, in the encoder, the effective implicit cross transformer is designed to fuse the guidance from the color feature with continuous coordinate mapping. In addition, the unrelated guidance is filtered out by correlation evaluation in the high-dimension feature space. Unlike the scale only introduced in the decoder, this paper additionally embeds the scale into the position encoding and the feed-forward network in the encoder to learn the scale-aware feature representation. In the decoder, the high-resolution depth feature is reconstructed by using the internal prior and the external guidance. The internal prior is implemented by implicit self-attention in the depth super-resolution branch, and the external guidance is exploited via implicit cross-attention between both branches. Finally, the above decoded features are complementary to generate the high-resolution depth map. The sufficient experiments on the synthetic and real datasets for in-distribution and out-of-distribution upsampling scales validate the improved performance. The code and the models are public via https://github.com/NaNRan13/GIDF.
Yifan Zuo 0001, Yuming Fang 0001, Jiebin Yan, Wenhui Jiang 0001, Yuxin Peng 0001, Yan Huang 0023
IEEE Trans. Image Process.6
2025 Subjective and Objective Quality Assessment of Non-Uniformly Distorted Omnidirectional Images
abstract
Omnidirectional image quality assessment (OIQA) has been one of the hot topics in IQA with the continuous development of VR techniques, and achieved much success in the past few years. However, most studies devote themselves to the uniform distortion issue, i.e., all regions of an omnidirectional image are perturbed by the “same amount” of noise, while ignoring the non-uniform distortion issue, i.e., partial regions undergo “different amount” of perturbation with the other regions in the same omnidirectional image. Additionally, nearly all OIQA models are verified on the platforms containing a limited number of samples, which largely increases the over-fitting risk and therefore impedes the development of OIQA. To alleviate these issues, we elaborately explore this topic from both subjective and objective perspectives. Specifically, we construct a large OIQA database containing 10,320 non-uniformly distorted omnidirectional images, each of which is generated by considering quality impairments on one or two camera len(s). Then we meticulously conduct psychophysical experiments and delve into the influence of both holistic and individual factors (i.e., distortion range and viewing condition) on omnidirectional image quality. Furthermore, we propose a perception-guided OIQA model for non-uniform distortion by adaptively simulating users' viewing behavior. Experimental results demonstrate that the proposed model outperforms state-of-the-art methods.
Jiebin Yan, Jiale Rao, Xuelin Liu, Yuming Fang 0001, Yifan Zuo 0001, Weide Liu
IEEE Trans. Multim.1
2025 Computational Analysis of Degradation Modeling in Blind Panoramic Image Quality Assessment
abstract
Blind panoramic image quality assessment (BPIQA) has recently brought a new challenge to the visual quality community, due to the complex interaction between immersive content and human behavior. Although many efforts have been made to advance BPIQA from both conducting psychophysical experiments and designing performance-driven objective algorithms, limited content and few samples in those closed sets inevitably would result in shaky conclusions, thereby hindering the development of BPIQA; we refer to it as the easy-database issue. In this article, we present a sufficient computational analysis of degradation modeling in BPIQA to thoroughly explore the easy-database issue , where we carefully design three types of experiments via investigating the gap between BPIQA and blind image quality assessment (BIQA), the necessity of specific design in BPIQA models, and the generalization ability of BPIQA models. From extensive experiments, we find that easy databases narrow the gap between the performance of BPIQA and BIQA models, which is unconducive to the development of BPIQA. And the easy databases make the BPIQA models be closed to saturation; therefore, the effectiveness of the associated specific designs cannot be well verified. Besides, the BPIQA models trained on our recently proposed databases with complicated degradation show better generalization ability. Thus, we believe that much more efforts are highly desired to put into BPIQA from both subjective viewpoint and objective viewpoint.
Jiebin Yan, Ziwen Tan, Jiale Rao, Yifan Zuo 0001, Yuming Fang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Viewport-Unaware Blind Omnidirectional Image Quality Assessment: A Flexible and Effective Paradigm
abstract
Most of the existing blind omnidirectional image quality assessment (BOIQA) models rely on viewport generation by modeling user viewing behavior or transforming omnidirectional images (OIs) into varying formats; however, these methods are either computationally expensive or less scalable. To solve these issues, in this article, we present a flexible and effective paradigm, which is viewport-unaware and can be easily adapted to 2D plane image quality assessment (2D-IQA). Specifically, the proposed BOIQA model includes an adaptive prior-equator sampling module for extracting a patch sequence from the equirectangular projection (ERP) image in a resolution-agnostic manner, a progressive deformation-unaware feature fusion module which is able to capture patch-wise quality degradation in a deformation-immune way, and a local-to-global quality aggregation module to adaptively map local perception to global quality. Extensive experiments across four OIQA databases (including uniformly distorted OIs and non-uniformly distorted OIs) demonstrate that the proposed model achieves competitive performance with low complexity against other state-of-the-art models, and we also verify its adaptive capacity to 2D-IQA. The source code is available at https://github.com/KangchengWu/OIQA .
Jiebin Yan, Kangcheng Wu, Junjie Chen 0008, Ziwen Tan, Yuming Fang 0001, Weide Liu
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Meta-Point Learning and Refining for Category-Agnostic Pose Estimation
abstract
Category-agnostic pose estimation (CAPE) aims to predict keypoints for arbitrary classes given a few support images annotated with keypoints. Existing methods only rely on the features extracted at support keypoints to predict or refine the keypoints on query image, but a few support feature vectors are local and inadequate for CAPE. Considering that human can quickly perceive potential keypoints of arbitrary objects, we propose a novel framework for CAPE based on such potential keypoints (named as meta-points). Specifically, we maintain learnable embeddings to capture inherent information of various keypoints, which interact with image feature maps to produce meta-points without any support. The produced meta-points could serve as meaningful potential keypoints for CAPE. Due to the inevitable gap between inherency and annotation, we finally utilize the identities and details offered by support key-points to assign and refine meta-points to desired keypoints in query image. In addition, we propose a progressive deformable point decoder and a slacked regression loss for better prediction and supervision. Our novel framework not only reveals the inherency of key points but also outperforms existing methods of CAPE. Comprehensive experiments and in-depth studies on large-scale MP-100 dataset demon-strate the effectiveness of our framework. Code is avaiable at https://github.com/chenbys/MetaPoint
Junjie Chen 0008, Jiebin Yan, Yuming Fang 0001, Li Niu 0002
CVPR2
2024 Quality of Experience of Viewport Adaptive Omnidirectional Video Streaming
abstract
With the explosive growth of multimedia streaming services and virtual reality devices, omnidirectional video (ODV) is becoming increasingly popular in practical applications. However, streaming the entire ODV with high definition and high frame rate induces a waste of bandwidth. The tile-based viewport adaptive streaming provides a solution to overcome volatile network conditions, while the scheme would lead to quality adaptation when the network changes dynamically. In this paper, we focus on investigating how the human visual quality of experience (QoE) changes with time-varying ODV quality. Specifically, we construct a new quality of experience database for viewport adaptive ODV streaming named JUFEOVQoE, which includes twelve original ODVs with diverse content, and corresponding 378 viewport videos generated by compressing the raw viewport videos using a variety of combinations of quantization parameter (QP), spatial (S), and temporal resolutions (T). We conduct a series of subjective experiments to collect the mean opinion scores of the viewport video sequences and the viewing direction data of the subjects. Furthermore, we test several state-of-the-art objective QoE models on the proposed database. Experimental results demonstrate that existing mainstream QoE methods cannot predict the QoE of the viewport adaptive streaming ODVs accurately. The database will be released to facilitate further research.
Xuelin Liu, Haoyun Zhang, Jiebin Yan, Yuming Fang 0001, Shiqi Wang 0001
ICIP3
2024 Blind Quality Assessment of Panoramic Images Based on Multiple Viewport Sequences
abstract
With the development of virtual reality (VR) technology, panoramic image (PI), which is an important digital form of immersive multimedia, has drawn much attention from researchers. However, distortions are inevitably introduced in the process of processing, encoding and compression, which damages their quality and affects the user’s experience. Therefore, assessing the quality of panoramic images is urgent. In this paper, with the consideration of viewing behavior, we propose a novel blind panoramic image quality assessment model, which consists of three parts, viewport generation, feature extraction and quality prediction. Specifically, inspired by the viewing process of PI, we first generate multiple viewport sequences according to the real viewing trajectory and then extract multilevel features with a pre-trained backbone. The concatenated features are taken as the input of a recurrent neural network to evaluate the perceptual quality of PI. To validate the effectiveness of the proposed method, objective experiments are conducted on the public subjective panoramic image quality database. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods.
Xuelin Liu, Jiebin Yan, Yuming Fang 0001, Hantao Liu
ISCAS2
2024 A foreground-context dual-guided network for light-field salient object detection
abstract
Light-field salient object detection (SOD) has become an emerging trend as it records comprehensive information about natural scenes that can benefit salient object detection in various ways. However, salient object detection models with light-field data as input have not been thoroughly explored. The existing methods cannot effectively suppress the noise, and it is difficult to distinguish the foreground and background under challenging conditions including self-similarity, complex backgrounds, large depth of field, and non-Lambertian scenarios. In order to extract the feature of light-field images effectively and suppress the noise in light-field, in this paper, we propose a foreground and context dual guided network. Specifically, we design a global context extraction module (GCEM) and a local foreground extraction module (LFEM). GCEM is used to suppress global noise and roughly predict saliency maps. GCEM also can extract global context information from deep-level features to guide decoding process. By extracting local information from shallow-level, LFEM refines the prediction obtained by GCEM. In addition, we use RGB images to enhance the light-field images before the input GCEM. Experimental results show that our proposed method is effective in suppressing global noise and achieves better results when dealing with transparent objects and complex backgrounds. The experimental results show that the proposed method outperforms several other state-of-the-art methods on three light-field datasets.
Xin Zheng 0006, Deyang Liu, Chengtao Lv, Jiebin Yan
Signal Process. Image Commun.5
2024 Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling
abstract
With the rapid development of multimedia processing and deep learning technologies, especially in the field of video understanding, video quality assessment (VQA) has achieved significant progress. Although researchers have moved from designing efficient video quality mapping models to various research directions, in-depth exploration of the effectiveness-efficiency trade-offs of spatio-temporal modeling in VQA models is still less sufficient. Considering the fact that videos have highly redundant information, this paper investigates this problem from the perspective of joint spatial and temporal sampling, aiming to seek the answer to how little information we should keep at least when feeding videos into the VQA models while with acceptable performance sacrifice. To this end, we drastically sample the video’s information from both spatial and temporal dimensions, and the heavily squeezed video is then fed into a stable VQA model. Comprehensive experiments regarding joint spatial and temporal sampling are conducted on six public video quality databases, and the results demonstrate the acceptable performance of the VQA model when throwing away most of the video information. Furthermore, with the proposed joint spatial and temporal sampling strategy, we make an initial attempt to design an online VQA model, which is instantiated by as simple as possible a spatial feature extractor, a temporal feature fusion module, and a global quality regression module. Through quantitative and qualitative experiments, we verify the feasibility of online VQA model by simplifying itself and reducing input.
Jiebin Yan, Yuming Fang 0001, Xuelin Liu, Xue Xia 0005, Weide Liu
IEEE Trans. Circuits Syst. Video Technol.1
2024 A2 GSTran: Depth Map Super-Resolution via Asymmetric Attention With Guidance Selection
abstract
Currently, Convolutional Neural Network (CNN) has dominated guided depth map super-resolution (SR). However, the inefficient receptive field growing and input-independent convolution limit the generalization of CNN. Motivated by vision transformer, this paper proposes an efficient transformer-based backbone A2GSTran for guided depth map SR, which resolves the above intrinsic defect of CNN. In addition, state-of-the-art (SOTA) models only refine depth features with the guidance which is implicitly selected without supervision. So, there is no explicit guarantee to mitigate the artifacts of texture copying and edge blurring. Accordingly, the proposed A2GSTran simultaneously solves two sub-problems,i.e., guided monocular depth estimation and guided depth SR, in separate branches. Specifically, the explicit supervision upon monocular depth estimation lifts the efficiency of guidance selection. The feature fusion between branches is designed via bi-directional cross attention. Moreover, since guidance domain is defined in high resolution (HR), we propose asymmetric cross attention to maintain the guidance information via pixel unshuffle instead of pooling which has unequal channel number to depth features. Based on the supervisions to depth reconstruction and guidance selection, the final depth features are refined by fusing the output features of the corresponding branches via channel attention to generate the HR depth map. Sufficient experimental results on synthetic and real datasets for multiple scales validate our contributions compared with SOTA models. The code and models are public via https://github.com/alex-cate/Depth_Map_Super-resolution_via_Asymmetric_Attention_with_Guidance_Selection
Yifan Zuo 0001, Yifeng Zeng, Yuming Fang 0001, Xiaoshui Huang, Jiebin Yan
IEEE Trans. Circuits Syst. Video Technol.6
2024 Saliency Guided Deep Neural Network for Color Transfer With Light Optimization
abstract
Color transfer aims to change the color information of the target image according to the reference one. Many studies propose color transfer methods by analysis of color distribution and semantic relevance, which do not take the perceptual characteristics for visual quality into consideration. In this study, we propose a novel color transfer method based on the saliency information with brightness optimization. First, a saliency detection module is designed to separate the foreground regions from the background regions for images. Then a dual-branch module is introduced to implement color transfer for images. Finally, a brightness optimization operation is designed during the fusion of foreground and background regions for color transfer. Experimental results show that the proposed method can implement the color transfer for images while keeping the color consistency well. Compared with other existing studies, the proposed method can obtain significant performance improvement. The source code and pre-trained models are available at https://github.com/PlanktonQAQ/SCTNet.
Yuming Fang 0001, Pengwei Yuan, Chenlei Lv, Jiebin Yan, Weisi Lin
IEEE Trans. Image Process.5
2024 Perceptual Quality Assessment of Omnidirectional Images: A Benchmark and Computational Model
abstract
Compared with traditional 2D images, omnidirectional images (also referred to as 360 ∘ images) have more complicated perceptual characteristics due to the particularities of imaging and display. How humans perceive omnidirectional images in an immersive environment and form the immersive quality of experience are important problems. Thus, it is crucial to measure the quality of omnidirectional images under different viewing conditions, which suffer from realistic distortions. In this article, we build a large-scale subjective assessment database for omnidirectional images and carry out a comprehensive psychophysical experiment to study the relationships between different factors (viewing conditions and viewing behaviors) and the perceptual quality of omnidirectional images. In addition, we collect both subjective ratings and head movement data. A thorough analysis of the collected subjective data is also provided, where we make several interesting findings. Moreover, with the proposed database, we propose a novel transformer-based omnidirectional image quality assessment model. To be consistent with the human viewing process, viewing conditions and behaviors are naturally incorporated into the proposed model. Specifically, the proposed model mainly consists of three parts: viewport sequence generation, multi-scale feature extraction, and perceptual quality prediction. Extensive experimental results conducted on the proposed database demonstrate the effectiveness of the proposed method over existing image quality assessment methods.
Xuelin Liu, Jiebin Yan, Yuming Fang 0001, Yang Liu 0293
ACM Trans. Multim. Comput. Commun. Appl.2
2023 UDNet: Uncertainty-aware deep network for salient object detection
Yuming Fang 0001, Jiebin Yan, Wenhui Jiang 0001, Yang Liu 0293
Pattern Recognit.3
2023 Study of Spatio-Temporal Modeling in Video Quality Assessment
abstract
Video quality assessment (VQA) has received remarkable attention recently. Most of the popular VQA models employ recurrent neural networks (RNNs) to capture the temporal quality variation of videos. However, each long-term video sequence is commonly labeled with a single quality score, with which RNNs might not be able to learn long-term quality variation well: What's the real role of RNNs in learning the visual quality of videos? Does it learn spatio-temporal representation as expected or just aggregating spatial features redundantly? In this study, we conduct a comprehensive study by training a family of VQA models with carefully designed frame sampling strategies and spatio-temporal fusion methods. Our extensive experiments on four publicly available in- the-wild video quality datasets lead to two main findings. First, the plausible spatio-temporal modeling module (i. e., RNNs) does not facilitate quality-aware spatio-temporal feature learning. Second, sparsely sampled video frames are capable of obtaining the competitive performance against using all video frames as the input. In other words, spatial features play a vital role in capturing video quality variation for VQA. To our best knowledge, this is the first work to explore the issue of spatio-temporal modeling in VQA.
Yuming Fang 0001, Zhaoqian Li, Jiebin Yan, Xiangjie Sui, Hantao Liu
IEEE Trans. Image Process.3
2022 Perceptual Quality Assessment of Omnidirectional Images
abstract
Omnidirectional images, also called 360◦images, have attracted extensive attention in recent years, due to the rapid development of virtual reality (VR) technologies. During omnidirectional image processing including capture, transmission, consumption, and so on, measuring the perceptual quality of omnidirectional images is highly desired, since it plays a great role in guaranteeing the immersive quality of experience (IQoE). In this paper, we conduct a comprehensive study on the perceptual quality of omnidirectional images from both subjective and objective perspectives. Specifically, we construct the largest so far subjective omnidirectional image quality database, where we consider several key influential elements, i.e., realistic non-uniform distortion, viewing condition, and viewing behavior, from the user view. In addition to subjective quality scores, we also record head and eye movement data. Besides, we make the first attempt by using the proposed database to train a convolutional neural network (CNN) for blind omnidirectional image quality assessment. To be consistent with the human viewing behavior in the VR device, we extract viewports from each omnidirectional image and incorporate the user viewing conditions naturally in the proposed model. The proposed model is composed of two parts, including a multi-scale CNN-based feature extraction module and a perceptual quality prediction module. The feature extraction module is used to incorporate the multi-scale features, and the perceptual quality prediction module is designed to regress them to perceived quality scores. The experimental results on our database verify that the proposed model achieves the competing performance compared with the state-of-the-art methods.
Yuming Fang 0001, Jiebin Yan, Xuelin Liu, Yang Liu 0293
AAAI3
2022 Benchmarking 360° Saliency Models by General-Purpose Metrics
abstract
How to effectively evaluate a model's capability to predict the visual attention of observers in 360° scenes gains interest along with the advancement of saliency prediction modeling of omnidirectional images (ODIs). So far, many general-purpose metrics from 2D saliency literature have been adopted to evaluate the 360° saliency models. However, whether they are still effective when being adopted to evaluate the 360° saliency models has not been explored. In this paper, we testify several standard saliency evaluation metrics on the 360° saliency models and comprehensively analyze their behaviors in the omnidirectional scenario. We find that 1) most metrics under-penalize false positives; 2) existing 360° datasets involve severe equator bias that few metrics can effectively penalize. We hope this case study can provide a guideline for benchmarking 360° image/video saliency models.
Xiangjie Sui, Jiebin Yan, Yuming Fang 0001
MMSP3
2022 Bilinear CNNs for Blind Quality Assessment of Fine-Grained Images
abstract
Most of the existing image quality assessment (IQA) studies focus on discriminable images, whose relative visual quality could be easily determined by human beings (we also call this issue coarse-grained (CG)-IQA). The effective models designed for CG-IQA struggle for quality assessment of the images with subtle differences (often exist in many real applications), which is also called fine-grained (FG) IQA problem. Thus, we make the first, to the best of our knowledge, attempt to build a novel blind IQA (BIQA) model for the images with FG distortion, aiming to fill the gap between objective IQA model and real applications. Specifically, the proposed model mainly consists of a feature extraction module (a sequence of convolution layers), a squeeze-and-excitation module, and a bilinear pooling module, whose objectives are extracting quality-aware features, enhancing features' representation ability, and discriminability. We conduct extensive experiments on a public FG-IQA database, and demonstrate the superiority of the proposed method and the effectiveness of each module.
Jiebin Yan, Yuming Fang 0001, Wenhui Jiang 0001
MMSP2
2022 Evaluating the Robustness of Depth Image Super-Resolution Models
abstract
Depth image super-resolution (DISR) is one of the hot topics in computer vision. Although great progress has been made in this research topic, the robustness of DISR models is not sufficiently investigated, which is of great importance in the real applications. Accordingly, in this paper, we make an initial attempt to investigate the robustness of DISR models. Specifically, we test their generalization ability when the input depth image suffers from visual quality degradation. To facilitate this study, we construct a large-scale depth image dataset in which the reference depth images are perturbed to generate the degraded depth images automatically. Then, we test six top-performing DISR models on the constructed dataset and then compare their strengths and weaknesses. By conducting comprehensive experiments, we find that depth image super-resolution models perform poorly on Gaussian noise, and that the higher the level, the lower the quality of the predicted depth map. Furthermore, some DISR models only outperform at lower magnifications (such as 2x and 4x).
Dengxiang Wang, Jiebin Yan, Xuelin Liu, Yifan Zuo 0001
MMSP2
2022 Subjective and Objective Quality of Experience of Free Viewpoint Videos
abstract
Free viewpoint videos (FVVs) provide immersive experiences for end-users, and they have been applied in many applications, such as movies, sports, and TV shows. However, the development of quantifying the quality of experience (QoE) of FVVs is still relatively slow due to the high costs of data collection and limited public databases. In this paper, we conduct a comprehensive study on FVV QoE. First, we construct the largest, to the best of our knowledge, FVV QoE database called Youku-FVV from two complex real scenarios, i. e., entertainment and sports. Specifically, Youku-FVV originates from the videos captured by dozens of real cameras arranged annularly. We use these videos to generate virtual viewpoints, which make up FVVs together with real views. In constructing the FVV QoE database, we consider both internal and external influencing factors of QoE, which correspond to FVV generation and playback, respectively. Besides, we make an initial attempt to train an efficient no reference FVV QoE prediction model using this database, where several sparse frame sampling strategies are validated. And we demonstrate the feasibility of striving for the balance between effectiveness and efficiency of FVV QoE prediction. The proposed FVV QoE database and source codes are publicly available at https://github.com/QTJiebin/FVV_QoE.
Jiebin Yan, Jing Li 0026, Yuming Fang 0001, Zhaohui Che, Xue Xia 0005, Yang Liu 0293
IEEE Trans. Image Process.1
2021 Exposing Semantic Segmentation Failures via Maximum Discrepancy Competition
Jiebin Yan, Yuming Fang 0001, Zhangyang Wang, Kede Ma
Int. J. Comput. Vis.1
2021 Asymmetrically distorted 3D video quality assessment: From the motion variation to perceived quality
Yuming Fang 0001, Xiangjie Sui, Jiebin Yan, Yifan Zuo 0001, Jiheng Wang, Zhaoqian Li
Signal Process.3
2021 Visual attention prediction for Autism Spectrum Disorder with hierarchical semantic fusion
Yuming Fang 0001, Yifan Zuo 0001, Wenhui Jiang 0001, Hanqin Huang, Jiebin Yan
Signal Process. Image Commun.6
2021 Objective quality assessment of synthesized images by local variation measurement
Xiangjie Sui, Mengna Ding, Jiebin Yan, Yuming Fang 0001, Yifan Zuo 0001, Zuowen Tan
Signal Process. Image Commun.3
2021 Perceptual Quality Assessment for Asymmetrically Distorted Stereoscopic Video by Temporal Binocular Rivalry
abstract
In this paper, we propose a two-stage weighting based perceptual quality assessment framework for asymmetrically distorted stereoscopic video (SV) sequences by temporal binocular rivalry. Firstly, a traditional 2D image quality assessment (IQA) method is employed to measure spatial distortion, and the temporal distortion is evaluated by the magnitude differences between motion vectors of distorted and reference video frames. Secondly, the structural strength (SS) computed by gradient map and the motion energy (ME) computed by frame difference map are used to estimate the intensity of visual stimulus in spatial and temporal domain respectively. Then, SS and ME are considered as the importance indexes to combine the quality scores of spatial and temporal distortion to estimate perceived distortion of single-view video sequences, which is denoted as the first-stage weighting. Finally, considering that the difference of intensity of visual stimulus between two eyes results in binocular rivalry, a novel temporal binocular rivalry inspired weighting method is designed to integrate the quality scores of left- and right-views for the final visual quality prediction of SV sequences, which is denoted as the second-stage weighting. Experimental results on Waterloo-IVC SV quality databases show that several specific examples of 2D-IQA methods within the proposed framework can obtain highly competitive performance over other existing ones.
Yuming Fang 0001, Xiangjie Sui, Jiheng Wang, Jiebin Yan, Jianjun Lei 0001, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.4
2021 Superpixel-Based Quality Assessment of Multi-Exposure Image Fusion for Both Static and Dynamic Scenes
abstract
Multi-exposure image fusion (MEF) algorithms have been used to merge a stack of low dynamic range images with various exposure levels into a well-perceived image. However, little work has been dedicated to predicting the visual quality of fused images. In this work, we propose a novel and efficient objective image quality assessment (IQA) model for MEF images of both static and dynamic scenes based on superpixels and an information theory adaptive pooling strategy. First, with the help of superpixels, we divide fused images into large- and small-changed regions using the structural inconsistency map between each exposure and fused images. Then, we compute the quality maps based on the Laplacian pyramid for large- and small-changed regions separately. Finally, an information theory induced adaptive pooling strategy is proposed to compute the perceptual quality of the fused image. Experimental results on three public databases of MEF images demonstrate the proposed model achieves promising performance and yields a relatively low computational complexity. Additionally, we also demonstrate the potential application for parameter tuning of MEF algorithms.
Yuming Fang 0001, Yan Zeng 0001, Wenhui Jiang 0001, Hanwei Zhu, Jiebin Yan
IEEE Trans. Image Process.5
2021 Blind Quality Assessment for Tone-Mapped Images by Analysis of Gradient and Chromatic Statistics
abstract
A tone-mapped image (TMI) obtained from the corresponding high dynamic range (HDR) image induces artifacts and distortion, which might result in the loss of structure information and impaired color. By analyzing the visual characteristics of TMIs, this work proposes a robust blind visual quality evaluation method for TMIs by using gradient and chromatic statistics (VQGC). First, motivated by the perceptual mechanism that the human visual system (HVS) is sensitive to image structure variation, we employ the gradient features to measure structure degradation in TMIs. To predict structure distortion accurately, we compute the gradient magnitude and orientation to measure image structure variation, and the relative gradient magnitude and orientation are also computed to capture microstructure change. Second, the color invariance descriptors are utilized to capture the visual degradation of colorfulness by local binary pattern (LBP) on four chromatic feature maps. Finally, the gradient and chromatic features are combined together as the final quality-aware feature vector, which is applied to assess the perceptual quality of TMIs by support vector regression (SVR). Comparison experiments show that the performance of the proposed method is better than other existing blind quality assessment methods on public databases.
Yuming Fang 0001, Jiebin Yan, Rengang Du, Yifan Zuo 0001, Wenying Wen, Yan Zeng 0001, Leida Li
IEEE Trans. Multim.2
2020 Blind Stereoscopic Image Quality Assessment By Deep Neural Network Of Multi-Level Feature Fusion
abstract
In this paper, we propose an effective blind image quality assessment (BIQA) method for stereoscopic images by deep neural network (DNN) of multi-level feature fusion (MLFF) inspired by the multi-scale characteristics and binocular properties of the human visual system (HVS). Specifically, we firstly feed the left- and right-view images into a weight sharing convolutional neural network (CNN) for jointly feature extraction. To aggregate multi-level features, we concatenate the low-, middle-, and high-level feature maps of stereoscopic images to simulate the complicated visual interaction processing in the HVS. Two fully connected layers are used to build the nonlinear mapping from the highly abstract features to the quality scores of stereoscopic images. The experiments conducted on two public databases prove the validity of the proposed MLFF method.
Jiebin Yan, Yuming Fang 0001, Xiongkuo Min, Yiru Yao, Guangtao Zhai
ICME1
2020 No Reference Quality Assessment for 3D Synthesized Views by Local Structure Variation and Global Naturalness Change
abstract
Depth image based rendering (DIBR) has been widely used to generate different virtual viewpoints of the same scene from the new perspective. However, DIBR tends to introduce annoying artifacts including blurring, discontinuity, blocking, and stretching, etc.. Thus, to improve DIBR performance, it is important to accurately measure the visual quality of synthesized views. In this paper, we propose a novel and effective no reference (NR) quality assessment method for 3D synthesized views by local variation and global change (LVGC). More specifically, we firstly compute the Gaussian derivatives for the input image to extract structure and chromatic features. Then, we use the local binary pattern (LBP) operator to encode the structure and chromatic feature maps, which are used to calculate quality-aware features to measure the local structural and chromatic distortion. Besides, we extract luminance features by global change to evaluate the naturalness of 3D synthesized views. With these extracted features, we utilize random forest regression (RFR) to train the quality prediction model from visual features to human ratings. Experimental results on three public benchmark databases demonstrate the effectiveness of our method on estimating visual quality of 3D synthesized views.
Jiebin Yan, Yuming Fang 0001, Rengang Du, Yan Zeng 0001, Yifan Zuo 0001
IEEE Trans. Image Process.1
2019 Stereoscopic image quality assessment by deep convolutional neural network
Yuming Fang 0001, Jiebin Yan, Xuelin Liu, Jiheng Wang
J. Vis. Commun. Image Represent.2
2018 No Reference Quality Assessment for Screen Content Images With Both Local and Global Feature Representation
abstract
In this paper, we propose a novel no reference quality assessment method by incorporating statistical luminance and texture features (NRLT) for screen content images (SCIs) with both local and global feature representation. The proposed method is designed inspired by the perceptual property of the human visual system (HVS) that the HVS is sensitive to luminance change and texture information for image perception. In the proposed method, we first calculate the luminance map through the local normalization, which is further used to extract the statistical luminance features in global scope. Second, inspired by existing studies from neuroscience that high-order derivatives can capture image texture, we adopt four filters with different directions to compute gradient maps from the luminance map. These gradient maps are then used to extract the second-order derivatives by local binary pattern. We further extract the texture feature by the histogram of high-order derivatives in global scope. Finally, support vector regression is applied to train the mapping function from quality-aware features to subjective ratings. Experimental results on the public large-scale SCI database show that the proposed NRLT can achieve better performance in predicting the visual quality of SCIs than relevant existing methods, even including some full reference visual quality assessment methods.
Yuming Fang 0001, Jiebin Yan, Leida Li, Jinjian Wu, Weisi Lin
IEEE Trans. Image Process.2
2017 No reference quality assessment for stereoscopic images by statistical features
abstract
In this paper, we propose a novel no reference (NR) quality assessment metric for stereoscopic images by statistical features. First, we calculate the luminance map through the local normalization, which is further used to extract the statistic luminance features. Second, we predict the disparity map of the stereoscopic image, which is further combined with the corresponding left and right views to extract the statistical structure and depth features for the stereoscopic image. The support vector regression (SVR) is employed as the mapping function from the quality-aware features to subjective quality scores. Experimental results on four publicly available large-scale stereoscopic image databases show that the proposed metric can obtain high-accuracy performance and is competitive with the state-of-the-art methods designed for visual quality prediction of stereoscopic images.
Yuming Fang 0001, Jiebin Yan, Jiheng Wang
QoMEX2
2017 Objective Quality Assessment of Screen Content Images by Uncertainty Weighting
abstract
In this paper, we propose a novel full-reference objective quality assessment metric for screen content images (SCIs) by structure features and uncertainty weighting (SFUW). The input SCI is first divided into textual and pictorial regions. The visual quality of textual regions is estimated based on perceptual structural similarity, where the gradient information is adopted as the structural feature. To predict the visual quality of pictorial regions in SCIs, we extract the structural features and luminance features for similarity computation between the reference and distorted pictorial patches. To obtain the final visual quality of SCI, we design an uncertainty weighting method by perceptual theories to fuse the visual quality of textual and pictorial regions effectively. Experimental results show that the proposed SFUW can obtain better performance of visual quality prediction for SCIs than other existing ones.
Yuming Fang 0001, Jiebin Yan, Jiaying Liu 0001, Shiqi Wang 0001, Qiaohong Li, Zongming Guo
IEEE Trans. Image Process.2