EDBT 2026 Demo / reviewers in the wild / expert
Feng Shao 0001
dblp:83/820-1
· DBLP profile ↗
151ranked-venue papers
22as first author
87since 2021 · last 2026
0000-0002-2495-9924ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 98 · 16 first-author · 47 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 23 since 2021Artificial intelligence and machine learning · 16 · 3 first-author · 9 since 2021Computer networks · 6 · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InvJND: Just Noticeable Difference Estimation via Deep Invertible Network
Qiuping Jiang, Zhihua Wang 0002, Shiqi Wang 0001, Feng Shao 0001, Guangtao Zhai, Weisi Lin |
Int. J. Comput. Vis. | 6 |
| 2026 | BigCounter: A Bidirectional-Guided Network With Scene-Semantics-Driven Fusion for RGB-Thermal Crowd CountingabstractAccurate crowd counting has become increasingly essential for public safety management and Internet of Video Things (IoVT) applications, driven by rapid population growth and urbanization. However, RGB-Thermal (RGB-T) crowd counting remains challenging due to poor recognition of small targets and degraded performance under extreme conditions such as low-light environments. To address these issues, we propose a bidirectional-guided network with scene-semantics driven fusion for RGB-T crowd counting (BigCounter) that enhances robustness and generalization in complex scenes. BigCounter comprises of three parallel branches: a primary branch, a dynamic illumination auxiliary enhancement branch (DIAEB), and a high-resolution auxiliary enhancement branch (HAEB), which respectively improve robustness under illumination variations and accuracy for small target detection. Moreover, a cross-layer scene-driven fusion module (CLSFM) and a cross-modal semantic-driven fusion module (CMSFM) are designed to strengthen structural consistency and explore semantic complementarity between modalities. Through multi-branch collaboration and semantic-aware fusion, BigCounter significantly enhances feature representation. Extensive experiments on two benchmark RGB-T datasets demonstrate that BigCounter achieves superior accuracy and generalization compared with state-of-the-art methods. Xiaomin Fan, Feng Shao 0001, Baoyang Mu, Xiongli Chai, Zhongjie Zhu, Zhiyi Mo |
IEEE Internet Things J. | 2 |
| 2026 | PU-TransMamba: A Hybrid Point Cloud Upsampling Framework With Detail-Aware Transformer and Spatially Coherent MambaabstractHigh-quality 3D point clouds are essential for high-fidelity perception in Internet of Things (IoT)-enabled intelligent systems. While Point Cloud Upsampling (PCU) is widely used to mitigate data sparsity, existing methods often struggle to balance the preservation of fine-grained local details with the maintenance of global topological consistency. Transformer-based approaches frequently suffer from excessive computational overhead and high-frequency detail loss, whereas emerging state space models like Mamba, despite their efficiency, inevitably sacrifice spatial coherence due to the 1D serialization of irregular 3D points. To address these critical bottlenecks, we introduce PU-TransMamba, a hybrid framework that synergistically leverages a Detail-Aware Transformer and a Spatially-Coherent Mamba. Each component is designed to resolve specific PCU limitations: a Complexity-Aware Bilateral Decoder is developed to adaptively recover sharp geometric edges by processing features across dual domains, while a Sequence-Aligned Mamba Encoder utilizes multiple spatial curvature descriptors to compensate for the spatial information loss inherent in serialization. Additionally, a Global Geometry Injector and a Local Neighbor Injector are designed to ensure structural integrity by infusing holistic skeletal priors and neighborhood context, respectively. To minimize feature discrepancies between the hybrid branches, we also propose a Self-Distillation Loss. Extensive experiments on five benchmark datasets demonstrate that PU-TransMamba outperforms state-of-the-art methods in both reconstruction accuracy and computational scalability. The results confirm its ability to recover intricate geometries, indicating significant potential for IoT-driven 3D perception and communication systems. Feng Shao 0001, Xiongli Chai, Hangwei Chen, Zhongjie Zhu, Zhiyi Mo |
IEEE Internet Things J. | 2 |
| 2026 | SGNet: A Structure-Guided Lightweight Network for VDT Salient Object DetectionabstractVisual-Depth-Thermal (VDT) salient object detection (SOD) aims to jointly exploit RGB, depth and thermal cues to segment the most visually significant regions. However, most existing VDT SOD models are heavy in parameters and computational cost, limiting their deployment on real-world and edge devices. To tackle this, we propose SGNet, a structure-guided lightweight network for efficient VDT SOD. Specifically, we design a lightweight Tri-modal Fusion Module (TFM) to integrate three modalities at the semantic level, and a Shared Structure Extraction Module (SSEM) to extract common structural information from depth and thermal modalities. A Structure Refine Module (SRM) further injects the extracted structure into the deepest semantic features, while a Multiscale Feature Refinement Module (MFRM) progressively decodes multi-level features under deep supervision to produce saliency maps with clear boundaries. Benefiting from these modules, SGNet achieves competitive performance on the VDT2048 benchmark with only 5.51 M parameters and a real-time speed of 120 FPS at 320 × 320 resolution, surpassing state-of-the-art methods while remaining deployment-friendly. Huizhi Wang, Feng Shao 0001, Xuebin Wei, Xiongli Chai, Hangwei Chen, Zhongjie Zhu |
IEEE Internet Things J. | 2 |
| 2026 | Multi-dimensional human preference assessment for AI-generated images with supervised contrastive learning
Xuebin Wei, Feng Cai, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | Interactive feature fusion for camera-radar-based vehicle segmentation in bird's-eye view
Chenyang Lu 0002, Liang Li 0010, Xiangchao Meng, Qiuping Jiang, Feng Shao 0001 |
Pattern Recognit. | 6 |
| 2026 | Edge-Embedded Bidirectional Interactive Network for Lightweight Surface Defect DetectionabstractConvolutional neural networks (CNN)-based surface defect detection methods have achieved remarkable results. However, most existing approaches face two key challenges: first, they typically require substantial computational resources to learn rich features, making it difficult to balance performance and computational cost; second, when dealing with complex backgrounds and various defect shapes, they often lose edge details. To address these challenges, we propose an Edge-embedded Bidirectional Interactive Network (EBINet) for lightweight surface defect detection. Specifically, we introduce an edge generation module (EGM), which enhances feature details by interacting with low-level and high-level features, providing high-quality edge details for defect regions. In addition, we design a bidirectional interactive decoder (BID) that consists of a self-recognition module (SRM), cross-attention fusion module (CAFM), and multiscale edge embedded module (MEEM). This decoder gradually integrates features from different stages and thus can effectively capture inter-layer feature relationships to generate high-quality saliency maps. Extensive experiments conducted on four public defect datasets demonstrate that the lightweight EBINet offers strong competitiveness and superior performance, requiring only 3.57M parameters and 2.30G FLOPs for a 256×256 input image. Feng Shao 0001, Dongze Jin, Baoyang Mu, Hangwei Chen |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | USformer: A U-Shaped Structure Transformer for RGB-Thermal Semantic Segmentation and Traffic Scene UnderstandingabstractRecent advancements in multimodal approaches, particularly RGB-thermal (RGB-T) segmentation, have significantly promote the development of Intelligent Transportation Systems (ITS). However, existing methods still encounter challenges related to modality discrepancy and the effective integration of multi-scale features. To address these issues, we propose the U-shaped Structure Transformer (USformer) for RGB-T semantic segmentation. We improve the feature flow of existing methods by designing a novel U-shaped encoding network that integrates inter-layer fusion and cross-modal fusion. Specifically, our method introduces an inter-layer interaction mechanism that facilitates the iterative fusion of high-level semantic and low-level detail features. For each layer, our fusion process is divided into two stages: the Cross-Modal and -Scale Auxiliary (CMSA) module enforces distribution alignment across modalities and scales, while the Cross-Attention Feature Merger (CAFM) allows each modality to refine its own feature selection by employing a multi-head cross-attention mechanism. These modules effectively adapt and integrate well-established attention designs into our U-shaped encoding architecture, thereby achieving efficient multi-modal feature alignment and fusion. Finally, we utilize the Mask2Former decoder to aggregate the fused features from multiple layers and improve the segmentation across various object sizes and complex scenes. Extensive experiments on four RGB-T datasets demonstrate that our proposed USformer achieves state-of-the-art performance. Feng Shao 0001, Baoyang Mu, Xiongli Chai, Qiuping Jiang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2026 | Quality Evaluation of AI-Generated Images: Subjective Study and Objective MethodologyabstractIn recent years, AI-Generated Images (AIGIs) have attracted significant attention and shown great potential in various applications, including entertainment, advertisement, education, and product design. Driven by this trend, various Text-to-Image (T2I) models are developed. However, the quality of AIGIs produced by these models varies widely, with many low-quality images failing to meet human aesthetic standards. Consequently, research into both subjective and objective Image Quality Assessment (IQA) methods for AIGIs is crucial. In this paper, we introduce a dataset called AIGI-IQAD, designed to enhance our understanding of human aesthetic preferences for AIGIs. The dataset contains 2,880 AIGIs generated by 8 T2I models using 360 deliberately designed text prompts. Further, we conducted subjective experiments to gather ratings from both aesthetic quality and text-image consistency. Building on this dataset, we propose a model named Question-guided Multimodal Interaction Network (QMI-Net) for evaluating AIGIs. QMI-Net assesses human preferences for AIGIs by focusing on both aesthetic quality and text-image consistency. Specifically, QMI-Net uses a question-answering approach to guide Multimodal Large Language Models (MLLMs) in generating detailed aesthetic and similarity information. The Visual and Aesthetic Feature Fusion Module (VAFFM) then fuses the aesthetic features with the visual features extracted by Contrastive Language-Image Pre-training (CLIP) to obtain more comprehensive aesthetic quality features. Comprehensive experiments demonstrate that state-of-the-art performance is achieved by QMI-Net on our AIGI-IQAD and three other public datasets.The AIGI-IQAD datasets and QMI-Net will be released athttps://github.com/ctxya1207/QMI-Net. Feng Shao 0001, Hangwei Chen, Xuejin Wang, Qiuping Jiang |
IEEE Trans. Multim. | 2 |
| 2025 | Rethinking Lightweight RGB-Thermal Salient Object Detection With Local and Global Perception NetworkabstractRGB–thermal salient object detection (RGB-T SOD) aims to segment the most intriguing parts through RGB and thermal images. However, the high computational costs and large model sizes of existing methods prevent its deployment in edge computing platforms. To cope with it, a new lightweight RGB-T SOD method is proposed, namely, local and global perception network (LGPNet), which relieves the pressure of data transmission and processing and improves robustness in complicated surroundings. Specifically, high-level lightweight fusion (HLF) blocks and low-level lightweight fusion (LLF) blocks are designed to reduce the difference between two modalities and perform multimodal feature fusion. Unlike existing convolutional neural network-based lightweight SOD methods, HLF and LLF blocks combine the spatial inductive bias of convolutional neural networks with the global perception of Transformer, which is able to perform feature extraction and fusion with fewer parameters and larger receptive field. Experimental results show that our method is able to compete with the state-of-the-art RGB-T SOD methods while having only 7.35 learnable parameters[M], 6.40G floating-point operations and real-time speed (33 frames per second (FPS) on PyTorch framework and 224 FPS on TensorRT framework). Dongze Jin, Feng Shao 0001, Zhengxuan Xie, Baoyang Mu, Hangwei Chen |
IEEE Internet Things J. | 2 |
| 2025 | RGBT-Booster: Detail-Boosted Fusion Network for RGB-Thermal Crowd Counting With Local Contrastive LearningabstractWith the swift development of the Internet of Video Things (IOVT), crowd counting has demerged as an indispensable technology in the domains of intelligent transportation and video surveillance. However, due to the insufficient extraction of detail head information and the limited ability to reduce the multimodality differences, the existing methods still have large errors in accurate RGB-thermal (RGB-T) crowd counting. To this end, we propose a novel RGB-T crowd counting network, i.e., RGBT-Booster, to effectively deal with the aforementioned challenges. In RGBT-Booster, by introducing additional detail auxiliary branches for RGB and thermal infrared images and the proposed enhanced detail fusion module (EDFM), we can obtain richer low-level head detail features. In addition, we also propose a local contrastive learning (LCL) to further reduce the multimodality differences for accurate crowd counting. Experimental results on two public RGB-T crowd counting datasets (i.e., RGBT crowd counting (RGBT-CC) and DroneRGBT) and one RGB-Depth (RGB-D) crowd counting dataset (i.e., ShanghaiTechRGBD) show that the proposed RGBT-Booster achieves effective and superior counting performance, compared with previous methods. The source code and datasets used in the experiments will be released athttps://github.com/QSBAOYANGMU/RGBT-Booster. Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Long Xu 0001, Qiuping Jiang |
IEEE Internet Things J. | 2 |
| 2025 | Self-supervised panoramic stitched image quality assessment based on contrastive learning
Xiaoer Li, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | MDGINet: Multi-frequency Dynamic Guidance and Interaction network for image denoising
Feng Shao 0001, Hangwei Chen, Xiongli Chai, Qiuping Jiang |
Knowl. Based Syst. | 2 |
| 2025 | Art Comes From Life: Artistic Image Aesthetics Assessment via Attribute Knowledge AmalgamationabstractAssessing the aesthetic quality and visual appeal of artworks has become one of the hotspots in current research. The existing artistic image aesthetics assessment (AIAA) methods directly learn aesthetics from images, while ignoring the impact of variations in visual attributes on human aesthetic perception, which hampers the further development of AIAA. To address this issue, this paper presents a new AIAA method based on attribute knowledge amalgamation, named AKA-Net. Specifically, we initially learn common attribute aesthetic rules (e.g., composition and color) through pre-training on natural aesthetic images. Then, we devise a multi-model amalgamation strategy based on contrastive learning to transfer different types of prior attribute knowledge into a single target model, enabling flexible and efficient aesthetic prediction. Finally, an attribute-aware feature enhancement module (AFEM) is introduced to better establish the relationship between aesthetic quality and attribute knowledge. Experimental results on three public benchmark AIAA databases demonstrate that the proposed AKA-Net outperforms the state-of-the-art AIAA metrics. Hangwei Chen, Feng Shao 0001, Xiongli Chai, Baoyang Mu, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | GADFNet: Geometric Priors Assisted Dual-Projection Fusion Network for Monocular Panoramic Depth EstimationabstractPanoramic depth estimation is crucial for acquiring comprehensive 3D environmental perception information, serving as a foundational basis for numerous panoramic vision tasks. The key challenge in panoramic depth estimation is how to address various distortions in 360° omnidirectional images. Most panoramic images are displayed as 2D equirectangular projections, which exhibit significant distortion, particularly with the severe fisheye effect near the equatorial regions. Traditional depth estimation methods for perspective images are unsuitable for such projections. On the other hand, cubemap projection consists of six distortion-free perspective images, allowing the use of existing depth estimation methods. However, the boundaries between faces of a cubemap projection introduce discontinuities, causing a loss of global information when using cube maps alone. In this work, we propose an innovative geometric priors assisted dual-projection fusion network (GADFNet) that leverages geometric priors of panoramic images and the strengths of both projection types to enhance the accuracy of panoramic depth estimation. Specifically, to better focus the network on key areas, we introduce a distortion perception module (DPM) and incorporate geometric information into the loss function. To more effectively extract global information from the equirectangular projection branch, we propose a scene understanding module (SUM), which captures features from different dimensions. Additionally, to achieve effective fusion of the two projections, we design a dual projection adaptive fusion module (DPAFM) to dynamically adjust the weights of the two branches during fusion. Extensive experiments conducted on four public datasets (including both virtual and real-world scenarios) demonstrate that our proposed GADFNet outperforms existing methods, achieving superior performance. Chengchao Huang, Feng Shao 0001, Hangwei Chen, Baoyang Mu, Long Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | A Mutual Head Knowledge Distillation Framework for Lightweight RGB-T Crowd CountingabstractAs an important technology in the fields of intelligent transportation and public safety, crowd counting that can obtain pedestrian flow information has attracted extensive attention from academic and industrial communities. However, existing RGB-T crowd counting methods cannot effectively balance the counting accuracy and computational complexity in practical applications. For this, we propose a Mutual Head Knowledge Distillation Framework (MHKDF) to obtain a lightweight RGB-T crowd counting network for efficient and accurate pedestrian number estimation. Specifically, to avoid the influence of parameter and structure differences between teacher and student networks on the distillation effect, we propose a Cooperative Mutual Knowledge Distillation (CMKD) strategy to comprehensively and dynamically transfer the crowd analysis ability of the complex teacher model (MHKDF-T) to the lightweight student model (MHKDF-S). In addition, the upper bound of the performance of the student network depends on the teacher model with high accuracy. Therefore, to take advantage of the complementary advantages of frequency domain and spatial domain feature fusion, we propose a Multi-Modal Spatial-Frequency Hybrid Fusion Module (MSFHFM) to futher improve counting accuracy of MHKDF-T. Comprehensive experiments on two RGB-T crowd counting datasets demonstrate that our MHKDF-S achieves competitive performance with only 5.68 FLOPs and 4.89M parameters. Our code will be released at https://github.com/BaoYangCC/MHKDF. Baoyang Mu, Feng Shao 0001, Hangwei Chen, Xuejin Wang, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Adversarial Robust Salient Object Detection in Optical Remote Sensing Images With Implicit Feature EnhancementabstractDeep neural networks (DNNs) have achieved significant progress in optical remote sensing images salient object detection (ORSI-SOD) and are widely applied to various remote sensing image analysis tasks. However, few SOD models demonstrate robust performance under adversarial perturbations, which ultimately leads to a decline in detection accuracy. Moreover, most existing defense methods inject fixed Gaussian noise globally into the image. Although such approaches are easy to implement, they have several limitations in inaccurate uncertainty estimation and neglecting the unique characteristics of local salient regions. Furthermore, existing adversarial defense research rarely addresses the challenges specific to the ORSI-SOD task, leaving a gap in effective defense strategies. To tackle these issues, we propose a novel defense method, which enhances the adversarial robustness of ORSI-SOD models through implicit feature enhancement. The algorithm first proposes a two-stage strategy of reverse local noise search and forward global noise optimization, enhancing generalization ability by implicitly enhancing features to better simulate network uncertainty. Then, the algorithm proposes a global-guided texture information enhancement (GTIE) module for low-level features and a global-guided semantics information enhancement (GSIE) module for high-level features, focusing on strengthening low-level texture information and enhancing the model’s understanding of high-level contextual semantic features, respectively. This dual-module design effectively weakens the impact of adversarial noise, significantly improving the robustness and accuracy of object detection. Extensive experiments on three ORSI-SOD datasets demonstrate that our defense strategy better estimates the uncertainty, resulting in an average performance improvement of 23.2% in$F_{\beta } ^{\mathrm { max}}$and 34.1% in$E_{\xi } ^{\mathrm { max}}$across six ORSI-SOD models under five different adversarial attack methods. Our code will be released in the public repository athttps://github.com/kexi0714/IFe. Feng Shao 0001, Xiangchao Meng, Hangwei Chen, Xiongli Chai, Zhiyi Mo |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Dual-Branch Cross-Resolution Interaction Learning Network for Change Detection at Different ResolutionsabstractChange detection (CD) plays a critical role in remote sensing (RS) image analysis. However, the detection accuracy is often compromised due to differences in imaging conditions between bitemporal images, especially in scenarios where images have varying spatial resolutions from multisource RS satellites. To address this challenge, we propose an innovative dual-branch cross-resolution interaction learning network (DCILNet). This network employs strategies for image spatial resolution alignment and feature space correction to achieve efficient CD. To fully leverage the information from images with different resolutions, we design two cross-resolution CD branches. These branches interact through a cross-resolution feature correction module (CRFCM) and a multiresolution feature fusion module (MRFM), facilitating feature interaction and learning across branches to maximize feature representation. We conduct both qualitative and quantitative experiments on three public datasets. The experimental results show that, compared to other comparative methods, the proposed DCILNet exhibits stronger competitiveness. Our research shows that by integrating image spatial resolution alignment and feature space correction strategies and adopting dual-branch interactive learning, the model effectively addresses the challenges posed by resolution discrepancies in CD tasks. Our code will be available athttps://github.com/Li738/DCILNet. Feng Shao 0001, Xiangchao Meng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | IM-CMDet: An Intramodal Enhancement and Cross-Modal Fusion Network for Small Object Detection in UAV Aerial Visible-Infrared ImageryabstractUAV aerial Visible-Infrared (RGBT) object detection has been widely applied in fields such as military operations and rescue missions. However, although numerous UAV aerial RGBT object detection methods exist, several challenges remain in this field. On the one hand, drones typically operate at high altitudes, and objects only occupy a small number of pixels in imaging, posing a significant challenge to object detection. On the other hand, spatial misalignment between modalities remains a major obstacle in cross-modal fusion—especially given the small size of the objects. To address the above issues, this paper proposes IM-CMDet, an intra-modal enhancement and cross-modal fusion network for small object detection in UAV-based RGBT imagery, which comprises three effective modules: the Detail-Semantics Joint Enhancement module (DSJE), the Differential-based Fusion Weight Generation module (DFWG) and the Feature Reconstruction Network (FRN). The DSJE module prevents small object features from being overwhelmed by background noise through optimizing feature representations across different levels. The FRN module is designed to overcome modality differences and build inter-modality information correlation via swin-Transformer architecture. To further enhance the network’s sensitivity to small objects, the DFWG combines differential and spatial attention to generate the final fusion weights while reducing the impact of background noise on detection performance. Extensive experiments on RGBTDronePerson and two additional benchmarks demonstrate that IM-CMDet achieves state-of-the-art performance through effective cross-modal fusion, significantly advancing small-object detection in complex aerial scenarios. The code is available at https://github.com/RS-Minchao/IM-CMDet. Minchao Luo, Rui Zhao 0003, Shenfu Zhang, Feng Shao 0001, Xiangchao Meng |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Dual-Task Cascaded Network for Spatial-Temporal-Spectral Remote Sensing Image FusionabstractSpatial-temporal-spectral fusion is dedicated to integrating the complementary advantages of multisource images to obtain fused image with all high spatial, high temporal and high spectral resolutions, which is promising but more challenging. On the one hand, traditional studies deployed on MODIS and Landsat data cannot be transferred to most spaceborne hyperspectral (HS) data with lower temporal resolution; on the other hand, the rigid time relation modeling in most existing studies exhibits weakness orienting to non-linear land-cover changes. In this paper, we propose a dual-task cascaded network for spatial-temporal-spectral fusion, with collaborative modeling on spatialspectral joint enhancement and temporal variation estimation in a unified framework. The spatial-spectral joint enhancement task was designed with an iterative alternating projection, meticulously crafted to address the scale variance among observation. Additionally, the spatial enhancement unit and error correction unit were coupled modeling to enhance the spatial and spectral fidelity. The temporal variation estimation on spectral fine tuning network was developed, to further enhance the temporal and spectral fidelity. Extensive experiments were implemented on Ziyuan(ZY)-1 02D HS data and Sentinel-2 multispectral (MS) data. Both qualitative and quantitative results demonstrated the competitive performance of the proposed method. Xiangchao Meng, Xu Chen 0041, Mengjing Zhang, Feng Shao 0001, Gang Yang 0006, Weiwei Sun 0005 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Bidirectional Spectral Attention Multiscale Aggregation Network for Spectral Super-ResolutionabstractSpectral super-resolution (SSR) is the computational process of generating a high-dimensional hyperspectral image from a low-dimensional image through spectral reconstruction techniques. Recently, deep learning has demonstrated remarkable potential in the field of SSR, achieving impressive results. However, existing deep learning-based approaches often fail to deliver high-fidelity SSR outcomes. These methods tend to focus primarily on spectral information while paying insufficient attention to the critical role of spatial features. Furthermore, they lack effective strategies for capturing inter-band relationships, resulting in suboptimal spectral information modeling. To address these limitations, we propose a novel network for SSR, termed Bidirectional Spectral Attention Multi-Scale Aggregation Network (BiSANet). BiSANet features three U-Net-like branches and integrates two advanced attention mechanisms. The bidirectional spectral attention modules dynamically model inter-spectral dependencies through forward and reverse spectral feature extraction, enhanced by a weight-sharing strategy. Specifically, we reverse the spectral order of feature maps to activate complementary global trends and local details, overcoming the limitations of unidirectional modeling in traditional methods. Additionally, an independent spatial reconstruction branch with a dedicated loss function ensures precise spatial detail preservation. Experimental results demonstrate that BiSANet outperforms state-of-the-art methods across three benchmarks. For instance, on the DFC2018 Houston dataset, it achieves a 4.26% PSNR improvement and an 11.52% SAM reduction, highlighting its robustness and accuracy in spectral-spatial reconstruction. Xintao Zhong, Shenfu Zhang, Gang Yang 0006, Weiwei Sun 0005, Feng Shao 0001, Xiangchao Meng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Deep Underwater Image Quality Assessment With Explicit Degradation Awareness EmbeddingabstractUnderwater Image Quality Assessment (UIQA) is currently an area of intensive research interest. Existing deep learning-based UIQA models always learn a deep neural network to directly map the input degraded underwater image into a final quality score via end-to-end training. However, a wide variety of image contents or distortion types may correspond to the same quality score, making it challenging to train such a deep model merely with a single subjective quality score as supervision. An intuitive idea to solve this problem is to exploit more detailed degradation-aware information as supplementary guidance to facilitate model learning. In this paper, we devise a novel deep UIQA model with Explicit Degradation Awareness embedding, i.e., EDANet. To train the EDANet, a two-stage training strategy is adopted. First, a tailored Degradation Information Discovery subnetwork (DIDNet) is pre-trained to infer a residual map between the input degraded underwater image and its pseudoreference counterpart. The inferred residual map explicitly characterizes the local degradation of the input underwater image. The intermediate feature representations on the decoder side of DIDNet are then embedded into the Degradation-guided Quality Evaluation subnetwork (DQENet), which significantly enhances the feature characterization capability with higher degradation awareness for quality prediction. The superiority of our EDANet against 18 state-of-the-art methods has been well demonstrated by extensive comparisons on two benchmark datasets. The source code of our EDANet is available at https://github.com/yia-yuese/EDANet. Qiuping Jiang, Yuese Gu, Zongwei Wu, Chongyi Li, Huan Xiong, Feng Shao 0001, Zhihua Wang 0002 |
IEEE Trans. Image Process. | 6 |
| 2025 | Remote sensing multi-view stereo using ConvLSTM guided iterative depth refinement
Xuebin Wei, Yunxin Ye, Feng Cai, Liyan Wu, Feng Shao 0001 |
J. Supercomput. | 5 |
| 2025 | No-Reference Point Cloud Quality Assessment via Graph Convolutional NetworkabstractThree-dimensional (3D) point cloud, as an emerging visual media format, is increasingly favored by consumers as it can provide more realistic visual information than two-dimensional (2D) data. Similar to 2D plane images and videos, point clouds inevitably suffer from quality degradation and information loss through multimedia communication systems. Therefore, automatic point cloud quality assessment (PCQA) is of critical importance. In this work, we propose a novel no-reference PCQA method by using a graph convolutional network (GCN) to characterize the mutual dependencies of multi-view 2D projected image contents. The proposed GCN-based PCQA (GC-PCQA) method contains three modules, i.e., multi-view projection, graph construction, and GCN-based quality prediction. First, multi-view projection is performed on the test point cloud to obtain a set of horizontally and vertically projected images. Then, a perception-consistent graph is constructed based on the spatial relations among different projected images. Finally, reasoning on the constructed graph is performed by GCN to characterize the mutual dependencies and interactions between different projected images, and aggregate feature information of multi-view projected images for final quality prediction. Experimental results on two publicly available benchmark databases show that our proposed GC-PCQA can achieve superior performance than state-of-the-art quality assessment metrics. Qiuping Jiang, Wei Zhou 0021, Feng Shao 0001, Guangtao Zhai, Weisi Lin |
IEEE Trans. Multim. | 4 |
| 2025 | Cross-Modal Hierarchical Knowledge Distillation for Image Aesthetics AssessmentabstractThe field of image aesthetics assessment (IAA) is rapidly advancing due to its wide applications. However, relying solely on single-modal information for aesthetic evaluation presents inherent limitations. While multimodal IAA models incorporating user comments have achieved significant advancements, these comments are often unavailable due to privacy concerns and practical considerations, and they also introduce additional computational overhead during inference. To address this issue, we propose a cross-modal hierarchical knowledge distillation method, termed HKD-IAA, to enhance the performance of unimodal image models effectively. Specifically, HKD-IAA comprises four components: feature extraction, feature decomposition, hierarchical knowledge distillation, and dynamic decay. During training, we first decompose the extracted features into a weighted sum of basic aesthetic elements and their corresponding weights, thereby reducing the learning difficulty for the student model. Building on this, we design a new hierarchical knowledge distillation framework, which aligns features at the feature, relation, and response levels to effectively transfer the knowledge from the teacher model. Finally, we introduce a dynamic decay strategy to adjust the weight of the distillation loss, thereby enhancing the student model's learning effectiveness during training. Extensive experiments on two benchmark datasets validate that the proposed method achieves state-of-the-art performance using only visual modal data. Our code is available athttps://github.com/Hangwei-Chen/HKD-IAA. Hangwei Chen, Feng Shao 0001, Weiyi Jing, Huizhi Wang, Qiuping Jiang |
IEEE Trans. Multim. | 2 |
| 2025 | Cross-Projection Distilling Knowledge for Omnidirectional Image Quality Assessment
Huixin Hu, Feng Shao 0001, Hangwei Chen, Xiongli Chai, Qiuping Jiang |
IEEE Trans. Multim. | 2 |
| 2025 | MISF-Net: Modality-Invariant and -Specific Fusion Network for RGB-T Crowd CountingabstractTo accurately perform crowd counting, utilizing the complementary relationship between RGB and thermal images to analyze the crowd has become the focus of current research. Due to different imaging principles, multi-modal images often contain different contents, which are their modality-specific information. For example, RGB images contain more texture and color details, while thermal images contain thermal radiation information. Meanwhile, they also describe the same target content, e.g., crowds, which are modality-invariant. However, existing methods only design different modules to directly fuse RGB and thermal image features, which did not fully consider the above facts. In this paper, by analyzing the similarities and differences between multi-modal images, we propose a Modality-Invariant and -Specific Fusion Network (MISF-Net) for RGB-T Crowd Counting. Specifically, we design a modality decomposition and fusion module (MDFM), which decomposes RGB and thermal image features into modality-invariant and -specific features by using the similarity and difference supervision between multi-modal features. Besides, reconstruction supervision is also used to prevent network learning from generating bias. After that, different fusion strategies are applied to the invariant and specific features, respectively. In addition, to adapt to the variations in size of different pedestrians, we design a modality-invariant fusion module (MIFM). Finally, after the fusion decoder, MISF-Net can obtain a more accurate crowd density map. Comprehensive experiments on the RGB-T crowd counting dataset show that our MISF-Net can achieve competitive performance. Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Hangwei Chen, Zhongjie Zhu, Qiuping Jiang |
IEEE Trans. Multim. | 2 |
| 2025 | Spatial-Spectral Heterogeneity-Aware Network for Hyperspectral and LiDAR Joint ClassificationabstractThe integration of hyperspectral (HS) imagery and light detection and ranging (LiDAR) data for land cover classification has emerged as a prominent research focus. Despite the satisfactory classification accuracies achieved by existing methodologies, several unaddressed issues that remain warrant consideration. First, current approaches overlook the pronounced spectral and spatial heterogeneities in remote sensing (RS) images designated for multiclassification tasks, limiting the performance of classification models. Moreover, most existing studies amalgamate elevation features with other characteristics through simple addition and interaction operations, and they do not delve deeply into exploiting elevation height information, leading to an imbalance in the representation of elevation height. In light of the aforementioned issues, this article introduces a spatial-spectral heterogeneity-aware network (S2HANet) for the joint classification of HS and LiDAR data. Specifically, a shared spectral correction module (SSCM) is designed in the spectral branch to preliminarily alleviate the problem of large intraclass variance, followed by the use of a contrastive learning framework to enhance the intraclass compactness and interclass separability of spectral features. A multichannel signed distance discrimination module (MCSDDM) is developed to learn the distance relationships between intra- and interclass pixels and boundaries, and using prior boundary information to improve spatial boundary information. In addition, an elevation boost module (EBM) and an elevation injection module (EIM) are meticulously designed to phase-in elevation height information, further enhancing the utilization of elevation data and better facilitating the fusion of the two modalities. The proposed S2HANet has demonstrated exceptional classification performance across three opening benchmark datasets. Shenfu Zhang, Qiang Liu 0035, Rui Zhao 0003, Feng Shao 0001, Xiangchao Meng |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | CAFCNet: Cross-modality asymmetric feature complement network for RGB-T salient object detection
Dongze Jin, Feng Shao 0001, Zhengxuan Xie, Baoyang Mu, Hangwei Chen, Qiuping Jiang |
Expert Syst. Appl. | 2 |
| 2024 | Hallucinated-PQA: No reference point cloud quality assessment via injecting pseudo-reference features
Baoyang Mu, Feng Shao 0001, Hangwei Chen, Qiuping Jiang, Long Xu 0001, Yo-Sung Ho |
Expert Syst. Appl. | 2 |
| 2024 | Multi-domain pseudo-reference quality evaluation for infrared and visible image fusionabstractAbstract Infrared and visible image fusion involves merging the advantages of infrared and visible images to generate a composite image that encompasses thermal radiation as well as intricate texture details. Infrared and visible image fusion has garnered increasing attention, with numerous fusion methods proposed. However, how to fairly perceive the performance of fused image remains a contentious topic. This paper is dedicated to solving this problem from two perspectives (e.g., subjective and objective aspects). Firstly, an infrared and visible fusion image quality assessment dataset was constructed, including 60 pairs of infrared and visible images captured in various scenes, along with 540 fusion images with different types and degrees of distortions. Additionally, a subjective evaluation dataset of 16,200 subjective scores by 30 participants was further provided for the fused image. Secondly, to overcome the challenging assessment for infrared and visible fusion images without a real reference image, an interesting multi‐domain pseudo‐reference image quality assessment model (MPIQAM) is proposed, by comprehensively considering the thermal radiation information distortion, texture information distortion, and overall naturalness of the fused image. The proposed MPIQAM was compared with 18 mainstream objective metrics, and the experimental findings showcased a commendable level of competitiveness. Xiangchao Meng, Chaoqi Chen, Qiang Liu 0035, Feng Shao 0001 |
IET Image Process. | 4 |
| 2024 | Visual Prompt Multibranch Fusion Network for RGB-Thermal Crowd CountingabstractAs population growth and urbanization continue, accurate crowd counting is increasingly important for public safety management and the Internet of Video Things (IOVT). However, RGB and thermal infrared (RGB-T) crowd counting still faces challenges in improving feature extraction capability for RGB streams and reducing multimodality differences. For this, we propose a visual prompt multibranch fusion network (VPMFNet) to tackle the above challenges. Specifically, to improve the ability of crowd analysis of the RGB stream in RGB-T crowd counting, through designing the prompt enhancement module, we take the prior features of head perception in the crowd as visual prompt cues to embed into the RGB stream. In terms of RGB and thermal image feature fusion, we fully reduce the modality differences from the perspectives of local fusion, global fusion, and multireceptive field fusion to accurately estimate the pedestrian number. Various experiments on two RGB-T crowd counting data sets demonstrate that our VPMFNet achieves a smaller estimation error in the number of pedestrians. Besides, our VPMFNet outperforms existing methods (i.e., multicolumn convolutional neural network, BL, SANet, UCNet, HDFNet, BBSNet, BL+IDAM, BL+CSCA, dual-branch enhanced feature fusion network, and GETANet) on the RGB-D data set. Our code will be released athttps://github.com/QSBAOYANGMU/VPMFNet. Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Hangwei Chen, Qiuping Jiang, Yo-Sung Ho |
IEEE Internet Things J. | 2 |
| 2024 | Blind cartoon image quality assessment based on local structure and chromatic statistics
Hangwei Chen, Xuejin Wang, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Blind quality assessment of light field image based on view and focus stacks
Fucui Li, Mengmeng Ye, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Hybrid CNN-transformer based meta-learning approach for personalized image aesthetics assessment
Xingao Yan, Feng Shao 0001, Hangwei Chen, Qiuping Jiang |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Plain-PCQA: No-Reference Point Cloud Quality Assessment by Analysis of Plain Visual and Geometrical ComponentsabstractIn reviewing the research progress in Point Cloud Quality Assessment (PCQA), two main pathways have emerged, i.e., 2D projections and 3D point descriptors. The former primarily focuses on visual information, while the latter concentrates on crucial geometrical information in three-dimensional space. However, the current studies lack a thorough investigation of the impact of visual components and seldom pay special attention to plane-point fusion strategies. To comprehensively represent features and effectively tackle various types of impairments, we propose an end-to-end learning paradigm, only considering plain visual and geometrical factors called Plain-PCQA, for quantitatively evaluating objective metrics of 3D dense point clouds associated with human perception. Firstly, we explore a sophisticated preprocessing technique. The entire point clouds are packaged into six projections by moving virtual cameras, which can conveniently increase the visual samples during the training stage. Given the high resolution of the projected image, we have opted for a relatively lightweight network, namely ResNet-18, as the backbone to enable higher resolution input data. Five cropped patches from the projected image are collectively fed into this network. In light of the presence of some invalid information in the projections, a mask weight is devised to calculate the significance of each patch based on its effective informational content. Secondly, dual neural networks, comprising of a No-Reference (NR) branch and a Degraded-Reference (DR) branch, are designed with fundamental visual components to provide quantitative quality metrics. Specifically, the NR branch utilizes the feature output of each block in the Vision Transformer (ViT) model to obtain long-range low-level and high-level visual NR quality. The DR branch employs KLT (Karhunen-Loève Transform) to acquire the principal component information of an image as the macro-structural image, and then feeds the difference between input images and macro-structural images into a network for DR quality extraction. Thirdly, a Plane-Point Interaction Transformer (P2IT) is presented by incorporating texture and semantic features in 2D projections and geometrical features in 3D spaces to characterize the complete features with a connected 2D-3D feature representation. With these elaborately designed deep features, the proposed model can achieve competitive performances relying solely on plain visual and geometrical components. The experimental results demonstrate the potential of the proposed approach in multiple representative databases, which surpasses existing state-of-the-art methods significantly. Xiongli Chai, Feng Shao 0001, Baoyang Mu, Hangwei Chen, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Dynamic Weighted Fusion and Progressive Refinement Network for Visible-Depth-Thermal Salient Object DetectionabstractThe introduction of depth/thermal modality has significantly enhanced the performance of dual-modal salient object detection (SOD) methods. However, depth maps and thermal images are prone to environmental interference, making them insufficient for providing salient information. To address this challenge, triple-modal SOD methods have been proposed. However, these methods often overlook the detrimental effects of defective modalities during fusion, leading to subpar performance. To tackle this issue, we present a novel dynamic weighted fusion and progressive refinement network (DWFPRNet) for Visible-Depth-Thermal (V-D-T) SOD. Specifically, we first use the dual-modal fusion module (DFM) to fuse dual modalities, thereby obtaining fused features. Subsequently, the modality selective fusion module (MSFM) mines complementary information between fused features, considering both fusion features and the quality of feature maps, to achieve weighted fusion. Finally, we design a progressive refinement decoder (PRD) to realize interaction and multi-scale learning among different scale features and generate high-quality saliency maps. Extensive experiments conducted on the VDT-2048 public dataset demonstrate that our method outperforms existing state-of-the-art multi-modal methods. Feng Shao 0001, Baoyang Mu, Hangwei Chen, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | CIG-STF: Change Information Guided Spatiotemporal Fusion for Remote Sensing ImagesabstractSpatiotemporal fusion has been attracting increasing attention in remote sensing applications, such as environmental monitoring and land cover change detection, due to its excellent ability to obtain high spatial and temporal resolution images. The land cover change has always been a great challenge in spatiotemporal fusion. Although most spatiotemporal fusion methods have demonstrated satisfactory performance in addressing phenological changes, the performance in terms of abrupt land cover type changes, such as floods or mudslides, falls short. To alleviate this issue, we propose a change information guided spatiotemporal fusion (CIG-STF) method. The proposed CIG-STF integrates change detection and spatiotemporal fusion in a unified framework, by taking advantage of change detection in capturing land cover changes to assist spatiotemporal fusion. Specifically, the CIG-STF comprises three modules: multiscale dilated feature extractor module (MDFE), spatiotemporal fusion-change detection integrated module (STF-CD), and reconstruction module. The MDFE employs multiscale dilated convolutions to comprehensively extract features, to prevent crucial information loss by increasing the convolutional receptive field. In the STF-CD, a change detection module on attention strategy is integrated into the spatiotemporal fusion task, by excavating land cover changes to further enhance the fusion performance. In addition, we design a dynamic decay loss function to further leverage change information, ensuring the accuracy of both change information and prediction results. The experiments were verified on the publicly available LGC and Daxing datasets with manual change labels. The experimental results demonstrate the superior performance of the proposed CIG-STF in both phenological variations and land cover type changes. Mingzhu You, Xiangchao Meng, Qiang Liu 0035, Feng Shao 0001, Randi Fu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Blind Quality Evaluator of Light Field Images by Group-Based Representations and Multiple Plane-Oriented Perceptual CharacteristicsabstractDue to the emergency of multi-view cameras and commercial Light Field (LF) cameras, the demand of high-performance LF quality evaluator is of great significance for guiding LF acquisition, processing and application and further promoting the visual perceived quality of LF visualizations. However, LF Images (LFIs), as high-dimensional data, suffer from various quality degradations not only in the spatial domain but also in the angular domain. Therefore, it is of great challenge to predict LF quality accurately. An effective LF evaluator should be able to represent these heterogeneous artifacts. In this paper, we provide a novel No-Reference LF Quality Assessment Evaluator (NR LF-QAE) to tackle this problem. Firstly, to measure angular consistency among viewports, we utilize group-based representations to character information similarity of aligned view stacks. Secondly, to better describe the texture information of LFIs, unifying spatial-angular texture statistic measurement is performed via Local Binary Patterns from Three Orthogonal Planes (LBP-TOP). Thirdly, we design 3D Log-Gabor filters to extract LF global structure information in Sub-Aperture Images (SAIs) as spatial feature characterizations and 2D Log-Gabor filters are adopted to characterize ray direction/depth information in Epipolar Plane Images (EPIs) as angular feature characterizations. By comprehensive LF information analyses in angular consistency and spatial-angular feature extraction with texture and structure descriptors, experimental results demonstrate the superiority of the proposed NR LF-QAE over the state-of-the-art comparative models in predicting the quality of LFIs on three available benchmark databases. The code will be released athttps://github.com/zerosola/NR-LF-QAE. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xuejin Wang, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2024 | SCFANet: Semantics and Context Feature Aggregation Network for 360° Salient Object DetectionabstractHow to solve the problem of geometric distortion is the key for salient object detection (SOD) in 360° omnidirectional images. Most of the current methods integrate global and local visual cues through the fusion of the 360° equirectangular images and corresponding 360° cube-map images. The fusion in a single level cannot effectively utilize the information between the 360° equirectangular images and corresponding 360° cube-map images. In this work, we innovatively propose a semantics and context feature aggregation network (SCFANet) by fully exploring the interactivity between the two projection data. Specifically, we use Vision Transformer (ViT) to capture global visual cues for 360° equirectangular images and Convolutional Neural Network (CNN) to capture local visual cues for 360° cube-map images. To achieve effective fusion of the two projection data, we design a semantic guidance module (SGM), in which semantic features are used to guide the information fusion of the 360° equirectangular images and corresponding 360° cube-map images at each level. Then, a context fusion module (CFM) containing one local input and two context inputs is designed to integrate multi-scale features, where the local input extracts its own multi-scale information, and the context inputs complements their fine details and location information. Finally, we use feature aggregation and refinement module (FARM) to aggregate semantics and context feature and adopt a deep supervision strategy for training. Extensive experiments on two public 360° datasets show that our SCFANet exhibits competitive performance compared to other state-of-the-art (SOTA) 360° salient object detection models. Feng Shao 0001, Xiongli Chai, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2024 | Progressive Bidirectional Feature Extraction and Enhancement Network for Quality Evaluation of Night-Time ImagesabstractBlind image quality assessment (BIQA) has received increasing attention in the past decades. However, it still remains inadequately researched on BIQA for night-time images suffering from the diverse authentic degradations. Since the intrinsic content degradations of night-time images are highly related to the illumination, how to use the connection between content and illumination to enhance the feature representation ability is the key issue in designing BIQA methods for night-time images. In this article, we first construct an ultra-high-definition night-time image dataset (UHD-NID) with high image resolution and abundant parameter settings. UHD-NID contains 1600 images with a high resolution of 5616 × 3744, and each group of images contains ten exposure levels. Then, we conduct subjective assessment and analyze the subjective data to obtain a mean opinion score to each image in UHD-NID. To enhance the feature representation ability in content and illumination, we propose a Progressive Bidirectional Feature Extraction and Enhancement Network (PBFEE-Net). In addition, we use a decomposition network to decompose the input image into the reflectance and illumination, which can facilitate the ability of feature extraction to some extent. The experimental results show that our proposed method achieves superior performance in evaluating the quality of night-time images. Jiangli Shi, Feng Shao 0001, Chongzhen Tian, Hangwei Chen, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2024 | Benchmark Dataset and Pair-Wise Ranking Method for Quality Evaluation of Night-Time Image EnhancementabstractNight-time image enhancement (NIE) aims at boosting the intensity of low-light regions while suppressing noises or light effects in night-time images, and numerous efforts have been made for this task. However, few explorations focus on the quality evaluation issue of enhanced night-time images (ENTIs), and how to fairly compare the performance of different NIE algorithms remains a challenging problem. In this paper, we firstly construct a new Real-world Night-Time Image Enhancement Quality Assessment (i.e., RNTIEQA) dataset that includes two typical types of night-time scenes (i.e., extremely low light and uneven light scenes), and carry out human subjective studies to compare the quality of ENTIs obtained by a set of representative NIE algorithms. Afterwards, a new objective ranking method that comprehensively considering image intrinsic and impairment attributes is proposed for automatically predicting the quality of ENTIs. Experimental results on our RNTIEQA dataset demonstrate that the proposed method outperforms the off-the-shelf competitors. Our dataset and code will be released athttps://github.com/Leilei-Huang-work/RNTIEQA-dataset. Xuejin Wang, Leilei Huang, Hangwei Chen, Qiuping Jiang, ShaoWei Weng, Feng Shao 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Collaborative Learning and Style-Adaptive Pooling Network for Perceptual Evaluation of Arbitrary Style TransferabstractAlthough the research of arbitrary style transfer (AST) has achieved great progress in recent years, few studies pay special attention to the perceptual evaluation of AST images that are usually influenced by complicated factors, such as structure-preserving, style similarity, and overall vision (OV). Existing methods rely on elaborately designed hand-crafted features to obtain quality factors and apply a rough pooling strategy to evaluate the final quality. However, the importance weights between the factors and the final quality will lead to unsatisfactory performances by simple quality pooling. In this article, we propose a learnable network, named collaborative learning and style-adaptive pooling network (CLSAP-Net) to better address this issue. The CLSAP-Net contains three parts, i.e., content preservation estimation network (CPE-Net), style resemblance estimation network (SRE-Net), and OV target network (OVT-Net). Specifically, CPE-Net and SRE-Net use the self-attention mechanism and a joint regression strategy to generate reliable quality factors for fusion and weighting vectors for manipulating the importance weights. Then, grounded on the observation that style type can influence human judgment of the importance of different factors, our OVT-Net utilizes a novel style-adaptive pooling strategy guiding the importance weights of factors to collaboratively learn the final quality based on the trained CPE-Net and SRE-Net parameters. In our model, the quality pooling process can be conducted in a self-adaptive manner because the weights are generated after understanding the style type. The effectiveness and robustness of the proposed CLSAP-Net are well validated by extensive experiments on the existing AST image quality assessment (IQA) databases. Our code will be released at https://github.com/Hangwei-Chen/CLSAP-Net. Hangwei Chen, Feng Shao 0001, Xiongli Chai, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Multi-layer and Multi-scale feature aggregation for DIBR-Synthesized image quality assessment
Xuejin Wang, Xiongli Chai, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | TCCL-Net: Transformer-Convolution Collaborative Learning Network for Omnidirectional Image Super-ResolutionabstractAs virtual reality and metaverse become more and more popular, the Omnidirectional Image (OI) has attracted extreme attention due to its immersive display characteristics. However, users only watch a portion of the content in a specific viewport extracted from a panoramic view, which will lead to a problem of resolution mismatch that requires High-Resolution (HR) for clear near-eye displays in viewports. Hence, it is necessary to exploit a Super Resolution (SR) solution for reconstructing Low-Resolution (LR) OIs. Different from 2D SR methods, the variation of pixel distributions along latitudes is a critical factor in designing an Omnidirectional Image Super-Resolution (OISR) scheme. In this paper, we put forward a novel end-to-end network with a Transformer and Convolution Collaborative Learning Network (TCCL-Net) for OISR. Firstly, Swin Transformer blocks and residual convolution blocks are employed to extract long-range and short-range dependencies, thereby digging into more rich and heterogeneous features from these two branches. Secondly, to better fuse these two features, cross-guided enhanced attention mechanisms are designed for bidirectional information enhancement onto both channel and spatial features . Thirdly, to alleviate nonuniformly pixel distributions across latitudes, we add an absolute positional encoding into Swin Transformer to represent patch weights at different positions and propose a tile-based panoramic reconstruction module to super-resolve various bands with different pixel sampling characteristics across latitudes. Experimental results on two available benchmark datasets demonstrate the superiority of the proposed approach over the state-of-the-art method in achieving OISR task. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Hongwei Ying |
Knowl. Based Syst. | 2 |
| 2023 | Jointly Texture Enhanced and Stereo Captured Network for Stereo Image Super-Resolution
Kangjun Jin, Xuejin Wang, Feng Shao 0001 |
Pattern Recognit. Lett. | 3 |
| 2023 | Modality-Induced Transfer-Fusion Network for RGB-D and RGB-T Salient Object DetectionabstractThe ability of capturing the complementary information of multi-modality data is critical to the development of multi-modality salient object detection (SOD). Most of existing studies attempt to integrate multi-modality information through various fusion strategies. However, most of these methods ignore the inherent differences in multi-modality data, resulting in poor performance when dealing with some challenging scenarios. In this paper, we propose a novel Modality-Induced Transfer-Fusion Network (MITF-Net) for RGB-D and RGB-T SOD by fully exploring the complementarity in multi-modality data. Specifically, we first deploy a modality transfer fusion (MTF) module to bridge the semantic gap between single and multi-modality data, and then mine the cross-modality complementarity based on point-to-point structural similarity information. Then, we design a cycle-separated attention (CSA) module to optimize the cross-layer information recurrently, and measure the effectiveness of cross-layer features through point-wise convolution-based multi-scale channel attention. Furthermore, we refine the boundaries in the decoding stage to obtain high-quality saliency maps with sharp boundaries. Extensive experiments on 13 RGB-D and RGB-T SOD datasets show that the proposed MITF-Net achieves a competitive and excellent performance. Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Quality Evaluation of Arbitrary Style Transfer: Subjective Study and Objective MetricabstractArbitrary neural style transfer is a vital topic with great research value and wide industrial application, which strives to render the structure of one image using the style of another. Recent researches have devoted great efforts on the task of arbitrary style transfer (AST) for improving the stylization quality. However, there are very few explorations about the quality evaluation of AST images, even it can potentially guide the design of different algorithms. In this paper, we first construct a new AST images quality assessment database (AST-IQAD), which consists 150 content-style image pairs and the corresponding 1200 stylized images produced by eight typical AST algorithms. Then, a subjective study is conducted on our AST-IQAD database, which obtains the subjective rating scores of all stylized images on the three subjective evaluations, i.e., content preservation (CP), style resemblance (SR), and overall vision (OV). To quantitatively measure the quality of AST image, we propose a new sparse representation-based method, which computes the quality according to the sparse feature similarity. Experimental results on our AST-IQAD have demonstrated the superiority of the proposed method. The dataset and source code will be released athttps://github.com/Hangwei-Chen/AST-IQAD-SRQE Hangwei Chen, Feng Shao 0001, Xiongli Chai, Yuese Gu, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Viewport-Sphere-Branch Network for Blind Quality Assessment of Stitched 360° Omnidirectional ImagesabstractCompared with conventional images/videos, omnidirectional data records rich information with higher resolution and wider Field-of-View. Moreover, the stitching distortions introduced in the panoramic content generation process make the quality assessment task more challenging. Targeting at designing an accurate and fast stitched 360° omnidirectional image quality evaluator, we propose a Viewport-Sphere-Branch Network (VSBNet) via dual-branch quality estimation. Specifically, for the viewport quality estimation, we extract distorted viewports around the stitching seams and conduct distortion rectification through a progressively complementary network to obtain pseudo-reference viewports. The qualitative and quantitative experiments validate that pseudo-reference viewports are reliable. Then, the differences between distorted and pseudo-reference viewports are quantified through transformer architecture to obtain quality scores of viewports. The introduction of pseudo-reference viewports can effectively improve the performance of the viewport quality prediction branch. To establish general scenario awareness and accurately evaluate the immersive experience, we extract feature representation through deformable convolutions to eliminate 2D-to-Sphere intrinsic sampling distortions and use multilayer perceptron to predict score of the whole sphere. The final prediction score is obtained by aggregating the quality scores from viewport and sphere branches. We evaluate the proposed VSBNet on two benchmark databases and results demonstrate that the combination of two branches can obtain more accurate results. Overall, our method is superior to existing full reference and no reference models designed for conventional images and 360° omnidirectional images. Chongzhen Tian, Feng Shao 0001, Xiongli Chai, Qiuping Jiang, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Cross-Modality Double Bidirectional Interaction and Fusion Network for RGB-T Salient Object DetectionabstractRGB-T salient object detection (SOD) aims to detect and segment saliency regions on RGB images and the corresponding thermal maps. The ability of alleviating the modality difference between RGB and thermal modality plays a vital role in the development of RGB-T SOD. However, most of the existing methods try to integrate multi-modal information through various fusion strategies, or reduce the modality difference via unidirectional or undifferentiated bidirectional interaction, but failing in some challenging scenes. To deal with the above question, a novel Cross-Modality Double Bidirectional Interaction and Fusion Network (CMDBIF-Net) for RGB-T SOD is proposed. Specifically, we construct an interactive branch to indirectly bridge the RGB and thermal modalities. In addition, we propose a double bidirectional interaction (DBI) module composed of a forward interaction block (FIB) and a backward interaction block (BIB) to reduce the cross-modality differences. Moreover, a multi-scale feature enhancement and fusion (MSFEF) module is introduced to integrate the multi-modal features with considering the internal gap of different modality. Finally, we use a cascaded decoder and a cross-level feature enhancement (CLFE) module to generate high-quality saliency map. Extensive experiments are conducted on three publicly available RGB-T SOD datasets shows that the proposed CMDBIF-Net achieves outstanding performance against the state-of-the-art (SOTA) RGB-T SOD methods. Zhengxuan Xie, Feng Shao 0001, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | FTDN: Multispectral and Hyperspectral Image Fusion With Diverse Temporal Difference SpansabstractMultispectral (MS)-hyperspectral (HS) image fusion, which aims to enhance the spatial resolution of low spatial resolution HS images with a high spatial resolution MS has provided a wide range of applications in remote sensing. However, relatively long revisit cycles of HS satellites and irresistible weather factors cause the acquisition of HS and MS images at the same time difficult. Most of the existing approaches neglect the temporal difference between MS and HS images, and perform weakness in the challenging case with diverse temporal difference spans. In this paper, we propose a novel image fusion strategy with embedding a stage of feature matching before interaction. On the one hand, we explore the role of spectral correlation modeling between HS and MS images, which accounts for the utilization of available spatial information from MS images. On the other hand, we design a feature aggregation module to fully exploit the nonlinear gaps and dependencies of heterogeneous data and utilize adaptive gains to realize complementary information projection and fusion. We build Dongying (DY) and Yellow River Estuary (YRE) remote sensing datasets based on Sentinel-2 and ZiYuan(ZY)-1 02D satellites with diverse temporal difference spans. The extensive experiments demonstrate that our method is robust to the span of temporal difference and shows superior performance over the existing methods visually and quantitatively. Xu Chen 0041, Xiangchao Meng, Qiang Liu 0035, Huiping Jiang, Gang Yang 0006, Weiwei Sun 0005, Feng Shao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | PSSTFN: A Progressive Spatial-Temporal-Spectral Fusion Network for Remote Sensing ImagesabstractSpatial-temporal-spectral fusion (STSF) is highly desirable to generate dense-time image series with high spatial and spectral resolution by integrating the complementary advantages of multi-source and multi-temporal observations. However, most existing STSF methods are still limited to the assumption of linear temporal, spatial and spectral relations. In addition, the STSF methods on Landsat and MODIS data are insufficient to characterize the inherent properties of the current spaceborne hyperspectral images with low spatial and temporal resolutions. For these, we propose a progressive STSF network (PSSTFN) by interestingly integrating spatial-spectral fusion and spectral-temporal fusion into a unified end-to-end STSF framework. Specifically, in the spatial-spectral fusion stage, we obtain the hierarchical features with different receptive fields and propose a multi-attention guided module for joint learning and refinement of spatial-spectral features. In the spectral-temporal fusion stage, a feature insertion module is presented to embed the difference images into the resulting spatial-spectral features, and the estimation from deeper layers is cascaded for more reliable spatial information. We build Dongying (DY) and Yellow River Estuary (YRE) remote sensing datasets based on Sentinel-2 and ZiYuan(ZY)-1 02D satellites for verification, and the experimental results on reduce- and full-resolution data demonstrate the superior performance of our method over the existing methods visually and quantitatively. Xu Chen 0041, Xiangchao Meng, Feng Shao 0001, Weiwei Sun 0005 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Dual-Task Interactive Learning for Unsupervised Spatio-Temporal-Spectral Fusion of Remote Sensing ImagesabstractSpatio-temporal-spectral fusion aims to produce high spatio-temporal-spectral resolution images by integrating the complementary spatial, temporal, and spectral advantages of multi-source remote sensing images. However, on one hand, existing spatio-temporal-spectral fusion methods are insufficient to exploit the inherent complex nonlinear spatial, temporal, and spectral relationship among multisource and multitemporal observations. On the other hand, since the unavailability of real high spatio-temporal-spectral resolution images, it is difficult to adopt deep learning methods with supervised training. In this paper, we propose an effective Unsupervised Spatio-Temporal-Spectral Fusion Model (USTSFM) with dual-task interactive learning to alleviate these problems. The proposed USTSFM has two branches: the Spatio-Temporal-Spectral Mapping (STSM) branch is to describe the temporal relationship, and the Spectral Super Resolution (SSR) branch is to model the spectral relationship. Moreover, the spatial-spectral interaction compensation block is designed to make the two branches compensate and benefited from each other. This intrinsically related and mutually facilitated strategy allows the USTSFM to sufficiently exploit the inherent spatial, temporal, and spectral relationship. In addition, a shared reconstruction module is meticulously designed for the two tasks, which not only reduces the parameters but also allows the supervised task to guide the convergence of the unsupervised task, boosting the stability of unsupervised training. The qualitative and quantitative results demonstrated the proposed USTSFM has richer spatial details and more accurate predictions than the other state-of-the-art methods. Qiang Liu 0035, Xu Chen 0041, Xiangchao Meng, Hangwei Chen, Feng Shao 0001, Weiwei Sun 0005 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Detail Injection-Based Spatio-Temporal Fusion for Remote Sensing Images With Land Cover ChangesabstractSpatio-temporal fusion can generate time-series images with high spatial resolution, and it is highly desirable in various applications, especially in monitoring fine dynamic changes of surface features on remote sensing images. Currently, most spatio-temporal fusion methods predict the target fine image by employing the auxiliary fine images on neighboring phases; however, they are generally limited in abrupt land cover changes between the target and the neighboring auxiliary images. In this paper, we propose a novel Detail Injection-based Spatio-Temporal Fusion (DISTF) model to alleviate this problem, by exploring the inherent relationship between the spatio-temporal fusion and spatio-spectral fusion. The proposed DISTF consists of three modules: a Three-branch Detail Injection (TDI) module, a Fine Detail Prediction (FDP) module, and a reconstruction module. The interpretable TDI module is inspired by spatio-spectral fusion, aiming to inject the non-changed detail information extracted from the neighboring fine images into the target coarse image, which can preserve the abrupt change information captured in the target coarse image. The FDP module is designed to further integrate the correlated information from the outputs of TDI and refine the spatial-spectral information to boost the fusion accuracy. Finally, the reconstruction module and the hybrid loss function are designed to more effective reconstruct the high-quality target fine image. The qualitative and quantitative experiment results on two datasets with different types of changes demonstrated that the proposed DISTF method achieves richer spatial detail and more accurate prediction than the eight existing methods. Qiang Liu 0035, Xiangchao Meng, Xinghua Li 0002, Feng Shao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Domain Adaptive Cross Reconstruction for Change Detection of Heterogeneous Remote Sensing Images via a Feedback Guidance MechanismabstractChange detection on heterogeneous optical and synthetic aperture radar (SAR) images is soaring and plays a crucial role in monitoring land cover changes, such as disaster emergencies and natural resource monitoring. This is commonly recognized as a promising but challenging work due to the intrinsic differences in imaging mechanisms between the optical and SAR images. Recently, deep learning-based change detection methods based on two-step processing have attracted attention, i.e., first image translation between optical and SAR images to alleviate their modality differences and then change detection based on the translated images. However, image translation itself is a trouble task for the heterogeneous optical and SAR images. The unreliable image translation results further limit the accuracy of change detection. In this paper, to mitigate this problem, we propose a change detection model on domain adaptation by novelty integrating change detection and image reconstruction into a unified framework. Specifically, we first transform the optical and SAR images into an intermediate common domain for comparison. Moreover, cross reconstruction for optical and SAR images is designed to maintain the characteristics of the images and improve the performance of domain adaptation. In addition, a feedback guidance mechanism is circumspectly designed to co-optimize change detection and image reconstruction tasks. Extensive experiments were conducted on four publicly available datasets, the results demonstrate the effectiveness of our proposed method. Qiang Liu 0035, Kai Ren 0003, Xiangchao Meng, Feng Shao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Perceptual Quality Assessment of Cartoon ImagesabstractIn the animation industry, automatically predicting the quality of cartoon images based on the inputs of general distortions and color change is an urgent task, while the existing no-reference (NR) methods usually measure the perceptual quality of the natural images. In this paper, based on the observation that structure and color are the main factors affecting cartoon images quality, we proposed a new NR quality prediction metric for cartoon images, which fully takes gradient and color information into account. The experimental results on our newly constructed NBU-CIQAD dataset with color change and other existing cartoon image dataset demonstrate that the proposed method significantly outperforms existing no-references methods for the task of cartoon image quality assessment. The database and code will be released athttps://github.com/1010075746/NBU-CIQAD. Hangwei Chen, Xiongli Chai, Feng Shao 0001, Xuejin Wang, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 3 |
| 2023 | Composition-Guided Neural Network for Image Cropping Aesthetic AssessmentabstractHow to explore the interaction between image aesthetic rules and crops is the key to finding views with good composition. Besides, it is subjective to evaluate candidate crops, which mainly depends on aesthetic knowledge, but it is not an easy task for people without extensive photography experience. However, existing methods mostly find good views by extracting general aesthetic features of crops without fully exploring the aesthetic rules. Motivated by this, we innovatively propose a composition-guided image cropping aesthetic assessment network (CGICAANet) for efficiently finding good crops and optimizing the cropping operation. Specifically, we adopt a direct and comprehensive composition pattern module, which adaptively mines suitable compositions for the images and emphasizes the dominant position of visual elements to contribute to optimizing the best crops in an interpretable way. Moreover, we designed a multi-task loss function to train the model. Particularly, to explore the commonality between predicted crops and labels, the complete intersection-over-union loss is adopted thoroughly considering the overlap area, central point distance and the consistency of aspect ratios for crops concurrently. Therefore, the predicted best crop can preserve the visual elements and have better composition. Experimental results with lightweight MobileNetV2 and ShuffleNetV2 as backbone networks demonstrate that our method can obtain comparable or better performance in terms of efficiency and accuracy. Shijia Ni, Feng Shao 0001, Xiongli Chai, Hangwei Chen, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2022 | M2OVQA: Multi-space signal characterization and multi-channel information aggregation for quality assessment of compressed omnidirectional videos
Xiongli Chai, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Deep network based stereoscopic image quality assessment via binocular summing and differencing
Jinbin Hu 0002, Xuejin Wang, Xiongli Chai, Feng Shao 0001, Qiuping Jiang |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | Integrated Fusion for Panchromatic, Multispectral, Hyperspectral Remote Sensing Images With Different Swath WidthsabstractZi Yuan (ZY)-1 02D satellite simultaneously provides the low spatial resolution (LR) and narrow swath-width hyperspectral (HS) image, the moderate spatial resolution (MR) multispectral (MS) image with a wider swath width, and the high spatial resolution (HR) panchromatic (PAN) image with the same wide swath width to the MR MS. How to comprehensively integrate their complementary advantages to obtain the wide swath-width and high-fidelity HR HS image is interesting but challenging. In this paper, we propose an integrated fusion method for the HR PAN, MR MS, and LR HS images with different swath widths, to generate the optimal wide swath-width HR HS image. The proposed method is based on the encoder-decoder learning framework. In the proposed fusion framework, a novel multi-branch encoder structure with an enhanced HS-encoder module and the multilevel spatial-spectral aggregation block is designed, by considering the difference in the spatial and spectral resolution among the multi-sensor images. The experiments on synthetic and real datasets from both qualitative and quantitative aspects demonstrated the competitive performance of the proposed method. Xiangjun Meng, Xiangchao Meng, Qiang Liu 0035, Jinfang Shu, Feng Shao 0001, Gang Yang 0006, Weiwei Sun 0005 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | SARF: A Simple, Adjustable, and Robust Fusion MethodabstractPansharpening aims to sharpen a low spatial resolution (LR) multispectral (MS) image using a high spatial resolution (HR) panchromatic (PAN) image to obtain the HR MS image. Though large numbers of pansharpening methods have been proposed, and many advanced methods have shown high quantitative results, few of them are widely used in real applications. This may be attributed to their instability for different images with different ground surface features, or the complexity to be implemented and the time-consuming process for some state-of-the-art methods. In this letter, we proposed a simple, adjustable, and robust fusion (SARF) method. In the proposed method, a spatial-spectral coenhanced strategy was proposed, and several details of the proposed fusion model were specifically designed for the “simple, adjustable, robust” features. It was tested and verified by four-band and eight-band MS images based on reduced resolution (RR) and full resolution (FR) experiments. The experimental results demonstrated the promising spatial visuality of the proposed method, and the spectral fidelity was more robust than most of component substitution (CS)-based and multiresolution analysis (MRA)-based methods. Xiangchao Meng, Gang Yang 0006, Feng Shao 0001, Weiwei Sun 0005, Huanfeng Shen, Shutao Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Monocular and Binocular Interactions Oriented Deformable Convolutional Networks for Blind Quality Assessment of Stereoscopic Omnidirectional ImagesabstractStereoscopic omnidirectional content, as a novel visual media, has drawn wide attention in recent years due to its ability in providing strong immersive experience. Since Stereoscopic Omnidirectional Images (SOIs) involve the properties from panoramic and stereoscopic visual perception, it is very challenging to establish an efficient and effective visual quality evaluation model for SOIs. To better measure the user’s experience in virtual reality, we put forward a novel deep learning framework to assess the quality of SOIs in this paper. Firstly, the deformable convolutions instead of standard convolutions are adopted to ensure the invariant receptive fields of convolutional kernels on Equi-Rectangular Projection (ERP). Secondly, according to the stereoscopic property, we use binocular-difference information and a coarse-to-fine mechanism to construct the binocular feature extraction network. Thirdly, a three-channel network involving left-view, right-view and binocular-difference channels is presented to simulate the process of monocular and binocular interactions, in which independent quality labels are provided for each channel to reflect the individual effect of monocular and binocular visions on the whole visual quality. Finally, experimental results on two available benchmark databases demonstrate the superiority of the proposed metric over the state-of-the-art blind quality assessment models in predicting the quality of SOIs. Moreover, our model is efficient in computational cost as the feature extraction is directly applied on ERP images. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | CGMDRNet: Cross-Guided Modality Difference Reduction Network for RGB-T Salient Object DetectionabstractHow to explore the interaction between the RGB and thermal modalities is the key success of the RGB-T saliency object detection (SOD). Most of the existing methods integrate multi-modality information by designing various fusion strategies. However, the modality gap between the RGB and thermal features will lead to unsatisfactory performances by simple feature concatenation. To solve this problem, we innovatively propose a cross-guided modality difference reduction network (CGMDRNet) to achieve intrinsic consistency feature fusion via reducing the modality differences. Specifically, we design a modality difference reduction (MDR) module, which is embedded in each layer of the backbone network. The module uses a cross-guided strategy to reduce the modality difference between the RGB and thermal features. Then, a cross-attention fusion (CAF) module is designed to fuse cross-modality features with small modality differences. In addition, we use a transformer-based feature enhancement (TFE) module to enhance the high-level feature representation that contributes more to performance. Finally, the high-level features guide the fusion of low-level features to obtain a saliency map with clear boundaries. Extensive experiments on three public RGB-T datasets show that the proposed CGMDRNet achieves competitive performance compared with state-of-the-art (SOTA) RGB-T SOD models. Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Underwater Image Enhancement Quality Evaluation: Benchmark Dataset and Objective MetricabstractDue to the attenuation and scattering of light by water, there are many quality defects in raw underwater images such as color casts, decreased visibility, reduced contrast,et al.. Many different underwater image enhancement (UIE) algorithms have been proposed to enhance underwater image quality. However, how to fairly compare the performance among UIE algorithms remains a challenging problem. So far, the lack of comprehensive human subjective user study with large-scale benchmark dataset and reliable objective image quality assessment (IQA) metric makes it difficult to fully understand the true performance of UIE algorithms. We in this paper make efforts in both subjective and objective aspects to fill these gaps. Firstly, we construct a new Subjectively-Annotated UIE benchmark Dataset (SAUD) which simultaneously provides real-world raw underwater images, readily available enhanced results by representative UIE algorithms, and subjective ranking scores of each enhanced result. Secondly, we propose an effective No-reference (NR) Underwater Image Quality metric (NUIQ) to automatically evaluate the visual quality of enhanced underwater images. Experiments on the constructed SAUD dataset demonstrate the superiority of our proposed NUIQ metric, achieving higher consistency with subjective rankings than 22 mainstream NR-IQA metrics. The dataset and source code will be made available athttps://github.com/yia-yuese/SAUD-Dataset. Qiuping Jiang, Yuese Gu, Chongyi Li, Runmin Cong, Feng Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | LGGD+: Image Retargeting Quality Assessment by Measuring Local and Global Geometric DistortionsabstractNumerous image retargeting algorithms have been proposed to achieve adaptive image resizing during the past years. To compare different image retargeting algorithms, reliable objective image retargeting quality assessment (IRQA) metrics are highly desired. Given that image retargeting usually introduces geometric distortions, this paper presents an objective IRQA metric by measuring both local and global geometric distortions (LGGD). Since human visual system perception is highly dependent on edges and the geometric distortions caused by image retargeting usually cause edge deformation, a sketch token-based local edge descriptor (ST-LED) is introduced to represent geometric-aware features in LGGD. First, ST-LED is first applied on both source and retargeted images for edge pattern representation. Second, pixel-level backward registration is conducted to enable estimating local geometric distortion (LGD) and a spatial pyramid-improved Bag-of-Token (BoT) model is built to enable estimating global geometric distortion (GGD). Since the proposed LGGD metric only focuses on geometric distortion while image retargeting quality is related with more aspects, we further fuse LGGD and an existing (EXT) IRQA metric to build a final version called LGGD+ for IRQA. Experiments on two benchmark databases demonstrate the superiority of LGGD+ and the excellent compatibility of our proposed LGGD for further improving a wide range of existing IRQA metrics (including both geometric distortion and non-geometric distortion metrics). In addition, the effectiveness of our LGGD metric is also demonstrated in another relevant task, i.e., quality evaluation of depth-image-based rendering (DIBR)-synthesized images, which also calls for accurate estimation of geometric distortion. Zhenyu Peng, Qiuping Jiang, Feng Shao 0001, Wei Gao 0003, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | VSOIQE: A Novel Viewport-Based Stitched 360° Omnidirectional Image Quality EvaluatorabstractWith the rapid development of virtual reality (VR), 360° omnidirectional images and videos have drawn wide attention. However, the quality assessment of 360° omnidirectional images is a challenging task, especially when the panoramic image contains multiple stitching distortions. We propose a viewport-based stitched 360° omnidirectional image quality evaluator (VSOIQE), by first extracting the features of salient and stitching viewports, and then inferring the overall perceptual quality via multiple linear regression (MLR). Comprehensive image attributes including edge, color, shape and information entropy are considered in the framework. Experimental results on two benchmark databases demonstrate the superiority of the proposed metric over both the state-of-the-art quality models designed for 2D images and the quality models developed for 360° omnidirectional images. Chongzhen Tian, Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | A Blind Full Resolution Assessment Method for Pansharpened Images Based on Multistream Collaborative LearningabstractPansharpening aims to fuse a high spatial resolution (HR) panchromatic (PAN) image and a low spatial resolution (LR) multispectral (MS) image to obtain an HR-MS image. However, due to the lack of the real HR-MS reference image, determining pansharpened image quality at full resolution has always been a contentious issue in the community. We propose a blind full resolution assessment method for pansharpened images based on multi-stream collaborative learning. The proposed method designs a Siamese framework to collaboratively learn the spatial, spectral, and overall quality of the fused image. The parameters of the feature extraction layer in the spatial and spectral evaluation models are frozen for the overall evaluation model, improving accuracy and convergence speed. The proposed method was comprehensively tested and verified based on a large-scale data set consisting of 13620 fused images obtained by six pansharpening methods with four different thematic data sets. Furthermore, a large-scale subjective evaluation data set in which each of the 13620 fused images was assessed by 28 participants, was utilized to comprehensively valid the proposed method. The experimental results demonstrated the superior performance of the proposed method to other state-of-the-art quality assessments. Kedi Bao, Xiangchao Meng, Xiongli Chai, Feng Shao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | PSTAF-GAN: Progressive Spatio-Temporal Attention Fusion Method Based on Generative Adversarial NetworkabstractSpatio-temporal fusion aims to integrate multisource remote sensing images with complementary high spatial and temporal resolutions, so as to obtain time-series high spatial resolution fused images. Currently, deep learning (DL)-based spatio-temporal fusion methods have received broad attention. However, on one hand, most of the existing DL-based methods train the model in a band-by-band manner, ignoring the correlations among bands. On the other hand, the general coarse spatio-temporal changes in low spatial resolution images (e.g., MODIS) calculated at the pixel domain cannot completely cover the fine spatio-temporal changes in high spatial resolution images (e.g., Landsat), due to complex surface features and the general large spatial resolution ratio between fine and coarse images. Besides, the existing DL-based spatio-temporal fusion methods are insufficient in exploring multiscale information by only stacking convolutional kernels with different sizes. To alleviate the above challenges, we propose a progressive spatio-temporal attention fusion model in a multiband training manner based on generative adversarial network (PSTAF-GAN). Specifically, we design a flexible multiscale feature extraction architecture to extract multiscale feature hierarchies. Then, spatio-temporal changes are calculated on the feature domain in different feature hierarchies. Besides, a spatio-temporal attention fusion architecture is proposed to fuse the spatio-temporal changes and ground details in a coarse-to-fine manner, which can explore multiscale information more sufficient and gradually recover the target image. The results of quantitative and qualitative experiments on two publicly available benchmark datasets show that the proposed PSTAF-GAN can achieve the best performance compared with the state-of-the-art methods. Qiang Liu 0035, Xiangchao Meng, Feng Shao 0001, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Blind Full-Resolution Quality Evaluation Method for PansharpeningabstractPansharpening methods have been developed for nearly 40 years; however, how to quantitatively evaluate the quality of pansharpened images at full resolution (FR) is probably the most debated topic in this field due to the inherent unavailable of the real HR MS reference image. In this article, a novel blind FR quality evaluation method for pansharpening is proposed. In the proposed method, spatial and spectral features that are sensitive to spatial and spectral distortions of fused images are comprehensively considered and jointly learned based on online multivariate Gaussian (MVG) to construct the evaluation model. It directly outputs the quality of fused images, rather than the stepwise evaluation of spectral score, spatial score, and final overall quality score by the weighted combination of them, which may introduce contradictory results. First, a pristine benchmark evaluation model is established on the spatial features from the original high-spatial-resolution (HR) panchromatic (PAN) image and the spectral invariant assumption between ideal fused and original multispectral (MS) images. Second, a testing evaluation model for the fused image is founded. Finally, the quality of the fused image is measured based on the distance between the testing and benchmark models. The experimental results demonstrated the superior performance of the proposed method. Furthermore, the proposed method can be generalized to other interesting tasks, such as the nonreference evaluation for pansharpening with missing information and the nonreference evaluation for hyperspectral image fusion. The source code is available onhttps://github.com/yyxhpkq/MQNR. Xiangchao Meng, Kedi Bao, Jinfang Shu, Bingzhong Zhou, Feng Shao 0001, Weiwei Sun 0005, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Spatio-Temporal-Spectral Collaborative Learning for Spatio-Temporal Fusion with Land Cover ChangesabstractSpatio-temporal fusion by combining the complementary spatial and temporal advantages of multi-source remote sensing images to obtain time-series high spatial resolution images is highly desirable in monitoring surface dynamics. Currently, deep learning (DL)-based fusion methods have received extensive attention. However, existing DL-based spatio-temporal fusion methods are generally limited in fusing the images with land cover changes. In this paper, we propose a spatio-temporal-spectral collaborative learning framework for spatio-temporal fusion to alleviate this problem. Specifically, the proposed method integrates the convolutional neural network and recurrent neural network into a unified framework, consisting of three sub-networks: multi-scale siamese convolutional neural network, multi-layer convolutional recurrent neural network, and adaptive weighting fusion network. The multi-scale siamese convolutional neural network has a flexible weight-sharing network to extract multi-scale spatial-spectral features from multi-source remote sensing images. The multi-layer convolutional recurrent neural network is constructed on the convolutional long-short term memory units to comprehensively learn the land cover changes by spatial, spectral, and temporal joint features. The adaptive weighting fusion network with a spatio-temporal-spectral change loss is proposed to further improve the interpretability and robustness. The experiments were performed on the publicly available benchmark datasets featured by phenology and land cover type changes, respectively. The experimental results demonstrated the competitive performance of the proposed method than other state-of-the-art fusion methods. Xiangchao Meng, Qiang Liu 0035, Feng Shao 0001, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Vision Transformer for PansharpeningabstractPansharpening is a fundamental and hot-spot research topic in remote sensing image fusion. In recent years, self-attention-based transformer has attracted considerable attention in natural language processing (NLP) and introduced to attend to computer vision (CV) tasks. Inspired by great success of the vision transformer (ViT) in image classification, we propose an improved and advanced purely transformer-based model for pansharpening. In the proposed method, stacked multispectral (MS) and panchromatic (PAN) images are cropped into patches (i.e., tokens), and after a three-layer self-attention-based encoder, these tokens contain rich information. After upsampled and stitched, a high spatial resolution (HR) MS image is finally obtained. Instead of convolutional neural networks (CNNs) pursuing a short-distance dependency, our proposed method aims to build up a long-distance dependency, to make full use of more useful features. The experiments were conducted on an opening benchmark dataset, including IKONOS with four-band MS/PAN images and WorldView-2 MS images featured by eight bands. In addition, the experiments were performed on reduced and full-resolution datasets from both qualitative and quantitative evaluation aspects. The experimental results indicate the competitive performance of the proposed model than other pansharpening methods, including the state-of-the-art pansharpening algorithms based on CNN. Xiangchao Meng, Feng Shao 0001, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Convolution-Embedded Vision Transformer With Elastic Positional Encoding for PansharpeningabstractTransformer, especially vision transformer (ViT), is attracting increasing attention in various computer vision (CV) tasks. However, two urgent problems exist for the ViT: 1) Owing to its attending to an image in the patch level, the vision transformer seems to have a better performance in fetching global representations but limited in extracting local features, which is an inherent advantage for the convolutional neural network (CNN); 2) the learnable positional encoding plays a positive role, but limits the cross-resolution ability of the network. Specifically, the pre-trained model could only generate images with the same size during training. To conquer the two problems, we propose a novel convolution-embedded vision transformer with elastic positional encoding in this paper. On one hand, we propose a joint CNN and self-attention network to collaboratively extract local and global features. On the other hand, we propose to integrate the elastic CNN-based positional encoder into the framework to solve the rigid limitation of the ViT in cross resolution issues and improve the performance. Extensive experiments were conducted on IKONOS and WorldView-2 with 4-band and 8-band multispectral images, respectively. The visual and numerical results show the competitive performance of the proposed method. Xiangjun Meng, Xiangchao Meng, Feng Shao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Single Image Super-Resolution Quality Assessment: A Real-World Dataset, Subjective Studies, and an Objective MetricabstractNumerous single image super-resolution (SISR) algorithms have been proposed during the past years to reconstruct a high-resolution (HR) image from its low-resolution (LR) observation. However, how to fairly compare the performance of different SISR algorithms/results remains a challenging problem. So far, the lack of comprehensive human subjective study on large-scale real-world SISR datasets and accurate objective SISR quality assessment metrics makes it unreliable to truly understand the performance of different SISR algorithms. We in this paper make efforts to tackle these two issues. Firstly, we construct a real-world SISR quality dataset (i.e., RealSRQ) and conduct human subjective studies to compare the performance of the representative SISR algorithms. Secondly, we propose a new objective metric, i.e., KLTSRQA, based on the Karhunen-Loéve Transform (KLT) to evaluate the quality of SISR images in a no-reference (NR) manner. Experiments on our constructed RealSRQ and the latest synthetic SISR quality dataset (i.e., QADS) have demonstrated the superiority of our proposed KLTSRQA metric, achieving higher consistency with human subjective scores than relevant existing NR image quality assessment (NR-IQA) metrics. The dataset and the code will be made available at https://github.com/Zhentao-Liu/RealSRQ-KLTSRQA. Qiuping Jiang, Ke Gu 0001, Feng Shao 0001, Xinfeng Zhang 0001, Hantao Liu, Weisi Lin |
IEEE Trans. Image Process. | 4 |
| 2022 | Toward Top-Down Just Noticeable Difference Estimation of Natural ImagesabstractJust noticeable difference (JND) of natural images refers to the maximum pixel intensity change magnitude that typical human visual system (HVS) cannot perceive. Existing efforts on JND estimation mainly dedicate to modeling the diverse masking effects in either/both spatial or/and frequency domains, and then fusing them into an overall JND estimate. In this work, we turn to a dramatically different way to address this problem with a top-down design philosophy. Instead of explicitly formulating and fusing different masking effects in a bottom-up way, the proposed JND estimation model dedicates to first predicting a critical perceptual lossless (CPL) counterpart of the original image and then calculating the difference map between the original image and the predicted CPL image as the JND map. We conduct subjective experiments to determine the critical points of 500 images and find that the distribution of cumulative normalized KLT coefficient energy values over all 500 images at these critical points can be well characterized by a Weibull distribution. Given a testing image, its corresponding critical point is determined by a simple weighted average scheme where the weights are determined by a fitted Weibull distribution function. The performance of the proposed JND model is evaluated explicitly with direct JND prediction and implicitly with two applications including JND-guided noise injection and JND-guided image compression. Experimental results have demonstrated that our proposed JND model can achieve better performance than several latest JND models. In addition, we also compare the proposed JND model with existing visual difference predicator (VDP) metrics in terms of the capability in distortion detection and discrimination. The results indicate that our JND model also has a good performance in this task. The code of this work are available at https://github.com/Zhentao-Liu/KLT-JND. Qiuping Jiang, Shiqi Wang 0001, Feng Shao 0001, Weisi Lin |
IEEE Trans. Image Process. | 4 |
| 2022 | Unsupervised Decomposition and Correction Network for Low-Light Image EnhancementabstractVision-based intelligent driving assistance systems and transportation systems can be improved by enhancing the visibility of the scenes captured in extremely challenging conditions. In particular, many low-image image enhancement (LIE) algorithms have been proposed to facilitate such applications in low-light conditions. While deep learning-based methods have achieved substantial success in this field, most of them require paired training data, which is difficult to be collected. This paper advocates a novel Unsupervised Decomposition and Correction Network (UDCN) for LIE without depending on paired data for training. Inspired by the Retinex model, our method first decomposes images into illumination and reflectance components with an image decomposition network (IDN). Then, the decomposed illumination is processed by an illumination correction network (ICN) and fused with the reflectance to generate a primary enhanced result. In contrast with fully supervised learning approaches, UDCN is an unsupervised one which is trained only with low-light images and corresponding histogram equalized (HE) counterparts (can be derived from the low-light image itself) as input. Both the decomposition and correction networks are optimized under the guidance of hybrid no-reference quality-aware losses and inter-consistency constraints between the low-light image and its HE counterpart. In addition, we also utilize an unsupervised noise removal network (NRN) to remove the noise previously hidden in the darkness for further improving the primary result. Qualitative and quantitative comparison results are reported to demonstrate the efficacy of UDCN and its superiority over several representative alternatives in the literature. The results and code will be made public available athttps://github.com/myd945/UDCN. Qiuping Jiang, Yudong Mao, Runmin Cong, Wenqi Ren, Chao Huang 0008, Feng Shao 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Cross-Modality Fusion and Progressive Integration Network for Saliency Prediction on Stereoscopic 3D ImagesabstractTraditional 2D image-based saliency prediction models suffer from unsatisfactory performance when dealing with stereoscopic 3D (S3D) images because eye movements in the case of freely viewing S3D images are demonstrated to be guided by both RGB and depth features. This paper studies the problem of saliency prediction on S3D images, where the interactions between RGB and depth modalities are both taken into account. Specifically, we design a novel deep neural network named Cross-modality Fusion and Progressive Integration Network (CFPI-Net) to address this problem. It consists of a Multi-level Cross-modality Feature Fusion (MCFF) module and a Multi-stage Progressive Feature Integration (MPFI) module. The MCFF module first captures hierarchical contexture features from each modality and then effectively fuses the hierarchical contexture features from different modalities at each level. The MPFI module involves multiple cascaded deeply supervised feature integration (DSFI) blocks in which the low-level and high-level cross-modality features are progressively integrated using the integrated features in the previous stage as a guidance. Our proposed CFPI-Net benefits from the advantages of multi-level feature representation, cross-modality feature fusion, and multi-stage progressive feature integration, which hereby fully boost the performance. Experimental results on two benchmark datasets demonstrate that CFPI-Net outperforms state-of-the-art saliency prediction methods both quantitatively and qualitatively. All the results and relevant codes will be made available to the public. Yudong Mao, Qiuping Jiang, Runmin Cong, Wei Gao 0003, Feng Shao 0001, Sam Kwong |
IEEE Trans. Multim. | 5 |
| 2022 | List-Wise Rank Learning for Stereoscopic Image Retargeting Quality AssessmentabstractStereoscopic imageretargeting (SIR) techniques attempt to display stereoscopic images on stereoscopic devices of various resolutions and aspect ratios to provide the users with better viewing experience. However, new quality perceptual problems emerge in the retargeted stereoscopic images generated by current SIR operators are quite different from those in the retargeted 2D images. In this paper, we dedicate to exploring the perceptual quality-related factors (e.g., shape preservation, object preservation and visual comfort.) of retargeted stereoscopic images, and propose a novel quality evaluation metric for SIR to achieve a more consistent evaluation with 3D perception and image degradation mechanism in the SIR process. Moreover, image quality features and 3D perceptual features are integrated into one representation for an overall perceptual quality prediction using a list-wise ranking approach, which gives priority to the ranking among the SIR results generated from the same stereoscopic source. Experimental results demonstrate that the proposed method outperforms most quality models developed for retargeted 2D/stereoscopic images. Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Xiongli Chai, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2022 | Combining Retargeting Quality and Depth Perception Measures for Quality Evaluation of Retargeted StereopairsabstractStereoscopic Image Retargeting (SIR) aims to adapt stereoscopic images and videos to 3D display devices with various aspect ratios by emphasizing the important content while retaining surrounding context with minimal visual distortion. To address the issue of SIR evaluation, this paper presents a new objective quality assessment method for retargeted stereopairs by combining image quality and depth perception measures. Specifically, the image quality measure is conducted between the source and retargeted intermediate views generated by the view synthesis method to characterize the geometric distortion and content loss of the retargeted stereopair, while several depth-aware features are extracted to measure the visual comfort/discomfort and depth sensation when human views a 3D scene. Then, the extracted features are integrated into an overall perceptual quality prediction. Experiment results on NBU SIRQA and SIRD databases verify the superiority of our method. Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Zhenqi Fu, Xiangchao Meng, Ke Gu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2021 | Stitched image quality assessment based on local measurement errors and global statistical properties
Chongzhen Tian, Xiongli Chai, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Quality assessment for color correction-based stitched images via bi-directional matching
Xuejin Wang, Xiongli Chai, Feng Shao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Two-Branch Deep Neural Network for Underwater Image Enhancement in HSV Color SpaceabstractDue to the influence of light absorption and scattering, underwater images usually suffer from quality deteriorations such as color cast and reduced contrast. The diverse quality degradations not only dissatisfy the user expectation but also lead to a significant performance drop in many underwater vision applications. This letter proposes a novel two-branch deep neural network for underwater image enhancement (UIE), which is capable of separately removing color cast and enhancing image contrast by fully leveraging useful properties of the HSV color space in disentangling chrominance and intensity. Specifically, the input underwater image is first converted into the HSV color space and disentangled into HS and V channels to serve as the input of the two branches, respectively. Then, the color cast removal branch enhances the H and S channels with a generative adversarial network architecture while the contrast enhancement branch enhances the V channel via a traditional convolutional neural network. The enhanced channels by the two branches are merged and converted back into RGB color space to obtain the final enhanced result. Experimental results demonstrate that, compared with state-of-the-art UIE methods, our method can produce much more visually pleasing enhanced results. Junkang Hu, Qiuping Jiang, Runmin Cong, Wei Gao 0003, Feng Shao 0001 |
IEEE Signal Process. Lett. | 5 |
| 2021 | Roundness-Preserving Warping for Aesthetic Enhancement-Based Stereoscopic Image EditingabstractImage editing is an effective solution to adapt contents for different applications. In this paper, we present a roundness-preserving warping model for stereoscopic image editing, in which energy constraints from image quality energy, aesthetics energy and depth adaptation energy are involved in the framework to solve the optimization. Specifically, to preserve object roundness during warping, the relationship between object's shape and disparity is established and is applied for depth adaptation. Different from the existing stereoscopic image editing methods, the main innovations of our method are to achieve a tradeoff in balancing information loss and reducing semantic distortion while providing a novel death adaptation model for recomposition and retargeting applications. Experimental results demonstrate the effectiveness of our method in enhancing the aesthetics of stereoscopic images. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | No-Reference Image Contrast Evaluation by Generating Bidirectional PseudoreferencesabstractThis article proposes a simple yet reliable no-reference image contrast evaluator (NICE) by generating bidirectional pseudoreferences (BPR). Different from the existing no-reference metrics that only operate on the contrast distorted image (CDI) itself, our proposed NICE-BPR measures the deviations of a CDI to its corresponding aggravated and enhanced counterparts (i.e., BPRs) in a hybrid feature space. Given a CDI, we first perform contrast aggravation and contrast enhancement using gamma correction and histogram equalization, respectively. Then, hybrid contrast-aware features are, respectively, extracted from the CDI and its corresponding BPRs via the analysis of histogram, entropy, and structure. The features obtained from the CDI are one-by-one compared with those from the BPRs to derive the bidirectional feature deviation vector. Finally, a quality predictor is built by learning a regression model to fuse the feature vector into a continuous quality score. Extensive experiments on several databases well-demonstrate the superiority of NICE-BPR. Qiuping Jiang, Zhenyu Peng, Guanghui Yue 0001, Feng Shao 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2021 | Subjective and Objective Quality Assessment for Stereoscopic Image RetargetingabstractBinocular stereoscopic image retargeting (SIR) aims to adjust 3D images into target aspect ratios. In recent years, various SIR methods have been proposed, but there are few researches on visual quality assessment. As a consequence, we construct a benchmark stereoscopic image retargeting quality assessment database (NBU-SIRQA), which contains 720 stereoscopic retargeted images generated by eight representative SIR operators. Subjective test is conducted to obtain the mean opinion score (MOS) for each stereoscopic retargeted image. Additionally, we propose an objective SIRQA metric based on grid deformation and information loss (GDIL). The main idea of GDIL is to decompose the SIR operator into two transformations: monocular image retargeting transformation and viewpoint transformation. In each transformation, grid deformation and information loss are extracted simultaneously to represent image quality and 3D perception quality. Experimental results validated on our established NBU-SIRQA database show the superiority of our metric in measuring the quality of stereoscopic retargeted images over the existing approaches. Zhenqi Fu, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2021 | Exploiting Local Degradation Characteristics and Global Statistical Properties for Blind Quality Assessment of Tone-Mapped HDR ImagesabstractTone mapping operators (TMOs) are developed to convert a high dynamic range (HDR) image into a low dynamic range (LDR) one for display with the goal of preserving as much visual information as possible. However, image quality degradation is inevitable due to the dynamic range compression during the tone-mapping process. This accordingly raises an urgent demand for effective quality evaluation methods to select a high-quality tone-mapped image (TMI) from a set of candidates generated by distinct TMOs or the same TMO with different parameter settings. A key element to the success of TMI quality evaluation is to extract effective features that are highly consistent with human perception. Towards this end, this paper proposes a novel blind TMI quality metric by exploiting both local degradation characteristics and global statistical properties for feature extraction. Several image attributes including texture, structure, colorfulness and naturalness are considered either locally or globally. The extracted local and global features are aggregated into an overall quality via regression. Experimental results on two benchmark databases demonstrate the superiority of the proposed metric over both the state-of-the-art blind quality models designed for synthetically distorted images (SDIs) and the blind quality models specifically developed for TMIs. Xuejin Wang, Qiuping Jiang, Feng Shao 0001, Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Measuring Coarse-to-Fine Texture and Geometric Distortions for Quality Assessment of DIBR-Synthesized ImagesabstractA synthesized view can be generated via Depth-Image-Based Rendering (DIBR) technique using one (or more) color images and the associated depth maps. However, several artifacts may occur in the synthesized views due to the imperfect color images, depth maps or texture inpainting techniques, which cannot be effectively estimated by the conventional quality metrics designed for natural images. In this paper, a new quality metric is proposed to evaluate DIBR-synthesized images by measuring texture and geometric distortions. The artifacts are first analyzed on different phases of the synthesis process, and the associated features are extracted to estimate the degree of texture and geometric distortions from both coarse and fine scales. Finally, individual quality scores are aggregated into an overall quality via regression. Experimental results on three publicly available DIBR datasets demonstrate the superiority of the proposed method over the state-of-the-art quality models. Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2021 | Transformation-Aware Similarity Measurement for Image Retargeting Quality Assessment via Bidirectional RewarpingabstractImage retargeting is an effective way to adapt images for target displays with different aspect ratios and sizes. Meanwhile, effective image retargeting quality assessment (IRQA) is important for optimizing the image retargeting operations. In this paper, we propose a transform-aware similarity (TRASIM) measurement metric for IRQA, including bidirectional geometric distortion measurement, bidirectional information loss measurement, and global salient structure distortion measurement. The main innovation of the TRASIM is to build a universal framework to establish the similarity transformation via bidirectional rewarping to simulate different types of retargeting operators. Based on the similarity transformation, geometric distortion and content loss are measured to determine the retargeting quality. Experimental results on two widely used databases (CUHK and RetargetMe) indicate that the proposed TRASIM has higher consistency with subjective ranks, compared with the state-of-the-art IRQA metrics. Feng Shao 0001, Zhenqi Fu, Qiuping Jiang, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Blind quality assessment for multiply distorted stereoscopic images towards IoT-based 3D capture systems
Xuejin Wang, Meiling Qi, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | A large-scale remote sensing database for subjective and objective quality assessment of pansharpened images
Yiming Xiong, Feng Shao 0001, Xiangchao Meng, Qiuping Jiang, Weiwei Sun 0005, Randi Fu, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | MSTGAR: Multioperator-Based Stereoscopic Thumbnail Generation With Arbitrary ResolutionabstractAt present, thumbnail generation for 2D images has been extensively studied, but the research in thumbnail generation for stereoscopic images is still relatively lacking. This paper presents a novel thumbnail generation technology for stereoscopic images based on multioperator with the following innovations: 1) The warping technique is used to retarget a stereopair into six-scale resolutions with different contexts, and the disparity is uniformly adjusted to a certain value based on just noticeable depth difference (JNDD) model, which overcomes the issues that 3D perception in stereoscopic thumbnail is uncontrollable and the sense of depth disappears in low-resolution stereoscopic images. 2) The six-scale images are cropped via cropping network, and are optimized to a target resolution based on the designed image visual representation energy. As a result, our method has better visual effect than state-of-the-art methods in generating thumbnail for stereoscopic display. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2019 | Simultaneous object size and depth adjustment for stereoscopic 3D images
Feng Shao 0001, Yanjia Fei, Randi Fu, Gangyi Jiang, Yo-Sung Ho |
Inf. Sci. | 1 |
| 2019 | Fast inter-frame prediction in multi-view video coding based on perceptual distortion threshold model
Gangyi Jiang, Baozhen Du, Shuqing Fang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
Signal Process. Image Commun. | 5 |
| 2019 | Authentically Distorted Image Quality Assessment by Learning From Empirical Score DistributionsabstractMost existing works on image quality assessment (IQA) focus on predicting a scalar quality score (SQS) based on the assumption that people can reach a consensus on the judgment of image quality. However, assigning a single scalar fails to reveal the subjective diversity that an image will probably receive divergent opinion scores from different subjects. This is particularly true for real-world authentically distorted images which usually involve composite mixtures of multiple distortions. To characterize such an property, this letter proposes to use a more informative vectorized label called empirical score distribution (ESD) to build an ESD-aided deep neural network (DNN) for authentically distorted image quality prediction. Our proposed network contains two streams: ESD prediction stream and SQS prediction stream. The whole DNN is optimized end-to-end with a combined loss so that both of the supervision information from ESD and SQS can be fully utilized in the training process. Experiments on two public authentically distorted image databases verify the superiority of our method. Qiuping Jiang, Zhenyu Peng, Sheng Yang 0006, Feng Shao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2019 | A Risk-Aware Pairwise Rank Learning Approach for Visual Discomfort Prediction of Stereoscopic 3DabstractFor visual discomfort prediction (VDP) of stereoscopic 3D images, a common two-stage framework is to first extract features that are predictive of the experienced visual discomfort level when viewing stereoscopic images and then use typical regression tools to learn the mapping from the extracted features to visual discomfort scores. Most existing approaches for stereoscopic 3D VDP focus on the former stage, i.e., feature extraction, while limited efforts have dedicated to exploiting more powerful and robust learning algorithms in this field. In this letter, inspired by the pairwise comparison-based subjective evaluation methodology, we propose a novel Risk-Aware Pairwise Rank Learning (RAPRL) approach to further improve the prediction accuracy. Unlike the traditional VDP approaches using different regression tools for feature-score mapping, our proposed RARL method addresses this problem based on a completely different pairwise rank learning framework with a risk-aware constraint. Experiments have verified the effectiveness and robustness of our proposed VDP model using RAPRL as the learning algorithm. Qiuping Jiang, Feng Shao 0001, Wei Gao 0003, Yo-Sung Ho |
IEEE Signal Process. Lett. | 2 |
| 2019 | BLIQUE-TMI: Blind Quality Evaluator for Tone-Mapped Images Based on Local and Global Feature AnalysesabstractHigh dynamic range (HDR) image, which has a powerful capacity to represent the wide dynamic range of real-world scenes, has been receiving attention from both academic and industrial communities. Although HDR imaging devices have become prevalent, the display devices for HDR images are still limited. To facilitate the visualization of HDR images in standard low dynamic range displays, many different tone mapping operators (TMOs) have been developed. To create a fair comparison of different TMOs, this paper proposes a BLInd QUality Evaluator to blindly predict the quality of Tone-Mapped Images (BLIQUE-TMI) without accessing the corresponding HDR versions. BLIQUE-TMI measures the quality of TMIs by considering the following aspects: 1) visual information; 2) local structure; and 3) naturalness. To be specific, quality-aware features related to the former two aspects are extracted in a local manner based on sparse representation, while quality-aware features related to the third aspect are derived based on global statistics modeling in both intensity and color domains. All the extracted local and global quality-aware features constitute a final feature vector. An emergent machine learning technique, i.e., extreme learning machine, is adopted to learn a quality predictor from feature space to quality space. The superiority of BLIQUE-TMI to several leading blind IQA metrics is well demonstrated on two benchmark databases. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Unified No-Reference Quality Assessment of Singly and Multiply Distorted Stereoscopic ImagesabstractA challenging problem in the no-reference quality assessment of multiply distorted stereoscopic images (MDSIs) is to simulate the monocular and binocular visual properties under a mixed type of distortions. Due to the joint effects of multiple distortions in MDSIs, the underlying monocular and binocular visual mechanisms have different manifestations with those of singly distorted stereoscopic images (SDSIs). This paper presents a unified no-reference quality evaluator for SDSIs and MDSIs by learning monocular and binocular local visual primitives (MB-LVPs). The main idea is to learn MB-LVPs to characterize the local receptive field properties of the visual cortex in response to SDSIs and MDSIs. Furthermore, we also consider that the learning of primitives should be performed in a task-driven manner. For this, two penalty terms including reconstruction error and quality inconsistency are jointly minimized within a supervised dictionary learning framework, generating a set of quality-oriented MB-LVPs for each single and multiple distortion modality. Given an input stereoscopic image, feature encoding is performed using the learned MB-LVPs as codebooks, resulting in the corresponding monocular and binocular responses. Finally, responses across all the modalities are fused with probabilistic weights which are determined by the modality-specific sparse reconstruction errors, yielding the final monocular and binocular features for quality regression. The superiority of our method has been verified on several SDSI and MDSI databases. Qiuping Jiang, Feng Shao 0001, Wei Gao 0003, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Image Process. | 2 |
| 2018 | 3D visual discomfort predictor based on subjective perceived-constraint sparse representation in 3D display system
Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Zongju Peng, Feng Shao 0001, Hao Jiang 0014 |
Future Gener. Comput. Syst. | 6 |
| 2018 | Perceptual stereoscopic image quality assessment method with tensor decomposition and manifold learningabstractPerceptual quality assessment of stereoscopic images is a challenge in three‐dimensional video systems. Existing studies suggest that simply averaging the quality of left and right views can effectively predict the quality of symmetrically distorted stereoscopic images, but prediction deviation occurs in the case of asymmetrically distorted stereoscopic images. Most previous stereoscopic image quality assessment (SIQA) methods have been based only on the luminance component of the images; in addition, the basis of human visual perception is critical to image quality assessment and lies on the low‐dimensional manifold. Inspired by this, a new perceptual SIQA method is proposed, which includes two stages: training stage and quality prediction stage. In the training stage, the authors apply Tucker decomposition to RGB images to reduce dimensions along colour channels to produce training sets, and the projection matrix is obtained through manifold learning. In the quality prediction stage, considering the binocular visual characteristics of visual perception, the overall stereoscopic estimate depends on the monocular image quality via a local energy ratio based pooling strategy and cyclopean based binocular quality. Extensive experiments on three available benchmark databases demonstrate that the proposed metric has better performance and achieves highly consistent alignment with subjective assessment compared with state‐of‐the‐art SIQA metrics. Gangyi Jiang, Meiling He, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
IET Image Process. | 4 |
| 2018 | Local and global sparse representation for no-reference quality assessment of stereoscopic images
Fucui Li, Feng Shao 0001, Qiuping Jiang, Randi Fu, Gangyi Jiang, Mei Yu 0001 |
Inf. Sci. | 2 |
| 2018 | No reference stereo video quality assessment based on motion feature in tensor decomposition domain
Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | Learning a referenceless stereopair quality engine with deep nonnegativity constrained sparse autoencoder
Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
Pattern Recognit. | 2 |
| 2018 | Toward Domain Transfer for No-Reference Quality Prediction of Asymmetrically Distorted Stereoscopic ImagesabstractWe have presented a no-reference quality prediction method for asymmetrically distorted stereoscopic images, which aims to transfer the information from source feature domain to its target quality domain using a label consistent K-singular value decomposition classification framework. To this end, we construct a category-deviation database for dictionary learning that assigns a label for each stereoscopic image to indicate if it is noticeable or unnoticeable by human eyes. Then, by incorporating a category consistent term into the objective function, we learn view-specific feature and quality dictionaries to establish a semantic framework between the source feature domain and the target quality domain. The quality pooling is comparatively simple and only needs to estimate the quality score based on the classification probability. The experimental results demonstrate the effectiveness of our blind metric. Feng Shao 0001, Zhuqing Zhang, Qiuping Jiang, Weisi Lin, Gangyi Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Learning Sparse Representation for Objective Image Retargeting Quality AssessmentabstractThe goal of image retargeting is to adapt source images to target displays with different sizes and aspect ratios. Different retargeting operators create different retargeted images, and a key problem is to evaluate the performance of each retargeting operator. Subjective evaluation is most reliable, but it is cumbersome and labor-consuming, and more importantly, it is hard to be embedded into online optimization systems. This paper focuses on exploring the effectiveness of sparse representation for objective image retargeting quality assessment. The principle idea is to extract distortion sensitive features from one image (e.g., retargeted image) and further investigate how many of these features are preserved or changed in another one (e.g., source image) to measure the perceptual similarity between them. To create a compact and robust feature representation, we learn two overcomplete dictionaries to represent the distortion sensitive features of an image. Features including local geometric structure and global context information are both addressed in the proposed framework. The intrinsic discriminative power of sparse representation is then exploited to measure the similarity between the source and retargeted images. Finally, individual quality scores are fused into an overall quality by a typical regression method. Experimental results on several databases have demonstrated the superiority of the proposed method. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
IEEE Trans. Cybern. | 2 |
| 2018 | Optimizing Multistage Discriminative Dictionaries for Blind Image Quality AssessmentabstractState-of-the-art algorithms for blind image quality assessment (BIQA) typically have two categories. The first category approaches extract natural scene statistics (NSS) as features based on the statistical regularity of natural images. The second category approaches extract features by feature encoding with respect to a learned codebook. However, several problems need to be addressed in existing codebook-based BIQA methods. First, the high-dimensional codebook-based features are memory-consuming and have the risk of over-fitting. Second, there is a semantic gap between the constructed codebook by unsupervised learning and image quality. To address these problems, we propose a novel codebook-based BIQA method by optimizing multistage discriminative dictionaries (MSDDs). To be specific, MSDDs are learned by performing the label consistent K-SVD (LC-KSVD) algorithm in a stage-by-stage manner. For each stage, a new quality consistency constraint called “quality-discriminative regularization” term is introduced and incorporated into the reconstruction error term to form a unified objective function, which can be effectively solved by LC-KSVD for discriminative dictionary learning. Then, the latter stage takes the reconstruction residual data in the former stage as input based on which LC-KSVD is repeatedly performed until the final stage is reached. Once the MSDDs are learned, multistage feature encoding is performed to extract feature codes. Finally, the feature codes are concatenated across all stages and aggregated over the entire image for quality prediction via regression. The proposed method has been evaluated on five databases and experimental results well confirm its superiority over existing relevant BIQA methods. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Ke Gu 0001, Gangyi Jiang, Huifang Sun |
IEEE Trans. Multim. | 2 |
| 2018 | Multistage Pooling for Blind Quality Prediction of Asymmetric Multiply-Distorted Stereoscopic ImagesabstractQuality prediction for asymmetric multiply-distorted stereoscopic images (MDSIs) confronts more challenges than previous stereoscopic image quality assessment (SIQA) issues, whereas the existing no-reference SIQA methods have been limited to understand the asymmetric distortions and multiple distortions simultaneously for general-purpose blind quality prediction. In this paper, we propose a multistage pooling (MUSP) model for quality prediction of asymmetric MDSIs. In the training stage, we establish multimodal sparse representation framework for phase and amplitude components, respectively. In the testing stage, we use an MUSP strategy to simulate the pooling procedure undergoing multimodal quality pooling, feature pooling, binocular pooling, and phase-amplitude quality pooling in order. Experimental results on our new established database (NBU-MDSID Phase-II) demonstrate the effectiveness of our blind metric. Feng Shao 0001, Qiuping Jiang, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Multim. | 1 |
| 2018 | No-Reference View Synthesis Quality Prediction for 3-D Videos Based on Color-Depth InteractionsabstractIn a 3-D video system, automatically predicting the quality of synthesized 3-D video based on the inputs of color and depth videos is an urgent but very difficult task, while the existing full-reference methods usually measure the perceptual quality of the synthesized video. In this paper, a high-efficiency view synthesis quality prediction (HEVSQP) metric for view synthesis is proposed. Based on the derived VSQP model that quantifies the influences of color and depth distortions and their interactions in determining the perceptual quality of 3-D synthesized video, color-involved VSQP and depth-involved VSQP indices are predicted, respectively, and are combined to yield an HEVSQP index. Experimental results on our constructed NBU-3D Synthesized Video Quality Database demonstrate that the proposed HEVSOP has good performance evaluated on the entire synthesized video-quality database, compared with other full-reference and no-reference video-quality assessment metrics. Feng Shao 0001, Qizheng Yuan, Weisi Lin, Gangyi Jiang |
IEEE Trans. Multim. | 1 |
| 2018 | Toward a Blind Quality Predictor for Screen Content ImagesabstractBlind quality assessment of screen content images (SCIs) is much challenging than traditional natural images. In this paper, we propose a blind quality predictor for SCIs to explore the issue from the perspective of sparse representation. Specifically, we conduct local sparse representation for the textual and pictorial regions, respectively, and conduct global sparse representation for the global SCIs. Subsequently, the underlying relationship between the feature and quality vectors is bridged in a universal sparse representation framework. The quality pooling is comparatively simply that only need to estimate the local and global quality scores and combine them to a total one. Our experimental results show that the proposed predictor can achieve better prediction performance to be in line with subjective assessment. Feng Shao 0001, Fucui Li, Gangyi Jiang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2017 | MSFE: Blind image quality assessment based on multi-stage feature encodingabstractBlind image quality assessment (BIQA) methods based on visual codebooks have received much attention due to its prominent generalization capacity across different image domains. Existing codebook-based BIQA methods depend on large-size codebooks and high-dimensional features, which are memory-consuming and have the risk of over-fitting. Thus, it is necessary to design quality metrics with much smaller codebooks. This paper presents a novel multistage feature encoding (MSFE)-based BIQA method which requires much lower dimensional features while preserving comparable or even better performance. To specify, MSFE is performed over multiple cascaded and much smaller sub-codebooks to generate more compact and discriminative features for quality prediction. The latter stage takes the encoding residuals in the former stage as input. We use KSVD and sparse coding for codebook training and feature encoding in the framework, respectively. Finally, the generated sparse feature codes in all stages are combined and aggregated over the entire image for quality prediction via support vector regression (SVR). We evaluate the proposed method on several natural and screen content image databases. The experimental results confirm its superiority in terms of both validity and universality. Qiuping Jiang, Feng Shao 0001, Gangyi Jiang |
ICIP | 2 |
| 2017 | A new tone-mapped image quality assessment approach for high dynamic range imaging systemabstractTone-mapping operators are designed to apply high dynamic range (HDR) images on widely-used low dynamic range (LDR) devices. Developing well-performed tone-mapped image quality assessment (IQA) method is highly desired because traditional IQA method cannot be adopted in cross dynamic range quality measuring. To this end, we proposed a quality assessment method based on image exposure property. Specifically, an image exposure property determination model is utilized to segment HDR image into different exposure region. Then, quality features are extracted according to the distortion characteristics of each exposure region. Finally, the quality of tone-mapped image can be acquired by a trained regression model. Validation experiments on public database show that the proposed method can accurately predict the quality of tone-mapped image. Yang Song 0015, Gangyi Jiang, Hao Jiang 0014, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
ICIP | 5 |
| 2017 | Visual comfort assessment for stereoscopic images based on sparse coding with multi-scale dictionaries
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
Neurocomputing | 2 |
| 2017 | Leveraging visual attention and neural activity for stereoscopic 3D visual comfort assessment
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
Multim. Tools Appl. | 2 |
| 2017 | Toward Simultaneous Visual Comfort and Depth Sensation Optimization for Stereoscopic 3-D ExperienceabstractVisual comfort and depth sensation are two important incongruent counterparts in determining the overall stereoscopic 3-D experience. In this paper, we proposed a novel simultaneous visual comfort and depth sensation optimization approach for stereoscopic images. The main motivation of the proposed optimization approach is to enhance the overall stereoscopic 3-D experience. Toward this end, we propose a two-stage solution to address the optimization problem. In the first layer-independent disparity adjustment process, we iteratively adjust the disparity range of each depth layer to satisfy with visual comfort and depth sensation constraints simultaneously. In the following layer-dependent disparity process, disparity adjustment is implemented based on a defined total energy function built with intra-layer data, inter-layer data and just noticeable depth difference terms. Experimental results on perceptually uncomfortable and comfortable stereoscopic images demonstrate that in comparison with the existing methods, the proposed method can achieve a reasonable performance balance between visual comfort and depth sensation, leading to promising overall stereoscopic 3-D experience. Feng Shao 0001, Weisi Lin, Zhutuan Li, Gangyi Jiang, Qionghai Dai |
IEEE Trans. Cybern. | 1 |
| 2017 | QoE-Guided Warping for Stereoscopic Image RetargetingabstractIn the field of stereoscopic 3D (S3D) display, it is an interesting as well as meaningful issue to retarget the stereoscopic images to the target resolution, while the existing stereoscopic image retargeting methods do not fully take user's Quality of Experience (QoE) into account. In this paper, we have presented a QoE-guided warping method for stereoscopic image retargeting, which retarget the stereoscopic image and adapt its depth range to the target display while promoting user's QoE. Our method takes shape preservation, visual comfort preservation, and depth perception preservation energies into account, and simultaneously optimizes the 2D coordinates and depth information in 3D space. It also considers the specific viewing configuration in the visual comfort and depth perception preservation energy constraints. Experimental results on visually uncomfortable and comfortable stereoscopic images demonstrate that in comparison with the existing stereoscopic image retargeting methods, the proposed method can achieve a reasonable performance optimization among the QoE's factors of image quality, visual comfort, and depth perception, leading to promising overall S3D experience. Feng Shao 0001, Wenchong Lin, Weisi Lin, Qiuping Jiang, Gangyi Jiang |
IEEE Trans. Image Process. | 1 |
| 2017 | Learning Sparse Representation for No-Reference Quality Assessment of Multiply Distorted Stereoscopic ImagesabstractBinocular combination under different distortion types poses a great challenge to three-dimensional image quality assessment (3D-IQA). However, the research works on 3D-IQA with multiple distortion types are very limited. In this paper, we first construct a new multiply distorted stereoscopic image database (NBU-MDSID), which is composed of 270 multiply distorted stereoscopic images and 90 singly distorted stereoscopic images that are corrupted simultaneously and independently by blurring, JPEG compression, and noise injection. We then propose a new multimodal blind metric for quality assessment of multiply distorted stereoscopic images. Inspired by multimodal sparse representation framework, modality-specific dictionaries and the corresponding projection matrices are learned from the singly distorted training database at the training stage, and the testing stage only needs to estimate the quality score based on the reconstruction errors. Experimental results demonstrate the effectiveness of our blind metric. Feng Shao 0001, Weijun Tian, Weisi Lin, Gangyi Jiang, Qionghai Dai |
IEEE Trans. Multim. | 1 |
| 2016 | Novel visibility threshold model for asymmetrically distorted stereoscopic imagesabstractExisting perceptual researches on stereoscopic images mainly focus on the threshold of whole image distortion, rather than the effect of texture feature on the so-called threshold of just-noticeable distortion. Obviously, it is unreasonable to use a single unified perception threshold for natural stereoscopic images as the texture complexity typically varies in different blocks of natural images. To solve this problem, we generated an asymmetrically distorted stereoscopic image database with different texture densities and conducted a large number of subjective experiments. A strong correlation between the asymmetrical visibility threshold and texture complexity was revealed from the subjective experiments. Finally, a nonlinear fitting model was designed to uncover this relationship, which can be applied to asymmetrical coding to control the perceived quality of stereoscopic images. Baozhen Du, Mei Yu 0001, Gangyi Jiang, Yun Zhang 0002, Feng Shao 0001, Zongju Peng, Tianzhi Zhu |
VCIP | 5 |
| 2016 | A depth video processing algorithm based on cluster dependent and corner-ware filtering
Zongju Peng, Mingsong Guo, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001 |
Neurocomputing | 6 |
| 2016 | Binocular perception based reduced-reference stereo video quality assessment method
Mei Yu 0001, Kaihui Zheng, Gangyi Jiang, Feng Shao 0001, Zongju Peng |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | A fast inter coding algorithm for HEVC based on texture and motion quad-tree models
Zongju Peng, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001 |
Signal Process. Image Commun. | 6 |
| 2016 | No-reference Stereoscopic Image Quality Assessment Using Binocular Self-similarity and Deep Neural Network
Yaqi Lv, Mei Yu 0001, Gangyi Jiang, Feng Shao 0001, Zongju Peng |
Signal Process. Image Commun. | 4 |
| 2016 | On Predicting Visual Comfort of Stereoscopic Images: A Learning to Rank Based ApproachabstractPredicting the degree of experienced visual comfort in the context of stereoscopic 3-D (S3D) viewing is particularly challenging. In this letter, a simple yet effective visual comfort assessment (VCA) approach for stereoscopic images is proposed from the perspective of learning to rank (L2R). The proposed L2R-based VCA (L2R-VCA) approach is inspired by the traditional absolute categorical rating (ACR) methodology in subjective study and is to characterize the qualitative description behavior of human subjective study. Experimental results on our recently built database confirm the promising performance of the proposed L2R-VCA approach, yielding higher consistency with human subject judgment results. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
IEEE Signal Process. Lett. | 2 |
| 2016 | Learning Receptive Fields and Quality Lookups for Blind Quality Assessment of Stereoscopic ImagesabstractBlind quality assessment of 3D images encounters more new challenges than its 2D counterparts. In this paper, we propose a blind quality assessment for stereoscopic images by learning the characteristics of receptive fields (RFs) from perspective of dictionary learning, and constructing quality lookups to replace human opinion scores without performance loss. The important feature of the proposed method is that we do not need a large set of samples of distorted stereoscopic images and the corresponding human opinion scores to learn a regression model. To be more specific, in the training phase, we learn local RFs (LRFs) and global RFs (GRFs) from the reference and distorted stereoscopic images, respectively, and construct their corresponding local quality lookups (LQLs) and global quality lookups (GQLs). In the testing phase, blind quality pooling can be easily achieved by searching optimal GRF and LRF indexes from the learnt LQLs and GQLs, and the quality score is obtained by combining the LRF and GRF indexes together. Experimental results on three publicly 3D image quality assessment databases demonstrate that in comparison with the existing methods, the devised algorithm achieves high consistent alignment with subjective assessment. Feng Shao 0001, Weisi Lin, Gangyi Jiang, Mei Yu 0001, Qionghai Dai |
IEEE Trans. Cybern. | 1 |
| 2016 | Toward a Blind Deep Quality Evaluator for Stereoscopic Images Based on Monocular and Binocular InteractionsabstractDuring recent years, blind image quality assessment (BIQA) has been intensively studied with different machine learning tools. Existing BIQA metrics, however, do not design for stereoscopic images. We believe this problem can be resolved by separating 3D images and capturing the essential attributes of images via deep neural network. In this paper, we propose a blind deep quality evaluator (DQE) for stereoscopic images (denoted by 3D-DQE) based on monocular and binocular interactions. The key technical steps in the proposed 3D-DQE are to train two separate 2D deep neural networks (2D-DNNs) from 2D monocular images and cyclopean images to model the process of monocular and binocular quality predictions, and combine the measured 2D monocular and cyclopean quality scores using different weighting schemes. Experimental results on four public 3D image quality assessment databases demonstrate that in comparison with the existing methods, the devised algorithm achieves high consistent alignment with subjective assessment. Feng Shao 0001, Weijun Tian, Weisi Lin, Gangyi Jiang, Qionghai Dai |
IEEE Trans. Image Process. | 1 |
| 2016 | Learning Blind Quality Evaluator for Stereoscopic Images Using Joint Sparse RepresentationabstractPerceptual quality prediction for stereoscopic images is of fundamental importance in determining the level of quality perceived by humans in terms of the 3D viewing experience. However, the existing no-reference quality assessment (NR-IQA) framework has its limitation in addressing binocular combination for stereoscopic images. In this paper, we propose a new NR-IQA for stereoscopic images using joint sparse representation. We analyze the relationship between left and right quality predictors, and formulate stereoscopic quality prediction as a combination of feature-prior and feature-distribution. Based on this finding, we extract feature vector that handles different features to be interacted by joint sparse representation, and use support vector regression to characterize feature-prior. Meanwhile, we implement feature-distribution using sparsity regularization as the basis of weights for binocular combination to derive the overall quality score. Experimental results on five public 3D IQA databases demonstrate that in comparison with the existing methods, the devised algorithm achieves high consistent alignment with subjective assessment. Feng Shao 0001, Kemeng Li, Weisi Lin, Gangyi Jiang, Qionghai Dai |
IEEE Trans. Multim. | 1 |
| 2015 | 3D Visual Comfort Assessment via Sparse Coding
Qiuping Jiang, Feng Shao 0001 |
ICIG (1) | 2 |
| 2015 | Difference of Gaussian statistical features based blind image quality assessment: A deep learning approachabstractNowadays, natural scene statistics (NSS) based blind image quality assessment (BIQA) models trained by machine learning, tend to achieve excellent performance. However, BIQA is still a very challenging research topic due to the lack of reference images. The key of further improvement lies in feature mining and pooling strategy decision. In this work, a new BIQA model is proposed to utilize local normalized multi-scale difference of Gaussian (DoG) response in distorted images as features which show a high correlation with perceptual quality. Then, a three-step-framework based deep neural network (DNN) is designed and employed as the pooling strategy. Compared with the support vector machine (SVM), the proposed three-step-framework DNN can excavate better feature representation, leading to more accurate predictions and stronger generalization ability. The proposed model achieves state-of-the-art performance on two authoritative databases and excellent generalization ability in cross database experiments. Yaqi Lv, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Feng Shao 0001 |
ICIP | 5 |
| 2015 | Supervised dictionary learning for blind image quality assessmentabstractIn this paper, we propose a supervised dictionary learning framework for blind image quality assessment (BIQA) by using quality-constraint sparse coding. Different with the traditional dictionary learning framework which only ensures the learnt dictionary accounting for image features, we add a quality-related regularization term in the framework to learn a feature-related dictionary and a quality-related dictionary jointly. Specifically, the feature-related and quality-related dictionaries share the same sparse coefficients, so that the reconstruction errors form the image feature vectors and quality score vectors are both minimized. Once the feature-related and quality-related dictionaries are learned, given a testing sample, we first abstract its feature vector and then compute the corresponding sparse coefficients w.r.t. the learnt feature-related dictionary, its quality score can be directly reconstructed based on the learnt quality-related dictionary and the estimated sparse coefficients. Experiment results on three publicly available IQA databases show the promising performance of the proposed model. Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
VCIP | 2 |
| 2015 | Supervised dictionary learning for blind image quality assessment using quality-constraint sparse coding
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Depth video spatial and temporal correlation enhancement algorithm based on just noticeable rendering distortion model
Zongju Peng, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 5 |
| 2015 | Binocular vision based objective quality assessment method for stereoscopic images
Gangyi Jiang, Junming Zhou, Mei Yu 0001, Yun Zhang 0002, Feng Shao 0001, Zongju Peng |
Multim. Tools Appl. | 5 |
| 2015 | A depth perception and visual comfort guided computational model for stereoscopic 3D visual saliency
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng, Changhong Yu |
Signal Process. Image Commun. | 2 |
| 2015 | Using Binocular Feature Combination for Blind Quality Assessment of Stereoscopic ImagesabstractThe quality assessment of 3D images is more challenging than its 2D counterparts, and little investigation has been dedicated to blind quality assessment of stereoscopic images. In this letter, we propose a novel blind quality assessment for stereoscopic images based on binocular feature combination. The prominent contribution of this work is that we simplify the process of binocular quality prediction as monocular feature encoding and binocular feature combination. Experimental results on two publicly available 3D image quality assessment databases demonstrate the promising performance of the proposed method. Feng Shao 0001, Kemeng Li, Weisi Lin, Gangyi Jiang, Mei Yu 0001 |
IEEE Signal Process. Lett. | 1 |
| 2015 | Full-Reference Quality Assessment of Stereoscopic Images by Learning Binocular Receptive Field PropertiesabstractQuality assessment of 3D images encounters more challenges than its 2D counterparts. Directly applying 2D image quality metrics is not the solution. In this paper, we propose a new full-reference quality assessment for stereoscopic images by learning binocular receptive field properties to be more in line with human visual perception. To be more specific, in the training phase, we learn a multiscale dictionary from the training database, so that the latent structure of images can be represented as a set of basis vectors. In the quality estimation phase, we compute sparse feature similarity index based on the estimated sparse coefficient vectors by considering their phase difference and amplitude difference, and compute global luminance similarity index by considering luminance changes. The final quality score is obtained by incorporating binocular combination based on sparse energy and sparse complexity. Experimental results on five public 3D image quality assessment databases demonstrate that in comparison with the most related existing methods, the devised algorithm achieves high consistency with subjective assessment. Feng Shao 0001, Kemeng Li, Weisi Lin, Gangyi Jiang, Mei Yu 0001, Qionghai Dai |
IEEE Trans. Image Process. | 1 |
| 2014 | Disparity based stereo image reversible data hidingabstractAs the popularity of three dimensional video, security of stereo image has become an evident issue to be solved. This paper presents a disparity based stereo image reversible data hiding by using histogram shifting, which can recover the original stereo image from marked stereo image without any distortion. Inter-correlations between left and right views of stereo image are utilized to predict pixels accurately. Then prediction error bins are constructed, and many points are around zero-valued bin for embedding data with low distortion of stereo images. The zero-valued bin is used twice to embed data, so that embedding capacity can reach more than 1 bit per pixel. Experimental results demonstrate that the proposed method outperforms the extended stereo image data hiding methods. Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
ICIP | 4 |
| 2014 | Stereo image watermarking scheme for authentication with self-recovery capability using inter-view reference sharing
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
Multim. Tools Appl. | 5 |
| 2014 | Reduced-reference stereoscopic image quality assessment based on view and disparity zero-watermarks
Wujie Zhou, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
Signal Process. Image Commun. | 4 |
| 2014 | PMFS: A Perceptual Modulated Feature Similarity Metric for Stereoscopic Image Quality AssessmentabstractStereoscopic image quality assessment (SIQA) is an important and challenging issue in three dimensional applications. In this letter, a perceptual modulated feature similarity (PMFS) metric for SIQA is proposed by considering the monocular and binocular perception properties. Specifically, stereoscopic image is first classified into monocular occlusion and binocular rivalry regions. Then, feature similarities between the original and distorted stereoscopic images are defined and measured for the monocular occlusion and binocular rivalry regions as the local monocular and binocular quality maps, respectively. Monocular and binocular just noticeable difference visual saliency models are presented to construct a modulation function to derive monocular and binocular quality scores. Finally, those scores are integrated into an overall quality score by support vector regression. Extensive experiments performed on LIVE phase II and MICT asymmetric databases demonstrate that the proposed PMFS metric can achieve much higher consistency with the subjective quality scores than some state-of-the-art SIQA metrics. Wujie Zhou, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
IEEE Signal Process. Lett. | 4 |
| 2013 | Perceptual Full-Reference Quality Assessment of Stereoscopic Images by Considering Binocular Visual CharacteristicsabstractPerceptual quality assessment is a challenging issue in 3D signal processing research. It is important to study 3D signal directly instead of studying simple extension of the 2D metrics directly to the 3D case as in some previous studies. In this paper, we propose a new perceptual full-reference quality assessment metric of stereoscopic images by considering the binocular visual characteristics. The major technical contribution of this paper is that the binocular perception and combination properties are considered in quality assessment. To be more specific, we first perform left-right consistency checks and compare matching error between the corresponding pixels in binocular disparity calculation, and classify the stereoscopic images into non-corresponding, binocular fusion, and binocular suppression regions. Also, local phase and local amplitude maps are extracted from the original and distorted stereoscopic images as features in quality assessment. Then, each region is evaluated independently by considering its binocular perception property, and all evaluation results are integrated into an overall score. Besides, a binocular just noticeable difference model is used to reflect the visual sensitivity for the binocular fusion and suppression regions. Experimental results show that compared with the relevant existing metrics, the proposed metric can achieve higher consistency with subjective assessment of stereoscopic images. Feng Shao 0001, Weisi Lin, Shanbo Gu, Gangyi Jiang, Thambipillai Srikanthan |
IEEE Trans. Image Process. | 1 |
| 2013 | Joint Bit Allocation and Rate Control for Coding Multi-View Video Plus Depth Based 3D VideoabstractIn three-dimensional (3D) video coding, distortion in texture video and depth maps can all affect the quality of the synthesized virtual views. Therefore, under the total bitrate constraint, effective bit allocation between texture and depth information is very important for 3D video coding. In this paper, the major technical contribution is to formulate view synthesis quality for optimal resource allocation in 3D video coding, since such quality is what that matters most to the ultimate user (i.e., the viewer) of the system; to be more specific, a new joint bit allocation and rate control method for multi-view video plus depth (MVD) based 3D video coding is proposed accordingly. We firstly derive a view synthesis distortion model to characterize the effect of coding distortion of texture video and depth maps on the synthesized virtual views. Based on this model, we derive a rate-distortion model to characterize the relationship between the bitrate and the view synthesis distortion, and the optimal bitrate ratio between texture and depth is established adaptively by solving the associated optimization problem. Finally, the rate control algorithm is performed on view level, texture/depth level and frame level. Experimental results show that compared with other methods, the proposed bit allocation method obtains higher performance of view synthesis. Moreover, the proposed rate control method can accurately control the bitrate to satisfy the total bitrate constraint. Feng Shao 0001, Gangyi Jiang, Weisi Lin, Mei Yu 0001, Qionghai Dai |
IEEE Trans. Multim. | 1 |
| 2012 | Depth map compression and depth-aided view rendering for a three-dimensional video systemabstractThree-dimensional (3D) video technologies are becoming increasingly popular, as they can provide high quality and immersive experience to end users, where depth maps are employed to generate the virtual views by depth-image-based rendering technique. However, how to reduce the compression and rendering complexities for depth maps while maintaining high rendering quality is still unresolved. In this study, a novel depth map compression and depth-aided view rendering method is proposed. In the proposed method, depth maps are represented with different layers and compressed with different macroblock-mode decision procedure, and several optimisation techniques, including spatio-temporal consistent warping, colour correction and temporal consistent hole filling are embedded into the view rendering framework. Experimental results show that compared with the traditional method, the proposed method can reduce more than 79% compression computational complexity and more than 45% rendering computational complexity, while maintaining high rendering quality. Feng Shao 0001, Mei Yu 0001, Gangyi Jiang, Fucui Li, Zongju Peng |
IET Signal Process. | 1 |
| 2012 | Asymmetric Coding of Multi-View Video Plus Depth Based 3-D Video for View RenderingabstractThe recent years have witnessed three-dimensional (3-D) video technology to become increasingly popular, as it can provide high-quality and immersive experience to end users, where view rendering with depth-image-based rendering (DIBR) technique is employed to generate the virtual views. Distortions in depth map may induce geometry changes in the virtual views, and distortions in texture video may be propagated to the virtual views. Thus, effective compression of both texture videos and depth maps is important for 3-D video system. From the perspective of bit allocation, asymmetric coding of the texture videos and depth maps is an effective way to get the optimal solution of 3-D video compression and view rendering problems. In this paper, a novel asymmetric coding method of multi-view video plus depth (MVD) based 3-D video is proposed on purpose of providing high-quality view rendering. In the proposed method, two models are proposed to characterize view rendering distortion and binocular suppression in 3-D video. Then, an asymmetric coding method of MVD-based 3-D video is proposed by combining two models in encoding framework. Finally, a chrominance reconstruction algorithm is presented to achieve accurate reconstruction. Experimental results show that compared with other methods, the proposed method can obtain higher performance of view rendering under the total bitrate constraint. Moreover, the perceptual visual quality of 3-D video is almost unaffected with the proposed method. Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Ken Chen 0003, Yo-Sung Ho |
IEEE Trans. Multim. | 1 |
| 2011 | A Novel Rate Control Algorithm for H.264/AVC Based on Human Visual System
Jiangying Zhu, Mei Yu 0001, Qiaoyan Zheng, Zongju Peng, Feng Shao 0001, Fucui Li, Gangyi Jiang |
PSIVT (2) | 5 |
| 2011 | Subjective quality analyses of stereoscopic images in 3DTV systemabstractSubjective quality evaluation is the basis of quality evaluation of stereoscopic images. As the lack of a public and diverse testing database currently, in this paper, a symmetric stereoscopic images database is built. And then the subjective quality of stereoscopic images is analyzed from two aspects, one is the effects of JPEG, JPEG2000, H.264. The other is the comparisons between symmetric and asymmetric stereoscopic images from Gaussian blurring, white Gaussian noise, JPEG and JPEG2000, respectively. The results show three compressions are quite different in the subjective quality of symmetric stereoscopic images at different bitrates, and the comparisons between symmetric and asymmetric stereoscopic images investigate the properties of binocular fusion, binocular suppression, and binocular summation. Junming Zhou, Gangyi Jiang, Xiangying Mao, Mei Yu 0001, Feng Shao 0001, Zongju Peng, Yun Zhang 0002 |
VCIP | 5 |
| 2010 | A Novel Rate Control Method for H.264/AVC Based on Frame Complexity and Importance
Haibing Chen, Mei Yu 0001, Feng Shao 0001, Zongju Peng, Fucui Li, Gangyi Jiang |
ACIVS (2) | 3 |
| 2010 | Asymmetric multi-view video coding based on chrominance reconstructionabstractThree-dimensional video (3DV) technology is becoming increasingly popular, as it can provide high quality and immersive experience to end users. Huge amount of data for storage and transmission is an important problem to be solved. In this paper, an asymmetric MVC method is proposed. Color correction is first performed as a preprocessing step to provide consistent color information among views. Then, all color corrected views are classified into color views and non-color views. The chrominance information in non-color views is all discarded and only preserved in color view in MVC codec. Thus, a large amount of coding bitrate can be saved. At the decoder, a chrominance reconstruction algorithm is presented to achieve accurate color reconstruction for those non-color views. Experimental results show that the proposed method can achieve large bitrate saving against the results compressed with the original JMVM codec. Moreover, the proposed method can obtain better reconstruction quality without noticeable quality degradation. Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Junyong You |
ICME | 1 |
| 2010 | Fast color correction for multi-view video by modeling spatio-temporal variation
Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 1 |
| 2007 | A Content-Adaptive Multi-View Video Color Correction AlgorithmabstractA content-adaptive color correction algorithm for multi-view video is proposed due to variation in lighting or camera parameters. We first establish color correction property between the target image and source image. Then color correction matrix can be obtained by global correction or preferred region matching correction. Finally, video tracking technique is used to correct multi-view video sequences. Experimental results show the proposed algorithm has better correction effect for different multi-view video sequences. Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Ken Chen 0003 |
ICASSP (1) | 1 |
| 2006 | Fast Multi-view Disparity Estimation for Multi-view Video Systems
Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, You Yang 0002 |
ACIVS | 3 |
| 2006 | Fast Adaptive Block Matching for Ray-Space Coding in FTV SystemabstractRay-space representation is the main approach to realizing free viewpoint television (FTV) with complicated scene. Data compression in ray-space is one of the key technologies in ray-space based FTV systems. In ray-space based FTV system, block matching complication is the most important factor to influence coding efficiency. In this paper, a fast adaptive block matching algorithm is proposed by using starting search point prediction, still block determination and search stop criteria strategies. Experimental results show that the search speed is improved greatly as well as the coding efficiency Mei Yu 0001, Feng Shao 0001, Gangyi Jiang |
ICASSP (2) | 2 |
| 2006 | Efficient Block Matching for Ray-Space Predictive Coding in Free-Viewpoint Television Systems
Gangyi Jiang, Feng Shao 0001, Mei Yu 0001, Ken Chen 0003, Tae Young Choi |
ICCSA (1) | 2 |
| 2006 | New Color Correction Approach to Multi-view Images with Region Correspondence
Gangyi Jiang, Feng Shao 0001, Mei Yu 0001, Ken Chen 0003, Xiexiong Chen |
ICIC (1) | 2 |
| 2006 | A New Image Correction Method for Multiview Video SystemabstractBecause of scene illumination or camera calibration, color appearance of the same object between different viewpoints may be different in multiview video system. Traditional illumination compensation algorithm for image is unable to solve this problem effectively. In this paper, a novel color correction method for multiview video system is proposed based on retinex color constancy theory. To eliminate influence of un-consistent light sources, histogram equalization, retinex processing and color restoration are performed for multiview images to extract reflectance that describes object intrinsic properties. Experimental results show that the proposed image correction method for multiview video system is effective Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Xiexiong Chen |
ICME | 1 |