Hangwei Chen

dblp:325/9991 · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
31since 2021 · last 2026
0000-0002-3756-2029ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 20 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 PU-TransMamba: A Hybrid Point Cloud Upsampling Framework With Detail-Aware Transformer and Spatially Coherent Mamba
abstract
High-quality 3D point clouds are essential for high-fidelity perception in Internet of Things (IoT)-enabled intelligent systems. While Point Cloud Upsampling (PCU) is widely used to mitigate data sparsity, existing methods often struggle to balance the preservation of fine-grained local details with the maintenance of global topological consistency. Transformer-based approaches frequently suffer from excessive computational overhead and high-frequency detail loss, whereas emerging state space models like Mamba, despite their efficiency, inevitably sacrifice spatial coherence due to the 1D serialization of irregular 3D points. To address these critical bottlenecks, we introduce PU-TransMamba, a hybrid framework that synergistically leverages a Detail-Aware Transformer and a Spatially-Coherent Mamba. Each component is designed to resolve specific PCU limitations: a Complexity-Aware Bilateral Decoder is developed to adaptively recover sharp geometric edges by processing features across dual domains, while a Sequence-Aligned Mamba Encoder utilizes multiple spatial curvature descriptors to compensate for the spatial information loss inherent in serialization. Additionally, a Global Geometry Injector and a Local Neighbor Injector are designed to ensure structural integrity by infusing holistic skeletal priors and neighborhood context, respectively. To minimize feature discrepancies between the hybrid branches, we also propose a Self-Distillation Loss. Extensive experiments on five benchmark datasets demonstrate that PU-TransMamba outperforms state-of-the-art methods in both reconstruction accuracy and computational scalability. The results confirm its ability to recover intricate geometries, indicating significant potential for IoT-driven 3D perception and communication systems.
Feng Shao 0001, Xiongli Chai, Hangwei Chen, Zhongjie Zhu, Zhiyi Mo
IEEE Internet Things J.4
2026 SGNet: A Structure-Guided Lightweight Network for VDT Salient Object Detection
abstract
Visual-Depth-Thermal (VDT) salient object detection (SOD) aims to jointly exploit RGB, depth and thermal cues to segment the most visually significant regions. However, most existing VDT SOD models are heavy in parameters and computational cost, limiting their deployment on real-world and edge devices. To tackle this, we propose SGNet, a structure-guided lightweight network for efficient VDT SOD. Specifically, we design a lightweight Tri-modal Fusion Module (TFM) to integrate three modalities at the semantic level, and a Shared Structure Extraction Module (SSEM) to extract common structural information from depth and thermal modalities. A Structure Refine Module (SRM) further injects the extracted structure into the deepest semantic features, while a Multiscale Feature Refinement Module (MFRM) progressively decodes multi-level features under deep supervision to produce saliency maps with clear boundaries. Benefiting from these modules, SGNet achieves competitive performance on the VDT2048 benchmark with only 5.51 M parameters and a real-time speed of 120 FPS at 320 × 320 resolution, surpassing state-of-the-art methods while remaining deployment-friendly.
Huizhi Wang, Feng Shao 0001, Xuebin Wei, Xiongli Chai, Hangwei Chen, Zhongjie Zhu
IEEE Internet Things J.5
2026 Edge-Embedded Bidirectional Interactive Network for Lightweight Surface Defect Detection
abstract
Convolutional neural networks (CNN)-based surface defect detection methods have achieved remarkable results. However, most existing approaches face two key challenges: first, they typically require substantial computational resources to learn rich features, making it difficult to balance performance and computational cost; second, when dealing with complex backgrounds and various defect shapes, they often lose edge details. To address these challenges, we propose an Edge-embedded Bidirectional Interactive Network (EBINet) for lightweight surface defect detection. Specifically, we introduce an edge generation module (EGM), which enhances feature details by interacting with low-level and high-level features, providing high-quality edge details for defect regions. In addition, we design a bidirectional interactive decoder (BID) that consists of a self-recognition module (SRM), cross-attention fusion module (CAFM), and multiscale edge embedded module (MEEM). This decoder gradually integrates features from different stages and thus can effectively capture inter-layer feature relationships to generate high-quality saliency maps. Extensive experiments conducted on four public defect datasets demonstrate that the lightweight EBINet offers strong competitiveness and superior performance, requiring only 3.57M parameters and 2.30G FLOPs for a 256×256 input image.
Feng Shao 0001, Dongze Jin, Baoyang Mu, Hangwei Chen
IEEE Trans Autom. Sci. Eng.5
2026 Toward a Completely Blind Attacker for No-Reference Image Quality Assessment Models
abstract
No-reference image quality assessment (NR-IQA) models are critically vulnerable to adversarial attacks, posing significant risks to downstream vision systems. However, existing attack methods suffer from high computational costs, reliance on Mean Opinion Score (MOS) annotations, and poor cross-model transferability. To overcome these limitations, we propose Degrade-to-OverReconstruct (DOR), a novel prior knowledge-driven black-box attack framework operating in a "completely blind" manner, requiring neither MOS labels nor surrogate models, inducing significant prediction bias solely based on distortion statistics. Specifically, DOR generates universal adversarial examples by first applying mild degradation to preserve global structure and then employing aggressive over-reconstruction using a Residual Denoising Diffusion Model (RDDM) to adaptively disrupt intrinsic Natural Scene Statistics (NSS)-a shared foundation across NR-IQA models. Extensive experiments on synthetic (LIVE, TID2013) and authentic (CLIVE) datasets demonstrate DOR's strong attack performance and superior transferability against leading NR-IQA models that cover diverse deep neural network architectures. Our work pioneers a diffusion model-based "completely blind" attack paradigm, offering a practical, MOS-free solution for adversarial robustness assessment of NR-IQA models in real-world deployments.
Xinyu Ruan, Hangwei Chen, Chao Huang 0008, Wenqi Ren, Qiuping Jiang
IEEE Trans. Image Process.2
2026 Quality Evaluation of AI-Generated Images: Subjective Study and Objective Methodology
abstract
In recent years, AI-Generated Images (AIGIs) have attracted significant attention and shown great potential in various applications, including entertainment, advertisement, education, and product design. Driven by this trend, various Text-to-Image (T2I) models are developed. However, the quality of AIGIs produced by these models varies widely, with many low-quality images failing to meet human aesthetic standards. Consequently, research into both subjective and objective Image Quality Assessment (IQA) methods for AIGIs is crucial. In this paper, we introduce a dataset called AIGI-IQAD, designed to enhance our understanding of human aesthetic preferences for AIGIs. The dataset contains 2,880 AIGIs generated by 8 T2I models using 360 deliberately designed text prompts. Further, we conducted subjective experiments to gather ratings from both aesthetic quality and text-image consistency. Building on this dataset, we propose a model named Question-guided Multimodal Interaction Network (QMI-Net) for evaluating AIGIs. QMI-Net assesses human preferences for AIGIs by focusing on both aesthetic quality and text-image consistency. Specifically, QMI-Net uses a question-answering approach to guide Multimodal Large Language Models (MLLMs) in generating detailed aesthetic and similarity information. The Visual and Aesthetic Feature Fusion Module (VAFFM) then fuses the aesthetic features with the visual features extracted by Contrastive Language-Image Pre-training (CLIP) to obtain more comprehensive aesthetic quality features. Comprehensive experiments demonstrate that state-of-the-art performance is achieved by QMI-Net on our AIGI-IQAD and three other public datasets.The AIGI-IQAD datasets and QMI-Net will be released athttps://github.com/ctxya1207/QMI-Net.
Feng Shao 0001, Hangwei Chen, Xuejin Wang, Qiuping Jiang
IEEE Trans. Multim.3
2025 Rethinking Lightweight RGB-Thermal Salient Object Detection With Local and Global Perception Network
abstract
RGB–thermal salient object detection (RGB-T SOD) aims to segment the most intriguing parts through RGB and thermal images. However, the high computational costs and large model sizes of existing methods prevent its deployment in edge computing platforms. To cope with it, a new lightweight RGB-T SOD method is proposed, namely, local and global perception network (LGPNet), which relieves the pressure of data transmission and processing and improves robustness in complicated surroundings. Specifically, high-level lightweight fusion (HLF) blocks and low-level lightweight fusion (LLF) blocks are designed to reduce the difference between two modalities and perform multimodal feature fusion. Unlike existing convolutional neural network-based lightweight SOD methods, HLF and LLF blocks combine the spatial inductive bias of convolutional neural networks with the global perception of Transformer, which is able to perform feature extraction and fusion with fewer parameters and larger receptive field. Experimental results show that our method is able to compete with the state-of-the-art RGB-T SOD methods while having only 7.35 learnable parameters[M], 6.40G floating-point operations and real-time speed (33 frames per second (FPS) on PyTorch framework and 224 FPS on TensorRT framework).
Dongze Jin, Feng Shao 0001, Zhengxuan Xie, Baoyang Mu, Hangwei Chen
IEEE Internet Things J.5
2025 MDGINet: Multi-frequency Dynamic Guidance and Interaction network for image denoising
Feng Shao 0001, Hangwei Chen, Xiongli Chai, Qiuping Jiang
Knowl. Based Syst.3
2025 Art Comes From Life: Artistic Image Aesthetics Assessment via Attribute Knowledge Amalgamation
abstract
Assessing the aesthetic quality and visual appeal of artworks has become one of the hotspots in current research. The existing artistic image aesthetics assessment (AIAA) methods directly learn aesthetics from images, while ignoring the impact of variations in visual attributes on human aesthetic perception, which hampers the further development of AIAA. To address this issue, this paper presents a new AIAA method based on attribute knowledge amalgamation, named AKA-Net. Specifically, we initially learn common attribute aesthetic rules (e.g., composition and color) through pre-training on natural aesthetic images. Then, we devise a multi-model amalgamation strategy based on contrastive learning to transfer different types of prior attribute knowledge into a single target model, enabling flexible and efficient aesthetic prediction. Finally, an attribute-aware feature enhancement module (AFEM) is introduced to better establish the relationship between aesthetic quality and attribute knowledge. Experimental results on three public benchmark AIAA databases demonstrate that the proposed AKA-Net outperforms the state-of-the-art AIAA metrics.
Hangwei Chen, Feng Shao 0001, Xiongli Chai, Baoyang Mu, Qiuping Jiang
IEEE Trans. Circuits Syst. Video Technol.1
2025 GADFNet: Geometric Priors Assisted Dual-Projection Fusion Network for Monocular Panoramic Depth Estimation
abstract
Panoramic depth estimation is crucial for acquiring comprehensive 3D environmental perception information, serving as a foundational basis for numerous panoramic vision tasks. The key challenge in panoramic depth estimation is how to address various distortions in 360° omnidirectional images. Most panoramic images are displayed as 2D equirectangular projections, which exhibit significant distortion, particularly with the severe fisheye effect near the equatorial regions. Traditional depth estimation methods for perspective images are unsuitable for such projections. On the other hand, cubemap projection consists of six distortion-free perspective images, allowing the use of existing depth estimation methods. However, the boundaries between faces of a cubemap projection introduce discontinuities, causing a loss of global information when using cube maps alone. In this work, we propose an innovative geometric priors assisted dual-projection fusion network (GADFNet) that leverages geometric priors of panoramic images and the strengths of both projection types to enhance the accuracy of panoramic depth estimation. Specifically, to better focus the network on key areas, we introduce a distortion perception module (DPM) and incorporate geometric information into the loss function. To more effectively extract global information from the equirectangular projection branch, we propose a scene understanding module (SUM), which captures features from different dimensions. Additionally, to achieve effective fusion of the two projections, we design a dual projection adaptive fusion module (DPAFM) to dynamically adjust the weights of the two branches during fusion. Extensive experiments conducted on four public datasets (including both virtual and real-world scenarios) demonstrate that our proposed GADFNet outperforms existing methods, achieving superior performance.
Chengchao Huang, Feng Shao 0001, Hangwei Chen, Baoyang Mu, Long Xu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 A Mutual Head Knowledge Distillation Framework for Lightweight RGB-T Crowd Counting
abstract
As an important technology in the fields of intelligent transportation and public safety, crowd counting that can obtain pedestrian flow information has attracted extensive attention from academic and industrial communities. However, existing RGB-T crowd counting methods cannot effectively balance the counting accuracy and computational complexity in practical applications. For this, we propose a Mutual Head Knowledge Distillation Framework (MHKDF) to obtain a lightweight RGB-T crowd counting network for efficient and accurate pedestrian number estimation. Specifically, to avoid the influence of parameter and structure differences between teacher and student networks on the distillation effect, we propose a Cooperative Mutual Knowledge Distillation (CMKD) strategy to comprehensively and dynamically transfer the crowd analysis ability of the complex teacher model (MHKDF-T) to the lightweight student model (MHKDF-S). In addition, the upper bound of the performance of the student network depends on the teacher model with high accuracy. Therefore, to take advantage of the complementary advantages of frequency domain and spatial domain feature fusion, we propose a Multi-Modal Spatial-Frequency Hybrid Fusion Module (MSFHFM) to futher improve counting accuracy of MHKDF-T. Comprehensive experiments on two RGB-T crowd counting datasets demonstrate that our MHKDF-S achieves competitive performance with only 5.68 FLOPs and 4.89M parameters. Our code will be released at https://github.com/BaoYangCC/MHKDF.
Baoyang Mu, Feng Shao 0001, Hangwei Chen, Xuejin Wang, Qiuping Jiang
IEEE Trans. Circuits Syst. Video Technol.3
2025 Adversarial Robust Salient Object Detection in Optical Remote Sensing Images With Implicit Feature Enhancement
abstract
Deep neural networks (DNNs) have achieved significant progress in optical remote sensing images salient object detection (ORSI-SOD) and are widely applied to various remote sensing image analysis tasks. However, few SOD models demonstrate robust performance under adversarial perturbations, which ultimately leads to a decline in detection accuracy. Moreover, most existing defense methods inject fixed Gaussian noise globally into the image. Although such approaches are easy to implement, they have several limitations in inaccurate uncertainty estimation and neglecting the unique characteristics of local salient regions. Furthermore, existing adversarial defense research rarely addresses the challenges specific to the ORSI-SOD task, leaving a gap in effective defense strategies. To tackle these issues, we propose a novel defense method, which enhances the adversarial robustness of ORSI-SOD models through implicit feature enhancement. The algorithm first proposes a two-stage strategy of reverse local noise search and forward global noise optimization, enhancing generalization ability by implicitly enhancing features to better simulate network uncertainty. Then, the algorithm proposes a global-guided texture information enhancement (GTIE) module for low-level features and a global-guided semantics information enhancement (GSIE) module for high-level features, focusing on strengthening low-level texture information and enhancing the model’s understanding of high-level contextual semantic features, respectively. This dual-module design effectively weakens the impact of adversarial noise, significantly improving the robustness and accuracy of object detection. Extensive experiments on three ORSI-SOD datasets demonstrate that our defense strategy better estimates the uncertainty, resulting in an average performance improvement of 23.2% in$F_{\beta } ^{\mathrm { max}}$and 34.1% in$E_{\xi } ^{\mathrm { max}}$across six ORSI-SOD models under five different adversarial attack methods. Our code will be released in the public repository athttps://github.com/kexi0714/IFe.
Feng Shao 0001, Xiangchao Meng, Hangwei Chen, Xiongli Chai, Zhiyi Mo
IEEE Trans. Geosci. Remote. Sens.4
2025 Cross-Modal Hierarchical Knowledge Distillation for Image Aesthetics Assessment
abstract
The field of image aesthetics assessment (IAA) is rapidly advancing due to its wide applications. However, relying solely on single-modal information for aesthetic evaluation presents inherent limitations. While multimodal IAA models incorporating user comments have achieved significant advancements, these comments are often unavailable due to privacy concerns and practical considerations, and they also introduce additional computational overhead during inference. To address this issue, we propose a cross-modal hierarchical knowledge distillation method, termed HKD-IAA, to enhance the performance of unimodal image models effectively. Specifically, HKD-IAA comprises four components: feature extraction, feature decomposition, hierarchical knowledge distillation, and dynamic decay. During training, we first decompose the extracted features into a weighted sum of basic aesthetic elements and their corresponding weights, thereby reducing the learning difficulty for the student model. Building on this, we design a new hierarchical knowledge distillation framework, which aligns features at the feature, relation, and response levels to effectively transfer the knowledge from the teacher model. Finally, we introduce a dynamic decay strategy to adjust the weight of the distillation loss, thereby enhancing the student model's learning effectiveness during training. Extensive experiments on two benchmark datasets validate that the proposed method achieves state-of-the-art performance using only visual modal data. Our code is available athttps://github.com/Hangwei-Chen/HKD-IAA.
Hangwei Chen, Feng Shao 0001, Weiyi Jing, Huizhi Wang, Qiuping Jiang
IEEE Trans. Multim.1
2025 Cross-Projection Distilling Knowledge for Omnidirectional Image Quality Assessment
Huixin Hu, Feng Shao 0001, Hangwei Chen, Xiongli Chai, Qiuping Jiang
IEEE Trans. Multim.3
2025 MISF-Net: Modality-Invariant and -Specific Fusion Network for RGB-T Crowd Counting
abstract
To accurately perform crowd counting, utilizing the complementary relationship between RGB and thermal images to analyze the crowd has become the focus of current research. Due to different imaging principles, multi-modal images often contain different contents, which are their modality-specific information. For example, RGB images contain more texture and color details, while thermal images contain thermal radiation information. Meanwhile, they also describe the same target content, e.g., crowds, which are modality-invariant. However, existing methods only design different modules to directly fuse RGB and thermal image features, which did not fully consider the above facts. In this paper, by analyzing the similarities and differences between multi-modal images, we propose a Modality-Invariant and -Specific Fusion Network (MISF-Net) for RGB-T Crowd Counting. Specifically, we design a modality decomposition and fusion module (MDFM), which decomposes RGB and thermal image features into modality-invariant and -specific features by using the similarity and difference supervision between multi-modal features. Besides, reconstruction supervision is also used to prevent network learning from generating bias. After that, different fusion strategies are applied to the invariant and specific features, respectively. In addition, to adapt to the variations in size of different pedestrians, we design a modality-invariant fusion module (MIFM). Finally, after the fusion decoder, MISF-Net can obtain a more accurate crowd density map. Comprehensive experiments on the RGB-T crowd counting dataset show that our MISF-Net can achieve competitive performance.
Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Hangwei Chen, Zhongjie Zhu, Qiuping Jiang
IEEE Trans. Multim.4
2024 CAFCNet: Cross-modality asymmetric feature complement network for RGB-T salient object detection
Dongze Jin, Feng Shao 0001, Zhengxuan Xie, Baoyang Mu, Hangwei Chen, Qiuping Jiang
Expert Syst. Appl.5
2024 Hallucinated-PQA: No reference point cloud quality assessment via injecting pseudo-reference features
Baoyang Mu, Feng Shao 0001, Hangwei Chen, Qiuping Jiang, Long Xu 0001, Yo-Sung Ho
Expert Syst. Appl.3
2024 Visual Prompt Multibranch Fusion Network for RGB-Thermal Crowd Counting
abstract
As population growth and urbanization continue, accurate crowd counting is increasingly important for public safety management and the Internet of Video Things (IOVT). However, RGB and thermal infrared (RGB-T) crowd counting still faces challenges in improving feature extraction capability for RGB streams and reducing multimodality differences. For this, we propose a visual prompt multibranch fusion network (VPMFNet) to tackle the above challenges. Specifically, to improve the ability of crowd analysis of the RGB stream in RGB-T crowd counting, through designing the prompt enhancement module, we take the prior features of head perception in the crowd as visual prompt cues to embed into the RGB stream. In terms of RGB and thermal image feature fusion, we fully reduce the modality differences from the perspectives of local fusion, global fusion, and multireceptive field fusion to accurately estimate the pedestrian number. Various experiments on two RGB-T crowd counting data sets demonstrate that our VPMFNet achieves a smaller estimation error in the number of pedestrians. Besides, our VPMFNet outperforms existing methods (i.e., multicolumn convolutional neural network, BL, SANet, UCNet, HDFNet, BBSNet, BL+IDAM, BL+CSCA, dual-branch enhanced feature fusion network, and GETANet) on the RGB-D data set. Our code will be released athttps://github.com/QSBAOYANGMU/VPMFNet.
Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Hangwei Chen, Qiuping Jiang, Yo-Sung Ho
IEEE Internet Things J.4
2024 Blind cartoon image quality assessment based on local structure and chromatic statistics
Hangwei Chen, Xuejin Wang, Feng Shao 0001
J. Vis. Commun. Image Represent.1
2024 Hybrid CNN-transformer based meta-learning approach for personalized image aesthetics assessment
Xingao Yan, Feng Shao 0001, Hangwei Chen, Qiuping Jiang
J. Vis. Commun. Image Represent.3
2024 Plain-PCQA: No-Reference Point Cloud Quality Assessment by Analysis of Plain Visual and Geometrical Components
abstract
In reviewing the research progress in Point Cloud Quality Assessment (PCQA), two main pathways have emerged, i.e., 2D projections and 3D point descriptors. The former primarily focuses on visual information, while the latter concentrates on crucial geometrical information in three-dimensional space. However, the current studies lack a thorough investigation of the impact of visual components and seldom pay special attention to plane-point fusion strategies. To comprehensively represent features and effectively tackle various types of impairments, we propose an end-to-end learning paradigm, only considering plain visual and geometrical factors called Plain-PCQA, for quantitatively evaluating objective metrics of 3D dense point clouds associated with human perception. Firstly, we explore a sophisticated preprocessing technique. The entire point clouds are packaged into six projections by moving virtual cameras, which can conveniently increase the visual samples during the training stage. Given the high resolution of the projected image, we have opted for a relatively lightweight network, namely ResNet-18, as the backbone to enable higher resolution input data. Five cropped patches from the projected image are collectively fed into this network. In light of the presence of some invalid information in the projections, a mask weight is devised to calculate the significance of each patch based on its effective informational content. Secondly, dual neural networks, comprising of a No-Reference (NR) branch and a Degraded-Reference (DR) branch, are designed with fundamental visual components to provide quantitative quality metrics. Specifically, the NR branch utilizes the feature output of each block in the Vision Transformer (ViT) model to obtain long-range low-level and high-level visual NR quality. The DR branch employs KLT (Karhunen-Loève Transform) to acquire the principal component information of an image as the macro-structural image, and then feeds the difference between input images and macro-structural images into a network for DR quality extraction. Thirdly, a Plane-Point Interaction Transformer (P2IT) is presented by incorporating texture and semantic features in 2D projections and geometrical features in 3D spaces to characterize the complete features with a connected 2D-3D feature representation. With these elaborately designed deep features, the proposed model can achieve competitive performances relying solely on plain visual and geometrical components. The experimental results demonstrate the potential of the proposed approach in multiple representative databases, which surpasses existing state-of-the-art methods significantly.
Xiongli Chai, Feng Shao 0001, Baoyang Mu, Hangwei Chen, Qiuping Jiang, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.4
2024 Dynamic Weighted Fusion and Progressive Refinement Network for Visible-Depth-Thermal Salient Object Detection
abstract
The introduction of depth/thermal modality has significantly enhanced the performance of dual-modal salient object detection (SOD) methods. However, depth maps and thermal images are prone to environmental interference, making them insufficient for providing salient information. To address this challenge, triple-modal SOD methods have been proposed. However, these methods often overlook the detrimental effects of defective modalities during fusion, leading to subpar performance. To tackle this issue, we present a novel dynamic weighted fusion and progressive refinement network (DWFPRNet) for Visible-Depth-Thermal (V-D-T) SOD. Specifically, we first use the dual-modal fusion module (DFM) to fuse dual modalities, thereby obtaining fused features. Subsequently, the modality selective fusion module (MSFM) mines complementary information between fused features, considering both fusion features and the quality of feature maps, to achieve weighted fusion. Finally, we design a progressive refinement decoder (PRD) to realize interaction and multi-scale learning among different scale features and generate high-quality saliency maps. Extensive experiments conducted on the VDT-2048 public dataset demonstrate that our method outperforms existing state-of-the-art multi-modal methods.
Feng Shao 0001, Baoyang Mu, Hangwei Chen, Qiuping Jiang
IEEE Trans. Circuits Syst. Video Technol.4
2024 Progressive Bidirectional Feature Extraction and Enhancement Network for Quality Evaluation of Night-Time Images
abstract
Blind image quality assessment (BIQA) has received increasing attention in the past decades. However, it still remains inadequately researched on BIQA for night-time images suffering from the diverse authentic degradations. Since the intrinsic content degradations of night-time images are highly related to the illumination, how to use the connection between content and illumination to enhance the feature representation ability is the key issue in designing BIQA methods for night-time images. In this article, we first construct an ultra-high-definition night-time image dataset (UHD-NID) with high image resolution and abundant parameter settings. UHD-NID contains 1600 images with a high resolution of 5616 × 3744, and each group of images contains ten exposure levels. Then, we conduct subjective assessment and analyze the subjective data to obtain a mean opinion score to each image in UHD-NID. To enhance the feature representation ability in content and illumination, we propose a Progressive Bidirectional Feature Extraction and Enhancement Network (PBFEE-Net). In addition, we use a decomposition network to decompose the input image into the reflectance and illumination, which can facilitate the ability of feature extraction to some extent. The experimental results show that our proposed method achieves superior performance in evaluating the quality of night-time images.
Jiangli Shi, Feng Shao 0001, Chongzhen Tian, Hangwei Chen, Long Xu 0001, Yo-Sung Ho
IEEE Trans. Multim.4
2024 Benchmark Dataset and Pair-Wise Ranking Method for Quality Evaluation of Night-Time Image Enhancement
abstract
Night-time image enhancement (NIE) aims at boosting the intensity of low-light regions while suppressing noises or light effects in night-time images, and numerous efforts have been made for this task. However, few explorations focus on the quality evaluation issue of enhanced night-time images (ENTIs), and how to fairly compare the performance of different NIE algorithms remains a challenging problem. In this paper, we firstly construct a new Real-world Night-Time Image Enhancement Quality Assessment (i.e., RNTIEQA) dataset that includes two typical types of night-time scenes (i.e., extremely low light and uneven light scenes), and carry out human subjective studies to compare the quality of ENTIs obtained by a set of representative NIE algorithms. Afterwards, a new objective ranking method that comprehensively considering image intrinsic and impairment attributes is proposed for automatically predicting the quality of ENTIs. Experimental results on our RNTIEQA dataset demonstrate that the proposed method outperforms the off-the-shelf competitors. Our dataset and code will be released athttps://github.com/Leilei-Huang-work/RNTIEQA-dataset.
Xuejin Wang, Leilei Huang, Hangwei Chen, Qiuping Jiang, ShaoWei Weng, Feng Shao 0001
IEEE Trans. Multim.3
2024 Collaborative Learning and Style-Adaptive Pooling Network for Perceptual Evaluation of Arbitrary Style Transfer
abstract
Although the research of arbitrary style transfer (AST) has achieved great progress in recent years, few studies pay special attention to the perceptual evaluation of AST images that are usually influenced by complicated factors, such as structure-preserving, style similarity, and overall vision (OV). Existing methods rely on elaborately designed hand-crafted features to obtain quality factors and apply a rough pooling strategy to evaluate the final quality. However, the importance weights between the factors and the final quality will lead to unsatisfactory performances by simple quality pooling. In this article, we propose a learnable network, named collaborative learning and style-adaptive pooling network (CLSAP-Net) to better address this issue. The CLSAP-Net contains three parts, i.e., content preservation estimation network (CPE-Net), style resemblance estimation network (SRE-Net), and OV target network (OVT-Net). Specifically, CPE-Net and SRE-Net use the self-attention mechanism and a joint regression strategy to generate reliable quality factors for fusion and weighting vectors for manipulating the importance weights. Then, grounded on the observation that style type can influence human judgment of the importance of different factors, our OVT-Net utilizes a novel style-adaptive pooling strategy guiding the importance weights of factors to collaboratively learn the final quality based on the trained CPE-Net and SRE-Net parameters. In our model, the quality pooling process can be conducted in a self-adaptive manner because the weights are generated after understanding the style type. The effectiveness and robustness of the proposed CLSAP-Net are well validated by extensive experiments on the existing AST image quality assessment (IQA) databases. Our code will be released at https://github.com/Hangwei-Chen/CLSAP-Net.
Hangwei Chen, Feng Shao 0001, Xiongli Chai, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Neural Networks Learn. Syst.1
2023 Modality-Induced Transfer-Fusion Network for RGB-D and RGB-T Salient Object Detection
abstract
The ability of capturing the complementary information of multi-modality data is critical to the development of multi-modality salient object detection (SOD). Most of existing studies attempt to integrate multi-modality information through various fusion strategies. However, most of these methods ignore the inherent differences in multi-modality data, resulting in poor performance when dealing with some challenging scenarios. In this paper, we propose a novel Modality-Induced Transfer-Fusion Network (MITF-Net) for RGB-D and RGB-T SOD by fully exploring the complementarity in multi-modality data. Specifically, we first deploy a modality transfer fusion (MTF) module to bridge the semantic gap between single and multi-modality data, and then mine the cross-modality complementarity based on point-to-point structural similarity information. Then, we design a cycle-separated attention (CSA) module to optimize the cross-layer information recurrently, and measure the effectiveness of cross-layer features through point-wise convolution-based multi-scale channel attention. Furthermore, we refine the boundaries in the decoding stage to obtain high-quality saliency maps with sharp boundaries. Extensive experiments on 13 RGB-D and RGB-T SOD datasets show that the proposed MITF-Net achieves a competitive and excellent performance.
Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.4
2023 Quality Evaluation of Arbitrary Style Transfer: Subjective Study and Objective Metric
abstract
Arbitrary neural style transfer is a vital topic with great research value and wide industrial application, which strives to render the structure of one image using the style of another. Recent researches have devoted great efforts on the task of arbitrary style transfer (AST) for improving the stylization quality. However, there are very few explorations about the quality evaluation of AST images, even it can potentially guide the design of different algorithms. In this paper, we first construct a new AST images quality assessment database (AST-IQAD), which consists 150 content-style image pairs and the corresponding 1200 stylized images produced by eight typical AST algorithms. Then, a subjective study is conducted on our AST-IQAD database, which obtains the subjective rating scores of all stylized images on the three subjective evaluations, i.e., content preservation (CP), style resemblance (SR), and overall vision (OV). To quantitatively measure the quality of AST image, we propose a new sparse representation-based method, which computes the quality according to the sparse feature similarity. Experimental results on our AST-IQAD have demonstrated the superiority of the proposed method. The dataset and source code will be released athttps://github.com/Hangwei-Chen/AST-IQAD-SRQE
Hangwei Chen, Feng Shao 0001, Xiongli Chai, Yuese Gu, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.1
2023 Cross-Modality Double Bidirectional Interaction and Fusion Network for RGB-T Salient Object Detection
abstract
RGB-T salient object detection (SOD) aims to detect and segment saliency regions on RGB images and the corresponding thermal maps. The ability of alleviating the modality difference between RGB and thermal modality plays a vital role in the development of RGB-T SOD. However, most of the existing methods try to integrate multi-modal information through various fusion strategies, or reduce the modality difference via unidirectional or undifferentiated bidirectional interaction, but failing in some challenging scenes. To deal with the above question, a novel Cross-Modality Double Bidirectional Interaction and Fusion Network (CMDBIF-Net) for RGB-T SOD is proposed. Specifically, we construct an interactive branch to indirectly bridge the RGB and thermal modalities. In addition, we propose a double bidirectional interaction (DBI) module composed of a forward interaction block (FIB) and a backward interaction block (BIB) to reduce the cross-modality differences. Moreover, a multi-scale feature enhancement and fusion (MSFEF) module is introduced to integrate the multi-modal features with considering the internal gap of different modality. Finally, we use a cascaded decoder and a cross-level feature enhancement (CLFE) module to generate high-quality saliency map. Extensive experiments are conducted on three publicly available RGB-T SOD datasets shows that the proposed CMDBIF-Net achieves outstanding performance against the state-of-the-art (SOTA) RGB-T SOD methods.
Zhengxuan Xie, Feng Shao 0001, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.4
2023 Dual-Task Interactive Learning for Unsupervised Spatio-Temporal-Spectral Fusion of Remote Sensing Images
abstract
Spatio-temporal-spectral fusion aims to produce high spatio-temporal-spectral resolution images by integrating the complementary spatial, temporal, and spectral advantages of multi-source remote sensing images. However, on one hand, existing spatio-temporal-spectral fusion methods are insufficient to exploit the inherent complex nonlinear spatial, temporal, and spectral relationship among multisource and multitemporal observations. On the other hand, since the unavailability of real high spatio-temporal-spectral resolution images, it is difficult to adopt deep learning methods with supervised training. In this paper, we propose an effective Unsupervised Spatio-Temporal-Spectral Fusion Model (USTSFM) with dual-task interactive learning to alleviate these problems. The proposed USTSFM has two branches: the Spatio-Temporal-Spectral Mapping (STSM) branch is to describe the temporal relationship, and the Spectral Super Resolution (SSR) branch is to model the spectral relationship. Moreover, the spatial-spectral interaction compensation block is designed to make the two branches compensate and benefited from each other. This intrinsically related and mutually facilitated strategy allows the USTSFM to sufficiently exploit the inherent spatial, temporal, and spectral relationship. In addition, a shared reconstruction module is meticulously designed for the two tasks, which not only reduces the parameters but also allows the supervised task to guide the convergence of the unsupervised task, boosting the stability of unsupervised training. The qualitative and quantitative results demonstrated the proposed USTSFM has richer spatial details and more accurate predictions than the other state-of-the-art methods.
Qiang Liu 0035, Xu Chen 0041, Xiangchao Meng, Hangwei Chen, Feng Shao 0001, Weiwei Sun 0005
IEEE Trans. Geosci. Remote. Sens.4
2023 Perceptual Quality Assessment of Cartoon Images
abstract
In the animation industry, automatically predicting the quality of cartoon images based on the inputs of general distortions and color change is an urgent task, while the existing no-reference (NR) methods usually measure the perceptual quality of the natural images. In this paper, based on the observation that structure and color are the main factors affecting cartoon images quality, we proposed a new NR quality prediction metric for cartoon images, which fully takes gradient and color information into account. The experimental results on our newly constructed NBU-CIQAD dataset with color change and other existing cartoon image dataset demonstrate that the proposed method significantly outperforms existing no-references methods for the task of cartoon image quality assessment. The database and code will be released athttps://github.com/1010075746/NBU-CIQAD.
Hangwei Chen, Xiongli Chai, Feng Shao 0001, Xuejin Wang, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Multim.1
2023 Composition-Guided Neural Network for Image Cropping Aesthetic Assessment
abstract
How to explore the interaction between image aesthetic rules and crops is the key to finding views with good composition. Besides, it is subjective to evaluate candidate crops, which mainly depends on aesthetic knowledge, but it is not an easy task for people without extensive photography experience. However, existing methods mostly find good views by extracting general aesthetic features of crops without fully exploring the aesthetic rules. Motivated by this, we innovatively propose a composition-guided image cropping aesthetic assessment network (CGICAANet) for efficiently finding good crops and optimizing the cropping operation. Specifically, we adopt a direct and comprehensive composition pattern module, which adaptively mines suitable compositions for the images and emphasizes the dominant position of visual elements to contribute to optimizing the best crops in an interpretable way. Moreover, we designed a multi-task loss function to train the model. Particularly, to explore the commonality between predicted crops and labels, the complete intersection-over-union loss is adopted thoroughly considering the overlap area, central point distance and the consistency of aspect ratios for crops concurrently. Therefore, the predicted best crop can preserve the visual elements and have better composition. Experimental results with lightweight MobileNetV2 and ShuffleNetV2 as backbone networks demonstrate that our method can obtain comparable or better performance in terms of efficiency and accuracy.
Shijia Ni, Feng Shao 0001, Xiongli Chai, Hangwei Chen, Yo-Sung Ho
IEEE Trans. Multim.4
2022 CGMDRNet: Cross-Guided Modality Difference Reduction Network for RGB-T Salient Object Detection
abstract
How to explore the interaction between the RGB and thermal modalities is the key success of the RGB-T saliency object detection (SOD). Most of the existing methods integrate multi-modality information by designing various fusion strategies. However, the modality gap between the RGB and thermal features will lead to unsatisfactory performances by simple feature concatenation. To solve this problem, we innovatively propose a cross-guided modality difference reduction network (CGMDRNet) to achieve intrinsic consistency feature fusion via reducing the modality differences. Specifically, we design a modality difference reduction (MDR) module, which is embedded in each layer of the backbone network. The module uses a cross-guided strategy to reduce the modality difference between the RGB and thermal features. Then, a cross-attention fusion (CAF) module is designed to fuse cross-modality features with small modality differences. In addition, we use a transformer-based feature enhancement (TFE) module to enhance the high-level feature representation that contributes more to performance. Finally, the high-level features guide the fusion of low-level features to obtain a saliency map with clear boundaries. Extensive experiments on three public RGB-T datasets show that the proposed CGMDRNet achieves competitive performance compared with state-of-the-art (SOTA) RGB-T SOD models.
Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.4