EDBT 2026 Demo / reviewers in the wild / expert
Qiuping Jiang
dblp:149/0439
· DBLP profile ↗
166ranked-venue papers
27as first author
135since 2021 · last 2026
0000-0002-6025-9343ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 108 · 20 first-author · 83 since 2021Artificial intelligence and machine learning · 33 · 5 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 15 since 2021Computer networks · 7 · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Codebook-Empowered Analysis-Friendly Extreme Underwater Image CompressionabstractWhile existing underwater image compression (UIC) methods optimize for human perception or basic redundancies, they neglect inter-image correlations and fail to prioritize machine-friendly features essential for automated analysis. This paper introduces a novel -quantized (VQ) codebook-driven framework for machine-centric UIC. We leverage VQ codebooks -- pre-trained as external priors on diverse underwater data -- to unify three critical stages: (1) Machine-friendly feature extraction via contrastive learning with high/low-quality codebooks, enhancing degradation robustness; (2) Compact compression using variable-size codebooks to map discriminative features to entropy-coded indices, enabling ultra-low bitrates (less than 0.04bpp); and (3) Feature refinement at the decoder, restoring semantic fidelity for downstream tasks. In addition, we contribute the first Underwater Visual Question Answering (UVQA) benchmark to holistically evaluate machine perception across object presence, counting, and localization. Extensive experiments demonstrate that our framework significantly outperforms state-of-the-art codecs in machine vision task performance at ultra-low bitrates. The VQ-codebook effectively harnesses inter-image redundancy, combats joint degradation, and delivers compact, analysis-friendly representations, establishing a new paradigm for machine-centric UIC. Jianhao Wu, Yudong Mao, Qiuping Jiang |
AAAI | 3 |
| 2026 | Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real PhotosabstractVision-Language Models (VLMs) are increasingly deployed in socially consequential settings, raising concerns about social bias driven by demographic cues. A central challenge in measuring such social bias is attribution under visual confounding: real-world images entangle race and gender with correlated factors such as background and clothing, obscuring attribution. We propose a \textbf{face-only counterfactual evaluation paradigm} that isolates demographic effects while preserving real-image realism. Starting from real photographs, we generate counterfactual variants by editing only facial attributes related to race and gender, keeping all other visual factors fixed. Based on this paradigm, we construct \textbf{FOCUS}, a dataset of 480 scene-matched counterfactual images across six occupations and ten demographic groups, and propose \textbf{REFLECT}, a benchmark comprising three decision-oriented tasks: two-alternative forced choice, multiple-choice socioeconomic inference, and numeric salary recommendation. Experiments on five state-of-the-art VLMs reveal that demographic disparities persist under strict visual control and vary substantially across task formulations. These findings underscore the necessity of controlled, counterfactual audits and highlight task design as a critical factor in evaluating social bias in multimodal models. Qiuping Jiang, Xiaojun Chang, Jun Yu 0002 |
ACL (1) | 4 |
| 2026 | InvJND: Just Noticeable Difference Estimation via Deep Invertible Network
Qiuping Jiang, Zhihua Wang 0002, Shiqi Wang 0001, Feng Shao 0001, Guangtao Zhai, Weisi Lin |
Int. J. Comput. Vis. | 1 |
| 2026 | Semantic Contrast for Domain-Robust Underwater Image Quality AssessmentabstractUnderwater image quality assessment (UIQA) is hindered by complex degradation and domain shifts across aquatic environments. Existing no-reference IQA methods rely on costly and subjective mean opinion scores (MOS), which limit their generalization to unseen domains. To overcome these challenges, we propose SCUIA, an unsupervised UIQA framework leveraging semantic contrastive learning for quality prediction without human annotations. Specifically, we introduce a vision-language contrastive learning strategy that aligns image features with textual embeddings in a unified semantic space, capturing implicit degradation-quality correlations. We further enhance quality discrimination with a hierarchical contrastive learning mechanism that combines image-specific statistical priors and semantic prompts. A triplet-based inter-group contrastive loss explicitly models relative quality relationships. To tackle cross-domain variations, we develop an unsupervised domain adaptation module that uses local statistical features to guide CLIP fine-tuning to disentangle domain-invariant quality representations from domain-specific noise. This enables zero-shot cross-domain quality prediction without labeled data. Extensive experiments on public UIQA benchmarks demonstrate significant improvements over existing methods, highlighting superior generalization and domain adaptability. Jingchun Zhou, Chunjiang Liu, Qiuping Jiang, Xianping Fu, Junhui Hou, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Joint luminance-chrominance learning for quality assessment of low-light image enhancement
Tuxin Guan, Qiuping Jiang, Xiongli Chai |
Pattern Recognit. | 2 |
| 2026 | FEAD: Frequency-enhanced cross-modal prompt learning for zero-shot anomaly detection
Qianyi Li, Xiaojun Chang, Qiuping Jiang |
Pattern Recognit. | 3 |
| 2026 | Omnidirectional image quality assessment using frequency-domain information
Lixiong Liu, Ruibo Cheng, Qingbing Sang, Qiuping Jiang |
Pattern Recognit. | 4 |
| 2026 | Interactive feature fusion for camera-radar-based vehicle segmentation in bird's-eye view
Chenyang Lu 0002, Liang Li 0010, Xiangchao Meng, Qiuping Jiang, Feng Shao 0001 |
Pattern Recognit. | 5 |
| 2026 | Unpaired overwater image defogging using inverted dark channel prior-guided cycle-consistent generative adversarial network
Yaozong Mo, Tuxin Guan, Qiuping Jiang, Wenqi Ren, Wenwu Wang 0001 |
Pattern Recognit. | 4 |
| 2026 | Underwater image compression for human and machine visions with hybrid priors embedding
Jianhao Wu, Zhanhong Lu, Yudong Mao, Qiuping Jiang |
Pattern Recognit. | 5 |
| 2026 | Amplitude exchanging network for unsupervised underwater image enhancement
Runhao Zeng, Xionglin Zhu, Wenfu Peng, Jiezhang Cao, Zhihua Wang 0002, Qiuping Jiang |
Pattern Recognit. | 7 |
| 2026 | Collision attack and error corrected multimodal-inspired framework for underwater video enhancementabstract• Propose a multimodal-inspired underwater video enhancement framework. • Introduce collision-aware noise injection to simulate feature-level conflicts. • Design a hybrid ViT-convolutional architecture with low computational cost. • Achieve improved clarity and temporal consistency across underwater frames. Underwater video enhancement addresses degradation from absorption, scattering, and turbidity, which hinders visual tasks such as detection and tracking. Unlike traditional single-frame methods, we treat temporal cues, occlusion patterns, and structured degradations as implicit modalities within video streams. To this end, we propose CAECNet (Collision-Attack Error Correction Network), a novel multimodal-inspired underwater video enhancement network that integrates the ‘Collision Attack’ training strategy with the ‘Error Correction’ mechanism. This network overcomes the limitations of traditional single-frame enhancement methods, implementing high-precision and real-time inference. It improves temporal perception through multi-frame fusion and utilizes previous frames to assist real-time inference, meeting the demands of dynamic processing. By incorporating a Vision Transformer (ViT) and a lightweight depthwise separable convolution module, the network enhances spatial feature representation and computational efficiency. A branching-based error correction upsampler is designed to correct feature representation errors and reduce information entropy loss, thereby improving video detail restoration quality. The “Collision Attack” training strategy injects structured noise to accelerate network feature learning and reduce computational costs. Experimental results show that CAECNet significantly outperforms existing methods on multiple underwater video datasets, improving image clarity, inter-frame consistency, and computational efficiency, making it suitable for underwater robotic intelligent perception tasks. Jingchun Zhou, Chunjiang Liu, Dehuan Zhang, Zongxin He, Zifan Lin, Qiuping Jiang |
Pattern Recognit. | 6 |
| 2026 | Sea-out NeRF: Spatial perception enhancement for underwater unmanned systems
Jingchun Zhou, Tianyu Liang, Dehuan Zhang, Gemine Vivone, Qiuping Jiang, Minyi Xu |
Pattern Recognit. | 5 |
| 2026 | Underwater image stitching via optimal seam estimation and multi-band fusion
Jingchun Zhou, Danny J. J. Wang, Bing Long, Dehuan Zhang, Qiuping Jiang |
Pattern Recognit. | 5 |
| 2026 | Self-Anchored Progressive Framework With Noise Mitigation for Unsupervised Camouflaged Object DetectionabstractUnsupervised Camouflaged Object Detection (UCOD) presents a significant challenge due to the inherent similarity between camouflaged objects and their backgrounds, compounded by the absence of manual annotations. Although pixel-level pseudo-labeling has proven effective for unsupervised salient object detection (USOD), it is far less reliable for COD, where the concealed and ambiguous nature of camouflaged objects frequently produces noisy pseudo-labels, causing misjudgments, missed detections, and imprecise boundaries. To overcome this, we propose SAPNet, a novel self-anchored progressive framework for UCOD. Rather than depending on noisy pixel-level supervision, we leverage semantically reliable foreground and background regions as high-confidence anchors. This effectively transforms the unsupervised problem into a more robust weakly supervised paradigm, reducing learning difficulty and mitigating overfitting to noise. SAPNet learns camouflaged objects progressively by first emphasizing these confident regions and then exploiting DINO's contextual awareness to recover complete structures. Central to our framework is the semantic-driven region detector (SDRD), which employs cascaded convolutions and a residual attention projection mechanism to suppress background noise, filter erroneous information, and enhance spatial context, ensuring reliable supervision signals. Furthermore, a region-based context inference module (RCIM) is introduced to iteratively refine object boundaries by integrating multi-level semantic features under the guidance of these refined region-level anchors. Extensive experiments on four benchmark COD datasets demonstrate that SAPNet significantly outperforms state-of-the-art unsupervised methods. The source code of our SAPNet is available at https://github.com/ArloJie/SAPNet. Binwei Xu, Tuo Shen, Guanghui Yue 0001, Qiuping Jiang |
IEEE Trans. Image Process. | 5 |
| 2026 | Toward a Completely Blind Attacker for No-Reference Image Quality Assessment ModelsabstractNo-reference image quality assessment (NR-IQA) models are critically vulnerable to adversarial attacks, posing significant risks to downstream vision systems. However, existing attack methods suffer from high computational costs, reliance on Mean Opinion Score (MOS) annotations, and poor cross-model transferability. To overcome these limitations, we propose Degrade-to-OverReconstruct (DOR), a novel prior knowledge-driven black-box attack framework operating in a "completely blind" manner, requiring neither MOS labels nor surrogate models, inducing significant prediction bias solely based on distortion statistics. Specifically, DOR generates universal adversarial examples by first applying mild degradation to preserve global structure and then employing aggressive over-reconstruction using a Residual Denoising Diffusion Model (RDDM) to adaptively disrupt intrinsic Natural Scene Statistics (NSS)-a shared foundation across NR-IQA models. Extensive experiments on synthetic (LIVE, TID2013) and authentic (CLIVE) datasets demonstrate DOR's strong attack performance and superior transferability against leading NR-IQA models that cover diverse deep neural network architectures. Our work pioneers a diffusion model-based "completely blind" attack paradigm, offering a practical, MOS-free solution for adversarial robustness assessment of NR-IQA models in real-world deployments. Xinyu Ruan, Hangwei Chen, Chao Huang 0008, Wenqi Ren, Qiuping Jiang |
IEEE Trans. Image Process. | 5 |
| 2026 | Harnessing Multi-Modal Large Language Models for Measuring and Interpreting Color DifferencesabstractThe accurate measurement of perceptual color differences (CDs) between two images plays an important role in modern smartphone photography. Although traditional CD metrics provide numerical scores to quantify color variations, they often lack the ability to offer intuitive insights or explanations that reflect the factors behind these differences in a way that aligns with human perception and reasoning. Here, we present CD-Reasoning, an innovative method designed not merely to compute numerical CD scores but also to provide a detailed rationale for the observed CDs between images. This method surpasses simple numerical quantification, delivering a more profound and explanatory analysis that bridges quantitative assessments with the qualitative reasoning characteristic of human perception. The development of the CD-Reasoning model begins with the compilation of a multi-modal CD dataset dubbed M-SPCD based on the existing SPCD, where we collect textual descriptions that detail the quantification of CDs across seven pivotal attributes: white balance, brightness contrast, color contrast, overall brightness, overall color, shadow detail, and highlight detail. Utilizing the newly curated M-SPCD dataset, we enhance the capabilities of cutting-edge Multimodal Large Language Models (MLLMs) to not only accurately assess numerical CD scores but also to provide in-depth reasoning that explains the CDs between two images. Extensive experiments demonstrate that the proposed CD-Reasoning not only achieves superior accuracy compared to state-of-the-art CD metrics but also significantly exceeds leading MLLMs in CD interpreting. Source codes will be available at https://github.com/LongYu-LY/CD-Reasoning. Zhihua Wang 0002, Qiuping Jiang, Chao Huang 0008, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2026 | SGNet: Style-Guided Network With Temporal Compensation for Unpaired Low-Light Colonoscopy Video EnhancementabstractA low-light colonoscopy video enhancement method is needed as poor illumination in colonoscopy can hinder accurate disease diagnosis and adversely affect surgical procedures. Existing low-light video enhancement methods usually apply a frame-by-frame enhancement strategy without considering the temporal correlation between them, which often causes a flickering problem. In addition, most methods are designed for endoscopic devices with fixed imaging styles and cannot be easily adapted to different devices. In this paper, we propose a Style-Guided Network (SGNet) for unpaired Low-Light Colonoscopy Video Enhancement (LLCVE). Given that collecting content-consistent paired videos is difficult, SGNet adopts a CycleGAN-based framework to convert low-light videos to normal-light videos, in which a Temporal Compensation (TC) module and a Style Guidance (SG) module are proposed to alleviate the flickering problem and achieve flexible style transfer, respectively. The TC module compensates for a low-light frame by learning the correlated feature of its adjacent frames, thereby improving the temporal smoothness of the enhanced video. The SG module encodes the text of the imaging style and adaptively explores its intrinsic relationships with video features to obtain style representations, which are then used to guide the subsequent enhancement process. Extensive experiments on a curated database show that SGNet achieves promising performance on the LLCVE task, outperforming state-of-the-art methods in both quantitative metrics and visual quality. Guanghui Yue 0001, Wanqing Liu, Jingfeng Du, Tianwei Zhou, Hanhe Lin, Qiuping Jiang, Wenqi Ren |
IEEE Trans. Image Process. | 7 |
| 2026 | Decouple-Then-Synergize: A Self-Paced Collaborative Learning Network for RGB-T Snowy Urban Scene ParsingabstractFusing RGB and thermal infrared images is essential for advancing urban scene analysis. However, both modalities exhibit severe performance degradation under snowy conditions. Although independent enhancement modules can partially mitigate this issue, stacking multiple modules with different functions increases model complexity and may cause intermodular interference. To address these limitations, we propose a "decouple-then-synergize" framework that decouples the task into frequency-oriented enhancement and spatial semantic fusion, implemented by FRENet (frequency restoration enhancement network) and SIFNet (spatial interactive fusion network), respectively. FRENet uses an asymmetric enhancement strategy that selectively sharpens RGB color gradients while amplifying faint thermal targets. It incorporates a precise spectral refinement module to restore high-frequency details. SIFNet introduces a Mamba zipper fusion module to achieve robust interaction of high-level semantics and performs a reconstruction task to implicitly integrate thermal features into the RGB stream. To ensure effective collaboration between the two networks, we design a self-paced curriculum that manages bidirectional knowledge exchange at both the sample and pixel levels. This approach enables the networks to evolve into their enhanced versions, namely FRENet-collaborative learning (CL) and SIFNet-CL. Extensive experiments on the SUS and PST900 datasets demonstrate that our framework outperforms state-of-the-art scene parsing methods. The code and associated results are available at https://github.com/Lyb-2001/SPCL. Wujie Zhou, Yiben Li, Qiuping Jiang, Runmin Cong, Weisi Lin |
IEEE Trans. Image Process. | 3 |
| 2026 | Turbidity-Similarity Decoupling: Feature-Consistent Mutual Learning for Underwater Salient Object DetectionabstractUnderwater salient object detection (USOD) faces two major challenges that hinder accurate detection: substantial image noise owing to water turbidity and low foreground-background contrast caused by high visual similarity. In this study, a dual-model architecture based on mutual learning is proposed to address these issues. First, DenoisedNet, which focuses on addressing water turbidity issues, is developed. Using a separation-denoising-enhancement processing framework, it suppresses noise while maintaining target feature integrity through domain separation and cleaning enhancement modules. Second, SearchNet is designed to address the foreground-background similarity issue. It achieves precise localization through pseudo-label generation and layer-by-layer search mechanisms. To enable both networks to address these challenges collaboratively, a feature-consistent mutual-learning strategy is proposed, which aligns encoded features and prediction results, via evaluation and cross modes, respectively. This strategy enables their respective strengths to be complemented and the challenges of USOD to be solved more comprehensively. Our DenoisedNet and SearchNet outperform the best existing methods on the USOD10K and USOD benchmarks, achieving MAE improvements of 4.52%/5.52% and 1.61%/8.94%, respectively. The source code is available at https://github.com/BeibeiIsFreshman/DSNet_CL. Wujie Zhou, Beibei Tang, Runmin Cong, Qiuping Jiang |
IEEE Trans. Image Process. | 4 |
| 2026 | USformer: A U-Shaped Structure Transformer for RGB-Thermal Semantic Segmentation and Traffic Scene UnderstandingabstractRecent advancements in multimodal approaches, particularly RGB-thermal (RGB-T) segmentation, have significantly promote the development of Intelligent Transportation Systems (ITS). However, existing methods still encounter challenges related to modality discrepancy and the effective integration of multi-scale features. To address these issues, we propose the U-shaped Structure Transformer (USformer) for RGB-T semantic segmentation. We improve the feature flow of existing methods by designing a novel U-shaped encoding network that integrates inter-layer fusion and cross-modal fusion. Specifically, our method introduces an inter-layer interaction mechanism that facilitates the iterative fusion of high-level semantic and low-level detail features. For each layer, our fusion process is divided into two stages: the Cross-Modal and -Scale Auxiliary (CMSA) module enforces distribution alignment across modalities and scales, while the Cross-Attention Feature Merger (CAFM) allows each modality to refine its own feature selection by employing a multi-head cross-attention mechanism. These modules effectively adapt and integrate well-established attention designs into our U-shaped encoding architecture, thereby achieving efficient multi-modal feature alignment and fusion. Finally, we utilize the Mask2Former decoder to aggregate the fused features from multiple layers and improve the segmentation across various object sizes and complex scenes. Extensive experiments on four RGB-T datasets demonstrate that our proposed USformer achieves state-of-the-art performance. Feng Shao 0001, Baoyang Mu, Xiongli Chai, Qiuping Jiang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2026 | Quality Evaluation of AI-Generated Images: Subjective Study and Objective MethodologyabstractIn recent years, AI-Generated Images (AIGIs) have attracted significant attention and shown great potential in various applications, including entertainment, advertisement, education, and product design. Driven by this trend, various Text-to-Image (T2I) models are developed. However, the quality of AIGIs produced by these models varies widely, with many low-quality images failing to meet human aesthetic standards. Consequently, research into both subjective and objective Image Quality Assessment (IQA) methods for AIGIs is crucial. In this paper, we introduce a dataset called AIGI-IQAD, designed to enhance our understanding of human aesthetic preferences for AIGIs. The dataset contains 2,880 AIGIs generated by 8 T2I models using 360 deliberately designed text prompts. Further, we conducted subjective experiments to gather ratings from both aesthetic quality and text-image consistency. Building on this dataset, we propose a model named Question-guided Multimodal Interaction Network (QMI-Net) for evaluating AIGIs. QMI-Net assesses human preferences for AIGIs by focusing on both aesthetic quality and text-image consistency. Specifically, QMI-Net uses a question-answering approach to guide Multimodal Large Language Models (MLLMs) in generating detailed aesthetic and similarity information. The Visual and Aesthetic Feature Fusion Module (VAFFM) then fuses the aesthetic features with the visual features extracted by Contrastive Language-Image Pre-training (CLIP) to obtain more comprehensive aesthetic quality features. Comprehensive experiments demonstrate that state-of-the-art performance is achieved by QMI-Net on our AIGI-IQAD and three other public datasets.The AIGI-IQAD datasets and QMI-Net will be released athttps://github.com/ctxya1207/QMI-Net. Feng Shao 0001, Hangwei Chen, Xuejin Wang, Qiuping Jiang |
IEEE Trans. Multim. | 6 |
| 2026 | PyUIE: A Coarse-to-Fine Deep Pyramid Network for Underwater Image EnhancementabstractUnderwater images often suffer from color distortion, reduced contrast, and blurriness due to light refraction, absorption, and scattering. In this paper, we propose a coarse-to-fine deepPyramid network forUnderwaterImageEnhancement (PyUIE). Specifically, PyUIE begins by decomposing the input image into high- and low-frequency components using a Laplacian pyramid. The low-frequency residual, which primarily contains lighting and color information, is processed with a lightweight deterministic color mapping network to correct global illumination and color distortions. Concurrently, the high-frequency components containing the fine details are enhanced in a coarse-to-fine manner, such that each higher scale is guided by the reconstruction from the adjacent lower scale. This hierarchical strategy effectively mitigates the risk of over-enhancement by avoiding excessive modifications to the high-frequency components. Additionally, we implement a multi-scale supervised training strategy, enabling the model to learn and reconstruct features across multiple scales, which enhances its ability to capture diverse details and improves its generalization and robustness. Extensive experiments demonstrate that our method successfully restores fine details and small structures in underwater images while producing vivid and visually appealing colors, thereby outperforming existing enhancement methods in both qualitative and quantitative evaluations. The code is available athttps://github.com/ttttllt/PyUIE.git. Wenchao Jiang, Yingqing Tan, Zhenxuan Qiu, Zhihua Wang 0002, Yang Yu 0014, Qiuping Jiang |
IEEE Trans. Multim. | 6 |
| 2025 | ICAA-Mamba: Vision Mamba for Image Color Aesthetics AssessmentabstractImage Color Aesthetics Assessment (ICAA) focuses on evaluating the aesthetic quality of color composition within images. This task involves analyzing and quantifying the visual appeal of color arrangements, taking into account factors such as harmony, contrast, and balance, to provide an objective assessment of color aesthetics. In this paper, we investigate the application of the State Space Model (Mamba) to the ICAA task, with a focus on exploring the perceptual capabilities of vision Mamba. To this end, we propose a novel framework that employs Mamba as the backbone to extract informative patterns from images, leveraging its global receptive field and linear complexity with respect to input length. Instead of relying solely on a final-layer feature, followed by fully connected layers for quality prediction, we exploit a comprehensive set of multi-scale features, thereby constructing a richer global representation of color information. To fuse these multi-scale features effectively, we introduce a Weighted Feature Fusion Module (WFFM), which adaptively assigns weights to emphasize salient color-related information. We conduct extensive experiments on two benchmark datasets, ICAA17K and SPAQ, successfully demonstrating its effectiveness for the ICAA task. The code is available at https://github.com/qinghaoya/icaa-mamba. Qinghao Xie, Wenchao Jiang, Zhihua Wang 0002, Qiuping Jiang |
ICASSP | 4 |
| 2025 | Multiple Glancing at Quality: Benchmark Dataset and Objective Quality Assessment Metric for Low-light Image EnhancementabstractQuality metrics play a crucial role in guiding the development of image enhancement algorithms, which have consistently sought effective quality assessment methodologies and comprehensive datasets. To address this need, we first built a large-scale dataset for low-light enhanced image quality assessment and gathered the corresponding subjective evaluation scores. Recently, vision-language pre-training models have demonstrated considerable potential in the realm of quality assessment. However, its efficacy is limited by the fine-grained perception in low-level quality assessment. As such, we further propose a novel quality assessment framework using contrastive prompt learning, which harnesses the robust priors of vision-language pre-training models to improve the perceptual capacity of deep networks for low-level quality features. Experiments on the proposed RSLE dataset show that our method outperforms existing SOTA image quality assessment methods. Our database and the source code will be made publicly available. Yudong Mao, Peilin Chen 0001, Zhao Wang 0004, Qiuping Jiang, Shiqi Wang 0001 |
ISCAS | 4 |
| 2025 | Underwater Camera: Improving Visual Perception Via Adaptive Dark Pixel Prior and Color Correction
Jingchun Zhou, Qiuping Jiang, Wenqi Ren, Kin-Man Lam 0001, Weishi Zhang |
Int. J. Comput. Vis. | 3 |
| 2025 | RGBT-Booster: Detail-Boosted Fusion Network for RGB-Thermal Crowd Counting With Local Contrastive LearningabstractWith the swift development of the Internet of Video Things (IOVT), crowd counting has demerged as an indispensable technology in the domains of intelligent transportation and video surveillance. However, due to the insufficient extraction of detail head information and the limited ability to reduce the multimodality differences, the existing methods still have large errors in accurate RGB-thermal (RGB-T) crowd counting. To this end, we propose a novel RGB-T crowd counting network, i.e., RGBT-Booster, to effectively deal with the aforementioned challenges. In RGBT-Booster, by introducing additional detail auxiliary branches for RGB and thermal infrared images and the proposed enhanced detail fusion module (EDFM), we can obtain richer low-level head detail features. In addition, we also propose a local contrastive learning (LCL) to further reduce the multimodality differences for accurate crowd counting. Experimental results on two public RGB-T crowd counting datasets (i.e., RGBT crowd counting (RGBT-CC) and DroneRGBT) and one RGB-Depth (RGB-D) crowd counting dataset (i.e., ShanghaiTechRGBD) show that the proposed RGBT-Booster achieves effective and superior counting performance, compared with previous methods. The source code and datasets used in the experiments will be released athttps://github.com/QSBAOYANGMU/RGBT-Booster. Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Long Xu 0001, Qiuping Jiang |
IEEE Internet Things J. | 5 |
| 2025 | SGUVE-Net: Semantic-Guided Underwater Video Enhancement Network for Real-Time IoT-Based Marine MonitoringabstractUnderwater video enhancement is crucial for marine research and monitoring applications, particularly in the context of the Internet of Things (IoT), where autonomous underwater vehicles (AUVs) and sensor networks are deployed for environmental monitoring and tracking marine life. However, the scarcity of undistorted underwater video data and distortions, such as motion blur and water turbidity, limit the effectiveness of enhancement models. Existing methods typically focus on frame-by-frame enhancement and overlook temporal coherence and computational efficiency. To address these issues, we propose SGUVE-Net (Semantic-Guided Underwater Video Enhancement Network), which combines a multi-scale feature-aware network with a semantic branch for localized enhancement. The main branch employs an encoder-decoder architecture, combining spatial group shifting and dual attention mechanisms to fully exploit contextual information for precise alignment. In contrast, the semantic branch focuses on enhancing key regions of the video frames by incorporating high-level semantic cues, which improves motion tracking accuracy and mitigates motion blur, thereby enhancing video quality for real-time IoT applications. They complement each other to achieve differentiated modeling of static background details and dynamic target features. Experimental results show that SGUVE-Net outperforms state-of-the-art methods across several metrics, providing an effective solution for underwater video enhancement in IoT systems. Jingchun Zhou, Wenyu Fan, Bing Long, Dehuan Zhang, Zongxin He, Qiuping Jiang, Muhammad Ghulam |
IEEE Internet Things J. | 6 |
| 2025 | Degradation-Decoupling Vision Enhancement for Intelligent Underwater Robot Vision Perception SystemabstractUnderwater robots rely on high-quality visual data for precise monitoring and manipulation, yet complex underwater environments often degrade image quality through color distortion, texture blurring, and detail loss. Existing enhancement methods partially address these issues, but fail to effectively decouple nonlinear relationships among degradation factors, leading to inconsistent performance. To address these challenges, we propose a degradation-content decoupling-based underwater image enhancement network (DCDN). The framework integrates a super-fusion cascade module for dynamic feature weighting, reducing artifacts, and combines multichannel color space transformation with texture-guided correction to decouple and optimize degradation factors. This approach improves color fidelity and texture detail restoration by refining color information and adapting local textures. Experiments on public datasets demonstrate that DCDN outperforms existing methods in various underwater scenarios. This work enhances the visual capabilities of underwater robots, supporting intelligent transportation applications, such as marine logistics and underwater inspections. Jingchun Zhou, Chunjiang Liu, Bing Long, Dehuan Zhang, Qiuping Jiang, Muhammad Ghulam |
IEEE Internet Things J. | 5 |
| 2025 | Hybrid Knowledge Distillation for RGB-T Crowd Density Estimation in Smart Surveillance SystemsabstractCrowd density estimation is a practical application task in which speed efficiency is as crucial as the accuracy of the results. Hence, we propose the hybrid knowledge distillation network (HKDNet) for RGB-thermal (RGB-T) crowd density estimation to address the limitations of computational cost and training time from the perspective of ensuring accuracy. We efficiently combine the advantages of traditional convolution and self-attention through a multimodal interactive transform. Subsequently, cross-graph convolution is extended to an interactive space for multimodal relationship reasoning. Finally, from the perspectives of channel and space, the crowd density map is obtained using channel synergy supplementation and spatial detail filling. In contrast to the computing resources required when using a heavyweight teacher network, the proposed HKDNet uses only approximately 6% of the parameters of the teacher network by abstracting the teacher network into a network hierarchy to generate a lightweight and efficient student network. Extensive experiments show that the proposed HKDNet performed well on two RGB-T crowd density estimation datasets. The code and models are available athttps://github.com/WBangG/HKDNet. Wujie Zhou, Weiqing Yan, Qiuping Jiang |
IEEE Internet Things J. | 4 |
| 2025 | MDGINet: Multi-frequency Dynamic Guidance and Interaction network for image denoising
Feng Shao 0001, Hangwei Chen, Xiongli Chai, Qiuping Jiang |
Knowl. Based Syst. | 5 |
| 2025 | Cross-Modality Interactive Attention Network for AI-generated image quality assessment
Tianwei Zhou, Songbai Tan, Leida Li, Baoquan Zhao, Qiuping Jiang, Guanghui Yue 0001 |
Pattern Recognit. | 5 |
| 2025 | Art Comes From Life: Artistic Image Aesthetics Assessment via Attribute Knowledge AmalgamationabstractAssessing the aesthetic quality and visual appeal of artworks has become one of the hotspots in current research. The existing artistic image aesthetics assessment (AIAA) methods directly learn aesthetics from images, while ignoring the impact of variations in visual attributes on human aesthetic perception, which hampers the further development of AIAA. To address this issue, this paper presents a new AIAA method based on attribute knowledge amalgamation, named AKA-Net. Specifically, we initially learn common attribute aesthetic rules (e.g., composition and color) through pre-training on natural aesthetic images. Then, we devise a multi-model amalgamation strategy based on contrastive learning to transfer different types of prior attribute knowledge into a single target model, enabling flexible and efficient aesthetic prediction. Finally, an attribute-aware feature enhancement module (AFEM) is introduced to better establish the relationship between aesthetic quality and attribute knowledge. Experimental results on three public benchmark AIAA databases demonstrate that the proposed AKA-Net outperforms the state-of-the-art AIAA metrics. Hangwei Chen, Feng Shao 0001, Xiongli Chai, Baoyang Mu, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Toward Dimension-Enriched Underwater Image Quality AssessmentabstractThe absorption and scattering of light in the water medium naturally impair the quality of underwater images, leading to multiple degradation effects including color casts, reduced visibility, and blurriness. Underwater Image Enhancement (UIE) techniques strive to mitigate these issues, yet the efficacy of different UIE algorithms remains highly variable. This variability underscores the necessity for an objective quality metric capable of precisely assessing the visual quality of underwater images. Traditional quality metrics, which primarily rely on a single score to depict the overall quality level, are insufficiently comprehensive to describe the complex degradation characteristics intrinsic to underwater environments and the multi-dimensional nature of underwater image quality. To address this issue, we construct the first UIE quality evaluation dataset with multi-dimensional quality annotations, broadening the subjective labels from a single overall quality score to multiple specific degradation-related scores. The dataset is known as an enhanced version of our previous Subjectively Annotated UIE Benchmark Dataset (SAUD) and is called SAUD2.0 hereinafter. Based on the SAUD2.0 dataset, we also introduce a Multi-stream COllaborative LEarning network (MCOLE) tailored for quality evaluation of enhanced underwater images. MCOLE capitalizes on the multi-dimensional quality annotations within SAUD2.0, facilitating the training of three specialized networks focused on extracting distinct sets of features: color, visibility, and semantic. These extracted features are then interacted and cohesively merged for quality prediction. Comprehensive experiments conducted on two benchmark datasets reveal that the proposed MCOLE outperforms current underwater image quality metrics. These results clearly validate the efficacy of exploring the multi-dimensional nature of underwater image quality and integrating such multi-dimensional quality annotations into underwater image quality evaluation. Our dataset and code are available athttps://github.com/0117Tzx/MCOLE. Qiuping Jiang, Xiao Yi, Li Ouyang, Jingchun Zhou, Zhihua Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Underwater Salient Object Detection via Dual-Stage Self-Paced Learning and Depth EmphasisabstractSalient object detection of underwater scenes (USOD) poses greater challenges than that of traditional terrestrial scenes due to the presence of diverse and complex underwater image degradation. Current deep learning-based USOD methods generally treat all samples equally while failing to account for the varying difficulty levels of different training samples, thus leading to a limited performance. To tackle this challenge, this paper introduces a novel deep USOD method which benefits from iterative Dual-stage Self-paced Learning (DSPL) and Salient Object Depth Emphasis (SODE). Specifically, a DSPL strategy, which enforces the network to only focus on simpler samples in the first stage and then shifts attention to more challenging samples in the second stage, is devised to imitate the learning process of humans. The whole network is iteratively trained with the DSPL strategy and thus gradually adapted to various underwater scenes with different difficulty levels. Additionally, the proposed method involves an SODE module, which adaptively enhances depth information to effectively locate salient objects, addressing the issue of unreliable depth data caused by underwater image quality degradation. Experimental results on two benchmark datasets demonstrate the superior performance of the proposed method against state-of-the-art methods. The source code of our method will be made available athttps://github.com/NIT-JJH/SPDE. Jianhui Jin, Qiuping Jiang, Qingyuan Wu, Binwei Xu, Runmin Cong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | A Mutual Head Knowledge Distillation Framework for Lightweight RGB-T Crowd CountingabstractAs an important technology in the fields of intelligent transportation and public safety, crowd counting that can obtain pedestrian flow information has attracted extensive attention from academic and industrial communities. However, existing RGB-T crowd counting methods cannot effectively balance the counting accuracy and computational complexity in practical applications. For this, we propose a Mutual Head Knowledge Distillation Framework (MHKDF) to obtain a lightweight RGB-T crowd counting network for efficient and accurate pedestrian number estimation. Specifically, to avoid the influence of parameter and structure differences between teacher and student networks on the distillation effect, we propose a Cooperative Mutual Knowledge Distillation (CMKD) strategy to comprehensively and dynamically transfer the crowd analysis ability of the complex teacher model (MHKDF-T) to the lightweight student model (MHKDF-S). In addition, the upper bound of the performance of the student network depends on the teacher model with high accuracy. Therefore, to take advantage of the complementary advantages of frequency domain and spatial domain feature fusion, we propose a Multi-Modal Spatial-Frequency Hybrid Fusion Module (MSFHFM) to futher improve counting accuracy of MHKDF-T. Comprehensive experiments on two RGB-T crowd counting datasets demonstrate that our MHKDF-S achieves competitive performance with only 5.68 FLOPs and 4.89M parameters. Our code will be released at https://github.com/BaoYangCC/MHKDF. Baoyang Mu, Feng Shao 0001, Hangwei Chen, Xuejin Wang, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Multidimensional Exploration of Segment Anything Model for Weakly Supervised Video Salient Object DetectionabstractFully supervised video salient object detection (VSOD) has made considerable breakthroughs using costly and time-consuming pixel-wise annotations. Recently, to achieve a trade-off between the annotation burden and the model performance, scribble-based VSOD tasks have attracted increasing attention. However, learning the complete object structure and precise boundary details from sparse scribble annotations remains challenging. In this paper, we propose a series of strategies to effectively explore valid information from the recently proposed segmentation foundation model “Segment Anything Model (SAM)” in various perspectives to address these challenges. Specifically, due to the limited performance of SAM on videos, we propose a SAM-guided label enhancement method instead of directly using the results of SAM, which can introduce edge information while reducing the interference of erroneous information. Moreover, we propose a SAM-driven spatiotemporal network guided by general semantic features from the SAM encoder to help the model be aware of global connections. Additionally, we propose a SAM-based global-aware loss, which further considers the affinity constraint between predicted results and foreground labels or background labels from a global perspective, guiding the model to perceive the complete salient objects. Experimental results demonstrate that our method outperforms state-of-the-art weakly supervised VSOD methods and is comparable to fully supervised VSOD methods. Binwei Xu, Qiuping Jiang, Xing Zhao 0001, Chenyang Lu 0002, Haoran Liang 0001, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | MDNet: Mamba-Effective Diffusion-Distillation Network for RGB-Thermal Urban Dense PredictionabstractIn recent years, significant progress has been achieved in urban dense prediction tasks, particularly with advancements in deep learning models and novel architectures that enhance segmentation accuracy and computational efficiency. However, the following challenges persist: i) Existing modal fusion methods typically adopt convolutional neural networks (CNNs) or transformer (Trans)-based methods, which lead to inadequate global modeling or excessive computation owing to the introduction of quadratic complexity modeling; and ii) existing dense prediction networks typically utilize discriminative networks (codecs), which result in networks with insufficient discriminative properties. To address these issues, we propose the Mamba-effective diffusion-distillation network (MDNet) for RGB-thermal urban dense prediction. First, a new Mamba-effective fusion module is proposed, which efficiently models long-range pixel-level features using Mamba and generates pixel-level adaptive weights to fully utilize complementary modal information. Second, inspired by human self-reflection, a new diffusion self-distillation (DSD) strategy is proposed. The DSD generates coarse-grained binary semantic information via conditional multimodal image diffusion, which serves as self-distillation labels to improve the discriminative properties of the network. Experimental results demonstrate that the proposed MDNet achieves state-of-the-art performance on the MFNet dataset with fewer parameters and reduced computational effort. Extended experiments on the PST900 dataset further illustrate the effectiveness and generalizability of MDNet. The source code and results are available athttps://github.com/Tortoisewhp/MDNet. Wujie Zhou, Hongping Wu, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Cross-Scale Style-Guided Enhancement for Underwater Remote Sensing ImageryabstractUnderwater imaging is essential for marine remote sensing tasks, such as environmental monitoring, resource exploration, and autonomous navigation. However, images captured in underwater environments often suffer from complex degradations, including wavelength-dependent color distortion, contrast attenuation, and structural detail loss. To address these challenges, we propose a Cross-Scale Style-Guided Network (CSG-Net) for robust underwater image enhancement. CSG-Net employs a dual-stage collaborative framework that decouples global degradation modeling from local detail refinement. In the first stage, a Style Extraction Network (SE-Net) extracts multi-scale degradation-aware style priors that implicitly encode large-scale physical degradation patterns, such as red-channel attenuation and spectral imbalance. In the second stage, a Style-Guided Enhancement Network (SG-Net) leverages these style features to guide spatially adaptive enhancement, enabling consistent color correction and fine-grained detail recovery. To alleviate semantic degradation during scale transitions, CSG-Net introduces the proposed multi-resolution feature-preserving cross-scale interaction (MFPCSI) module, which enhances the preservation and integration of hierarchical features. Combined with the Multi-Stream Information Fusion (MSIF) module, this design enables the effective fusion of semantic and structural information across spatial scales. The proposed components enable the preservation of fine-grained details while adaptively integrating semantic and structural cues across multiple scales. Comprehensive experiments conducted on diverse and challenging underwater image datasets demonstrate that CSG-Net consistently surpasses state-of-the-art approaches in terms of PSNR, SSIM, and UIQM. Furthermore, the model exhibits strong cross-domain generalization and delivers high-fidelity visual results, underscoring its suitability for deployment in practical vision systems operating under complex, real-world environments. Jingchun Zhou, Dehuan Zhang, Xingcheng Han, Qiuping Jiang, Gemine Vivone, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Deep Underwater Image Quality Assessment With Explicit Degradation Awareness EmbeddingabstractUnderwater Image Quality Assessment (UIQA) is currently an area of intensive research interest. Existing deep learning-based UIQA models always learn a deep neural network to directly map the input degraded underwater image into a final quality score via end-to-end training. However, a wide variety of image contents or distortion types may correspond to the same quality score, making it challenging to train such a deep model merely with a single subjective quality score as supervision. An intuitive idea to solve this problem is to exploit more detailed degradation-aware information as supplementary guidance to facilitate model learning. In this paper, we devise a novel deep UIQA model with Explicit Degradation Awareness embedding, i.e., EDANet. To train the EDANet, a two-stage training strategy is adopted. First, a tailored Degradation Information Discovery subnetwork (DIDNet) is pre-trained to infer a residual map between the input degraded underwater image and its pseudoreference counterpart. The inferred residual map explicitly characterizes the local degradation of the input underwater image. The intermediate feature representations on the decoder side of DIDNet are then embedded into the Degradation-guided Quality Evaluation subnetwork (DQENet), which significantly enhances the feature characterization capability with higher degradation awareness for quality prediction. The superiority of our EDANet against 18 state-of-the-art methods has been well demonstrated by extensive comparisons on two benchmark datasets. The source code of our EDANet is available at https://github.com/yia-yuese/EDANet. Qiuping Jiang, Yuese Gu, Zongwei Wu, Chongyi Li, Huan Xiong, Feng Shao 0001, Zhihua Wang 0002 |
IEEE Trans. Image Process. | 1 |
| 2025 | Perception-Oriented Bidirectional Attention Network for Image Super-Resolution Quality AssessmentabstractMany super-resolution (SR) algorithms have been proposed to increase image resolution. However, full-reference (FR) image quality assessment (IQA) metrics for comparing and evaluating different SR algorithms are limited. In this work, we propose the Perception-oriented Bidirectional Attention Network (PBAN) for image SR FR-IQA, which is composed of three modules: an image encoder module, a perception-oriented bidirectional attention (PBA) module, and a quality prediction module. First, we encode the input images for feature representations. Inspired by the characteristics of the human visual system, we then construct the perception-oriented PBA module. Specifically, different from existing attention-based SR IQA methods, we conceive a Bidirectional Attention to bidirectionally construct visual attention to distortion, which is consistent with the generation and evaluation processes of SR images. To further guide the quality assessment towards the perception of distorted information, we propose Grouped Multi-scale Deformable Convolution, enabling the proposed method to adaptively perceive distortion. Moreover, we design Sub-information Excitation Convolution to direct visual perception to both sub-pixel and sub-channel attention. Finally, the quality prediction module is exploited to integrate quality-aware features and regress quality scores. Extensive experiments demonstrate that our proposed PBAN outperforms state-of-the-art quality assessment methods. Xiaoyuan Yang 0003, Guanghui Yue 0001, Jun Fu 0007, Qiuping Jiang, Xu Jia 0012, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
IEEE Trans. Image Process. | 5 |
| 2025 | Toward Better Than Pseudo-Reference in Underwater Image EnhancementabstractSince degraded underwater images are not always accompanied with distortion-free counterparts in real-world situations, existing underwater image enhancement (UIE) methods are mostly learned on a paired set consisting of raw underwater images and their corresponding pseudo-reference labels. Although the existing UIE datasets manually select the best model-generated results as pseudo-References, such pseudo-reference labels do not always exhibit perfect visual quality. Therefore, it would be interesting to investigate whether it is possible to break through the performance bottleneck of UIE networks trained with imperfect pseudo-references. Motivated by these facts, this paper focuses on innovating more advanced loss functions rather than designing more complex network architectures. Specifically, a plug-and-play hybrid Performance SurPassing Loss (PSPL), consisting of a Quality Score Comparison Loss (QSCL) and a scene Depth-aware Unpaired Contrastive Loss (DUCL), is formulated to guide the training of UIE network. Functionally, QSCL aims to guide the UIE network to generate enhanced results with better visual quality than pseudo-references by constructing image quality score comparison losses from both image-level and region-level. Nevertheless, only using QSCL cannot guarantee obtaining desired results for those severely degraded distant regions. Therefore, we also design a tailored DUCL to handle this challenging issue from the scene depth perspective, i.e., DUCL encourages the distant regions of the enhanced results to be closer to the high-quality nearby regions (pull) and far away from the low-quality distant regions (push) of the pseudo-references. Extensive experimental results demonstrate the advantage of using PSPL over the state-of-the-arts even with an extremely simple and lightweight UIE network. The source code will be released at https://github.com/lewis081/PSPL. Yi Liu 0085, Qiuping Jiang, Xingbo Li, Ting Luo 0001, Wenqi Ren |
IEEE Trans. Image Process. | 2 |
| 2025 | High-Resolution Underwater Creature SegmentationabstractUnderwater creature segmentation (UCS) is critical for marine research and robotics but faces unique challenges: environmental distortions and biological traits that distinguish it from terrestrial segmentation. While deep learning advances exist, current UCS models are constrained to low-resolution inputs, losing critical details when processing high-resolution (HR) imagery and degrading segmentation precision. To bridge this gap, we introduce UCS4K, the first large-scale HR dataset for UCS, containing 4,096 images with pixel-wise annotations. UCS4K offers 4 times higher average resolution than existing datasets, covering diverse species, habitats, and environmental complexities essential for robust model training. Additionally, we propose a Resolution-Asymmetric Dual-branch Alignment and Refinement (RADAR) network to address the efficiency-receptiveness trade-off in HR-UCS. RADAR decouples context and detail processing: a CNN branch preserves HR spatial details, while a Transformer branch models global semantics on downsampled inputs to avoid quadratic complexity. Crucially, it resolves the inherent semantic misalignment issue between branches via the Global Semantic Alignment (GSA) module in the encoder and the Bidirectional Collaborative Refinement (BCR) module-embedded decoder that progressively integrates multi-scale encoding features to sharpen boundaries. This asymmetric design ensures efficient long-range context capture without sacrificing spatial precision. Extensive benchmarks demonstrate that RADAR sets new state-of-the-art performance on UCS4K and other existing datasets. Our contributions establish the first HR benchmark for UCS and deliver a scalable framework for high-precision segmentation. Dataset, code, and models are available at https://github.com/WHYfromNUT/RADAR. Huiyang Wu, Qiuping Jiang, Zongwei Wu, Runmin Cong, Cédric Demonceaux, Yi Yang 0001, Xiangyang Ji |
IEEE Trans. Image Process. | 2 |
| 2025 | Text-Guided Semantic Alignment Network With Spatial-Frequency Interaction for Infrared-Visible Image Fusion Under Extreme IlluminationabstractAlthough text-guided infrared-visible image fusion helps improve content understanding under extreme illumination, existing methods usually ignore semantic differences between textual and visual features, resulting in limited improvement. To address this challenge, we propose a Text-Guided Semantic Alignment Network, termed TSANet, for extreme-illumination infrared-visible image fusion. The network follows an encoder-decoder structure, with two image encoders, two text encoders, and one decoder. It uses a Semantic Alignment and Fusion (SAF) block to bridge the two image encoders in each layer. Specifically, the SAF block consists of two parallel Semantic Alignment (SA) modules, corresponding to the infrared and visible modalities, respectively, and a Spatial-Frequency Interaction (SFI) module. The SA module aligns the visual feature from the image encoder with its corresponding textual feature from the text encoder, to guide the network focus on key semantic regions of infrared and visible images. The SFI module aggregates the spatial and frequency information extracted from the modality-aligned features of two SA modules for complementary representation learning. The network progressively complements two image modalities by connecting the SAF blocks from top to down, and finally provides a visually pleasing fusion effect by feeding the output of the last block into the decoder. Recognizing that existing datasets lack illumination diversity, we contribute a new dataset specifically designed for extreme-illumination image fusion. Extensive experiments show the effectiveness and superiority of TSANet over seven state-of-the-art methods. The source code and dataset are available at https://github.com/WentaoLi-CV/TSANet. Guanghui Yue 0001, Cheng Zhao 0003, Zhiliang Wu, Tianwei Zhou, Qiuping Jiang, Runmin Cong |
IEEE Trans. Image Process. | 6 |
| 2025 | Multi-Prior Fusion Transfer Plugin for Adapting In-Air Models to Underwater Image Enhancement and DetectionabstractUnderwater data is inherently scarce and exhibits complex distributions, making it challenging to train high-performance models from scratch. In contrast, in-air models are structurally mature, resource-rich, and offer strong potential for transfer. However, significant discrepancies in visual characteristics and feature distributions between underwater and in-air environments often lead to severe performance degradation when applying in-air models directly. To address this issue, we propose IA2U, a lightweight plugin designed for efficient underwater adaptation without modifying the original model architecture. IA2U can be flexibly integrated into arbitrary in-air networks, offering high generalizability and low deployment costs. Specifically, IA2U incorporates three types of prior knowledge-water type, degradation pattern, and sample semantics-which are embedded into intermediate layers through feature injection and channel-wise modulation to guide the network's response to underwater-specific features. Furthermore, a multi-scale feature alignment module is introduced to dynamically balance information across different resolution paths, enhancing consistency and contextual representation. Extensive experiments demonstrate that IA2U significantly improves both image enhancement and object detection performance. Specifically, on the UIEB dataset, IA2U boosts Shallow-UWNet by 5.2 dB in PSNR and reduces LPIPS by 52%; on the RUOD dataset, it increases AP by 1.8% when applied to the PAA detector. IA2U provides an effective and scalable solution for building robust underwater perception systems with minimal adaptation costs. Our code is available at https://github.com/zhoujingchun03/IA2U. Jingchun Zhou, Dehuan Zhang, Zongxin He, Qilin Gai, Qiuping Jiang |
IEEE Trans. Image Process. | 5 |
| 2025 | Adaptive Cross-Feature Fusion Network With Inconsistency Guidance for Multi-Modal Brain Tumor SegmentationabstractIn the context of contemporary artificial intelligence, increasing deep learning (DL) based segmentation methods have been recently proposed for brain tumor segmentation (BraTS) via analysis of multi-modal MRI. However, known DL-based works usually directly fuse the information of different modalities at multiple stages without considering the gap between modalities, leaving much room for performance improvement. In this paper, we introduce a novel deep neural network, termed ACFNet, for accurately segmenting brain tumor in multi-modal MRI. Specifically, ACFNet has a parallel structure with three encoder-decoder streams. The upper and lower streams generate coarse predictions from individual modality, while the middle stream integrates the complementary knowledge of different modalities and bridges the gap between them to yield fine prediction. To effectively integrate the complementary information, we propose an adaptive cross-feature fusion (ACF) module at the encoder that first explores the correlation information between the feature representations from upper and lower streams and then refines the fused correlation information. To bridge the gap between the information from multi-modal data, we propose a prediction inconsistency guidance (PIG) module at the decoder that helps the network focus more on error-prone regions through a guidance strategy when incorporating the features from the encoder. The guidance is obtained by calculating the prediction inconsistency between upper and lower streams and highlights the gap between multi-modal data. Extensive experiments on the BraTS 2020 dataset show that ACFNet is competent for the BraTS task with promising results and outperforms six mainstream competing methods. Guanghui Yue 0001, Guibin Zhuo, Tianwei Zhou, Weide Liu, Tianfu Wang 0001, Qiuping Jiang |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | AGFNet: Adaptive Gated Fusion Network for RGB-T Semantic SegmentationabstractRGB-T semantic segmentation can effectively pop-out objects from challenging scenarios (e.g., low illumination and low contrast environments) by combining RGB and thermal infrared images. However, the existing cutting-edge RGB-T semantic segmentation methods often present insufficient exploration of multi-modal feature fusion, where they overlook the differences between the two modalities. In this paper, we propose an adaptive gated fusion network (AGFNet) to conduct RGB-T semantic segmentation, where the multi-modal features are combined via the gating mechanisms and the spatial details are enhanced via the introduction of edge information. Specifically, the AGFNet employs a cross-modal adaptive gated-attention fusion (CAGF) module to aggregate the RGB and thermal features, where we give a sufficient exploration of the complementarity between the two-modal features via the gated attention unit (GAU). Particularly, in GAU, the gates can be used to purify the features, and the channel and spatial attention mechanisms are further employed to enhance the two-modal features interactively. Then, we design an edge detection (ED) module to learn the object-related edge cues, which simultaneously incorporates local detail information from low-level features and global location information from high-level features. After that, we deploy the edge guidance (EG) module to emphasize the spatial details of the fused features. Next, we deploy the contextual elevation (CE) module to enrich the contextual information of features by iteratively introducing the sine and cosine functions. Finally, considering that the quality of thermal images is usually lower than that of RGB images, we progressively integrate the multi-level RGB encoder features with multi-level decoder features, thereby focusing more on appearance information. Following this way, we can acquire the final high-quality segmentation result. Extensive experiments are performed on three public datasets including MFNet, PST900 and FMB datasets, and the experimental results show that our method achieves competitive performance when compared with the 22 state-of-the-art methods. Xiaofei Zhou 0003, Liuxin Bao, Haibing Yin, Qiuping Jiang, Jiyong Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | No-Reference Point Cloud Quality Assessment via Graph Convolutional NetworkabstractThree-dimensional (3D) point cloud, as an emerging visual media format, is increasingly favored by consumers as it can provide more realistic visual information than two-dimensional (2D) data. Similar to 2D plane images and videos, point clouds inevitably suffer from quality degradation and information loss through multimedia communication systems. Therefore, automatic point cloud quality assessment (PCQA) is of critical importance. In this work, we propose a novel no-reference PCQA method by using a graph convolutional network (GCN) to characterize the mutual dependencies of multi-view 2D projected image contents. The proposed GCN-based PCQA (GC-PCQA) method contains three modules, i.e., multi-view projection, graph construction, and GCN-based quality prediction. First, multi-view projection is performed on the test point cloud to obtain a set of horizontally and vertically projected images. Then, a perception-consistent graph is constructed based on the spatial relations among different projected images. Finally, reasoning on the constructed graph is performed by GCN to characterize the mutual dependencies and interactions between different projected images, and aggregate feature information of multi-view projected images for final quality prediction. Experimental results on two publicly available benchmark databases show that our proposed GC-PCQA can achieve superior performance than state-of-the-art quality assessment metrics. Qiuping Jiang, Wei Zhou 0021, Feng Shao 0001, Guangtao Zhai, Weisi Lin |
IEEE Trans. Multim. | 2 |
| 2025 | Cross-Modal Hierarchical Knowledge Distillation for Image Aesthetics AssessmentabstractThe field of image aesthetics assessment (IAA) is rapidly advancing due to its wide applications. However, relying solely on single-modal information for aesthetic evaluation presents inherent limitations. While multimodal IAA models incorporating user comments have achieved significant advancements, these comments are often unavailable due to privacy concerns and practical considerations, and they also introduce additional computational overhead during inference. To address this issue, we propose a cross-modal hierarchical knowledge distillation method, termed HKD-IAA, to enhance the performance of unimodal image models effectively. Specifically, HKD-IAA comprises four components: feature extraction, feature decomposition, hierarchical knowledge distillation, and dynamic decay. During training, we first decompose the extracted features into a weighted sum of basic aesthetic elements and their corresponding weights, thereby reducing the learning difficulty for the student model. Building on this, we design a new hierarchical knowledge distillation framework, which aligns features at the feature, relation, and response levels to effectively transfer the knowledge from the teacher model. Finally, we introduce a dynamic decay strategy to adjust the weight of the distillation loss, thereby enhancing the student model's learning effectiveness during training. Extensive experiments on two benchmark datasets validate that the proposed method achieves state-of-the-art performance using only visual modal data. Our code is available athttps://github.com/Hangwei-Chen/HKD-IAA. Hangwei Chen, Feng Shao 0001, Weiyi Jing, Huizhi Wang, Qiuping Jiang |
IEEE Trans. Multim. | 5 |
| 2025 | Cross-Projection Distilling Knowledge for Omnidirectional Image Quality Assessment
Huixin Hu, Feng Shao 0001, Hangwei Chen, Xiongli Chai, Qiuping Jiang |
IEEE Trans. Multim. | 5 |
| 2025 | Multimodal Evidential Learning for Open-World Weakly-Supervised Video Anomaly DetectionabstractEfforts in weakly-supervised video anomaly detection center on detecting abnormal events within videos by coarse-grained labels, which has been successfully applied to many real-world applications. However, a significant limitation of most existing methods is that they are only effective for specific objects in specific scenarios, which makes them prone to misclassification or omission when confronted with previously unseen anomalies. Relative to conventional anomaly detection tasks, Open-world Weakly-supervised Video Anomaly Detection (OWVAD) poses greater challenges due to the absence of labels and fine-grained annotations for unknown anomalies. To address the above problem, we propose a multi-scale evidential vision-language model to achieve open-world video anomaly detection. Specifically, we leverage generalized visual-language associations derived from CLIP to harness the full potential of large pre-trained models in addressing the OWVAD task. Subsequently, we integrate a multi-scale temporal modeling module with a multimodal evidence collector to achieve precise frame-level detection of both seen and unseen anomalies. Extensive experiments on two widely-utilized benchmarks have conclusively validated the effectiveness of our method. The code will be made publicly available. Chao Huang 0008, Weiliang Huang, Qiuping Jiang, Wei Wang 0335, Jie Wen 0001, Bob Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Dataset and Metric for Quality Assessment of HDR Tone Mapping: Detail Visibility, Color Naturalness, and Overall QualityabstractTone-Mapping Operators (TMOs) aim at converting high dynamic range (HDR) images into standard dynamic range (SDR) ones that are suitable for being displayed on standard screens. As the visual quality of tone-mapped image (TMI) is paramount, conducting quality assessment of TMIs becomes crucial. Despite the growing body of research on TMI quality assessment, the existing metrics are often limited to a narrow selection of hand-picked examples generated by a restricted range of TMOs. Consequently, their ability of generalizing to the wide array of TMIs encountered in practical scenarios remains unclear. Moreover, the quality degradation in practical TMIs can be intricate, diverse, and complex. To overcome these limitations, we construct so far the largest subjective-annotated TMI quality assessment dataset which comprises a total number of 14,000 TMIs generated by applying 20 representative TMOs to 700 HDR images. The dataset is accompanied by subjective scores that encompass multiple quality dimensions, i.e., TMI quality dataset in terms of Detail visibility, Color naturalness, and overall Quality (TDCQ). In addition, we also design a multi-branch deep neural network tailored to characterize the multi-dimensional quality perception of TMIs, i.e., Color naturalness-, Detail visibility-aware TMI Quality (CDTIQ) metric, allowing for a comprehensive and multifaceted quality assessment of TMIs. Through extensive experiments, we demonstrate the superiority of our proposed metric, showcasing a higher correlation with subjective rating results compared to other relevant no-reference image quality metrics. Qiuping Jiang, Xiwen Li, Zhihua Wang 0002, Guangtao Zhai |
IEEE Trans. Multim. | 1 |
| 2025 | Underwater Image Enhancement With Cascaded Contrastive LearningabstractUnderwater image enhancement (UIE) is a highly challenging task due to the complexity of underwater environment and the diversity of underwater image degradation. Due to the application of deep learning, current UIE methods have made significant progress. Most of the existing deep learning-based UIE methods follow a single-stage network which cannot effectively address the diverse degradations simultaneously. In this paper, we propose to address this issue by designing a two-stage deep learning framework and taking advantage of cascaded contrastive learning to guide the network training of each stage. The proposed method is called CCL-Net in short. Specifically, the proposed CCL-Net involves two cascaded stages, i.e., a color correction stage tailored to the color deviation issue and a haze removal stage tailored to improve the visibility and contrast of underwater images. To guarantee the underwater image can be progressively enhanced, we also apply contrastive loss as an additional constraint to guide the training of each stage. In the first stage, the raw underwater images are used as negative samples for building the first contrastive loss, ensuring the enhanced results of the first color correction stage are better than the original inputs. While in the second stage, the enhanced results rather than the raw underwater images of the first color correction stage are used as the negative samples for building the second contrastive loss, thus ensuring the final enhanced results of the second haze removal stage are better than the intermediate color corrected results. Extensive experiments on multiple benchmark datasets demonstrate that our CCL-Net can achieve superior performance compared to many state-of-the-art methods. In addition, a series of ablation studies also verify the effectiveness of each key component involved in the proposed CCL-Net. Yi Liu 0085, Qiuping Jiang, Ting Luo 0001, Jingchun Zhou |
IEEE Trans. Multim. | 2 |
| 2025 | MISF-Net: Modality-Invariant and -Specific Fusion Network for RGB-T Crowd CountingabstractTo accurately perform crowd counting, utilizing the complementary relationship between RGB and thermal images to analyze the crowd has become the focus of current research. Due to different imaging principles, multi-modal images often contain different contents, which are their modality-specific information. For example, RGB images contain more texture and color details, while thermal images contain thermal radiation information. Meanwhile, they also describe the same target content, e.g., crowds, which are modality-invariant. However, existing methods only design different modules to directly fuse RGB and thermal image features, which did not fully consider the above facts. In this paper, by analyzing the similarities and differences between multi-modal images, we propose a Modality-Invariant and -Specific Fusion Network (MISF-Net) for RGB-T Crowd Counting. Specifically, we design a modality decomposition and fusion module (MDFM), which decomposes RGB and thermal image features into modality-invariant and -specific features by using the similarity and difference supervision between multi-modal features. Besides, reconstruction supervision is also used to prevent network learning from generating bias. After that, different fusion strategies are applied to the invariant and specific features, respectively. In addition, to adapt to the variations in size of different pedestrians, we design a modality-invariant fusion module (MIFM). Finally, after the fusion decoder, MISF-Net can obtain a more accurate crowd density map. Comprehensive experiments on the RGB-T crowd counting dataset show that our MISF-Net can achieve competitive performance. Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Hangwei Chen, Zhongjie Zhu, Qiuping Jiang |
IEEE Trans. Multim. | 6 |
| 2025 | Learning Video Salient Object Detection Progressively From Unlabeled VideosabstractRecently, deep learning-based video salient object detection (VSOD) has achieved some breakthroughs, but these methods rely on expensive annotated videos with pixel-wise annotations or weak annotations. In this paper, based on the similarities and differences between VSOD and image salient object detection (SOD), we propose a novel VSOD method via a progressive framework that locates and segments salient objects in sequence without utilizing any video annotation. To efficiently use the knowledge learned in the SOD dataset for VSOD efficiently, we introduce dynamic saliency to compensate for the lack of motion information of SOD during the locating process while maintaining the same fine segmenting process. Specifically, we utilize the coarse locating model trained on the image dataset, to identify frames with both static and dynamic saliency. Locating results of these frames are selected as spatiotemporal location labels. Moreover, by tracking salient objects in adjacent frames, the number of spatiotemporal location labels is increased. On the basis of these location labels, a two-stream locating network with an optical flow branch is proposed to capture salient objects in videos. The results with respect to five public benchmarks demonstrate that our method outperforms the state-of-the-art weakly and unsupervised methods. Binwei Xu, Qiuping Jiang, Haoran Liang 0001, Dingwen Zhang, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Multim. | 2 |
| 2025 | Progressive Region-to-Boundary Exploration Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) aims to segment targeted objects that have similar colors, textures, or shapes to their background environment. Due to the limited ability in distinguishing highly similar patterns, existing COD methods usually produce inaccurate predictions, especially around the boundary areas, when coping with complex scenes. This paper proposes a Progressive Region-to-Boundary Exploration Network (PRBE-Net) to accurately detect camouflaged objects. PRBE-Net follows an encoder-decoder framework and includes three key modules. Specifically, firstly, both high-level and low-level features of the encoder are integrated by a region and boundary exploration module to explore their complementary information for extracting the object's coarse region and fine boundary cues simultaneously. Secondly, taking the region cues as the guidance information, a Region Enhancement (RE) module is used to adaptively localize and enhance the region information at each layer of the encoder. Subsequently, considering that camouflaged objects usually have blurry boundaries, a Boundary Refinement (BR) decoder is used after the RE module to better detect the boundary areas with the assistance of boundary cues. Through top-down deep supervision, PRBE-Net can progressively refine the prediction. Extensive experiments on four datasets indicate that our PRBE-Net achieves superior results over 21 state-of-the-art COD methods. Additionally, it also shows good results on polyp segmentation, a COD-related task in the medical field. Guanghui Yue 0001, Shangjie Wu, Tianwei Zhou, Jie Du 0001, Yu Luo 0004, Qiuping Jiang |
IEEE Trans. Multim. | 7 |
| 2025 | RSUIA: Dynamic No-Reference Underwater Image Assessment via Reinforcement SequencesabstractUnderwater image quality assessment (UIQA) is a challenging task due to the complexities of underwater environments. Traditional UIQA methods primarily rely on fitting mean opinion scores (MOS), which are limited by human visual biases. To address the above limitation, we propose a no-reference underwater image quality assessment paradigm using reinforcement sequences. Our paradigm leverages reinforcement learning to iteratively merge the input image with the corresponding ground truth, generating an optimized sequence of images. A classifier generates probability arrays for the optimized sequence, which are converted into objective scores by a regression model. Unlike existing methods that focus solely on the final quality score, our paradigm emphasizes dynamic quality changes throughout the image-enhancement process. By employing objective mixing ratio labels, our reinforcement sequence dataset reduces subjective bias. The multiscale classifier captures local and global information differences between the input and ground truth images, effectively preserving the contrast and detail in diverse lighting conditions. Our paradigm combines multi-source data classification with support vector regression, optimizing the mapping of feature vectors to quality scores through fine-tuning libsvm kernel parameters. Experimental results on multiple benchmark datasets demonstrate that our paradigm outperforms the state-of-the-art UIQA methods, providing an effective solution for Underwater Image quality Assessment via Reinforcement Sequences (RSUIA). Jingchun Zhou, Chunjiang Liu, Dehuan Zhang, Zongxin He, Ferdous Sohel, Qiuping Jiang |
IEEE Trans. Multim. | 6 |
| 2025 | High-Precision Dichotomous Image Segmentation With Frequency and Scale AwarenessabstractDichotomous image segmentation (DIS) with rich fine-grained details within a single image is a challenging task. Despite the plausible results achieved by deep learning-based methods, most of them fail to segment generic objects when the boundary is cluttered with the background. In fact, the gradual decrease in feature map resolution during the encoding stage and the misleading texture clue may be the main issues. To handle these issues, we devise a novel frequency- and scale-aware deep neural network (FSANet) for high-precision DIS. The core of our proposed FSANet is twofold. First, a multimodality fusion (MF) module that integrates the information in spatial and frequency domains is adopted to enhance the representation capability of image features. Second, a collaborative scale fusion module (CSFM) which deviates from the traditional serial structures is introduced to maintain high resolution during the entire feature encoding stage. In the decoder side, we introduce hierarchical context fusion (HCF) and selective feature fusion (SFF) modules to infer the segmentation results from the output features of the CSFM module. We conduct extensive experiments on several benchmark datasets and compare our proposed method with existing state-of-the-art (SOTA) methods. The experimental results demonstrate that our FSANet achieves superior performance both qualitatively and quantitatively. The code will be made available at https://github.com/chasecjg/FSANet. Qiuping Jiang, Jinguang Cheng, Zongwei Wu, Runmin Cong, Radu Timofte |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Multiattentive Perception and Multilayer Transfer Network Using Knowledge Distillation for RGB-D Indoor Scene ParsingabstractScene parsing has gained wide attention in the field of computer vision, with emerging methods and techniques providing superior solutions. Although some methods have improved performance, they tend to neglect the number of model parameters and computational size, which makes achieving real-time operation in practical applications challenging. To address these limitations, we propose a multiattentive perception and multilayer transfer network that employs knowledge distillation (MPMTNet-KD), which is generated by a student network (MPMTNet-S) under the guidance of a teacher network (MPMTNet-T) with the aid of our proposed multilayer transfer knowledge distillation (KD) methods. To capture complete information from different modalities, a multiattentive perception module (MAPM) is introduced to mine features from various perspectives, and hetero-oriented sensing (HOS) convolution is utilized to integrate cross-layer features in a single and holistic manner. Importantly, we introduce multilayer transfer KD to explore the different knowledge types between layers, as well as intraclass and interclass correlations. In addition, we use the discrete cosine transform (DCT) approach combined with filtering during the KD process to mitigate noise that may be induced by the depth map, thereby improving the depth information and further enhancing the knowledge transfer effect. We conducted comprehensive experiments on two challenging indoor benchmark datasets, namely NYUDv2 and SUN RGB-D. Compared with existing methods, the proposed MPMTNet-KD reduces the number of parameters from 125.8 M in MPMTNet-T to 28.3 M in MPMTNet-S, achieving a mean intersection over union (mIoU) of 54.9% in the indoor scene parsing task. MPMTNet-KD was also evaluated on two additional public datasets, namely MFNet and PST900, to demonstrate its generalization capacity. The source code is available at https://github.com/XUEXIKUAIL/MPMTNet. Wujie Zhou, Bitao Jian, Yuanyuan Liu 0004, Qiuping Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | MorphQPV: Exploiting Isomorphism in Quantum Programs to Facilitate Confident VerificationabstractUnlike classical computing, quantum program verification (QPV) is much more challenging due to the non-duplicability of quantum states that collapse after measurement. Prior approaches rely on deductive verification that shows poor scalability. Or they require exhaustive assertions that cannot ensure the program is correct for all inputs. In this paper, we propose MorphQPV, a confident assertion-based verification methodology. Our key insight is to leverage the isomorphism in quantum programs, which implies a structure-preserve relation between the program runtime states. In the assertion statement, we define a tracepoint pragma to label the verified quantum state and an assume-guarantee primitive to specify the expected relation between states. Then, we characterize the ground-truth relation between states using an isomorphism-based approximation, which can effectively obtain the program states under various inputs while avoiding repeated executions. Finally, the verification is formulated as a constraint optimization problem with a confidence estimation model to enable rigorous analysis. Experiments suggest that MorphQPV reduces the number of program executions by 107.9× when verifying the 27-qubit quantum lock algorithm and improves the probability of success by 3.3×-9.9× when debugging five benchmarks. Siwei Tan, Debin Xiang, Liqiang Lu, Junlin Lu, Qiuping Jiang, Mingshuai Chen, Jianwei Yin |
ASPLOS (3) | 5 |
| 2024 | Multimodal Representation Distribution Learning for Medical Image Segmentation
Chao Huang 0008, Weichao Cai, Qiuping Jiang, Zhihua Wang 0002 |
IJCAI | 3 |
| 2024 | Progressive Point Cloud Denoising with Cross-Stage Cross-Coder Adaptive Edge Graph Convolution NetworkabstractDue to the limitation of collection device and unstable scanning process, point cloud data is usually noisy. This noise deforms the underlying structures of point clouds and inevitably affects downstream tasks such as rendering, reconstruction and classification. In this paper, we propose a Cross-stage Cross-coder Adaptive Edge Graph Convolution Network (C2AENet) to denoise point clouds. Our network uses multiple stages to progressively and iteratively denoise points. To improve the effectiveness, we add connections between two stages and between the encoder and decoder, leading to the cross-stage cross-coder architecture. Additionally, existing graph-based point cloud learning methods tend to capture the local structure. They typically construct a semantic graph based on semantic distance, which may ignore Euclidean neighbors and lead to insufficient geometry perception. Therefore, we introduce a geometric graph and adaptively calculate edge attention based on the local and global structural information of the points. This results in a novel graph convolution module that allows the network to capture richer contextual information and focus on more important parts. Extensive experiments demonstrate that the proposed method is competitive compared with other state-of-the-art methods. The code is available at: https://github.com/chenwuwq/C2AENet. Hehe Fan, Qiuping Jiang, Chao Huang 0008, Yi Yang 0001 |
ACM Multimedia | 3 |
| 2024 | CAFCNet: Cross-modality asymmetric feature complement network for RGB-T salient object detection
Dongze Jin, Feng Shao 0001, Zhengxuan Xie, Baoyang Mu, Hangwei Chen, Qiuping Jiang |
Expert Syst. Appl. | 6 |
| 2024 | Hallucinated-PQA: No reference point cloud quality assessment via injecting pseudo-reference features
Baoyang Mu, Feng Shao 0001, Hangwei Chen, Qiuping Jiang, Long Xu 0001, Yo-Sung Ho |
Expert Syst. Appl. | 4 |
| 2024 | CD-iNet: Deep Invertible Network for Perceptual Image Color Difference Measurement
Zhihua Wang 0002, Keshuo Xu, Keyan Ding, Qiuping Jiang, Yifan Zuo 0001, Zhangkai Ni, Yuming Fang 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | HCLR-Net: Hybrid Contrastive Learning Regularization with Locally Randomized Perturbation for Underwater Image EnhancementabstractUnderwater image enhancement presents a significant challenge due to the complex and diverse underwater environments that result in severe degradation phenomena such as light absorption, scattering, and color distortion. More importantly, obtaining paired training data for these scenarios is a challenging task, which further hinders the generalization performance of enhancement models. To address these issues, we propose a novel approach, the Hybrid Contrastive Learning Regularization (HCLR-Net). Our method is built upon a distinctive hybrid contrastive learning regularization strategy that incorporates a unique methodology for constructing negative samples. This approach enables the network to develop a more robust sample distribution. Notably, we utilize non-paired data for both positive and negative samples, with negative samples are innovatively reconstructed using local patch perturbations. This strategy overcomes the constraints of relying solely on paired data, boosting the model’s potential for generalization. The HCLR-Net also incorporates an Adaptive Hybrid Attention module and a Detail Repair Branch for effective feature extraction and texture detail restoration, respectively. Comprehensive experiments demonstrate the superiority of our method, which shows substantial improvements over several state-of-the-art methods in terms of quantitative metrics, significantly enhances the visual quality of underwater images, establishing its innovative and practical applicability. Our code is available at: https://github.com/zhoujingchun03/HCLR-Net . Jingchun Zhou, Chongyi Li, Qiuping Jiang, Man Zhou 0003, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu |
Int. J. Comput. Vis. | 4 |
| 2024 | Correction: HCLR-Net: Hybrid Contrastive Learning Regularization with Locally Randomized Perturbation for Underwater Image Enhancement
Jingchun Zhou, Chongyi Li, Qiuping Jiang, Man Zhou 0003, Kin-Man Lam 0001, Weishi Zhang, Xianping Fu |
Int. J. Comput. Vis. | 4 |
| 2024 | Visual Prompt Multibranch Fusion Network for RGB-Thermal Crowd CountingabstractAs population growth and urbanization continue, accurate crowd counting is increasingly important for public safety management and the Internet of Video Things (IOVT). However, RGB and thermal infrared (RGB-T) crowd counting still faces challenges in improving feature extraction capability for RGB streams and reducing multimodality differences. For this, we propose a visual prompt multibranch fusion network (VPMFNet) to tackle the above challenges. Specifically, to improve the ability of crowd analysis of the RGB stream in RGB-T crowd counting, through designing the prompt enhancement module, we take the prior features of head perception in the crowd as visual prompt cues to embed into the RGB stream. In terms of RGB and thermal image feature fusion, we fully reduce the modality differences from the perspectives of local fusion, global fusion, and multireceptive field fusion to accurately estimate the pedestrian number. Various experiments on two RGB-T crowd counting data sets demonstrate that our VPMFNet achieves a smaller estimation error in the number of pedestrians. Besides, our VPMFNet outperforms existing methods (i.e., multicolumn convolutional neural network, BL, SANet, UCNet, HDFNet, BBSNet, BL+IDAM, BL+CSCA, dual-branch enhanced feature fusion network, and GETANet) on the RGB-D data set. Our code will be released athttps://github.com/QSBAOYANGMU/VPMFNet. Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Hangwei Chen, Qiuping Jiang, Yo-Sung Ho |
IEEE Internet Things J. | 5 |
| 2024 | Hybrid CNN-transformer based meta-learning approach for personalized image aesthetics assessment
Xingao Yan, Feng Shao 0001, Hangwei Chen, Qiuping Jiang |
J. Vis. Commun. Image Represent. | 4 |
| 2024 | Vision-Language Consistency Guided Multi-Modal Prompt Learning for Blind AI Generated Image Quality AssessmentabstractRecently, textual prompt tuning has shown inspirational performance in adapting Contrastive Language-Image Pre-training (CLIP) models to natural image quality assessment. However, such uni-modal prompt learning method only tunes the language branch of CLIP models. This is not enough for adapting CLIP models to AI generated image quality assessment (AGIQA) since AGIs visually differ from natural images. In addition, the consistency between AGIs and user input text prompts, which correlates with the perceptual quality of AGIs, is not investigated to guide AGIQA. In this letter, we propose vision-language consistency guided multi-modal prompt learning for blind AGIQA, dubbed CLIP-AGIQA. Specifically, we introduce learnable textual and visual prompts in language and vision branches of CLIP models, respectively. Moreover, we design a text-to-image alignment quality prediction task, whose learned vision-language consistency knowledge is used to guide the optimization of the above multi-modal prompts. Experimental results on two public AGIQA datasets demonstrate that the proposed method outperforms state-of-the-art quality assessment models. Jun Fu 0007, Wei Zhou 0021, Qiuping Jiang, Hantao Liu, Guangtao Zhai |
IEEE Signal Process. Lett. | 3 |
| 2024 | Plain-PCQA: No-Reference Point Cloud Quality Assessment by Analysis of Plain Visual and Geometrical ComponentsabstractIn reviewing the research progress in Point Cloud Quality Assessment (PCQA), two main pathways have emerged, i.e., 2D projections and 3D point descriptors. The former primarily focuses on visual information, while the latter concentrates on crucial geometrical information in three-dimensional space. However, the current studies lack a thorough investigation of the impact of visual components and seldom pay special attention to plane-point fusion strategies. To comprehensively represent features and effectively tackle various types of impairments, we propose an end-to-end learning paradigm, only considering plain visual and geometrical factors called Plain-PCQA, for quantitatively evaluating objective metrics of 3D dense point clouds associated with human perception. Firstly, we explore a sophisticated preprocessing technique. The entire point clouds are packaged into six projections by moving virtual cameras, which can conveniently increase the visual samples during the training stage. Given the high resolution of the projected image, we have opted for a relatively lightweight network, namely ResNet-18, as the backbone to enable higher resolution input data. Five cropped patches from the projected image are collectively fed into this network. In light of the presence of some invalid information in the projections, a mask weight is devised to calculate the significance of each patch based on its effective informational content. Secondly, dual neural networks, comprising of a No-Reference (NR) branch and a Degraded-Reference (DR) branch, are designed with fundamental visual components to provide quantitative quality metrics. Specifically, the NR branch utilizes the feature output of each block in the Vision Transformer (ViT) model to obtain long-range low-level and high-level visual NR quality. The DR branch employs KLT (Karhunen-Loève Transform) to acquire the principal component information of an image as the macro-structural image, and then feeds the difference between input images and macro-structural images into a network for DR quality extraction. Thirdly, a Plane-Point Interaction Transformer (P2IT) is presented by incorporating texture and semantic features in 2D projections and geometrical features in 3D spaces to characterize the complete features with a connected 2D-3D feature representation. With these elaborately designed deep features, the proposed model can achieve competitive performances relying solely on plain visual and geometrical components. The experimental results demonstrate the potential of the proposed approach in multiple representative databases, which surpasses existing state-of-the-art methods significantly. Xiongli Chai, Feng Shao 0001, Baoyang Mu, Hangwei Chen, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Dynamic Hypergraph Convolutional Network for No-Reference Point Cloud Quality AssessmentabstractWith the rapid advancement of three-dimensional (3D) sensing technology, point cloud has emerged as one of the most important approaches for representing 3D data. However, quality degradation inevitably occurs during the acquisition, transmission, and process of point clouds. Therefore, point cloud quality assessment (PCQA) with automatic visual quality perception is particularly critical. In the literature, the graph convolutional networks (GCNs) have achieved certain performance in point cloud-related tasks. However, they cannot fully characterize the nonlinear high-order relationship of such complex data. In this paper, we propose a novel no-reference (NR) PCQA method with hypergraph learning. Specifically, a dynamic hypergraph convolutional network (DHCN) composing of a projected image encoder, a point group encoder, a dynamic hypergraph generator, and a perceptual quality predictor, is devised. First, a projected image encoder and a point group encoder are used to extract feature representations from projected images and point groups, respectively. Then, using the feature representations obtained by the two encoders, dynamic hypergraphs are generated during each iteration, aiming to constantly update the interactive information between the vertices of hypergraphs. Finally, we design the perceptual quality predictor to conduct quality reasoning on the generated hypergraphs. By leveraging the interactive information among hypergraph vertices, feature representations are well aggregated, resulting in a notable improvement in the accuracy of quality pediction. Experimental results on several point cloud quality assessment databases demonstrate that our proposed DHCN can achieve state-of-the-art performance. The code will be available at:https://github.com/chenwuwq/DHCN. Qiuping Jiang, Wei Zhou 0021, Long Xu 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Rethinking and Conceptualizing Just Noticeable Difference Estimation by Residual LearningabstractThe human visual system (HVS) cannot perceive the pixel intensity change below a certain threshold which is also known as the just noticeable difference (JND). Conventional JND prediction models mainly follow a two-step pipeline by first modeling the diverse masking effects based on the findings of the HVS and then fusing the results of different masking effect models into an overall JND map. However, due to the insufficient understanding of the HVS properties at the current stage, it is difficult to devise accurate computational models to characterize the complex masking effects. Moreover, the reasonability of the manually designed fusion schemes also lacks justification. In this work, we rethink the JND estimation problem from a fresh perspective by conceptualizing the JND as the difference map between the pristine image and its corresponding Critical Perceptual Lossless (CPL) counterpart. Building on this insight, we introduce a deep residual learning framework called ResJND to learn the discrepancies between the pristine image and its CPL counterpart, aiming to predict JND map implicitly. To support the training of our proposed ResJND model, we construct a dedicated CPL image dataset called CPL-Set which comprises a collection of pristine images and their corresponding CPL images selected by thorough subjective experiments. Comprehensive experiments have conclusively shown that our ResJND model excels at accurately predicting the JND map. Additionally, it demonstrates superior performance in associated applications, such as JND-guided noise injection, JND-guided image compression, and distortion visibility prediction. Codes are available at: https://github.com/Knife646/ResJND. Qiuping Jiang, Zhihua Wang 0002, Shiqi Wang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Dynamic Weighted Fusion and Progressive Refinement Network for Visible-Depth-Thermal Salient Object DetectionabstractThe introduction of depth/thermal modality has significantly enhanced the performance of dual-modal salient object detection (SOD) methods. However, depth maps and thermal images are prone to environmental interference, making them insufficient for providing salient information. To address this challenge, triple-modal SOD methods have been proposed. However, these methods often overlook the detrimental effects of defective modalities during fusion, leading to subpar performance. To tackle this issue, we present a novel dynamic weighted fusion and progressive refinement network (DWFPRNet) for Visible-Depth-Thermal (V-D-T) SOD. Specifically, we first use the dual-modal fusion module (DFM) to fuse dual modalities, thereby obtaining fused features. Subsequently, the modality selective fusion module (MSFM) mines complementary information between fused features, considering both fusion features and the quality of feature maps, to achieve weighted fusion. Finally, we design a progressive refinement decoder (PRD) to realize interaction and multi-scale learning among different scale features and generate high-quality saliency maps. Extensive experiments conducted on the VDT-2048 public dataset demonstrate that our method outperforms existing state-of-the-art multi-modal methods. Feng Shao 0001, Baoyang Mu, Hangwei Chen, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Dual-Constraint Coarse-to-Fine Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) is an important yet challenging task, with great application values in industrial defect detection, medical care, etc. The challenges mainly come from the high intrinsic similarities between target objects and background. In this paper, inspired by the biological studies that object detection consists of two steps, i.e., search and identification, we propose a novel framework, named DCNet, for accurate COD. DCNet explores candidate objects and extra object-related edges through two constraints (object area and boundary) and detects camouflaged objects in a coarse-to-fine manner. Specifically, we first exploit an area-boundary decoder (ABD) to obtain initial region cues and boundary cues simultaneously by fusing multi-level features of the backbone. Then, an area search module (ASM) is embedded into each level of the backbone to adaptively search coarse regions of objects with the assistance of region cues from the ABD. After the ASM, an area refinement module (ARM) is utilized to identify fine regions of objects by fusing adjacent-level features with the guidance of boundary cues. Through the deep supervision strategy, DCNet can finally localize the camouflaged objects precisely. Extensive experiments on three benchmark COD datasets demonstrate that our DCNet is superior to 12 state-of-the-art COD methods. In addition, DCNet shows promising results on two COD-related tasks, i.e., industrial defect detection and polyp segmentation. Guanghui Yue 0001, Houlu Xiao, Hai Xie, Tianwei Zhou, Wei Zhou 0021, Weiqing Yan, Baoquan Zhao, Tianfu Wang 0001, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2024 | Modal Evaluation Network via Knowledge Distillation for No-Service Rail Surface Defect DetectionabstractDeep learning techniques have largely solved the problem of rail surface defect detection (SDD), however, two aspects have yet to be addressed. In most existing approaches, two red–green–blue and depth (RGB-D) streams are indiscriminately fused across modalities, ignoring the fact that RGB and depth images produce different feature qualities in different scenes. Additionally, in their focus on performance, previous studies have overlooked the fact that models produce several parameters, resulting in unrealistic practical applications. To address these challenges, we designed a modal evaluation network (MENet) via knowledge distillation (KD) (MENet-S*) for a no-service rail SDD to adaptively manage information in each scenario and achieve model compression. First, to dynamically adjust the feature distribution and quality, dynamic and static feature coding ideas are introduced. Second, modal evaluation distillation is introduced, which allows a compact model (MENet-S) to learn the feature evaluation process of a complex model (MENet-T). Third, to enable MENet-S to learn the dynamic encoding process of MENet-T and to improve the feature representation of MENet-S, we propose accessible knowledge distillation. Furthermore, multitiered KD is introduced to facilitate the learning of MENet-S. Based on extensive experiments using the industrial RGB-D dataset NEU RSDDS-AUG, we observed that MENet-S* (MENet-S with KD) outperformed 16 state-of-the-art methods. In addition, to demonstrate the generalization capability of MENet-S*, we evaluated the proposed network on three additional public datasets, and MENet-S* achieved competitive results. The source codes and results are available at https://github.com/hjklearn/MENet-KD. Wujie Zhou, Jiankang Hong, Weiqing Yan, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | DGPINet-KD: Deep Guided and Progressive Integration Network With Knowledge Distillation for RGB-D Indoor Scene AnalysisabstractSignificant advancements in RGB-D semantic segmentation have been made owing to the increasing availability of robust depth information. Most researchers have combined depth with RGB data to capture complementary information in images. Although this approach improves segmentation performance, it requires excessive model parameters. To address this problem, we propose DGPINet-KD, a deep-guided and progressive integration network with knowledge distillation (KD) for RGB-D indoor scene analysis. First, we used branching attention and depth guidance to capture coordinated, precise location information and extract more complete spatial information from the depth map to complement the semantic information for the encoded features. Second, we trained the student network (DGPINet-S) with a well-trained teacher network (DGPINet-T) using a multilevel KD. Third, an integration unit was developed to explore the contextual dependencies of the decoding features and to enhance relational KD. Comprehensive experiments on two challenging indoor benchmark datasets, NYUDv2 and SUN RGB-D, demonstrated that DGPINet-KD achieved improved performance in indoor scene analysis tasks compared with existing methods. Notably, on the NYUDv2 dataset, DGPINet-KD (DGPINet-S with KD) achieves a pixel accuracy gain of 1.7% and a class accuracy gain of 2.3% compared with DGPINet-S. In addition, compared with DGPINet-T, the proposed DGPINet-KD (DGPINet-S with KD) utilizes significantly fewer parameters (29.3M) while maintaining accuracy. The source code is available at https://github.com/XUEXIKUAIL/DGPINet. Wujie Zhou, Bitao Jian, Meixin Fang, Xiena Dong, Yuanyuan Liu 0004, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Weakly Supervised Video Anomaly Detection via Self-Guided Temporal Discriminative TransformerabstractWeakly supervised video anomaly detection is generally formulated as a multiple instance learning (MIL) problem, where an anomaly detector learns to generate frame-level anomaly scores under the supervision of MIL-based video-level classification. However, most previous works suffer from two drawbacks: 1) they lack ability to model temporal relationships between video segments and 2) they cannot extract sufficient discriminative features to separate normal and anomalous snippets. In this article, we develop a weakly supervised temporal discriminative (WSTD) paradigm, that aims to leverage both temporal relation and feature discrimination to mitigate the above drawbacks. To this end, we propose a transformer-styled temporal feature aggregator (TTFA) and a self-guided discriminative feature encoder (SDFE). Specifically, TTFA captures multiple types of temporal relationships between video snippets from different feature subspaces, while SDFE enhances the discriminative powers of features by clustering normal snippets and maximizing the separability between anomalous snippets and normal centers in embedding space. Experimental results on three public benchmarks indicate that WSTD outperforms state-of-the-art unsupervised and weakly supervised methods, which verifies the superiority of the proposed method. Chao Huang 0008, Chengliang Liu 0003, Jie Wen 0001, Lian Wu, Yong Xu 0001, Qiuping Jiang, Yaowei Wang 0001 |
IEEE Trans. Cybern. | 6 |
| 2024 | MSTNet-KD: Multilevel Transfer Networks Using Knowledge Distillation for the Dense Prediction of Remote-Sensing ImagesabstractRecently, methods based on convolutional neural networks have achieved good results in the dense prediction of remote-sensing images, particularly when employing normalized digital surface models. However, most existing methods use multiscale convolution and attention methods to mine multimodal feature information without considering the differences and complementarities between the two features. Moreover, previous studies have prioritized model segmentation performance and ignored parametric issues, which makes it difficult to deploy the model in practical applications. To address this challenge, we designed a multilevel semantic transfer network (MSTNet) for the dense prediction of remote-sensing images using a knowledge-distillation approach to adaptively select useful semantic information for the transfer network. We designed a multilevel semantic knowledge alignment distillation framework (MSKA) to enable a compact student model to learn the semantic information extracted from a complex model. The MSKA framework comprises three main components: cross-layer semantic alignment, dynamic semantic aggregation, and softening learning for semantic information transfer and predictive label softening. Experiments on the Vaihingen and Potsdam datasets showed that the student network employing the MSKA framework achieved excellent segmentation performance with only 8.88M parameters and 2.09 gigaFLOPs in terms of computational costs compared with current state-of-the-art methods. The source code and results are available at https://github.com/LYZ00918/MSKANet. Wujie Zhou, Yangzhen Li, Juan Huan, Yuanyuan Liu 0004, Qiuping Jiang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Remote Sensing Image Scene Classification via Graph Template Enhancement and Supplementation Network With Dual-Teacher Knowledge DistillationabstractImage-processing techniques used for remote sensing images (RSIs) have addressed the issues associated with scene classification. However, two challenges remain; first, most bi-modal methods integrate the extracted features indiscriminately while neglecting certain more relevant features, and bi-modal feature qualities may vary in different scenarios. Second, previous studies have improved the feature extraction performance at the expense of computational speed and complexity, hindering the wide application of improved versions. To address these issues, we constructed a graph template enhancement and supplementation network (GTESNet) with dual-teacher knowledge distillation (KD), GTESNet-S$^{\ast }$, to emphasize and manage the extracted features adaptively and compress the model. First, we developed a graph template enhancement (GTE) module that combined long-range contextual information with clustered template features. Second, an exchange-correlation fusion (ECF) module was introduced to allocate feature distribution dynamically and achieve integration. Third, we constructed feedback complementary decoders (FCDs) consisting of two subdecoders with a cascade connection that used a feedback mechanism to supplement the final output. A dual-teacher distillation method that utilized simple confidence perceptions to assign weights to both teachers was implemented during KD. Furthermore, we introduced clustered templates and graph-relation distillations (GRDs), allowing the student GTESNet-S to learn the clustering ability and semantic association process of the teacher GTESNet-T. Finally, we introduced simple progressive decoder distillation (PDD) to facilitate the learning of GTESNet-S. Extensive experiments on two publicly available datasets demonstrated that the GTESNet outperformed similar state-of-the-art (SOTA) methods. The code is available at:https://github.com/MAXHAN22/GTESNet. Wujie Zhou, Penghan Yang, Yuanyuan Liu 0004, Runmin Cong, Qiuping Jiang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | DTKD-Net: Dual-Teacher Knowledge Distillation Lightweight Network for Water-Related Optics Image EnhancementabstractWater-related optics images are often degraded by absorption and scattering effects. Current underwater image enhancement (UIE) methods improve image quality but neglect the constraints of underwater imaging environments. To address this issue, we propose a double-teacher knowledge distilling network (DTKD-Net), which uses a dynamic teaching strategy within a dual-teacher framework to enhance knowledge distillation (KD), improving the student network’s ability to capture complex underwater features. Specifically, DTKD-Net focuses on clear-to-clear and blurry-to-clear image learning to enhance underwater images. It aims to preserve details in clear images and restore blurred ones. The dual-teacher network uses an intermediate layer with the middle layer of the student network to compute feature differences for feature guidance. The network uses a dynamic strategy where a Teacher-Sub stops guidance when its output matches the student’s, which helps with contrastive learning and improves the network’s ability to handle complex underwater scenes. Extensive experiments and visual comparisons show that DTKD-Net reduces the model size, demonstrating superior efficiency and effectiveness in enhancing underwater images. Jingchun Zhou, Dehuan Zhang, Gemine Vivone, Qiuping Jiang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | DSANet-KD: Dual Semantic Approximation Network via Knowledge Distillation for Rail Surface Defect DetectionabstractOwing to the development of convolutional neural networks (CNNs), the detection of defects on rail surfaces has significantly improved. Although existing methods achieve good results, they incur huge computational and parameter costs associated with CNNs. The usual approach to this problem is to design lightweight models that meet the needs of real-world applications; however, the performance is often compromised. To address the aforementioned problems, we designed a dual semantic approximation network via knowledge distillation (DSANet-KD, a student model with knowledge distillation) for rail surface defect detection; it focuses on both foreground and background knowledge and obtains more accurate prediction results. This model comprises an adaptive 3D spatial integration module, feature-optimization decoding module, and dual semantic approximation knowledge-distillation framework. Specifically, we employed a thoroughly trained teacher defect detection network equipped with dual semantic approximation information as an experienced teacher to guide the training of a student defect detection network. Experimental results showed that the proposed DSANet-KD achieved better accuracy with a smaller number of parameters than the state-of-the-art methods. To demonstrate the generalizability of DSANet-KD, experiments were conducted on a publicly available RGBD-SOD dataset, whose source code is available at: https://github.com/hjklearn/DSANet-KD. Wujie Zhou, Jiankang Hong, Xiaoxiao Ran, Weiqing Yan, Qiuping Jiang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Blind Quality Evaluator of Light Field Images by Group-Based Representations and Multiple Plane-Oriented Perceptual CharacteristicsabstractDue to the emergency of multi-view cameras and commercial Light Field (LF) cameras, the demand of high-performance LF quality evaluator is of great significance for guiding LF acquisition, processing and application and further promoting the visual perceived quality of LF visualizations. However, LF Images (LFIs), as high-dimensional data, suffer from various quality degradations not only in the spatial domain but also in the angular domain. Therefore, it is of great challenge to predict LF quality accurately. An effective LF evaluator should be able to represent these heterogeneous artifacts. In this paper, we provide a novel No-Reference LF Quality Assessment Evaluator (NR LF-QAE) to tackle this problem. Firstly, to measure angular consistency among viewports, we utilize group-based representations to character information similarity of aligned view stacks. Secondly, to better describe the texture information of LFIs, unifying spatial-angular texture statistic measurement is performed via Local Binary Patterns from Three Orthogonal Planes (LBP-TOP). Thirdly, we design 3D Log-Gabor filters to extract LF global structure information in Sub-Aperture Images (SAIs) as spatial feature characterizations and 2D Log-Gabor filters are adopted to characterize ray direction/depth information in Epipolar Plane Images (EPIs) as angular feature characterizations. By comprehensive LF information analyses in angular consistency and spatial-angular feature extraction with texture and structure descriptors, experimental results demonstrate the superiority of the proposed NR LF-QAE over the state-of-the-art comparative models in predicting the quality of LFIs on three available benchmark databases. The code will be released athttps://github.com/zerosola/NR-LF-QAE. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xuejin Wang, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 3 |
| 2024 | Perception-Driven Deep Underwater Image Enhancement Without Paired SupervisionabstractUnderwater image enhancement (UIE) aims to improve the visual quality of raw underwater images. Current UIE algorithms primarily train a deep neural network (DNN) on synthetic datasets or datasets with pseudo labels by minimizing the reconstruction loss between enhanced images and ground truth images. However, there is a domain gap between synthetic and real-world underwater images, and the widely used$\ell _{1}$or$\ell _{2}$loss tends to overlook the importance of human perception, resulting in unsatisfactory perceptual quality of the final enhanced results. In this paper, we propose an unsupervised perception-driven DNN called PDD-Net for generalizable UIE. Instead of relying on paired images for training, we resort to an unsupervised generative adversarial network (GAN) with a large-scale set of easily available natural images as the target domain. This enables training on larger image sets collected from various domains while avoiding over-fitted to any specific data generation protocol. Additionally, to make the visual quality of enhanced underwater images more in line with human perception, we pre-train a DNN-based pairwise quality ranking (PQR) model based on which a PQR loss is formulated to progressively guides the enhancement of raw underwater image toward the higher quality direction. In addition, we introduce a global attention module (GAM) that integrates modulation and attention mechanisms to enable capturing rich global and local information, leading to improvements in both brightness and contrast. Extensive experiments demonstrate that our proposed PDD-Net exhibits excellent generalization capabilities and outperforms existing methods in terms of both visual perception quality and quantitative indicators across different datasets. Qiuping Jiang, Yaozu Kang, Zhihua Wang 0002, Wenqi Ren, Chongyi Li |
IEEE Trans. Multim. | 1 |
| 2024 | Benchmark Dataset and Pair-Wise Ranking Method for Quality Evaluation of Night-Time Image EnhancementabstractNight-time image enhancement (NIE) aims at boosting the intensity of low-light regions while suppressing noises or light effects in night-time images, and numerous efforts have been made for this task. However, few explorations focus on the quality evaluation issue of enhanced night-time images (ENTIs), and how to fairly compare the performance of different NIE algorithms remains a challenging problem. In this paper, we firstly construct a new Real-world Night-Time Image Enhancement Quality Assessment (i.e., RNTIEQA) dataset that includes two typical types of night-time scenes (i.e., extremely low light and uneven light scenes), and carry out human subjective studies to compare the quality of ENTIs obtained by a set of representative NIE algorithms. Afterwards, a new objective ranking method that comprehensively considering image intrinsic and impairment attributes is proposed for automatically predicting the quality of ENTIs. Experimental results on our RNTIEQA dataset demonstrate that the proposed method outperforms the off-the-shelf competitors. Our dataset and code will be released athttps://github.com/Leilei-Huang-work/RNTIEQA-dataset. Xuejin Wang, Leilei Huang, Hangwei Chen, Qiuping Jiang, ShaoWei Weng, Feng Shao 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Perceptual Quality Assessment of Retouched Face ImagesabstractNowadays, it is a common practice to retouch face images before sharing them on websites, social media, and even identification cards. In response, increased criticisms have appeared about taking photo retouching to an extreme. This naturally leads to the necessity of designing perceptual quality assessment methods that can measure how much a retouched face image has strayed from reality. However, such an issue has seldom been considered. In this paper, we conduct both subjective and objective studies to advance this field. Firstly, we construct a benchmark database (termed SZU-RFD) via subjective experiments. SZU-RFD consists of 200 high-quality images with Asian faces and 1,600 retouched images generated by three popular photo-editing tools under different settings. Secondly, considering that retouching usually distorts the image texture, we propose a novel no-reference (NR) quality assessment method, named TANet, for retouched face images by taking the textural artifact into account. Specifically, a texture enhancement module is embedded into the shallow layer to help the network focus on textural information, and a multi-task learning strategy is applied to improve the performance of the main task with the assistance of an auxiliary task, i.e., texture recognition. Extensive experiments on the constructed SZU-RFD show that our proposed TANet correlates well with subjective perceptual judgments and is superior to 19 mainstream NR image quality assessment methods in evaluating retouched face images. Guanghui Yue 0001, Honglv Wu, Qiuping Jiang, Tianwei Zhou, Weiqing Yan, Tianfu Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | A Pixel Distribution Remapping and Multi-Prior Retinex Variational Model for Underwater Image EnhancementabstractHigh-quality underwater imaging is crucial for underwater exploration. However, particle scattering and light absorption by seawater significantly degrade image clarity. To address these issues, we propose a novel underwater image enhancement (UIE) method that combines pixel distribution remapping (PDR) with a multi-priority Retinex variational model. We design a pre-compensation method for severely attenuated channels that effectively prevents new color artifacts during color correction. By combining the inter-channel coupling relationships, we compute a limiting factor to remap pixel distribution curves to improve image contrast. In addition, considering the significant noise interference, we introduce the prior knowledge, including underwater noise and texture priors, while constructing the variational model, and design penalty terms that match the underwater characteristics to remove excessive noise in the reflectance component. Our approach efficiently decouples the illumination and reflectance components using a rapid solver. Subsequently, gamma correction adjusts the illumination component, and the corrected illumination and reflectance components are fused to reconstruct the final natural output image. Comprehensive evaluations across various datasets reveal that our approach significantly surpasses current state-of-the-art (SOTA) methods. These results demonstrate the effectiveness of our method in correcting color bias and compensating for luminance losses in underwater imagery. Our code is available at:https://github.com/zhoujingchun03/PDRMRV. Jingchun Zhou, Shiyin Wang, Zifan Lin, Qiuping Jiang, Ferdous Sohel |
IEEE Trans. Multim. | 4 |
| 2024 | Collaborative Learning and Style-Adaptive Pooling Network for Perceptual Evaluation of Arbitrary Style TransferabstractAlthough the research of arbitrary style transfer (AST) has achieved great progress in recent years, few studies pay special attention to the perceptual evaluation of AST images that are usually influenced by complicated factors, such as structure-preserving, style similarity, and overall vision (OV). Existing methods rely on elaborately designed hand-crafted features to obtain quality factors and apply a rough pooling strategy to evaluate the final quality. However, the importance weights between the factors and the final quality will lead to unsatisfactory performances by simple quality pooling. In this article, we propose a learnable network, named collaborative learning and style-adaptive pooling network (CLSAP-Net) to better address this issue. The CLSAP-Net contains three parts, i.e., content preservation estimation network (CPE-Net), style resemblance estimation network (SRE-Net), and OV target network (OVT-Net). Specifically, CPE-Net and SRE-Net use the self-attention mechanism and a joint regression strategy to generate reliable quality factors for fusion and weighting vectors for manipulating the importance weights. Then, grounded on the observation that style type can influence human judgment of the importance of different factors, our OVT-Net utilizes a novel style-adaptive pooling strategy guiding the importance weights of factors to collaboratively learn the final quality based on the trained CPE-Net and SRE-Net parameters. In our model, the quality pooling process can be conducted in a self-adaptive manner because the weights are generated after understanding the style type. The effectiveness and robustness of the proposed CLSAP-Net are well validated by extensive experiments on the existing AST image quality assessment (IQA) databases. Our code will be released at https://github.com/Hangwei-Chen/CLSAP-Net. Hangwei Chen, Feng Shao 0001, Xiongli Chai, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Blind Quality Assessment of Dense 3D Point Clouds with Structure Guided ResamplingabstractObjective quality assessment of three-dimensional (3D) point clouds is essential for the development of immersive multimedia systems in real-world applications. Despite the success of perceptual quality evaluation for 2D images and videos, blind/no-reference metrics are still scarce for 3D point clouds with large-scale irregularly distributed 3D points. Therefore, in this article, we propose an objective point cloud quality index with Structure Guided Resampling (SGR) to automatically evaluate the perceptually visual quality of dense 3D point clouds. The proposed SGR is a general-purpose blind quality assessment method without the assistance of any reference information. Specifically, considering that the human visual system is highly sensitive to structure information, we first exploit the unique normal vectors of point clouds to execute regional pre-processing that consists of keypoint resampling and local region construction. Then, we extract three groups of quality-related features, including (1) geometry density features, (2) color naturalness features, and (3) angular consistency features. Both the cognitive peculiarities of the human brain and naturalness regularity are involved in the designed quality-aware features that can capture the most vital aspects of distorted 3D point clouds. Extensive experiments on several publicly available subjective point cloud quality databases validate that our proposed SGR can compete with state-of-the-art full-reference, reduced-reference, and no-reference quality assessment algorithms. Wei Zhou 0021, Qi Yang 0003, Qiuping Jiang, Guangtao Zhai, Weisi Lin |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Object Segmentation by Mining Cross-Modal SemanticsabstractMulti-sensor clues have shown promise for object segmentation, but inherent noise in each sensor, as well as the calibration error in practice, may bias the segmentation accuracy. In this paper, we propose a novel approach by mining the Cross-Modal Semantics to guide the fusion and decoding of multimodal features, with the aim of controlling the modal contribution based on relative entropy. We explore semantics among the multimodal inputs in two aspects: the modality-shared consistency and the modality-specific variation. Specifically, we propose a novel network, termed XMSNet, consisting of (1) all-round attentive fusion (AF), (2) coarse-to-fine decoder (CFD), and (3) cross-layer self-supervision. On the one hand, the AF block explicitly dissociates the shared and specific representation and learns to weight the modal contribution by adjusting the proportion, region, and pattern, depending upon the quality. On the other hand, our CFD initially decodes the shared feature and then refines the output through specificity-aware querying. Further, we enforce semantic consistency across the decoding layers to enable interaction across network hierarchies, improving feature discriminability. Exhaustive comparison on eleven datasets with depth or thermal clues, and on two challenging tasks, namely salient and camouflage object segmentation, validate our effectiveness in terms of both performance and robustness. The source code is publicly available at https://github.com/Zongwei97/XMSNet. Zongwei Wu, Zhuyun Zhou, Zhaochong An, Qiuping Jiang, Cédric Demonceaux, Guolei Sun, Radu Timofte |
ACM Multimedia | 5 |
| 2023 | TCCL-Net: Transformer-Convolution Collaborative Learning Network for Omnidirectional Image Super-ResolutionabstractAs virtual reality and metaverse become more and more popular, the Omnidirectional Image (OI) has attracted extreme attention due to its immersive display characteristics. However, users only watch a portion of the content in a specific viewport extracted from a panoramic view, which will lead to a problem of resolution mismatch that requires High-Resolution (HR) for clear near-eye displays in viewports. Hence, it is necessary to exploit a Super Resolution (SR) solution for reconstructing Low-Resolution (LR) OIs. Different from 2D SR methods, the variation of pixel distributions along latitudes is a critical factor in designing an Omnidirectional Image Super-Resolution (OISR) scheme. In this paper, we put forward a novel end-to-end network with a Transformer and Convolution Collaborative Learning Network (TCCL-Net) for OISR. Firstly, Swin Transformer blocks and residual convolution blocks are employed to extract long-range and short-range dependencies, thereby digging into more rich and heterogeneous features from these two branches. Secondly, to better fuse these two features, cross-guided enhanced attention mechanisms are designed for bidirectional information enhancement onto both channel and spatial features . Thirdly, to alleviate nonuniformly pixel distributions across latitudes, we add an absolute positional encoding into Swin Transformer to represent patch weights at different positions and propose a tile-based panoramic reconstruction module to super-resolve various bands with different pixel sampling characteristics across latitudes. Experimental results on two available benchmark datasets demonstrate the superiority of the proposed approach over the state-of-the-art method in achieving OISR task. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Hongwei Ying |
Knowl. Based Syst. | 3 |
| 2023 | Modality-Induced Transfer-Fusion Network for RGB-D and RGB-T Salient Object DetectionabstractThe ability of capturing the complementary information of multi-modality data is critical to the development of multi-modality salient object detection (SOD). Most of existing studies attempt to integrate multi-modality information through various fusion strategies. However, most of these methods ignore the inherent differences in multi-modality data, resulting in poor performance when dealing with some challenging scenarios. In this paper, we propose a novel Modality-Induced Transfer-Fusion Network (MITF-Net) for RGB-D and RGB-T SOD by fully exploring the complementarity in multi-modality data. Specifically, we first deploy a modality transfer fusion (MTF) module to bridge the semantic gap between single and multi-modality data, and then mine the cross-modality complementarity based on point-to-point structural similarity information. Then, we design a cycle-separated attention (CSA) module to optimize the cross-layer information recurrently, and measure the effectiveness of cross-layer features through point-wise convolution-based multi-scale channel attention. Furthermore, we refine the boundaries in the decoding stage to obtain high-quality saliency maps with sharp boundaries. Extensive experiments on 13 RGB-D and RGB-T SOD datasets show that the proposed MITF-Net achieves a competitive and excellent performance. Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Quality Evaluation of Arbitrary Style Transfer: Subjective Study and Objective MetricabstractArbitrary neural style transfer is a vital topic with great research value and wide industrial application, which strives to render the structure of one image using the style of another. Recent researches have devoted great efforts on the task of arbitrary style transfer (AST) for improving the stylization quality. However, there are very few explorations about the quality evaluation of AST images, even it can potentially guide the design of different algorithms. In this paper, we first construct a new AST images quality assessment database (AST-IQAD), which consists 150 content-style image pairs and the corresponding 1200 stylized images produced by eight typical AST algorithms. Then, a subjective study is conducted on our AST-IQAD database, which obtains the subjective rating scores of all stylized images on the three subjective evaluations, i.e., content preservation (CP), style resemblance (SR), and overall vision (OV). To quantitatively measure the quality of AST image, we propose a new sparse representation-based method, which computes the quality according to the sparse feature similarity. Experimental results on our AST-IQAD have demonstrated the superiority of the proposed method. The dataset and source code will be released athttps://github.com/Hangwei-Chen/AST-IQAD-SRQE Hangwei Chen, Feng Shao 0001, Xiongli Chai, Yuese Gu, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Bidirectional Collaborative Mentoring Network for Marine Organism Detection and BeyondabstractOrganism detection plays a vital role in marine resource exploitation and marine economy. How to accurately locate the target organism object within the camouflaged and dark light oceanic scene has recently drawn great attention in the research community. Existing learning-based works usually leverage local texture details within a neighboring area, with few methods explicitly exploring the usage of contextualized awareness for accurate object detection. From a novel perspective, we in this work present a Bidirectional Collaborative Mentoring Network (BCMNet) which fully explores both texture and context clues during the encoding and decoding stages, making the cross-paradigm interaction bidirectional and improving the scene understanding at all stages. Specifically, we first extract texture and context features through a dual-branch encoder and attentively fuse them through our adjacent feature fusion (AFF) block. Then, we propose a structure-aware module (SAM) and a detail-enhanced module (DEM) to form our two-stage decoding pipeline. On the one hand, our SAM leverages both local and global clues to preserve morphological integrity and generate an initial prediction of the target object. On the other hand, the DEM explicitly explores long-range dependencies to refine the initially predicted object mask further. The combination of SAM and DEM enables better extracting, preserving, and enhancing the object morphology, making it easier to segment the target object from the camouflaged background with sharp contour. Extensive experiments on three benchmark datasets show that our proposed BCMNet performs favorably over state-of-the-art models. The code will be made available athttps://github.com/chasecjg/BCMNet. Jinguang Cheng, Zongwei Wu, Shuo Wang 0010, Cédric Demonceaux, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | A Weakly Supervised Learning Framework for Salient Object Detection via Hybrid LabelsabstractFully-supervised salient object detection (SOD) methods have made great progress, but such methods often rely on a large number of pixel-level annotations, which are time-consuming and labour-intensive. In this paper, we focus on a new weakly-supervised SOD task under hybrid labels, where the supervision labels include a large number of coarse labels generated by the traditional unsupervised method and a small number of real labels. To address the issues of label noise and quantity imbalance in this task, we design a new pipeline framework with three sophisticated training strategies. In terms of model framework, we decouple the task into label refinement sub-task and salient object detection sub-task, which cooperate with each other and train alternately. Specifically, the R-Net is designed as a two-stream encoder-decoder model equipped with Blender with Guidance and Aggregation Mechanisms (BGA), aiming to rectify the coarse labels for more reliable pseudo-labels, while the S-Net is a replaceable SOD network supervised by the pseudo labels generated by the current R-Net. Note that, we only need to use the trained S-Net for testing. Moreover, in order to guarantee the effectiveness and efficiency of network training, we design three training strategies, including alternate iteration mechanism, group-wise incremental mechanism, and credibility verification mechanism. Experiments on five SOD benchmarks show that our method achieves competitive performance against weakly-supervised/unsupervised methods both qualitatively and quantitatively. The code and results can be found from the link ofhttps://rmcong.github.io/proj_Hybrid-Label-SOD.html. Runmin Cong, Chen Zhang 0013, Qiuping Jiang, Shiqi Wang 0001, Yao Zhao 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | A Perception-Aware Decomposition and Fusion Framework for Underwater Image EnhancementabstractThis paper presents a perception-aware decomposition and fusion framework for underwater image enhancement (UIE). Specifically, a general structural patch decomposition and fusion (SPDF) approach is introduced. SPDF is built upon the fusion of two complementary pre-processed inputs in a perception-aware and conceptually independent image space. First, a raw underwater image is pre-processed to produce two complementary versions including a contrast-corrected image and a detail-sharpened image. Then, each of them is decomposed into three conceptually independent components, i.e., mean intensity, contrast, and structure, via structural patch decomposition (SPD). Afterwards, the corresponding components are fused using tailored strategies. The three components after fusion are finally integrated via inverting the decomposition to reconstruct a final enhanced underwater image. The main advantage of SPDF is that two complementary pre-processed images are fused in a perception-aware and conceptually independent image space and the fusions of different components can be performed separately without any interactions and information loss. Comprehensive comparisons on two benchmark datasets demonstrate that SPDF outperforms several state-of-the-art UIE algorithms qualitatively and quantitatively. Moreover, the effectiveness of SPDF is also verified on another two relevant tasks, i.e., low-light image enhancement and single image dehazing. The code will be made available soon. Yaozu Kang, Qiuping Jiang, Chongyi Li, Wenqi Ren, Hantao Liu, Pengjun Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Viewport-Sphere-Branch Network for Blind Quality Assessment of Stitched 360° Omnidirectional ImagesabstractCompared with conventional images/videos, omnidirectional data records rich information with higher resolution and wider Field-of-View. Moreover, the stitching distortions introduced in the panoramic content generation process make the quality assessment task more challenging. Targeting at designing an accurate and fast stitched 360° omnidirectional image quality evaluator, we propose a Viewport-Sphere-Branch Network (VSBNet) via dual-branch quality estimation. Specifically, for the viewport quality estimation, we extract distorted viewports around the stitching seams and conduct distortion rectification through a progressively complementary network to obtain pseudo-reference viewports. The qualitative and quantitative experiments validate that pseudo-reference viewports are reliable. Then, the differences between distorted and pseudo-reference viewports are quantified through transformer architecture to obtain quality scores of viewports. The introduction of pseudo-reference viewports can effectively improve the performance of the viewport quality prediction branch. To establish general scenario awareness and accurately evaluate the immersive experience, we extract feature representation through deformable convolutions to eliminate 2D-to-Sphere intrinsic sampling distortions and use multilayer perceptron to predict score of the whole sphere. The final prediction score is obtained by aggregating the quality scores from viewport and sphere branches. We evaluate the proposed VSBNet on two benchmark databases and results demonstrate that the combination of two branches can obtain more accurate results. Overall, our method is superior to existing full reference and no reference models designed for conventional images and 360° omnidirectional images. Chongzhen Tian, Feng Shao 0001, Xiongli Chai, Qiuping Jiang, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Cross-Modality Double Bidirectional Interaction and Fusion Network for RGB-T Salient Object DetectionabstractRGB-T salient object detection (SOD) aims to detect and segment saliency regions on RGB images and the corresponding thermal maps. The ability of alleviating the modality difference between RGB and thermal modality plays a vital role in the development of RGB-T SOD. However, most of the existing methods try to integrate multi-modal information through various fusion strategies, or reduce the modality difference via unidirectional or undifferentiated bidirectional interaction, but failing in some challenging scenes. To deal with the above question, a novel Cross-Modality Double Bidirectional Interaction and Fusion Network (CMDBIF-Net) for RGB-T SOD is proposed. Specifically, we construct an interactive branch to indirectly bridge the RGB and thermal modalities. In addition, we propose a double bidirectional interaction (DBI) module composed of a forward interaction block (FIB) and a backward interaction block (BIB) to reduce the cross-modality differences. Moreover, a multi-scale feature enhancement and fusion (MSFEF) module is introduced to integrate the multi-modal features with considering the internal gap of different modality. Finally, we use a cascaded decoder and a cross-level feature enhancement (CLFE) module to generate high-quality saliency map. Extensive experiments are conducted on three publicly available RGB-T SOD datasets shows that the proposed CMDBIF-Net achieves outstanding performance against the state-of-the-art (SOTA) RGB-T SOD methods. Zhengxuan Xie, Feng Shao 0001, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Graph Attention Guidance Network With Knowledge Distillation for Semantic Segmentation of Remote Sensing ImagesabstractDeep learning has become a popular method for studying the semantic segmentation of high-resolution remote sensing images (HRRSIs). Existing methods have adopted convolutional neural networks to achieve better segmentation accuracy of HRRSIs, and the success of these models often depends on the model complexity and parameter quantity. However, the deployment of these models on equipment with limited resources is a significant challenge. To solve this problem, a lightweight student network framework—a graph attention guidance network (GAGNet) with knowledge distillation, called GAGNet-S*—is proposed in this study, which distills knowledge from pretrained large teacher network (GAGNet-T) and builds reliable weak labels to optimize untrained student network (GAGNet-S). Inspired by the graph convolution network, this study designs a graph convolution module called the attention-graph decoder, which combines attention mechanisms with graph convolution to optimize image features and improve segmentation accuracy in the semantic segmentation task of HRRSIs. In addition, a dense cross-decoder was designed for multiscale dense fusion, which utilizes rich semantic information in the high-level features to guide and refine the low-level features from the bottom up. Extensive experiments showed that GAGNet-S* (GAGNet-S with knowledge distillation) achieved excellent segmentation performance on two widely used datasets: Potsdam and Vaihingen. The code and models are available at https://github.com/F8AoMn/GAGNet-KD. Wujie Zhou, Xiaomin Fan, Weiqing Yan, Shengdao Shan, Qiuping Jiang, Jenq-Neng Hwang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | GSGNet-S*: Graph Semantic Guidance Network via Knowledge Distillation for Optical Remote Sensing Image Scene AnalysisabstractIn recent years, optical remote sensing image (ORSI) scene analysis has attracted increasing interest. However, existing networks show a trend of bifurcation. Lightweight networks have very high inference speed but poor inference of contextual information in highly complex backgrounds. In contrast, networks with high-performance contextual information reasoning capability require many parameters and are computationally expensive. Since the knowledge distillation method can greatly lighten the model, we propose a graph semantic guided network (GSGNet) that utilizes knowledge refinement for ORSI scenario analysis, which has a high inference speed while maintaining practical contextual inference capability. Rich semantic and detailed information facilitates semantic segmentation of optical remote sensing images. We design adjacent dynamic capture and local-global map inference modules that can effectively extract low-level spatial details and high-level contextual semantics. To improve the attention map relearning performance of the distillation method, we designed semantically guided fusion modules to locate spatial information and refine edge information. We also employed a structural relationship transfer distillation method in which the structural relationship knowledge of the teacher model (GSGNet-T) was used to guide the student model (GSGNet-S). We compared the performances of GSGNet-T and the GSGNet-S with knowledge distillation (GSGNet-S*) with those of several state-of-the-art methods on the Vaihingen and Potsdam datasets. Extensive experiments showed that GSGNet-S* outperformed most advanced methods with only 19.61M parameters and a computation cost of 2.9G FLOPs. The experimental results and code of our network can be accessed at the following URL: https://github.com/LYZ00918/GSGNet-KD. Wujie Zhou, Yangzhen Li, Weiqing Yan, Meixin Fang, Qiuping Jiang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | WaveNet: Wavelet Network With Knowledge Distillation for RGB-T Salient Object DetectionabstractIn recent years, various neural network architectures for computer vision have been devised, such as the visual transformer and multilayer perceptron (MLP). A transformer based on an attention mechanism can outperform a traditional convolutional neural network. Compared with the convolutional neural network and transformer, the MLP introduces less inductive bias and achieves stronger generalization. In addition, a transformer shows an exponential increase in the inference, training, and debugging times. Considering a wave function representation, we propose the WaveNet architecture that adopts a novel vision task-oriented wavelet-based MLP for feature extraction to perform salient object detection in RGB (red-green-blue)-thermal infrared images. In addition, we apply knowledge distillation to a transformer as an advanced teacher network to acquire rich semantic and geometric information and guide WaveNet learning with this information. Following the shortest-path concept, we adopt the Kullback-Leibler distance as a regularization term for the RGB features to be as similar to the thermal infrared features as possible. The discrete wavelet transform allows for the examination of frequency-domain features in a local time domain and time-domain features in a local frequency domain. We apply this representation ability to perform cross-modality feature fusion. Specifically, we introduce a progressively cascaded sine-cosine module for cross-layer feature fusion and use low-level features to obtain clear boundaries of salient objects through the MLP. Results from extensive experiments indicate that the proposed WaveNet achieves impressive performance on benchmark RGB-thermal infrared datasets. The results and code are publicly available at https://github.com/nowander/WaveNet. Wujie Zhou, Qiuping Jiang, Runmin Cong, Jenq-Neng Hwang |
IEEE Trans. Image Process. | 3 |
| 2023 | Toward Multicenter Skin Lesion Classification Using Deep Neural Network With Adaptively Weighted Balance LossabstractRecently, deep neural network-based methods have shown promising advantages in accurately recognizing skin lesions from dermoscopic images. However, most existing works focus more on improving the network framework for better feature representation but ignore the data imbalance issue, limiting their flexibility and accuracy across multiple scenarios in multi-center clinics. Generally, different clinical centers have different data distributions, which presents challenging requirements for the network's flexibility and accuracy. In this paper, we divert the attention from framework improvement to the data imbalance issue and propose a new solution for multi-center skin lesion classification by introducing a novel adaptively weighted balance (AWB) loss to the conventional classification network. Benefiting from AWB, the proposed solution has the following advantages: 1) it is easy to satisfy different practical requirements by only changing the backbone; 2) it is user-friendly with no tuning on hyperparameters; and 3) it adaptively enables small intraclass compactness and pays more attention to the minority class. Extensive experiments demonstrate that, compared with solutions equipped with state-of-the-art loss functions, the proposed solution is more flexible and more competent for tackling the multi-center imbalanced skin lesion classification task with considerable performance on two benchmark datasets. In addition, the proposed solution is proved to be effective in handling the imbalanced gastrointestinal disease classification task and the imbalanced DR grading task. Code is available at https://github.com/Weipeishan2021. Guanghui Yue 0001, Peishan Wei, Tianwei Zhou, Qiuping Jiang, Weiqing Yan, Tianfu Wang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Perceptual Quality Assessment of Cartoon ImagesabstractIn the animation industry, automatically predicting the quality of cartoon images based on the inputs of general distortions and color change is an urgent task, while the existing no-reference (NR) methods usually measure the perceptual quality of the natural images. In this paper, based on the observation that structure and color are the main factors affecting cartoon images quality, we proposed a new NR quality prediction metric for cartoon images, which fully takes gradient and color information into account. The experimental results on our newly constructed NBU-CIQAD dataset with color change and other existing cartoon image dataset demonstrate that the proposed method significantly outperforms existing no-references methods for the task of cartoon image quality assessment. The database and code will be released athttps://github.com/1010075746/NBU-CIQAD. Hangwei Chen, Xiongli Chai, Feng Shao 0001, Xuejin Wang, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2023 | Deep Blind Image Quality Assessment Powered by Online Hard Example MiningabstractRecently, blind image quality assessment (BIQA) models based on deep neural networks (DNNs) have achieved impressive performance on existing datasets. However, due to the intrinsic imbalance property of the training set, not all distortions or images are handled equally well. Online hard example mining (OHEM) is a promising way to alleviate this issue. Inspired by the recent finding that network pruning disproportionately hampers the model's memorization of a tractable subset, atypical, low-quality, long-tailed samples, that are hard-to-memorize during training and easily “forgotten” during pruning, we propose an effective “plug-and-play” OHEM pipeline, especially for generalizable deep BIQA. Specifically, we train two parallel weight-sharing branches simultaneously, where one is full model and other is a “self-competitor” generated from the full model online by network pruning. Then, we leverage the prediction disagreement between the full model and its pruned variant (i.e., the self-competitor) to expose easily “forgettable” samples, which are therefore regarded as the hard ones. We then enforce the prediction consistency between the full model and its pruned variant to implicitly put more focus on these hard samples, which benefits the full model to recover forgettable information introduced by pruning. Extensive experiments across multiple datasets and BIQA models demonstrate that the proposed OHEM can improve the model performance and generalizability as measured by correlation numbers and group maximum differentiation (gMAD) competition. Our code are available at:https://github.com/wangzhihua520/IQA_with_OHEM Zhihua Wang 0002, Qiuping Jiang, Shanshan Zhao 0001, Wensen Feng, Weisi Lin |
IEEE Trans. Multim. | 2 |
| 2023 | Self-Supervised Attentive Generative Adversarial Networks for Video Anomaly DetectionabstractVideo anomaly detection (VAD) refers to the discrimination of unexpected events in videos. The deep generative model (DGM)-based method learns the regular patterns on normal videos and expects the learned model to yield larger generative errors for abnormal frames. However, DGM cannot always do so, since it usually captures the shared patterns between normal and abnormal events, which results in similar generative errors for them. In this article, we propose a novel self-supervised framework for unsupervised VAD to tackle the above-mentioned problem. To this end, we design a novel self-supervised attentive generative adversarial network (SSAGAN), which is composed of the self-attentive predictor, the vanilla discriminator, and the self-supervised discriminator. On the one hand, the self-attentive predictor can capture the long-term dependences for improving the prediction qualities of normal frames. On the other hand, the predicted frames are fed to the vanilla discriminator and self-supervised discriminator for performing true-false discrimination and self-supervised rotation detection, respectively. Essentially, the role of the self-supervised task is to enable the predictor to encode semantic information into the predicted normal frames via adversarial training, in order for the angles of rotated normal frames can be detected. As a result, our self-supervised framework lessens the generalization ability of the model to abnormal frames, resulting in larger detection errors for abnormal frames. Extensive experimental results indicate that SSAGAN outperforms other state-of-the-art methods, which demonstrates the validity and advancement of SSAGAN. Chao Huang 0008, Jie Wen 0001, Yong Xu 0001, Qiuping Jiang, Jian Yang 0003, Yaowei Wang 0001, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Image Quality Assessment-driven Reinforcement Learning for Mixed Distorted Image RestorationabstractDue to the diversity of the degradation process that is difficult to model, the recovery of mixed distorted images is still a challenging problem. The deep learning model trained under certain degradation declines significantly in other degradation situations. In this article, we explore ways to use a combination of tools to deal with the mixed distortion. First, we illustrate the limitations of a single deep network in dealing with multiple distortion types and then introduce a hierarchical toolkit with distinguished powerful tools. Second, we investigate how an efficient representation of images combined with a reinforcement learning (RL) paradigm helps to deal with tool noise in continuous restoration. The proposed method can accurately capture the distortion preferences for selecting the optimal recovery tools by RL agent. Finally, to fully utilize random tools for unknown distortion combinations, we adopt the exploration scheme with various quality evaluation methods to achieve more quality improvements. Experimental results demonstrate that the peak signal-to-noise ratio of the proposed method is 3.30 dB higher than other state-of-the-art RL-based methods on the CSIQ single distortion dataset and 0.95 dB higher on the DIV2K mixed distortion dataset. Xiaoyu Zhang 0002, Wei Gao 0003, Ge Li 0002, Qiuping Jiang, Runmin Cong |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Improving IQA Performance Based on Deep Mutual LearningabstractIn this paper, we propose a novel solution, termed DML-IQA, for the image quality assessment (IQA) tasks. DML-IQA holds a dual-branch network architecture and builds the IQA model through a deep mutual learning (DML) strategy. Specifically, the two branches extract stable feature representations by feeding different transformed images into the classical CNNs. The DML strategy first calculates the prediction loss of each branch and the consistency loss across two branches, followed by updating the network iteratively to converge. Overall, DML-IQA has the following advantages: 1) It is flexible to adapt to diverse backbones for tackling the IQA issues in both the laboratory and wild; 2) It improves the baseline’s performance by approximately 1%~2%, especially performs well in the case of small samples. Extensive experiments on four public datasets show that the proposed DML-IQA can handle the IQA tasks with considerable effectiveness and generalization. Guanghui Yue 0001, Honglv Wu, Qiuping Jiang, Tianfu Wang 0001 |
ICIP | 4 |
| 2022 | Pixel-Level Anomaly Detection via Uncertainty-aware Prototypical TransformerabstractPixel-level visual anomaly detection, which aims to recognize the abnormal areas from images, plays an important role in industrial fault detection and medical diagnosis. However, it is a challenging task due to the following reasons: i) the large variation of anomalies; and ii) the ambiguous boundary between anomalies and their normal surroundings. In this work, we present an uncertainty-aware prototypical transformer (UPformer), which takes into account both the diversity and uncertainty of anomaly to achieve accurate pixel-level visual anomaly detection. To this end, we first design a memory-guided prototype learning transformer encoder to learn and memorize the prototypical representations of anomalies for enabling the model to capture the diversity of anomalies. Additionally, an anomaly detection uncertainty quantizer is designed to learn the distributions of anomaly detection for measuring the anomaly detection uncertainty. Furthermore, an uncertainty-aware transformer decoder is proposed to leverage the detection uncertainties to guide the model to focus on the uncertain areas and generate the final detection results. As a result, our method achieves more accurate anomaly detection by combining the benefits of prototype learning and uncertainty estimation. Experimental results on five datasets indicate that our method achieves state-of-the-art anomaly detection performance. Chao Huang 0008, Chengliang Liu 0003, Zheng Zhang 0006, Zhihao Wu 0002, Jie Wen 0001, Qiuping Jiang, Yong Xu 0001 |
ACM Multimedia | 6 |
| 2022 | Deep network based stereoscopic image quality assessment via binocular summing and differencing
Jinbin Hu 0002, Xuejin Wang, Xiongli Chai, Feng Shao 0001, Qiuping Jiang |
J. Vis. Commun. Image Represent. | 5 |
| 2022 | A brief survey on adaptive video streaming quality assessment
Wei Zhou 0021, Xiongkuo Min, Qiuping Jiang |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | Monocular and Binocular Interactions Oriented Deformable Convolutional Networks for Blind Quality Assessment of Stereoscopic Omnidirectional ImagesabstractStereoscopic omnidirectional content, as a novel visual media, has drawn wide attention in recent years due to its ability in providing strong immersive experience. Since Stereoscopic Omnidirectional Images (SOIs) involve the properties from panoramic and stereoscopic visual perception, it is very challenging to establish an efficient and effective visual quality evaluation model for SOIs. To better measure the user’s experience in virtual reality, we put forward a novel deep learning framework to assess the quality of SOIs in this paper. Firstly, the deformable convolutions instead of standard convolutions are adopted to ensure the invariant receptive fields of convolutional kernels on Equi-Rectangular Projection (ERP). Secondly, according to the stereoscopic property, we use binocular-difference information and a coarse-to-fine mechanism to construct the binocular feature extraction network. Thirdly, a three-channel network involving left-view, right-view and binocular-difference channels is presented to simulate the process of monocular and binocular interactions, in which independent quality labels are provided for each channel to reflect the individual effect of monocular and binocular visions on the whole visual quality. Finally, experimental results on two available benchmark databases demonstrate the superiority of the proposed metric over the state-of-the-art blind quality assessment models in predicting the quality of SOIs. Moreover, our model is efficient in computational cost as the feature extraction is directly applied on ERP images. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | CGMDRNet: Cross-Guided Modality Difference Reduction Network for RGB-T Salient Object DetectionabstractHow to explore the interaction between the RGB and thermal modalities is the key success of the RGB-T saliency object detection (SOD). Most of the existing methods integrate multi-modality information by designing various fusion strategies. However, the modality gap between the RGB and thermal features will lead to unsatisfactory performances by simple feature concatenation. To solve this problem, we innovatively propose a cross-guided modality difference reduction network (CGMDRNet) to achieve intrinsic consistency feature fusion via reducing the modality differences. Specifically, we design a modality difference reduction (MDR) module, which is embedded in each layer of the backbone network. The module uses a cross-guided strategy to reduce the modality difference between the RGB and thermal features. Then, a cross-attention fusion (CAF) module is designed to fuse cross-modality features with small modality differences. In addition, we use a transformer-based feature enhancement (TFE) module to enhance the high-level feature representation that contributes more to performance. Finally, the high-level features guide the fusion of low-level features to obtain a saliency map with clear boundaries. Extensive experiments on three public RGB-T datasets show that the proposed CGMDRNet achieves competitive performance compared with state-of-the-art (SOTA) RGB-T SOD models. Feng Shao 0001, Xiongli Chai, Hangwei Chen, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Underwater Image Enhancement Quality Evaluation: Benchmark Dataset and Objective MetricabstractDue to the attenuation and scattering of light by water, there are many quality defects in raw underwater images such as color casts, decreased visibility, reduced contrast,et al.. Many different underwater image enhancement (UIE) algorithms have been proposed to enhance underwater image quality. However, how to fairly compare the performance among UIE algorithms remains a challenging problem. So far, the lack of comprehensive human subjective user study with large-scale benchmark dataset and reliable objective image quality assessment (IQA) metric makes it difficult to fully understand the true performance of UIE algorithms. We in this paper make efforts in both subjective and objective aspects to fill these gaps. Firstly, we construct a new Subjectively-Annotated UIE benchmark Dataset (SAUD) which simultaneously provides real-world raw underwater images, readily available enhanced results by representative UIE algorithms, and subjective ranking scores of each enhanced result. Secondly, we propose an effective No-reference (NR) Underwater Image Quality metric (NUIQ) to automatically evaluate the visual quality of enhanced underwater images. Experiments on the constructed SAUD dataset demonstrate the superiority of our proposed NUIQ metric, achieving higher consistency with subjective rankings than 22 mainstream NR-IQA metrics. The dataset and source code will be made available athttps://github.com/yia-yuese/SAUD-Dataset. Qiuping Jiang, Yuese Gu, Chongyi Li, Runmin Cong, Feng Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | LGGD+: Image Retargeting Quality Assessment by Measuring Local and Global Geometric DistortionsabstractNumerous image retargeting algorithms have been proposed to achieve adaptive image resizing during the past years. To compare different image retargeting algorithms, reliable objective image retargeting quality assessment (IRQA) metrics are highly desired. Given that image retargeting usually introduces geometric distortions, this paper presents an objective IRQA metric by measuring both local and global geometric distortions (LGGD). Since human visual system perception is highly dependent on edges and the geometric distortions caused by image retargeting usually cause edge deformation, a sketch token-based local edge descriptor (ST-LED) is introduced to represent geometric-aware features in LGGD. First, ST-LED is first applied on both source and retargeted images for edge pattern representation. Second, pixel-level backward registration is conducted to enable estimating local geometric distortion (LGD) and a spatial pyramid-improved Bag-of-Token (BoT) model is built to enable estimating global geometric distortion (GGD). Since the proposed LGGD metric only focuses on geometric distortion while image retargeting quality is related with more aspects, we further fuse LGGD and an existing (EXT) IRQA metric to build a final version called LGGD+ for IRQA. Experiments on two benchmark databases demonstrate the superiority of LGGD+ and the excellent compatibility of our proposed LGGD for further improving a wide range of existing IRQA metrics (including both geometric distortion and non-geometric distortion metrics). In addition, the effectiveness of our LGGD metric is also demonstrated in another relevant task, i.e., quality evaluation of depth-image-based rendering (DIBR)-synthesized images, which also calls for accurate estimation of geometric distortion. Zhenyu Peng, Qiuping Jiang, Feng Shao 0001, Wei Gao 0003, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | VSOIQE: A Novel Viewport-Based Stitched 360° Omnidirectional Image Quality EvaluatorabstractWith the rapid development of virtual reality (VR), 360° omnidirectional images and videos have drawn wide attention. However, the quality assessment of 360° omnidirectional images is a challenging task, especially when the panoramic image contains multiple stitching distortions. We propose a viewport-based stitched 360° omnidirectional image quality evaluator (VSOIQE), by first extracting the features of salient and stitching viewports, and then inferring the overall perceptual quality via multiple linear regression (MLR). Comprehensive image attributes including edge, color, shape and information entropy are considered in the framework. Experimental results on two benchmark databases demonstrate the superiority of the proposed metric over both the state-of-the-art quality models designed for 2D images and the quality models developed for 360° omnidirectional images. Chongzhen Tian, Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | No-Reference Quality Assessment for 360-Degree Images by Analysis of Multifrequency Information and Local-Global Naturalnessabstract360-degree/omnidirectional images (OIs) have received remarkable attention due to the increasing applications of virtual reality (VR). Compared to conventional 2D images, OIs can provide more immersive experiences to consumers, benefiting from the higher resolution and plentiful field of views (FoVs). Moreover, observing OIs is usually in a head-mounted display (HMD) without references. Therefore, an efficient blind quality assessment method, which is specifically designed for 360-degree images, is urgently desired. In this paper, motivated by the characteristics of the human visual system (HVS) and the viewing process of VR visual content, we propose a novel and effective no-reference omnidirectional image quality assessment (NR OIQA) algorithm by MultiFrequency Information and Local-Global Naturalness (MFILGN). Specifically, inspired by the frequency-dependent property of the visual cortex, we first decompose the projected equirectangular projection (ERP) maps into wavelet subbands by using discrete Haar wavelet transform (DHWT). Then, the entropy intensities of low-frequency and high-frequency subbands are exploited to measure the multifrequency information of OIs. In addition to considering the global naturalness of ERP maps, owing to the browsed FoVs, we extract the natural scene statistics (NSS) features from each viewport image as the measure of local naturalness. With the proposed multifrequency information measurement and local-global naturalness measurement, we utilize support vector regression (SVR) as the final image quality regressor to train the quality evaluation model from visual quality-related features to human ratings. To our knowledge, the proposed model is the first no-reference quality assessment method for 360-degree images that combines multifrequency information and image naturalness. Experimental results on two publicly available OIQA databases demonstrate that our proposed MFILGN outperforms state-of-the-art full-reference (FR) and NR approaches. Wei Zhou 0021, Jiahua Xu 0001, Qiuping Jiang, Zhibo Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Self-Supervision-Augmented Deep Autoencoder for Unsupervised Visual Anomaly DetectionabstractDeep autoencoder (AE) has demonstrated promising performances in visual anomaly detection (VAD). Learning normal patterns on normal data, deep AE is expected to yield larger reconstruction errors for anomalous samples, which is utilized as the criterion for detecting anomalies. However, this hypothesis cannot be always tenable since the deep AE usually captures the low-level shared features between normal and abnormal data, which leads to similar reconstruction errors for them. To tackle this problem, we propose a self-supervised representation-augmented deep AE for unsupervised VAD, which can enlarge the gap of anomaly scores between normal and abnormal samples by introducing autoencoding transformation (AT). Essentially, AT is introduced to facilitate AE to learn the high-level visual semantic features of normal images by introducing a self-supervision task (transformation reconstruction). In particular, our model inputs the original and transformed images into the encoder for obtaining latent representations; afterward, they are fed to the decoder for reconstructing both the original image and applied transformation. In this way, our model can utilize both image and transformation reconstruction errors to detect anomaly. Extensive experiments indicate that the proposed method outperforms other state-of-the-art methods, which demonstrates the validity and advancement of our model. Chao Huang 0008, Zehua Yang, Jie Wen 0001, Yong Xu 0001, Qiuping Jiang, Jian Yang 0003, Yaowei Wang 0001 |
IEEE Trans. Cybern. | 5 |
| 2022 | Consistent Quality Oriented Rate Control in HEVC Via Balancing Intra and Inter Frame CodingabstractConsistent quality oriented rate control (RC) in video coding has attracted much more attention. However, the existing efforts only focus on decreasing variations between every two adjacent frames, but neglect coding tradeoff problem between intraframes and interframes. In this article, we deal with it from a new perspective, where intraframe quantization parameter (IQP) and RC are optimized for balanced coding. First, due to the importance of intraframes, a new framework is proposed for consistent quality oriented IQP prediction, and then we remove unqualified IQP candidates using the proposed penalty term. Second, we extensively evaluate possible features, and select target bits per pixel for all remaining frames, average and standard variance of frame QPs, where equivalent acquisition methods for QP features are given. Third, predicted IQPs are clipped effectively according to bandwidth and previous information for better bit rate accuracy. Compared with high efficiency video coding reference baseline, experiments demonstrate that our method reduces quality fluctuation greatly by 37.2% on frame-level standard variance of peak-signal-noise-ratio and 45.1% on that of structural similarity. Moreover, it also can have satisfactory results on rate-distortion performance, bit accuracy and buffer control. Wei Gao 0003, Qiuping Jiang, Ronggang Wang, Siwei Ma 0001, Ge Li 0002, Sam Kwong |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Abnormal Event Detection Using Deep Contrastive Learning for Intelligent Video Surveillance SystemabstractThe continuous developments of urban and industrial environments have increased the demand for intelligent video surveillance. Deep learning has achieved remarkable performance for anomaly detection in surveillance videos. Previous approaches achieve anomaly detection with a single-pretext task (image reconstruction or prediction) and detect anomalies by larger reconstruction error or poor prediction. However, they cannot fully exploit the discriminative semantics and temporal context information. Moreover, tackling anomaly detection with a single pretext task is suboptimal due to the nonalignment between the pretext task and anomaly detection. In this article, we propose a temporal-aware contrastive network (TAC-Net) to address the abovementioned problems of anomaly detection for intelligence video surveillance. TAC-Net is an unsupervised method that utilizes deep contrastive self-supervised learning to capture the high-level semantic features and tackles anomaly detection with multiple self-supervised tasks. During inference phase, the multiple task losses and contrastive similarity are utilized to calculate the anomaly score. Experimental results show that our method is superior to state-of-the-art approaches on three benchmarks, which demonstrates the validity and advancement of TAC-Net. Chao Huang 0008, Zhihao Wu 0002, Jie Wen 0001, Yong Xu 0001, Qiuping Jiang, Yaowei Wang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | Single Image Super-Resolution Quality Assessment: A Real-World Dataset, Subjective Studies, and an Objective MetricabstractNumerous single image super-resolution (SISR) algorithms have been proposed during the past years to reconstruct a high-resolution (HR) image from its low-resolution (LR) observation. However, how to fairly compare the performance of different SISR algorithms/results remains a challenging problem. So far, the lack of comprehensive human subjective study on large-scale real-world SISR datasets and accurate objective SISR quality assessment metrics makes it unreliable to truly understand the performance of different SISR algorithms. We in this paper make efforts to tackle these two issues. Firstly, we construct a real-world SISR quality dataset (i.e., RealSRQ) and conduct human subjective studies to compare the performance of the representative SISR algorithms. Secondly, we propose a new objective metric, i.e., KLTSRQA, based on the Karhunen-Loéve Transform (KLT) to evaluate the quality of SISR images in a no-reference (NR) manner. Experiments on our constructed RealSRQ and the latest synthetic SISR quality dataset (i.e., QADS) have demonstrated the superiority of our proposed KLTSRQA metric, achieving higher consistency with human subjective scores than relevant existing NR image quality assessment (NR-IQA) metrics. The dataset and the code will be made available at https://github.com/Zhentao-Liu/RealSRQ-KLTSRQA. Qiuping Jiang, Ke Gu 0001, Feng Shao 0001, Xinfeng Zhang 0001, Hantao Liu, Weisi Lin |
IEEE Trans. Image Process. | 1 |
| 2022 | Toward Top-Down Just Noticeable Difference Estimation of Natural ImagesabstractJust noticeable difference (JND) of natural images refers to the maximum pixel intensity change magnitude that typical human visual system (HVS) cannot perceive. Existing efforts on JND estimation mainly dedicate to modeling the diverse masking effects in either/both spatial or/and frequency domains, and then fusing them into an overall JND estimate. In this work, we turn to a dramatically different way to address this problem with a top-down design philosophy. Instead of explicitly formulating and fusing different masking effects in a bottom-up way, the proposed JND estimation model dedicates to first predicting a critical perceptual lossless (CPL) counterpart of the original image and then calculating the difference map between the original image and the predicted CPL image as the JND map. We conduct subjective experiments to determine the critical points of 500 images and find that the distribution of cumulative normalized KLT coefficient energy values over all 500 images at these critical points can be well characterized by a Weibull distribution. Given a testing image, its corresponding critical point is determined by a simple weighted average scheme where the weights are determined by a fitted Weibull distribution function. The performance of the proposed JND model is evaluated explicitly with direct JND prediction and implicitly with two applications including JND-guided noise injection and JND-guided image compression. Experimental results have demonstrated that our proposed JND model can achieve better performance than several latest JND models. In addition, we also compare the proposed JND model with existing visual difference predicator (VDP) metrics in terms of the capability in distortion detection and discrimination. The results indicate that our JND model also has a good performance in this task. The code of this work are available at https://github.com/Zhentao-Liu/KLT-JND. Qiuping Jiang, Shiqi Wang 0001, Feng Shao 0001, Weisi Lin |
IEEE Trans. Image Process. | 1 |
| 2022 | Unsupervised Decomposition and Correction Network for Low-Light Image EnhancementabstractVision-based intelligent driving assistance systems and transportation systems can be improved by enhancing the visibility of the scenes captured in extremely challenging conditions. In particular, many low-image image enhancement (LIE) algorithms have been proposed to facilitate such applications in low-light conditions. While deep learning-based methods have achieved substantial success in this field, most of them require paired training data, which is difficult to be collected. This paper advocates a novel Unsupervised Decomposition and Correction Network (UDCN) for LIE without depending on paired data for training. Inspired by the Retinex model, our method first decomposes images into illumination and reflectance components with an image decomposition network (IDN). Then, the decomposed illumination is processed by an illumination correction network (ICN) and fused with the reflectance to generate a primary enhanced result. In contrast with fully supervised learning approaches, UDCN is an unsupervised one which is trained only with low-light images and corresponding histogram equalized (HE) counterparts (can be derived from the low-light image itself) as input. Both the decomposition and correction networks are optimized under the guidance of hybrid no-reference quality-aware losses and inter-consistency constraints between the low-light image and its HE counterpart. In addition, we also utilize an unsupervised noise removal network (NRN) to remove the noise previously hidden in the darkness for further improving the primary result. Qualitative and quantitative comparison results are reported to demonstrate the efficacy of UDCN and its superiority over several representative alternatives in the literature. The results and code will be made public available athttps://github.com/myd945/UDCN. Qiuping Jiang, Yudong Mao, Runmin Cong, Wenqi Ren, Chao Huang 0008, Feng Shao 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Cross-Modality Fusion and Progressive Integration Network for Saliency Prediction on Stereoscopic 3D ImagesabstractTraditional 2D image-based saliency prediction models suffer from unsatisfactory performance when dealing with stereoscopic 3D (S3D) images because eye movements in the case of freely viewing S3D images are demonstrated to be guided by both RGB and depth features. This paper studies the problem of saliency prediction on S3D images, where the interactions between RGB and depth modalities are both taken into account. Specifically, we design a novel deep neural network named Cross-modality Fusion and Progressive Integration Network (CFPI-Net) to address this problem. It consists of a Multi-level Cross-modality Feature Fusion (MCFF) module and a Multi-stage Progressive Feature Integration (MPFI) module. The MCFF module first captures hierarchical contexture features from each modality and then effectively fuses the hierarchical contexture features from different modalities at each level. The MPFI module involves multiple cascaded deeply supervised feature integration (DSFI) blocks in which the low-level and high-level cross-modality features are progressively integrated using the integrated features in the previous stage as a guidance. Our proposed CFPI-Net benefits from the advantages of multi-level feature representation, cross-modality feature fusion, and multi-stage progressive feature integration, which hereby fully boost the performance. Experimental results on two benchmark datasets demonstrate that CFPI-Net outperforms state-of-the-art saliency prediction methods both quantitatively and qualitatively. All the results and relevant codes will be made available to the public. Yudong Mao, Qiuping Jiang, Runmin Cong, Wei Gao 0003, Feng Shao 0001, Sam Kwong |
IEEE Trans. Multim. | 2 |
| 2022 | List-Wise Rank Learning for Stereoscopic Image Retargeting Quality AssessmentabstractStereoscopic imageretargeting (SIR) techniques attempt to display stereoscopic images on stereoscopic devices of various resolutions and aspect ratios to provide the users with better viewing experience. However, new quality perceptual problems emerge in the retargeted stereoscopic images generated by current SIR operators are quite different from those in the retargeted 2D images. In this paper, we dedicate to exploring the perceptual quality-related factors (e.g., shape preservation, object preservation and visual comfort.) of retargeted stereoscopic images, and propose a novel quality evaluation metric for SIR to achieve a more consistent evaluation with 3D perception and image degradation mechanism in the SIR process. Moreover, image quality features and 3D perceptual features are integrated into one representation for an overall perceptual quality prediction using a list-wise ranking approach, which gives priority to the ranking among the SIR results generated from the same stereoscopic source. Experimental results demonstrate that the proposed method outperforms most quality models developed for retargeted 2D/stereoscopic images. Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Xiongli Chai, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 3 |
| 2022 | Combining Retargeting Quality and Depth Perception Measures for Quality Evaluation of Retargeted StereopairsabstractStereoscopic Image Retargeting (SIR) aims to adapt stereoscopic images and videos to 3D display devices with various aspect ratios by emphasizing the important content while retaining surrounding context with minimal visual distortion. To address the issue of SIR evaluation, this paper presents a new objective quality assessment method for retargeted stereopairs by combining image quality and depth perception measures. Specifically, the image quality measure is conducted between the source and retargeted intermediate views generated by the view synthesis method to characterize the geometric distortion and content loss of the retargeted stereopair, while several depth-aware features are extracted to measure the visual comfort/discomfort and depth sensation when human views a 3D scene. Then, the extracted features are integrated into an overall perceptual quality prediction. Experiment results on NBU SIRQA and SIRD databases verify the superiority of our method. Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Zhenqi Fu, Xiangchao Meng, Ke Gu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 3 |
| 2021 | Corrections to "Blind quality assessment for image superresolution using deep two-stream convolutional networks"
Wei Zhou 0021, Qiuping Jiang, Yuwang Wang, Zhibo Chen 0001, Weiping Li 0003 |
Inf. Sci. | 2 |
| 2021 | Two-Branch Deep Neural Network for Underwater Image Enhancement in HSV Color SpaceabstractDue to the influence of light absorption and scattering, underwater images usually suffer from quality deteriorations such as color cast and reduced contrast. The diverse quality degradations not only dissatisfy the user expectation but also lead to a significant performance drop in many underwater vision applications. This letter proposes a novel two-branch deep neural network for underwater image enhancement (UIE), which is capable of separately removing color cast and enhancing image contrast by fully leveraging useful properties of the HSV color space in disentangling chrominance and intensity. Specifically, the input underwater image is first converted into the HSV color space and disentangled into HS and V channels to serve as the input of the two branches, respectively. Then, the color cast removal branch enhances the H and S channels with a generative adversarial network architecture while the contrast enhancement branch enhances the V channel via a traditional convolutional neural network. The enhanced channels by the two branches are merged and converted back into RGB color space to obtain the final enhanced result. Experimental results demonstrate that, compared with state-of-the-art UIE methods, our method can produce much more visually pleasing enhanced results. Junkang Hu, Qiuping Jiang, Runmin Cong, Wei Gao 0003, Feng Shao 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Roundness-Preserving Warping for Aesthetic Enhancement-Based Stereoscopic Image EditingabstractImage editing is an effective solution to adapt contents for different applications. In this paper, we present a roundness-preserving warping model for stereoscopic image editing, in which energy constraints from image quality energy, aesthetics energy and depth adaptation energy are involved in the framework to solve the optimization. Specifically, to preserve object roundness during warping, the relationship between object's shape and disparity is established and is applied for depth adaptation. Different from the existing stereoscopic image editing methods, the main innovations of our method are to achieve a tradeoff in balancing information loss and reducing semantic distortion while providing a novel death adaptation model for recomposition and retargeting applications. Experimental results demonstrate the effectiveness of our method in enhancing the aesthetics of stereoscopic images. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | No-Reference Image Contrast Evaluation by Generating Bidirectional PseudoreferencesabstractThis article proposes a simple yet reliable no-reference image contrast evaluator (NICE) by generating bidirectional pseudoreferences (BPR). Different from the existing no-reference metrics that only operate on the contrast distorted image (CDI) itself, our proposed NICE-BPR measures the deviations of a CDI to its corresponding aggravated and enhanced counterparts (i.e., BPRs) in a hybrid feature space. Given a CDI, we first perform contrast aggravation and contrast enhancement using gamma correction and histogram equalization, respectively. Then, hybrid contrast-aware features are, respectively, extracted from the CDI and its corresponding BPRs via the analysis of histogram, entropy, and structure. The features obtained from the CDI are one-by-one compared with those from the BPRs to derive the bidirectional feature deviation vector. Finally, a quality predictor is built by learning a regression model to fuse the feature vector into a continuous quality score. Extensive experiments on several databases well-demonstrate the superiority of NICE-BPR. Qiuping Jiang, Zhenyu Peng, Guanghui Yue 0001, Feng Shao 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Online Learning-Based Multi-Stage Complexity Control for Live Video CodingabstractHigh Efficiency Video Coding (HEVC) can significantly improve the compression efficiency in comparison with the preceding H.264/Advanced Video Coding (AVC) but at the cost of extremely high computational complexity. Hence, it is challenging to realize live video applications on low-delay and power-constrained devices, such as the smart mobile devices. In this article, we propose an online learning-based multi-stage complexity control method for live video coding. The proposed method consists of three stages: multi-accuracy Coding Unit (CU) decision, multi-stage complexity allocation, and Coding Tree Unit (CTU) level complexity control. Consequently, the encoding complexity can be accurately controlled to correspond with the computing capability of the video-capable device by replacing the traditional brute-force search with the proposed algorithm, which properly determines the optimal CU size. Specifically, the multi-accuracy CU decision model is obtained by an online learning approach to accommodate the different characteristics of input videos. In addition, multi-stage complexity allocation is implemented to reasonably allocate the complexity budgets to each coding level. In order to achieve a good trade-off between complexity control and rate distortion (RD) performance, the CTU-level complexity control is proposed to select the optimal accuracy of the CU decision model. The experimental results show that the proposed algorithm can accurately control the coding complexity from 100% to 40%. Furthermore, the proposed algorithm outperforms the state-of-the-art algorithms in terms of both accuracy of complexity control and RD performance. Chao Huang 0008, Zongju Peng, Yong Xu 0001, Qiuping Jiang, Yun Zhang 0002, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Image Process. | 5 |
| 2021 | Progressive Self-Guided Loss for Salient Object DetectionabstractWe present a simple yet effective progressive self-guided loss function to facilitate deep learning-based salient object detection (SOD) in images. The saliency maps produced by the most relevant works still suffer from incomplete predictions due to the internal complexity of salient objects. Our proposed progressive self-guided loss simulates a morphological closing operation on the model predictions for progressively creating auxiliary training supervisions to step-wisely guide the training process. We demonstrate that this new loss function can guide the SOD model to highlight more complete salient objects step-by-step and meanwhile help to uncover the spatial dependencies of the salient object pixels in a region growing manner. Moreover, a new feature aggregation module is proposed to capture multi-scale features and aggregate them adaptively by a branch-wise attention mechanism. Benefiting from this module, our SOD framework takes advantage of adaptively aggregated multi-scale features to locate and detect salient objects effectively. Experimental results on several benchmark datasets show that our loss function not only advances the performance of existing SOD models without architecture modification but also helps our proposed framework to achieve state-of-the-art performance. Sheng Yang 0006, Weisi Lin, Guosheng Lin, Qiuping Jiang, Zichuan Liu |
IEEE Trans. Image Process. | 4 |
| 2021 | Subjective and Objective Quality Assessment for Stereoscopic Image RetargetingabstractBinocular stereoscopic image retargeting (SIR) aims to adjust 3D images into target aspect ratios. In recent years, various SIR methods have been proposed, but there are few researches on visual quality assessment. As a consequence, we construct a benchmark stereoscopic image retargeting quality assessment database (NBU-SIRQA), which contains 720 stereoscopic retargeted images generated by eight representative SIR operators. Subjective test is conducted to obtain the mean opinion score (MOS) for each stereoscopic retargeted image. Additionally, we propose an objective SIRQA metric based on grid deformation and information loss (GDIL). The main idea of GDIL is to decompose the SIR operator into two transformations: monocular image retargeting transformation and viewpoint transformation. In each transformation, grid deformation and information loss are extracted simultaneously to represent image quality and 3D perception quality. Experimental results validated on our established NBU-SIRQA database show the superiority of our metric in measuring the quality of stereoscopic retargeted images over the existing approaches. Zhenqi Fu, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 3 |
| 2021 | Exploiting Local Degradation Characteristics and Global Statistical Properties for Blind Quality Assessment of Tone-Mapped HDR ImagesabstractTone mapping operators (TMOs) are developed to convert a high dynamic range (HDR) image into a low dynamic range (LDR) one for display with the goal of preserving as much visual information as possible. However, image quality degradation is inevitable due to the dynamic range compression during the tone-mapping process. This accordingly raises an urgent demand for effective quality evaluation methods to select a high-quality tone-mapped image (TMI) from a set of candidates generated by distinct TMOs or the same TMO with different parameter settings. A key element to the success of TMI quality evaluation is to extract effective features that are highly consistent with human perception. Towards this end, this paper proposes a novel blind TMI quality metric by exploiting both local degradation characteristics and global statistical properties for feature extraction. Several image attributes including texture, structure, colorfulness and naturalness are considered either locally or globally. The extracted local and global features are aggregated into an overall quality via regression. Experimental results on two benchmark databases demonstrate the superiority of the proposed metric over both the state-of-the-art blind quality models designed for synthetically distorted images (SDIs) and the blind quality models specifically developed for TMIs. Xuejin Wang, Qiuping Jiang, Feng Shao 0001, Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Measuring Coarse-to-Fine Texture and Geometric Distortions for Quality Assessment of DIBR-Synthesized ImagesabstractA synthesized view can be generated via Depth-Image-Based Rendering (DIBR) technique using one (or more) color images and the associated depth maps. However, several artifacts may occur in the synthesized views due to the imperfect color images, depth maps or texture inpainting techniques, which cannot be effectively estimated by the conventional quality metrics designed for natural images. In this paper, a new quality metric is proposed to evaluate DIBR-synthesized images by measuring texture and geometric distortions. The artifacts are first analyzed on different phases of the synthesis process, and the associated features are extracted to estimate the degree of texture and geometric distortions from both coarse and fine scales. Finally, individual quality scores are aggregated into an overall quality via regression. Experimental results on three publicly available DIBR datasets demonstrate the superiority of the proposed method over the state-of-the-art quality models. Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 3 |
| 2021 | Transformation-Aware Similarity Measurement for Image Retargeting Quality Assessment via Bidirectional RewarpingabstractImage retargeting is an effective way to adapt images for target displays with different aspect ratios and sizes. Meanwhile, effective image retargeting quality assessment (IRQA) is important for optimizing the image retargeting operations. In this paper, we propose a transform-aware similarity (TRASIM) measurement metric for IRQA, including bidirectional geometric distortion measurement, bidirectional information loss measurement, and global salient structure distortion measurement. The main innovation of the TRASIM is to build a universal framework to establish the similarity transformation via bidirectional rewarping to simulate different types of retargeting operators. Based on the similarity transformation, geometric distortion and content loss are measured to determine the retargeting quality. Experimental results on two widely used databases (CUHK and RetargetMe) indicate that the proposed TRASIM has higher consistency with subjective ranks, compared with the state-of-the-art IRQA metrics. Feng Shao 0001, Zhenqi Fu, Qiuping Jiang, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | MMNet: Multi-Stage and Multi-Scale Fusion Network for RGB-D Salient Object DetectionabstractMost existing RGB-D salient object detection (SOD) methods directly extract and fuse raw features from RGB and depth backbones. Such methods can be easily restricted by low-quality depth maps and redundant cross-modal features. To effectively capture multi-scale cross-modal fusion features, this paper proposes a novel Multi-stage and Multi-Scale Fusion Network (MMNet), which consists of a cross-modal multi-stage fusion module (CMFM) and a bi-directional multi-scale decoder (BMD). Similar to the mechanism of visual color stage doctrine in human visual system, the proposed CMFM aims to explore the useful and important feature representations in feature response stage, and effectively integrate them into available cross-modal fusion features in adversarial combination stage. Moreover, the proposed BMD learns the combination of cross-modal fusion features from multiple levels to capture both local and global information of salient objects and further reasonably boost the performance of the proposed method. Comprehensive experiments demonstrate that the proposed method can achieve consistently superior performance over the other 14 state-of-the-art methods on six popular RGB-D datasets when evaluated by 8 different metrics. Guibiao Liao, Wei Gao 0003, Qiuping Jiang, Ronggang Wang, Ge Li 0002 |
ACM Multimedia | 3 |
| 2020 | Blind quality assessment for image superresolution using deep two-stream convolutional networks
Wei Zhou 0021, Qiuping Jiang, Yuwang Wang, Zhibo Chen 0001, Weiping Li 0003 |
Inf. Sci. | 2 |
| 2020 | Blind quality assessment for multiply distorted stereoscopic images towards IoT-based 3D capture systems
Xuejin Wang, Meiling Qi, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | A large-scale remote sensing database for subjective and objective quality assessment of pansharpened images
Yiming Xiong, Feng Shao 0001, Xiangchao Meng, Qiuping Jiang, Weiwei Sun 0005, Randi Fu, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | Learning a Unified Blind Image Quality Metric via On-Line and Off-Line Big Training InstancesabstractIn this work, we resolve a big challenge that most current image quality metrics (IQMs) are unavailable across different image contents, especially simultaneously coping with natural scene (NS) images or screen content (SC) images. By comparison with existing works, this paper deploys on-line and off-line data for proposing a unified no-reference (NR) IQM, not only applied to different distortion types and intensities but also to various image contents including classical NS images and prevailing SC images. Our proposed NR IQM is developed with two data-driven learning processes following feature extraction, which is based on scene statistic models, free-energy brain principle, and human visual system (HVS) characteristics. In the first process, the scene statistic models and an image retrieve technique are combined, based on on-line and off-line training instances, to derive a novel loose classifier for retrieving clean images and helping to infer the image content. In the second process, the features extracted by incorporating the inferred image content, free-energy and low-level perceptual characteristics of the HVS are learned by utilizing off-line training samples to analyze the distortion types and intensities and thereby to predict the image quality. The two processes mentioned above depend on a gigantic quantity of training data, much exceeding the number of images applied to performance validation, and thus make our model's performance more reliable. Through extensive experiments, it has been validated that the proposed blind IQM is capable of simultaneously inferring the quality of NS and SC images, and it has attained superior performance as compared with popular and state-of-the-art IQMs on the subjective NS and SC image quality databases. The source code of our model will be released with the publication of the paper at https://kegu.netlify.com. Ke Gu 0001, Junfei Qiao 0001, Qiuping Jiang, Weisi Lin, Daniel Thalmann |
IEEE Trans. Big Data | 4 |
| 2020 | MSTGAR: Multioperator-Based Stereoscopic Thumbnail Generation With Arbitrary ResolutionabstractAt present, thumbnail generation for 2D images has been extensively studied, but the research in thumbnail generation for stereoscopic images is still relatively lacking. This paper presents a novel thumbnail generation technology for stereoscopic images based on multioperator with the following innovations: 1) The warping technique is used to retarget a stereopair into six-scale resolutions with different contexts, and the disparity is uniformly adjusted to a certain value based on just noticeable depth difference (JNDD) model, which overcomes the issues that 3D perception in stereoscopic thumbnail is uncontrollable and the sense of depth disappears in low-resolution stereoscopic images. 2) The six-scale images are cropped via cropping network, and are optimized to a target resolution based on the designed image visual representation energy. As a result, our method has better visual effect than state-of-the-art methods in generating thumbnail for stereoscopic display. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Multim. | 3 |
| 2020 | A Dilated Inception Network for Visual Saliency PredictionabstractRecently, with the advent of deep convolutional neural networks (DCNN), the improvements in visual saliency prediction research are impressive. One possible direction to approach the next improvement is to fully characterize the multi-scale saliency-influential factors with a computationally-friendly module in DCNN architectures. In this work, we propose an end-to-end dilated inception network (DINet) for visual saliency prediction. It captures multi-scale contextual features effectively with very limited extra parameters. Instead of utilizing parallel standard convolutions with different kernel sizes as the existing inception module, our proposed dilated inception module (DIM) uses parallel dilated convolutions with different dilation rates which can significantly reduce the computation load while enriching the diversity of receptive fields in feature maps. Moreover, the performance of our saliency model is further improved by using a set of linear normalization-based probability distribution distance metrics as loss functions. As such, we can formulate saliency prediction as a global probability distribution prediction task for better saliency inference instead of a pixel-wise regression problem. Experimental results on several challenging saliency benchmark datasets demonstrate that our DINet with proposed loss functions can achieve state-of-the-art performance with shorter inference time. Sheng Yang 0006, Guosheng Lin, Qiuping Jiang, Weisi Lin |
IEEE Trans. Multim. | 3 |
| 2019 | Encoding Complexity Control for Live Video Applications: An Interpretable Machine Learning ApproachabstractIn this paper, we propose an interpretable machine learning-based complexity control method for efficiently im-plementing HEVC on live video applications with different computing capacities and limited powers. Specifically, a complexity allocation method is designed to reasonably assign the complexity resources. Then, a multi-accuracy Coding Unit (CU) decision model is obtained by interpret-ably adjusting the parameters to efficiently and flexibly achieve a tradeoff between encoding complexity and rate distortion performance. Finally, a coding tree unit-level complexity control method is proposed to select appropri-ate accuracy of the CU decision model for making the en-coding complexity approach the target. The experimental results show that the proposed method outperforms state-of-the-art methods in terms of accuracy and encoding efficiency. Chao Huang 0008, Zongju Peng, Qiuping Jiang, Gangyi Jiang |
ICME | 4 |
| 2019 | SGDNet: An End-to-End Saliency-Guided Deep Neural Network for No-Reference Image Quality AssessmentabstractWe propose an end-to-end saliency-guided deep neural network (SGDNet) for no-reference image quality assessment (NR-IQA). Our SGDNet is built on an end-to-end multi-task learning framework in which two sub-tasks including visual saliency prediction and image quality prediction are jointly optimized with a shared feature extractor. The existing multi-task CNN-based NR-IQA methods which usually consider distortion identification as the auxiliary sub-task cannot accurately identify the complex mixtures of distortions exist in authentically distorted images. By contrast, our saliency prediction sub-task is more universal because visual attention always exists when viewing every image, regardless of its distortion type. More importantly, related works have reported that saliency information is highly correlated with image quality while this property is fully utilized in our proposed SGNet by training the model with more informative labels including saliency maps and quality scores simultaneously. In addition, the outputs of the saliency prediction sub-task are transparent to the primary quality regression sub-task by providing a kind of spatial attention masks for a more perceptually-consistent feature fusion. By training the whole network with the two sub-tasks together, more discriminant features can be learned and a more accurate mapping from feature representations to quality scores can be established. Experimental results on both authentically and synthetically distorted IQA datasets demonstrate the superiority of our SGDNet, as compared to the state-of-the-art approaches. Sheng Yang 0006, Qiuping Jiang, Weisi Lin, Yongtao Wang |
ACM Multimedia | 2 |
| 2019 | Learning content-specific codebooks for blind quality assessment of screen content images
Yongqiang Bai, Mei Yu 0001, Qiuping Jiang, Gangyi Jiang, Zhongjie Zhu |
Signal Process. | 3 |
| 2019 | Authentically Distorted Image Quality Assessment by Learning From Empirical Score DistributionsabstractMost existing works on image quality assessment (IQA) focus on predicting a scalar quality score (SQS) based on the assumption that people can reach a consensus on the judgment of image quality. However, assigning a single scalar fails to reveal the subjective diversity that an image will probably receive divergent opinion scores from different subjects. This is particularly true for real-world authentically distorted images which usually involve composite mixtures of multiple distortions. To characterize such an property, this letter proposes to use a more informative vectorized label called empirical score distribution (ESD) to build an ESD-aided deep neural network (DNN) for authentically distorted image quality prediction. Our proposed network contains two streams: ESD prediction stream and SQS prediction stream. The whole DNN is optimized end-to-end with a combined loss so that both of the supervision information from ESD and SQS can be fully utilized in the training process. Experiments on two public authentically distorted image databases verify the superiority of our method. Qiuping Jiang, Zhenyu Peng, Sheng Yang 0006, Feng Shao 0001 |
IEEE Signal Process. Lett. | 1 |
| 2019 | A Risk-Aware Pairwise Rank Learning Approach for Visual Discomfort Prediction of Stereoscopic 3DabstractFor visual discomfort prediction (VDP) of stereoscopic 3D images, a common two-stage framework is to first extract features that are predictive of the experienced visual discomfort level when viewing stereoscopic images and then use typical regression tools to learn the mapping from the extracted features to visual discomfort scores. Most existing approaches for stereoscopic 3D VDP focus on the former stage, i.e., feature extraction, while limited efforts have dedicated to exploiting more powerful and robust learning algorithms in this field. In this letter, inspired by the pairwise comparison-based subjective evaluation methodology, we propose a novel Risk-Aware Pairwise Rank Learning (RAPRL) approach to further improve the prediction accuracy. Unlike the traditional VDP approaches using different regression tools for feature-score mapping, our proposed RARL method addresses this problem based on a completely different pairwise rank learning framework with a risk-aware constraint. Experiments have verified the effectiveness and robustness of our proposed VDP model using RAPRL as the learning algorithm. Qiuping Jiang, Feng Shao 0001, Wei Gao 0003, Yo-Sung Ho |
IEEE Signal Process. Lett. | 1 |
| 2019 | Deep Road Scene UnderstandingabstractRoad scene understanding is a difficult task in autonomous driving. In this letter, we propose a novel deep encoder-decoder architecture for road scene understanding in an end-to-end manner. This core trainable understanding engine includes an encoder network, a decoder network with two streams, and a pixel-level fusion network with classification layer. The encoder network is composed of the front-end model of the classical convolution neural network, VGGNet. The decoder network with two streams includes multi-scale skip connection modules to reduce the down-scaling effect. Finally, a fusion network fuses the two-level information from the two streams of the decoder network for precise pixel-level classification. Additionally, the convolution layer is added to each skip connection module to increase the depth of the architecture. Our architecture achieves outstanding performance on the publicly available CamVid dataset and significantly outperforms previous architectures. This deep architecture is ideal for road scene understanding. Wujie Zhou, Sijia Lv, Qiuping Jiang, Lu Yu 0003 |
IEEE Signal Process. Lett. | 3 |
| 2019 | BLIQUE-TMI: Blind Quality Evaluator for Tone-Mapped Images Based on Local and Global Feature AnalysesabstractHigh dynamic range (HDR) image, which has a powerful capacity to represent the wide dynamic range of real-world scenes, has been receiving attention from both academic and industrial communities. Although HDR imaging devices have become prevalent, the display devices for HDR images are still limited. To facilitate the visualization of HDR images in standard low dynamic range displays, many different tone mapping operators (TMOs) have been developed. To create a fair comparison of different TMOs, this paper proposes a BLInd QUality Evaluator to blindly predict the quality of Tone-Mapped Images (BLIQUE-TMI) without accessing the corresponding HDR versions. BLIQUE-TMI measures the quality of TMIs by considering the following aspects: 1) visual information; 2) local structure; and 3) naturalness. To be specific, quality-aware features related to the former two aspects are extracted in a local manner based on sparse representation, while quality-aware features related to the third aspect are derived based on global statistics modeling in both intensity and color domains. All the extracted local and global quality-aware features constitute a final feature vector. An emergent machine learning technique, i.e., extreme learning machine, is adopted to learn a quality predictor from feature space to quality space. The superiority of BLIQUE-TMI to several leading blind IQA metrics is well demonstrated on two benchmark databases. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Unified No-Reference Quality Assessment of Singly and Multiply Distorted Stereoscopic ImagesabstractA challenging problem in the no-reference quality assessment of multiply distorted stereoscopic images (MDSIs) is to simulate the monocular and binocular visual properties under a mixed type of distortions. Due to the joint effects of multiple distortions in MDSIs, the underlying monocular and binocular visual mechanisms have different manifestations with those of singly distorted stereoscopic images (SDSIs). This paper presents a unified no-reference quality evaluator for SDSIs and MDSIs by learning monocular and binocular local visual primitives (MB-LVPs). The main idea is to learn MB-LVPs to characterize the local receptive field properties of the visual cortex in response to SDSIs and MDSIs. Furthermore, we also consider that the learning of primitives should be performed in a task-driven manner. For this, two penalty terms including reconstruction error and quality inconsistency are jointly minimized within a supervised dictionary learning framework, generating a set of quality-oriented MB-LVPs for each single and multiple distortion modality. Given an input stereoscopic image, feature encoding is performed using the learned MB-LVPs as codebooks, resulting in the corresponding monocular and binocular responses. Finally, responses across all the modalities are fused with probabilistic weights which are determined by the modality-specific sparse reconstruction errors, yielding the final monocular and binocular features for quality regression. The superiority of our method has been verified on several SDSI and MDSI databases. Qiuping Jiang, Feng Shao 0001, Wei Gao 0003, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Image Process. | 1 |
| 2018 | Local and global sparse representation for no-reference quality assessment of stereoscopic images
Fucui Li, Feng Shao 0001, Qiuping Jiang, Randi Fu, Gangyi Jiang, Mei Yu 0001 |
Inf. Sci. | 3 |
| 2018 | Learning a referenceless stereopair quality engine with deep nonnegativity constrained sparse autoencoder
Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
Pattern Recognit. | 1 |
| 2018 | Blind stereoscopic 3D image quality assessment via analysis of naturalness, structure, and binocular asymmetry
Guanghui Yue 0001, Chunping Hou, Qiuping Jiang, Yang Yang 0045 |
Signal Process. | 3 |
| 2018 | Toward Domain Transfer for No-Reference Quality Prediction of Asymmetrically Distorted Stereoscopic ImagesabstractWe have presented a no-reference quality prediction method for asymmetrically distorted stereoscopic images, which aims to transfer the information from source feature domain to its target quality domain using a label consistent K-singular value decomposition classification framework. To this end, we construct a category-deviation database for dictionary learning that assigns a label for each stereoscopic image to indicate if it is noticeable or unnoticeable by human eyes. Then, by incorporating a category consistent term into the objective function, we learn view-specific feature and quality dictionaries to establish a semantic framework between the source feature domain and the target quality domain. The quality pooling is comparatively simple and only needs to estimate the quality score based on the classification probability. The experimental results demonstrate the effectiveness of our blind metric. Feng Shao 0001, Zhuqing Zhang, Qiuping Jiang, Weisi Lin, Gangyi Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Learning Sparse Representation for Objective Image Retargeting Quality AssessmentabstractThe goal of image retargeting is to adapt source images to target displays with different sizes and aspect ratios. Different retargeting operators create different retargeted images, and a key problem is to evaluate the performance of each retargeting operator. Subjective evaluation is most reliable, but it is cumbersome and labor-consuming, and more importantly, it is hard to be embedded into online optimization systems. This paper focuses on exploring the effectiveness of sparse representation for objective image retargeting quality assessment. The principle idea is to extract distortion sensitive features from one image (e.g., retargeted image) and further investigate how many of these features are preserved or changed in another one (e.g., source image) to measure the perceptual similarity between them. To create a compact and robust feature representation, we learn two overcomplete dictionaries to represent the distortion sensitive features of an image. Features including local geometric structure and global context information are both addressed in the proposed framework. The intrinsic discriminative power of sparse representation is then exploited to measure the similarity between the source and retargeted images. Finally, individual quality scores are fused into an overall quality by a typical regression method. Experimental results on several databases have demonstrated the superiority of the proposed method. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
IEEE Trans. Cybern. | 1 |
| 2018 | Optimizing Multistage Discriminative Dictionaries for Blind Image Quality AssessmentabstractState-of-the-art algorithms for blind image quality assessment (BIQA) typically have two categories. The first category approaches extract natural scene statistics (NSS) as features based on the statistical regularity of natural images. The second category approaches extract features by feature encoding with respect to a learned codebook. However, several problems need to be addressed in existing codebook-based BIQA methods. First, the high-dimensional codebook-based features are memory-consuming and have the risk of over-fitting. Second, there is a semantic gap between the constructed codebook by unsupervised learning and image quality. To address these problems, we propose a novel codebook-based BIQA method by optimizing multistage discriminative dictionaries (MSDDs). To be specific, MSDDs are learned by performing the label consistent K-SVD (LC-KSVD) algorithm in a stage-by-stage manner. For each stage, a new quality consistency constraint called “quality-discriminative regularization” term is introduced and incorporated into the reconstruction error term to form a unified objective function, which can be effectively solved by LC-KSVD for discriminative dictionary learning. Then, the latter stage takes the reconstruction residual data in the former stage as input based on which LC-KSVD is repeatedly performed until the final stage is reached. Once the MSDDs are learned, multistage feature encoding is performed to extract feature codes. Finally, the feature codes are concatenated across all stages and aggregated over the entire image for quality prediction via regression. The proposed method has been evaluated on five databases and experimental results well confirm its superiority over existing relevant BIQA methods. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Ke Gu 0001, Gangyi Jiang, Huifang Sun |
IEEE Trans. Multim. | 1 |
| 2018 | Multistage Pooling for Blind Quality Prediction of Asymmetric Multiply-Distorted Stereoscopic ImagesabstractQuality prediction for asymmetric multiply-distorted stereoscopic images (MDSIs) confronts more challenges than previous stereoscopic image quality assessment (SIQA) issues, whereas the existing no-reference SIQA methods have been limited to understand the asymmetric distortions and multiple distortions simultaneously for general-purpose blind quality prediction. In this paper, we propose a multistage pooling (MUSP) model for quality prediction of asymmetric MDSIs. In the training stage, we establish multimodal sparse representation framework for phase and amplitude components, respectively. In the testing stage, we use an MUSP strategy to simulate the pooling procedure undergoing multimodal quality pooling, feature pooling, binocular pooling, and phase-amplitude quality pooling in order. Experimental results on our new established database (NBU-MDSID Phase-II) demonstrate the effectiveness of our blind metric. Feng Shao 0001, Qiuping Jiang, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Multim. | 3 |
| 2017 | MSFE: Blind image quality assessment based on multi-stage feature encodingabstractBlind image quality assessment (BIQA) methods based on visual codebooks have received much attention due to its prominent generalization capacity across different image domains. Existing codebook-based BIQA methods depend on large-size codebooks and high-dimensional features, which are memory-consuming and have the risk of over-fitting. Thus, it is necessary to design quality metrics with much smaller codebooks. This paper presents a novel multistage feature encoding (MSFE)-based BIQA method which requires much lower dimensional features while preserving comparable or even better performance. To specify, MSFE is performed over multiple cascaded and much smaller sub-codebooks to generate more compact and discriminative features for quality prediction. The latter stage takes the encoding residuals in the former stage as input. We use KSVD and sparse coding for codebook training and feature encoding in the framework, respectively. Finally, the generated sparse feature codes in all stages are combined and aggregated over the entire image for quality prediction via support vector regression (SVR). We evaluate the proposed method on several natural and screen content image databases. The experimental results confirm its superiority in terms of both validity and universality. Qiuping Jiang, Feng Shao 0001, Gangyi Jiang |
ICIP | 1 |
| 2017 | Visual comfort assessment for stereoscopic images based on sparse coding with multi-scale dictionaries
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
Neurocomputing | 1 |
| 2017 | Leveraging visual attention and neural activity for stereoscopic 3D visual comfort assessment
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
Multim. Tools Appl. | 1 |
| 2017 | QoE-Guided Warping for Stereoscopic Image RetargetingabstractIn the field of stereoscopic 3D (S3D) display, it is an interesting as well as meaningful issue to retarget the stereoscopic images to the target resolution, while the existing stereoscopic image retargeting methods do not fully take user's Quality of Experience (QoE) into account. In this paper, we have presented a QoE-guided warping method for stereoscopic image retargeting, which retarget the stereoscopic image and adapt its depth range to the target display while promoting user's QoE. Our method takes shape preservation, visual comfort preservation, and depth perception preservation energies into account, and simultaneously optimizes the 2D coordinates and depth information in 3D space. It also considers the specific viewing configuration in the visual comfort and depth perception preservation energy constraints. Experimental results on visually uncomfortable and comfortable stereoscopic images demonstrate that in comparison with the existing stereoscopic image retargeting methods, the proposed method can achieve a reasonable performance optimization among the QoE's factors of image quality, visual comfort, and depth perception, leading to promising overall S3D experience. Feng Shao 0001, Wenchong Lin, Weisi Lin, Qiuping Jiang, Gangyi Jiang |
IEEE Trans. Image Process. | 4 |
| 2016 | On Predicting Visual Comfort of Stereoscopic Images: A Learning to Rank Based ApproachabstractPredicting the degree of experienced visual comfort in the context of stereoscopic 3-D (S3D) viewing is particularly challenging. In this letter, a simple yet effective visual comfort assessment (VCA) approach for stereoscopic images is proposed from the perspective of learning to rank (L2R). The proposed L2R-based VCA (L2R-VCA) approach is inspired by the traditional absolute categorical rating (ACR) methodology in subjective study and is to characterize the qualitative description behavior of human subjective study. Experimental results on our recently built database confirm the promising performance of the proposed L2R-VCA approach, yielding higher consistency with human subject judgment results. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
IEEE Signal Process. Lett. | 1 |
| 2015 | 3D Visual Comfort Assessment via Sparse Coding
Qiuping Jiang, Feng Shao 0001 |
ICIG (1) | 1 |
| 2015 | Supervised dictionary learning for blind image quality assessmentabstractIn this paper, we propose a supervised dictionary learning framework for blind image quality assessment (BIQA) by using quality-constraint sparse coding. Different with the traditional dictionary learning framework which only ensures the learnt dictionary accounting for image features, we add a quality-related regularization term in the framework to learn a feature-related dictionary and a quality-related dictionary jointly. Specifically, the feature-related and quality-related dictionaries share the same sparse coefficients, so that the reconstruction errors form the image feature vectors and quality score vectors are both minimized. Once the feature-related and quality-related dictionaries are learned, given a testing sample, we first abstract its feature vector and then compute the corresponding sparse coefficients w.r.t. the learnt feature-related dictionary, its quality score can be directly reconstructed based on the learnt quality-related dictionary and the estimated sparse coefficients. Experiment results on three publicly available IQA databases show the promising performance of the proposed model. Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
VCIP | 1 |
| 2015 | Supervised dictionary learning for blind image quality assessment using quality-constraint sparse coding
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | A depth perception and visual comfort guided computational model for stereoscopic 3D visual saliency
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng, Changhong Yu |
Signal Process. Image Commun. | 1 |