EDBT 2026 Demo / reviewers in the wild / expert
Bo Hu 0008
dblp:04/2380-8
· DBLP profile ↗
35ranked-venue papers
18as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 15 first-author · 17 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HVS-inspired blind image quality index with prominent perception learning and multi-level progressive integration
Taiyang Chen, Bo Hu 0008, Chunyi Li 0001, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001 |
Neurocomputing | 2 |
| 2026 | Perceptual Quality Assessment of Low-Light Enhanced Images: A Multi-Annotated Subjective Dataset and a Multimodal Objective MethodabstractLow-light Image Enhancement Algorithms (LIEAs) aim to improve the visibility and visual quality of images captured in low-light environments. However, none of the existing LIEAs can comprehensively restore all visual contents, which makes it inevitable for the Enhanced Low-light Images (ELIs) to have different degrees of distortion, thereby affecting the visual quality. Currently, there is little research focusing on the quality assessment of these ELIs, partly due to the lack of publicly available datasets. Moreover, existing quality assessment methods primarily focus on a single visual modality and fail to sufficiently exploit the structural information across multiple image attributes, consequently resulting in suboptimal prediction performance. To this end, this paper conducts a systematic study on both subjective and objective quality assessment of ELIs. Firstly, we construct the first Multi-annotated and multi-modal Low-light image Enhancement quality dataset (MLE), which contains 1,000 ELIs, along with subjective studies to obtain multiple attribute annotations, quality scores, and textual descriptions. Based on this, we further propose an Attribute-guided Vision-Language Graph Reasoning Network (AVGR-Net) for ELI quality prediction, which effectively integrates multi-attribute visual and textual information through cross-modal graph reasoning and alignment. Extensive data analysis and experimental results validate both the reliability of the MLE dataset and the superior performance of the AVGR-Net compared to state-of-the-art methods. Bo Hu 0008, Leida Li, Ke Gu 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | IBCL-VQA: Video Quality Assessment with In-Batch Contrastive Learning and Two-Phase Feature FusionabstractVideo Quality Assessment (VQA) technology is of significant importance for improving video transmission, storage, and processing. Although Convolutional Neural Networks (CNNs)-based and Transformer-based methods have achieved significant progress, they still suffer from some drawbacks. The previous methods treated video data as independent samples, thereby neglecting the close or distant relationships between different quality levels and consequently constraining the model’s discriminative ability; meanwhile, most existing fusion strategies utilize fixed architectures that lack the ability to adapt to the data for optimal integration, resulting in insufficient utilization of spatio-temporal information and limited expression capabilities of the fused features. To address the above issues, this article proposes a VQA method with In-Batch Contrastive Learning and Two-Phase Feature Fusion (IBCL-VQA). Firstly, the spatial features are extracted through two branches, which not only preserve the global semantics but also focus on the local regions. The temporal characteristics are obtained through a pre-trained video recognition model. Secondly, we propose an in-batch contrastive learning mechanism which, through the principles of maximizing intra-class similarity and minimizing inter-class similarity, combined with a dynamically adjusted penalty strategy, models the correlation between video quality levels. Thirdly, a two-phase feature fusion strategy, consisting of the Gated Spatio-temporal Attention Unit (GSTU) and the Adaptive Fusion Cell (AFC), is further proposed. The former achieves spatio-temporal feature fusion through dynamic weight allocation, and the latter adaptively integrates the features from the two branches based on data characteristics. Finally, a regression module outputs the quality score. Experimental results on five real-world VQA datasets demonstrate the superior performance of the IBCL-VQA. Furthermore, the strong generalizability is verified through cross-database testing. The code and pre-trained weights will be publicly available at: https://github.com/BoHu90/IBCL-VQA . Bo Hu 0008, Leida Li, Lihuo He, Xinbo Gao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Bidirectional Reference Image Quality Assessment via Content-Quality Correlation ModelingabstractThe emphasis on no-reference image quality assessment has often overshadowed the significance of Full-Reference Image Quality Assessment (FR-IQA), which generally better reflects human contrastive perception mechanism. However, FRIQA presents challenges in obtaining content-aligned reference images. To tackle these issues, a novel Bidirectional Reference Image Quality Assessment (BRIQA) method is proposed, centering on leveraging bidirectional reference images and content-quality correlation modeling. First, triplets of content-aligned low-quality and content-non-aligned high-quality reference images are generated using two easily accessible approaches. To prevent the extraction of redundant information, two feature extractors pretrained through unsupervised contrastive learning are utilized to independently extract content and quality features for the triplet images. Then, an attention-mixer is introduced to further mine quality difference information and enhance content feature. Finally, a content-quality correlation modeler is proposed to model the relationship between quality differences and visual contents. Experimental results on benchmark datasets demonstrate that the BRIQA outperforms existing state-of-the-art methods. Bo Hu 0008, Wenzhi Chen, Chunyi Li 0001, Jiaxu Leng, Weisheng Li 0001, Xinbo Gao 0001 |
ICASSP | 1 |
| 2025 | AGIAA-2K: A Fine-grained Dataset for Aesthetic and Alignment Evaluation of AI-Generated ImagesabstractWith the advancement of AI-generated content technologies, AI-generated images (AGIs) have become increasingly influential in artistic creation and visual communication. However, the aesthetic quality of AGIs varies significantly due to technical limitations and the influence of user input, underscoring the urgent need for systematic aesthetic evaluation of AGIs. In addition, it is difficult to ensure the consistency of text-to-image, which compresses the application space of AGIs. To address these issues, a fine-grained dataset for Aesthetic and Alignment evaluation of AGIs (AGIAA-2K) is presented. This dataset contains 2,064 images generated using 172 well-designed prompts across six different AGI models, with each image annotated based on the subjective experiment. Then, the rationality of the dataset is verified by data analysis. Finally, the performances of the existing algorithms are evaluated in terms of image aesthetic assessment and text-to-image alignment of AGIs. The results demonstrate that these algorithms cannot effectively evaluate these two aspects. The AGIAA-2K is available at https://github.com/BoHu90/AGIAA-2K. Bo Hu 0008, Nanxiang Li, Lihuo He, Wen Lu 0004, Leida Li, Xinbo Gao 0001 |
ICASSP | 1 |
| 2025 | A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality AssessmentabstractWide-angle video is favored for its wide viewing angle and ability to capture a large area of scenery, making it an ideal choice for sports and adventure recording. However, wide-angle video is prone to deformation, exposure and other distortions, resulting in poor video quality and affecting the perception and experience, which may seriously hinder its application in fields such as competitive sports. Up to now, few explorations focus on the quality assessment issue of wide-angle video. This deficiency primarily stems from the absence of a specialized dataset for wide-angle videos. To bridge this gap, we construct the first Multi-annotated and multi-modal Wide-angle Video quality assessment (MWV) dataset. Then, the performances of state-of-the-art video quality methods on the MWV dataset are investigated by inter-dataset testing and intra-dataset testing. Experimental results show that these methods impose significant limitations on their applicability. Bo Hu 0008, Chunyi Li 0001, Lihuo He, Leida Li, Xinbo Gao 0001 |
ICASSP | 1 |
| 2025 | A Two-Stage AIGC Image Quality Assessment with T2I Correspondence and Visual PerceptionabstractImage quality assessment (IQA) of artificial intelligence-generated content (AIGC) has recently attracted significant research attention. Unlike general-purpose IQA, which primarily focuses on evaluating image content, AIGCIQA often requires addressing both the Text-to-Image (T2I) correspondence and the perceptual quality of images. To address this requirement, this paper proposes a novel two-stage AIGCIQA method. The first stage evaluates the alignment of the AI-generated images (AIGIs) with their corresponding descriptions, serving as an indicator of overall image quality. Specifically, positive and negative prompts are constructed to describe the T2I correspondence degree, and then a CLIP model is employed to predict the degree based on these image-prompt pairs. The second stage refines the perceptual quality assessment by integrating both global and local degradation features of AIGIs. Importantly, the contribution of local features is measured according to their correlation with the overall image, ensuring key regions are adequately represented in the quality prediction. Experimental results on AGIQA-1K, AGIQA-3K, and AIGCIQA2023 demonstrate the superior performance of the proposed method. Jili Xia, Lihuo He, Bo Hu 0008, Bo Han 0004, Xinbo Gao 0001 |
ICASSP | 3 |
| 2025 | MACA-VQA: Quality Assessment of UGC Videos via Multi-level Distortion Adaptation and Spatiotemporal Cross-Attention FusionabstractUser-generated content (UGC) videos often exhibit complex distortions and diverse content, posing significant challenges for traditional video quality assessment (VQA) methods. Approaches that directly merge distortion and semantic information risk feature conflicts and the loss of details. In addition, simple concatenation of spatiotemporal features fails to capture vital interactions, limiting predictive accuracy. Motivated by these challenges, this paper proposes a Multi-level Distortion Adaptation and Spatiotemporal Cross-Attention Fusion framework for VQA, named MACA-VQA. Specifically, a novel multi-level adaptive strategy progressively incorporates distortion information into each Transformer layer of the CLIP model, enabling layer-wise fusion of semantic and distortion features. Furthermore, a newly introduced cross-attention fusion mechanism dynamically integrates spatiotemporal features, capturing complex, multidimensional interactions. Extensive experiments demonstrate that MACA-VQA achieves state-of-the-art performance on multiple public datasets, validating its effectiveness and robustness in both intra-dataset and inter-dataset scenarios. The source code is available at https://github.com/BoHu90/MACA-VQA Bo Hu 0008, Yimeng Zhao, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001 |
ICME | 1 |
| 2025 | Low-light Image Enhancement Quality Assessment: A Real-World Dataset and An Objective MethodabstractLow-light Image Enhancement (LIE) technology adaptively improves brightness while preserving texture details and suppressing noise artifacts, thereby reducing visual degradation caused by insufficient illumination. While deep learning-based image enhancement algorithms have made significant progress, a key gap remains in establishing standardized methods for fairly evaluating and comparing their performance. To bridge this gap, this paper systematically investigates enhanced low-light image quality assessment from both subjective and objective dimensions. First, we introduce a Real-world Low-light Image Enhancement quality assessment dataset (RLIE), which contains 1540 images from 154 scenarios, each with a subjective score given by the subjects. Based on this, we propose a low light enhanced image quality assessment method based on Multi-level Illumination Injection and Hierarchical Discrepancy Perception (MIIHDP). The core idea of this method is to hierarchically inject separated illumination information into the feature extraction process, then tailor the processing of difference information at different scales to obtain a more comprehensive representation. Finally, extensive statistical analyses demonstrate the rationality of the proposed RLIE dataset, and experimental results show the superior performance of the proposed MIIHDP compared with state-of-the-arts. Our dataset and code are released at: https://github.com/BoHu90/RLIE. Chunyi Li 0001, Bo Hu 0008, Taiyang Chen, Leida Li, Lihuo He, Xinbo Gao 0001 |
ACM Multimedia | 2 |
| 2025 | Blind image quality assessment for in-the-wild images by integrating distorted patch selection and multi-scale-and-granularity fusion
Jili Xia, Lihuo He, Xinbo Gao 0001, Bo Hu 0008 |
Knowl. Based Syst. | 4 |
| 2025 | Blind Quality Assessment of Wide-Angle Videos Based on Deformation Representation Learning and Multi-Dimensional Feature FusionabstractWide-angle videos shot with short-focus lenses often exhibit deformation distortions, which poses significant challenges for video quality assessment (VQA). Although current VQA methods focus primarily on video content and distortion perception, there has been little explicit research on the impact of deformation characteristics on the perception of wide-angle video quality. To this end, this paper makes the first attempt to construct a novel wide-angle video quality assessment method based on deformation representation learning and multi-dimensional feature fusion, termed DRLMF. Specifically, we first analyze the deformation distribution characteristics of wide-angle videos based on the deformation camera model. Based on this, a three-stream video perception and assessment network is proposed. The first branch extracts global semantics using the image encoder of CLIP. The second branch introduces an effective deformation region selection strategy and proposes an interpretable deformation representation learning module. This module leverages the perception advantages of convolutional neural networks (CNNs) in local distortions and considers the correlation between patch size and distortion perception. The third branch extracts motion features using an action recognition network. Finally, an effective multi-dimensional feature fusion module is proposed to integrate more refined and richer semantic, deformation, and motion features. Extensive experiments on wide-angle VQA datasets and standard video datasets show that the DRLMF outperforms the state-of-the-arts in terms of prediction monotonicity and accuracy. The codes will be available at https://github.com/BoHu90/DRLMF. Bo Hu 0008, Leida Li, Lihuo He, Wen Lu 0004, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Diffusion Model-Based Visual Compensation Guidance and Visual Difference Analysis for No-Reference Image Quality AssessmentabstractExisting free-energy guided No-Reference Image Quality Assessment (NR-IQA) methods continue to face challenges in effectively restoring complexly distorted images. The features guiding the main network for quality assessment lack interpretability, and efficiently leveraging high-level feature information remains a significant challenge. As a novel class of state-of-the-art (SOTA) generative model, the diffusion model exhibits the capability to model intricate relationships, enhancing image restoration effectiveness. Moreover, the intermediate variables in the denoising iteration process exhibit clearer and more interpretable meanings for high-level visual information guidance. In view of these, we pioneer the exploration of the diffusion model into the domain of NR-IQA. We design a novel diffusion model for enhancing images with various types of distortions, resulting in higher quality and more interpretable high-level visual information. Our experiments demonstrate that the diffusion model establishes a clear mapping relationship between image reconstruction and image quality scores, which the network learns to guide quality assessment. Finally, to fully leverage high-level visual information, we design two complementary visual branches to collaboratively perform quality evaluation. Extensive experiments are conducted on seven public NR-IQA datasets, and the results demonstrate that the proposed model outperforms SOTA methods for NR-IQA. The codes will be available at https://github.com/handsomewzy/DiffV2IQA. Zhaoyang Wang 0003, Bo Hu 0008, Mingyang Zhang 0002, Jie Li 0001, Leida Li, Maoguo Gong, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Visual-Language Multi-Task Blind Image Quality Assessment With Local Quality WeightingabstractThe objective of blind image quality assessment (BIQA) is to develop a model capable of automatically evaluating image quality without requiring any reference knowledge. While multi-task learning has been widely utilized in BIQA, it has predominantly remained unimodal. This paper delves into the Visual-Language multi-task BIQA model, where distortion knowledge can be captured through image-text contrastive learning. Specifically, Visual-Language auxiliary tasks targeting distortion type and quality level are introduced, respectively, where both positive and negative image-text pairs are constructed for the target distorted image. Subsequently, image-text correspondences are learned in the embedding space while simultaneously evaluating image quality. Notably, in the auxiliary task learning, the proposed method not only brings the image and its corresponding positive text prompt closer but also pushes away the image from its negative text prompts, thereby facilitating the extraction of pertinent distortion features. In the quality assessment task, a patch-wise strategy is employed during the training phase. Differing from conventional BIQA methods, a novel NSS-guided quality weighting is introduced to gauge the correlation between patch quality and global quality, thereby enabling precise quality prediction. Extensive experiments are conducted on six IQA datasets, and the experimental results verify the superiority of the proposed method. Jili Xia, Lihuo He, Bo Hu 0008, Leida Li, Xinbo Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Blind image quality index with high-level Semantic Guidance and low-level fine-grained Representation
Bo Hu 0008, Leida Li, Ke Gu 0001, Shuaijian Wang, Weisheng Li 0001, Xinbo Gao 0001 |
Neurocomputing | 1 |
| 2024 | Blind image quality assessment based on hierarchical dependency learning and quality aggregationabstractImage quality assessment (IQA) aims to build a quality prediction model to assess image quality automatically rather than artificially. Due to a lack of reference images, blind image quality assessment (BIQA) has become an attractive yet challenging research topic. Inspired by the hierarchical perception mechanism in the human visual system , some existing BIQA methods aggregate multi-stage features of a convolutional neural network (CNN). However, they are regardless of the latent dependencies. To solve this problem, we propose a novel BIQA method based on hierarchical dependency learning and quality aggregation (HDLaQA). The proposed method includes multi-stage feature extraction, hierarchical dependency learning, and quality aggregation. In multi-stage feature extraction, a CNN is used as the feature extractor and multi-stage features are output for further learning. In hierarchical dependency learning, spatial and channel dependencies among the multi-stage features are modeled. To this end, a dual-head spatial dependency (DSD) module is designed to harvest the spatial dependencies between the adjacent-stage features and deliver these dependencies to the next stage. Moreover, exponential bilinear pooling (EBP) is presented to learn the channel dependencies, which is more stable than commonly used BP. In quality aggregation, multiple quality scores are predicted based on the learned dependencies, and multiple learnable weights are used to measure the importance of the predicted scores for final quality evaluation. Experimental results on seven IQA databases demonstrate the competitiveness of the proposed method on both synthetic and authentic distortions. Jili Xia, Lihuo He, Xinbo Gao 0001, Bo Hu 0008 |
Neurocomputing | 4 |
| 2024 | Multi-branch progressive embedding network for crowd counting
Lifang Zhou, Songlin Rao, Weisheng Li 0001, Bo Hu 0008 |
Image Vis. Comput. | 4 |
| 2024 | Confidence-based dynamic cross-modal memory network for image aesthetic assessment
Xiaodan Zhang 0005, Jinye Peng 0001, Xinbo Gao 0001, Bo Hu 0008 |
Pattern Recognit. | 5 |
| 2024 | Blind Image Quality Index With Cross-Domain Interaction and Cross-Scale IntegrationabstractWith the assistance of Convolutional Neural Networks (CNNs), Image Quality Assessment (IQA) models have made great progress in evaluating both simulated distortion and authentic distortion. However, most of the existing IQA models only learn the features of distorted images, and thus do not make full use of the available feature representation of other domains. Furthermore, the common multi-scale fusion strategies are relatively simple, such as downsampling and concatenating, which further limits the prediction performance. To this end, we propose a novel blind image quality index with cross-domain interaction and cross-scale integration, which is designed based on the combination of CNN and Transformer. First, the hierarchical spatial-domain and gradient-domain representations are obtained through a typical CNN architecture. Then, based on the proposed gradient-query cross-attention, these two types of features are fully interacted in the Cross-Domain Interaction (CDI) module. To represent the distortion information more comprehensively, the Cross-Scale Integration (CSI) module is proposed to combine the information between different scales progressively. Finally, the quality score is obtained through a simple regression module. The experimental results on five public IQA databases of both simulated and authentic scenes show that the proposed model outperforms the compared state-of-the-art metrics. In addition, cross-database experiments show that the proposed model has strong generalization performance. Bo Hu 0008, Leida Li, Ji Gan, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | BMI-Net: A Brain-inspired Multimodal Interaction Network for Image Aesthetic AssessmentabstractImage aesthetic assessment (IAA) has drawn wide attention in recent years as more and more users post images and texts on the Internet to share their views. The intense subjectivity and complexity of IAA make it extremely challenging. Text triggers the subjective expression of human aesthetic experience based on human implicit memory, so incorporating the textual information and identifying the relationship with the image is of great importance for IAA. However, IAA with the image as input fails to fully consider subjectivity, while existing multimodal IAA ignores the interrelationship among modalities. To this end, we propose a brain-inspired multimodal interaction network (BMI-Net) that simulates how the association area of the cerebral cortex processes sensory stimuli. In particular, the knowledge integration LSTM (KI-LSTM) is proposed to learn the image-text interaction relation. The proposed scalable multimodal fusion (SMF) based on low-rank decomposition fuses image, text and interaction modalities to predict the aesthetic distribution. Extensive experiments show that the proposed BMI-Net outperforms existing state-of-the-art methods on three IAA tasks. Xixi Nie, Bo Hu 0008, Xinbo Gao 0001, Leida Li, Xiaodan Zhang 0005, Bin Xiao 0002 |
ACM Multimedia | 2 |
| 2023 | Reduced-reference image deblurring quality assessment based on multi-scale feature enhancement and aggregation
Bo Hu 0008, Shuaijian Wang, Xinbo Gao 0001, Leida Li, Ji Gan, Xixi Nie |
Neurocomputing | 1 |
| 2023 | Characters as graphs: Interpretable handwritten Chinese character recognition via Pyramid Graph Transformer
Ji Gan, Yuyan Chen, Bo Hu 0008, Jiaxu Leng, Weiqiang Wang 0001, Xinbo Gao 0001 |
Pattern Recognit. | 3 |
| 2023 | MLNet: A Multi-Domain Lightweight Network for Multi-Focus Image FusionabstractExisting multi-focus image fusion (MFIF) methods are difficult to achieve satisfactory results in both fusion performance and rate simultaneously. The spatial domain methods are hard to determine the focus/defocus boundary (FDB), and the transform domain methods are likely to damage the content information of the source images. Moreover, the deep learning-based MFIF methods are usually confronted with low rate due to complex models and enormous learnable parameters. To address these issues, we propose a multi-domain lightweight network (MLNet) for MFIF, which can achieve competitive results in both performance and rate. The proposed MLNet mainly includes three modules, namely focus extraction (FE), focus measure (FM) and image fusion (IF). In the interpretable FE module, the image features extracted by discrete cosine transform-based convolution (DCTConv) and local binary pattern-based convolution (LBPConv) are concatenated and fed into the FM module. DCTConv based on transform domain takes DCT coefficients to construct a fixed convolution kernel without parameter learning, which can effectively capture the high/low frequency content of the image. LBPConv based on spatial domain can achieve structure features and gradient information from source images. In the FM module, a 3-layer 1 × 1 convolution with a few learnable parameters is employed to generate the initial decision map, which has the properties of flexible input. The fused image is obtained by the IF module according to the final decision map. In terms of quantitative and qualitative evaluations, extensive experiments validate that the proposed method outperforms existing state-of-the-art methods on three public datasets. In addition, the proposed MLNet contains only 0.01 M parameters, which is 0.2% of the first CNN-based MFIF method [25]. Xixi Nie, Bo Hu 0008, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Boundary-Aware Bias Loss for Transformer-Based Aerial Image Segmentation ModelabstractInspired by the tremendous success of the transformer-based model in natural language processing (NLP), many efforts introduce the transformer-based model into the image processing tasks. However, naive transformer models have to down-sample the image resolution to satisfy computational restrictions, thus discarding the local information, which is catastrophic for high-performance remote sensing image segmentation. Hence, this paper proposes a novel trainable boundary-aware bias loss function to enhance transformer-based models of extracting local information. On the Challenging ISPRS Potsdam dataset, two representative transformer-based models achieve remarkable performance improvements, proving the effectiveness of the proposed method. Yan Zhang 0108, Siqi Liu 0008, Bo Hu 0008, Xinbo Gao 0001 |
ICASSP | 4 |
| 2022 | ICNet: Joint Alignment and Reconstruction via Iterative Collaboration for Video Super-ResolutionabstractMost previous frameworks either cost too much time or adopt some fixed modules resulting in alignment error in video super-resolution (VSR). In this paper, we propose a novel many-to-many VSR framework with Iterative Collaboration (ICNet), which employs the concurrent operation by iterative collaboration between alignment and reconstruction proving to be more efficient and effective than existing recurrent and sliding-window frameworks. With the proposed iterative collaboration, alignment can be conducted on super-resolved features from reconstruction while accurate alignment boosts reconstruction in return. In each iteration, the features of low-resolution video frames are first fed into the alignment and reconstruction subnetworks, which can generate temporal aligned features and spatial super-resolved features. Then, both outputs are fed into the proposed Tidy Two-stream Fusion (TTF) subnetwork that shares inter-frame temporal information and intra-frame spatial information without redundancy. Moreover, we design the Frequency Separation Reconstruction (FSR) subnetwork to not only model high-frequency and low-frequency information separately but also take benefit of each other for better reconstruction. Extensive experiments on benchmark datasets demonstrate that the proposed ICNet outperforms state-of-the-art VSR methods in terms of PSNR/SSIM values and visual quality, respectively. Jiaxu Leng, Jia Wang 0036, Xinbo Gao 0001, Bo Hu 0008, Ji Gan, Chenqiang Gao |
ACM Multimedia | 4 |
| 2022 | Hierarchical discrepancy learning for image restoration quality assessment
Bo Hu 0008, Shuaijian Wang, Leida Li, Jiaxu Leng, Yuzhe Yang 0001, Xinbo Gao 0001 |
Signal Process. | 1 |
| 2020 | No-reference quality assessment for live broadcasting videos in temporal and spatial domainsabstractNowadays, live broadcasting video has become increasingly popular and high‐quality live broadcasting video is highly needed. In practice, live broadcasting videos usually undergo several processing stages, which inevitably introduce multiple distortions, e. g. frame freezing and intensity mutation, causing the degraded quality of experience. However, little work has been done to the quality evaluation of live broadcasting videos, which may hinder the further development of more advanced live broadcasting video delivery systems. Motivated by this, this study presents a no‐reference quality evaluation model for live broadcasting videos (LBVQA) in temporal and spatial domains. In the temporal domain, statistic features are extracted to measure the frame freezing and intensity mutation, and the entropy‐based feature is extracted to describe the global jitter. In the spatial domain, blurring is measured based on phase coherence, and abnormal exposure ratio is calculated based on an adaptive threshold. Finally, all features are fed into a backpropagation neural network to train the quality prediction model. Experimental results on the Live Broadcasting Video Database demonstrate the advantages of the proposed metric over the state‐of‐the‐art image and video quality metrics. Yipo Huang, Leida Li, Yu Zhou 0009, Bo Hu 0008 |
IET Image Process. | 4 |
| 2020 | Subjective and objective quality assessment for image restoration: A critical survey
Bo Hu 0008, Leida Li, Jinjian Wu, Jiansheng Qian |
Signal Process. Image Commun. | 1 |
| 2020 | Perceptual quality assessment for multimodal medical image fusion
Lu Tang 0001, Chuangeng Tian, Leida Li, Bo Hu 0008 |
Signal Process. Image Commun. | 4 |
| 2020 | Blind Quality Index of Depth Images Based on Structural Statistics for View SynthesisabstractThe quality of depth images is crucial for virtual view synthesis. However, the quality assessment of depth images is still largely unexplored. This letter presents a blind quality metric of Depth image based on Structural Statistics (DSS). The design philosophy is inspired by the fact that structural distortion in the depth images usually leads to geometric distortion, which is the main cause for degraded quality of synthesized views. Specifically, the statistical features for shape and orientation are calculated based on discrete orthogonal moments and gradients, generating two groups of quality-aware features. Then, the quality model is built from the extracted statistical features using a regression module. The experimental results demonstrate the effectiveness of the proposed metric. Yipo Huang, Leida Li, Hancheng Zhu, Bo Hu 0008 |
IEEE Signal Process. Lett. | 4 |
| 2019 | Internal generative mechanism driven blind quality index for deblocked images
Bo Hu 0008, Leida Li, Jiansheng Qian |
Multim. Tools Appl. | 1 |
| 2019 | Pairwise-Comparison-Based Rank Learning for Benchmarking Image Restoration AlgorithmsabstractImage restoration has attracted substantial attention recently and many image restoration algorithms have been proposed for restoring latent clear images from degraded images. However, determining how to objectively evaluate the performances of these algorithms remains an open problem, which may hinder the further development of advanced image restoration techniques. Most image restoration-quality metrics are designed for specific restoration applications; hence, their generalization ability is limited. For benchmarking image restoration algorithms, the ranking of restored images that are generated via various algorithms, is the most heavily considered factor. Inspired by this, this paper presents a pairwise-comparison-based rank learning framework for benchmarking the performances of image restoration algorithms, which focuses on the relative quality ranking of restored images. Under the proposed framework, we further propose a general image restoration quality metric by integrating quality-aware features in both the spatial and frequency domains. The proposed metric exhibits good generalization performance, and it is applicable to various restoration applications. The results of extensive experiments that were conducted on eight public databases of five restoration scenarios demonstrate the superior performance of the proposed method over the existing quality metrics. Moreover, the proposed framework is used to improve the existing quality metrics for benchmarking image restoration algorithms and highly encouraging results are obtained. Bo Hu 0008, Leida Li, Hantao Liu, Weisi Lin, Jiansheng Qian |
IEEE Trans. Multim. | 1 |
| 2018 | Internal Generative Mechanism Driven Blind Quality Index for Deblocked ImagesabstractImage deblocking has been widely studied. However, the relevant quality evaluation of deblocked images remains an open problem. The deblocked images are usually contaminated by multiple distortions, typically blocking artifacts and blur. Although various quality metrics have been reported, they are not designed specially for deblocked images, so they cannot accurately predict the quality of deblocked images. To fill this gap, we propose a new quality metric for deblocked images. With the guidance of the internal generative mechanism (IG- M) theory, a deblocked image is first decomposed into two portions, i.e., the predicted and disorderly portions. Then the distortions in the predicted portion are evaluated. Specifically, the distortion-specific features are extracted to evaluate blocking artifacts and blur in the spatial domain, separately. The joint effect of blocking artifacts and blur is evaluated by extracting energy-based features in the Curvelet domain. Finally, all features are combined to train a random forest model for quality prediction of deblocked images. Experimental results conducted on a newly released DeBlocked Image Database (DBID) demonstrate that the proposed metric outperforms the existing relevant quality metrics. Bo Hu 0008, Leida Li, Jiansheng Qian |
ICIP | 1 |
| 2018 | Perceptual quality evaluation for motion deblurringabstractMotion deblurring has been widely studied. However, the relevant quality evaluation of motion deblurred images remains an open problem. The motion deblurred images are usually contaminated by noise, ringing and residual blur (NRRB) simultaneously. Unfortunately, most of the existing quality metrics are not designed for multiply distorted images, so they are limited in predicting the quality of motion deblurred images. In this study, the authors propose a new quality metric for motion deblurred images by measuring NRRB. For a motion deblurred image, the noise level is first estimated. Then the ringing effect is measured by incorporating visual saliency model to adapt to the characteristic of the human visual system. A reblurring‐based method is proposed to extract similarity features between a motion deblurred image and its re‐blurred version for evaluating the residual blur. Finally, the overall quality score of a motion deblurred image is obtained by pooling the scores of noise, ringing and blur. Experimental results conducted on a motion deblurring database demonstrate that the proposed metric significantly outperforms the existing quality metrics. In addition, the proposed NRRB metric is used for improving the existing general‐purpose no‐reference metrics, and very encouraging results are achieved. Bo Hu 0008, Leida Li, Jiansheng Qian |
IET Comput. Vis. | 1 |
| 2017 | No-reference quality assessment of compressive sensing image recovery
Bo Hu 0008, Leida Li, Jinjian Wu, Shiqi Wang 0001, Lu Tang 0001, Jiansheng Qian |
Signal Process. Image Commun. | 1 |
| 2016 | Perceptual evaluation of Compressive Sensing Image RecoveryabstractCompressive sensing (CS) has been attracting tremendous attention in recent years. Extensive CS recovery algorithms have been proposed for effective image reconstruction. However, little work has been dedicated to the perceptual evaluation of CS image recovery algorithms and the corresponding recovered images. In this paper, we first build a Compressive Sensing Recovered Image Database (CSRID), which contains images generated by ten popular CS image recovery algorithms at different sensing rates. We then carry out a subjective experiment using the single-stimulus method to obtain the subjective qualities of the images. The subjective scores are then used to evaluate the performances of the CS image recovery algorithms. Finally, the performances of general-purpose no-reference (NR) quality metrics and image blur metrics are investigated on the CSRID database. Experimental results show that the state-of-the-art quality metrics are very limited in predicting the quality of CS recovered images. Bo Hu 0008, Leida Li, Jiansheng Qian, Yuming Fang 0001 |
QoMEX | 1 |