VLDB 2026 Research / reviewers in the wild / expert
Huasheng Wang
dblp:179/1717
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0009-0003-9290-8445ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIQANet: A Novel Dual-Branch Deep Learning Framework for MRI Image Quality AssessmentabstractImage quality assessment (IQA) algorithms have significantly advanced over the past two decades, primarily focusing on natural images. However, applying these methods directly to medical imaging often yields suboptimal performance due to inherent differences such as the structural complexity of medical images and the limited availability of annotated databases. In this study, we conduct a comprehensive evaluation of state-of-the-art IQA methods, including 29 traditional full-reference (FR), 4 traditional no-reference (NR), and 9 deep learning-based approaches, to assess their effectiveness in the context of medical imaging. Our evaluation is performed on a recently developed MRI image quality assessment benchmark, revealing critical performance gaps in existing methods. Building on these findings, we propose a novel dual-branch deep learning framework specifically designed for medical IQA (MIQANet). The proposed approach effectively combines global contextual information with local structural details, enhancing the model’s ability to capture subtle degradations and structural inconsistencies in MRI scans. Experiential results demonstrate the superiority of our approach over existing methods, providing valuable theoretical and practical insights for enhancing quality assessment of medical images. Yueran Ma, Huasheng Wang, Jean-Yves Tanguy, Phillip Wardle, Elizabeth A. Krupinski, Padraig Corcoran, Hantao Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | KSIQA: A Knowledge-Sharing Model for No-Reference Image Quality AssessmentabstractNo-reference image quality assessment (NR-IQA) aims to quantitatively measure human perception of visual quality without comparing a distorted image to a reference. Despite recent advances, existing NR-IQR approaches often demonstrate insufficient ability to capture perceptual cues in the absence of a reference, limiting their generalisability across diverse and complex real-world image degradations. These limitations hinder their ability to match the reliability of full-reference IQA (FR-IQA) counterparts. A key challenge, therefore, is to enable NR-IQA models to emulate the reference-aware reasoning exhibited by humans and FR-IQA methods. To address this challenge, we propose a novel NR-IQA model based on a knowledge-sharing (KS) strategy to simulate this capability and predict image quality more effectively. Specifically, we designate an FR-IQA model as the teacher and an NR-IQA model as the student. Unlike conventional knowledge distillation (KD), our proposed architecture enables the NR-IQA student and FR-IQA teacher to share a decoder rather than being independent models. Furthermore, the student model contains a Mental Imagery Generation (MIG) module to learn mental imagery as the reference. To fully exploit local and global information, we adopt a vision transformer (ViT) branch and a convolutional neural network branch for feature extraction (FE). Finally, a quality-aware regressor (QAR) combined with deep ordinal regression is constructed to infer the quality score. Experiments show that our proposed NR-IQA model, KSIQA, has class-leading performance against current no-reference (NR) techniques across widespread benchmark datasets. Huasheng Wang, Hongchen Tan, Jianxun Lou, Xiaochang Liu, Wei Zhou 0021, Ying Chen 0011, Roger M. Whitaker, Walter Colombo, Hantao Liu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | CLIP-DQA: Blindly Evaluating Dehazed Images from Global and Local Perspectives Using CLIPabstractBlind dehazed image quality assessment (BDQA), which aims to accurately predict the visual quality of dehazed images without any reference information, is essential for the evaluation, comparison, and optimization of image dehazing algorithms. Existing learning-based BDQA methods have achieved remarkable success, while the small scale of DQA datasets limits their performance. To address this issue, in this paper, we propose to adapt Contrastive Language-Image Pre-Training (CLIP), pre-trained on large-scale image-text pairs, to the BDQA task. Specifically, inspired by the fact that the human visual system understands images based on hierarchical features, we take global and local information of the dehazed image as the input of CLIP. To accurately map the input hierarchical information of dehazed images into the quality score, we tune both the vision branch and language branch of CLIP with prompt learning. Experimental results on two authentic DQA datasets demonstrate that our proposed approach, named CLIP-DQA, achieves more accurate quality predictions over existing BDQA methods. The code is available at https://github.com/JunFu1995/CLIP-DQA. Yirui Zeng, Jun Fu 0007, Hadi Amirpour, Huasheng Wang, Guanghui Yue 0001, Hantao Liu, Ying Chen 0011, Wei Zhou 0021 |
ISCAS | 4 |
| 2025 | Vision-based human action quality assessment: A systematic reviewabstractHuman Action Quality Assessment (AQA), which aims to automatically evaluate the performance of actions executed by humans, is an emerging field of human action analysis. Although many review articles have been conducted for human action analysis fields such as action recognition and action prediction, there is a lack of up-to-date and systematic reviews related to AQA. This paper aims to provide a systematic literature review of existing papers on vision-based human AQA. This systematic review was rigorously conducted following the PRISMA guideline through the databases of Scopus , IEEE Xplore, and Web of Science in July 2024. Ninety-six research articles were selected for the final analysis after applying inclusion and exclusion criteria. This review presents an overview of various aspects of AQA, including existing applications, data acquisition methods, public datasets, state-of-the-art methods and evaluation metrics . We observe an increase in the number of studies in AQA since 2019, primarily due to the advent of deep learning methods and motion capture devices. We categorize these AQA methods into skeleton-based and video-based methods based on the data modality used. There are different evaluation metrics for various AQA tasks. SRC is the most commonly used evaluation metric, with fifty-six out of ninety-six selected papers using it to evaluate their models. Sports event scoring, surgical skill evaluation and rehabilitation assessment are the most popular three scenarios in this direction based on existing papers and there are more new scenarios being explored such as piano skill assessment. Furthermore, the existing challenges and future research directions are provided, which can be a helpful guide for researchers to explore AQA. Huasheng Wang, Katarzyna Stawarz, Shiyin Li, Hantao Liu |
Expert Syst. Appl. | 2 |
| 2025 | Adaptive Spatiotemporal Graph Transformer Network for Action Quality AssessmentabstractLong video action quality assessment (AQA) aims to evaluate the performance of long-term actions depicted in a video and produce an overall assessment for action quality. A video of long-term actions often contains more complicated temporal and spatial information than that of short-term actions. However, existing approaches that segment a video into individual clips for independent analysis potentially disrupt the narrative flow and diminish contextual details within and across clips, impeding comprehensive video understanding. To address this challenge, we propose an adaptive spatiotemporal graph transformer network (ASGTN) that combines multiple graph structures and transformer attention mechanisms to capture both local and global contextual information within and across clips in a long video. Specifically, the adaptive spatiotemporal graph (ASG) combines a spatial graph branch, designed to enrich the local nuanced spatiotemporal relations within an individual clip, and a temporal graph branch, tailored to dynamically learn the semantic context across different clips. Furthermore, a transformer encoder is integrated to amplify the global dependencies across clips in the entire video. This structure is designed to preserve narrative coherence and maintain essential contextual details in video-level features. Finally, we employ a level-focused decoder to predict the action quality score distribution. Experiments demonstrate that our model achieves state-of-the-art results on popular AQA datasets. Our code is available athttps://github.com/jiangliu5/ASGTN_AQA. Huasheng Wang, Wei Zhou 0021, Katarzyna Stawarz, Padraig Corcoran, Ying Chen 0011, Hantao Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Chest X-Ray Visual Saliency Modeling: Eye-Tracking Dataset and Saliency Prediction ModelabstractRadiologists' eye movements during medical image interpretation reflect their perceptual-cognitive processes of diagnostic decisions. The eye movement data can be modeled to represent clinically relevant regions in a medical image and potentially integrated into an artificial intelligence (AI) system for automatic diagnosis in medical imaging. In this article, we first conduct a large-scale eye-tracking study involving 13 radiologists interpreting 191 chest X-ray (CXR) images, establishing a best-of-its-kind CXR visual saliency benchmark. We then perform analysis to quantify the reliability and clinical relevance of saliency maps (SMs) generated for CXR images. We develop CXR image saliency prediction method (CXRSalNet), a novel saliency prediction model that leverages radiologists' gaze information to optimize the use of unlabeled CXR images, enhancing training and mitigating data scarcity. We also demonstrate the application of our CXR saliency model in enhancing the performance of AI-powered diagnostic imaging systems. Jianxun Lou, Huasheng Wang, Xinbo Wu, John Cho Hui Ng, Kaveri A. Thakoor, Padraig Corcoran, Ying Chen 0011, Hantao Liu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | A Bioinspired Deep Learning Framework for Saliency-Based Image Quality AssessmentabstractAdvancements in deep learning have led to significant progress in no-reference (NR) image quality assessment (NR-IQA) for evaluating the perceived quality of digital images without relying on a reference. However, existing NR-IQA models remain suboptimal in handling complex and diverse natural images. Visual saliency constitutes a critical element for enhancing the reliability of NR-IQA, but the optimal use of saliency in deep learning-based NR-IQA has not heretofore been significantly explored. In this article, we present a novel method for integrating saliency in NR-IQA, which is motivated by the saliency-based visual search mechanism that different parts of the visual input are visited by the focus of attention (FOA) in the order of decreasing saliency. By dividing saliency into the high and low levels of FOA, we build a bioinspired deep neural network-BioSIQNet-based on a multitask learning (MTL) framework. The network architecture consists of two saliency-specific tasks and one primary image quality assessment (IQA) task. The low and high saliency (HS) are separately encoded and integrated into the early and deeper layers of the IQA network, respectively, analogous to the hierarchical processing in the visual cortex of the brain that allocates low attentional resources to process the simple patterns and high resources to learn intricate representations. We demonstrate that leveraging the synergy between visual attention and image quality perception and joint learning of these interconnected visual tasks can enhance the overall learning capabilities of the primary IQA model. Experiments validate the effectiveness of our proposed BioSIQNet for NR-IQA. Huasheng Wang, Yueran Ma, Hongchen Tan, Xiaochang Liu, Ying Chen 0011, Hantao Liu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Attention-Bridged Modal Interaction for Text-to-Image GenerationabstractWe propose a novel Text-to-Image Generation Network, Attention-bridged Modal Interaction Generative Adversarial Network (AMI-GAN), to better explore modal interaction and perception for high-quality image synthesis. The AMI-GAN contains two novel designs: an Attention-bridged Modal Interaction (AMI) module and a Residual Perception Discriminator (RPD). In AMI, we mainly design a multi-scale attention mechanism to exploit semantics alignment, fusion, and enhancement between text and image, to better refine details and context semantics of the synthesized image. In RPD, we design a multi-scale information perception mechanism with our proposed novel information adjustment function, to encourage the discriminator to better perceive visual differences between the real and synthesized image. Consequently, the discriminator will drive the generator to improve the visual quality of the synthesized image. Besides, based on these novel designs, we can design two versions, a single-stage generation framework (AMI-GAN-S), and a multi-stage generation framework (AMI-GAN-M), respectively. The former can synthesize high-resolution images because of its low computational cost; the latter can synthesize images with realistic detail. Experimental results on two widely used T2I datasets showed that our AMI-GANs achieve competitive performance in T2I task. Hongchen Tan, Kaiqiang Xu, Huasheng Wang, Xiuping Liu, Xin Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Blind Image Quality Assessment via Adaptive Graph AttentionabstractRecent advancements in blind image quality assessment (BIQA) are primarily propelled by deep learning technologies. While leveraging transformers can effectively capture long-range dependencies and contextual details in images, the significance of local information in image quality assessment can be undervalued. To address this challenging problem, we propose a novel feature enhancement framework tailored for BIQA. Specifically, we devise an Adaptive Graph Attention (AGA) module to simultaneously augment both local and contextual information. It not only refines the post-transformer features into an adaptive graph, facilitating local information enhancement, but also exploits interactions amongst diverse feature channels. The proposed technique can better reduce redundant information introduced during feature updates compared to traditional convolution layers, streamlining the self-updating process for feature maps. Experimental results show that our proposed model outperforms state-of-the-art BIQA models in predicting the perceived quality of images. The code of the model will be made publicly available. Huasheng Wang, Hongchen Tan, Jianxun Lou, Xiaochang Liu, Wei Zhou 0021, Hantao Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | SSPNet: Predicting Visual Saliency ShiftsabstractWhen images undergo quality degradation caused by editing, compression or transmission, their saliency tends to shift away from its original position. Saliency shifts indicate visual behaviour change and therefore contain vital information regarding perception of visual content and its distortions. Given a pristine image and its distorted format, we want to be able to detect saliency shifts induced by distortions. The resulting saliency shift map (SSM) can be used to identify the region and degree of visual distraction caused by distortions, and consequently to perceptually optimise image coding or enhancement algorithms. To this end, we first create a largest-of-its-kind eye-tracking database, comprising 60 pristine images and their associated 540 distorted formats viewed by 96 subjects. We then propose a computational model to predict the saliency shift map (SSM), utilising transformers and convolutional neural networks. Experimental results demonstrate that the proposed model is highly effective in detecting distortion-induced saliency shifts in natural images. Huasheng Wang, Jianxun Lou, Xiaochang Liu, Hongchen Tan, Roger M. Whitaker, Hantao Liu |
IEEE Trans. Multim. | 1 |
| 2024 | Global and Local Interactive Perception Network for Referring Image SegmentationabstractThe effective modal fusion and perception between the language and the image are necessary for inferring the reference instance in the referring image segmentation (RIS) task. In this article, we propose a novel RIS network, the global and local interactive perception network (GLIPN), to enhance the quality of modal fusion between the language and the image from the local and global perspectives. The core of GLIPN is the global and local interactive perception (GLIP) scheme. Specifically, the GLIP scheme contains the local perception module (LPM) and the global perception module (GPM). The LPM is designed to enhance the local modal fusion by the correspondence between word and image local semantics. The GPM is designed to inject the global structured semantics of images into the modal fusion process, which can better guide the word embedding to perceive the whole image's global structure. Combined with the local-global context semantics fusion, extensive experiments on several benchmark datasets demonstrate the advantage of the proposed GLIPN over most state-of-the-art approaches. Jing Liu 0059, Hongchen Tan, Yongli Hu, Huasheng Wang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Deep Ordinal Regression Framework for No-Reference Image Quality AssessmentabstractDue to the rapid development of deep learning techniques, no-reference image quality assessment (NR-IQA) has achieved significant improvement. NR-IQA aims to predict a real-valued variable for image quality, using the image in question as the sole input. Existing deep learning-based NR-IQA models are formulated as a regression problem and trained by minimising the mean squared error. The error measurement does not consider the relative ordering between different ratings on the quality scale, which consequently affects the efficacy of the model. To account for this problem, we reformulate NR-IQA learning as an ordinal regression problem and propose a simple yet effective framework using deep convolutional neural networks (DCNN) and Transformers. NR-IQA learning is achieved by a deep ordinal loss and using a soft ordinal inference to transform the predicted probabilities to a continuous variable for image quality. Experimental results demonstrate the superiority of our proposed NR-IQA model based on deep ordinal regression. In addition, this framework can be easily extended with various DCNN architectures to build advanced IQA models. Huasheng Wang, Yulin Tu, Xiaochang Liu, Hongchen Tan, Hantao Liu |
IEEE Signal Process. Lett. | 1 |
| 2022 | Incomplete Descriptor Mining With Elastic Loss for Person Re-IdentificationabstractIn this paper, we propose a novel person Re-ID model, Consecutive Batch DropBlock Network (CBDB-Net), to capture the attentive and robust person descriptor for the person Re-ID task. The CBDB-Net contains two novel designs: the Consecutive Batch DropBlock Module (CBDBM) and the Elastic Loss (EL). In the Consecutive Batch DropBlock Module (CBDBM), we firstly conduct uniform partition on the feature maps. And then, we independently and continuously drop each patch from top to bottom on the feature maps, which can output multiple incomplete feature maps. In the training stage, these multiple incomplete features can better encourage the Re-ID model to capture the robust person descriptor for the Re-ID task. In the Elastic Loss (EL), we design a novel weight control item to help the Re-ID model adaptively balance hard sample pairs and easy sample pairs in the whole training process. Through an extensive set of ablation studies, we verify that the Consecutive Batch DropBlock Module (CBDBM) and the Elastic Loss (EL) each contribute to the performance boosts of CBDB-Net. We demonstrate that our CBDB-Net can achieve the competitive performance on the three standard person Re-ID datasets (the Market-1501, the DukeMTMC-Re-ID, and the CUHK03 dataset), three occluded Person Re-ID datasets (the Occluded DukeMTMC, the Partial-REID, and the Partial iLIDS dataset), and a general image retrieval dataset (In-Shop Clothes Retrieval dataset). Hongchen Tan, Xiuping Liu, Yuhao Bian, Huasheng Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |