Jianxun Lou

dblp:304/2331 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-2982-595XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KSIQA: A Knowledge-Sharing Model for No-Reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) aims to quantitatively measure human perception of visual quality without comparing a distorted image to a reference. Despite recent advances, existing NR-IQR approaches often demonstrate insufficient ability to capture perceptual cues in the absence of a reference, limiting their generalisability across diverse and complex real-world image degradations. These limitations hinder their ability to match the reliability of full-reference IQA (FR-IQA) counterparts. A key challenge, therefore, is to enable NR-IQA models to emulate the reference-aware reasoning exhibited by humans and FR-IQA methods. To address this challenge, we propose a novel NR-IQA model based on a knowledge-sharing (KS) strategy to simulate this capability and predict image quality more effectively. Specifically, we designate an FR-IQA model as the teacher and an NR-IQA model as the student. Unlike conventional knowledge distillation (KD), our proposed architecture enables the NR-IQA student and FR-IQA teacher to share a decoder rather than being independent models. Furthermore, the student model contains a Mental Imagery Generation (MIG) module to learn mental imagery as the reference. To fully exploit local and global information, we adopt a vision transformer (ViT) branch and a convolutional neural network branch for feature extraction (FE). Finally, a quality-aware regressor (QAR) combined with deep ordinal regression is constructed to infer the quality score. Experiments show that our proposed NR-IQA model, KSIQA, has class-leading performance against current no-reference (NR) techniques across widespread benchmark datasets.
Huasheng Wang, Hongchen Tan, Jianxun Lou, Xiaochang Liu, Wei Zhou 0021, Ying Chen 0011, Roger M. Whitaker, Walter Colombo, Hantao Liu
IEEE Trans. Neural Networks Learn. Syst.4
2025 MMP-2k: A Benchmark Multi-Labeled Macro Photography Image Quality Assessment Database
abstract
Macro photography (MP) is a specialized field of photography that captures objects at an extremely close range, revealing tiny details. Although an accurate macro photography image quality assessment (MPIQA) metric can benefit macro photograph capturing, which is vital in some domains such as scientific research and medical applications, the lack of MPIQA data limits the development of MPIQA metrics. To address this limitation, we conducted a large-scale MPIQA study. Specifically, to ensure diversity both in content and quality, we sampled 2,000 MP images from 15,700 MP images, collected from three public image websites. For each MP image, 17 (out of 21 after outlier removal) quality ratings and a detailed quality report of distortion magnitudes, types, and positions are gathered by a lab study. The images, quality ratings, and quality reports form our novel multi-labeled MPIQA database, MMP-2k. Experimental results showed that the state-of-the-art generic IQA metrics underperform on MP images. The database and supplementary materials are available at https://github.com/Future-IQA/MMP-2k.
Jiashuo Chang, Jianxun Lou, Zhen Qiu 0001, Hanhe Lin
ICIP3
2025 AL-SSFGait: A Dual-Branch Framework with Aligned Silhouette-Skeleton for Gait Recognition
abstract
Gait recognition has emerged as a promising biometric technology in computer vision. While state-of-the-art methods achieve high accuracy on laboratory datasets, their performance deteriorates significantly in real-world scenarios. This performance drop is primarily attributed to spatio-temporal distribution inconsistencies caused by varying viewpoints, occlusions, and background clutter. Combining silhouette and skeleton modalities has the potential to enhance the robustness of gait representation in such challenging scenarios. However, the effectiveness of multimodal approaches is hindered by the significant gap between silhouette and skeleton data. To address these issues, we propose the Aligned Silhouette and Skeleton Feature Fusion for Gait Recognition (AL-SSFGait) framework, a dual-branch framework designed to enable robust cross-modal representation. First, an affine alignment strategy is applied to normalize the spatial distributions of silhouette and skeleton inputs. Next, a Skeleton Map Generation method is constructed by converting discrete joint coordinates into continuous heatmaps, thereby reducing the semantic gap between the two modalities. Finally, we introduce SSFGait, a co-attention-based dual-branch network that performs dynamic spatio-temporal feature fusion across both modalities. Experimental results demonstrate that AL-SSFGait achieves 91.0% accuracy on the CASIA-B dataset and a Rank-1 accuracy of 72.3% on the GREW dataset. Ablation studies further validate the effectiveness of each proposed component.
Jian-lou Lou, Guiping Zhang, Jianxun Lou
MMAsia3
2025 Image Manipulation Quality Assessment
abstract
Image quality assessment (IQA) and its computational models play a vital role in modern computer vision applications. Research has traditionally focused on signal distortions arising during image compression and transmission, and their impact on perceived image quality. However, little attention is paid to image manipulation that alters an image using various filters. With the prevalence of image manipulation in real-life scenarios, it is critical to understand how humans perceive filter-altered images and to develop reliable IQA models capable of automatically assessing the quality of filtered images. In this paper, we build a new IQA database for filter-altered images, comprised of 360 images manipulated by various filters. To ensure the subjective IQA faithfully reflects human visual perception, we conduct a fully-controlled psychovisual experiment. Building upon the ground truth, we propose an innovative deep learning-based no-reference IQA (NR-IQA) model named IMQA that can accurately predict the perceived quality of filter-altered images. This model involves constructing an image filtering-aware module to learn discriminatory features for filter-altered images; and fuses these features with the representations generated by an image quality-aware module. Experimental results demonstrate the superior performance of the proposed IMQA model.
Xinbo Wu, Jianxun Lou, Wan'an Liu, Paul L. Rosin, Gualtiero Colombo 0001, Stuart M. Allen, Roger M. Whitaker, Hantao Liu
IEEE Trans. Circuits Syst. Video Technol.2
2025 Distortion-Induced Saliency Shifts in Video
abstract
Visual saliency modelling is of fundamental importance in modern video processing and its applications. Our previous eye-tracking study revealed that signal distortions caused by editing, compression, or transmission alter gaze patterns and consequently induce saliency shifts in both spatial and temporal domains. Saliency shifts provide crucial insights into viewers’ behavioural responses to video distortions, facilitating the perception-based optimisation of video algorithms. However, the spatio-temporal saliency shifts and their measurable effects on perception related applications remain largely unexplored. In this paper, we first investigate the measurement of distortion-induced saliency shifts (DSS) in videos and analyse DSS behaviours as functions of video content, time order and critical distortion disruption. Second, based on our findings, we construct three vision models to quantitatively simulate distinct DSS behaviours and integrate them into a comprehensive DSS behaviour model. Finally, we demonstrate that the computational DSS model can enhance emerging video technologies.
Xinbo Wu, Jianxun Lou, Zhengyan Dong, Fan Zhang 0017, Paul L. Rosin, Hantao Liu
IEEE Trans. Multim.2
2025 Chest X-Ray Visual Saliency Modeling: Eye-Tracking Dataset and Saliency Prediction Model
abstract
Radiologists' eye movements during medical image interpretation reflect their perceptual-cognitive processes of diagnostic decisions. The eye movement data can be modeled to represent clinically relevant regions in a medical image and potentially integrated into an artificial intelligence (AI) system for automatic diagnosis in medical imaging. In this article, we first conduct a large-scale eye-tracking study involving 13 radiologists interpreting 191 chest X-ray (CXR) images, establishing a best-of-its-kind CXR visual saliency benchmark. We then perform analysis to quantify the reliability and clinical relevance of saliency maps (SMs) generated for CXR images. We develop CXR image saliency prediction method (CXRSalNet), a novel saliency prediction model that leverages radiologists' gaze information to optimize the use of unlabeled CXR images, enhancing training and mitigating data scarcity. We also demonstrate the application of our CXR saliency model in enhancing the performance of AI-powered diagnostic imaging systems.
Jianxun Lou, Huasheng Wang, Xinbo Wu, John Cho Hui Ng, Kaveri A. Thakoor, Padraig Corcoran, Ying Chen 0011, Hantao Liu
IEEE Trans. Neural Networks Learn. Syst.1
2024 Time-Interval Visual Saliency Prediction in Mammogram Reading
abstract
Radiologists’ eye movements during medical image interpretation reflect their perceptual-cognitive behaviour and correlate with diagnostic decisions. Previous study has shown the significance of gaze behaviour of different time intervals for the decision-making process. Being able to automatically predict the visual attention of radiologists for different reading phases would enhance the reliability and explainability of artificial intelligence (AI) in diagnostic imaging. In this paper, we investigate the time-interval visual saliency in mammogram reading. We propose a novel visual saliency prediction model based on deep learning, which predicts a sequence of time-interval saliency maps for an input mammogram. Experimental results demonstrate the efficacy of the proposed time-interval saliency model.
Jianxun Lou, Xinbo Wu, Hantao Liu
ICASSP1
2024 A Benchmark of Variance of Opinion Scores in Image Quality Assessment
abstract
Mean opinion score (MOS) has been used as the benchmark to measure the perceived quality of digital images. However, the usefulness of MOS diminishes when a substantial variation between individual opinions occurs. It is critical to measure the stimulus-driven variance of opinion scores (VOS) and scrutinise images that evoke a large VOS, and consequently, use VOS to inform our interpretation of MOS. In this paper, we create a VOS benchmark for individual differences in image quality assessment and analyse the importance of VOS classification as a function of distortion intensity, distortion type and scene content. In addition, a simple yet effective deep learning-based model is built, aiming to identify images with a large variation in viewers’ quality judgements.
Jianxun Lou, Xinbo Wu, Padraig Corcoran, Gualtiero Colombo 0001, Roger M. Whitaker, Hantao Liu
ICIP1
2024 TranSalNet+: Distortion-aware saliency prediction
abstract
Predicting the saliency of images affected by distortion is a challenging but emerging research problem. Given a distorted image, we wish to accurately predict saliency as perceived by humans. A recent distortion-aware saliency benchmark – the CUDAS database – reveals the inadequacy of existing saliency models in handling distorted images. In this paper, we devise a deep learning Distortion-Aware Saliency Module (DASM) that enables capturing saliency features related to image distortions, and integrates this module into a saliency prediction architecture. To achieve the high expressive capability of DASM using supervised learning, we create a dedicated dataset that draws upon a large-scale saliency dataset and machine-generated image quality assessments . Experimental results demonstrate the superior performance of the proposed model in predicting the saliency of distorted images.
Jianxun Lou, Xinbo Wu, Padraig Corcoran, Paul L. Rosin, Hantao Liu
Neurocomputing1
2024 RAD-IQMRI: A benchmark for MRI image quality assessment
abstract
Magnetic resonance imaging (MRI) is susceptible to visual artifacts that can degrade the perceptual image quality, potentially leading to inaccurate or inefficient diagnoses in clinical practice. It is critical to evaluate the perceptual image quality and build this technique into clinical solutions. In a previous study, an MRI database was created for image quality assessment (IQA), where various types of MRI artifacts with different degrees of degradation were simulated. Application specialists assessed the image quality; however, radiologists’ perception of MRI image quality remains unknown. To make IQA clinically relevant, in this paper we conduct a new subjective experiment where 13 radiologists rated the quality of images contained in the MRI database. Based on this subjective IQA benchmark named RAD-IQMRI, we evaluate the performance of state-of-the-art objective IQA models, providing insights into their application for MRI image quality assessment in clinical settings.
Yueran Ma, Jianxun Lou, Jean-Yves Tanguy, Padraig Corcoran, Hantao Liu
Neurocomputing2
2024 Blind Image Quality Assessment via Adaptive Graph Attention
abstract
Recent advancements in blind image quality assessment (BIQA) are primarily propelled by deep learning technologies. While leveraging transformers can effectively capture long-range dependencies and contextual details in images, the significance of local information in image quality assessment can be undervalued. To address this challenging problem, we propose a novel feature enhancement framework tailored for BIQA. Specifically, we devise an Adaptive Graph Attention (AGA) module to simultaneously augment both local and contextual information. It not only refines the post-transformer features into an adaptive graph, facilitating local information enhancement, but also exploits interactions amongst diverse feature channels. The proposed technique can better reduce redundant information introduced during feature updates compared to traditional convolution layers, streamlining the self-updating process for feature maps. Experimental results show that our proposed model outperforms state-of-the-art BIQA models in predicting the perceived quality of images. The code of the model will be made publicly available.
Huasheng Wang, Hongchen Tan, Jianxun Lou, Xiaochang Liu, Wei Zhou 0021, Hantao Liu
IEEE Trans. Circuits Syst. Video Technol.4
2024 Predicting Radiologists' Gaze With Computational Saliency Models in Mammogram Reading
abstract
Previous studies have shown that there is a strong correlation between radiologists' diagnoses and their gaze when reading medical images. The extent to which gaze is attracted by content in a visual scene can be characterised as visual saliency. There is a potential for the use of visual saliency in computer-aided diagnosis in radiology. However, little is known about what methods are effective for diagnostic images, and how these methods could be adapted to address specific applications in diagnostic imaging. In this study, we investigate 20 state-of-the-art saliency models including 10 traditional models and 10 deep learning-based models in predicting radiologists' visual attention while reading 196 mammograms. We found that deep learning-based models represent the most effective type of methods for predicting radiologists' gaze in mammogram reading; and that the performance of these saliency models can be significantly improved by transfer learning. In particular, an enhanced model can be achieved by pre-training the model on a large-scale natural image saliency dataset and then fine-tuning it on the target medical image dataset. In addition, based on a systematic selection of backbone networks and network architectures, we proposed a parallel multi-stream encoded model which outperforms the state-of-the-art approaches for predicting saliency of mammograms.
Jianxun Lou, Hanhe Lin, Philippa Young, Zelei Yang, Susan Cheng Shelmerdine, David Marshall 0001, Emiliano Spezi, Marco Palombo, Hantao Liu
IEEE Trans. Multim.1
2024 SSPNet: Predicting Visual Saliency Shifts
abstract
When images undergo quality degradation caused by editing, compression or transmission, their saliency tends to shift away from its original position. Saliency shifts indicate visual behaviour change and therefore contain vital information regarding perception of visual content and its distortions. Given a pristine image and its distorted format, we want to be able to detect saliency shifts induced by distortions. The resulting saliency shift map (SSM) can be used to identify the region and degree of visual distraction caused by distortions, and consequently to perceptually optimise image coding or enhancement algorithms. To this end, we first create a largest-of-its-kind eye-tracking database, comprising 60 pristine images and their associated 540 distorted formats viewed by 96 subjects. We then propose a computational model to predict the saliency shift map (SSM), utilising transformers and convolutional neural networks. Experimental results demonstrate that the proposed model is highly effective in detecting distortion-induced saliency shifts in natural images.
Huasheng Wang, Jianxun Lou, Xiaochang Liu, Hongchen Tan, Roger M. Whitaker, Hantao Liu
IEEE Trans. Multim.2
2022 Predicting Radiologist Attention During Mammogram Reading with Deep and Shallow High-Resolution Encoding
abstract
Radiologists’ eye-movement during diagnostic image reading reflects their personal training and experience, which means that their diagnostic decisions are related to their perceptual processes. For training, monitoring, and performance evaluation of radiologists, it would be beneficial to be able to automatically predict the spatial distribution of the radiologist’s visual attention on the diagnostic images. The measurement of visual saliency is a well-studied area that allows for prediction of a person’s gaze attention. However, compared with the extensively studied natural image visual saliency (in free viewing tasks), the saliency for diagnostic images is less studied; there could be fundamental differences in eye-movement behaviours between these two domains. Most current saliency prediction models have been optimally developed for natural images, which could lead them to be less adept at predicting the visual attention of radiologists during the diagnosis. In this paper, we propose a method specifically for automatically capturing the visual attention of radiologists during mammogram reading. By adopting high-resolution image representations from both deep and shallow encoders, the proposed method avoids potential detail losses and achieves superior results on multiple evaluation metrics in a large mammogram eye-movement dataset.
Jianxun Lou, Hanhe Lin, David Marshall 0001, Young Yang, Susan Cheng Shelmerdine, Hantao Liu
ICIP1
2022 TranSalNet: Towards perceptually relevant visual saliency prediction
abstract
Convolutional neural networks (CNNs) have significantly advanced computational modelling for saliency prediction. However, accurately simulating the mechanisms of visual attention in the human cortex remains an academic challenge. It is critical to integrate properties of human vision into the design of CNN architectures, leading to perceptually more relevant saliency prediction. Due to the inherent inductive biases of CNN architectures, there is a lack of sufficient long-range contextual encoding capacity. This hinders CNN-based saliency models from capturing properties that emulate viewing behaviour of humans. Transformers have shown great potential in encoding long-range information by leveraging the self-attention mechanism. In this paper, we propose a novel saliency model that integrates transformer components to CNNs to capture the long-range contextual visual information. Experimental results show that the transformers provide added value to saliency prediction, enhancing its perceptual relevance in the performance. Our proposed saliency model using transformers has achieved superior results on public benchmarks and competitions for saliency prediction models. The source code of our proposed saliency model TranSalNet is available at: https://github.com/LJOVO/TranSalNet.
Jianxun Lou, Hanhe Lin, David Marshall 0001, Dietmar Saupe, Hantao Liu
Neurocomputing1
2021 Study of Saccadic Eye Movements in Diagnostic Imaging
abstract
Eye movements reflect the visual process of humans’ perception and cognition. In the field of medical imaging, the diagnosis rendered by radiologists is closely related to their eye movements when reading radiological images. It is beneficial to study the eye movements of radiologists to improve the diagnostic performance. However, existing studies are mainly focused on the radiologists’ fixations but rarely on their saccade patterns. Moreover, these studies are almost based on limited datasets. In this paper, we present a quantitative study of the gaze behavior of radiologists from the perspective of saccade patterns on a large-scale dataset. The dataset comprises of the eye-tracking data of 10 expert radiologists reading 196 mammograms. By analyzing the saccade amplitude, direction, and bias of radiologists, we found that radiologists have specific saccade patterns in image reading and the saccade patterns are significantly affected by the different reading phases, working experience, and orientations of the mammograms.
Jianxun Lou, Philippa Young, Hantao Liu
ICIP1