VLDB 2026 Research / reviewers in the wild / expert
Guanghui Yue 0001
dblp:179/3244
· DBLP profile ↗
118ranked-venue papers
34as first author
99since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 74 · 24 first-author · 59 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 9 first-author · 18 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A hybrid Mamba-Transformer model with progressive feature enhancement for medical image segmentation
Huanhuan Lv, Wanqi Ma, Guanghui Yue 0001, Songru Jiang |
Comput. Vis. Image Underst. | 3 |
| 2026 | Explainable AI in Medicine: A Comprehensive Narrative Review of Methods, Applications, and Future DirectionsabstractABSTRACT Artificial Intelligence (AI) is increasingly utilised in medicine; however, its “black‐box” nature continues to hinder clinical trust, adoption, and validation. Explainable AI (XAI) has emerged as a critical field to address these transparency challenges by making AI‐driven decisions more interpretable and actionable. This narrative review examines the progress of XAI in medicine over the past decade. We first introduce fundamental XAI concepts and describe our review methodology, followed by a comprehensive analysis of key application domains, including medical imaging, electronic health records (EHRs), and multi‐omics data. Methodologically, we categorise XAI techniques into model‐agnostic approaches (e.g., SHAP, LIME, Anchors) and model‐specific approaches (e.g., Grad‐CAM, LRP, TreeSHAP). Beyond summarising their principles, advantages, and limitations, we further provide a systematic analysis of clinical reliability and failure modes associated with each class of methods, highlighting how explanation techniques may produce misleading, unstable, or non‐causal interpretations in real‐world clinical settings. The review then discusses the demonstrated benefits of XAI, including result validation, bias detection, and improved patient–clinician communication, while critically examining persistent challenges such as limited clinical deployment, inconsistent evaluation standards, and the lack of prospective validation. Finally, we outline future research directions, emphasising the need to adapt XAI to large‐scale foundation models and conversational AI systems, as well as to extend its applicability in biomedical and multi‐omics interpretation. We argue that while XAI is essential for improving transparency, trust, and clinical adoption, its reliable and scalable integration into clinical workflows requires both methodological innovation and rigorous, clinically grounded validation frameworks. Helin Wang, Xueyu Liu, Jiashuo Shi, Yu-ang Li, Guanghui Yue 0001, Yongfei Wu |
Expert Syst. J. Knowl. Eng. | 8 |
| 2026 | TriS -Net: A Progressive Learning Framework for Medical Image Segmentation With Multigranularity SupervisionabstractABSTRACT Accurate medical image segmentation is essential for disease diagnosis, treatment planning and outcome monitoring. However, current segmentation methods heavily rely on large‐scale, pixel‐level annotations, which are costly and labour‐intensive to obtain. To address this challenge, we propose TriS‐Net (Triple‐Supervision Segmentation Network) , a progressive framework that integrates image‐level, bounding box‐level and pixel‐level labels into a multigranularity supervision pipeline under limited annotation settings. In the first stage, TriS‐Net uses image‐level labels to train a classification branch, enabling the network to learn discriminative features and localise potential lesion regions. In the second stage, a box‐guided mask refinement strategy (BMR) is proposed, which combines Soft‐NMS filtering and a one‐to‐one matching mechanism to obtain reliable candidate regions. CIoU is further employed to derive image‐level quality metrics that impose quality‐aware weighted constraints on segmentation learning, thereby improving spatial localisation and structural consistency. In the third stage, a small number of pixel‐level labels are used for fine‐grained supervision, further enhancing segmentation accuracy and boundary details. The proposed method is validated on the BraTS 2019 and LiTS 2017 datasets, on which it outperforms several existing methods under limited annotation settings. Additional experiments on the BUSI ultrasound dataset further demonstrate its good generalisation capability across different imaging modalities. Xueyu Liu, Junxin Chen 0001, Guanghui Yue 0001, Yongfei Wu |
Expert Syst. J. Knowl. Eng. | 6 |
| 2026 | BUT-Net: Boundary-Aware U-Net structure with Two-Path Transformers for lesion segmentation in mCNV using OCT images
Hai Xie, Zhenquan Wu, Shaobin Chen, Guanghui Yue 0001, Tianfu Wang 0001, Bai Ying Lei |
Expert Syst. Appl. | 4 |
| 2026 | BeatDance: Generating beat-consistent 3D dance with hierarchical spatial-temporal modeling
Xiaojian Shen, Dahu Shi, Jianrong Zhang, Yunzhi Zhuge, Zhiliang Wu, Guanghui Yue 0001, Wei Zhou 0021 |
Pattern Recognit. | 9 |
| 2026 | Hybrid explicit and implicit encoding for multi-view representation learning
Shuochen Yao, Yusheng Zhang, Weiqing Yan, Chang Tang, Guanghui Yue 0001, Kaile Su |
Pattern Recognit. | 5 |
| 2026 | Expressive Human Volumetric Video Generation With Rich TextabstractPlain text has become the dominant interactive interface for text-driven human volumetric video generation. However, its limited customization options hinder users from expressing motion effects with accuracy. For example, plain text struggles to specify continuous variables such as motion amplitude, speed, and joint trajectories with precision, and it fails to convey stylized motion characteristics. Additionally, crafting detailed textual prompts for complex motion sequences is cumbersome, while excessively long prompts strain text encoders. To address these limitations, we propose a rich text-based framework that supports font styles, sizes, and trajectory sketching. By extracting motion-related attributes from rich text, our method enables fine-grained control over motion styles, precise speed regulation, and accurate joint trajectory manipulation. These capabilities are realized through gradient-guided noise editing and ControlNet-based motion optimization, which operate within the latent motion diffusion process. Specifically, we design a unified gradient-guided adaptation mechanism to ensure that the generated motion video adheres strictly to the specified constraints. Furthermore, we introduce realism-oriented optimization for stylistic and joint-level control, refining motion synthesis at a granular level to produce smoother, more natural movements. We present multiple comparative evaluations showcasing volumetric video generation from both rich text and plain text. Through quantitative analysis, we demonstrate that our method surpasses strong plain-text baselines, producing expressive, customizable human volumetric motion videos. Guanghui Yue 0001, Wei Zhou 0021, Xudong Mao, Ruomei Wang 0001, Baoquan Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Perception-Inspired Network for Stereo Image Quality AssessmentabstractExisting stereo image quality assessment (SIQA) methods generally have limitations in binocular fusion and fine-grained perception modeling. To address these issues, we propose a Perception-Inspired Network for SIQA that simulates binocular difference-guided fusion, high-frequency sensitivity, and hierarchical perception mechanisms of the human visual system (HVS). First, a difference-guided binocular fusion (DGBF) module is designed to mimic the binocular difference sensitivity mechanism, which exploits difference information at both the feature-level and image-level to optimize binocular fusion. Furthermore, the image distortion primarily affects the high-frequency components, which are critical for perceptual quality. To reflect this, we propose a high-frequency enhancement module (HFEM) to simulate the human eye's sensitivity to edge and texture distortions. Finally, to better achieve fine-grained perception modeling, we propose a hierarchical quality regression strategy that simulates the human perceptual process, from perceiving local details to forming a global quality judgment, thereby achieving a quality prediction more aligned with human subjective evaluation. Experimental results demonstrate that the proposed method outperforms mainstream approaches, achieving a PLCC of 0.9734 on the LIVE I database, and a PLCC of 0.9632 on the LIVE II database. Yongli Chang, Guanghui Yue 0001, Li Yu 0004, Yakun Ju, Hadi Amirpour, Moncef Gabbouj, Wei Zhou 0021 |
IEEE Trans. Image Process. | 2 |
| 2026 | Self-Supervised Unfolding Network With Shared Reflectance Learning for Low-Light Image EnhancementabstractRecently, incorporating Retinex theory with unfolding networks has attracted increasing attention in the low-light image enhancement field. However, existing methods have two limitations, i.e., ignoring the modeling of the physical prior of Retinex theory and relying on a large amount of paired data. To advance this field, we propose a novel self-supervised unfolding network, named S2UNet, for the LIE task. Specifically, we formulate a novel optimization model based on the principle that content-consistent images under different illumination should share the same reflectance. The model simultaneously decomposes two illumination-different images into a shared reflectance component and two independent illumination components. Due to the absence of the normal-light image, we process the low-light image with gamma correction to create the illumination-different image pair. Then, we translate this model into a multi-stage unfolding network, in which each stage alternately optimizes the shared reflectance component and the respective illumination components of the two images. During progressive multi-stage optimization, the network inherently encodes the reflectance consistency prior by jointly estimating an optimal reflectance across varying illumination conditions. Finally, considering the presence of noise in low-light images and to suppress noise amplification, we propose a self-supervised denoising mechanism. Extensive experiments on nine benchmark datasets demonstrate that our proposed S2UNet outperforms state-of-the-art unsupervised methods in terms of both quantitative metrics and visual quality, while achieving competitive performance compared to supervised methods. The source code will be available at https://github.com/J-Liu-DL/S2UNet. Jia Liu 0025, Yu Luo 0004, Guanghui Yue 0001, Jie Ling 0002, Chia-Wen Lin, Guangtao Zhai, Wei Zhou 0021 |
IEEE Trans. Image Process. | 3 |
| 2026 | Self-Anchored Progressive Framework With Noise Mitigation for Unsupervised Camouflaged Object DetectionabstractUnsupervised Camouflaged Object Detection (UCOD) presents a significant challenge due to the inherent similarity between camouflaged objects and their backgrounds, compounded by the absence of manual annotations. Although pixel-level pseudo-labeling has proven effective for unsupervised salient object detection (USOD), it is far less reliable for COD, where the concealed and ambiguous nature of camouflaged objects frequently produces noisy pseudo-labels, causing misjudgments, missed detections, and imprecise boundaries. To overcome this, we propose SAPNet, a novel self-anchored progressive framework for UCOD. Rather than depending on noisy pixel-level supervision, we leverage semantically reliable foreground and background regions as high-confidence anchors. This effectively transforms the unsupervised problem into a more robust weakly supervised paradigm, reducing learning difficulty and mitigating overfitting to noise. SAPNet learns camouflaged objects progressively by first emphasizing these confident regions and then exploiting DINO's contextual awareness to recover complete structures. Central to our framework is the semantic-driven region detector (SDRD), which employs cascaded convolutions and a residual attention projection mechanism to suppress background noise, filter erroneous information, and enhance spatial context, ensuring reliable supervision signals. Furthermore, a region-based context inference module (RCIM) is introduced to iteratively refine object boundaries by integrating multi-level semantic features under the guidance of these refined region-level anchors. Extensive experiments on four benchmark COD datasets demonstrate that SAPNet significantly outperforms state-of-the-art unsupervised methods. The source code of our SAPNet is available at https://github.com/ArloJie/SAPNet. Binwei Xu, Tuo Shen, Guanghui Yue 0001, Qiuping Jiang |
IEEE Trans. Image Process. | 4 |
| 2026 | SGNet: Style-Guided Network With Temporal Compensation for Unpaired Low-Light Colonoscopy Video EnhancementabstractA low-light colonoscopy video enhancement method is needed as poor illumination in colonoscopy can hinder accurate disease diagnosis and adversely affect surgical procedures. Existing low-light video enhancement methods usually apply a frame-by-frame enhancement strategy without considering the temporal correlation between them, which often causes a flickering problem. In addition, most methods are designed for endoscopic devices with fixed imaging styles and cannot be easily adapted to different devices. In this paper, we propose a Style-Guided Network (SGNet) for unpaired Low-Light Colonoscopy Video Enhancement (LLCVE). Given that collecting content-consistent paired videos is difficult, SGNet adopts a CycleGAN-based framework to convert low-light videos to normal-light videos, in which a Temporal Compensation (TC) module and a Style Guidance (SG) module are proposed to alleviate the flickering problem and achieve flexible style transfer, respectively. The TC module compensates for a low-light frame by learning the correlated feature of its adjacent frames, thereby improving the temporal smoothness of the enhanced video. The SG module encodes the text of the imaging style and adaptively explores its intrinsic relationships with video features to obtain style representations, which are then used to guide the subsequent enhancement process. Extensive experiments on a curated database show that SGNet achieves promising performance on the LLCVE task, outperforming state-of-the-art methods in both quantitative metrics and visual quality. Guanghui Yue 0001, Wanqing Liu, Jingfeng Du, Tianwei Zhou, Hanhe Lin, Qiuping Jiang, Wenqi Ren |
IEEE Trans. Image Process. | 1 |
| 2026 | Uncertainty-Guided Spatiotemporal Consistency Fusion Network for Infrared-Visible Video Fusion Under Extremely Low-Light ConditionsabstractInfrared-visible video fusion under extremely low-light conditions is critically important yet remains underexplored, largely due to the scarcity of high-quality datasets and challenges posed by spatiotemporal uncertainty and modality bias. To address the dataset shortage, we built a dataset of 4,739 infrared and visible registration video pairs captured under extremely low-light conditions, spanning 5 scene types and 17 subcategories. Further, we proposed an Uncertainty-guided Spatiotemporal Consistency Fusion Network, termed USCFNet, for the infrared-visible video fusion. At each layer of the encoder, an Entropy-Gated SpatioTemporal Attention (EGSTA) module is introduced to capture temporal instability and spatial reliability variations through entropy-aware attention modulation, thereby enhancing feature spatiotemporal consistency. The refined infrared and visible features are then fused via a Difference-Guided Fusion (DGF) module, which adaptively exploits their content and edge differences to improve structural integrity and detail clarity. By progressively connecting DGF modules from shallow to deep layers, the network achieves the synergistic fusion of shallow textures and deep semantics. Subsequently, the output of the last DGF module is fused with the modality features of the last layer through a hierarchical mixture-of-experts fusion module. This module enables the balanced integration of modality information while preserving fine local details. Finally, the fusion feature is fed into the decoder to produce the final fused video. Extensive experiments on our dataset and two public datasets show that USCFNet outperforms competing methods, achieving lower distortion and stronger spatiotemporal consistency. The source code and dataset are available at https://github.com/Zhaocheng1/ELVID. Cheng Zhao 0003, Tianyun Song, Zhiliang Wu, Tianfu Wang 0001, Moncef Gabbouj, Guanghui Yue 0001, Bai Ying Lei, Wei Zhou 0021 |
IEEE Trans. Image Process. | 6 |
| 2026 | Contrastive Representation Learning for Cross-Domain Blood Cell Image Classification With Denoising MechanismabstractAccurate identification and classification of white blood cells are essential for diagnosing hematological malignancies and analyzing blood disorders. Existing approaches predominantly leverage masked autoencoders (MAEs) to extract intrinsic blood cell features through image reconstruction as a pretext task. However, these methods encounter two critical challenges: (1) their generalization performance deteriorates under domain shifts caused by variations in staining techniques, illumination conditions, and microscope settings, and (2) the learned data distribution often deviates from the true distribution of blood cell features. To overcome these limitations, we propose CD-CBC, a novel framework for cross-domain blood cell image classification that integrates contrastive representation learning with a denoising mechanism. CD-CBC consists of two key components: a LoRA-based segmentation anything model (LoRA-SAM) and a contrastive masked autoencoder (CMAE). LoRA-SAM mitigates shortcut learning in contrastive learning by eliminating background noise and platelet interference, while CMAE captures fine-grained semantic features and models spatial relationships, enhancing cross-domain robustness. Additionally, we introduce a denoising mechanism in the latent space, which guides the model to focus on unmasked patches during reconstruction, allowing it to better capture the true distribution of blood cell features. Extensive experiments on two benchmark blood cell datasets demonstrate that CD-CBC achieves superior cross-domain performance, reaching an average accuracy of 62.47%, which is 3.17% higher than the current state-of-the-art, thereby confirming its strong generalization capability. Renyu Fu, Chengfu Ji, Sen Xiang, Guanghui Yue 0001, Tianyi Wang 0006, Chang Tang |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | TFFN: Three-Branch Feature Fusion Network for Stereoscopic Omnidirectional Image Quality AssessmentabstractStereoscopic omnidirectional image (SOI) has both omnidirectional and stereoscopic perception features. Many previous models have proved the viewport characteristics and stereoscopic visual features are crucial for quality perception of SOI. However, effective monocular and binocular visual features extraction and fusion are difficult due to the size of SOI and inaccuracy of feature representation. In this paper, we proposed a three-branch feature fusion network (TFFN) by fusing two-stream binocular visual features and the important monocular features based on the viewport perspective. The hierarchical fusion module is first designed to fuse effective binocular visual features from different semantic scales, and the pseudo-difference information extraction module is built to obtain the accuracy monocular visual features to complement the binocular visual features. Finally, the above monocular and binocular visual features are fused together to measure the quality of SOI. The comparison experiments are conducted on three public datasets and the analysis of the results demonstrate the effectiveness of the proposed method. Yun Liu 0009, Daoxin Fan, Huiyu Duan, Peiguang Jing, Guanghui Yue 0001, Guangtao Zhai |
IEEE Trans. Multim. | 6 |
| 2025 | DGLL: A Hybrid Global-Local Feature Learning Network for Precise Tooth Landmark DetectionabstractThe precise identification of key landmarks on three-dimensional tooth mesh models is paramount for computer-aided orthodontic treatment. However, existing methodologies exhibit limitations with respect to the integration of global and local features, which undermines accuracy in complex scenarios and excessively emphasizes relative landmark positions, resulting in displacement errors. To mitigate these issues, this study introduces DGLL, a hybrid feature learning network characterized by a dual-branch architecture that amalgamates global and local features. DGLL integrates a Cascaded Topological Relation Module (CTRM) to stabilize the extraction of global features and a Pan-scale Feature Modulation Module (PFMM) to balance relative and absolute positional accuracy. Empirical evaluations across various tooth types demonstrate that DGLL consistently enhances the accuracy of landmark localization. This research provides an effective approach to the automated analysis of tooth data, thereby improving the precision and efficacy of orthodontic treatment. Jianwen Huang, Guoheng Huang, Fuchen Zheng, Chi-Man Pun, Ka-Cheng Choi, Lianglun Cheng, Guanghui Yue 0001 |
BIBM | 7 |
| 2025 | GameMLD: A Game-Sourced Motion-Language Dataset for Stylized Motion GenerationabstractText-guided character animation generation has emerged as a significant research area with broad applications in gaming, film, interactive media, and beyond. However, existing motion-language datasets face limitations in motion quality, stylistic diversity, and annotation depth, particularly for professional applications. In contrast to existing datasets based on motion capture or video reconstruction techniques, our dataset leverages professionally crafted game animations and employs a structured annotation framework that incorporates standardized game design terminology. The dataset contains 8,700 high-fidelity motion sequences paired with 26,100 multi-level textual descriptions, generated through our proposed annotation pipeline that combines domain expertise with large language models. Through comprehensive experiments and user studies, we demonstrate GameMLD’s advantages in motion quality, style expressiveness, and annotation quality. Additionally, we showcase its practical value by developing a text-driven character animation generation system that effectively supports game production pipelines. Our experiments with state-of-the-art motion synthesis models demonstrate significant improvements in both animation quality and style control. The GameMLD dataset and source code can be reached via this link. Yiyu Fu, Ziming Cheng, Yihao Liao, Jiangfeiyang Wang, Ruomei Wang 0001, Guanghui Yue 0001, Chenlei Lv, Baoquan Zhao |
ICME | 6 |
| 2025 | MSPoint-Gait: Multi-Scale Point Cloud Analysis for 3D Gait Recognition via Cross-Modal LearningabstractRecent advances in LiDAR technology have enabled privacy-preserving gait recognition using 3D point cloud data. However, existing approaches struggle with the inherent challenges of point cloud processing and understanding such as spatial sparsity, irregular sampling, and complex temporal dynamics. In this paper, we present MSPoint-Gait, a novel framework that addresses these challenges through multi-scale analysis and cross-modal learning. At the core of our framework lies a Depth-Aware Attention Module (DAAM) that leverages rich 3D geometric information to generate attention-weighted depth representations, enabling fine-grained feature extraction from point cloud sequences. We further introduce a Multi-Scale Spatio-Temporal (MSST) network that hierarchically captures both local and global gait patterns through adaptive convolution kernels across multiple spatial and temporal scales. These components are unified through a novel cross-modal learning strategy that effectively bridges the semantic gap between raw point clouds and structured depth representations. The proposed frame-work achieves state-of-the-art performance on the challenging SUSTech1K dataset, with 91.9% Rank-1 and 98.0% Rank-5 accuracy, demonstrating significant improvements over existing methods across various walking conditions and viewpoints. Xinzhu Li, Yikun Chen, Guanghui Yue 0001, Wei Zhou 0021, Ruomei Wang 0001, Xudong Mao, Juepeng Zheng, Fan Zhou 0001, Ziqi Qiu, Baoquan Zhao |
ICME | 4 |
| 2025 | Prior-Guided Test Time Adaptation for Blind Image Quality AssessmentabstractCurrent blind image quality assessment (BIQA) models usually lack adaptability to the test data with distribution shifts to the training data. This inspires an investigation into test time adaptation (TTA) methods to address distribution shifts between training and test data. However, existing methods mainly focus on simple feature alignment strategies, which may lead to incorrect knowledge generalization. To this issue, we propose a prior-guided test time adaptation (PGTA-IQA) for blind image quality assessment. Concretely, we extract the quality prior knowledge from the pre-trained BIQA model through clustering. The extracted quality prior knowledge forms the foundation for subsequent optimizations. These optimizations are carried out from two complementary perspectives: inter-cluster and intra-cluster. From the inter-cluster perspective, we propose a confident rank learning approach which consists of a relative quality matrix (RQM) and a confidence filtering strategy (CFS) to generate the high-confident quality rankings. From the intra-cluster perspective, we propose a selective feature alignment approach by only aligning the closest neighboring samples within the same cluster to reduce the impact of noisy labels. The experimental results demonstrate the effectiveness of the proposed approaches. Shishun Tian, Fangjie Hou, Guanghui Yue 0001, Yuanhao Gong, Wenbin Zou, Ting Su 0004 |
ICME | 3 |
| 2025 | MCSMoG: Multi-Conditional Diffusion for Stylized Motion Generation with Parametric ControlabstractStylized human motion synthesis remains a fundamental challenge in computer animation and graphics, with a wide spectrum of applications spanning gaming, film production, virtual reality, and beyond. While recent advances in text-driven motion generation have shown promise, existing approaches face critical limitations including the inability to maintain consistent trajectory control, the lack of fine-grained stylization intensity adjustment, and inadequate generalization across diverse motion styles. To address these challenges, We introduce MCSMoG, a novel framework for controllable stylized motion synthesis through multi-conditional guidance. First, a new Multi-Conditional Motion Latent Diffusion (MC-MLD) model is proposed to introduce additional trajectory guidance and achieve trajectory decoupling. Second, we develop a Style and Non-Style Feature Fusion Module that dynamically blends motion features through an adjustable parameter, providing control over stylization intensity. Third, we integrate MotionCLIP as our style encoder, enhancing the model’s generalization capability across diverse and unseen motion styles. Extensive experiments conducted on the combined HumanML3D and 100STYLE datasets demonstrate that our approach outperforms state-of-the-art methods, achieving a 4.6% reduction in FID scores and a 4.1% increase in motion diversity. User studies further confirm the superiority of our method in style fidelity, semantic consistency, and motion naturalness. Xinzhu Li, Guanghui Yue 0001, Wei Zhou 0021, Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao |
ICME | 4 |
| 2025 | CLIP-DQA: Blindly Evaluating Dehazed Images from Global and Local Perspectives Using CLIPabstractBlind dehazed image quality assessment (BDQA), which aims to accurately predict the visual quality of dehazed images without any reference information, is essential for the evaluation, comparison, and optimization of image dehazing algorithms. Existing learning-based BDQA methods have achieved remarkable success, while the small scale of DQA datasets limits their performance. To address this issue, in this paper, we propose to adapt Contrastive Language-Image Pre-Training (CLIP), pre-trained on large-scale image-text pairs, to the BDQA task. Specifically, inspired by the fact that the human visual system understands images based on hierarchical features, we take global and local information of the dehazed image as the input of CLIP. To accurately map the input hierarchical information of dehazed images into the quality score, we tune both the vision branch and language branch of CLIP with prompt learning. Experimental results on two authentic DQA datasets demonstrate that our proposed approach, named CLIP-DQA, achieves more accurate quality predictions over existing BDQA methods. The code is available at https://github.com/JunFu1995/CLIP-DQA. Yirui Zeng, Jun Fu 0007, Hadi Amirpour, Huasheng Wang, Guanghui Yue 0001, Hantao Liu, Ying Chen 0011, Wei Zhou 0021 |
ISCAS | 5 |
| 2025 | DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait RecognitionabstractRobust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent literature, these 2D representations fall short of capturing sufficient cues that can be exploited to handle viewpoint variations, and capture finer and meaningful details of gait. In this paper, we introduce a novel framework, termed DepthGait, that incorporates RGB-derived depth maps and silhouettes for enhanced gait recognition. Specifically, apart from the 2D silhouette representation of the human body, the proposed pipeline explicitly estimates depth maps from a given RGB image sequence and uses them as a new modality to capture discriminative features inherent in human locomotion. In addition, a novel multi-scale and cross-level fusion scheme has also been developed to bridge the modality gap between depth maps and silhouettes. Extensive experiments on standard benchmarks demonstrate that the proposed DepthGait achieves state-of-the-art performance compared to peer methods and attains an impressive mean rank-1 accuracy on the challenging datasets. Xinzhu Li, Juepeng Zheng, Yikun Chen, Xudong Mao, Guanghui Yue 0001, Wei Zhou 0021, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao |
ACM Multimedia | 5 |
| 2025 | VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question AnsweringabstractCross-video question answering presents significant challenges beyond traditional single-video understanding, particularly in establishing meaningful connections across video streams and managing the complexity of multi-source information retrieval. We introduce VideoForest, a novel framework that addresses these challenges through person-anchored hierarchical reasoning, enabling effective cross-video understanding without requiring end-to-end training. VideoForest integrates three key innovations: 1) a human-anchored feature extraction mechanism that employs ReID and tracking algorithms to establish robust spatiotemporal relationships across multiple video sources; 2) a multi-granularity spanning tree structure that hierarchically organizes visual content around person-level trajectories; and 3) a multi-agent reasoning framework that efficiently traverses this hierarchical structure to answer complex queries. To evaluate our method, we develop CrossVideoQA, a comprehensive benchmark specifically designed for person-centric cross-video analysis. Experimental results demonstrate VideoForest's superior performance in cross-video reasoning tasks, achieving 71.93% accuracy in person recognition, 83.75% in behavior analysis, and 51.67% in summarization and reasoning. Yiran Meng, Junhong Ye, Wei Zhou 0021, Guanghui Yue 0001, Xudong Mao, Ruomei Wang 0001, Baoquan Zhao |
ACM Multimedia | 4 |
| 2025 | Learning Content-enhanced Tokens for Domain Generalized Semantic SegmentationabstractVisual foundation models (VFMs) have demonstrated impressive generalization capabilities in computer vision tasks. Previous studies show that fine-tuning VFMs with learnable tokens can achieve better generalization performance than full-parameter fine-tuning. The problem we need to address is how to learn the tokens that focus on the content information while ignoring the influence of style. For this purpose, we propose a novel Dual-Branch Content-enhanced Token (DBCT) learning framework. Specifically, we construct a style-suppressing branch, which contains a Style-sensitive Channel Suppression (SCS) module to transform the frozen VFM features into style-suppressed features, enabling the learning of style-invariant tokens. In addition, to compensate for the content degradation caused by the style-suppressing branch, we introduce a content-preserving branch that directly takes the frozen VFM features as input to learn content-focused tokens. Meanwhile, we propose a Token-query Linking (TLink) strategy to connect the two sets of tokens with the queries in the decoder. Through extensive experiments, our method achieves advanced results on various benchmarks. Shishun Tian, Wenbin Zou, Yuanhao Gong, Guanghui Yue 0001, Ting Su 0004 |
MMAsia | 5 |
| 2025 | Distortion-Aware Network for Zero-Reference Retinal Image EnhancementabstractCaptured retinal images usually have quality issues, manifested as containing multiple distortions (e.g., low light and blurring). Low-quality images bring a challenge to the screening and diagnosis of ophthalmic diseases. Existing image enhancement methods typically neglect the analysis of distortions and require high-quality reference images for model learning, making them unsuitable for clinical applications. In this paper, we propose a Distortion-Aware Network (DANet) for retinal image enhancement in a zero-reference way. DANet consists of three parallel branches by incorporating atmospheric scattering theory, which decomposes the low-quality image into a clean image, a transmission map, and an atmospheric light map. The upper branch utilizes a dark channel prior module to estimate the atmospheric light map, and the middle branch uses a transmission map generation module to estimate the transmission map. In contrast, the lower branch uses a deblurring module and a low-light enhancement module to obtain a deblurred image and an illumination-enhanced image and fuses these two images using a fusion block to generate the final enhanced image. Taking into account the limited publicly available datasets, we curate two datasets for the retinal image enhancement task. Experimental results show that our DANet can greatly improve the visual quality of the image with good interpretability, achieving superior performance over seven state-of-the-art methods. Tianwei Zhou, Yuhang Feng, Shaoping Zhang, Linling Li, Guanghui Yue 0001, Shishun Tian, Tianfu Wang 0001 |
MMAsia | 5 |
| 2025 | Multi-task cyclical consistency learning based medical image segmentation
Le Han, Xueyu Liu, Guanghui Yue 0001, Mingqiang Wei, Yongfei Wu |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | GLMKD: Joint global and local mutual knowledge distillation for weakly supervised lesion segmentation in histopathology images
Hangbei Cheng, Xueyu Liu, Jun Zhang 0095, Xiaorong Dong, Xuetao Ma 0001, Xing Chen 0017, Guanghui Yue 0001, Yidi Li 0001, Yongfei Wu |
Expert Syst. Appl. | 9 |
| 2025 | Improving multi-modal brain tumor segmentation via pre-training and knowledge distillation based post-training
Weide Liu, Jingwen Hou, Xiaoyang Zhong, Huijing Zhan, Jun Cheng 0003, Yuming Fang 0001, Guanghui Yue 0001 |
Neurocomputing | 7 |
| 2025 | DFedMQ: Decentralized Federated Learning Based on Dynamic Selection Collaboration and Topology OptimizationabstractCentralized federated learning is being widely researched and applied. However, centralized federated learning is prone to problems such as single point of failure and privacy disclosure because it relies too much on the central server. Focusing on decentralized federated learning, this paper innovatively constructs a decentralized federated learning framework based on dynamic selection collaboration and topology optimization. Firstly, we propose a dynamic client selection algorithm based on node training quality. Then, a global network topology for data communication is constructed by us based on the Watts-Strogatz(WS) model. Finally, we design a temporary topology algorithm to realize synchronization and model update in training. In the process of decentralized federated learning, the global network topology based on WS model cooperates with the current network topology constructed by temporary topology algorithm. The two network topologies work together to realize a dynamic client selection algorithm based on node training quality. A large number of experiments verify that DFedMQ can accelerate the model convergence and improve the training effect under the premise of privacy protection. Bin Jiang 0003, Guanghui Yue 0001, Xue-rong Cui, Jian Wang 0061, Houbing Song |
IEEE Internet Things J. | 3 |
| 2025 | Multi-instance curriculum learning for histopathology image classification with bias reduction
Zihao Mi, Xueyu Liu, Guanghui Yue 0001, Junhong Yue, Mingqiang Wei, Yidi Li 0001, Yongfei Wu |
Medical Image Anal. | 4 |
| 2025 | Cross-Modality Interactive Attention Network for AI-generated image quality assessment
Tianwei Zhou, Songbai Tan, Leida Li, Baoquan Zhao, Qiuping Jiang, Guanghui Yue 0001 |
Pattern Recognit. | 6 |
| 2025 | CLIP-DQA V2: Exploring CLIP for Dehazed Image Quality Assessment From a Fragment-Level Perspective
Yirui Zeng, Jun Fu 0007, Guanghui Yue 0001, Hantao Liu, Wei Zhou 0021 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Boundary-Supplementary Network for Carotid Plaque Segmentation in Ultrasound ImagesabstractAccurate measurement of carotid plaques in ultrasound images is essential for the early prevention and intervention of stroke. However, carotid plaques exhibit complex and diverse shapes, with blurry and irregular boundaries, making this task very challenging. This paper proposes a novel Boundary-Supplementary Network (BSNet) for carotid plaque segmentation. BSNet follows an encoder-decoder structure and incorporates two essential modules: the Multi-Scale Attention Module (MSAM) and the Boundary-Supplementary Module (BSM). MSAM is deployed at each layer to comprehensively analyze the feature at different spatial scales and dynamically fuse multi-scale features using the attention mechanism. It aids the network in better recognizing carotid plaques with irregular shapes. BSM incorporates the outputs of the MSAM at the current and adjacent high layer with the prediction map of the adjacent high layer to enhance boundary regions. Through topdown deep supervision, BSNet can accurately localize the carotid plaque regions with clear boundaries. Extensive experiments on a collected dataset demonstrate that BSNet outperforms 8 state of-the-art methods in carotid plaque segmentation. Xinjian Zhu, Bingbing Cheng, Yuqiang Shen, Guanghui Yue 0001 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Progressive Feature Enhancement Network for Automated Colorectal Polyp SegmentationabstractIn recent years, colorectal polyp segmentation has attracted increasing attention in academia and industry. Although most existing methods can achieve commendable outcomes, they often confront difficulty when localizing challenging polyps with complex background, variable shape/size, and ambiguous boundary, because of the limitations in modeling global context and in cross-layer feature interaction. To cope with these challenges, this paper proposes a novel Progressive Feature Enhancement Network (PFENet) for polyp segmentation. Specifically, PFENet follows an encoder-decoder structure and utilizes the pyramid vision transformer as the encoder to capture multi-scale long-term dependencies at different stages. A cross-stage feature enhancement (CFE) module is embedded in each stage. The CFE module enhances the feature representation ability from interaction among adjacent stages, which helps integrate scale information for recognizing polyps with complex background and variable shape/size. In addition, a foreground boundary co-enhancement (FBC) module is used at each decoder to simultaneously enhance the foreground and boundary information by incorporating the output of the adjacent high stage and the coarse segmentation map, which is generated by fusing features of all four stages via a coarse map generation module. Through top-down connections of FBC modules, PFENet can progressively refine the prediction in a coarse-to-fine manner. Extensive experiments show the effectiveness of our PFENet in the polyp segmentation task, with the mIoU and mDic values over 0.886 and 0.931 tested on two in-domain datasets and over 0.735 and 0.809 tested on three out-of-domain datasets.Note to Practitioners—Automated and accurate polyp segmentation in colonoscopy images is a critical prerequisite for subsequent detection, removal, and diagnosis of polyps in clinical practice. This paper proposes a novel deep neural network for polyp segmentation, termed PFENet, with a CFE module to enhance the feature representation ability for better capturing polyps with complex background and variable shape/size, and a FBC module to simultaneously enhance the foreground and boundary information on the feature representation provided by the CFE module. Qualitative and quantitative results on five public datasets show that our PFENet yields accurate predictions and is superior to 9 state-of-the-art polyp segmentation methods. The proposed PFENet will facilitate potential computer-aided diagnosis systems in clinical practice, in which it can better promote medical decision-making than competing methods in polyp detection and removal. Guanghui Yue 0001, Houlu Xiao, Tianwei Zhou, Songbai Tan, Yun Liu 0009, Weiqing Yan |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Boundary-Guided Feature-Aligned Network for Colorectal Polyp SegmentationabstractColorectal polyp segmentation in endoscopic images is very important for the prevention and treatment of colorectal cancer. Because of the high similarity between polyps and their surrounding tissues, most deep neural network (DNN) based methods often struggle with blurry boundaries and result in inaccurate segmentation. In this paper, we propose a Boundary-guided Feature-aligned Network (BFNet) for polyp segmentation by taking a boundary prediction task as an auxiliary. Firstly, BFNet aggregates multi-layer features extracted from the backbone to mine boundary cues. Secondly, a flexible feature aggregation (FFA) module is used at each layer to adaptively fuse cross-layer features for coarse polyp localization. In the FFA module, considering the spatial misalignment between features at different layers, the feature of the high layer is aligned to and fused with that of the current layer using the deformable convolution and flexible merge block. After that, a boundary-guided feature enhancement (BFE) module is applied to refine the localization at boundary areas. In the BFE module, the boundary information is extracted and highlighted in both channel and spatial dimensions using the attention mechanisms with the assistance of boundary cues. By applying deep supervision to the BFE modules, BFNet can produce accurate polyp segmentation. Experimental results show that our BFNet outperforms 14 state-of-the-art DNN-based polyp segmentation methods on both in-domain and out-of-domain tests. Guanghui Yue 0001, Shangjie Wu, Cheng Zhao 0003, Tianwei Zhou, Baoquan Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Energy-Efficient Wireless Resource Allocation for Heterogeneous Federated Multitask Networks Based on Evolutionary LearningabstractWith the continuous development of 6G technology and the Internet of Things, small terminal devices are gradually joining deep model training through wireless networks, leading to the evolution of federated learning. In comparison to traditional centralized learning, federated learning not only leverages the computational power of individual terminals but also ensures the security of terminal data. However, the increasing number of devices poses new requirements on resource utilization in federated learning at scale. In this paper, we aim to address these challenges by proposing an energy-efficient and adaptive resource allocation strategy for wireless heterogeneous layered federated learning model (HLFLM). Specifically, we deploy both macro base stations and multiple micro base stations to construct a HLFLM, and perform resource allocation for subcarriers and power optimization. This approach focuses on optimizing energy consumption in federated learning networks while enhancing scalability and real-time performance of wireless communication. Experimental results demonstrate the effectiveness of the proposed method in medium-sized scenarios. Bin Jiang 0003, Lixin Cai, Guanghui Yue 0001, Fei Luo 0003, Shibao Li, Jian Wang 0061 |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Perception-Oriented Bidirectional Attention Network for Image Super-Resolution Quality AssessmentabstractMany super-resolution (SR) algorithms have been proposed to increase image resolution. However, full-reference (FR) image quality assessment (IQA) metrics for comparing and evaluating different SR algorithms are limited. In this work, we propose the Perception-oriented Bidirectional Attention Network (PBAN) for image SR FR-IQA, which is composed of three modules: an image encoder module, a perception-oriented bidirectional attention (PBA) module, and a quality prediction module. First, we encode the input images for feature representations. Inspired by the characteristics of the human visual system, we then construct the perception-oriented PBA module. Specifically, different from existing attention-based SR IQA methods, we conceive a Bidirectional Attention to bidirectionally construct visual attention to distortion, which is consistent with the generation and evaluation processes of SR images. To further guide the quality assessment towards the perception of distorted information, we propose Grouped Multi-scale Deformable Convolution, enabling the proposed method to adaptively perceive distortion. Moreover, we design Sub-information Excitation Convolution to direct visual perception to both sub-pixel and sub-channel attention. Finally, the quality prediction module is exploited to integrate quality-aware features and regress quality scores. Extensive experiments demonstrate that our proposed PBAN outperforms state-of-the-art quality assessment methods. Xiaoyuan Yang 0003, Guanghui Yue 0001, Jun Fu 0007, Qiuping Jiang, Xu Jia 0012, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
IEEE Trans. Image Process. | 3 |
| 2025 | Text-Guided Semantic Alignment Network With Spatial-Frequency Interaction for Infrared-Visible Image Fusion Under Extreme IlluminationabstractAlthough text-guided infrared-visible image fusion helps improve content understanding under extreme illumination, existing methods usually ignore semantic differences between textual and visual features, resulting in limited improvement. To address this challenge, we propose a Text-Guided Semantic Alignment Network, termed TSANet, for extreme-illumination infrared-visible image fusion. The network follows an encoder-decoder structure, with two image encoders, two text encoders, and one decoder. It uses a Semantic Alignment and Fusion (SAF) block to bridge the two image encoders in each layer. Specifically, the SAF block consists of two parallel Semantic Alignment (SA) modules, corresponding to the infrared and visible modalities, respectively, and a Spatial-Frequency Interaction (SFI) module. The SA module aligns the visual feature from the image encoder with its corresponding textual feature from the text encoder, to guide the network focus on key semantic regions of infrared and visible images. The SFI module aggregates the spatial and frequency information extracted from the modality-aligned features of two SA modules for complementary representation learning. The network progressively complements two image modalities by connecting the SAF blocks from top to down, and finally provides a visually pleasing fusion effect by feeding the output of the last block into the decoder. Recognizing that existing datasets lack illumination diversity, we contribute a new dataset specifically designed for extreme-illumination image fusion. Extensive experiments show the effectiveness and superiority of TSANet over seven state-of-the-art methods. The source code and dataset are available at https://github.com/WentaoLi-CV/TSANet. Guanghui Yue 0001, Cheng Zhao 0003, Zhiliang Wu, Tianwei Zhou, Qiuping Jiang, Runmin Cong |
IEEE Trans. Image Process. | 1 |
| 2025 | Benchmarking Laryngeal Neoplasm Segmentation: A Multicenter Dataset and an Effective MethodabstractWhile accurate and automatic Laryngeal Neoplasm Segmentation (LNS) can benefit the diagnosis and prevention of laryngeal cancers, existing LNS-related works are very limited due to the lack of public datasets. This paper conducts systematic research to take the research field a step further. Firstly, we create a multicenter LNS dataset, named as MLN-Seg. Collecting from four hospitals, it has 2,273 laryngeal images with a diversity in resolutions and modalities, where each image is pixel-wise annotated by experienced physicians. Secondly, considering the scarcity of LNS methods and similarity between LNS and Colorectal Polyp Segmentation (CPS) tasks, we collect 15 CPS methods and validate their performance on MLN-Seg. It shows that despite the similarity between the two tasks, existing CPS methods underperform on LNS, especially those with blurry boundaries and camouflaged characteristics. Lastly, considering the LNS challenges, we propose an effective segmentation method, termed Scale-Sensitive Network (S2Net). S2Net scales the feature at each layer of the network up and down and integrates all the scaled features to coarsely localize neoplasm regions. In addition, a Localization Calibration (LC) module is used to refine uncertain areas. By connecting the LC modules from top to down, S2Net can finally accurately segment the laryngeal neoplasms. Extensive tests on MLN-Seg shows that S2Net has better learning ability and generalizability than competing methods. In addition, evaluation on five public datasets shows that S2Net achieves comparable performance in the CPS task. Guanghui Yue 0001, Shangjie Wu, Ruxian Tian, Hanhe Lin, Huaiqing Lv, Zhenkun Yu, Xicheng Song |
IEEE Trans. Image Process. | 1 |
| 2025 | Adaptive Cross-Feature Fusion Network With Inconsistency Guidance for Multi-Modal Brain Tumor SegmentationabstractIn the context of contemporary artificial intelligence, increasing deep learning (DL) based segmentation methods have been recently proposed for brain tumor segmentation (BraTS) via analysis of multi-modal MRI. However, known DL-based works usually directly fuse the information of different modalities at multiple stages without considering the gap between modalities, leaving much room for performance improvement. In this paper, we introduce a novel deep neural network, termed ACFNet, for accurately segmenting brain tumor in multi-modal MRI. Specifically, ACFNet has a parallel structure with three encoder-decoder streams. The upper and lower streams generate coarse predictions from individual modality, while the middle stream integrates the complementary knowledge of different modalities and bridges the gap between them to yield fine prediction. To effectively integrate the complementary information, we propose an adaptive cross-feature fusion (ACF) module at the encoder that first explores the correlation information between the feature representations from upper and lower streams and then refines the fused correlation information. To bridge the gap between the information from multi-modal data, we propose a prediction inconsistency guidance (PIG) module at the decoder that helps the network focus more on error-prone regions through a guidance strategy when incorporating the features from the encoder. The guidance is obtained by calculating the prediction inconsistency between upper and lower streams and highlights the gap between multi-modal data. Extensive experiments on the BraTS 2020 dataset show that ACFNet is competent for the BraTS task with promising results and outperforms six mainstream competing methods. Guanghui Yue 0001, Guibin Zhuo, Tianwei Zhou, Weide Liu, Tianfu Wang 0001, Qiuping Jiang |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Subjective and Objective Quality Assessment of Colonoscopy VideosabstractCaptured colonoscopy videos usually suffer from multiple real-world distortions, such as motion blur, low brightness, abnormal exposure, and object occlusion, which impede visual interpretation. However, existing works mainly investigate the impacts of synthesized distortions, which differ from real-world distortions greatly. This research aims to carry out an in-depth study for colonoscopy Video Quality Assessment (VQA). In this study, we advance this topic by establishing both subjective and objective solutions. Firstly, we collect 1,000 colonoscopy videos with typical visual quality degradation conditions in practice and construct a multi-attribute VQA database. The quality of each video is annotated by subjective experiments from five distortion attributes (i.e., temporal-spatial visibility, brightness, specular reflection, stability, and utility), as well as an overall perspective. Secondly, we propose a Distortion Attribute Reasoning Network (DARNet) for automatic VQA. DARNet includes two streams to extract features related to spatial and temporal distortions, respectively. It adaptively aggregates the attribute-related features through a multi-attribute association module to predict the quality score of each distortion attribute. Motivated by the observation that the rating behaviors for all attributes are different, a behavior guided reasoning module is further used to fuse the attribute-aware features, resulting in the overall quality. Experimental results on the constructed database show that our DARNet correlates well with subjective ratings and is superior to nine state-of-the-art methods. Guanghui Yue 0001, Jingfeng Du, Tianwei Zhou, Wei Zhou 0021, Weisi Lin |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Pyramid Network With Quality-Aware Contrastive Loss for Retinal Image Quality AssessmentabstractCaptured retinal images vary greatly in quality. Low-quality images increase the risk of misdiagnosis. This motivates to design effective retinal image quality assessment (RIQA) methods. Current deep learning-based methods usually classify the image into three levels of "Good", "Usable", and "Reject", while ignoring the quantitative feedback for more detailed quality scores. This study proposes a unified RIQA framework, named QAC-Net, that can evaluate the quality of retinal images in both qualitative and quantitative manners. To improve the prediction accuracy, QAC-Net focuses on extracting discriminative features by using two strategies. On the one hand, it adopts a pyramid network structure that simultaneously inputs the scaled images to learn quality-aware features at different scales and purify the feature representation through a consistency loss. On the other hand, to improve feature representation, it utilizes a quality-aware contrastive (QAC) loss that considers quality relationships between different images. The QAC losses for qualitative and quantitative evaluation tasks have different forms in view of the task differences. Considering the shortage of datasets for the quantitative evaluation task, we construct a dataset with 2,300 authentically distorted retinal images, each of which is annotated with a numerical quality score through subjective experiments. Experimental results on public and our constructed datasets show that our QAC-Net is competent for the RIQA tasks with considerable performance. Guanghui Yue 0001, Shaoping Zhang, Tianwei Zhou, Bin Jiang 0003, Weide Liu, Tianfu Wang 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Unsupervised Low-Light Image Enhancement With Self-Paced LearningabstractLow-light image enhancement (LIE) aims to restore images taken under poor lighting conditions, thereby extracting more information and details to robustly support subsequent visual tasks. While past deep learning (DL)-based techniques have achieved certain restoration effects, these existing methods treat all samples equally, ignoring the fact that difficult samples may be detrimental to the network's convergence at the initial training stages of network training. In this paper, we introduce a self-paced learning (SPL)-based LIE method named SPNet, which consists of three key components: the feature extraction module (FEM), the low-light image decomposition module (LIDM), and a pre-trained denoise module. Specifically, for a given low-light image, we first input the image, its pseudo-reference image, and its histogram-equalized version into the FEM to obtain preliminary features. Second, to avoid ambiguities during the early stages of training, these features are then adaptively fused via an SPL strategy and processed for retinex decomposition via LIDM. Third, we enhance the network performance by constraining the gradient prior relationship between the illumination components of the images. Finally, a pre-trained denoise module reduces noise inherent in LIE. Extensive experiments on nine public datasets reveal that the proposed SPNet outperforms eight state-of-the-art DL-based methods in both qualitative and quantitative evaluations and outperforms three conventional methods in quantitative assessments. Yu Luo 0004, Xuanrong Chen, Jie Ling 0002, Chao Huang 0001, Wei Zhou 0021, Guanghui Yue 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Progressive Region-to-Boundary Exploration Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) aims to segment targeted objects that have similar colors, textures, or shapes to their background environment. Due to the limited ability in distinguishing highly similar patterns, existing COD methods usually produce inaccurate predictions, especially around the boundary areas, when coping with complex scenes. This paper proposes a Progressive Region-to-Boundary Exploration Network (PRBE-Net) to accurately detect camouflaged objects. PRBE-Net follows an encoder-decoder framework and includes three key modules. Specifically, firstly, both high-level and low-level features of the encoder are integrated by a region and boundary exploration module to explore their complementary information for extracting the object's coarse region and fine boundary cues simultaneously. Secondly, taking the region cues as the guidance information, a Region Enhancement (RE) module is used to adaptively localize and enhance the region information at each layer of the encoder. Subsequently, considering that camouflaged objects usually have blurry boundaries, a Boundary Refinement (BR) decoder is used after the RE module to better detect the boundary areas with the assistance of boundary cues. Through top-down deep supervision, PRBE-Net can progressively refine the prediction. Extensive experiments on four datasets indicate that our PRBE-Net achieves superior results over 21 state-of-the-art COD methods. Additionally, it also shows good results on polyp segmentation, a COD-related task in the medical field. Guanghui Yue 0001, Shangjie Wu, Tianwei Zhou, Jie Du 0001, Yu Luo 0004, Qiuping Jiang |
IEEE Trans. Multim. | 1 |
| 2025 | Smooth Multiple Kernel k-Means via Underlying Graph FilteringabstractClustering has attracted more and more attention as one of the most fundamental techniques in the field of unsupervised learning. To deal with nonlinear problems, clustering methods have been extended to the kernel version. As a traditional kernel clustering algorithm, multiple kernel k-means (MKKM) aims to learn clustering results from a consensus kernel obtained by combining a set of predefined kernels optimally. However, we observe that the existing MKKM algorithm and its variants insufficiently consider the noise that existed in kernel space and the underlying structure of kernelized data points. To this end, we propose a novel smooth MKKM via underlying graph filtering (SMKKM-UGF) to learn the smooth representations of kernelized data points through their nearby nodes in the underlying graph. In particular, different from the common graph filter, we jointly update the graph filter while learning the smooth kernel, so that the graph filter can be guaranteed to adapt to the updating kernel space constantly. Besides, an iterative algorithm with proven convergence is designed to solve the resultant optimization problem. Extensive experiments have been performed on numerous benchmark datasets, whose results prove the superiority of the proposed SMKKM-UGF compared to the other state-of-the-art clustering methods. The demo code of this work is publicly available at https://github.com/wqyang23/SMKKM-UGF.git. Wenqi Yang, Chang Tang, Xinwang Liu 0002, Guanghui Yue 0001, Yuanyuan Liu 0004, Changqing Zhang 0002, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Improved Text-Driven Human Motion Generation via Out-of-Distribution Detection and Rectification
Yiyu Fu, Baoquan Zhao, Chenlei Lv, Guanghui Yue 0001, Ruomei Wang 0001, Fan Zhou 0001 |
CVM (1) | 4 |
| 2024 | Dynamic Elite Individual Setting Based Heterogeneous Comprehensive Learning Particle Swarm Optimization
Tianwei Zhou, Yunbao Pan, Guanghui Yue 0001, Ben Niu 0002 |
ICIC (2) | 4 |
| 2024 | Subjective Quality Assessment of Thermal Infrared ImagesabstractThermal infrared images (TIIs) can be distorted by multiple factors, resulting in noise, low contrast, limited dynamic range, and fuzziness, which greatly impede their usefulness. It is crucial to evaluate the quality of TIIs. Unfortunately, there have been very few attempts to study this problem. In this study, we collected 1,000 authentically distorted TIIs using thermal infrared acquisition equipment and conducted strict subjective experiments to obtain a thermal infrared image quality assessment (IQA) database. Each image’s quality score was obtained under strict scoring rules. Finally, we investigated the feasibility of several no-reference (NR) IQA methods in quality assessment of TIIs. We found that existing NR-IQA methods achieve ordinary performance in such a task, and there is an urgent need to develop a specific IQA methods for TIIs. The findings together with the constructed database are expected to pave the way for the development of more advanced IQA methods for further development of this field. Guanghui Yue 0001, Jinxia Zhang, Zhaofei Xu, Shuigen Wang, Tianwei Zhou, Yuanhao Gong, Wei Zhou 0021 |
ICIP | 1 |
| 2024 | Clip-Medfake: Synthetic Data Augmentation With AI-Generated Content for Improved Medical Image ClassificationabstractData augmentation is serving as a critical and fundamental technology to improve model generalization and performance in a wide spectrum of machine learning tasks. Despite the increasing interest in developing various pathways to artificially generate new data to reduce the overfitting issue during model training, enriching the diversity of training data in the field of medicine remains facing enormous challenges. By virtue of recent advancements in generative artificial intelligence, we present a novel data augmentation framework, CLIP-MedFake, to address the shortage of training data used in medical image classification. The proposed method first employs the Stable Diffusion model to generate new fake data based on a small amount of training data, and then adopts the paradigm of few-shot learning and uses the CLIP architecture as the backbone to pre-train the model with synthetic data and then fine-tune it with real medical images. Extensive experiment results on two publicly available datasets demonstrate the effectiveness of the proposed method in promoting medical image classification. Honghui Chen, Baoquan Zhao, Guanghui Yue 0001, Weide Liu, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001 |
ICIP | 3 |
| 2024 | Start-Tv: A Closed-Form Initialization For Total Variation ModelsabstractAlthough there are many iterative solvers for total variation models, few attention has been paid on the fast and effective approximation to their optimal solutions. In this paper, we propose a closed-form filter that can efficiently and effectively approximate the optimal solution of total variation models. This filter has linear computation complexity $O(n)$ with respect to the total number of pixels and constant computation complexity $O(1)$ with respect to the window radius. Taking such filter as an initialization, our method can significantly accelerate all previous iterative solvers. Numerical experiments confirms that our initialization is roughly equivalent to $\mathbf{5 0}$ iterations in the iterative method but $\mathbf{1 0} \times$ faster. The proposed method can be applied in all total variation models to accelerate the optimization process, such as image smoothing, image reconstruction and optical flow estimation. Yuanhao Gong, Guanghui Yue 0001 |
ICIP | 2 |
| 2024 | Full-Reference Motion Quality Assessment Based on Efficient Monocular Parametric 3D Human Body ReconstructionabstractHuman motion capture and analysis are pivotal to a wide spectrum of killer applications in various domains such as sports, performing arts, diagnostic tests in physical medicine, rehabilitation, and figure training. However, automatic reconstruction, assessment, and visualization of human motions from a monocular video are still suffering from grand challenges that are inadequately addressed by existing studies. In this paper, we present a novel full-reference human motion quality assessment and visualization system based on monocular parametric 3D human body reconstruction. Specifically, our method first reconstructs a 3D parametric model from each sampled frame of a monocular video and harvests a physically plausible motion sequence using the proposed optimization scheme; Secondly, a full-reference assessment metric is designed to evaluate the consistency between the reconstructed motion and the reference; Finally, a new interactive visualization system is developed to facilitate multi-grained motion quality evaluation and visual analysis. Extensive quantitative and qualitative experiments demonstrate the effectiveness and superiority of the proposed method. Yiwei Yuan, Xiangyu Zeng 0005, Ling Xie, Yiyu Fu, Guanghui Yue 0001, Baoquan Zhao |
ICME | 6 |
| 2024 | Deep Bi-directional Attention Network for Image Super-Resolution Quality AssessmentabstractThere has emerged a growing interest in exploring efficient quality assessment algorithms for image super-resolution (SR). However, employing deep learning techniques, especially dual-branch algorithms, to automatically evaluate the visual quality of SR images remains challenging. Existing SR image quality assessment (IQA) metrics based on two-stream networks lack interactions between branches. To address this, we propose a novel full-reference IQA (FR-IQA) method for SR images. Specifically, producing SR images and evaluating how close the SR images are to the corresponding HR references are separate processes. Based on this consideration, we construct a deep Bidirectional Attention Network (BiAtten-Net) that dynamically deepens visual attention to distortions in both processes, which aligns well with the human visual system (HVS). Experiments on public SR quality databases demonstrate the superiority of our proposed BiAtten-Net over state-of-the-art quality assessment methods. In addition, the visualization results and ablation study show the effectiveness of bi-directional attention. Xiaoyuan Yang 0003, Jun Fu 0007, Guanghui Yue 0001, Wei Zhou 0021 |
ICME | 4 |
| 2024 | Predicting Plain Text Imageability for Faithful Prompt-Conditional Image Generation
Guanghui Yue 0001, Weide Liu, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao |
PRICAI (3) | 2 |
| 2024 | Cross-Device Image Saliency Detection: Database and Comparative AnalysisabstractWhile saliency detection for images has been extensively studied during the past decades, only a little work explores the influence of different viewing devices (i.e., tablet computer, mobile phone) towards human visual attention behavior. The lack of research in this field hinders the research progress in cross-device image saliency detection. In this paper, we first establish a novel cross-device saliency detection (CDSD) database based on eye-tracking experiments and investigate subjects’ visual attention behavior when using different viewing devices. Then, we evaluate several classic saliency detection models using the CDSD database and the evaluation results indicate that the cross-device performance of these models need further improvement. Finally, some meaningful discussions are provided which might enlighten the design of cross-device saliency detection model. The proposed CDSD database will be made publicly available. Xiaoying Ding, Guanghui Yue 0001 |
VCIP | 2 |
| 2024 | Image-Prompt Integration Network with Self-Ranking and Inter-Ranking Loss for AI-Generated Image Quality AssessmentabstractAI-generated images (AGIs) are increasingly utilized across diverse domains due to their ability to quickly produce high-quality visuals. However, assessing the quality of AGIs re-mains challenging due to their inherent variability and distinctive distortions. To address these challenges, we propose a novel AGI quality assessment method named SIRQA, which enhances feature representation by integrating visual features with textual prompts, effectively measureing the alignment between the generated images and the described content to improve the precision of quality assessment. Specifically, SIRQA employs self-ranking and inter-ranking mechanisms to refine feature representation. The self-ranking mechanism maintains consistency between feature distances and sampling scales, making sure that features from similar sampling scales are positioned closer together. Addition-ally, inter-ranking mechanism sorts the weighted similarity scores between images and prompts to align with the ranking in the label space. Extensive experiments on the AGIQA3K and PKUI2IQA datasets show that our SIRQA outperforms eight state-of-the-art algorithms in terms of both Spearman’s rank correlation coefficient (SRCC) and Pearson linear correlation coefficient (PLCC). Tianwei Zhou, Xizhang Yao, Songbai Tan, Xiaoying Ding, Guanghui Yue 0001 |
VCIP | 5 |
| 2024 | Parameter Control Framework for Multiobjective Evolutionary Computation Based on Deep Reinforcement LearningabstractTo address the challenge of parameter adjustment in complex environments, this paper introduces a transfer learning-based parameter control framework via deep reinforcement learning for multiobjective evolutionary algorithms (MOEAs). To avoid the requirement for accurate Pareto front information, this framework is proposed with comprehensive global-state information, including basic problem features, the relative position of individuals, the distribution of fitness value, and the grid-IGD. Building on this framework, four reinforced multiobjective evolutionary algorithms (r-MOEAs) are proposed and tested on four DTLZ benchmarks and eight WFG benchmarks. The results of the comparative analyses reveal that compared with the original MOEAs, the four r-MOEAs exhibit faster convergence and stronger robustness. It is also confirmed that our proposed parameter control framework has the capability to learn knowledge from different experiences and improve the performance of MOEAs. Tianwei Zhou, Ben Niu 0002, Guanghui Yue 0001 |
Int. J. Intell. Syst. | 5 |
| 2024 | Colorectal endoscopic image enhancement via unsupervised deep learning
Guanghui Yue 0001, Lvyin Duan, Jingfeng Du, Weiqing Yan, Shuigen Wang, Tianfu Wang 0001 |
Multim. Tools Appl. | 1 |
| 2024 | Boundary uncertainty aware network for automated polyp segmentation
Guanghui Yue 0001, Guibin Zhuo, Weiqing Yan, Tianwei Zhou, Chang Tang, Peng Yang 0011, Tianfu Wang 0001 |
Neural Networks | 1 |
| 2024 | Boundary Refinement Network for Colorectal Polyp Segmentation in Colonoscopy ImagesabstractPrecise polyp segmentation is vitally essential for detection and diagnosis of early colorectal cancer. Recent advances in artificial intelligence have brought infinite possibilities for this task. However, polyps usually vary greatly in shape and size and contain ambiguous boundary, bringing tough challenges to precise segmentation. In this letter, we introduce a novel Boundary Refinement Network (BRNet) for polyp segmentation. To be specific, we first introduce a boundary generation module (BGM) to generate boundary map by fusing both low-level spatial details and high-level concepts. Then, we utilize the boundary-guided refinement module to refine the polyp-aware features at each layer with the help of boundary cues from the BGM and the prediction from the adjacent high layer. Through top-down deep supervision, our BRNet can localize the polyp regions accurately with clear boundary. Extensive experiments are carried out on five datasets, and the results indicate the effectiveness of our BRNet over seven recently reported methods. Guanghui Yue 0001, Yuanyan Li, Wenchao Jiang, Wei Zhou 0021, Tianwei Zhou |
IEEE Signal Process. Lett. | 1 |
| 2024 | Pseudo-Supervised Low-Light Image Enhancement With Mutual LearningabstractLow-light image enhancement (LIE) is important for many high-level vision tasks as the poor visibility of underexposed images can severely degrade the performance of the subsequent image recognition, analysis, etc. Although recent deep-learning-based LIE methods exhibit promising performance, most of them require a large number of paired training images, thereby limiting the practicability to real scenarios. In this paper, we propose a pseudo-supervised LIE method with the integration of mutual learning. Specifically, for the given low-light image, we first use a quadratic curve to generate a pseudo-clear image, which is served as the auxiliary ground truth for supervision, then the pseudo-paired images are simultaneously input to two parallel homogeneous branches to learn the expected enhanced result through the knowledge distillation of two branches via mutual learning. As both the generated image and the input low-light image underlies the desired solution, the mutual learning strategy enables the two branches learn from each other and produce the final results. Extensive experiments demonstrate that the proposed method outperforms most existing unsupervised LIE methods in terms of both qualitative and quantitative evaluations, and also achieves competitive performance against many supervised and semi-supervised methods. Yu Luo 0004, Bijia You, Guanghui Yue 0001, Jie Ling 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Deep Pyramid Network for Low-Light Endoscopic Image EnhancementabstractEndoscopic images captured under low-light enclosed intestinal environment usually have poor visibility (manifested as uneven illumination and noise), affecting the work efficiency of physicians and the accuracy of lesion detection. To improve the image quality, the literature has reported many low-light image enhancement (LIE) methods. However, most methods do not perform well in handling the low-light endoscopic image enhancement (LEIE) task, usually bringing additional artifacts or amplifying noise. In this paper, we propose a novel deep pyramid enhancement network (DPENet) to enhance endoscopic images from both global and local perspectives. Specifically, considering the uneven illumination of endoscopic images, DPENet utilizes an image pyramid framework with three parallel branches to explore and integrate both global and local features at different scales. To suppress noise, DPENet sets multiple scale-space feature extraction blocks (SFEBs) in each branch. SFEB consists of a contextual feature extraction module (CFEM) and a spatial residual attention module (SRAM). CFEM mines contextual information to help the network understand semantic information while suppress the isolated noise. SRAM leverages the spatial attention mechanism to help the network adaptively focus on dim regions. Experimental results on a public dataset and our collected dataset show that DPENet is competent for the LEIE task with promising results, and outperforms 9 state-of-the-art LIE methods in both qualitative and quantitative aspects. Guanghui Yue 0001, Runmin Cong, Tianwei Zhou, Leida Li, Tianfu Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Dual-Constraint Coarse-to-Fine Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) is an important yet challenging task, with great application values in industrial defect detection, medical care, etc. The challenges mainly come from the high intrinsic similarities between target objects and background. In this paper, inspired by the biological studies that object detection consists of two steps, i.e., search and identification, we propose a novel framework, named DCNet, for accurate COD. DCNet explores candidate objects and extra object-related edges through two constraints (object area and boundary) and detects camouflaged objects in a coarse-to-fine manner. Specifically, we first exploit an area-boundary decoder (ABD) to obtain initial region cues and boundary cues simultaneously by fusing multi-level features of the backbone. Then, an area search module (ASM) is embedded into each level of the backbone to adaptively search coarse regions of objects with the assistance of region cues from the ABD. After the ASM, an area refinement module (ARM) is utilized to identify fine regions of objects by fusing adjacent-level features with the guidance of boundary cues. Through the deep supervision strategy, DCNet can finally localize the camouflaged objects precisely. Extensive experiments on three benchmark COD datasets demonstrate that our DCNet is superior to 12 state-of-the-art COD methods. In addition, DCNet shows promising results on two COD-related tasks, i.e., industrial defect detection and polyp segmentation. Guanghui Yue 0001, Houlu Xiao, Hai Xie, Tianwei Zhou, Wei Zhou 0021, Weiqing Yan, Baoquan Zhao, Tianfu Wang 0001, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Multitask Deep Neural Network With Knowledge-Guided Attention for Blind Image Quality AssessmentabstractBlind image quality assessment (BIQA) targets predict the perceptual quality of an image without any reference information. However, known methods have considerable room for performance improvement due to limited efforts in distortion knowledge usage. This paper proposes a novel multitask learning based BIQA method termed KGANet, which takes image distortion classification as an auxiliary task and uses the knowledge learned from the auxiliary task to assist accurate quality prediction. Different from existing CNN-based methods, KGANet adopts a transformer as the backbone for feature extraction, which can learn more powerful and robust representations. Specifically, it comprises two essential components: a cross-layer information fusion (CIF) module and a knowledge-guided attention (KGA) module. Considering that both global and local distortions appear in an image, CIF fuses the features of the adjacent layers extracted by the backbone to obtain a multiscale feature representation. KGA incorporates the distortion probability estimated by the auxiliary task with the distortion embeddings, which are selected from subword unit embeddings based on a textual template, to form distortion knowledge. This knowledge further serves as guidance to enhance the features of each layer and strengthen the connection between the main and auxiliary task. We demonstrate the effectiveness of the proposed KGANet through extensive experiments on benchmark databases. Experimental results show that KGANet correlates well with subjective perceptual judgments and achieves superior performance over 12 state-of-the-art BIQA methods. Tianwei Zhou, Songbai Tan, Baoquan Zhao, Guanghui Yue 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Self-Supervised Hyperspectral Anomaly Detection Based on Finite Spatialwise AttentionabstractHyperspectral anomaly detection (HAD) is of great value in both practical and theoretical terms. However, due to the lack of available semantic labels, previous works mainly relied on unsupervised or semi-supervised methods to construct learning models, which inevitably lacked semantic guidance and led to limited anomaly detection (AD) effectiveness. Besides, few previous methods jointly mine spectral and spatial global dependencies, which limits their effectiveness in practical scenarios. To address the above problems, we design a novel self-supervised HAD method, named the Self-Supervised Hyperspectral Anomaly Detection method based on the Finite Spatial-wise Attention. The core of proposed method is the designed Self-Supervised Hyperspectral Anomaly Detection transFormer (SSHADFormer). It explores the specific spectral attributes of hyperspectral images (HSIs) to reconstruct background HSI from a given RGB image, which solves the difficulty of acquiring semantic information and enhances the agility of AD models. In addition, we propose a Finite Spatial-wise Attention mechanism. The mechanism mines the cluster structure of the background spectrum in a data-driven manner, enhancing the discriminative ability between background and anomalous targets while avoiding anomalous targets interference during training. Extensive experiments on six public datasets demonstrate the effectiveness and agility of the proposed method. Dan Ma 0003, Guanghui Yue 0001, Beichen Li 0002, Runmin Cong, Zhiqiang Wu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | HT-RCM: Hashimoto's Thyroiditis Ultrasound Image Classification Model Based on Res-FCT and Res-CAMabstractThe early lesions of Hashimoto's thyroiditis are inconspicuous, and the ultrasonic features of these early lesions are indistinguishable from other thyroid diseases. This paper proposes a Hashimoto Thyroiditis ultrasound image classification model HT-RCM which consists of a Residual Full Convolution Transformer (Res-FCT) model and a Residual Channel Attention Module (Res-CAM). To collect the low-order information caused by hypoechoic signals accurately, the residual connection is injected between FCTs to form Res-FCT which helps HT-RCM superimpose the low-order input information and high-order output information together. Res-FCT can make HT-RCM focus more on hypoechoic information while avoiding gradient dispersion. The initial feature map is inserted into Res-FCT again through a down-sampling component, which further helps HT-RCM exact multi-level original semantic information in the ultrasound image. Res-CAM is constructed by implementing a residual connection between a channel attention module and a convolution layer. Res-CAM can effectively increase the weights of the lesion channels while suppressing the weights of the noise channels, which makes HT-RCM focus more on the lesion regions. The experimental results on our collected dataset show that HT-RCM outperforms the mainstream models and obtains state-of-the-art performance in HT ultrasound image classification. Wenchao Jiang, Tianchun Luo, Guanghui Yue 0001, Zhiming Zhao, Jianxuan Wen |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Specificity-Aware Federated Learning With Dynamic Feature Fusion Network for Imbalanced Medical Image ClassificationabstractRecently, federated learning has become a powerful technique for medical image classification due to its ability to utilize datasets from multiple clinical clients while satisfying privacy constraints. However, there are still some obstacles in federated learning. Firstly, most existing methods directly average the model parameters collected by medical clients on the server, ignoring the specificities of the local models. Secondly, class imbalance is a common issue in medical datasets. In this article, to handle these two challenges, we propose a novel specificity-aware federated learning framework that benefits from an Adaptive Aggregation Mechanism (AdapAM) and a Dynamic Feature Fusion Strategy (DFFS). Considering the specificity of each local model, we set the AdapAM on the server. The AdapAM utilizes reinforcement learning to adaptively weight and aggregate the parameters of local models based on their data distribution and performance feedback for obtaining the global model parameters. For the class imbalance in local datasets, we propose the DFFS to dynamically fuse the features of majority classes based on the imbalance ratio in the min-batch and collaborate the rest of features. We conduct extensive experiments on a dermoscopic dataset and a fundus image dataset. Experimental results show that our method can achieve state-of-the-art results in these two real-world medical applications. Guanghui Yue 0001, Peishan Wei, Tianwei Zhou, Youyi Song, Cheng Zhao 0003, Tianfu Wang 0001, Bai Ying Lei |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Perceptual Quality Assessment of Retouched Face ImagesabstractNowadays, it is a common practice to retouch face images before sharing them on websites, social media, and even identification cards. In response, increased criticisms have appeared about taking photo retouching to an extreme. This naturally leads to the necessity of designing perceptual quality assessment methods that can measure how much a retouched face image has strayed from reality. However, such an issue has seldom been considered. In this paper, we conduct both subjective and objective studies to advance this field. Firstly, we construct a benchmark database (termed SZU-RFD) via subjective experiments. SZU-RFD consists of 200 high-quality images with Asian faces and 1,600 retouched images generated by three popular photo-editing tools under different settings. Secondly, considering that retouching usually distorts the image texture, we propose a novel no-reference (NR) quality assessment method, named TANet, for retouched face images by taking the textural artifact into account. Specifically, a texture enhancement module is embedded into the shallow layer to help the network focus on textural information, and a multi-task learning strategy is applied to improve the performance of the main task with the assistance of an auxiliary task, i.e., texture recognition. Extensive experiments on the constructed SZU-RFD show that our proposed TANet correlates well with subjective perceptual judgments and is superior to 19 mainstream NR image quality assessment methods in evaluating retouched face images. Guanghui Yue 0001, Honglv Wu, Qiuping Jiang, Tianwei Zhou, Weiqing Yan, Tianfu Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | CDINet: Content Distortion Interaction Network for Blind Image Quality AssessmentabstractPerceptual image quality is related to content and distortion. Distortion classification is a common way to learn distortion information. How to extract distortion information consistent with human perception is a problem to be solved. Besides, the joint effect on image quality caused by the interplay of content and distortion has not been fully studied. In this paper, a novel Content Distortion Interaction Network (CDINet) is proposed for blind image quality assessment. Distortion representation are guided by content representation to learn quality-aware representation. CDINet consists of four components: a Distortion-Aware Module (DAM), a Content-Aware Module (CAM), an Asymmetric Content-Distortion Interaction (ACDI) module, and a quality regression module. The content representation and distortion representation are extracted respectively and fused interactively in CDINet. Specifically, with the assistance of image restoration, distortion representation consistent with human perception is learned. To further improve the ability in distortion representation, the DAM is used to construct the differences between the distorted image and its reference image. The proposed ACDI module enables the interaction of content and distortion representations to occur at different levels with less computational cost. Since the proposed CDINet considers the joint impact on image quality caused by the interplay of content and distortion, the predicted image qualities highly align with human perception. Comprehensive experiments on 8 benchmark datasets demonstrate that the proposed CDINet effectively extracts quality-aware representation, achieving state-of-the-art performance in evaluating both synthetically and authentically distorted images. Limin Zheng, Yu Luo 0004, Zihan Zhou 0007, Jie Ling 0002, Guanghui Yue 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | MHW-GAN: Multidiscriminator Hierarchical Wavelet Generative Adversarial Network for Multimodal Image FusionabstractImage fusion technology aims to obtain a comprehensive image containing a specific target or detailed information by fusing data of different modalities. However, many deep learning-based algorithms consider edge texture information through loss functions instead of specifically constructing network modules. The influence of the middle layer features is ignored, which leads to the loss of detailed information between layers. In this article, we propose a multidiscriminator hierarchical wavelet generative adversarial network (MHW-GAN) for multimodal image fusion. First, we construct a hierarchical wavelet fusion (HWF) module as the generator of MHW-GAN to fuse feature information at different levels and scales, which avoids information loss in the middle layers of different modalities. Second, we design an edge perception module (EPM) to integrate edge information from different modalities to avoid the loss of edge information. Third, we leverage the adversarial learning relationship between the generator and three discriminators for constraining the generation of fusion images. The generator aims to generate a fusion image to fool the three discriminators, while the three discriminators aim to distinguish the fusion image and edge fusion image from two source images and the joint edge image, respectively. The final fusion image contains both intensity information and structure information via adversarial learning. Experiments on public and self-collected four types of multimodal image datasets show that the proposed algorithm is superior to the previous algorithms in terms of both subjective and objective evaluation. Cheng Zhao 0003, Peng Yang 0011, Feng Zhou 0003, Guanghui Yue 0001, Shuigen Wang, Huisi Wu, Guoliang Chen 0005, Tianfu Wang 0001, Bai Ying Lei |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | GCFAgg: Global and Cross-View Feature Aggregation for Multi-View ClusteringabstractMulti-view clustering can partition data samples into their categories by learning a consensus representation in unsupervised way and has received more and more attention in recent years. However, most existing deep clustering methods learn consensus representation or view-specific representations from multiple views via view-wise aggregation way, where they ignore structure relationship of all samples. In this paper, we propose a novel multi-view clustering network to address these problems, called Global and Cross-view Feature Aggregation for Multi-View Clustering (GCFAggMVC). Specifically, the consensus data presentation from multiple views is obtained via cross-sample and cross-view feature aggregation, which fully explores the complementary of similar samples. Moreover, we align the consensus representation and the view-specific representation by the structure-guided contrastive learning module, which makes the view-specific representations from different samples with high structure relationship similar. The proposed module is a flexible multi-view data representation module, which can be also embedded to the incomplete multi-view data clustering task via plugging our module into other frameworks. Extensive experiments show that the proposed method achieves excellent performance in both complete multi-view data clustering tasks and incomplete multi-view data clustering tasks. Weiqing Yan, Yuanyang Zhang, Chenlei Lv, Chang Tang, Guanghui Yue 0001, Weisi Lin |
CVPR | 5 |
| 2023 | Subjective Quality Assessment of Enhanced Retinal ImagesabstractMany retinal images sometimes suffer from uneven illumination, which influences the analysis and diagnosis of retinal diseases. To improve the image quality of those retinal images, one feasible solution is to utilize low-light image enhancement (LIE) algorithms. However, how to evaluate the perceptual quality of enhanced retinal images (ERIs) generated by different LIE algorithms remains a challenging problem. In this paper, we conduct subjective experiments to investigate the quality assessment of ERIs. First, we collect 250 retinal images with the authentic low-light distortion, and then adopt eight LIE algorithms to produce 2000 ERIs. Second, a subjective experiment is conducted, resulting in the proposed Enhanced Retinal Image Quality Assessment Database (ERIQAD). Finally, we test some well-known no reference image quality assessment (NR IQA) methods on our proposed ERIQAD. Experimental results demonstrate that existing mainstream NR IQA methods merely achieve ordinary performance to predict the perceptual quality of ERIs. Guanghui Yue 0001, Shaoping Zhang, Tianwei Zhou, Wei Zhou 0021 |
ICIP | 1 |
| 2023 | No Reference Image Quality Assessment Via Quality Difference LearningabstractFor human beings, there is a natural preference for judging the relative quality rather than directly predicting the quality score of an image. Based on this view, we propose an image quality difference learning network (IQDLNet) for evaluating image quality in a no-reference manner. Specifically, the proposed IQDLNet consists of a quality difference-aware network (QDAN) and a quality assessment network (QAN). The QDAN aims to predict the score difference between two randomly matched images and the QAN aims to predict the quality score of these two images. To further enhance the mutual understanding of image semantics, a semantic interaction module (SIM) is proposed with a dual regressor set up to carry out competitive learning in combination with the quality difference-aware feature. Experimental results on five IQA datasets demonstrate the superior performance of the proposed method over eight state-of-the-arts. Jiaming Xie, Yu Luo 0004, Jie Ling 0002, Guanghui Yue 0001 |
ICME | 4 |
| 2023 | Detection of GAN generated image using color gradient representation
Yun Liu 0009, Zuliang Wan, Xiaohua Yin, Guanghui Yue 0001, Aiping Tan, Zhi Zheng 0006 |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | A no-reference panoramic image quality assessment with hierarchical perception and color features
Yun Liu 0009, Xiaohua Yin, Chang Tang, Guanghui Yue 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Blind omnidirectional image quality assessment with representative features and viewport oriented statistical features
Yun Liu 0009, Xiaohua Yin, Guanghui Yue 0001, Zhi Zheng 0006, Jinhe Jiang, Quangui He, Xinzhuang Li |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | Adversarial learning-based multi-level dense-transmission knowledge distillation for AP-ROP detection
Hai Xie, Yaling Liu, Haijun Lei, Tiancheng Song, Guanghui Yue 0001, Yueshanyi Du, Tianfu Wang 0001, Bai Ying Lei |
Medical Image Anal. | 5 |
| 2023 | Multi-scale enhanced graph convolutional network for mild cognitive impairment detection
Bai Ying Lei, Yun Zhu 0006, Shuangzhi Yu, Huoyou Hu, Yanwu Xu 0001, Guanghui Yue 0001, Tianfu Wang 0001, Cheng Zhao 0003, Shaobin Chen, Peng Yang 0011, Xuegang Song, Xiaohua Xiao, Shuqiang Wang |
Pattern Recognit. | 6 |
| 2023 | Reduced-Reference Quality Assessment of Point Clouds via Content-Oriented Saliency ProjectionabstractMany dense 3D point clouds have been exploited to represent visual objects instead of traditional images or videos. To evaluate the perceptual quality of various point clouds, in this letter, we propose a novel and efficient Reduced-Reference quality metric for point clouds, which is based on Content-oriented sAliency Projection (RR-CAP). Specifically, we make the first attempt to simplify reference and distorted point clouds into projected saliency maps with a downsampling operation. Through this process, we tackle the issue of transmitting large-volume original point clouds to end-users for quality assessment. Then, motivated by the characteristics of the human visual system (HVS), the objective quality scores of distorted point clouds are produced by combining content-oriented similarity and statistical correlation measurements. Finally, extensive experiments are conducted on SJTU-PCQA and WPC databases. The experiment results demonstrate that our proposed algorithm outperforms existing reduced-reference and no-reference quality metrics, and significantly reduces the performance gap between state-of-the-art full-reference quality assessment methods. In addition, we show the performance variation of each proposed technical component by ablation tests. Wei Zhou 0021, Guanghui Yue 0001, Ruizeng Zhang, Yipeng Qin, Hantao Liu |
IEEE Signal Process. Lett. | 2 |
| 2023 | Perceptual Quality Assessment of Enhanced Colonoscopy Images: A Benchmark Dataset and an Objective MethodabstractIn colonoscopy, the captured images are usually with low-quality appearance, such as non-uniform illumination, low contrast, etc., due to the specialized imaging environment, which may provide poor visual feedback and bring challenges to subsequent disease analysis. Many low-light image enhancement (LIE) algorithms have recently proposed to improve the perceptual quality. However, how to fairly evaluate the quality of enhanced colonoscopy images (ECIs) generated by different LIE algorithms remains a rarely-mentioned and challenging problem. In this study, we carry out a pioneering investigation on perceptual quality assessment of ECIs. Firstly, considering the lack of specific datasets, we collect 300 low-light images with diverse contents during the real-world colonoscopy and conduct rigorous subjective studies to compare the performance of 8 popular LIE methods, resulting in a benchmark dataset (named ECIQAD) for ECIs. Secondly, in view of the distinctive distortion characteristics of ECIs, we propose an effective no-reference Enhanced Colonoscopy Image Quality (ECIQ) method to automatically evaluate the perceptual quality of ECIs via analysis of brightness, contrast, colorfulness, naturalness, and noise. Extensive experiments on ECIQAD demonstrate the superiority of our proposed ECIQ method over 14 mainstream no-reference image quality assessment methods. Guanghui Yue 0001, Tianwei Zhou, Jingwen Hou, Weide Liu, Long Xu 0001, Tianfu Wang 0001, Jun Cheng 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Benchmarking Polyp Segmentation Methods in Narrow-Band Imaging Colonoscopy ImagesabstractIn recent years, there has been significant progress in polyp segmentation in white-light imaging (WLI) colonoscopy images, particularly with methods based on deep learning (DL). However, little attention has been paid to the reliability of these methods in narrow-band imaging (NBI) data. NBI improves visibility of blood vessels and helps physicians observe complex polyps more easily than WLI, but NBI images often include polyps with small/flat appearances, background interference, and camouflage properties, making polyp segmentation a challenging task. This paper proposes a new polyp segmentation dataset (PS-NBI2K) consisting of 2,000 NBI colonoscopy images with pixel-wise annotations, and presents benchmarking results and analyses for 24 recently reported DL-based polyp segmentation methods on PS-NBI2K. The results show that existing methods struggle to locate polyps with smaller sizes and stronger interference, and that extracting both local and global features improves performance. There is also a trade-off between effectiveness and efficiency, and most methods cannot achieve the best results in both areas simultaneously. This work highlights potential directions for designing DL-based polyp segmentation methods in NBI colonoscopy images, and the release of PS-NBI2K aims to drive further development in this field. Guanghui Yue 0001, Guibin Zhuo, Tianwei Zhou, Jingfeng Du, Weiqing Yan, Jingwen Hou, Weide Liu, Tianfu Wang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2023 | Toward Multicenter Skin Lesion Classification Using Deep Neural Network With Adaptively Weighted Balance LossabstractRecently, deep neural network-based methods have shown promising advantages in accurately recognizing skin lesions from dermoscopic images. However, most existing works focus more on improving the network framework for better feature representation but ignore the data imbalance issue, limiting their flexibility and accuracy across multiple scenarios in multi-center clinics. Generally, different clinical centers have different data distributions, which presents challenging requirements for the network's flexibility and accuracy. In this paper, we divert the attention from framework improvement to the data imbalance issue and propose a new solution for multi-center skin lesion classification by introducing a novel adaptively weighted balance (AWB) loss to the conventional classification network. Benefiting from AWB, the proposed solution has the following advantages: 1) it is easy to satisfy different practical requirements by only changing the backbone; 2) it is user-friendly with no tuning on hyperparameters; and 3) it adaptively enables small intraclass compactness and pays more attention to the minority class. Extensive experiments demonstrate that, compared with solutions equipped with state-of-the-art loss functions, the proposed solution is more flexible and more competent for tackling the multi-center imbalanced skin lesion classification task with considerable performance on two benchmark datasets. In addition, the proposed solution is proved to be effective in handling the imbalanced gastrointestinal disease classification task and the imbalanced DR grading task. Code is available at https://github.com/Weipeishan2021. Guanghui Yue 0001, Peishan Wei, Tianwei Zhou, Qiuping Jiang, Weiqing Yan, Tianfu Wang 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2023 | Interaction-Matrix Based Personalized Image Aesthetics AssessmentabstractPersonalized image aesthetics assessment (IAA) aims to estimate aesthetic experiences subject to the preferences of individual users, contrary to generic IAA that estimates aesthetic experiences subject to average preferences. Most existing personalized IAA methods treat personalized aesthetic experiences as deviations from a generic aesthetic experience, and therefore, personalized IAA models are designed to build upon the prior knowledge on generic IAA. However, we propose that acquiring knowledge on generic IAA is not necessary for building a personalized IAA model. Instead of modeling personalized IAA on the basis of generic IAA, this work proposes to directly estimate personalized aesthetic experiences from the interactions between image contents and user preferences (i.e., preference-content interaction), where interaction-matrices representing preference-content interactions are constructed without needs for prior generic IAA knowledge. To this end, we construct interaction-matrices from content features constructed from pre-trained image classification features and latent preference features. To realize a robust interaction-matrix based personalized IAA model, we discuss in detail on different strategies for constructing interaction-matrices and estimating personalized aesthetic scores from the interaction-matrices. Besides the personalized IAA scenario, we further propose strategies to adapt the proposed personalized IAA model to different scenarios of generic IAA. Extensive experiments show that: 1) our method significantly outperforms 5 previous relevant personalized IAA methods on FLICKR-AES dataset, especially the methods that require generic IAA knowledge as the basis; 2) in terms of generic IAA, the proposed approach also outperforms 13 generic IAA methods on AVA dataset. Jingwen Hou, Weisi Lin, Guanghui Yue 0001, Weide Liu, Baoquan Zhao |
IEEE Trans. Multim. | 3 |
| 2023 | Semi-Supervised Authentically Distorted Image Quality Assessment With Consistency-Preserving Dual-Branch Convolutional Neural NetworkabstractRecently, convolutional neural networks (CNNs) have provided a favoured prospect for authentically distorted image quality assessment (IQA). For good performance, most existing CNN-based methods rely on a large amount of labeled data for training, which is time-consuming and cumbersome to collect. By simultaneously exploiting few labeled data and many unlabeled data, we make a pioneering attempt to propose a semi-supervised framework (termed SSLIQA) with consistency-preserving dual-branch CNN for authentically distorted IQA in this paper. The proposed SSLIQA introduces a consistency-preserving strategy and transfers two kinds of consistency knowledge from the teacher branch to the student branch. Concretely, SSLIQA utilizes the sample prediction consistency to train the student to mimic output activations of individual examples represented by the teacher. Considering that subjects often refer to previous analogous cases to make scoring decisions, SSLIQA computes the semantic relation among different samples in a batch and encourages the consistency of sample semantic relation between two branches to explore extra quality-related information. Benefiting from the consistency-preserving strategy, we can exploit numerous unlabeled data to improve network's effectiveness and generalization. Experimental results on three authentically distorted IQA databases show that the proposed SSLIQA is stably effective under different student-teacher combinations and different labeled-to-unlabeled data ratios. In addition, it points out a new way on how to achieve higher performance with a smaller network. Guanghui Yue 0001, Leida Li, Tianwei Zhou, Hantao Liu, Tianfu Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Toward A No-reference Omnidirectional Image Quality Evaluation by Using Multi-perceptual FeaturesabstractCompared to ordinary images, omnidirectional image (OI) usually has a broader view and a higher resolution, and image quality assessment (IQA) can help people to understand and improve their visual experience. However, the current IQA works cannot achieve good performance. To address this, we proposed a novel visual perception-based no-reference/blind omnidirectional image quality assessment (NR/B-OIQA) model. The gradient-based global structural features and gray-level co-occurrence matrix-based local structural features are combined together to highlight the rich quality-aware structural information. And a novel steganalysis real model-based color descriptor is extracted to reflect the color information that ignored in most IQA models. With a multi-scale visual perception, we take image entropy and the natural scene statistics features to convey the high-level semantics and quantify the unnaturalness of omnidirectional images. Finally, we apply support vector regression to predict the objective quality value based on the subjective scores and extracted all features. Experiments are conducted on OIQA and CVIQD2018 Databases, and the results illustrate that our model has more reliable performance and stronger competitiveness and receives better conformity with the subjective values. Yun Liu 0009, Xiaohua Yin, Zuliang Wan, Guanghui Yue 0001, Zhi Zheng 0006 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | A genetic algorithm for backlight dimming for HDR displays
Lvyin Duan, Kurt Debattista, Guanghui Yue 0001, Demetris Marnerides, Alan Chalmers |
Vis. Comput. | 3 |
| 2022 | Improving IQA Performance Based on Deep Mutual LearningabstractIn this paper, we propose a novel solution, termed DML-IQA, for the image quality assessment (IQA) tasks. DML-IQA holds a dual-branch network architecture and builds the IQA model through a deep mutual learning (DML) strategy. Specifically, the two branches extract stable feature representations by feeding different transformed images into the classical CNNs. The DML strategy first calculates the prediction loss of each branch and the consistency loss across two branches, followed by updating the network iteratively to converge. Overall, DML-IQA has the following advantages: 1) It is flexible to adapt to diverse backbones for tackling the IQA issues in both the laboratory and wild; 2) It improves the baseline’s performance by approximately 1%~2%, especially performs well in the case of small samples. Extensive experiments on four public datasets show that the proposed DML-IQA can handle the IQA tasks with considerable effectiveness and generalization. Guanghui Yue 0001, Honglv Wu, Qiuping Jiang, Tianfu Wang 0001 |
ICIP | 1 |
| 2022 | Efficient Multiple Kernel Clustering via Spectral PerturbationabstractClustering is a fundamental task in the machine learning and data mining community. Among existing clustering methods, multiple kernel clustering (MKC) has been widely investigated due to its effectiveness to capture non-linear relationships among samples. However, most of the existing MKC methods bear intensive computational complexity in learning an optimal kernel and seeking the final clustering partition. In this paper, based on the spectral perturbation theory, we propose an efficient MKC method that reduces the computational complexity from O(n3) to O(nk2 + k3), with n and k denoting the number of data samples and the number of clusters, respectively. The proposed method recovers the optimal clustering partition from base partitions by maximizing the eigen gaps to approximate the perturbation errors. An equivalent optimization objective function is introduced to obtain base partitions. Furthermore, a kernel weighting scheme is embedded to capture the diversity among multiple kernels. Finally, the optimal partition, base partitions, and kernel weights are jointly learned in a unified framework. An efficient alternate iterative optimization algorithm is designed to solve the resultant optimization problem. Experimental results on various benchmark datasets demonstrate the superiority of the proposed method when compared to other state-of-the-art ones in terms of both clustering efficacy and efficiency. Chang Tang, Zhenglai Li, Weiqing Yan, Guanghui Yue 0001, Wei Zhang 0049 |
ACM Multimedia | 4 |
| 2022 | Bipartite Graph-based Discriminative Feature Learning for Multi-View ClusteringabstractMulti-view clustering is an important technique in machine learning research. Existing methods have improved in clustering performance, most of them learn graph structure depending on all samples, which are high complexity. Bipartite graph-based multi-view clustering can obtain clustering result by establishing the relationship between the sample points and small anchor points, which improve the efficiency of clustering. Most bipartite graph-based clustering methods only focus on topological graph structure learning depending on sample nodes, ignore the influence of node features. In this paper, we propose bipartite graph-based discriminative feature learning for multi-view clustering, which combines bipartite graph learning and discriminative feature learning to a unified framework. Specifically, the bipartite graph learning is proposed via multi-view subspace representation with manifold regularization terms. Meanwhile, our feature learning utilizes data pseudo-labels obtained by fused bipartite graph to seek projection direction, which make the same label be closer and make data points with different labels be far away from each other. At last, the proposed manifold regularization terms establish the relationship between constructed bipartite graph and new data representation. By leveraging the interactions between structure learning and discriminative feature learning, we are able to select more informative features and capture more accurate structure of data for clustering. Extensive experimental results on different scale datasets demonstrate our method achieves better or comparable clustering performance than the results of state-of-the-art methods. Weiqing Yan, Jinglei Liu, Guanghui Yue 0001, Chang Tang |
ACM Multimedia | 4 |
| 2022 | Multi-information Aggregation Network for Fundus Image Quality AssessmentabstractFundus image quality assessment (IQA) is essential for controlling the quality of retinal imaging and guaranteeing the reliability of diagnoses by ophthalmologists. Existing fundus IQA methods mainly explore local information to consider local distortions from convolutional neural networks (CNNs), yet ignoring global distortions. In this paper, we propose a novel multi-information aggregation network, termed MA-Net, for fundus IQA by extracting both local and global information. Specifically, MA-Net adopts an asymmetric dual-branch structure. For an input image, it uses the ResNet50 and vision transformer (ViT) to obtain the local and global representations from the upper and lower branches, respectively. In addition, MA-Net separately feed different images into the two branches to rank their quality for supplementing the feature representations. Thanks to the exploration of intra- and inter-class information between images, our MA-Net is competent for the fundus IQA task. Experiment results on the EyeQ dataset show that our MA-Net outperforms the baselines (i.e., ResNet50 and ViT) by 3.06% and 7.61% in Acc, and is superior to the mainstream methods. Guanghui Yue 0001, Lvyin Duan, Honglv Wu, Tianfu Wang 0001 |
VCIP | 2 |
| 2022 | Dynamic-boosting attention for self-supervised video representation learning
Chunping Hou, Guanghui Yue 0001 |
Appl. Intell. | 3 |
| 2022 | Quantization level based event-triggered control with measurement uncertainties
Tianwei Zhou, Guanghui Yue 0001, Ben Niu 0002 |
Inf. Sci. | 2 |
| 2022 | Two-stream interactive network based on local and global information for No-Reference Stereoscopic Image Quality Assessment
Yun Liu 0009, Baoqing Huang, Guanghui Yue 0001, Jingkai Wu, Zhi Zheng 0006 |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | Longitudinal study of early mild cognitive impairment via similarity-constrained group learning and self-attention based SBi-LSTM
Bai Ying Lei, Yanwu Xu 0001, Guanghui Yue 0001, Jiuwen Cao, Huoyou Hu, Shuangzhi Yu, Peng Yang 0011, Tianfu Wang 0001, Yali Qiu, Xiaohua Xiao, Shuqiang Wang |
Knowl. Based Syst. | 5 |
| 2022 | Unsupervised Domain Adaptation Based Image Synthesis and Feature Alignment for Joint Optic Disc and Cup SegmentationabstractDue to the discrepancy of different devices for fundus image collection, a well-trained neural network is usually unsuitable for another new dataset. To solve this problem, the unsupervised domain adaptation strategy attracts a lot of attentions. In this paper, we propose an unsupervised domain adaptation method based image synthesis and feature alignment (ISFA) method to segment optic disc and cup on fundus images. The GAN-based image synthesis (IS) mechanism along with the boundary information of optic disc and cup is utilized to generate target-like query images, which serves as the intermediate latent space between source domain and target domain images to alleviate the domain shift problem. Specifically, we use content and style feature alignment (CSFA) to ensure the feature consistency among source domain images, target-like query images and target domain images. The adversarial learning is used to extract domain-invariant features for output-level feature alignment (OLFA). To enhance the representation ability of domain-invariant boundary structure information, we introduce the edge attention module (EAM) for low-level feature maps. Eventually, we train our proposed method on the training set of the REFUGE challenge dataset and test it on Drishti-GS and RIM-ONE_r3 datasets. On the Drishti-GS dataset, our method achieves about 3% improvement of Dice on optic cup segmentation over the next best method. We comprehensively discuss the robustness of our method for small dataset domain adaptation. The experimental results also demonstrate the effectiveness of our method. Our code is available at https://github.com/thinkobj/ISFA. Haijun Lei, Weixin Liu 0002, Hai Xie, Benjian Zhao, Guanghui Yue 0001, Bai Ying Lei |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Boundary Constraint Network With Cross Layer Feature Integration for Polyp SegmentationabstractClinically, proper polyp localization in endoscopy images plays a vital role in the follow-up treatment (e.g., surgical planning). Deep convolutional neural networks (CNNs) provide a favoured prospect for automatic polyp segmentation and evade the limitations of visual inspection, e.g., subjectivity and overwork. However, most existing CNNs-based methods often provide unsatisfactory segmentation performance. In this paper, we propose a novel boundary constraint network, namely BCNet, for accurate polyp segmentation. The success of BCNet benefits from integrating cross-level context information and leveraging edge information. Specifically, to avoid the drawbacks caused by simple feature addition or concentration, BCNet applies a cross-layer feature integration strategy (CFIS) in fusing the features of the top-three highest layers, yielding a better performance. CFIS consists of three attention-driven cross-layer feature interaction modules (ACFIMs) and two global feature integration modules (GFIMs). ACFIM adaptively fuses the context information of the top-three highest layers via the self-attention mechanism instead of direct addition or concentration. GFIM integrates the fused information across layers with the guidance from global attention. To obtain accurate boundaries, BCNet introduces a bilateral boundary extraction module that explores the polyp and non-polyp information of the shallow layer collaboratively based on the high-level location information and boundary supervision. Through joint supervision of the polyp area and boundary, BCNet is able to get more accurate polyp masks. Experimental results on three public datasets show that the proposed BCNet outperforms seven state-of-the-art competing methods in terms of both effectiveness and generalization. Guanghui Yue 0001, Wanwan Han, Bin Jiang 0003, Tianwei Zhou, Runmin Cong, Tianfu Wang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | CEID: Benchmark Dataset for Designing Segmentation Algorithms of Instruments Used in Colorectal Endoscopy
Wanwan Han, Guanghui Yue 0001, Lvyin Duan, Jingfeng Du, Tianwei Zhou, Tianfu Wang 0001 |
ICIG (2) | 2 |
| 2021 | Semi-supervised Yolo Network for Induced Pluripotent Stem Cells Detection
Xinglie Wang, Jinqi Liao, Guanghui Yue 0001, Liangge He, Mingzhu Li, Enmin Liang, Tianfu Wang 0001, Guangqian Zhou, Bai Ying Lei |
ICIG (2) | 3 |
| 2021 | Differential Privacy for Industrial Internet of Things: Opportunities, Applications, and ChallengesabstractThe development of Internet of Things (IoT) brings new changes to various fields. Particularly, industrial IoT (IIoT) is promoting a new round of industrial revolution. With more applications of IIoT, privacy protection issues are emerging. Especially, some common algorithms in IIoT technology, such as deep models, strongly rely on data collection, which leads to the risk of privacy disclosure. Recently, differential privacy has been used to protect user-terminal privacy in IIoT, so it is necessary to make in-depth research on this topic. In this article, we conduct a comprehensive survey on the opportunities, applications, and challenges of differential privacy in IIoT. We first review related papers on IIoT and privacy protection, respectively. Then, we focus on the metrics of industrial data privacy, and analyze the contradiction between data utilization for deep models and individual privacy protection. Several valuable problems are summarized and new research ideas are put forward. In conclusion, this survey is dedicated to complete comprehensive summary and lay foundation for the follow-up research on industrial differential privacy. Bin Jiang 0003, Jianqiang Li 0001, Guanghui Yue 0001, Houbing Song |
IEEE Internet Things J. | 3 |
| 2021 | Global context guided hierarchically residual feature refinement network for defocus blur detection
Yongping Zhai, Jinsheng Deng, Guanghui Yue 0001, Wei Zhang 0049, Chang Tang |
Signal Process. | 4 |
| 2021 | No-Reference Image Contrast Evaluation by Generating Bidirectional PseudoreferencesabstractThis article proposes a simple yet reliable no-reference image contrast evaluator (NICE) by generating bidirectional pseudoreferences (BPR). Different from the existing no-reference metrics that only operate on the contrast distorted image (CDI) itself, our proposed NICE-BPR measures the deviations of a CDI to its corresponding aggravated and enhanced counterparts (i.e., BPRs) in a hybrid feature space. Given a CDI, we first perform contrast aggravation and contrast enhancement using gamma correction and histogram equalization, respectively. Then, hybrid contrast-aware features are, respectively, extracted from the CDI and its corresponding BPRs via the analysis of histogram, entropy, and structure. The features obtained from the CDI are one-by-one compared with those from the BPRs to derive the bidirectional feature deviation vector. Finally, a quality predictor is built by learning a regression model to fuse the feature vector into a continuous quality score. Extensive experiments on several databases well-demonstrate the superiority of NICE-BPR. Qiuping Jiang, Zhenyu Peng, Guanghui Yue 0001, Feng Shao 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Multi-scale Enhanced Graph Convolutional Network for Early Mild Cognitive Impairment Detection
Shuangzhi Yu, Shuqiang Wang, Xiaohua Xiao, Jiuwen Cao, Guanghui Yue 0001, Tianfu Wang 0001, Yanwu Xu 0001, Bai Ying Lei |
MICCAI (7) | 5 |
| 2020 | No-Reference Stereoscopic Image Quality Assessment Considering Multi-loss ConstraintsabstractIn this paper, a three-channel convolutional neural network (CNN) constrained by multiple loss functions is designed for stereoscopic image quality assessment (SIQA). Given that both monocular and binocular information are crucial for SIQA, we take the patches of left images, right images and difference images as the inputs of the three channels respectively. Since using the ground truth as the labels of image patches cannot accurately characterize their quality, we propose to individually label each image patch to preserve the quality difference among different regions and views. Moreover, the multi-loss structure is adopted in the proposed method to consider both local features and global features simultaneously, which can constrain the feature learning from multiple perspectives. And the additional adaptive loss weights make the multi-loss network more flexible and universal. The experimental results show that the proposed method is superior to other existing SIQA methods with state-of-the-art performance. Yongtian Han, Sumei Li, Guanghui Yue 0001, Yongli Chang |
VCIP | 3 |
| 2020 | Shape-optimizing mesh warping method for stereoscopic panorama stitching
Weiqing Yan, Guanghui Yue 0001, Yanwei Yu, Kai Wang 0014, Chang Tang, Xiangrong Tong |
Inf. Sci. | 2 |
| 2020 | AMD-GAN: Attention encoder and multi-branch structure based generative adversarial networks for fundus disease detection from scanning laser ophthalmoscopy images
Hai Xie, Haijun Lei, Xianlu Zeng, Yejun He, Guozhen Chen, Ahmed El-Azab, Guanghui Yue 0001, Bai Ying Lei |
Neural Networks | 7 |
| 2020 | Perceptual objective quality assessment of stereoscopic stitched images
Weiqing Yan, Guanghui Yue 0001, Yuming Fang 0001, Hua Chen 0004, Chang Tang, Gangyi Jiang |
Signal Process. | 2 |
| 2020 | Referenceless Quality Evaluation of Tone-Mapped HDR and Multiexposure Fused ImagesabstractNowadays, the standard dynamic range (SDR) image acquired at a fixed exposure exposes weakness in portraying fine-grained details of real scenes. The high dynamic range (HDR) image and other types of SDR images generated by multiexposure fusion techniques provide us new choices for scene representation. To display on SDR screens, an HDR image must be tone-mapped to an SDR one. Since different tone-mapping/fusion algorithms produce images with varying visual quality levels, it naturally desires a quality evaluation model for comparison. This article proposes an effective model in the absence of the reference image. By analyzing the characteristics of tone-mapped HDR and multiexposure fused images, we first extract multiple quality-sensitive features from the following aspects: 1) colorfulness; 2) exposure; and 3) naturalness. Then, the model is built by bridging all extracted features and associated subjective ratings via support vector regression. Extensive experiments on publicly available databases prove the superiority of our model over the state-of-the-art referenceless quality evaluation ones. Guanghui Yue 0001, Weiqing Yan, Tianwei Zhou |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | Blind Quality Evaluator for Screen Content Images via Analysis of StructureabstractExisting blind evaluators for screen content images (SCIs) are mainly learning-based and require a number of training images with co-registered human opinion scores. However, the size of existing databases is small, and it is labor-, time-consuming and expensive to largely generate human opinion scores. In this study, we propose a novel blind quality evaluator without training. Specifically, the proposed method first calculates the gradient similarity between a distorted image and its translated versions in four directions to estimate the structural distortion, the most obvious distortion in SCIs. Given that the edge region is easier to be distorted, the inter-scale gradient similarity is then calculated as the weighting map. Finally, the proposed method is derived by incorporating the gradient similarity map with the weighting map. Experimental results demonstrate its effectiveness and efficiency on a public available SCI database. Guanghui Yue 0001, Chunping Hou, Weisi Lin |
ICASSP | 1 |
| 2019 | Stereoscopic Video Quality Assessment Based on The Two-step-training Binocular Fusion NetworkabstractIn this paper, we propose a novel binocular fusion network for stereoscopic video quality assessment (SVQA). In this network, we construct a long-term fusion, competition, and processing process by simulating the long-term complex process of the whole visual pathway. And we employ a two-step-training strategy for this network, which solves the problem that the network is difficult to fit caused by using the same value to label the different quality regions and views of the same stereoscopic video. In the first step, we use the computed quality scores of different patches to train the local network, namely local regression. And then the global regression is performed by using MOS value based on the first step trained model. Besides, considering temporal information, we take spatiotemporal saliency feature flows as the inputs of the proposed network. The proposed method is tested on public stereoscopic video databases, and results show that our method outperforms any other methods. Sumei Li, Jianwei Xue, Yixiu Ding, Guanghui Yue 0001 |
VCIP | 5 |
| 2019 | Combining Local and Global Measures for DIBR-Synthesized Image Quality EvaluationabstractDepth-Image-Based-Rendering (DIBR) techniques are significant for three-dimensional (3D) video applications, e.g., 3D television and free viewpoint video (FVV). Unfortunately, the DIBR-synthesized image suffers from various distortions, which induce an annoying viewing experience for the entire FVV. Proposing a quality evaluator for DIBR-synthesized images is fundamental for the design of perceptual friendly FVV systems. Since the associated reference image is usually not accessible, full-reference (FR) methods cannot be directly applied for quality evaluation of the synthesized image. In addition, most traditional no-reference (NR) methods fail to effectively measure the specifically DIBR-related distortions. In this paper, we propose a novel NR quality evaluation method accounting for two categories of DIBR-related distortions, i.e., geometric distortions and sharpness. First, the disoccluded regions, as one of the most obvious geometric distortions, are captured by analyzing local similarity. Then, another typical geometric distortion (i.e., stretching) is detected and measured by calculating the similarity between it and its equal-size adjacent region. Second, considering the property of scale invariance, the global sharpness is measured as the distance between the distorted image and its downsampled version. Finally, the perceptual quality is estimated by linearly pooling the scores of two geometric distortions and sharpness together. Experimental results verify the superiority of the proposed method over the prevailing FR and NR metrics. More specifically, it is superior to all competing methods except APT in terms of effectiveness, but greatly outmatches APT in terms of implementation time. Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Tianwei Zhou, Guangtao Zhai |
IEEE Trans. Image Process. | 1 |
| 2019 | No-Reference Quality Evaluator of Transparently Encrypted ImagesabstractIn past years, various encrypted algorithms have been proposed to fully or partially protect the multimedia content in view of practical applications. In the context of digital TV broadcasting, transparent encryption only protects partial content and fulfills both security and quality requirements. To date, only a few reference-based works have been reported to evaluate the quality of transparently encrypted images. However, these works are incapable of reference-unavailable conditions. In this paper, we conduct the first attempt that proposes a novel quality evaluator in the absence of reference images. The key strategy of the proposed metric lies in extracting features by considering the motivation of transparently encrypted images. Specifically, given that encrypted images prevent content from being easily recognized, several features, including correlation coefficient, information entropy, and intensity statistic, are preliminarily extracted to estimate visual recognizability. Meanwhile, considering that encrypted images are avoided since they are of extremely low quality, we also capture many features to measure the distortions on multiple quality-sensitive image attributes, such as naturalness, structure, and texture. Finally, the quality evaluator is built by bridging all extracted features and corresponding quality scores via a regression module. Experimental results demonstrate that the proposed method is superior to the mainstream no-reference quality evaluation methods designed for synthetically distorted images and possesses a close approximation to state-of-the-art reference-based methods designed for encrypted images. Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Tianwei Zhou, Hantao Liu |
IEEE Trans. Multim. | 1 |
| 2019 | Subtitle Region Selection of S3D Images in Consideration of Visual Discomfort and Viewing HabitabstractSubtitles, serving as a linguistic approximation of the visual content, are an essential element in stereoscopic advertisement and the film industry. Due to the vergence accommodation conflict, the stereoscopic 3D (S3D) subtitle inevitably causes visual discomfort. To meet the viewing experience, the subtitle region should be carefully arranged. Unfortunately, very few works have been dedicated to this area. In this article, we propose a method for S3D subtitle region selection in consideration of visual discomfort and viewing habit. First, we divide the disparity map into multiple depth layers according to the disparity value. The preferential processed depth layer is determined by considering the disparity value of the foremost object. Second, the optimal region and coarse disparity value for S3D subtitle insertion are chosen by convolving the selective depth layer with the mean filter. Specifically, the viewing habit is considered during the region selection. Finally, after region selection, the disparity value of the subtitle is further modified by using the just noticeable depth difference (JNDD) model. Given that there is no public database reported for the evaluation of S3D subtitle insertion, we collect 120 S3D images as the test platform. Both objective and subjective experiments are conducted to evaluate the comfort degree of the inserted subtitle. Experimental results demonstrate that the proposed method can obtain promising performance in improving the viewing experience of the inserted subtitle. Guanghui Yue 0001, Chunping Hou, Tianwei Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | A two-channel convolutional neural network for image super-resolution
Sumei Li, Ru Fan, Guoqing Lei, Guanghui Yue 0001, Chunping Hou |
Neurocomputing | 4 |
| 2018 | Blind stereoscopic 3D image quality assessment via analysis of naturalness, structure, and binocular asymmetry
Guanghui Yue 0001, Chunping Hou, Qiuping Jiang, Yang Yang 0045 |
Signal Process. | 1 |
| 2018 | Optimal Region Selection for Stereoscopic Video Subtitle InsertionabstractStereoscopic subtitle insertion is a fundamental and essential element in stereoscopic film and TV industry. However, little work has been dedicated to the optimal region selection for stereoscopic subtitle insertion. In addition, there is no public database reported for the performance evaluation of it. In this paper, we build the first large-scale video database (TJU3D) for stereoscopic video subtitle insertion, which includes 50 video sequences with rich screen scenes. Compared with 2D subtitle region selection, there are several problems we have to consider in stereoscopic subtitle region selection: 1) the subtitle should avoid depth cue collision and occlusion from objects in stereoscopic video sequences; 2) the disparity value of the subtitle must be minimized to reduce visual discomfort; and 3) the temporal coherence constraint must be considered during region selection for subtitles in video sequences. By considering these constraints, we propose an optimal region selection algorithm for stereoscopic subtitle insertion. First, we compute the disparity map of each video frame in video sequences. For each frame, the optimal position and disparity value of the subtitle are determined by a subtitle region selection algorithm, which contains two parts (i.e., the coarse selection and fine selection). After that, by considering the temporal consistency between adjacent frames, the position and disparity value of each frame are further classified and processed in order to avoid the subtitle jitter. We evaluate the proposed method on TJU3D video database through two visual discomfort prediction metrics and one subjective experiment. To further verify the effectiveness of the proposed method, we also validate the performance of the proposed method on video comfort assessment database, i.e., IEEE-SA Stereo Database. Experimental results demonstrate that the visual discomfort is greatly reduced when using the proposed method compared with the basic method. Guanghui Yue 0001, Chunping Hou, Jianjun Lei 0001, Yuming Fang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Analysis of Structural Characteristics for Quality Assessment of Multiply Distorted ImagesabstractPerceptual image quality assessment (IQA) plays an important role in numerous applications, including image restoration, compression, enhancement, and others. Although many works have been conducted on individually distorted IQA problems and have achieved encouraging results, few studies have been conducted on multiple distorted (MD) IQA problems. Thus, limited progress has been made. In this paper, we propose a novel no reference image quality assessment (NR-IQA) method, named improved multiscale local binary pattern (IMLBP), for addressing multiply distorted IQA problems. The image structures are sensitive to image distortions, which motivates us to utilize the structural characteristics for overall image quality prediction. We improved the local binary pattern (LBP) by considering the human visual mechanism to better extract the structural information. The IMLBP contains two parts, the LBP and the radius difference LBP (DLBP). The DLBP reflects the values' changes in the radial direction. Specifically, when the radius value is small, the proposed descriptor is computed to represent microstructural information. Conversely, it represents macrostructural information when the radius becomes large. Moreover, to better mimick the human visual mechanism, the IMLBP is computed with the multiscale strategy and the operation is based on a patch unit whose size is proportional to the radius value. The frequency histogram of feature maps is transformed to feature vectors. Subsequently, a predictable function trained by the support vector regression is used to infer the overall quality score. Experimental results show that the proposed method outperforms most state-of-the-art IQA metrics on publicly available multiply distorted image databases. Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Nam Ling, Beichen Li 0002 |
IEEE Trans. Multim. | 1 |
| 2018 | Evaluating Quality of Screen Content Images Via Structural Variation AnalysisabstractWith the quick development and popularity of computers, computer-generated signals have drastically invaded into our daily lives. Screen content image is a typical example, since it also includes graphic and textual images as components as compared with natural scene images which have been deeply explored, and thus screen content image has posed novel challenges to current researches, such as compression, transmission, display, quality assessment, and more. In this paper, we focus our attention on evaluating the quality of screen content images based on the analysis of structural variation, which is caused by compression, transmission, and more. We classify structures into global and local structures, which correspond to basic and detailed perceptions of humans, respectively. The characteristics of graphic and textual images, e.g., limited color variations, and the human visual system are taken into consideration. Based on these concerns, we systematically combine the measurements of variations in the above-stated two types of structures to yield the final quality estimation of screen content images. Thorough experiments are conducted on three screen content image quality databases, in which the images are corrupted during capturing, compression, transmission, etc. Results demonstrate the superiority of our proposed quality model as compared with state-of-the-art relevant methods. Ke Gu 0001, Junfei Qiao 0001, Xiongkuo Min, Guanghui Yue 0001, Weisi Lin, Daniel Thalmann |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | Shallow and deep convolutional networks for image super-resolutionabstractA shallow and deep convolutional neural network is presented for the single-image super-resolution (SISR). The proposed method doesn't need hand-designed procedures, directly learning an end-to-end mapping between low-resolution (LR) and high-resolution (HR) images. The upsampling of the network by deconvolution leads to much more efficient and effective training, reducing the computational complexity of the overall SR operation. However, most existing methods based on CNNs for super resolution need preprocessing like bicubic interpolating LR images to the size of HR images. This method can restore more details by multi-scale manner, and has strong adaptability whether on images or videos. Our model is evaluated on different datasets, outperforming the existing methods in accuracy and visual impression. Ru Fan, Sumei Li, Guoqing Lei, Guanghui Yue 0001 |
ICIP | 4 |
| 2017 | Subjective quality assessment of animation imagesabstractIn the past few decades, many attempts have been maken to evaluate the image quality assessment (IQA) of natural scene images. However, the IQA research of animation images (AIs) has been highly overlooked. In this article, we carry out in-depth study on perceptual quality assessment of AIs. As the lack of a public and diverse testing database currently, this paper builds a large-scale Animation Images Quality Assessment Database (AIQAD). This database totally includes 1050 distorted images derived from 30 source images by corrupting seven distortion types with multiple distortion levels. Then, a subjective experiment, which is the basic and accurate quality evaluation measurement, is conducted to obtain the mean opinion score (MOS) for each image. Furthermore, we also investigate the feasibility of utilizing existing mainstream full reference (FR) IQA metrics to solve the IQA problem of AIs. Experimental results demonstrate that existing mainstream FR IQA metrics merely achieve fair performance on the proposed database. Guanghui Yue 0001, Chunping Hou, Ke Gu 0001 |
VCIP | 1 |
| 2017 | No reference image blurriness assessment with local binary patterns
Guanghui Yue 0001, Chunping Hou, Ke Gu 0001, Nam Ling |
J. Vis. Commun. Image Represent. | 1 |