Xuhang Chen 0002

dblp:308/1093-2 · DBLP profile ↗
← Back
49ranked-venue papers
3as first author
49since 2021 · last 2026
0000-0001-6000-3914ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 17 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DDG-MoE: A Dual-Driven Mixture-of-Experts Framework for Boundary-Ambiguous Radiological Pattern Disentanglement
Siyue Xie, Chun Peng, Qingyu Zhuang, Guoheng Huang, Xuhang Chen 0002
ICIC (1)5
2026 AHRS-Net: Anatomical Hierarchical Refining Semantic Network for Anatomy-Aware Spatially Grounded Radiology Report Generation
Siyue Xie, Kaijun Shen, Qingyu Zhuang, Guoheng Huang, Xuhang Chen 0002
ICIC (20)5
2026 ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis
abstract
High-fidelity character voice synthesis is a cornerstone of immersive multimedia applications, particularly for interacting with anime avatars and digital humans. However, existing systems struggle to maintain consistent persona traits across diverse emotional contexts. To bridge this gap, we present ATRIE, a unified framework utilizing a Persona-Prosody Dual-Track (P2-DT) architecture. Our system disentangles generation into a static Timbre Track (via Scalar Quantization) and a dynamic Prosody Track (via Hierarchical Flow-Matching), distilled from a 14B LLM teacher. This design enables robust identity preservation (Zero-Shot Speaker Verification EER: 0.04) and rich emotional expression. Evaluated on our extended AnimeTTS-Bench (50 characters), ATRIE achieves state-of-the-art performance in both generation and cross-modal retrieval (mAP: 0.75), establishing a new paradigm for persona-driven multimedia content creation. The code is available at Github.
Aoduo Li, Hongjian Xu, Shengmin Li, Sihao Qin, Zimeng Li 0001, Chi-Man Pun, Xuhang Chen 0002
ICMR8
2026 QWNet: A quaternion wavelet network for spatial-frequency aware multi-modal image fusion
Jietao Yang, Miaoshan Lin, Guoheng Huang, Xuhang Chen 0002, Xiaofeng Zhang 0006, Xiaochen Yuan, Chi-Man Pun, Bingo Wing-Kuen Ling
Neural Networks4
2026 Fre-QNet: Quaternion Progressive Perception Mechanism with Frequency-Guided Prompt for Blind Image Quality Assessment
abstract
Blind Image Quality Assessment faces challenges in enabling computational models to mimic the hierarchical progressive perception mechanisms of the Human Visual System (HVS). Existing methods often neglect the two-stage process of HVS—global distortion identification followed by local quality evaluation—and its distinct sensitivity to distortion types. To address this, we propose Fre-QNet, a novel framework integrating two key components: (1) A Quaternion Progressive Perception (QPP) module that hierarchically extracts multi-scale spatial features using quaternion convolution, explicitly simulating the global-to-local observation process of HVS while enhancing cross-scale interactions; (2) A Frequency Prompting (FP) module that quantifies distortion types and severity in the Fourier domain by leveraging frequency patterns of common distortions and the sensitivity variations of HVS. The QPP and FP modules collaboratively embed biological vision principles into computational modeling through dual-domain feature learning, with the QPP module directly anchoring the core logic of progressive perception. Experiments on TID2013 and CSIQ benchmarks demonstrate Fre-QNet’s superiority over state-of-the-art methods, validating its effectiveness in matching human perceptual quality judgments. Our source code is available at: https://github.com/hhsda/Fre-QNet .
Shize Li, Guoheng Huang, Yisen Zheng, Xiaochen Yuan, Xuhang Chen 0002, Lianglun Cheng, Chi-Man Pun
ACM Trans. Multim. Comput. Commun. Appl.6
2026 WAQNIQA: Wavelet-Augmented Quaternion Network for No-Reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) plays a pivotal role in computer vision by enabling image quality evaluation without reference images. While recent CNN and Transformer-based methods have advanced feature extraction, they face significant limitations. CNNs exhibit local feature bias, limiting their ability to capture global dependencies and complex structures critical to understanding diverse distortions. Transformers, despite modeling nonlocal dependencies through multihead attention, suffer from quadratic computational complexity with spatial dimensions, hindering efficient multiscale analysis. Moreover, their attention mechanisms frequently overlook critical interchannel dependencies, which are vital for capturing fine details in texture-rich images. Coupled with difficulties in handling high-noise environments and complex textures, this results in limited real-world accuracy and poor generalization across diverse datasets and unknown distortions. To bridge these gaps, we propose WAQNIQA, a novel wavelet-augmented quaternion network for NR-IQA. Distinct from conventional architectures, WAQNIQA integrates two synergistic modules: the wavelet-infused adaptive attention (WIAA) module, which leverages wavelet transforms (WTs) to achieve robust multiscale spatial-frequency analysis with linear complexity, and the quaternion collaborative feature enhancement (QCFE) module, which holistically models interchannel correlations to preserve fine texture details. Furthermore, we introduce PowerGridIQ, the first NR-IQA dataset specifically tailored for power grid scenarios. Extensive experiments demonstrate that WAQNIQA consistently surpasses state-of-the-art CNN and Transformer-based methods on PowerGridIQ and six public benchmarks. Notably, WAQNIQA exhibits superior cross-domain generalization, achieving competitive performance on the AGIQA-1K dataset for AI-generated content (AIGC) without explicit semantic alignment training, thereby validating its robustness against diverse and unknown distortions. Our code is available athttps://github.com/king-huoye/WAQNIQA
Yejing Huo, Guoheng Huang, Zhiwen Yu 0002, Xiaochen Yuan, Chi-Man Pun, Lianglun Cheng, Xuhang Chen 0002, Zehong Chen
IEEE Trans. Syst. Man Cybern. Syst.8
2025 FS-RWKV: Leveraging Frequency Spatial-Aware RWKV for 3T-to-7T MRI Translation
abstract
Ultra-high-field 7T MRI offers enhanced spatial resolution and tissue contrast that enable the detection of subtle pathological changes in neurological disorders. However, the limited availability of 7T scanners restricts widespread clinical adoption due to substantial infrastructure costs and technical demands. Computational approaches for synthesizing 7T-quality images from accessible 3T acquisitions present a viable solution to this accessibility challenge. Existing CNN approaches suffer from limited spatial coverage, while Transformer models demand excessive computational overhead. RWKV architectures offer an efficient alternative for global feature modeling in medical image synthesis, combining linear computational complexity with strong long-range dependency capture. Building on this foundation, we propose Frequency Spatial-RWKV (FS-RWKV), an RWKV-based framework for 3T-to-7T MRI translation. To better address the challenges of anatomical detail preservation and global tissue contrast recovery, FS- RWKV incorporates two key modules: (1) Frequency-Spatial Omnidirectional Shift (FSO-Shift), which performs discrete wavelet decomposition followed by omnidirectional spatial shifting on the low-frequency branch to enhance global contextual representation while preserving high-frequency anatomical details; and (2) Structural Fidelity Enhancement Block (SFEB), a module that adaptively reinforces anatomical structure through frequency-aware feature fusion. Comprehensive experiments on UNC and BNU datasets demonstrate that FS- RWKV consistently outperforms existing CNN-, Transformer-, GAN-, and RWKV-based baselines across both T1 wand T2w modalities, achieving superior anatomical fidelity and perceptual quality.
Yingtie Lei, Zimeng Li 0001, Yupeng Liu 0003, Xuhang Chen 0002, Chi-Man Pun
BIBM4
2025 DTEA: Dynamic Topology Weaving and Instability-Driven Entropic Attenuation for Medical Image Segmentation
abstract
In medical image segmentation, skip connections are used to merge global context and reduce the semantic gap between encoder and decoder. Current methods often struggle with limited structural representation and insufficient contextual modeling, affecting generalization in complex clinical scenarios. We propose the DTEA model, featuring a new skip connection framework with the Semantic Topology Reconfiguration (STR) and Entropic Perturbation Gating (EPG) modules. STR reorganizes multi-scale semantic features into a dynamic hypergraph to better model cross-resolution anatomical dependencies, enhancing structural and semantic representation. EPG assesses channel stability after perturbation and filters high-entropy channels to emphasize clinically important regions and improve spatial attention. Extensive experiments on three benchmark datasets show our framework achieves superior segmentation accuracy and better generalization across various clinical settings. The code is available at https://github.com/LWX-Research/DTEA.
Quanjun Li, Zimeng Li 0001, Chi-Man Pun, Yupeng Liu 0003, Xuhang Chen 0002
BIBM8
2025 HAFT: Hierarchical Attentional Fusion Transformer for Adaptive Feature Fusion in Medical Image Segmentation
abstract
Medical image segmentation remains fundamentally challenged by the complexity of anatomical structures and severe class imbalance. Existing methods often fall short in multi-scale feature integration and long-tailed distribution modeling, limiting their ability to simultaneously capture fine-grained local structures and holistic contextual semantics. To overcome these issues, we propose HAFT, a hierarchical attentional fusion transformer for adaptive feature fusion in medical image segmentation. We further introduce a Bayesian Adaptive Loss (BAL), which incorporates Bayesian uncertainty modeling to effectively alleviate the long-tailed distribution prevalent in medical datasets. Extensive experiments on multiple public benchmarks demonstrate that our method consistently outperforms existing approaches, particularly in segmenting intricate anatomical regions and rare pathological lesions. The code is available at https://github.com/QuincyQAQ/HAFT.
Quanjun Li, Fuchen Zheng, Junhua Zhou, Changwei Gong, Zimeng Li 0001, Yihua Shao, Xuhang Chen 0002
BIBM9
2025 Elevating Medical Image Security: A Cryptographic Framework Integrating Hyperchaotic Map and GRU
abstract
Chaotic systems play a key role in modern image encryption due to their sensitivity to initial conditions, ergodicity, and complex dynamics. However, many existing chaos-based encryption methods suffer from vulnerabilities, such as inade-quate permutation and diffusion, and suboptimal pseudorandom properties. To address these issues, this paper presents the Knot-like Unique Novel-Scan Image Encryption (Kun-IE). The framework comprises two main components: The 2D Sin–Cos Pi Hyperchaotic Map (2D-SCPHM), which provides a broader chaotic range and superior pseudorandom sequence generation, and the Knot-like Unique Novel-Scan Algorithm (Kun-SCAN), a novel permutation mechanism that markedly reduces pixel correlations and enhances resistance to statistical attacks. Kun-IE is flexible and supports encryption for images of any size. Experimental results and security analyses demonstrate its robustness against various cryptanalytic attacks, making it a strong solution for secure image communication. The code is available at this link.
Quanjun Li, Junhua Zhou, Yihang Dong, Mengqian Wang, Zimeng Li 0001, Changwei Gong, Xuhang Chen 0002
BIBM11
2025 PC-UNet: An Enforcing Poisson Statistics U-Net for Positron Emission Tomography Denoising
Jingchao Wang 0002, Liangsi Lu, Mingxuan Huang, Ruixin He, Yifeng Xie, Hanqian Liu, Minzhe Guo, Yangyang Liang, Zimeng Li 0001, Xuhang Chen 0002
BIBM12
2025 EEMS: Edge-Prompt Enhanced Medical Image Segmentation Based on Learnable Gating Mechanism
abstract
Medical image segmentation is vital for diagnosis, treatment planning, and disease monitoring but is challenged by complex factors like ambiguous edges and background noise. We introduce EEMS, a new model for segmentation, combining an Edge-Aware Enhancement Unit (EAEU) and a Multi-scale Prompt Generation Unit (MSPGU). EAEU enhances edge perception via multi-frequency feature extraction, accurately defining boundaries. MSPGU integrates high-level semantic and low-level spatial features using a prompt-guided approach, ensuring precise target localization. The Dual-Source Adaptive Gated Fusion Unit (DAGFU) merges edge features from EAEU with semantic features from MSPGU, enhancing segmentation accuracy and robustness. Tests on datasets like ISIC2018 confirm EEMS's superior performance and reliability as a clinical tool.
Quanjun Li, Zimeng Li 0001, Hongbin Ye, Yupeng Liu 0003, Haolun Li 0001, Xuhang Chen 0002
BIBM8
2025 HBFormer: A Hybrid-Bridge Transformer for Microtumor and Miniature Organ Segmentation
abstract
Medical image segmentation is a cornerstone of modern clinical diagnostics. While Vision Transformers that leverage shifted window-based self-attention have established new benchmarks in this field, they are often hampered by a critical limitation: their localized attention mechanism struggles to effectively fuse local details with global context. This deficiency is particularly detrimental to challenging tasks such as the segmentation of microtumors and miniature organs, where both finegrained boundary definition and broad contextual understanding are paramount. To address this gap, we propose HBFormer, a novel Hybrid-Bridge Transformer architecture. The 'Hybrid' design of HBFormer synergizes a classic U-shaped encoder-decoder framework with a powerful Swin Transformer backbone for robust hierarchical feature extraction. The core innovation lies in its 'Bridge' mechanism, a sophisticated nexus for multi-scale feature integration. This bridge is architecturally embodied by our novel Multi-Scale Feature Fusion (MFF) decoder. Departing from conventional symmetric designs, the MFF decoder is engineered to fuse multi-scale features from the encoder with global contextual information. It achieves this through a synergistic combination of channel and spatial attention modules, which are constructed from a series of dilated and depth-wise convolutions. These components work in concert to create a powerful feature bridge that explicitly captures long-range dependencies and refines object boundaries with exceptional precision. Comprehensive experiments on challenging medical image segmentation datasets, including multi-organ, liver tumor, and bladder tumor benchmarks, demonstrate that HBFormer achieves state-of-theart results, showcasing its outstanding capabilities in microtumor and miniature organ segmentation. Code and models are available at: https://github.com/lzeeorno/HBFormer.
Fuchen Zheng, Quanjun Li, Junhua Zhou, Xiaojiao Guo, Xuhang Chen 0002, Chi-Man Pun, Shoujun Zhou
BIBM7
2025 TDADL-IE: A Deep Learning-Driven Cryptographic Architecture for Medical Image Security
abstract
The rise of digital medical imaging, like MRI and CT, demands strong encryption to protect patient data in telemedicine and cloud storage. Chaotic systems are popular for image encryption due to their sensitivity and unique characteristics, but existing methods often lack sufficient security. This paper presents the Three-dimensional Diffusion Algorithm and Deep Learning Image Encryption system (TDADL-IE), built on three key elements. First, we propose an enhanced chaotic generator using an LSTM network with a 1D-Sine Quadratic Chaotic Map (1D-SQCM) for better pseudorandom sequence generation. Next, a new three-dimensional diffusion algorithm (TDA) is applied to encrypt permuted images. TDADL-IE is versatile for images of any size. Experiments confirm its effectiveness against various security threats. The code is available at https://github.com/QuincyQAQ/TDADL-IE.
Junhua Zhou, Quanjun Li, Yihua Shao, Yihang Dong, Mengqian Wang, Zimeng Li 0001, Changwei Gong, Xuhang Chen 0002
BIBM10
2025 Enhancing the Transferability of Adversarial Examples Against No-Reference Image Quality Assessment Models
Yijia Zeng, Xuhang Chen 0002, Chi-Man Pun
CGI (3)2
2025 SWAN: A Synergistic Wavelet Attention Network for Enhanced Underwater Image Enhancement
Junhua Zhou, Quanjun Li, Yihang Dong, Zimeng Li 0001, Zexiao Liang, Xuhang Chen 0002
CGI (3)7
2025 Underwater Image Restoration via Polymorphic Large Kernel CNNs
abstract
Underwater Image Restoration (UIR) remains a challenging task in computer vision due to the complex degradation of images in underwater environments. While recent approaches have leveraged various deep learning techniques, including Transformers and complex, parameter-heavy models to achieve significant improvements in restoration effects, we demonstrate that pure CNN architectures with lightweight parameters can achieve comparable results. In this paper, we introduce UIR-PolyKernel, a novel method for underwater image restoration that leverages Polymorphic Large Kernel CNNs. Our approach uniquely combines large kernel convolutions of diverse sizes and shapes to effectively capture long-range dependencies within underwater imagery. Additionally, we introduce a Hybrid Domain Attention module that integrates frequency and spatial domain attention mechanisms to enhance feature importance. By leveraging the frequency domain, we can capture hidden features that may not be perceptible to humans but are crucial for identifying patterns in both underwater and on-air images. This approach enhances the generalization and robustness of our UIR model. Extensive experiments on benchmark datasets demonstrate that UIR-PolyKernel achieves state-of-the-art performance in underwater image restoration tasks, both quantitatively and qualitatively. Our results show that well-designed pure CNN architectures can effectively compete with more complex models, offering a balance between performance and computational efficiency. This work provides new insights into the potential of CNN-based approaches for challenging image restoration tasks in underwater environments. The code is available at https://github.com/CXH-Research/UIR-PolyKernel.
Xiaojiao Guo, Yihang Dong, Xuhang Chen 0002, Weiwen Chen, Zimeng Li 0001, Fuchen Zheng, Chi-Man Pun
ICASSP3
2025 Superpixel-Enhanced Quaternion Feature Fusion and Contextualization Graph Contrastive Learning for Cervical Cancer Diagnosis
Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Lianglun Cheng, Chi-Man Pun, Guo Zhong, Qingjian Ye
ICONIP (2)4
2025 LensNet: An End-to-End Learning Framework for Empirical Point Spread Function Modeling and Lensless Imaging Reconstruction
abstract
Lensless imaging stands out as a promising alternative to conventional lens-based systems, particularly in scenarios demanding ultracompact form factors and cost-effective architectures. However, such systems are fundamentally governed by the Point Spread Function (PSF), which dictates how a point source contributes to the final captured signal. Traditional lensless techniques often require explicit calibrations and extensive pre-processing, relying on static or approximate PSF models. These rigid strategies can result in limited adaptability to real-world challenges, including noise, system imperfections, and dynamic scene variations, thus impeding high-fidelity reconstruction. In this paper, we propose LensNet, an end-to-end deep learning framework that integrates spatial-domain and frequency-domain representations in a unified pipeline. Central to our approach is a learnable Coded Mask Simulator (CMS) that enables dynamic, data-driven estimation of the PSF during training, effectively mitigating the shortcomings of fixed or sparsely calibrated kernels. By embedding a Wiener filtering component, LensNet refines global structure and restores fine-scale details, thus alleviating the dependency on multiple handcrafted pre-processing steps. Extensive experiments demonstrate LensNet's robust performance and superior reconstruction quality compared to state-of-the-art methods, particularly in preserving high-frequency details and attenuating noise. The proposed framework establishes a novel convergence between physics-based modeling and data-driven learning, paving the way for more accurate, flexible, and practical lensless imaging solutions for applications ranging from miniature sensors to medical diagnostics. The link of code is https://github.com/baijiesong/Lensnet.
Jiesong Bai, Yuhao Yin, Yihang Dong, Xiaofeng Zhang 0006, Chi-Man Pun, Xuhang Chen 0002
IJCAI6
2025 Multi-Scale Adaptively-Aware and Recalibration Network for Brain Tumor Segmentation with Missing Modalities
abstract
Accurate segmentation of brain tumor regions from multi-modal magnetic resonance imaging (MRI) is critical for clinical diagnosis. However, missing modalities is a common issue in clinical practice, where the unavailability of certain imaging modalities complicates the extraction and integration of complementary information across multiple modalities, leading to a decline in segmentation accuracy. Many existing models fail to adequately address subtle structural changes and boundary information within tumor regions when faced with missing modalities, limiting their ability to effectively adapt to complex tumor morphologies. To tackle these issues, The Multi-Scale Adaptively-Aware and Recalibration Network (MARNet) proposed in this paper can adaptively and fully explore the potential of multi-modal data under different combinations in the presence of missing modalities. MARNet incorporates a Feature Recalibration and Enhancement Module (FREM) that recalibrates and enhances the three-dimensional feature representation, emphasizing important fine-grained features of brain tumors. Subsequently, the Adaptive Shape-Aware Fusion Module (ASFM) fully exploits available modality information, achieving adaptive feature fusion for varying tumor locations and shapes, thereby compensating for information loss due to missing modalities. Furthermore, the Global and Multiscale Feature Integration Module (GMFIM) is designed to effectively capture long-range dependencies of tumors, particularly under conditions of missing modalities, aiding in the restoration and reconstruction of complete tumor structures. Extensive experiments on the BraTS2020, BraTS2018 and BraTS2015 datasets demonstrate that the proposed method surpasses several advanced brain tumor segmentation approaches in the context of missing modalities.
Guoheng Huang, Zhipeng Zheng, Xuhang Chen 0002, Lianglun Cheng
IJCNN4
2025 Code Retrieval with Mixture of Experts Prototype Learning Based on Classification
abstract
The semantic connection between code and queries is crucial for code retrieval, but many human-written queries fail to accurately capture the code's core intent, leading to ambiguity.This ambiguity complicates the code search process, as the queries do not provide a clear overview of the code's purpose.Our analysis reveals that while ambiguous queries may not precisely summarize the intent of the code, they often share the same general topics as the corresponding code.In light of this discovery, we propose Code Retrieval with Mixture of Experts Prototype Learning Based on Classification (CRME), a novel approach that combines classification for prototype-based representation learning and result ensembling.CRME utilizes specialized pre-trained models focused on the specific domains of ambiguous queries.It consists of two key components: Multiple Classification Prototype and Representation Learning with a Prototype-based Multi-model Contrastive (PMC) Loss during training, and Multi-Prototype Mixture of Experts Integration (MP-MoE) module for fine-grained ensemble inference.Our method can effectively address the issue of query ambiguity and improves search precision.Experimental results on the CodeSearchNet dataset, covering six sub-datasets, show that CRME outperforms existing methods, achieving an average MRR score of * Corresponding authors.
Feng Ling 0002, Guoheng Huang, Jingchao Wang 0002, Xiaochen Yuan, Xuhang Chen 0002, XueYong Zhang, Fanlong Zhang, Chi-Man Pun
Internetware5
2025 MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
abstract
Hallucinations pose a significant challenge in Large Vision Language Models (LVLMs), with misalignment between multimodal features identified as a key contributing factor. This paper reveals the negative impact of the long-term decay in Rotary Position Encoding (RoPE), used for positional modeling in LVLMs, on multimodal alignment. Concretely, under long-term decay, instruction tokens exhibit uneven perception of image tokens located at different positions within the two-dimensional space: prioritizing image tokens from the bottom-right region since in the one-dimensional sequence, these tokens are positionally closer to the instruction tokens. This biased perception leads to insufficient image-instruction interaction and suboptimal multimodal alignment. We refer to this phenomenon as ''image alignment bias.'' To enhance instruction's perception of image tokens at different spatial locations, we propose MCA-LLaVA, based on Manhattan distance, which extends the long-term decay to a two-dimensional, multi-directional spatial decay. MCA-LLaVA integrates the one-dimensional sequence order and two-dimensional spatial position of image tokens for positional modeling, mitigating hallucinations by alleviating image alignment bias. Experimental results of MCA-LLaVA across various hallucination and general benchmarks demonstrate its effectiveness and generality. The code can be accessed in https://github.com/ErikZ719/MCA-LLaVA.
Qiyan Zhao, Xiaofeng Zhang 0006, Yun Xing 0001, Xiaosong Yuan, Sinan Fan, Xuhang Chen 0002, Dahan Wang, Xu-Yao Zhang
ACM Multimedia8
2025 Cross-View Geo-Localization via Learning Correspondence Semantic Similarity Knowledge
Guanli Chen, Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun
MMM (1)4
2025 SFormer: SNR-Guided Transformer for Underwater Image Enhancement from the Frequency Domain
Yingtie Lei, Zimeng Li 0001, Chi-Man Pun, Xuhang Chen 0002
PRICAI (5)6
2025 MAC-Lookup: Multi-Axis Conditional Lookup Model for Underwater Image Enhancement
abstract
Enhancing underwater images is crucial for exploration. These images face Enhancing underwater images is crucial for exploration. These images face visibility and color issues due to light changes, water turbidity, and bubbles. Traditional prior-based methods and pixel-based methods often fail, while deep learning lacks sufficient high-quality datasets. We introduce the Multi-Axis Conditional Lookup (MAC-Lookup) model, which enhances visual quality by improving color accuracy, sharpness, and contrast. It includes Conditional 3D Lookup Table Color Correction (CLTCC) for preliminary color and quality correction and Multi-Axis Adaptive Enhancement (MAAE) for detail refinement. This model prevents over-enhancement and saturation while handling underwater challenges. Extensive experiments show that MAC-Lookup excels in enhancing underwater images by restoring details and colors better than existing methods. The code is https://github.com/onlycatdoraemon/MAC-Lookup.
Fanghai Yi, Zehong Zheng, Zexiao Liang, Yihang Dong, Xiyang Fang, Wangyu Wu, Xuhang Chen 0002
SMC7
2025 High-Fidelity Document Stain Removal via A Large-Scale Real-World Dataset and A Memory-Augmented Transformer
abstract
Document images are often degraded by various stains, significantly impacting their readability and hindering downstream applications such as document digitization and analysis. The absence of a comprehensive stained document dataset has limited the effectiveness of existing document enhancement methods in removing stains while preserving fine-grained details. To address this challenge, we construct StainDoc, the first large-scale, high-resolution (2145 x 2245) dataset specifically designed for document stain removal. StainDoc comprises over 5,000 pairs of stained and clean document images across multiple scenes. This dataset encompasses a diverse range of stain types, severities, and document backgrounds, facilitating robust training and evaluation of document stain removal algorithms. Furthermore, we propose StainRestorer, a Transformer-based document stain removal approach. StainRestorer employs a memory-augmented Transformer architecture that captures hierarchical stain representations at part, instance, and semantic levels via the DocMemory module. The Stain Removal Transformer (SRTransformer) leverages these feature representations through a dual attention mechanism: an enhanced spatial attention with an expanded receptive field, and a channel attention captures channel-wise feature importance. This combination enables precise stain removal while preserving document content integrity. Extensive experiments demonstrate StainRestorer's superior performance over state-of-the-art methods on the Stain-Doc dataset and its variants StainDoc.Mark and Stain-Doc.Seal, establishing a new benchmark for document stain removal. Our work highlights the potential of memory-augmented Transformers for this task and contributes a valuable dataset to advance future research.
Mingxian Li, Yingtie Lei, Xiaofeng Zhang 0006, Yihang Dong, Zimeng Li 0001, Xuhang Chen 0002
WACV8
2025 Underwater Image Restoration Through a Prior Guided Hybrid Sense Approach and Extensive Benchmark Analysis
abstract
Underwater imaging grapples with challenges from light-water interactions, leading to color distortions and reduced clarity. In response to these challenges, we propose a novel Color Balance Prior Guided Hybrid Sense Underwater Image Restoration framework (GuidedHybSensUIR). This framework operates on multiple scales, employing the proposed Detail Restorer module to restore low-level detailed features at finer scales and utilizing the proposed Feature Contextualizer module to capture long-range contextual relations of high-level general features at a broader scale. The hybridization of these different scales of sensing results effectively addresses color casts and restores blurry details. In order to effectively point out the evolutionary direction for the model, we propose a novel Color Balance Prior as a strong guide in the feature contextualization step and as a weak guide in the final decoding phase. We construct a comprehensive benchmark using paired training data from three real-world underwater datasets and evaluate on six test sets, including three paired and three unpaired, sourced from four real-world underwater datasets. Subsequently, we tested 14 traditional and retrained 23 deep learning existing underwater image restoration methods on this benchmark, obtaining metric results for each approach. This effort aims to furnish a valuable benchmarking dataset for standard basis for comparison. The extensive experiment results demonstrate that our method outperforms 37 other state-of-the-art methods overall on various benchmark datasets and metrics, despite not achieving the best results in certain individual cases. The code and dataset are available at https://github.com/CXH-Research/GuidedHybSensUIR.
Xiaojiao Guo, Xuhang Chen 0002, Shuqiang Wang, Chi-Man Pun
IEEE Trans. Circuits Syst. Video Technol.2
2025 An asymmetric calibrated transformer network for underwater image restoration
Xiaojiao Guo, Shenghong Luo, Yihang Dong, Zexiao Liang, Zimeng Li 0001, Xuhang Chen 0002
Vis. Comput.7
2025 Psanet: prototype-guided salient attention for few-shot segmentation
Guoheng Huang, Xiaochen Yuan, Zewen Zheng, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun
Vis. Comput.5
2025 Weakly supervised semantic segmentation via saliency perception with uncertainty-guided noise suppression
Guoheng Huang, Xiaochen Yuan, Zewen Zheng, Guo Zhong, Xuhang Chen 0002, Chi-Man Pun
Vis. Comput.6
2025 QEAN: quaternion-enhanced attention network for visual dance generation
Zhizhen Zhou, Yejing Huo, Guoheng Huang, An Zeng, Xuhang Chen 0002, Lian Huang, Zinuo Li
Vis. Comput.5
2024 Devignet: High-Resolution Vignetting Removal via a Dual Aggregated Fusion Transformer with Adaptive Channel Expansion
abstract
Vignetting commonly occurs as a degradation in images resulting from factors such as lens design, improper lens hood usage, and limitations in camera sensors. This degradation affects image details, color accuracy, and presents challenges in computational photography. Existing vignetting removal algorithms predominantly rely on ideal physics assumptions and hand-crafted parameters, resulting in the ineffective removal of irregular vignetting and suboptimal results. Moreover, the substantial lack of real-world vignetting datasets hinders the objective and comprehensive evaluation of vignetting removal. To address these challenges, we present VigSet, a pioneering dataset for vignetting removal. VigSet includes 983 pairs of both vignetting and vignetting-free high-resolution (over 4k) real-world images under various conditions. In addition, We introduce DeVigNet, a novel frequency-aware Transformer architecture designed for vignetting removal. Through the Laplacian Pyramid decomposition, we propose the Dual Aggregated Fusion Transformer to handle global features and remove vignetting in the low-frequency domain. Additionally, we propose the Adaptive Channel Expansion Module to enhance details in the high-frequency domain. The experiments demonstrate that the proposed model outperforms existing state-of-the-art methods. The code, models, and dataset are available at https://github.com/CXH-Research/DeVigNet.
Shenghong Luo, Xuhang Chen 0002, Weiwen Chen, Zinuo Li, Shuqiang Wang, Chi-Man Pun
AAAI2
2024 IMAN: An Adaptive Network for Robust NPC Mortality Prediction with Missing Modalities
abstract
Accurate prediction of mortality in nasopharyngeal carcinoma (NPC), a complex malignancy particularly challenging in advanced stages, is crucial for optimizing treatment strategies and improving patient outcomes. However, this predictive process is often compromised by the high-dimensional and heterogeneous nature of NPC-related data, coupled with the pervasive issue of incomplete multi-modal data, manifesting as missing radiological images or incomplete diagnostic reports. Traditional machine learning approaches suffer significant performance degradation when faced with such incomplete data, as they fail to effectively handle the high-dimensionality and intricate correlations across modalities. Even advanced multi-modal learning techniques like Transformers struggle to maintain robust performance in the presence of missing modalities, as they lack specialized mechanisms to adaptively integrate and align the diverse data types, while also capturing nuanced patterns and contextual relationships within the complex NPC data. To address these problem, we introduce IMAN: an adaptive network for robust NPC mortality prediction with missing modalities. IMAN features three integrated modules: the Dynamic Cross-Modal Calibration (DCMC) module employs adaptive, learnable parameters to scale and align medical images and field data; the Spatial-Contextual Attention Integration (SCAI) module enhances traditional Transformers by incorporating positional information within the self-attention mechanism, improving multi-modal feature integration; and the Context-Aware Feature Acquisition (CAFA) module adjusts convolution kernel positions through learnable offsets, allowing for adaptive feature capture across various scales and orientations in medical image modalities. Extensive experiments on our proprietary NPC dataset demonstrate IMAN’s robustness and high predictive accuracy, even with missing data. Compared to existing methods, IMAN consistently outperforms in scenarios with incomplete data, representing a significant advancement in mortality prediction for medical diagnostics and treatment planning. Our code is available at https://github.com/king-huoye/BIBM-2024/tree/master.
Yejing Huo, Guoheng Huang, Lianglun Cheng, Jianbin He, Xuhang Chen 0002, Xiaochen Yuan, Guo Zhong, Chi-Man Pun
BIBM5
2024 SMAFormer: Synergistic Multi-Attention Transformer for Medical Image Segmentation
abstract
In medical image segmentation, specialized computer vision techniques, notably transformers grounded in attention mechanisms and residual networks employing skip connections, have been instrumental in advancing performance. Nonetheless, previous models often falter when segmenting small, irregularly shaped tumors. To this end, we introduce SMAFormer, an efficient, Transformer-based architecture that fuses multiple attention mechanisms for enhanced segmentation of small tumors and organs. SMAFormer can capture both local and global features for medical image segmentation. The architecture comprises two pivotal components. First, a Synergistic Multi-Attention (SMA) Transformer block is proposed, which has the benefits of Pixel Attention, Channel Attention, and Spatial Attention for feature enrichment. Second, addressing the challenge of information loss incurred during attention mechanism transitions and feature fusion, we design a Feature Fusion Modulator. This module bolsters the integration between the channel and spatial attention by mitigating reshaping-induced information attrition. To evaluate our method, we conduct extensive experiments on various medical image segmentation tasks, including multi-organ, liver tumor, and bladder tumor segmentation, achieving state-of-the-art results. Code and models are available at: https://github.com/lzeeorno/SMAFormer.
Fuchen Zheng, Xuhang Chen 0002, Weihuang Liu, Haolun Li 0001, Yingtie Lei, Chi-Man Pun, Shoujun Zhou
BIBM2
2024 FAQNet: Frequency-Aware Quaternion Network for Endoscopic Highlight Removal
abstract
Due to the built-in light source within the endoscope, the illumination of bodily mucous can cause the formation of highlight regions due to reflection. This not only interferes with the diagnosis conducted by doctors but also poses a challenge to subsequent computer vision tasks. To tackle this issue, we introduce FAQNet, a network specifically designed for endoscopic image highlight removal. FAQNet seamlessly integrates multi-channel information leveraging quaternion convolution and spatial channel attention within our Quaternion Multi-Channel Fusion (QMCF) Module. This allows it to capture intricate details of color, texture, spatial information, and highlight characteristics within the imaged organ. Additionally, by employing frequency domain transformation and dilated convolution, the Contextual Information Integration (CII) Module effectively enlarges the receptive field, organizing contextual information between highlight regions and their surrounding areas. Lastly, the PixelShuffle Upsampling (PSU) Module generates the restored image. We validate our model’s performance on two benchmark datasets, demonstrating its superiority over existing highlight removal methodologies.
Dingzhou Zhu, Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun
BIBM4
2024 PDGC: Properly Disentangle by Gating and Contrasting for Cross-Domain Few-Shot Classification
Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Yan Li 0122, Chi-Man Pun, Junbing Quan
CGI (2)4
2024 FOPS-V: Feature-Aware Optimization and Parallel Scale Fusion for 3D Human Reconstruction in Video
Guoheng Huang, Lianglun Cheng, Yejing Huo, Xuhang Chen 0002, Xiaochen Yuan, Guo Zhong, Chi-Man Pun
ICONIP (8)5
2024 ROSAL: Semi-supervised Active Learning with Representation Aggregation and Outlier for Endoscopy Image Classification
Xiaocong Huang, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Xuhang Chen 0002, Chi-Man Pun, Jianwu Chen
ICONIP (11)5
2024 Test-Time Intensity Consistency Adaptation for Shadow Detection
Leyi Zhu, Weihuang Liu, Zimeng Li 0001, Xuhang Chen 0002, Chi-Man Pun
ICONIP (7)5
2024 Dual-Hybrid Attention Network for Specular Highlight Removal
Xiaojiao Guo, Xuhang Chen 0002, Shenghong Luo, Shuqiang Wang, Chi-Man Pun
ACM Multimedia2
2024 MedPrompt: Cross-modal Prompting for Multi-task Medical Image Translation
Xuhang Chen 0002, Shenghong Luo, Chi-Man Pun, Shuqiang Wang
PRCV (14)1
2024 DocDeshadower: Frequency-Aware Transformer for Document Shadow Removal
abstract
Shadows in scanned documents pose significant challenges for document analysis and recognition tasks due to their negative impact on visual quality and readability. Current shadow removal techniques, including traditional methods and deep learning approaches, face limitations in handling varying shadow intensities and preserving document details. To address these issues, we propose DocDeshadower, a novel multi-frequency Transformer-based model built upon the Laplacian Pyramid. By decomposing the shadow image into multiple frequency bands and employing two critical modules: the Attention-Aggregation Network for low-frequency shadow removal and the Gated Multi-scale Fusion Transformer for global refinement. DocDeshadower effectively removes shadows at different scales while preserving document content. Extensive experiments demonstrate DocDe-shadower's superior performance compared to state-of-the-art methods, highlighting its potential to significantly improve document shadow removal techniques. The code is available at https://github.com/leiyingtie/DocDeshadower.
Ziyang Zhou 0001, Yingtie Lei, Xuhang Chen 0002, Shenghong Luo, Chi-Man Pun
SMC3
2024 Cross-domain visual prompting with spatial proximity knowledge distillation for histological image classification
Guoheng Huang, Lianglun Cheng, Guo Zhong, Weihuang Liu, Xuhang Chen 0002, Muyan Cai
J. Biomed. Informatics6
2024 WavEnhancer: Unifying Wavelet and Transformer for Image Enhancement
Zinuo Li, Xuhang Chen 0002, Shu-Na Guo, Shuqiang Wang, Chi-Man Pun
J. Comput. Sci. Technol.2
2024 MFDNet: Multi-Frequency Deflare Network for efficient nighttime flare removal
Yiguo Jiang, Xuhang Chen 0002, Chi-Man Pun, Shuqiang Wang, Wei Feng 0005
Vis. Comput.2
2023 Shadocnet: Learning Spatial-Aware Tokens in Transformer for Document Shadow Removal
abstract
Shadow removal improves the visual quality and legibility of digital copies of documents. However, document shadow removal remains an unresolved subject. Traditional techniques rely on heuristics that vary from situation to situation. Given the quality and quantity of current public datasets, the majority of neural network models are ill-equipped for this task. In this paper, we propose a Transformer-based model for document shadow removal that utilizes shadow context encoding and decoding in both shadow and shadow-free regions. Additionally, shadow detection and pixel-level enhancement are included in the whole coarse-to-fine process. On the basis of comprehensive benchmark evaluations, it is competitive with state-of-the-art methods.
Xuhang Chen 0002, Xiaodong Cun, Chi-Man Pun, Shuqiang Wang
ICASSP1
2023 High-Resolution Document Shadow Removal via A Large-Scale Real-World Dataset and A Frequency-Aware Shadow Erasing Net
abstract
Shadows often occur when we capture the document with casual equipment, which influences the visual quality and readability of the digital copies. Different from the algorithms for natural shadow removal, the algorithms in document shadow removal need to preserve the details of fonts and figures in high-resolution input. Previous works ignore this problem and remove the shadows via approximate attention and small datasets, which might not work in real-world situations. We handle high-resolution document shadow removal directly via a larger-scale real-world dataset and a carefully-designed frequency-aware network. As for the dataset, we acquire over 7k couples of high-resolution (2462 × 3699) images of real-world documents pairs with various samples under different lighting circumstances, which is 10 times larger than existing datasets. As for the design of the network, we decouple the high-resolution images in the frequency domain, where the low-frequency details and high-frequency boundaries can be effectively learned via the carefully designed network structure. Powered by our network and dataset, the proposed method shows a clearly better performance than previous methods in terms of visual quality and numerical results. The code, models, and dataset are available at https://github.com/CXH-Research/DocShadow-SD7K.
Zinuo Li, Xuhang Chen 0002, Chi-Man Pun, Xiaodong Cun
ICCV2
2023 A Large-Scale Film Style Dataset for Learning Multi-frequency Driven Film Enhancement
abstract
Film, a classic image style, is culturally significant to the whole photographic industry since it marks the birth of photography. However, film photography is time-consuming and expensive, necessitating a more efficient method for collecting film-style photographs. Numerous datasets that have emerged in the field of image enhancement so far are not film-specific. In order to facilitate film-based image stylization research, we construct FilmSet, a large-scale and high-quality film style dataset. Our dataset includes three different film types and more than 5000 in-the-wild high resolution images. Inspired by the features of FilmSet images, we propose a novel framework called FilmNet based on Laplacian Pyramid for stylizing images across frequency bands and achieving film style outcomes. Experiments reveal that the performance of our model is superior than state-of-the-art techniques. The link of our dataset and code is https://github.com/CXH-Research/FilmNet.
Zinuo Li, Xuhang Chen 0002, Shuqiang Wang, Chi-Man Pun
IJCAI2
2023 Brain Diffuser: An End-to-End Brain Image to Brain Network Pipeline
Xuhang Chen 0002, Bai Ying Lei, Chi-Man Pun, Shuqiang Wang
PRCV (13)1