Xu Zhang 0044

dblp:98/5660-44 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0001-7685-7500ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Any2RSI: Controllable Remote Sensing Text-to-Image Generation via Any Control and Enriched Description
abstract
Recent advances in controllable text-to-image (T2I) generation have achieved impressive results in natural images, but remote sensing (RS) T2I remains challenging due to the unique nature of geospatial data. Existing methods struggle to integrate diverse spatial controls and model complex spatial relationships, often failing to maintain semantic consistency with typically vague or incomplete textual descriptions. Moreover, limited by small-scale, low-quality datasets, these models produce outputs with inconsistent layouts and unrealistic content. To address these issues, we propose Any2RSI, a flexible framework for controllable RS T2I generation. It features a Cross-Modal Multi-Control Adapter that extracts modality-agnostic embeddings from heterogeneous spatial inputs, enabling precise spatial guidance. To compensate for sparse or ambiguous text prompts, we introduce a VLM-Empowered Enriched Description Generation module that enhances input descriptions with cross-modal semantics for more coherent image generation. Furthermore, we present RST2I-110K, a new large-scale dataset with over 115,000 high-quality RS image-text pairs across diverse scenes, alleviating data scarcity in this domain. Extensive experiments show that Any2RSI achieves state-of-the-art performance on both existing and new datasets, improving the realism and structural accuracy of generated RS imagery.
Xu Zhang 0044, Lefei Zhang
AAAI1
2026 ClearAIR: A Human-Visual-Perception-Inspired All-in-One Image Restoration
abstract
Recently, All-in-One image restoration (AiOIR) has advanced significantly, offering promising solutions for complex real-world degradations. However, most existing approaches heavily rely on degradation-specific representation learning, which can lead to oversmoothing and artifacts in the restored images. To address this limitation, we propose ClearAIR, a novel AiOIR framework inspired by human visual perception and designed with a hierarchical restoration strategy in a coarse-to-fine manner. First, leveraging the global priority characteristic of early human visual perception, we employ an image quality assessment model to evaluate the overall image structure and degradation level. Next, we introduce a Semantic Guidance Unit to provide coarse semantic region guidance and a Task Identifier to predict local degradation types, enabling a more informed characterization of local degradation patterns. Finally, aiming at the challenge of local detail restoration, we propose an Internal Clue Reuse Mechanism that deeply mines the internal information of the image in a self-supervised manner to enhance the model’s capacity for fine-detail recovery. Experimental results demonstrate that ClearAIR achieves superior restoration performance across diverse synthetic and real-world datasets.
Xu Zhang 0044, Huan Zhang 0008, Guoli Wang 0004, Qian Zhang 0009, Lefei Zhang
AAAI1
2026 Perceive-IR: Learning to Perceive Degradation Better for All-in-One Image Restoration
abstract
Existing All-in-One image restoration methods often fail to perceive degradation types and severity levels simultaneously, overlooking the importance of fine-grained quality perception. Moreover, these methods often utilize highly customized backbones, which hinder their adaptability and integration into more advanced restoration networks. To address these limitations, we propose Perceive-IR, a novel backbone-agnostic All-in-One image restoration framework designed for fine-grained quality control across various degradation types and severity levels. Its modular structure allows core components to function independently of specific backbones, enabling seamless integration into advanced restoration models without significant modifications. Specifically, Perceive-IR operates in two key stages: 1) multi-level quality-driven prompt learning stage, where a fine-grained quality perceiver is meticulously trained to discern three-tier quality levels by optimizing the alignment between prompts and images within the CLIP perception space. This stage ensures a nuanced understanding of image quality, laying the groundwork for subsequent restoration; 2) restoration stage, where the quality perceiver is seamlessly integrated with a difficulty-adaptive perceptual loss, forming a quality-aware learning strategy. This strategy not only dynamically differentiates sample learning difficulty but also achieves fine-grained quality control by driving the restored image toward the ground truth while pulling it away from both low- and medium-quality samples. Furthermore, Perceive-IR incorporates a Semantic Guidance Module (SGM) and Compact Feature Extraction (CFE). The SGM leverages semantic information from pre-trained vision models to provide high-level contextual guidance, while the CFE focuses on extracting degradation-specific features, ensuring accurate handling of diverse image degradations. Extensive experiments demonstrate that Perceive-IR not only surpasses state-of-the-art methods but also generalizes reliably to zero-shot real-world and unknown degraded scenes, while adapting seamlessly to different backbone networks. This versatility underscores the framework's robustness and backbone-agnostic design. Project page at https://house-yuyu.github.io/Perceive-IR/.
Xu Zhang 0044, Jiaqi Ma 0002, Guoli Wang 0004, Qian Zhang 0009, Huan Zhang 0008, Lefei Zhang
IEEE Trans. Image Process.1
2026 Leveraging Multi-Text Joint Prompts in SAM for Robust Medical Image Segmentation
abstract
The Segment Anything Model (SAM) has attracted considerable attention due to its impressive performance and demonstrates potential in medical image segmentation. Compared to SAM's native point andbounding box prompts, text prompts offer a simpler and more efficient alternative in the medical field, yet this approach remains relatively underexplored. In this paper, we propose a SAM-based framework that integrates a pre-trained vision-language model to generate referring prompts, with SAM handling the segmentation task. The outputs from multimodal models such as CLIP serve as input to SAM's prompt encoder. A critical challenge stems from the inherent complexity of medical text descriptions: they typically encompass anatomical characteristics, imaging modalities, and diagnostic priorities, resulting in information redundancy and semantic ambiguity. To address this, we propose a text decomposition-recomposition strategy. First, clinical narratives are parsed into atomic semantic units (appearance, location, pathology, and so on). These elements are then recombined into optimized text expressions. We employ a cross-attention module among multiple texts to interact with the joint features, ensuring that the model focuses on features corresponding to effective descriptions. To validate the effectiveness of our method, we conducted experiments on several datasets. Compared to the native SAM based on geometric prompts, our model shows improved performance and usability.
Xu Zhang 0044, Huangxuan Zhao, Lefei Zhang, Yuan Xiong
IEEE J. Biomed. Health Informatics1
2026 Enhanced Quality-Aware Scalable Underwater Image Compression
abstract
Underwater imaging plays a pivotal role in marine exploration and ecological monitoring. However, it faces significant challenges of limited transmission bandwidth and severe distortion in the aquatic environment. In this work, to achieve the target of both underwater image compression and enhancement simultaneously, an enhanced quality-aware scalable underwater image compression framework is presented, which comprises a Base Layer (BL) and an Enhancement Layer (EL). In the BL, the underwater image is represented by a controllable number of non-zero sparse coefficients for coding bits saving. Furthermore, the underwater image enhancement dictionary is derived with shared sparse coefficients to make reconstruction close to the enhanced version. In the EL, a dual-branch filter comprising rough filtering and detail refinement branches is designed to produce a pseudo-enhanced version for residual redundancy removal and to improve the quality of final reconstruction. Extensive experimental results demonstrate that the proposed scheme outperforms the state-of-the-art works under five large-scale underwater image datasets in terms of Underwater Image Quality Measure (UIQM).
Linwei Zhu, Xu Zhang 0044, Huan Zhang 0008, Ye Li 0002, Runmin Cong, Sam Kwong
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Joint multi-dimensional dynamic attention and transformer for general image restoration
Huan Zhang 0008, Xu Zhang 0044, Nian Cai, Jianglei Di, Yun Zhang 0002
Comput. Vis. Image Underst.2
2025 Enhancing 3D video watching experiences: Tackling compression and 3D warping distortions in synthesized view with perceptual guidance
Huan Zhang 0008, Xu Zhang 0044, Linwei Zhu, Yun Zhang 0002, Jiang-Zhong Cao, Bingo Wing-Kuen Ling
Expert Syst. Appl.2
2025 Unveiling the underwater world: CLIP perception model-guided underwater image enhancement
Jiang-Zhong Cao, Zekai Zeng, Xu Zhang 0044, Huan Zhang 0008, Chunling Fan, Gangyi Jiang, Weisi Lin
Pattern Recognit.3
2025 UniUIR: Considering Underwater Image Restoration as an All-in-One Learner
abstract
Existing underwater image restoration (UIR) methods generally only handle color distortion or jointly address color and haze issues, but they often overlook the more complex degradations that can occur in underwater scenes. To address this limitation, we propose a Universal Underwater Image Restoration method, termed as UniUIR, considering the complex scenario of real-world underwater mixed distortions as an all-in-one manner. To disentangle degradation-specific effects and capture their inter-correlations, we propose the Mamba Mixture-of-Experts module (MMoEM). Each expert specializes in distinct aspects of degradation, while gating mechanism dynamically routes features to appropriate experts. This design enables collaborative prior extraction and preserves global context, all within linear computational complexity. Building upon this foundation, to enhance degradation representation and address the task conflicts that arise when handling multiple types of degradation, we introduce the spatial-frequency prior generator. This module extracts degradation prior information in both spatial and frequency domains, and adaptively selects the most appropriate task-specific prompts based on image content, thereby improving the accuracy of image restoration. Finally, to more effectively address complex, region-dependent distortions in UIR task, we incorporate depth information derived from a large-scale pre-trained depth prediction model, thereby enabling the network to perceive and leverage depth variations across different image regions to handle localized degradation. Extensive experiments demonstrate that UniUIR can produce more attractive results across qualitative and quantitative comparisons, and shows strong generalization than state-of-the-art methods. Project page at https://house-yuyu.github.io/UniUIR.
Xu Zhang 0044, Huan Zhang 0008, Guoli Wang 0004, Qian Zhang 0009, Lefei Zhang, Bo Du 0001
IEEE Trans. Image Process.1
2024 MAdapter: A Better Interaction Between Image and Language for Medical Image Segmentation
Xu Zhang 0044, Bo Ni, Lefei Zhang
MICCAI (9)1
2024 Saliency guided progressive fusion of infrared and polarization for military images with complex backgrounds$^{\star }$
Yukai Lao, Huan Zhang 0008, Xu Zhang 0044, Jiazhen Dou, Jianglei Di
Multim. Tools Appl.3
2023 AFD-Former: A Hybrid Transformer With Asymmetric Flow Division for Synthesized View Quality Enhancement
abstract
Recently, CNN-based post-processing has shown great potential in Synthesized View Quality Enhancement (SVQE). However, due to the limited receptive field of convolution, it is ineffective in explicitly modeling long-range dependencies, which are critical to eliminate the distortion induced by Depth Image Based Rendering (DIBR) in synthesized views. Although transformers exhibit tremendous success at learning global contextual information, it is weak at extracting local texture information. To take full advantages of the CNN and transformer, we present a novel U-shaped hybrid transformer with asymmetric flow division to collaboratively capture global-local information for SVQE, termed as AFD-former. Specifically, the AFD-former utilizes the Transformer-CNN Block (TCB) as encoder and decoder, in which several Dynamic Hybrid Attention Blocks (DHABs) are designed to simultaneously model long-range interactions and retain texture details. Then, considering that the deeper layers of the U-shaped network play more roles in capturing global information while shallow layers more in extracting local information, an Asymmetric Flow Division Unit (AFDU) is embedded into each DHAB to assign different contributions of global-local contextual information to the transformer and CNN branches across different layers. Finally, a dynamic learnable modulator is incorporated into two branches to help model effectively feature representation learning. That can be viewed as the dynamic process of adjusting the weight for each channel of the input feature based on contextual cues. Extensive experiments demonstrate that the proposed AFD-former can significantly enhance perceptual quality of synthesized views with similar SVQE speed compared with the related state-of-the-art SVQE methods. The source code will be available athttps://github.com/House-yuyu/AFD-former.
Xu Zhang 0044, Nian Cai, Huan Zhang 0008, Yun Zhang 0002, Jianglei Di, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.1