Xiaoming Zhang 0008

dblp:86/2120-8 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0000-3986-885XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Non-contrast CT esophageal varices grading through clinical prior-enhanced multi-organ analysis
Xiaoming Zhang 0008, Chunli Li, Jiacheng Hao, Yuan Gao 0017, Danyang Tu, Jianyi Qiao, Xiaoli Yin, Le Lu 0001, Ling Zhang 0002, Ke Yan 0006
Medical Image Anal.1
2025 Diff-Shadow: Global-guided Diffusion Model for Shadow Removal
abstract
We propose Diff-Shadow, a global-guided diffusion model for high-quality shadow removal. Previous transformer-based approaches can utilize global information to relate shadow and non-shadow regions but are limited in their synthesis ability and recover images with obvious boundaries. In contrast, diffusion-based methods can generate better content but they are not exempt from issues related to inconsistent illumination. In this work, we combine the advantages of diffusion models and global guidance to realize shadow-free restoration. Specifically, we propose a parallel UNets architecture: 1) the local branch performs the patch-based noise estimation in the diffusion process, and 2) the global branch recovers the low-resolution shadow-free images. A Reweight Cross Attention (RCA) module is designed to integrate global contextual information of non-shadow regions into the local branch. We further design a Global-guided Sampling Strategy (GSS) that mitigates patch boundary issues and ensures consistent illumination across shaded and unshaded regions in the recovered image. Comprehensive experiments on three publicly standard datasets ISTD, ISTD+, and SRD have demonstrated the effectiveness of Diff-Shadow. Compared to state-of-the-art methods, our method achieves a significant improvement in terms of PSNR, increasing from 32.33dB to 33.69dB on the ISTD dataset.
Jinting Luo, Ru Li 0002, Chengzhi Jiang, Xiaoming Zhang 0008, Mingyan Han, Ting Jiang 0005, Haoqiang Fan, Shuaicheng Liu
AAAI4
2025 PLUS: Plug-and-Play Enhanced Liver Lesion Diagnosis Model on Non-contrast CT Scans
Jiacheng Hao, Xiaoming Zhang 0008, Wei Liu 0127, Xiaoli Yin, Yuan Gao 0017, Chunli Li, Ling Zhang 0002, Le Lu 0001, Xu Han 0023, Ke Yan 0006
MICCAI (15)2
2025 MAT: Multi-Range Attention Transformer for Efficient Image Super-Resolution
abstract
Image super-resolution (SR) has significantly advanced through the adoption of Transformer architectures. However, conventional techniques aimed at enlarging the self-attention window to capture broader contexts come with inherent drawbacks, especially the significantly increased computational demands. Moreover, the feature perception within a fixed-size window of existing models restricts the effective receptive field (ERF) and the intermediate feature diversity. We demonstrate that a flexible integration of attention across diverse spatial extents can yield significant performance enhancements. In line with this insight, we introduce Multi-Range Attention Transformer (MAT) for SR tasks. MAT leverages the computational advantages inherent in dilation operation, in conjunction with self-attention mechanism, to facilitate both multi-range attention (MA) and sparse multi-range attention (SMA), enabling efficient capture of both regional and sparse global features. Combined with local feature extraction, MAT adeptly capture dependencies across various spatial ranges, improving the diversity and efficacy of its feature representations. We also introduce the MSConvStar module, which augments the model’s ability for multi-range representation learning. Comprehensive experiments show that our MAT exhibits superior performance to existing state-of-the-art SR models with remarkable efficiency ($\sim 3.3\times $faster than SRFormer-light). The codes are available athttps://github.com/stella-von/MAT.
Chengxing Xie, Xiaoming Zhang 0008, Linze Li 0001, Yuqian Fu, Biao Gong, Tianrui Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 LIDIA: Precise Liver Tumor Diagnosis on Multi-Phase Contrast-Enhanced CT via Iterative Fusion and Asymmetric Contrastive Learning
Wei Liu 0127, Xiaoming Zhang 0008, Xiaoli Yin, Xu Han 0023, Chunli Li, Yuan Gao 0017, Le Lu 0001, Ling Zhang 0002, Lei Zhang 0006, Ke Yan 0006
MICCAI (9)3
2024 Improved Esophageal Varices Assessment from Non-contrast CT Scans
Chunli Li, Xiaoming Zhang 0008, Yuan Gao 0017, Xiaoli Yin, Le Lu 0001, Ling Zhang 0002, Ke Yan 0006
MICCAI (5)2
2024 Efficient Single Image Super-Resolution with Entropy Attention and Receptive Field Augmentation
abstract
Transformer-based deep models for single image super-resolution (SISR) have greatly improved the performance of lightweight SISR tasks in recent years. However, they often suffer from heavy computational burden and slow inference due to the complex calculation of multi-head self-attention (MSA), seriously hindering their practical application and deployment. In this work, we present an efficient SR model to mitigate the dilemma between model efficiency and SR performance, which is dubbed Entropy Attention and Receptive Field Augmentation network (EARFA), and composed of a novel entropy attention (EA) and a shifting large kernel attention (SLKA). From the perspective of information theory, EA increases the entropy of intermediate features conditioned on a Gaussian distribution, providing more informative input for subsequent reasoning. On the other hand, SLKA extends the receptive field of SR models with the assistance of channel shifting, which also favors to boost the diversity of hierarchical features. Since the implementation of EA and SLKA does not involve complex computations (such as extensive matrix multiplications), the proposed method can achieve faster nonlinear inference than Transformer-based SR models while maintaining better SR performance. Extensive experiments show that the proposed model can significantly reduce the delay of model inference while achieving the SR performance comparable with other advanced models.
Xiaole Zhao, Linze Li 0001, Chengxing Xie, Xiaoming Zhang 0008, Ting Jiang 0005, Shuaicheng Liu, Tianrui Li 0001
ACM Multimedia4
2023 Boosting Single Image Super-Resolution via Partial Channel Shifting
abstract
Although deep learning has significantly facilitated the progress of single image super-resolution (SISR) in recent years, it still hits bottlenecks to further improve SR performance with the continuous growth of model scale. Therefore, one of the hotspots in the field is to construct efficient SISR models by elevating the effectiveness of feature representation. In this work, we present a straightforward and generic approach for feature enhancement that can effectively promote the performance of SR models, dubbed partial channel shifting (PCS). Specifically, it is inspired by the temporal shifting in video understanding and displaces part of the channels along the spatial dimensions, thus allowing the effective receptive field to be amplified and the feature diversity to be augmented at almost zero cost. Also, it can be assembled into off-the-shelf models as a plug-and-play component for performance boosting without extra network parameters and computational overhead. However, regulating the features with PCS encounters some issues, like shifting directions and amplitudes, proportions, patterns of shifted channels, etc. We impose some technical constraints on the issues to simplify the general channel shifting. Extensive and throughout experiments illustrate that the PCS indeed enlarges the effective receptive field, augments the feature diversity for efficiently enhancing SR recovery, and can endow obvious performance gains to existing models.
Xiaoming Zhang 0008, Tianrui Li 0001, Xiaole Zhao
ICCV1
2023 Lightweight Remote-Sensing Image Super-Resolution via Re-Parameterized Feature Distillation Network
abstract
Recent deep learning based-works have made remarkable progress in Remote-Sensing Image Super-Resolution (RSISR). However, the complicated network architecture as well as a huge amount of parameters increase computational cost, hindering their practical deployment. To alleviate this problem, we propose a novel Re-parameterized Feature Distillation Network (ReFDN) for lightweight and efficient RSISR tasks. Feature distillation, refinement, condensation, and enhancement are efficiently integrated into the re-parameterized feature distillation block named ReFDB for lighter and stronger feature extraction. With the help of elaborate re-parameterized convolution (ReConv) design, we further boost the feature refinement capability without extra inference costs. Additionally, we design an efficient channel and spatial attention module (ECSA) to enhance the important objects and regions of the intermediate features adaptively. Conducted on both commonly used datasets and additional Google Earth data, the experimental results demonstrate our method can achieve a good trade-off between SR performance and network complexity. Our code will be publicly available at https://github.com/DaxingZ/ReFDN.
Chunjiang Bian, Xiaoming Zhang 0008, Hongzhen Chen
IEEE Geosci. Remote. Sens. Lett.3
2022 Speech Emotion Recognition with Complementary Acoustic Representations
abstract
Since CNNs promote local features and Transformers capture long-range dependencies, we explore both models as encoders for acoustic representations in a parallel framework for speech emotion recognition. We choose logMels as input to the CNN encoder and MFCCs to the Transformer encoder. The complementary acoustic representations generated by the two encoders are then fused to predict the frequency distribution of emotions. To further improve the performance, we conduct data augmentation based on vocal tract length perturbation and pretrain the Transformer encoder. The proposed framework is evaluated under the speaker-independent (SI) setting on the improvisation part of the Interactive Emotional Dyadic Motion Capture (IEMOCAP) dataset. Our weighted and unweighted accuracies reached 81.6% and 79.8%, respectively. To the best of our knowledge, this is the state-of-the-art result reported so far on this dataset in the SI scenario.
Xiaoming Zhang 0008
SLT1