Zheyu Zhang 0002

dblp:172/0655-2 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-7487-4673ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Incomplete Multi-modal Brain Tumor Segmentation via Learnable Sorting State Space Model
abstract
Brain tumor segmentation plays a crucial role in clinical diagnosis, yet the frequent unavailability of certain MRI modalities poses a significant challenge. In this paper, we introduce the Learnable Sorting State Space Model (LS3M), a novel framework designed to maximize the utilization of available modalities for brain tumor segmentation. LS3M excels at efficiently modeling long-range dependencies based on the Mamba design, while incorporating differentiable permutation matrices that reorder input sequences based on modality-specific characteristics. This dynamic reordering ensures that critical spatial inductive biases and long-range semantic correlations inherent in 3D brain MRI are preserved, which is crucial for imcomplete multi-modal brain tumor segmentation. Once the input sequences are reordered using the generated permutation matrix, the Series State Space Model (S3M) block models the relationships between them, capturing both local and long-range dependencies. This enables effective representation of intra-modal and inter-modal relationships, significantly improving segmentation accuracy. Extensive experiments on the BraTS2018 and BraTS2020 datasets demonstrate that LS3M outperforms existing methods, offering a robust solution for brain tumor segmentation, particularly in scenarios with missing modalities.
Zheyu Zhang 0002, Yayuan Lu, Feipeng Ma, Yueyi Zhang 0001, Huanjing Yue, Xiaoyan Sun 0001
CVPR1
2025 Efficient Spiking Point Mamba for Point Cloud Analysis
Peixi Wu, Bosong Chai, Menghua Zheng, Zhangchi Hu, Jie Chen 0001, Zheyu Zhang 0002, Hebei Li, Xiaoyan Sun 0001
ICCV7
2025 MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language Models
abstract
The rapid evolution of vision-language models (VLMs), such as CLIP, has demonstrated remarkable zero-shot capabilities in downstream tasks. Prompt learning paradigms like Context Optimization (CoOp) refine learnable prompts for efficient adaptation. However, their application in the biomedical domain remains limited due to the insufficient utilization of specialized biomedical knowledge and cross-modality structural relationships. To address these, we introduce MeDKCoOp, a Medical Dual Knowledge-guided graph adaptation method that leverages systematic integration of knowledge through three aspects: exploit domain-specific knowledge from both textual and visual branches, formalize it into graph-structured representations, and leverage knowledge-guided relation transfer for learning cross-modality fusion. By dynamically optimizing learnable prompts through relation learning process, our method achieves disentangled visual representation and enhances transferability to downstream tasks. Evaluations across 8 biomedical datasets spanning 7 imaging modalities demonstrate state-of-the-art cross-domain generalization, with an average 15.12% accuracy improvement over baselines. Our work establishes a new paradigm through graph prompt learning in medical vision-language models, advancing robust diagnostic AI in data-scarce clinical scenarios. Our code is available at: https://github.com/WangYijun-OUC/MeDKCoOp.
Siying Wu, Lubin Gan, Zheyu Zhang 0002, Jing Zhang 0165, Zhangchi Hu, Huyue Zhu, Peixi Wu, Xiaoyan Sun 0001
ACM Multimedia4
2024 TMFormer: Token Merging Transformer for Brain Tumor Segmentation with Missing Modalities
abstract
Numerous techniques excel in brain tumor segmentation using multi-modal magnetic resonance imaging (MRI) sequences, delivering exceptional results. However, the prevalent absence of modalities in clinical scenarios hampers performance. Current approaches frequently resort to zero maps as substitutes for missing modalities, inadvertently introducing feature bias and redundant computations. To address these issues, we present the Token Merging transFormer (TMFormer) for robust brain tumor segmentation with missing modalities. TMFormer tackles these challenges by extracting and merging accessible modalities into more compact token sequences. The architecture comprises two core components: the Uni-modal Token Merging Block (UMB) and the Multi-modal Token Merging Block (MMB). The UMB enhances individual modality representation by adaptively consolidating spatially redundant tokens within and outside tumor-related regions, thereby refining token sequences for augmented representational capacity. Meanwhile, the MMB mitigates multi-modal feature fusion bias, exclusively leveraging tokens from present modalities and merging them into a unified multi-modal representation to accommodate varying modality combinations. Extensive experimental results on the BraTS 2018 and 2020 datasets demonstrate the superiority and efficacy of TMFormer compared to state-of-the-art methods when dealing with missing modalities.
Zheyu Zhang 0002, Yueyi Zhang 0001, Huanjing Yue, Aiping Liu, Yunwei Ou, Xiaoyan Sun 0001
AAAI1
2024 Multi-modal Diffusion Network with Controllable Variability for Medical Image Segmentation
abstract
In diffusion-based medical segmentation models, stochastic sampling is commonly used to generate multiple masks. However, the inherent variability in diffusion models can lead to significant biases in some masks, resulting in the fused mask deviating from the true mask. In this study, we propose a novel multi-modal diffusion segmentation network (MMDSN) with controllable variability, specifically designed to address the issue of variability in diffusion models. MMDSN achieves multi-modal conditional control through medical text annotations, thereby enhancing consistency of visual semantic representation and establishing a correspondence between vision and language for diffusion models. Additionally, MMDSN constrains the uncertainty distributions of multiple timesteps within the latent Gaussian space, controlling the variability at each denoising timestep. Extensive experiments on the Qata-Covid19 and MosMed datasets demonstrate that our proposed method surpasses existing state-of-the-art diffusion networks, producing a high-quality, controllable segmentation map with just a single reverse diffusion step and one sampling.
Zheyu Zhang 0002, Yueyi Zhang 0001, Jing Zhang 0165, Yunwei Ou, Xiaoyan Sun 0001
BIBM2
2024 Anatomical Consistency Distillation and Inconsistency Synthesis for Brain Tumor Segmentation with Missing Modalities
abstract
Multi-modal Magnetic Resonance Imaging (MRI) is imperative for accurate brain tumor segmentation, offering indispensable complementary information. Nonetheless, the absence of modalities poses significant challenges in achieving precise segmentation. Recognizing the shared anatomical structures between mono-modal and multi-modal representations, it is noteworthy that mono-modal images typically exhibit limited features in specific regions and tissues. In response to this, we present Anatomical Consistency Distillation and Inconsistency Synthesis (ACDIS), a novel framework designed to transfer anatomical structures from multi-modal to mono-modal representations and synthesize modality-specific features. ACDIS consists of two main components: Anatomical Consistency Distillation (ACD) and Modality Feature Synthesis Block (MFSB). ACD incorporates the Anatomical Feature Enhancement Block (AFEB), meticulously mining anatomical information. Simultaneously, Anatomical Consistency ConsTraints (ACCT) are employed to facilitate the consistent knowledge transfer, i.e., the richness of information and the similarity in anatomical structure, ensuring precise alignment of structural features across mono-modality and multi-modality. Complementarily, MFSB produces modality-specific features to rectify anatomical inconsistencies, thereby compensating for missing information in the segmented features. Through validation on the BraTS2018 and BraTS2020 datasets, ACDIS substantiates its efficacy in the segmentation of brain tumors with missing MRI modalities.
Zheyu Zhang 0002, Xinzhao Liu, Yueyi Zhang 0001, Huanjing Yue, Yunwei Ou, Xiaoyan Sun 0001
ECAI1
2024 ESTME: Event-driven Spatio-temporal Motion Enhancement for Micro-Expression Recognition
abstract
The inherently rapid and subtle changes in micro-expressions pose significant challenges for micro-expression recognition (MER). Previous methods, typically relying on frame aggregation or optical flow, struggle to accurately capture subtle changes because of low frame rate. In this paper, we propose an Event-driven Spatio-temporal Motion Enhancement Network, which incorporates event signals captured by an event camera, to assist MER. Specifically, we introduce an Event-Enhanced Motion Extractor module to exploit event signals’ high temporal resolution property, enhancing subtle motion details. We also propose an Event-Guided Attention module to focus on subtle changes in specific areas, capturing more precise spatial features of micro-expressions. Experimental results on synthetic and real-world datasets demonstrate the superiority of our method on MER, showcasing its strong ability to capture subtle motion changes.
Peilin Xiao, Yueyi Zhang 0001, Dachun Kai, Yansong Peng, Zheyu Zhang 0002, Xiaoyan Sun 0001
ICME5
2024 Hue Guidance Network for Single Image Reflection Removal
abstract
Reflection from glasses is ubiquitous in daily life, but it is usually undesirable in photographs. To remove these unwanted noises, existing methods utilize either correlative auxiliary information or handcrafted priors to constrain this ill-posed problem. However, due to their limited capability to describe the properties of reflections, these methods are unable to handle strong and complex reflection scenes. In this article, we propose a hue guidance network (HGNet) with two branches for single image reflection removal (SIRR) by integrating image information and corresponding hue information. The complementarity between image information and hue information has not been noticed. The key to this idea is that we found that hue information can describe reflections well and thus can be used as a superior constraint for the specific SIRR task. Accordingly, the first branch extracts the salient reflection features by directly estimating the hue map. The second branch leverages these effective features, which can help locate salient reflection regions to obtain a high-quality restored image. Furthermore, we design a new cyclic hue loss to provide a more accurate optimization direction for the network training. Experiments substantiate the superiority of our network, especially its excellent generalization ability to various reflection scenes, as compared with state-of-the-arts both qualitatively and quantitatively. Source codes are available at https://github.com/zhuyr97/HGRR.
Yurui Zhu, Xueyang Fu, Zheyu Zhang 0002, Aiping Liu, Zhiwei Xiong, Zhengjun Zha
IEEE Trans. Neural Networks Learn. Syst.3
2023 EoFormer: Edge-Oriented Transformer for Brain Tumor Segmentation
Dong She, Yueyi Zhang 0001, Zheyu Zhang 0002, Hebei Li, Xiaoyan Sun 0001
MICCAI (4)3
2021 Multifocal Attention-Based Cross-Scale Network for Image De-raining
abstract
Albeit existing deep learning-based image de-raining methods have achieved promising results, most of them only extract single scale features, and neglect the fact that similar rain streaks appear repeatedly across different scales. Therefore, this paper aims to explore the cross-scale cues in a multi-scale fashion. Specifically, we first introduce an adaptive-kernel pyramid to provide effective multi-scale information. Then, we design two cross-scale similarity attention blocks (CSSABs) to search spatial and channel relationships between two scales, respectively. The spatial CSSAB explores the spatial similarity between pixels of cross-scale features, while the channel CSSAB emphasizes the interdependencies among cross-scale features. To further improve the diversity of features, we adopt the wavelet transformation and multi-head mechanism in CSSABs to generate multifocal features which focus on different areas. Finally, based on our CSSABs, we construct an effective multifocal attention-based cross-scale network, which exhaustively utilizes the cross-scale correlations of both rain streaks and background, to achieve image de-raining. Experiments show the superiority of our network over state-of-the-art image de-raining approaches both qualitatively and quantitatively. The source code and pre-trained models are available at https://github.com/zhangzheyu0/Multifocal_derain.
Zheyu Zhang 0002, Yurui Zhu, Xueyang Fu, Zhiwei Xiong, Zhengjun Zha, Feng Wu 0001
ACM Multimedia1