Yulan Dai

dblp:308/6299 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0003-3143-7347ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Alzheimer's disease classification based on multimodal consistent distribution and trusted fusion
Xiaoyan Kui, Yulan Dai, Beiji Zou 0001, Chengzhang Zhu, Yang Li 0111, Zexin Ji, Liming Chen 0002, Miguel Bordallo López
Neural Networks2
2025 Robust Multimodal Representation Learning with Information Bottleneck and Balanced Fusion for Alzheimers Disease Classification
abstract
Given the capability of multimodal data to provide information from multiple perspectives, it is beneficial for improving the accuracy of Alzheimer’s disease (AD) classification. However, during practical multimodal learning, there is a phenomenon where certain modalities dominate the decision-making, leading to insufficient learning from other modalities. Moreover, redundant information within multimodal data can also hinder accurate classification decisions. Therefore, we propose a robust multimodal representation learning method for AD classification. Specifically, we first construct dedicated encoders for each multimodal data, including structural Magnetic Resonance Imaging (sMRI) images, Positron Emission Tomography (PET) images, and Mini-Mental State Examination (MMSE) scores, to extract their respective representations. Then, we employ the information bottleneck (IB) theory to guide the model to retain classification-related information in multimodal representations while reducing redundancy among modalities. Furthermore, to promote a balanced fusion of multimodal data, we redefine the classification confidence of each modality’s representation using an orthogonal weight classifier and then introduce a regularization term to amplify the prediction score differences for modalities with lower confidence. The experimental results on the Alzheimer’s Disease Neuroimaging Initiative (ADNI) dataset demonstrate that our method enhances the robustness of multimodal representations and achieves promising performance in AD-related classification tasks.
Yulan Dai, Beiji Zou 0001, Xiaoyan Kui, Zexin Ji, Chengzhang Zhu
ICIP1
2024 Strong Multimodal Representation Learner through Cross-domain Distillation for Alzheimer's Disease Classification
abstract
Vision-language foundational models have achieved commendable results on related tasks. However, their application to medical tasks is still limited due to issues arising from data biases. Currently, leveraging existing foundational models to improve medical tasks remains a challenge. To this end, this paper proposes a strong multimodal representation learning method based on cross-domain distillation handling structural Magnetic Resonance Imaging (sMRI), Positron Emission Computed Tomograph (PET) images, and mini-mental state examination (MMSE) score for Alzheimer’s disease (AD) classification. Specifically, we establish a text-to-image cross-domain distillation learning framework, enabling a text encoder pre-trained on general visual recognition tasks to guide the training of sMRI and PET image feature extractors. Simultaneously, positional encoding is used to extract the magnitude features of MMSE scores. Based on the multimodal representations extracted from sMRI, PET images, and MMSE scores, we perform a self-attention operation equipped with a gating mechanism for multimodal feature fusion. This mechanism controls the contribution of each modality representation to the classification decision, dynamically strengthening or weakening specific modality representations and helping construct stronger fused features for AD classification. Our method undergoes 5-fold cross-validation on the widely used ADNI dataset, and comparative experimental results demonstrate that our method achieves advanced performance in two AD-related binary classification tasks.
Yulan Dai, Beiji Zou 0001, Xiaoyan Kui, Qinsong Li, Wei Zhao 0040, Jun Liu 0075, Miguel Bordallo López
BIBM1
2024 Deep learning-based magnetic resonance image super-resolution: a survey
Zexin Ji, Beiji Zou 0001, Xiaoyan Kui, Jun Liu 0075, Wei Zhao 0040, Chengzhang Zhu, Peishan Dai, Yulan Dai
Neural Comput. Appl.8
2023 Wavelet-aware Transformer Network for Multi-contrast Knee MRI Super-resolution
abstract
In this paper, we propose a wavelet-aware transformer network (WATNet) for multi-contrast knee MRI super-resolution. Unlike conventional image domain-based super-resolution methods that can not explicitly model the lost high-frequency information, our WATNet endeavors to adaptively fuse the complementary frequency information of the multi-contrast image in the wavelet domain and further refine it in the image domain. The proposed WATNet consists of the multi-scale wavelet transformation (MSWT) module, wavelet-aware transformer (WAT) module, and reconstruction (Rec) module. Specifically, the MSWT module learns to transform the MR image to multi-scale wavelet domain features by the wavelet transformation. The WAT module can adaptively search and transfer similar wavelet domain reference information to the low-resolution one. The Rec module can restore high-quality images in the image domain. To further capture more high-frequency details, we also design the wavelet-based high-frequency loss. The qualitative and quantitative experimental results indicate that our proposed WATNet outperforms most state-of-the-art methods.
Zexin Ji, Xiaoyan Kui, Chengzhang Zhu, Yang Li 0111, Yulan Dai, Beiji Zou 0001
BIBM6
2022 A dynamic multi-modal fusion network for ovarian tumor differentiation
abstract
Accurate ovarian tumor differentiation is a challenging task where the benign and malignant tumors share similar T1C and T2WI MRI appearances. Therefore, it is necessary to leverage additional multi-modal data, e.g., the age, CA125level, and other clinical information, which are helpful but rarely exploited. In this paper, we propose a dynamic fusion network that can adaptively make full use of multi-modal data, including MRI and clinical information, to realize precise ovarian tumor differentiation. Specifically, we design a dynamic nonlinear module (D-Non-L module) on the top of the image representation. The D-Non-L module is formulated as an iterative nonlinear projection parameterized by the learned features of the patient-wise clinical information. With the help of this module, the interaction between clinical features and image features could be achieved to adaptively improve the discrimination of visual representations. Moreover, we construct a dual-path-based architecture to fully exploit the complementary information from T1C and T2WI MRIs. Extensive experimental results on the locally organized ovarian tumor dataset demonstrate that our methods are superior to the single-modal and single-path-based methods. And the proposed dynamic non-linear module obtains the best performance compared with other multi-modal fusion strategies.
Yang Li 0111, Beiji Zou 0001, Yulan Dai, Harrison X. Bai, Zhicheng Jiao
BIBM4
2022 Parameter-Free Latent Space Transformer for Zero-Shot Bidirectional Cross-modality Liver Segmentation
Yang Li 0111, Beiji Zou 0001, Yulan Dai, Chengzhang Zhu, Fan Yang 0054, Xin Li 0079, Harrison X. Bai, Zhicheng Jiao
MICCAI (4)3
2022 Improved PSP-Net Segmentation Network for Automatic Detection of Neovascularization in Color Fundus Images
abstract
Proliferative Diabetic Retinopathy (PDR) is a seri-ous retinal disease threatening diabetic patients. Intense retinal neovascularization in the retinal image is the most important clinical symptom of PDR, leading to visual distortion if not controlled. Accurate and timely detection of neovascularization from retinal images allows patients to receive adequate treatment to avoid further vision loss. In this work, we propose a retinal neovascularization automatic segmentation model based on im-proved Pyramid Scene Parsing Network (PSP-Net). To improve the accuracy of the model, we introduce the proposed channel attention module into the model. The network is evaluated with color fundus images from practice. Evaluation results show the network is superior to FCN, SegNet, U-Net and PSP-Net in accuracy and sensitivity. The model could achieve accuracy, sensitivity, specificity, precision and Jaccard similarity score of 0.9832,0.9265,0.9897,0.9116 and 0.8501, respectively. This paper proves through plenty of experimental results that the network model is able to improve the accuracy of segmentation, relieve the workload of doctors, and is worthy of further clinical promotion.
Qiuming Liu, Yulan Dai, Ruoxuan Zhou
VCIP3
2022 OVS-Net: An effective feature extraction network for optical coherence tomography angiography vessel segmentation
abstract
Abstract Optical coherence tomography angiography (OCTA), as a noninvasive imaging modality, has been widely used in clinical ophthalmology. However, the segmentation of retinal vessels in OCTA is under‐studied due to OCTA is a relatively new technology. In this article, an effective feature extraction network, OVS‐Net, is proposed for OCTA vessel segmentation. The OVS‐Net is divided into coarse stage and refine stage which structures are basically the same. In each stage, we utilize OctaveResBlock as the basic block to better extract the hierarchical multifrequency features of OCTA and capture the multiscale semantic features of the vessels. In order to improve the feature characterization, feature enhanced attention block is introduced into the network, which is proved to be more conducive for microvessel segmentation in our experiments. Multiscale feature blocks are introduced into the network to promote the deep integration of semantic features at different scales. Experiments on OCTA‐SS and OCTA‐500 datasets show that our proposed OVS‐Net achieves more competitive segmentation results than the existing methods, especially for microvessel segmentation.
Chengzhang Zhu, Han Wang 0064, Yalong Xiao, Yulan Dai, Beiji Zou 0001
Comput. Animat. Virtual Worlds4
2021 Multi-Label Classification Scheme Based on Local Regression for Retinal Vessel Segmentation
abstract
Segmenting small retinal vessels with width less than 2 pixels in fundus images is a challenging task. In this paper, in order to effectively segment the vessels, especially the narrow parts, we propose a local regression scheme to enhance the narrow parts, along with a novel multi-label classification method based on this scheme. We consider five labels for blood vessels and background in particular: the center of big vessels, the edge of big vessels, the center as well as the edge of small vessels, the center of background, and the edge of background. We first determine the multi-label by the local de-regression model according to the vessel pattern from the ground truth images. Then, we train a convolutional neural network (CNN) for multi-label classification. Next, we perform a local regression method to transform the previous multi-label into binary label to better locate small vessels and generate an entire retinal vessel image. Our method is evaluated using two publicly available datasets and compared with several state-of-the-art studies. The experimental results have demonstrated the effectiveness of our method in segmenting retinal vessels.
Beiji Zou 0001, Yulan Dai, Qi He 0008, Chengzhang Zhu
IEEE ACM Trans. Comput. Biol. Bioinform.2