VLDB 2026 Research / reviewers in the wild / expert
Xiaowei Yu 0001
dblp:69/10144-1
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2026
0009-0009-4583-0669ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical ImagingabstractMedical images exhibit latent anatomical groupings, such as organs, tissues, and pathological regions, that standard Vision Transformers (ViTs) fail to exploit. While recent work like SBM-Transformer attempts to incorporate such structures through stochastic binary masking, they suffer from non-differentiability, training instability, and the inability to model complex community structure. We present DCMM-Transformer, a novel ViT architecture for medical image analysis that incorporates a Degree-Corrected Mixed-Membership (DCMM) model as an additive bias in self-attention. Unlike prior approaches that rely on multiplicative masking and binary sampling, our method introduces community structure and degree heterogeneity in a fully differentiable and interpretable manner. Comprehensive experiments across diverse medical imaging datasets, including brain, chest, breast, and ocular modalities, demonstrate the superior performance and generalizability of the proposed approach. Furthermore, the learned group structure and structured attention modulation substantially enhance interpretability by yielding attention maps that are anatomically meaningful and semantically coherent. Huimin Cheng, Xiaowei Yu 0001, Shushan Wu, Luyang Fang, Jing Zhang 0010, Tianming Liu 0001, Dajiang Zhu, Wenxuan Zhong, Ping Ma 0001 |
AAAI | 2 |
| 2026 | Triplet longitudinal masked autoencoder for predicting individualized functional connectome development during infancy
Weiran Xia, Xin Zhang 0013, Dan Hu 0004, Xiaowei Yu 0001, Weiyan Yin, Zhengwang Wu, Li Wang 0026, Weili Lin, Gang Li 0001 |
Medical Image Anal. | 5 |
| 2025 | ECHOPulse: ECG Controlled Echocardio-gram Video GenerationabstractEchocardiography (ECHO) is essential for cardiac assessments, but its video quality and interpretation heavily relies on manual expertise, leading to inconsistent results from clinical and portable devices. ECHO video generation offers a solution by improving automated monitoring through synthetic data and generating high-quality videos from routine health data. However, existing models often face high computational costs, slow inference, and rely on complex conditional prompts that require experts' annotations. To address these challenges, we propose ECHOPulse, an ECG-conditioned ECHO video generation model. ECHOPulse introduces two key advancements: (1) it accelerates ECHO video generation by leveraging VQ-VAE tokenization and masked visual token modeling for fast decoding, and (2) it conditions on readily accessible ECG signals, which are highly coherent with ECHO videos, bypassing complex conditional prompts. To the best of our knowledge, this is the first work to use time-series prompts like ECG signals for ECHO video generation. ECHOPulse not only enables controllable synthetic ECHO data generation but also provides updated cardiac function information for disease monitoring and prediction beyond ECG alone. Evaluations on three public and private datasets demonstrate state-of-the-art performance in ECHO video generation across both qualitative and quantitative measures. Additionally, ECHOPulse can be easily generalized to other modality generation tasks, such as cardiac MRI, fMRI, and 3D CT generation. We will make the synthetic ECHO dataset, along with the code and model, publicly available upon acceptance. Yiwei Li 0002, Sekeun Kim, Zihao Wu 0001, Hanqi Jiang, Yi Pan 0001, Pengfei Jin, Sifan Song, Xiaowei Yu 0001, Tianze Yang, Tianming Liu 0001, Quanzheng Li, Xiang Li 0001 |
ICLR | 9 |
| 2025 | A Unified Continuous Staging Framework for Alzheimer's Disease and Lewy Body Dementia via Hierarchical Anatomical Features
Minheng Chen, Jing Zhang 0010, Xiaowei Yu 0001, Yanjun Lyu, Lu Zhang 0050, Tianming Liu 0001, Dajiang Zhu |
MICCAI (3) | 6 |
| 2025 | Core-Periphery Principle Guided State Space Model for Functional Connectome Classification
Minheng Chen, Xiaowei Yu 0001, Jing Zhang 0010, Yanjun Lyu, Lu Zhang 0050, Tianming Liu 0001, Dajiang Zhu |
MICCAI (12) | 2 |
| 2025 | Domain-Adaptive Diagnosis of Lewy Body Disease with Transferability Aware Transformer
Xiaowei Yu 0001, Jing Zhang 0010, Minheng Chen, Yanjun Lyu, Lu Zhang 0050, Tianming Liu 0001, Dajiang Zhu |
MICCAI (7) | 1 |
| 2025 | Feature Fusion Transferability Aware Transformer for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims to leverage the knowledge learned from labeled source domains to improve performance on the unlabeled target domains. While Convolutional Neural Networks (CNNs) have been dominant in previous UDA methods, recent research has shown promise in applying Vision Transformers (ViTs) to this task. In this study, we propose a novel Feature Fusion Transferability Aware Transformer (FFTAT) to enhance ViT performance in UDA tasks. Our method introduces two key innovations: First, we introduce a patch discriminator to evaluate the transferability of patches, generating a transferability matrix. We integrate this matrix into self-attention, directing the model to focus on transferable patches. Second, we propose a feature fusion technique to fuse embeddings in the latent space, enabling each embedding to incorporate information from all others, thereby improving generalization. These two components work in synergy to enhance feature representation learning. Extensive experiments on widely used benchmarks demonstrate that our method significantly improves UDA performance, achieving state-of-the-art (SOTA) results. Xiaowei Yu 0001, Zao Zhang |
WACV | 1 |
| 2025 | Learning lifespan brain anatomical correspondence via cortical developmental continuity transfer
Lu Zhang 0050, Zhengwang Wu, Xiaowei Yu 0001, Yanjun Lyu, Zihao Wu 0001, Haixing Dai, Lin Zhao 0004, Li Wang 0026, Gang Li 0001, Xianqiao Wang, Tianming Liu 0001, Dajiang Zhu |
Medical Image Anal. | 3 |
| 2025 | Exploring the Trade-Offs: Unified Large Language Models vs Local Fine-Tuned Models for Highly-Specific Radiology NLI TaskabstractRecently, ChatGPT and GPT-4 have emerged and gained immense global attention due to their unparalleled performance in language processing. Despite demonstrating impressive capability in various open-domain tasks, their adequacy in highly specific fields like radiology remains untested. Radiology presents unique linguistic phenomena distinct from open-domain data due to its specificity and complexity. Assessing the performance of large language models (LLMs) in such specific domains is crucial not only for a thorough evaluation of their overall performance but also for providing valuable insights into future model design directions: whether model design should be generic or domain-specific. To this end, in this study, we evaluate the performance of ChatGPT/GPT-4 on a radiology natural language inference (NLI) task and compare it to other models fine-tuned specifically on task-related data samples. We also conduct a comprehensive investigation on ChatGPT/GPT-4’s reasoning ability by introducing varying levels of inference difficulty. Our results show that 1) ChatGPT and GPT-4 outperform other LLMs in the radiology NLI task and 2) other specifically fine-tuned Bert-based models require significant amounts of data samples to achieve comparable performance to ChatGPT/GPT-4. These findings not only demonstrate the feasibility and promise of constructing a generic model capable of addressing various tasks across different domains, but also highlight several key factors crucial for developing a unified model, particularly in a medical context, paving the way for future artificial general intelligence (AGI) systems. We release our code and data to the research community. Zihao Wu 0001, Lu Zhang 0050, Xiaowei Yu 0001, Zhengliang Liu, Lin Zhao 0004, Yiwei Li 0002, Haixing Dai, Chong Ma 0004, Gang Li 0001, Wei Liu 0146, Quanzheng Li, Dinggang Shen, Xiang Li 0001, Dajiang Zhu, Tianming Liu 0001 |
IEEE Trans. Big Data | 4 |
| 2025 | Core-Periphery Multi-Modality Feature Alignment for Zero-Shot Medical Image AnalysisabstractMulti-modality learning, exemplified by the language-image pair pre-trained CLIP model, has demonstrated remarkable performance in enhancing zero-shot capabilities and has gained significant attention recently. However, simply applying language-image pre-trained CLIP to medical image analysis encounters substantial domain shifts, resulting in severe performance degradation due to inherent disparities between natural (non-medical) and medical image characteristics. To address this challenge and uphold or even enhance CLIP's zero-shot capability in medical image analysis, we develop a novel approach, Core-Periphery feature alignment for CLIP (CP-CLIP), to model medical images and corresponding clinical text jointly. To achieve this, we design an auxiliary neural network whose structure is organized by the core-periphery (CP) principle. This auxiliary CP network not only aligns medical image and text features into a unified latent space more efficiently but also ensures alignment driven by principles of brain network organization. In this way, our approach effectively mitigates and further enhances CLIP's zero-shot performance in medical image analysis. More importantly, the proposed CP-CLIP exhibits excellent explanatory capability, enabling the automatic identification of critical disease-related regions in clinical analysis. Extensive experiments and evaluation across five public datasets covering different diseases underscore the superiority of our CP-CLIP in zero-shot medical image prediction and critical features detection, showing its promising utility in multimodal feature alignment in current medical applications. Xiaowei Yu 0001, Lu Zhang 0050, Zihao Wu 0001, Dajiang Zhu |
IEEE Trans. Medical Imaging | 1 |
| 2024 | InterLUDE: Interactions between Labeled and Unlabeled Data to Enhance Semi-Supervised LearningabstractSemi-supervised learning (SSL) seeks to enhance task performance by training on both labeled and unlabeled data. Mainstream SSL image classification methods mostly optimize a loss that additively combines a supervised classification objective with a regularization term derived *solely* from unlabeled data. This formulation often neglects the potential for interaction between labeled and unlabeled images. In this paper, we introduce InterLUDE, a new approach to enhance SSL made of two parts that each benefit from labeled-unlabeled interaction. The first part, embedding fusion, interpolates between labeled and unlabeled embeddings to improve representation learning. The second part is a new loss, grounded in the principle of consistency regularization, that aims to minimize discrepancies in the model's predictions between labeled versus unlabeled inputs. Experiments on standard closed-set SSL benchmarks and a medical SSL task with an uncurated unlabeled set show clear benefits to our approach. On the STL-10 dataset with only 40 labels, InterLUDE achieves **3.2%** error rate, while the best previous method reports 6.3%. Xiaowei Yu 0001, Dajiang Zhu, Michael C. Hughes |
ICML | 2 |
| 2024 | CP-CLIP: Core-Periphery Feature Alignment CLIP for Zero-Shot Medical Image Analysis
Xiaowei Yu 0001, Zihao Wu 0001, Lu Zhang 0050, Jing Zhang 0010, Yanjun Lyu, Dajiang Zhu |
MICCAI (3) | 1 |
| 2024 | Gyri vs. Sulci: Core-Periphery Organization in Functional Brain Networks
Xiaowei Yu 0001, Lu Zhang 0050, Yanjun Lyu, Jing Zhang 0010, Tianming Liu 0001, Dajiang Zhu |
MICCAI (12) | 1 |
| 2024 | Eye-gaze Guided Multi-modal Alignment for Medical Representation LearningabstractIn the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly aligned from the data, without considering the explicit relationships in the medical context. This data-reliance may lead to low generalization of the learned alignment relationships. In this work, we propose the Eye-gaze Guided Multi-modal Alignment (EGMA) framework to harness eye-gaze data for better alignment of medical visual and textual features. We explore the natural auxiliary role of radiologists' eye-gaze data in aligning medical images and text, and introduce a novel approach by using eye-gaze data, collected synchronously by radiologists during diagnostic evaluations. We conduct downstream tasks of image classification and image-text retrieval on four medical datasets, where EGMA achieved state-of-the-art performance and stronger generalization across different datasets. Additionally, we explore the impact of varying amounts of eye-gaze data on model performance, highlighting the feasibility and utility of integrating this auxiliary data into multi-modal alignment framework. Chong Ma 0004, Hanqi Jiang, Wenting Chen, Yiwei Li 0002, Zihao Wu 0001, Xiaowei Yu 0001, Zhengliang Liu, Lei Guo 0002, Dajiang Zhu, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001 |
NeurIPS | 6 |
| 2024 | Real-time Core-Periphery Guided ViT with Smart Data Layout Selection on Mobile DevicesabstractMobile devices have become essential enablers for AI applications, particularly in scenarios that require real-time performance. Vision Transformer (ViT) has become a fundamental cornerstone in this regard due to its high accuracy. Recent efforts have been dedicated to developing various transformer architectures that offer im- proved accuracy while reducing the computational requirements. However, existing research primarily focuses on reducing the theoretical computational complexity through methods such as local attention and model pruning, rather than considering realistic performance on mobile hardware. Although these optimizations reduce computational demands, they either introduce additional overheads related to data transformation (e.g., Reshape and Transpose) or irregular computation/data-access patterns. These result in significant overhead on mobile devices due to their limited bandwidth, which even makes the latency worse than vanilla ViT on mobile. In this paper, we present ECP-ViT, a real-time framework that employs the core-periphery principle inspired by the brain functional networks to guide self-attention in ViTs and enable the deployment of ViT models on smartphones. We identify the main bottleneck in transformer structures caused by data transformation and propose a hardware-friendly core-periphery guided self-attention to decrease computation demands. Additionally, we design the system optimizations for intensive data transformation in pruned models. ECP-ViT, with the proposed algorithm-system co-optimizations, achieves a speedup of 4.6× to 26.9× on mobile GPUs across four datasets: STL-10, CIFAR100, TinyImageNet, and ImageNet. Zhihao Shu, Xiaowei Yu 0001, Zihao Wu 0001, Wenqi Jia 0003, Yinchen Shi, Miao Yin, Tianming Liu 0001, Dajiang Zhu, Wei Niu 0002 |
NeurIPS | 2 |
| 2022 | Longitudinal Infant Functional Connectivity Prediction via Conditional Intensive Triplet Network
Xiaowei Yu 0001, Dan Hu 0004, Lu Zhang 0050, Ying Huang 0007, Zhengwang Wu, Tianming Liu 0001, Li Wang 0026, Weili Lin, Dajiang Zhu, Gang Li 0001 |
MICCAI (8) | 1 |