EDBT 2026 Demo / reviewers in the wild / expert
Xikai Yang
dblp:264/9423
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-1762-9684ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SurgPub-Video: A Comprehensive Surgical Video Framework for Enhanced Surgical Intelligence in Vision-Language ModelabstractVision-Language Models (VLMs) have shown significant potential in surgical scene analysis, yet existing models are limited by frame-level datasets and lack high-quality video data with procedural surgical knowledge. To address these challenges, we make the following contributions: (i) SurgPub-Video, a comprehensive dataset of over 3,000 surgical videos and 25 million annotated frames across 11 specialities, sourced from peer-reviewed clinical journals, (ii) SurgLLaVA-Video, a specialized VLM for surgical video understanding, built upon the TinyLLaVA-Video architecture that supports both video-level and frame-level inputs, and (iii) a video-level surgical Visual Question Answering (VQA) benchmark, covering diverse 11 surgical specialities, such as vascular, cardiology, and thoracic. Extensive experiments, conducted on the proposed benchmark and three additional surgical downstream tasks (action recognition, skill assessment, and triplet recognition), show that SurgLLaVA-Video significantly outperforms both general-purpose and surgical-specific VLMs with only three billion parameters. Yaoqian Li, Xikai Yang, Dunyuan Xu, Litao Zhao, Xiaowei Hu 0001, Jinpeng Li 0004, Pheng-Ann Heng |
AAAI | 2 |
| 2026 | Revolutionizing Turn-by-Turn Navigation With Cloud-Edge Deep LearningabstractTurn-by-turn (TBT) navigation systems are integral to modern driving experiences, providing real-time audio instructions to guide drivers safely to destinations. However, existing audio instruction policy often relies on rule-based approaches that struggle to balance informational content with cognitive load, potentially leading to driver confusion or missed turns in complex environments. To overcome these difficulties, we first model the generation of navigation instructions as a multi-task learning problem by decomposing the audio content into combinations of modular elements. Then, we propose a novel deep learning framework that leverages the powerful spatiotemporal information processing capabilities of Transformers and the strong multi-task learning abilities of Mixture of Experts (MoE) to generate real-time, context-aware audio instructions for TBT driving navigation. A cloud-edge collaborative architecture is implemented to handle the computational demands of the model, ensuring scalability and real-time performance for practical applications. Experimental results in the real world demonstrate that the proposed method significantly reduces the yaw rate (the proportion of vehicles deviating from navigation routes) compared to traditional methods, delivering clearer and more effective audio instructions. This is the first large-scale application of deep learning in driving audio navigation, marking a substantial advancement in intelligent transportation and driving assistance technologies. Fanxiang Zeng, Xikai Yang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Sequence-Independent Continual Test-Time Adaptation with Mixture of Incremental Experts for Cross-Domain Segmentation
Dunyuan Xu, Yuchen Yuan, Xikai Yang, Jingyang Zhang, Jinpeng Li 0004, Pheng-Ann Heng |
MICCAI (16) | 4 |
| 2025 | Medical Large Vision Language Models with Multi-image Visual Ability
Xikai Yang, Juzheng Miao, Yuchen Yuan, Qi Dou 0001, Jinpeng Li 0004, Pheng-Ann Heng |
MICCAI (5) | 1 |
| 2025 | Multi-Scale Spatio-Temporal Transformer-Based Imbalanced Longitudinal Learning for Glaucoma Forecasting From Irregular Time Series ImagesabstractGlaucoma is one of the major eye diseases that leads to progressive optic nerve fiber damage and irreversible blindness, afflicting millions of individuals. Glaucoma forecast is a good solution to early screening and intervention of potential patients, which is helpful to prevent further deterioration of the disease. It leverages a series of historical fundus images of an eye and forecasts the likelihood of glaucoma occurrence in the future. However, the irregular sampling nature and the imbalanced class distribution are two challenges in the development of disease forecasting approaches. To this end, we introduce the Multi-scale Spatio-temporal Transformer Network (MST-former) based on the transformer architecture tailored for sequential image inputs, which can effectively learn representative semantic information from sequential images on both temporal and spatial dimensions. Specifically, we employ a multi-scale structure to extract features at various resolutions, which can largely exploit rich spatial information encoded in each image. Besides, we design a time distance matrix to scale time attention in a non-linear manner, which could effectively deal with the irregularly sampled data. Furthermore, we introduce a temperature-controlled Balanced Softmax Cross-entropy loss to address the class imbalance issue. Extensive experiments on the Sequential fundus Images for Glaucoma Forecast (SIGF) dataset demonstrate the superiority of the proposed MST-former method, achieving an AUC of 96.6% for glaucoma forecasting. Besides, our method shows excellent generalization capability on the Alzheimer's Disease Neuroimaging Initiative (ADNI) MRI dataset, with an accuracy of 88.2% for mild cognitive impairment and Alzheimer's disease prediction, outperforming the compared method by a large margin. A series of ablation studies further verify the contribution of our proposed components in addressing the irregular sampled and class imbalanced problems. Xikai Yang, Xi Wang 0013, Yuchen Yuan, Jinpeng Li 0004, Guangyong Chen, Ning Li Wang, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Effective Semi-Supervised Medical Image Segmentation With Probabilistic Representations and Prototype LearningabstractLabel scarcity, class imbalance and data uncertainty are three primary challenges that are commonly encountered in the semi-supervised medical image segmentation. In this work, we focus on the data uncertainty issue that is overlooked by previous literature. To address this issue, we propose a probabilistic prototype-based classifier that introduces uncertainty estimation into the entire pixel classification process, including probabilistic representation formulation, probabilistic pixel-prototype proximity matching, and distribution prototype update, leveraging principles from probability theory. By explicitly modeling data uncertainty at the pixel level, model robustness of our proposed framework to tricky pixels, such as ambiguous boundaries and noises, is greatly enhanced when compared to its deterministic counterpart and other uncertainty-aware strategy. Empirical evaluations on three publicly available datasets that exhibit severe boundary ambiguity show the superiority of our method over several competitors. Moreover, our method also demonstrates a stronger model robustness to simulated noisy data. Code is available at https://github.com/IsYuchenYuan/PPC. Yuchen Yuan, Xi Wang 0013, Xikai Yang, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Decoupling Feature Representations of Ego and Other Modalities for Incomplete Multi-modal Brain Tumor SegmentationabstractMulti-modal brain tumor segmentation typically involves four magnetic resonance imaging (MRI) modalities, while incomplete modalities significantly degrade performance. Existing solutions employ explicit or implicit modality adaptation, aligning features across modalities or learning a fused feature robust to modality incompleteness. They share a common goal of encouraging each modality to express both itself and the others. However, the two expression abilities are entangled as a whole in a seamless feature space, resulting in prohibitive learning burdens. In this paper, we propose DeMoSeg to enhance the modality adaptation by Decoupling the task of representing the ego and other Modalities for robust incomplete multi-modal Segmentation. The decoupling is super lightweight by simply using two convolutions to map each modality onto four feature sub-spaces. The first sub-space expresses itself (Self-feature), while the remaining sub-spaces substitute for other modalities (Mutual-features). The Self- and Mutual-features interactively guide each other through a carefully-designed Channel-wised Sparse Self-Attention (CSSA). After that, a Radiologist-mimic Cross-modality expression Relationships (RCR) is introduced to have available modalities provide Self-feature and also ‘lend’ their Mutual-features to compensate for the absent ones by exploiting the clinical prior knowledge. The benchmark results on BraTS2020, BraTS2018 and BraTS2015 verify the DeMoSeg’s superiority thanks to the alleviated modality adaptation difficulty. Concretely, for BraTS2020, DeMoSeg increases Dice by at least 0.92%, 2.95% and 4.95% on whole tumor, tumor core and enhanced tumor regions, respectively, compared to other state-of-the-arts. Codes are at https://github.com/kk42yy/DeMoSeg. Kaixiang Yang 0004, Wenqi Shan, Xikai Yang, Xi Wang 0013, Pheng-Ann Heng, Qiang Li 0018, Zhiwei Wang 0002 |
BIBM | 5 |
| 2024 | Coarse-to-Fine Latent Diffusion Model for Glaucoma Forecast on Sequential Fundus Images
Yuhan Zhang 0001, Xikai Yang, Xiao Ma 0011, Ningli Wang, Xi Wang 0013, Pheng-Ann Heng |
MICCAI (5) | 3 |
| 2023 | Semi-supervised Class Imbalanced Deep Learning for Cardiac MRI Segmentation
Yuchen Yuan, Xi Wang 0013, Xikai Yang, Ruijiang Li, Pheng-Ann Heng |
MICCAI (4) | 3 |