EDBT 2026 Demo / reviewers in the wild / expert
Xiaobing Yang
dblp:57/370
· DBLP profile ↗
8ranked-venue papers
0as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CIMB-MVQA: Causal intervention on modality-specific biases for medical visual question answering
Jiaman Ding, Xiaobing Yang |
Medical Image Anal. | 4 |
| 2026 | Beyond Static Knowledge: Dynamic Context-Aware Cross-Modal Contrastive Learning for Medical Visual Question AnsweringabstractMedical Visual Question Answering (Med-VQA) aims to analyze medical images and accurately respond to natural language queries, thereby optimizing clinical workflows and improving diagnostic and therapeutic outcomes. Although medical images contain rich visual information, the corresponding textual queries frequently lack sufficient descriptive content. This imbalance of information and modality differences leads to significant semantic bias. Furthermore, existing approaches integrate external medical knowledge to enhance model performance, they primarily rely on static knowledge that lacks dynamic adaptation to specific input samples, leading to redundant information and noise interference. To address these challenges, we propose a Contextual Knowledge-Aware Dynamic Perception for the Cross-Modal Reasoning and Alignment (CKRA) Model. To mitigate knowledge redundancy, CKRA employs a dynamic perception mechanism that leverages semantic cues from the query to selectively filter relevant medical knowledge specific to the current sample's context. To alleviate cross-modal semantic bias, CKRA bridges the distance between visual and linguistic features through knowledge-image contrastive learning, optimizing knowledge feature representation and directing the model's attention to key image regions. Further, we design a dual-stream guided attention network that facilitates cross-modal interaction and alignment across multiple dimensions. Experimental results show that the proposed CKRA model outperforms the state-of-the-art method on SLAKE and VQA-RAD datasets. In addition, ablation studies validate the effectiveness of each module, while Grad-CAM maps further demonstrate the feasibility of CKRA for medical visual questioning tasks. The source code and weights of the model are available at https://github.com/cloneiq/CKRA-MedVQA. Xupeng Feng, Wei Peng 0004, Xiaobing Yang |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Polyp-MoE: Modeling Temporal Consistency and Structural Priors via Mixture-of-Experts for Video Polyp SegmentationabstractAutomatic video polyp segmentation is vital for early colorectal cancer screening and assisted diagnosis. Yet, endoscopic videos often suffer from motion blur, specular reflections, and occlusions, which disrupt cross-frame continuity and structural integrity. We propose Polyp-MoE, a novel framework that unifies temporal modeling with a structure-aware mixture-of-experts (MoE) design. By jointly exploiting temporal consistency and expert-prior guidance, Polyp-MoE effectively localizes lesions and restores complete structures in degraded frames. Specifically, a Temporal Consistency Guided Memory (TCGM) module establishes dynamic spatiotemporal correspondences between adjacent frames for robust temporal alignment, while a Structural Prior Mixture-of-Experts (SPMoE) module extends the MoE paradigm to medical video segmentation by encoding location, size, and boundary priors through three learnable experts and adaptively fusing them via a lightweight router. Experiments on the SUN-SEG dataset show that Polyp-MoE consistently surpasses state-of-the-art methods, achieving Dice gains of +2.4% and +2.7% on the challenging Seen-Hard and Unseen-Hard subsets, confirming its strong temporal modeling and structural reconstruction capability. Code is available at https://github.com/cloneiq/Polyp-MoE Juntong Ti, Xiaobing Yang |
BIBM | 3 |
| 2025 | Body Part-Aware Cross-Modal Feature Interaction Learning for Medical Vision Question Answering
Shuting Dai, Hao Cong, Xiaobing Yang |
ICIC (5) | 4 |
| 2024 | Cross-Guided Attention and Combined Loss for Optimizing Gastrointestinal Visual Question AnsweringabstractMedical Visual Question Answering (Med-VQA) aims at answering clinical questions related to medical imaging, providing important reference for medical imaging diagnosis. Gastrointestinal visual question answering can deeply understand and analyze gastrointestinal medical images, provides supporting information to answer questions about the diagnostic process, and improves diagnostic accuracy and efficiency. However, most existing Med-VQA can only handle simple and general questions, struggling to tackle visual question answering related to complex diseases. Therefore, this paper proposes a Med-VQA model called CACL, which uses cross-guidance and improved combined loss. The model focuses on gastrointestinal image analysis and achieves efficient and accurate feature extraction for gastrointestinal images. A cross-guided attention module is also designed to enhance the model’s reasoning ability when dealing with complex cross-modal tasks. In addition, a multi-task composite loss function is designed to balance the loss of segmentation task and classification task, and improve the overall performance. Experimental results show that the proposed method can effectively improve the accuracy of gastrointestinal visual question answering. Xiaobing Yang |
BIBM | 3 |
| 2024 | Multimodal Online Knowledge Distillation Framework for Land Use/Cover Classification Using Full or Missing ModalitiesabstractMultimodal land use/cover classification using optical and synthetic aperture radar (SAR) images has attracted significant attention because the unique radiation and geometric characteristics of these images provide complementary information regarding land properties. However, the significant differences between these modalities create a large semantic gap, posing challenges for effective feature fusion in multimodal learning. Moreover, missing modalities often occur in practical applications due to weather constraints or sensor malfunctions, posing challenges to achieving high performance in cross-modal learning. In this study, we proposed a multimodal online knowledge distillation (MMOKD) framework, designed for land use/cover classification of optical and SAR images using either full or missing modalities. This framework trains one modality-fusion network alongside two modality-specific networks in an end-to-end manner, facilitating both multimodal and cross-modal learning. More specifically, we developed a multimodal feature fusion (MFF) module for integrating heterogeneous features, and a single-modal feature generation (SFG) module for encapsulating cross-modal complementary information. Additionally, we proposed the joint distillation with multitype fusion knowledge (JD-MFK) method, guiding the modality-specific student networks to comprehensively learn the modality-fusion teacher network. Notably, we adopted an online distillation strategy for real-time feedback and synchronous updates of both modality-fusion and modality-specific networks. Finally, we conducted extensive experiments on two multimodal land use/classification datasets with advanced multimodal fusion, cross-modal distillation, and specific baseline networks for comparison. The results demonstrate the effectiveness of the proposed MMODD, which not only outperforms the other networks in both full- and missing-modality scenarios, but also significantly improves model training efficiency. Xiao Liu 0050, Fei Jin, Shuxiang Wang, Jie Rui, Xibing Zuo, Xiaobing Yang, Chuanxiang Cheng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2016 | Research on autonomous moving robot path planning based on improved particle swarm optimizationabstractTwo improved particle swarm optimization algorithms are given to overcome the defects in the commonly used particle swarm optimization. These are particle swarm optimization with nonlinear inertia weight and simulated annealing particle swarm optimization. The global search ability and local search accuracy can be optimized by introducing nonlinear inertia weight coefficients. It is well known that the particle swarm optimization has a problem that the algorithm is easily trapped into the local optimum. This paper shows that such a problem can be solved partially by combining the particle swarm optimization with simulated annealing algorithm. Autonomous moving robot path planning is given based on improved particle swarm optimization. The simulation results show the validity of the proposed improved algorithm in moving robot path planning. Zhibin Nie, Xiaobing Yang, Shihong Gao |
CEC | 2 |
| 2005 | Using Latent Class Models for Neighbors Selection in Collaborative Filtering
Fansheng Kong, Xiaobing Yang |
ADMA | 3 |