EDBT 2026 Demo / reviewers in the wild / expert
Zilin Lu
dblp:331/5190
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0003-2437-283XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decoupling Target Semantics via Text-Anchored Visual Contrast for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning (SSL) provides an effective means of reducing reliance on large-scale annotated datasets by leveraging unlabeled data. However, existing SSL methods often struggle with semantic ambiguity, especially under limited supervision. Recent studies have incorporated textual information to provide contextual guidance, yet most focus on feature fusion rather than emphasizing target semantics critical for segmentation. In this paper, we proposed a novel Text-anchored Visual Decoupling (TeViD) framework for semi-supervised medical image segmentation. TeViD is built upon a teacher-student architecture with a dual-decoder design that explicitly disentangles target and background representations using both labeled and unlabeled data. For unlabeled data, a reversed cross-supervision mechanism is introduced to enhance decoder diversity and semantic separation. Furthermore, two contrastive learning objectives are proposed: a teacher-guided visual contrastive loss and a text-anchored contrastive loss, both designed to reinforce semantic disentanglement from visual and textual perspectives. Extensive experiments on five public datasets (covering X-ray, pathology, ultrasound, MRI, and CT) demonstrate that TeViD consistently outperforms both standard SSL and text-enhanced SSL methods, achieving average improvements of 5.72% in Dice and 8.15% in mIoU over the second-best competitor. The code is available at: https://github.com/jgfiuuuu/TeViD. Qingjie Zeng, Xinke Ma, Zilin Lu, Mengkang Lu, Yanning Zhang 0001, Yong Xia 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Harnessing Text Insights With Visual Alignment for Medical Image SegmentationabstractPre-trained vision-language models (VLMs) and language models (LMs) have recently garnered significant attention due to their remarkable ability to represent textual concepts, opening up new avenues in vision tasks. In medical image segmentation, efforts are being made to integrate text and image data using VLMs and LMs. However, current text-enhanced approaches face several challenges. First, using separate pre-trained vision and text models to encode image and text data can result in semantic shifts. Second, while VLMs can establish the correspondence between visual and textual features when pre-trained on paired image-text data, this alignment often deteriorates during segmentation tasks due to misalignment between the text and vision components in ongoing learning. In this paper, we propose TeViA, a novel approach that seamlessly integrates with various vision and text models, irrespective of their pre-training relationships. This integration is achieved through a segmentation-specific text-to-vision alignment design, ensuring both information gain and semantic consistency. Specifically, for each training data, a foreground visual representation is extracted from the segmentation head and used to supervise projection layers, thereby adjusting the textual features to better contribute to the segmentation task. Additionally, a historic visual prototype is created by aggregating target semantics from all training data and is updated using a momentum-based manner. This prototype aims to enhance the visual representation of each data instance by establishing feature-level connections, which in turn refines the textual features. The superiority of TeViA is validated on five public datasets, exhibiting over 6% Dice improvements compared to vision-only methods. Code is available at: https://github.com/jgfiuuuu/TeViA. Qingjie Zeng, Zilin Lu, Yutong Xie 0001, Zhiyong Wang 0001, Yanning Zhang 0001, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Exploring Text-Enhanced Mixture-of-Experts for Semi-supervised Medical Image Segmentation with Composite Data
Qingjie Zeng, Xinke Ma, Zilin Lu, Yong Xia 0001 |
MICCAI (6) | 4 |
| 2025 | PICK: Predict and Mask for Semi-supervised Medical Image Segmentation
Qingjie Zeng, Zilin Lu, Yutong Xie 0001, Yong Xia 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | PathBot: A Foundation Model for Pathological Image AnalysisabstractComputational pathology has emerged as a transformative paradigm by leveraging artificial intelligence to automate and enhance diagnostic procedures. However, existing models often target narrow tasks or specific tumor types, missing opportunities to unify diverse datasets and tasks through joint learning. In this work, we introduce PathBot, a foundation model tailored for comprehensive pathological image analysis. Central to PathBot is a ViT-Giant encoder with one billion parameters, the largest model to date trained on publicly available pathological data. We pre-train this encoder using a novel Masked Distillation Network (MDN) and an integrated learning strategy that combines contrastive and generative objectives. The pre-training leverages over 30 million image patches derived from 11,765 whole slide images (WSIs) across 32 cancer types in the Cancer Genome Atlas (TCGA). To evaluate its versatility, we pair the encoder with task-specific decoders for segmentation, detection, classification, and regression. Extensive experiments across 20 downstream tasks demonstrate that PathBot achieves state-of-the-art performance in most cases, showcasing its robustness and generalizability. Mengkang Lu, Qingjie Zeng, Zilin Lu, Zhe Li 0006, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Bridging the Semantic Gap in Medical Visual Question Answering With Prompt LearningabstractMedical Visual Question Answering (Med-VQA) aims to answer questions regarding the content of medical images, crucial for enhancing diagnostics and education in healthcare. However, progress in this field is hindered by data scarcity due to the resource-intensive nature of medical data annotation. While existing Med-VQA approaches often rely on pre-training to mitigate this issue, bridging the semantic gap between pre-trained models and specific tasks remains a significant challenge. This paper presents the Dynamic Semantic-Adaptive Prompting (DSAP) framework, leveraging prompt learning to enhance model performance in Med-VQA. To this end, we introduce two prompting strategies: Semantic Alignment Prompting (SAP) and Dynamic Question-Aware Prompting (DQAP). SAP prompts multi-modal inputs during fine-tuning, reducing the semantic gap by aligning model outputs with domain-specific contexts. Simultaneously, DQAP enhances answer selection by leveraging grammatical relationships between questions and answers, thereby improving accuracy and relevance. The DSAP framework was pre-trained on three datasets-ROCO, MedICaT, and MIMIC-CXR-and comprehensively evaluated against 15 existing Med-VQA models on three public datasets: VQA-RAD, SLAKE, and PathVQA. Our results demonstrate a substantial performance improvement, with DSAP achieving a 1.9% enhancement in average results across benchmarks. These findings underscore DSAP's effectiveness in addressing critical challenges in Med-VQA and suggest promising avenues for future developments in medical AI. Zilin Lu, Qingjie Zeng, Mengkang Lu, Geng Chen 0001, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Segment Together: A Versatile Paradigm for Semi-Supervised Medical Image SegmentationabstractThe scarcity of annotations has become a significant obstacle in training powerful deep-learning models for medical image segmentation, limiting their clinical application. To overcome this, semi-supervised learning that leverages abundant unlabeled data is highly desirable to enhance model training. However, most existing works still focus on specific medical tasks and underestimate the potential of learning across diverse tasks and datasets. In this paper, we propose a Versatile Semi-supervised framework (VerSemi) to present a new perspective that integrates various SSL tasks into a unified model with an extensive label space, exploiting more unlabeled data for semi-supervised medical image segmentation. Specifically, we introduce a dynamic task-prompted design to segment various targets from different datasets. Next, this unified model is used to identify the foreground regions from all labeled data, capturing cross-dataset semantics. Particularly, we create a synthetic task with a CutMix strategy to augment foreground targets within the expanded label space. To effectively utilize unlabeled data, we introduce a consistency constraint that aligns aggregated predictions from various tasks with those from the synthetic task, further guiding the model to accurately segment foreground regions during training. We evaluated our VerSemi framework against seven established SSL methods on four public benchmarking datasets. Our results suggest that VerSemi consistently outperforms all competing methods, beating the second-best method with a 2.69% average Dice gain on four datasets and setting a new state of the art for semi-supervised medical image segmentation. Code is available at https://github.com/maxwell0027/VerSemi. Qingjie Zeng, Yutong Xie 0001, Zilin Lu, Mengkang Lu, Yicheng Wu 0001, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Consistency-Guided Differential Decoding for Enhancing Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning (SSL) has been proven beneficial for mitigating the issue of limited labeled data, especially on volumetric medical image segmentation. Unlike previous SSL methods which focus on exploring highly confident pseudo-labels or developing consistency regularization schemes, our empirical findings suggest that differential decoder features emerge naturally when two decoders strive to generate consistent predictions. Based on the observation, we first analyze the treasure of discrepancy in learning towards consistency, under both pseudo-labeling and consistency regularization settings, and subsequently propose a novel SSL method called LeFeD, which learns the feature-level discrepancies obtained from two decoders, by feeding such information as feedback signals to the encoder. The core design of LeFeD is to enlarge the discrepancies by training differential decoders, and then learn from the differential features iteratively. We evaluate LeFeD against eight state-of-the-art (SOTA) methods on three public datasets. Experiments show LeFeD surpasses competitors without any bells and whistles, such as uncertainty estimation and strong constraints, as well as setting a new state of the art for semi-supervised medical image segmentation. Code has been released at https://github.com/maxwell0027/LeFeD. Qingjie Zeng, Yutong Xie 0001, Zilin Lu, Mengkang Lu, Jingfeng Zhang, Yong Xia 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Spot the Difference: Difference Visual Question Answering with Residual Alignment
Zilin Lu, Yutong Xie 0001, Qingjie Zeng, Mengkang Lu, Qi Wu 0001, Yong Xia 0001 |
MICCAI (5) | 1 |
| 2024 | Reciprocal Collaboration for Semi-supervised Medical Image Classification
Qingjie Zeng, Zilin Lu, Yutong Xie 0001, Mengkang Lu, Xinke Ma, Yong Xia 0001 |
MICCAI (11) | 2 |
| 2023 | PEFAT: Boosting Semi-Supervised Medical Image Classification via Pseudo-Loss Estimation and Feature Adversarial TrainingabstractPseudo-labeling approaches have been proven beneficial for semi-supervised learning (SSL) schemes in computer vision and medical imaging. Most works are dedicated to finding samples with high-confidence pseudo-labels from the perspective of model predicted probability. Whereas this way may lead to the inclusion of incorrectly pseudo-labeled data if the threshold is not carefully adjusted. In addition, low-confidence probability samples are frequently disregarded and not employed to their full potential. In this paper, we propose a novel Pseudo-loss Estimation and Feature Adversarial Training semi-supervised framework, termed as PEFAT, to boost the performance of multi-class and multi-label medical image classification from the point of loss distribution modeling and adversarial training. Specifically, we develop a trustworthy data selection scheme to split a high-quality pseudo-labeled set, inspired by the dividable pseudo-loss assumption that clean data tend to show lower loss while noise data is the opposite. Instead of directly discarding these samples with low-quality pseudo-labels, we present a novel regularization approach to learn discriminate information from them via injecting adversarial noises at the feature-level to smooth the decision boundary. Experimental results on three medical and two natural image benchmarks validate that our PEFAT can achieve a promising performance and surpass other state-of-the-art methods. The code is available at https://github.com/maxwell0027/PEFAT. Qingjie Zeng, Yutong Xie 0001, Zilin Lu, Yong Xia 0001 |
CVPR | 3 |
| 2022 | Robot control with multitasking of brain-computer interfaceabstractBrain-computer interfaces (BCI) have been extensively researched to assist people with motor paralysis in controlling external devices such as a robotic limb. However, most BCI systems required participants to focus on a single task, limiting their ability to generate other mental or physical activities. Therefore, people's performance of the BCI-based robotic control in multitasking was discussed, as eight healthy subjects performed motor-related tasks of motor imagery and two-handed balancing ball movement, while simultaneously performing visuospatial attention to asynchronously trigger “drinking” actions of a humanoid robot arm with accuracies of 90% and 87.5%, respectively. The online results indicate that the BCI-based robot control system developed for multi-task conditions has a high potential for human augmentation. Yajun Zhou, Zilin Lu, Yuanqing Li 0001 |
ICARCV | 2 |