EDBT 2026 Demo / reviewers in the wild / expert
Qiushi Yang
dblp:89/8437
· DBLP profile ↗
22ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0001-5737-5653ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Security and privacy · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MRM++: Enhanced Masked Relation Modeling for Multi-Modal Medical Pre-trainingabstractAbstract Recent progress in deep learning for automated multi-modal medical diagnosis heavily depend on extensive expert annotations, which is time-intensive and impractical. To mitigate this, masked image modeling (MIM)-based pre-training strategies have emerged, effectively learning generalized representations from unlabelled data for various downstream tasks. Nevertheless, these approaches are tailored for natural images while neglect the distinct characteristics of medical data, resulting in suboptimal generalization in medical diagnosis applications. In this work, we attempt to harness the complementary information of multi-modal medical data to perform self-supervised pre-training and propose MRM++, an enhanced masked relation modeling paradigm. Different from the previous MIM methods that randomly mask input data, causing potentially missing of disease-relevant semantics, we devise prior-guided relation masking to break token-wise feature relation guided by anatomy-aware prior in both self- and cross-modal aspects. This can preserve complete input semantics and enable the model to learn abundant disease-related knowledge. Furthermore, to boost semantic relation modeling, the relation matching is introduced, which aligns sample-wise relations among unmasked and masked features. By exploiting inter-sample relations, the relation matching imposes the global constraints in the feature space, ensuring ample semantic relation for robust feature representation. Additionally, considering that the model may overfit to the pre-training dataset and lead to inherent gap between pre-training and downstream fine-tuning, we conceive task-oriented adapting as a pre-stage before fine-tuning to simultaneously perform self-supervised and task-supervised learning on downstream dataset. It can adaptively transform knowledge from the pre-trained model to be compatible with downstream tasks while maintaining transferable information. Extensive experiments on medical image-text and image-genome benchmarks validate the effectiveness and transfer ability of the proposed framework, outperforming state-of-the-art methods across various downstream diagnostic tasks. Source codes are made publicly available on https://github.com/CUHK-AIM-Group/MRM_plus . Qiushi Yang, Wuyang Li, Zhe Peng, Fangxiao Cheng, Yixuan Yuan |
Int. J. Comput. Vis. | 1 |
| 2026 | Order-Optimal Byzantine-Robust Learning Under Heterogeneity via Fair Gradient ClippingabstractByzantine-robust distributed or federated learning (FL) refers to providing reliable performance under Byzantine attacks, which violate the prescribed protocols and transmit arbitrary information to the server to hamper the convergence of machine learning (ML) algorithms, via designing resilient aggregation rules to combat attacks. Although numerous robust rules have been suggested, their performance degrades for heterogeneous data. A few techniques have been exploited to handle this problem, but they either require preaggregation operations, hence increasing the computational load, or lack breakdown point analysis of their rules. This article proposes a new aggregation rule, which clips the gradients received from all workers according to the distance between the gradient and the aggregation center. That is, when the distance is larger than the radius $\gamma $ , the gradient will be clipped, and the longer the distance, the closer the clipped gradient is to the center. We theoretically analyze that the breakdown point of the developed rule is 0.5, the maximum value for robust aggregators. Moreover, our rule achieves order-optimal Byzantine-robust training error under data heterogeneity, while the median-based schemes, such as coordinate-wise median (CM) and geometric median (GM), are suboptimal. Experimental results demonstrate that the devised aggregation mechanism can handle different attacks well and outperforms the existing rules. Zhi-Yong Wang, Hao Nan Sheng, Qiushi Yang, Hing-Cheung So |
IEEE Trans. Cybern. | 3 |
| 2025 | Towards Fine-Grained Interactive Segmentation in Images and VideosabstractThe recent Segment Anything Models (SAMs) have emerged as foundational visual models for general interactive segmentation. Despite demonstrating robust generalization abilities, they still suffer performance degradations in scenarios demanding accurate masks. Existing methods for high-precision interactive segmentation face a trade-off between the ability to perceive intricate local details and maintaining stable prompting capability, which hinders the applicability and effectiveness of foundational segmentation models. To this end, we present an SAM2Refiner framework built upon the SAM2 backbone. This architecture allows SAM2 to generate fine-grained segmentation masks for both images and videos while preserving its inherent strengths. Specifically, we design a localization augment module, which incorporates local contextual cues to enhance global features via a cross-attention mechanism, thereby exploiting potential detailed patterns and maintaining semantic information. Moreover, to strengthen the prompting ability toward the enhanced object embedding, we introduce a prompt retargeting module to renew the embedding with spatially aligned prompt features. In addition, to obtain accurate high resolution segmentation masks, a mask refinement module is devised by employing a multi-scale cascaded structure to fuse mask features with hierarchical representations from the encoder. Extensive experiments demonstrate the effectiveness of our approach, revealing that the proposed method can produce highly precise masks for both images and videos, surpassing state-of-the-art methods. Qiushi Yang, Miaomiao Cui, Liefeng Bo |
ICCV | 2 |
| 2025 | Polyp-Gen: Realistic and Diverse Polyp Image Generation for Endoscopic Dataset ExpansionabstractAutomated diagnostic systems (ADS) have shown significant potential in the early detection of polyps during endoscopic examinations, thereby reducing the incidence of colorectal cancer. However, due to high annotation costs and strict privacy concerns, acquiring high-quality endoscopic images poses a considerable challenge in the development of ADS. Despite recent advancements in generating synthetic images for dataset expansion, existing endoscopic image generation algorithms failed to accurately generate the details of polyp boundary regions and typically required medical priors to specify plausible locations and shapes of polyps, which limited the realism and diversity of the generated images. To address these limitations, we present Polyp-Gen, the first full-automatic diffusion-based endoscopic image generation framework. Specifically, we devise a spatial-aware diffusion training scheme with a lesion-guided loss to enhance the structural context of polyp boundary regions. Moreover, to capture medical priors for the localization of potential polyp areas, we introduce a hierarchical retrieval-based sampling strategy to match similar fine-grained spatial features. In this way, our Polyp-Gen can generate realistic and diverse endoscopic images for building reliable ADS. Extensive experiments demonstrate the state-of-the-art generation quality, and the synthetic images can improve the downstream polyp detection task. Additionally, our Polyp-Gen has shown remarkable zeroshot generalizability on other datasets. The source code is available at https://github.com/CUHK-AIM-Group/Polyp-Gen. Zhen Chen 0013, Qiushi Yang, Weihao Yu 0005, Di Dong, Jiancong Hu, Yixuan Yuan |
ICRA | 3 |
| 2025 | Relation-Guided Versatile Regularization for Federated Semi-Supervised LearningabstractAbstract Federated semi-supervised learning (FSSL) target to address the increasing privacy concerns for the practical scenarios, where data holders are limited in labeling capability. Latest FSSL approaches leverage the prediction consistency between the local model and global model to exploit knowledge from partially labeled or completely unlabeled clients. However, they merely utilize data-level augmentation for prediction consistency and simply aggregate model parameters through the weighted average at the server, which leads to biased classifiers and suffers from skewed unlabeled clients. To remedy these issues, we present a novel FSSL framework, Relation-guided Versatile Regularization (FedRVR), consisting of versatile regularization at clients and relation-guided directional aggregation strategy at the server. In versatile regularization, we propose the model-guided regularization together with the data-guided one, and encourage the prediction of the local model invariant to two extreme global models with different abilities, which provides richer consistency supervision for local training. Moreover, we devise a relation-guided directional aggregation at the server, in which a parametric relation predictor is introduced to yield pairwise model relation and obtain a model ranking. In this manner, the server can provide a superior global model by aggregating relative dependable client models, and further produce an inferior global model via reverse aggregation to promote the versatile regularization at clients. Extensive experiments on three FSSL benchmarks verify the superiority of FedRVR over state-of-the-art counterparts across various federated learning settings. Qiushi Yang, Zhen Chen 0013, Zhe Peng, Yixuan Yuan |
Int. J. Comput. Vis. | 1 |
| 2025 | FedBM: Stealing knowledge from pre-trained language models for heterogeneous federated learning
Meilu Zhu, Qiushi Yang, Zhifan Gao, Yixuan Yuan, Jun Liu 0007 |
Medical Image Anal. | 2 |
| 2025 | Progressive Distillation With Optimal Transport for Federated Incomplete Multi-Modal Learning of Brain Tumor SegmentationabstractMulti-modal Magnetic Resonance Imaging (MRI) provide sufficient complementary information for brain tumor segmentation, however, most current approaches rely on complete modalities and may collapse with incomplete modalities. Moreover, most existing endeavors focus on training with centralized databases, failing to make full use of distributed multi-silo datasets with rich patient data to learn a more robust brain tumor segmentation model. In this paper, considering the distributed training scenarios, we formulate Federated Incomplete Multi-modal Learning (FedIML) for brain tumor segmentation, and propose Progressive distiLlation with Optimal Transport (PLOT) framework to gradually train a modality robust segmentation model at each client and achieve compatible model aggregation at the server. Specifically, to remedy the issue of unstable local training caused by the random modality input, we present Modality Progressive Distillation (MPD), a multi-level knowledge distillation strategy guided by a modality routing mechanism. At each client, MPD provides a gradually learning course for a student model in an easy-to-hard manner to achieve a stable local training process. Moreover, to address the problem that the layer-wise knowledge from different models may contradict, at the server, we design Optimal Transport-guided Model Aggregation (OTMA) strategy, which yields a global alignment solution for model parameters via solving an optimal transport problem. OTMA can achieve a compatible parameter aggregation and boost the distributed training. Extensive experiments on the BraTS-2021 dataset demonstrate the effectiveness of the proposed framework over state-of-the-art methods. Qiushi Yang, Meilu Zhu, Yat Ming Peter Woo, Leanne Lai Chan, Yixuan Yuan |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Unified Multi-Modal Diagnostic Framework With Reconstruction Pre-Training and Heterogeneity-Combat TuningabstractMedical multi-modal pre-training has revealed promise in computer-aided diagnosis by leveraging large-scale unlabeled datasets. However, existing methods based on masked autoencoders mainly rely on data-level reconstruction tasks, but lack high-level semantic information. Furthermore, two significant heterogeneity challenges hinder the transfer of pre-trained knowledge to downstream tasks, i.e., the distribution heterogeneity between pre-training data and downstream data, and the modality heterogeneity within downstream data. To address these challenges, we propose a Unified Medical Multi-modal Diagnostic (UMD) framework with tailored pre-training and downstream tuning strategies. Specifically, to enhance the representation abilities of vision and language encoders, we propose the Multi-level Reconstruction Pre-training (MR-Pretrain) strategy, including a feature-level and data-level reconstruction, which guides models to capture the semantic information from masked inputs of different modalities. Moreover, to tackle two kinds of heterogeneities during the downstream tuning, we present the heterogeneity-combat downstream tuning strategy, which consists of a Task-oriented Distribution Calibration (TD-Calib) and a Gradient-guided Modality Coordination (GM-Coord). In particular, TD-Calib fine-tunes the pre-trained model regarding the distribution of downstream datasets, and GM-Coord adjusts the gradient weights according to the dynamic optimization status of different modalities. Extensive experiments on five public medical datasets demonstrate the effectiveness of our UMD framework, which remarkably outperforms existing approaches on three kinds of downstream tasks. Li Pan 0004, Qiushi Yang, Tan Li 0002, Zhen Chen 0013 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma GradingabstractRecently, multimodal deep learning, which integrates histopathology slides and molecular biomarkers, has achieved a promising performance in glioma grading. Despite great progress, due to the intra-modality complexity and intermodality heterogeneity, existing studies suffer from inadequate histopathology representation learning and inefficient molecular-pathology knowledge alignment. These two issues hinder existing methods to precisely interpret diagnostic molecular-pathology features, thereby limiting their grading performance. Moreover, the real-world applicability of existing multimodal approaches is significantly restricted as molecular biomarkers are not always available during clinical deployment. To address these problems, we introduce a novel Focus on Focus (FoF) framework with paired pathology-genomic training and applicable pathology-only inference, enhancing molecular-pathology representation effectively. Specifically, we propose a Focus-oriented Representation Learning (FRL) module to encourage the model to identify regions positively or negatively related to glioma grading and guide it to focus on the diagnostic areas with a consistency constraint. To effectively link the molecular biomarkers to morphological features, we propose a Multi-view Cross-modal Alignment (MCA) module that projects histopathology representations into molecular subspaces, aligning morphological features with corresponding molecular biomarker status by supervised contrastive learning. Experiments on the TCGA GBMLGG dataset demonstrate that our FoF framework significantly improves the glioma grading. Remarkably, our FoF achieves superior performance using only histopathology slides compared to existing multimodal methods. The source code is available at https://github.com/peterlipan/FoF. Li Pan 0004, Qiushi Yang, Tan Li 0002, Xiaohan Xing, Maximus C. F. Yeung, Zhen Chen 0013 |
BIBM | 3 |
| 2024 | Enhancing Clinical Information for Zero-Shot Medical Diagnosis by Prompting Large Language ModelabstractIn real clinical diagnosis workflow, unseen disease categories are commonly encountered, where most existing supervised deep learning methods are invalid to accurately recognize. Recent works utilizing large-scale image-report datasets to train vision-language models have witnessed impressive zero-shot capabilities, while they rely on high-quality diagnosis reports that are difficult to collect, especially on some rare diseases. In this work, we propose Bidirectional vision-language Clinical information Exploitation (BCE), a new paradigm towards superior generalized zero-shot learning for medical diagnosis by multi-modal information mining. To harvest sparse disease semantics in medical images, the Cross-modal Knowledge Interaction (CKI) is designed by matching the global textual information towards local visual representations, which encourages the model to capture dense correspondence from visual to textual information. Furthermore, instead of using category keywords as text prompts to yield fixed descriptions from large language models (LLM) in previous works, we propose a Modality-Guided model Tuning (MGT) to encourage the LLM to produce fine-grained clinical information conditioned on input visual information. MGT can efficiently update additional learnable parameters inserted into the LLM and dynamically adapt them to yield instance-aware clinical information. Finally, a Fine-grained text-image Alignment (FA) is present to provide reliable constraint for superior discrimination. Extensive experiments on various medical generalized zero-shot learning benchmarks demonstrate the superiority of the proposed framework. Qiushi Yang, Meilu Zhu, Yixuan Yuan |
BIBM | 1 |
| 2024 | From Static to Dynamic Diagnostics: Boosting Medical Image Analysis via Motion-Informed Generative Videos
Wuyang Li, Xinyu Liu 0001, Qiushi Yang, Yixuan Yuan |
MICCAI (3) | 3 |
| 2024 | Stealing Knowledge from Pre-trained Language Models for Federated Classifier Debiasing
Meilu Zhu, Qiushi Yang, Zhifan Gao, Jun Liu 0007, Yixuan Yuan |
MICCAI (10) | 2 |
| 2023 | MRM: Masked Relation Modeling for Medical Image Pre-Training with GeneticsabstractModern deep learning techniques on automatic multi-modal medical diagnosis rely on massive expert annotations, which is time-consuming and prohibitive. Recent masked image modeling (MIM)-based pre-training methods have witnessed impressive advances for learning meaningful representations from unlabeled data and transferring to downstream tasks. However, these methods focus on natural images and ignore the specific properties of medical data, yielding unsatisfying generalization performance on downstream medical diagnosis. In this paper, we aim to leverage genetics to boost image pre-training and present a masked relation modeling (MRM) framework. Instead of explicitly masking input data in previous MIM methods leading to loss of disease-related semantics, we design relation masking to mask out token-wise feature relation in both self- and cross-modality levels, which preserves intact semantics within the input and allows the model to learn rich disease-related information. Moreover, to enhance semantic relation modeling, we propose relation matching to align the sample-wise relation between the intact and masked features. The relation matching exploits inter-sample relation by encouraging global constraints in the feature space to render sufficient semantic relation for feature representation. Extensive experiments demonstrate that the proposed framework is simple yet powerful, achieving state-of-the-art transfer performance on various downstream diagnosis tasks. Codes are available at https://github.com/CityU-AIM-Group/MRM. Qiushi Yang, Wuyang Li, Baopu Li, Yixuan Yuan |
ICCV | 1 |
| 2023 | Combat Long-Tails in Medical Classification with Relation-Aware Consistency and Virtual Features Compensation
Li Pan 0004, Qiushi Yang, Tan Li 0002, Zhen Chen 0013 |
MICCAI (6) | 3 |
| 2023 | Hierarchical Bias Mitigation for Semi-Supervised Medical Image ClassificationabstractSemi-supervised learning (SSL) has demonstrated remarkable advances on medical image classification, by harvesting beneficial knowledge from abundant unlabeled samples. The pseudo labeling dominates current SSL approaches, however, it suffers from intrinsic biases within the process. In this paper, we retrospect the pseudo labeling and identify three hierarchical biases: perception bias, selection bias and confirmation bias, at feature extraction, pseudo label selection and momentum optimization stages, respectively. In this regard, we propose a HierArchical BIas miTigation (HABIT) framework to amend these biases, which consists of three customized modules including Mutual Reconciliation Network (MRNet), Recalibrated Feature Compensation (RFC) and Consistency-aware Momentum Heredity (CMH). Firstly, in the feature extraction, MRNet is devised to jointly utilize convolution and permutator-based paths with a mutual information transfer module to exchanges features and reconcile spatial perception bias for better representations. To address pseudo label selection bias, RFC adaptively recalibrates the strong and weak augmented distributions to be a rational discrepancy and augments features for minority categories to achieve the balanced training. Finally, in the momentum optimization stage, in order to reduce the confirmation bias, CMH models the consistency among different sample augmentations into network updating process to improve the dependability of the model. Extensive experiments on three semi-supervised medical image classification datasets demonstrate that HABIT mitigates three biases and achieves state-of-the-art performance. Our codes are available at https://github.com/CityU-AIM-Group/HABIT. Qiushi Yang, Zhen Chen 0013, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2022 | Towards Robust Adaptive Object Detection under Noisy AnnotationsabstractDomain Adaptive Object Detection (DAOD) models a joint distribution of images and labels from an annotated source domain and learns a domain-invariant transformation to estimate the target labels with the given target domain images. Existing methods assume that the source domain labels are completely clean, yet large-scale datasets often contain error-prone annotations due to instance ambiguity, which may lead to a biased source distribution and severely degrade the performance of the domain adaptive detector de facto. In this paper, we represent the first effort to formulate noisy DAOD and propose a Noise Latent Transferability Exploration (NLTE) framework to address this issue. It is featured with 1) Potential Instance Mining (PIM), which leverages eligible proposals to recapture the miss-annotated instances from the background; 2) Morphable Graph Relation Module (MGRM), which models the adaptation feasibility and transition probability of noisy samples with relation matrices; 3) Entropy-Aware Gradient Reconcilement (EAGR), which incorporates the semantic information into the discrimination process and enforces the gradients provided by noisy and clean samples to be consistent towards learning domain-invariant representations. A thorough evaluation on benchmark DAOD datasets with noisy source annotations validates the effectiveness of NLTE. In particular, NLTE improves the mAP by 8.4% under 60% corrupted annotations and even approaches the ideal upper bound of training on a clean source dataset.11Code is available at https://github.com/CityU-AIM-Group/NLTE. Xinyu Liu 0001, Wuyang Li, Qiushi Yang, Baopu Li, Yixuan Yuan |
CVPR | 3 |
| 2022 | Semi-supervised Medical Image Classification with Temporal Knowledge-Aware Regularization
Qiushi Yang, Xinyu Liu 0001, Zhen Chen 0013, Bulat Ibragimov, Yixuan Yuan |
MICCAI (8) | 1 |
| 2022 | PMTUD is not Panacea: Revisiting IP Fragmentation Attacks against TCP
Xuewei Feng, Qi Li 0002, Kun Sun 0001, Ke Xu 0002, Baojun Liu 0002, Qiushi Yang, Hai-Xin Duan, Zhiyun Qian |
NDSS | 7 |
| 2022 | D2-Net: Dual Disentanglement Network for Brain Tumor Segmentation With Missing ModalitiesabstractMulti-modal Magnetic Resonance Imaging (MRI) can provide complementary information for automatic brain tumor segmentation, which is crucial for diagnosis and prognosis. While missing modality data is common in clinical practice and it can result in the collapse of most previous methods relying on complete modality data. Current state-of-the-art approaches cope with the situations of missing modalities by fusing multi-modal images and features to learn shared representations of tumor regions, which often ignore explicitly capturing the correlations among modalities and tumor regions. Inspired by the fact that modality information plays distinct roles to segment different tumor regions, we aim to explicitly exploit the correlations among various modality-specific information and tumor-specific knowledge for segmentation. To this end, we propose a Dual Disentanglement Network (D2-Net) for brain tumor segmentation with missing modalities, which consists of amodality disentanglement stage(MD-Stage) and atumor-region disentanglement stage(TD-Stage). In the MD-Stage, a spatial-frequency joint modality contrastive learning scheme is designed to directly decouple the modality-specific information from MRI data. To decompose tumor-specific representations and extract discriminative holistic features, we propose an affinity-guided dense tumor-region knowledge distillation mechanism in the TD-Stage through aligning the features of a disentangled binary teacher network with a holistic student network. By explicitly discovering relations among modalities and tumor regions, our model can learn sufficient information for segmentation even if some modalities are missing. Extensive experiments on the public BraTS-2018 database demonstrate the superiority of our framework over state-of-the-art methods in missing modalities situations. Codes are available athttps://github.com/CityU-AIM-Group/D2Net. Qiushi Yang, Xiaoqing Guo, Zhen Chen 0013, Yat Ming Peter Woo, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Poison Over Troubled Forwarders: A Cache Poisoning Attack Targeting DNS Forwarding Devices
Chaoyi Lu, Qiushi Yang, Dongjie Zhou, Baojun Liu 0002, Keyu Man, Shuang Hao 0001, Hai-Xin Duan, Zhiyun Qian |
USENIX Security Symposium | 4 |
| 2011 | Secure Communication in Multicast Graphs
Qiushi Yang, Yvo Desmedt |
ASIACRYPT | 1 |
| 2010 | General Perfectly Secure Message Transmission Using Linear Codes
Qiushi Yang, Yvo Desmedt |
ASIACRYPT | 1 |