EDBT 2026 Demo / reviewers in the wild / expert
Zhaorui Tan
dblp:332/0953
· DBLP profile ↗
17ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0001-5054-8275ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MedMAP: Promoting Incomplete Multi-Modal Brain Tumor Segmentation With AlignmentabstractBrain tumor segmentation is often based on multiple magnetic resonance imaging (MRI). However, in clinical practice, certain modalities of MRI may be missing, which presents a more difficult scenario. To cope with this challenge, Knowledge Distillation, Domain Adaption, and Shared Latent Space have emerged as commonly promising strategies. However, recent efforts to address the missing modality problem in brain tumor segmentation typically overlook the modality gaps and thus fail to learn important invariant feature representations across different modalities. Such drawback consequently leads to limited performance for missing modality models. To ameliorate these problems, pre-trained models are used in natural visual segmentation tasks to minimize the gaps. However, promising pre-trained models are difficult to obtain in the brain tumor segmentation task due to the lack of sufficient data. Along this line, in this paper, we propose a novel paradigm that aligns latent features of involved modalities to a well-defined distribution anchor as the substitution of the pre-trained model. As a major contribution, we prove that our novel training paradigm ensures a tight evidence lower bound, thus theoretically certifying its effectiveness. Extensive experiments on different backbones validate that the proposed paradigm can enable invariant feature representations and produce models with narrowed modality gaps. Models with our alignment paradigm show their superior performance on both BraTS2018, BraTS2020 and Brain Metastasis datasets. Zhaorui Tan, Muyin Chen, Xi Yang 0008, Haochuan Jiang, Kaizhu Huang |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | ScaleOT: Privacy-utility-scalable Offsite-tuning with Dynamic LayerReplace and Selective Rank CompressionabstractOffsite-tuning is a privacy-preserving method for tuning large language models (LLMs) by sharing a lossy compressed emulator from the LLM owners with data owners for downstream task tuning. This approach protects the privacy of both the model and data owners. However, current offsite tuning methods often suffer from adaptation degradation, high computational costs, and limited protection strength due to uniformly dropping LLM layers or relying on expensive knowledge distillation. To address these issues, we propose ScaleOT, a novel privacy-utility-scalable offsite-tuning framework that effectively balances privacy and utility. ScaleOT introduces a novel layerwise lossy compression algorithm that uses reinforcement learning to obtain the importance of each layer. It employs lightweight networks, termed harmonizers, to replace the raw LLM layers. By combining important original LLM layers and harmonizers in different ratios, ScaleOT generates emulators tailored for optimal performance with various model scales for enhanced privacy protection. Additionally, we present a rank reduction method to further compress the original LLM layers, significantly enhancing privacy with negligible impact on utility. Comprehensive experiments show that ScaleOT can achieve nearly lossless offsite tuning performance compared with full fine-tuning while obtaining better model privacy. Zhaorui Tan, Tiandi Ye, Lichun Li, Yuan Zhao 0015, Wenyan Liu 0001, Wei Wang 0002, Jianke Zhu |
AAAI | 2 |
| 2025 | Disentangling Tabular Data Towards Better One-Class Anomaly DetectionabstractTabular anomaly detection under the one-class classification setting poses a significant challenge, as it involves accurately conceptualizing "normal" derived exclusively from a single category to discern anomalies from normal data variations. Capturing the intrinsic correlation among attributes within normal samples presents one promising method for learning the concept. To do so, the most recent effort relies on a learnable mask strategy with a reconstruction task. However, this wisdom may suffer from the risk of producing uniform masks, i.e., essentially nothing is masked, leading to less effective correlation learning. To address this issue, we presume that attributes related to others in normal samples can be divided into two non-overlapping and correlated subsets, defined as CorrSets, to capture the intrinsic correlation effectively. Accordingly, we introduce an innovative method that disentangles CorrSets from normal tabular data. To our knowledge, this is a pioneering effort to apply the concept of disentanglement for one-class anomaly detection on tabular data. Extensive experiments on 20 tabular datasets show that our method substantially outperforms the state-of-the-art methods and leads to an average performance improvement of 6.1% on AUC-PR and 2.1% on AUC-ROC. Jianan Ye, Zhaorui Tan, Yijie Hu, Xi Yang 0008, Kaizhu Huang |
AAAI | 2 |
| 2025 | GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language ModelsabstractKai Yao, Zhaorui Tan, Penglei Gao, Lichun Li, Kaixin Wu, Yinggui Wang, Yuan Zhao, Yixin Ji, Jianke Zhu, Wei Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhaorui Tan, Penglei Gao, Lichun Li, Kaixin Wu, Yinggui Wang, Yuan Zhao 0015, Yixin Ji, Jianke Zhu, Wei Wang 0002 |
ACL (1) | 2 |
| 2025 | Structure-Aware Semantic Discrepancy and Consistency for 3D Medical Image Self-Supervised Learning
Tan Pan, Zhaorui Tan, Kaiyu Guo, Dongli Xu, Weidi Xu, Chen Jiang 0006, Xin Guo 0010, Yuan Qi 0001 |
ICCV | 2 |
| 2025 | Towards a Universal 3D Medical Multi-Modality Generalization via Learning Personalized Invariant Representation
Zhaorui Tan, Xi Yang 0008, Tan Pan, Chen Jiang 0006, Xin Guo 0010, Qiufeng Wang 0001, Anh Nguyen 0003, Yuan Qi 0001, Kaizhu Huang |
ICCV | 1 |
| 2025 | Minimal Semantic Sufficiency Meets Unsupervised Domain GeneralizationabstractThe generalization ability of deep learning has been extensively studied in supervised settings, yet it remains less explored in unsupervised scenarios. Recently, the Unsupervised Domain Generalization (UDG) task has been proposed to enhance the generalization of models trained with prevalent unsupervised learning techniques, such as Self-Supervised Learning (SSL). UDG confronts the challenge of distinguishing semantics from variations without category labels. Although some recent methods have employed domain labels to tackle this issue, such domain labels are often unavailable in real-world contexts. In this paper, we address these limitations by formalizing UDG as the task of learning a Minimal Sufficient Semantic Representation: a representation that (i) preserves all semantic information shared across augmented views (sufficiency), and (ii) maximally removes information irrelevant to semantics (minimality). We theoretically ground these objectives from the perspective of information theory, demonstrating that optimizing representations to achieve sufficiency and minimality directly reduces out-of-distribution risk. Practically, we implement this optimization through Minimal-Sufficient UDG (MS-UDG), a learnable model by integrating (a) an InfoNCE-based objective to achieve sufficiency; (b) two complementary components to promote minimality: a novel semantic-variation disentanglement loss and a reconstruction-based mechanism for capturing adequate variation. Empirically, MS-UDG sets a new state-of-the-art on popular unsupervised domain-generalization benchmarks, consistently outperforming existing SSL and UDG methods, without category or domain labels during representation learning. Tan Pan, Kaiyu Guo, Dongli Xu, Zhaorui Tan, Chen Jiang 0006, Deshu Chen, Xin Guo 0010, Brian C. Lovell, Limei Han, Mahsa Baktash |
NeurIPS | 4 |
| 2025 | Covariance-Based Space Regularization for Few-Shot Class Incremental Learning
Yijie Hu, Guanyu Yang 0002, Zhaorui Tan, Xiaowei Huang 0001, Kaizhu Huang, Qiufeng Wang 0001 |
WACV | 3 |
| 2025 | Heterogeneous graph contrastive learning for integration and alignment of spatial transcriptomics dataabstractSpatial transcriptomics (ST) technology enables the simultaneous capture of gene expression profile and spatial information within 2D tissue slices. However, conventional analyses that process each individual slice independently often overlook shared features across multiple slices, limiting comprehensive biological insights. To address this, we introduce GRASS, a deep graph representation learning-based framework designed for the integration and alignment of multislice ST data. GRASS consists of two core modules: GRASS_Integration, which employs a heterogeneous graph architecture integrating contrastive learning and a multi-expert collaboration strategy to fully utilize both shared and unique information, enabling multislice integration, clustering, and various downstream analyses; and GRASS_Alignment, which uses a dual-perception similarity metric to guide spot-level alignment, supporting downstream tasks such as imputation and 3D reconstruction. Experimental results on seven ST datasets from five different platforms demonstrate that GRASS consistently outperforms eight state-of-the-art methods in both integration and alignment tasks. By comprehensively addressing multi-level information integration, GRASS emerges as an ideal solution for the joint analysis of multislice ST data. Yang Gui, Zhaorui Tan, Chunzhong Li |
Briefings Bioinform. | 2 |
| 2025 | Stagger Network: Rethinking information loss in medical image segmentation with various-sized targets
Zhaorui Tan, Haochuan Jiang, Kaizhu Huang |
Neural Networks | 2 |
| 2025 | SCMix: Stochastic Compound Mixing for Open Compound Domain Adaptation in Semantic SegmentationabstractOpen compound domain adaptation (OCDA) aims to transfer knowledge from a labeled source domain to a mix of unlabeled homogeneous compound target domains while generalizing to open unseen domains. Existing OCDA methods solve the intradomain gaps by a divide-and-conquer strategy, which decomposes the problem into several individual and parallel domain adaptation (DA) tasks. In this work, starting from the general DA theory, we establish a novel generalization bound for the setting of OCDA. Built upon this, we argue that conventional OCDA approaches may substantially underestimate the inherent variance inside the compound target domains for model generalization, constraining the model's performance. We subsequently present stochastic compound mixing (SCMix), an augmentation strategy with the primary objective of mitigating the divergence between the source and mixed target distributions. Theoretical analyses are conducted to substantiate the superiority of SCMix, proving that single-target mixing is a subgroup of our method. Extensive experiments show that our method attains a lower empirical risk on OCDA semantic segmentation tasks, thus supporting our theories. In particular, combining the transformer architecture, SCMix achieves a notable performance boost compared to SoTA results. Zhaorui Tan, Zixian Su, Xi Yang 0008, Jie Sun 0024, Kaizhu Huang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Semantic-Aware Data Augmentation for Text-to-Image SynthesisabstractData augmentation has been recently leveraged as an effective regularizer in various vision-language deep neural networks. However, in text-to-image synthesis (T2Isyn), current augmentation wisdom still suffers from the semantic mismatch between augmented paired data. Even worse, semantic collapse may occur when generated images are less semantically constrained. In this paper, we develop a novel Semantic-aware Data Augmentation (SADA) framework dedicated to T2Isyn. In particular, we propose to augment texts in the semantic space via an Implicit Textual Semantic Preserving Augmentation, in conjunction with a specifically designed Image Semantic Regularization Loss as Generated Image Semantic Conservation, to cope well with semantic mismatch and collapse. As one major contribution, we theoretically show that Implicit Textual Semantic Preserving Augmentation can certify better text-image consistency while Image Semantic Regularization Loss regularizing the semantics of generated images would avoid semantic collapse and enhance image quality. Extensive experiments validate that SADA enhances text-image consistency and improves image quality significantly in T2Isyn models across various backbones. Especially, incorporating SADA during the tuning process of Stable Diffusion models also yields performance improvements. Zhaorui Tan, Xi Yang 0008, Kaizhu Huang |
AAAI | 1 |
| 2024 | Mind the Gap: Promoting Missing Modality Brain Tumor Segmentation with AlignmentabstractBrain tumor segmentation is often based on multiple magnetic resonance imaging (MRI). However, in clinical practice, certain modalities of MRI may be missing, which presents an even more difficult scenario. To cope with this challenge, knowledge distillation has emerged as one promising strategy. However, recent efforts typically overlook the modality gaps and thus fail to learn invariant feature representations across different modalities. Such drawback consequently leads to limited performance for both teachers and students. To ameliorate these problems, in this paper, we propose a novel paradigm that aligns latent features of involved modalities to a well-defined distribution anchor. As a major contribution, we prove that our novel training paradigm ensures a tight evidence lower bound, thus theoretically certifying its effectiveness. Extensive experiments on different backbones validate that the proposed paradigm can enable invariant feature representations and produce a teacher with narrowed modality gaps. This further offers superior guidance for missing modality students, achieving an average improvement of 1.75 on dice score. Zhaorui Tan, Haochuan Jiang, Xi Yang 0008, Kaizhu Huang |
BIBM | 2 |
| 2024 | Rethinking Multi-Domain Generalization with A General Learning ObjectiveabstractMulti-domain generalization$(mDG)$is universally aimed to minimize the discrepancy between training and testing distributions to enhance marginal-to-label distribution mapping. However, existing$mDG$literature lacks a general learning objective paradigm and often imposes constraints on static target marginal distributions. In this paper, we propose to leverage a Y-mapping to relax the constraint. We rethink the learning objective for$mDG$and design a new general learning objective to interpret and analyze most existing$mDG$wisdom. This general objective is bifurcated into two synergistic amis: learning domain-independent conditional features and maximizing a posterior. Explorations also extend to two effective regularization terms that incorporate prior information and suppress invalid causality, alleviating the issues that come with relaxed constraints. We theoretically contribute an upper bound for the domain alignment of domain-independent conditional features, disclosing that many previous$mDG$endeavors actually optimize partially the objective and thus lead to limited performance. As such, our study distills a general learning objective into four practical components, providing a general, robust, and flexible mechanism to handle complex domain shifts. Extensive empirical results indicate that the proposed objective with Y -mapping leads to substantially better$mDG$performance in various downstream tasks, including regression, segmentation, and classification. Code is available at htttps://github.com/zhaorui-t.an/GMDG/tree/main. Zhaorui Tan, Xi Yang 0008, Kaizhu Huang |
CVPR | 1 |
| 2024 | Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual ClassificationabstractVision models excel in image classification but struggle to generalize to unseen data, such as classifying images from unseen domains or discovering novel categories. In this paper, we explore the relationship between logical reasoning and deep learning generalization in visual classification. A logical regularization termed L-Reg is derived which bridges a logical analysis framework to image classification. Our work reveals that L-Reg reduces the complexity of the model in terms of the feature distribution and classifier weights. Specifically, we unveil the interpretability brought by L-Reg, as it enables the model to extract the salient features, such as faces to persons, for classification. Theoretical analysis and experiments demonstrate that L-Reg enhances generalization across various scenarios, including multi-domain generalization and generalized category discovery. In complex real-world scenarios where images span unknown classes and unseen domains, L-Reg consistently improves generalization, highlighting its practical efficacy. Zhaorui Tan, Xi Yang 0008, Qiufeng Wang 0001, Anh Nguyen 0003, Kaizhu Huang |
NeurIPS | 1 |
| 2023 | PAG: Protecting Artworks from Personalizing Image Generative Models
Zhaorui Tan, Siyuan Wang 0017, Xi Yang 0008, Kaizhu Huang |
ICONIP (4) | 1 |
| 2023 | Semantic Similarity Distance: Towards better text-image consistency metric in text-to-image generation
Zhaorui Tan, Xi Yang 0008, Zihan Ye, Qiufeng Wang 0001, Yuyao Yan, Anh Nguyen 0003, Kaizhu Huang |
Pattern Recognit. | 1 |