EDBT 2026 Demo / reviewers in the wild / expert
Tiehang Duan
dblp:184/7734
· DBLP profile ↗
20ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0003-4323-642XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Release the Potential of Memory Buffer in Continual Learning: A Dynamic System PerspectiveabstractContinual learning (CL) focuses on learning non-stationary data distribution without forgetting previous knowledge. The most widely used memory-replay approaches are often prone to memory overfitting due to the limited memory diversity and hardness. Existing work mitigating memory overfitting either lacks data diversity or hardness or is hard to train. To address the above limitations and release the memory buffer potential, we view the memory buffer transformation from a new dynamic system perspective and propose a continuous and reversible memory transformation method. We introduce an adversarial optimization objective that jointly learns the CL model and memory transformer. Specifically, we present a deterministic continuous memory transformer (DCMT) to generate diverse memory data. Furthermore, we inject uncertainty into the transformation function and develop a stochastic continuous memory transformer (SCMT), which substantially enhances the diversity of the transformed memory buffer. The presented neural transformation approaches have significant advantages over existing ones: (1) they significantly increase the memory buffer diversity and hardness to overfit; (2) they are memory efficient without needing to make a replica of the memory data. Extensive experiments show a significant improvement with our approach compared to strong baselines. Zhenyi Wang 0001, Li Shen 0008, Tiehang Duan, Yanjun Zhu, Tongliang Liu, Mingchen Gao, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Leveraging Vulnerabilities in Temporal Graph Neural Networks via Strategic High-Impact AssaultsabstractTemporal Graph Neural Networks (TGNNs) have become indispensable for analyzing dynamic graphs in critical applications such as social networks, communication systems, and financial networks. However, the robustness of TGNNs against adversarial attacks, particularly sophisticated attacks that exploit the temporal dimension, remains a significant challenge. Existing attack methods for Spatio-Temporal Dynamic Graphs (STDGs) often rely on simplistic, easily detectable perturbations (e.g., random edge additions/deletions) and fail to strategically target the most influential nodes and edges for maximum impact. We introduce the High Impact Attack (HIA), a novel restricted black-box attack framework specifically designed to overcome these limitations and expose critical vulnerabilities in TGNNs. HIA leverages a data-driven surrogate model to identify structurally important nodes (central to network connectivity) and dynamically important nodes (critical for the graph's temporal evolution). It then employs a hybrid perturbation strategy, combining strategic edge injection (to create misleading connections) and targeted edge deletion (to disrupt essential pathways), maximizing TGNN performance degradation. Importantly, HIA minimizes the number of perturbations to enhance stealth, making it more challenging to detect. Comprehensive experiments on five real-world datasets and four representative TGNN architectures (TGN, JODIE, DySAT, and TGAT) demonstrate that HIA significantly reduces TGNN accuracy on the link prediction task, achieving up to a 35.55% decrease in Mean Reciprocal Rank (MRR) - a substantial improvement over state-of-the-art baselines. These results highlight fundamental vulnerabilities in current STDG models and underscore the urgent need for robust defenses that account for both structural and temporal dynamics. Code and Data are available at https://github.com/ryandhjeon/hia. Donghyun Jeon, Lijing Zhu, Haifang Li 0003, Pengze Li, Jingna Feng, Tiehang Duan, Houbing Song, Cui Tao, Shuteng Niu |
CIKM | 6 |
| 2025 | ETT-CKGE: Efficient Task-Driven Tokens for Continual Knowledge Graph Embedding
Lijing Zhu, Qizhen Lan, Qing Tian 0003, Xi Xiao 0003, Tiehang Duan, Cui Tao, Shuteng Niu |
ECML/PKDD (6) | 9 |
| 2024 | Training A Secure Model Against Data-Free Model Extraction
Zhenyi Wang 0001, Li Shen 0008, Tiehang Duan, Siyu Luan, Tongliang Liu, Mingchen Gao |
ECCV (79) | 4 |
| 2024 | RefAI: a GPT-powered retrieval-augmented generative tool for biomedical literature recommendation and summarizationabstractOBJECTIVES: Precise literature recommendation and summarization are crucial for biomedical professionals. While the latest iteration of generative pretrained transformer (GPT) incorporates 2 distinct modes-real-time search and pretrained model utilization-it encounters challenges in dealing with these tasks. Specifically, the real-time search can pinpoint some relevant articles but occasionally provides fabricated papers, whereas the pretrained model excels in generating well-structured summaries but struggles to cite specific sources. In response, this study introduces RefAI, an innovative retrieval-augmented generative tool designed to synergize the strengths of large language models (LLMs) while overcoming their limitations. MATERIALS AND METHODS: RefAI utilized PubMed for systematic literature retrieval, employed a novel multivariable algorithm for article recommendation, and leveraged GPT-4 turbo for summarization. Ten queries under 2 prevalent topics ("cancer immunotherapy and target therapy" and "LLMs in medicine") were chosen as use cases and 3 established counterparts (ChatGPT-4, ScholarAI, and Gemini) as our baselines. The evaluation was conducted by 10 domain experts through standard statistical analyses for performance comparison. RESULTS: The overall performance of RefAI surpassed that of the baselines across 5 evaluated dimensions-relevance and quality for literature recommendation, accuracy, comprehensiveness, and reference integration for summarization, with the majority exhibiting statistically significant improvements (P-values <.05). DISCUSSION: RefAI demonstrated substantial improvements in literature recommendation and summarization over existing tools, addressing issues like fabricated papers, metadata inaccuracies, restricted recommendations, and poor reference integration. CONCLUSION: By augmenting LLM with external resources and a novel ranking algorithm, RefAI is uniquely capable of recommending high-quality literature and generating well-structured summaries, holding the potential to meet the critical needs of biomedical professionals in navigating and synthesizing vast amounts of scientific literature. Jeff Zhao, Manqi Li, Yifang Dang, Evan Yu, Jianfu Li, Zenan Sun, Usama Hussein, Jianguo Wen, Ahmed M. Abdelhameed, Junhua Mai, Shenduo Li, Yue Yu 0012, Xinyue Hu 0002, Daowei Yang, Jingna Feng, Zehan Li, Jianping He 0002, Tiehang Duan, Yanyan Lou, Fang Li 0011, Cui Tao |
J. Am. Medical Informatics Assoc. | 20 |
| 2024 | Online continual decoding of streaming EEG signal with a balanced and informative memory buffer
Tiehang Duan, Zhenyi Wang 0001, Fang Li 0011, Gianfranco Doretto, Donald A. Adjeroh, Yiyi Yin, Cui Tao |
Neural Networks | 1 |
| 2023 | Replay with Stochastic Neural Transformation for Online Continual EEG ClassificationabstractBrain computer interface (BCI) systems used for clinical assistance purposes such as wheelchair control require decoding of streaming brain signals i.e. electroencephalography (EEG) signals over a long period of time with subject shift in the middle. Numerous challenges arise during this online continual brain signal decoding process: 1) the EEG decoder needs to deal with streaming EEG signals from sequentially arriving subjects, with no data available beforehand for large-scale pretraining; 2) the EEG decoder should avoid catastrophic forgetting on previous subjects after learning on a new subject; 3) the EEG decoder should perform well on noisy signals with high variance across subjects. We proposed a principled replay-based approach for this general decoding scenario, forming a bi-level optimization framework with stochastic neural transformation for dynamic memory evolution, making them representative in feature space and encouraging the model to generalize well. The evolved signal segments are stored and replayed during later decoding stages to achieve optimal model performance on all previous subjects. The stochastic neural transformation performed in inner sup of bi-level optimization significantly enhances the diversity of stored signal segments and improves model robustness during online continual decoding. We perform detailed theoretical analysis on model’s generalization ability in addition to the empirical evaluations. We construct multiple new benchmarks to mimic real-world online sequential EEG decoding scenarios with underlying subject shifts. The extensive evaluation of the proposed approach shows it outperforms related strong baselines by a large margin. Tiehang Duan, Zhenyi Wang 0001, Gianfranco Doretto, Fang Li 0011, Cui Tao, Donald A. Adjeroh |
BIBM | 1 |
| 2023 | MetaMix: Towards Corruption-Robust Continual Learning with Temporally Self-Adaptive Data TransformationabstractContinual Learning (CL) has achieved rapid progress in recent years. However, it is still largely unknown how to determine whether a CL model is trustworthy and how to foster its trustworthiness. This work focuses on evaluating and improving the robustness to corruptions of existing CL models. Our empirical evaluation results show that existing state-of-the-art (SOTA) CL models are particularly vulnerable to various data corruptions during testing. To make them trustworthy and robust to corruptions deployed in safety-critical scenarios, we propose a meta-learning framework of self-adaptive data augmentation to tackle the corruption robustness in CL. The proposed framework, MetaMix, learns to augment and mix data, automatically transforming the new task data or memory data. It directly optimizes the generalization performance against data corruptions during training. To evaluate the corruption robustness of our proposed approach, we construct several CL corruption datasets with different levels of severity. We perform comprehensive experiments on both task- and class-continual learning. Extensive experiments demonstrate the effectiveness of our proposed method compared to SOTA baselines. Zhenyi Wang 0001, Li Shen 0008, Donglin Zhan, Qiuling Suo, Yanjun Zhu, Tiehang Duan, Mingchen Gao |
CVPR | 6 |
| 2023 | Distributionally Robust Cross Subject EEG DecodingabstractRecently, deep learning has shown to be effective for Electroencephalography (EEG) decoding tasks. Yet, its performance can be negatively influenced by two key factors: 1) the high variance and different types of corruption that are inherent in the signal, 2) the EEG datasets are usually relatively small given the acquisition cost, annotation cost and amount of effort needed. Data augmentation approaches for alleviation of this problem have been empirically studied, with augmentation operations on spatial domain, time domain or frequency domain handcrafted based on expertise of domain knowledge. In this work, we propose a principled approach to perform dynamic evolution on the data for improvement of decoding robustness. The approach is based on distributionally robust optimization and achieves robustness by optimizing on a family of evolved data distributions instead of the single training data distribution. We derived a general data evolution framework based on Wasserstein gradient flow (WGF) and provides two different forms of evolution within the framework. Intuitively, the evolution process helps the EEG decoder to learn more robust and diverse features. It is worth mentioning that the proposed approach can be readily integrated with other data augmentation approaches for further improvements. We performed extensive experiments on the proposed approach and tested its performance on different types of corrupted EEG signals. The model significantly outperforms competitive baselines on challenging decoding scenarios. Tiehang Duan, Zhenyi Wang 0001, Gianfranco Doretto, Fang Li 0011, Cui Tao, Donald A. Adjeroh |
ECAI | 1 |
| 2023 | Defending against Data-Free Model Extraction by Distributionally Robust Defensive TrainingabstractData-Free Model Extraction (DFME) aims to clone a black-box model without knowing its original training data distribution, making it much easier for attackers to steal commercial models. Defense against DFME faces several challenges: (i) effectiveness; (ii) efficiency; (iii) no prior on the attacker's query data distribution and strategy. However, existing defense methods: (1) are highly computation and memory inefficient; or (2) need strong assumptions about attack data distribution; or (3) can only delay the attack or prove a model theft after the model stealing has happened. In this work, we propose a Memory and Computation efficient defense approach, named MeCo, to prevent DFME from happening while maintaining the model utility simultaneously by distributionally robust defensive training on the target victim model. Specifically, we randomize the input so that it: (1) causes a mismatch of the knowledge distillation loss for attackers; (2) disturbs the zeroth-order gradient estimation; (3) changes the label prediction for the attack query data. Therefore, the attacker can only extract misleading information from the black-box model. Extensive experiments on defending against both decision-based and score-based DFME demonstrate that MeCo can significantly reduce the effectiveness of existing DFME methods and substantially improve running efficiency. Zhenyi Wang 0001, Li Shen 0008, Tongliang Liu, Tiehang Duan, Yanjun Zhu, Donglin Zhan, David S. Doermann, Mingchen Gao |
NeurIPS | 4 |
| 2023 | UNCER: A framework for uncertainty estimation and reduction in neural decoding of EEG signals
Tiehang Duan, Zhenyi Wang 0001, Sheng Liu 0001, Yiyi Yin, Sargur N. Srihari |
Neurocomputing | 1 |
| 2023 | Distributionally Robust Memory Evolution With Generalized Divergence for Continual LearningabstractContinual learning (CL) aims to learn a non-stationary data distribution and not forget previous knowledge. The effectiveness of existing approaches that rely on memory replay can decrease over time as the model tends to overfit the stored examples. As a result, the model's ability to generalize well is significantly constrained. Additionally, these methods often overlook the inherent uncertainty in the memory data distribution, which differs significantly from the distribution of all previous data examples. To overcome these issues, we propose a principled memory evolution framework that dynamically adjusts the memory data distribution. This evolution is achieved by employing distributionally robust optimization (DRO) to make the memory buffer increasingly difficult to memorize. We consider two types of constraints in DRO: f-divergence and Wasserstein ball constraints. For f-divergence constraint, we derive a family of methods to evolve the memory buffer data in the continuous probability measure space with Wasserstein gradient flow (WGF). For Wasserstein ball constraint, we directly solve it in the euclidean space. Extensive experiments on existing benchmarks demonstrate the effectiveness of the proposed methods for alleviating forgetting. As a by-product of the proposed framework, our method is more robust to adversarial examples than compared CL methods. Zhenyi Wang 0001, Li Shen 0008, Tiehang Duan, Qiuling Suo, Le Fang 0002, Wei Liu 0005, Mingchen Gao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Learning to Learn and Remember Super Long Multi-Domain Task SequenceabstractCatastrophic forgetting (CF) frequently occurs when learning with non-stationary data distribution. The CF issue remains nearly unexplored and is more challenging when meta-learning on a sequence of domains (datasets), called sequential domain meta-learning (SDML). In this work, we propose a simple yet effective learning to learn approach, i.e., meta optimizer, to mitigate the CF problem in SDML. We first apply the proposed meta optimizer to the simplified setting of SDML, domain-aware meta-learning, where the domain labels and boundaries are known during the learning process. We propose dynamically freezing the network and incorporating it with the proposed meta optimizer by considering the domain nature during meta training. In addition, we extend the meta optimizer to the more general setting of SDML, domain-agnostic meta-learning, where domain labels and boundaries are unknown during the learning process. We propose a domain shift detection technique to capture latent domain change and equip the meta optimizer with it to work in this setting. The proposed meta optimizer is versatile and can be easily integrated with several existing meta-learning algorithms. Finally, we construct a challenging and large-scale benchmark consisting of 10 heterogeneous domains with a super long task sequence consisting of 100K tasks. We perform extensive experiments on the proposed benchmark for both settings and demonstrate the effectiveness of our proposed method, outperforming current strong baselines by a large margin. Zhenyi Wang 0001, Li Shen 0008, Tiehang Duan, Donglin Zhan, Le Fang 0002, Mingchen Gao |
CVPR | 3 |
| 2022 | Meta-Learning with Less Forgetting on Large-Scale Non-Stationary Task Distributions
Zhenyi Wang 0001, Li Shen 0008, Le Fang 0002, Qiuling Suo, Donglin Zhan, Tiehang Duan, Mingchen Gao |
ECCV (20) | 6 |
| 2022 | Improving Task-free Continual Learning by Distributionally Robust Memory EvolutionabstractTask-free continual learning (CL) aims to learn a non-stationary data stream without explicit task definitions and not forget previous knowledge. The widely adopted memory replay approach could gradually become less effective for long data streams, as the model may memorize the stored examples and overfit the memory buffer. Second, existing methods overlook the high uncertainty in the memory data distribution since there is a big gap between the memory data distribution and the distribution of all the previous data examples. To address these problems, for the first time, we propose a principled memory evolution framework to dynamically evolve the memory data distribution by making the memory buffer gradually harder to be memorized with distributionally robust optimization (DRO). We then derive a family of methods to evolve the memory buffer data in the continuous probability measure space with Wasserstein gradient flow (WGF). The proposed DRO is w.r.t the worst-case evolved memory data distribution, thus guarantees the model performance and learns significantly more robust features than existing memory-replay-based methods. Extensive experiments on existing benchmarks demonstrate the effectiveness of the proposed methods for alleviating forgetting. As a by-product of the proposed framework, our method is more robust to adversarial examples than existing task-free CL methods. Zhenyi Wang 0001, Li Shen 0008, Le Fang 0002, Qiuling Suo, Tiehang Duan, Mingchen Gao |
ICML | 5 |
| 2021 | Meta Learning on a Sequence of Imbalanced Domains with Difficulty AwarenessabstractRecognizing new objects by learning from a few labeled examples in an evolving environment is crucial to obtain excellent generalization ability for real-world machine learning systems. A typical setting across current meta learning algorithms assumes a stationary task distribution during meta training. In this paper, we explore a more practical and challenging setting where task distribution changes over time with domain shift. Particularly, we consider realistic scenarios where task distribution is highly imbalanced with domain labels unavailable in nature. We propose a kernel-based method for domain change detection and a difficulty-aware memory management mechanism that jointly considers the imbalanced domain size and domain importance to learn across domains continuously. Furthermore, we introduce an efficient adaptive task sampling method during meta training, which significantly reduces task gradient variance with theoretical guarantees. Finally, we propose a challenging benchmark with imbalanced domain sequences and varied domain difficulty. We have performed extensive evaluations on the proposed benchmark, demonstrating the effectiveness of our method. Zhenyi Wang 0001, Tiehang Duan, Le Fang 0002, Qiuling Suo, Mingchen Gao |
ICCV | 2 |
| 2020 | Attention based Writer Independent VerificationabstractThe task of writer verification is to provide a likelihood score for whether the queried and known handwritten image samples belong to the same writer or not. Such a task calls for the neural network to make it's outcome interpretable, i.e. provide a view into the network's decision making process. We implement and integrate cross-attention and soft-attention mechanisms to capture the highly correlated and salient points in feature space of 2D inputs. The attention maps serve as an explanation premise for the network's output likelihood score. The attention mechanism also allows the network to focus more on relevant areas of the input, thus improving the classification performance. Our proposed approach achieves a precision of 86% for detecting intra-writer cases in CEDAR cursive “AND” dataset. Furthermore, we generate meaningful explanations for the provided decision by extracting attention maps from multiple levels of the network. Mohammad Abuzar Shaikh, Tiehang Duan, Mihir Chauhan, Sargur N. Srihari |
ICFHR | 2 |
| 2019 | Sequential Embedding Induced Text Clustering, a Non-parametric Bayesian Approach
Tiehang Duan, Qi Lou, Sargur N. Srihari, Xiaohui Xie |
PAKDD (3) | 1 |
| 2019 | Parallel clustering of single cell transcriptomic data with split-merge sampling on Dirichlet process mixturesabstractMOTIVATION: With the development of droplet based systems, massive single cell transcriptome data has become available, which enables analysis of cellular and molecular processes at single cell resolution and is instrumental to understanding many biological processes. While state-of-the-art clustering methods have been applied to the data, they face challenges in the following aspects: (i) the clustering quality still needs to be improved; (ii) most models need prior knowledge on number of clusters, which is not always available; (iii) there is a demand for faster computational speed. RESULTS: We propose to tackle these challenges with Parallelized Split Merge Sampling on Dirichlet Process Mixture Model (the Para-DPMM model). Unlike classic DPMM methods that perform sampling on each single data point, the split merge mechanism samples on the cluster level, which significantly improves convergence and optimality of the result. The model is highly parallelized and can utilize the computing power of high performance computing (HPC) clusters, enabling massive inference on huge datasets. Experiment results show the model outperforms current widely used models in both clustering quality and computational speed. AVAILABILITY AND IMPLEMENTATION: Source code is publicly available on https://github.com/tiehangd/Para_DPMM/tree/master/Para_DPMM_package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tiehang Duan, José P. Pinto, Xiaohui Xie |
Bioinform. | 1 |
| 2016 | Pseudo Boosted Deep Belief Network
Tiehang Duan, Sargur N. Srihari |
ICANN (2) | 1 |