EDBT 2026 Demo / reviewers in the wild / expert
Xiao Tan 0005
dblp:116/7143-5
· DBLP profile ↗
9ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-3874-9557ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiGen: Active Lipid Generation via a Molecular Language ModelabstractLipid nanoparticles (LNPs) can deliver cargos to both tumor and immune cells, playing a crucial role in biomedicine.Traditional approaches rely on experimental screening and expert knowledge, which can be costly and time-consuming.Recent methods based on language models have accelerated this process using deep learning.Although these methods can retrieve molecules for fusion or rank candidates from existing libraries, they are still limited by the scope of known formulations.In this work, we propose LiGen to generate lipid molecules efficiently and actively, facilitating the discovery of high-performing LNP formulations.We first train a lipid-specific molecular language model, LiCore, to learn hidden representations of lipid molecules.We then explore the learned latent space to generate improved candidate formulations.This process is guided by a trained predictor, which evaluates delivery efficiency and provides directional signals.In reconstruction task, LiCore achieves near-perfect reconstruction performance output with a low invalid ratio on both the LNP-Virtual900k and LNP-Exp12k datasets.The predictor consistently improves ranking-oriented metrics across multiple cell lines, with our method outperforming the best baselines by an average of 4.1%, 10.8%, and 8.1% in Top-50, Top-10, and Top-5 identification accuracy, respectively.Guided by predictor, LiGen generates novel lipid candidates that achieve a 30.7% relative improvement over baseline methods in predicted delivery efficiency, with some candidates exceeding 50% improvement. Ying Zhan 0001, Xiuqi Tang, Yan Zhang 0100, Xiao Tan 0005, Dian Shen, Beilun Wang |
ACL (1) | 4 |
| 2026 | Identification of Influential Node Group in Attributed Graph through Explaining Graph Neural NetworkabstractIdentification of influential groups of nodes in attributed graphs has applications in a wide range of real-world problems, for instance, collecting important proceedings in citation networks, or identifying essential genes for diagnosing disease in Protein-Protein Interaction networks. Previous approaches for influence maximization manipulated on the graph structure, despite their proliferation, neglect the node attribute information containing additional knowledge. In this work, we introduce Global Graph UNderstanding (GGUN), a perturbation-based framework leveraging the explanatory power of Graph Neural Networks. It takes into account the entire graph structure and node attributes simultaneously and fuses knowledge through GNN layers. Following the perturbation-based explanation, GGUN fills the gap between Deep Neural Network gradient-based feature importance analysis and discrete structure in the graph, which is formulated as a combinatorial optimization problem. Moreover, GGUN obtains an efficient solution by relaxing the infeasible combinatorial optimization problem with performance guaranteed. Evaluations of synthetic and real-world datasets show that GGUN outperforms baselines on both quantitative metrics and human-intelligible analysis. Xiao Tan 0005, Tongtong Su, Yan Zhang 0100, Binghui Xu, Dian Shen, Meng Wang 0009, Beilun Wang |
WWW | 1 |
| 2025 | Information-Agnostic Model Poisoning Attacks Against Byzantine-Robust Federated Learning
Yan Zhang 0100, Yueyao Chen, Xiao Tan 0005, Dian Shen, Meng Wang 0009, Beilun Wang |
DASFAA (4) | 3 |
| 2025 | Bi-perspective Splitting Defense: Achieving Clean-Seed-Free Backdoor SecurityabstractBackdoor attacks have seriously threatened deep neural networks (DNNs) by embedding concealed vulnerabilities through data poisoning. To counteract these attacks, training benign models from poisoned data garnered considerable interest from researchers. High-performing defenses often rely on additional clean subsets/seeds, which is untenable due to increasing privacy concerns and data scarcity. In the absence of additional clean subsets/seeds, defenders resort to complex feature extraction and analysis, resulting in excessive overhead and compromised performance. To address these challenges, we identify the key lies in sufficient utilization of both the easier-to-obtain target labels and clean hard samples. In this work, we propose a Bi-perspective Splitting Defense (BSD). BSD distinguishes clean samples using both semantic and loss statistics characteristics through open set recognition-based splitting (OSS) and altruistic model-based data splitting (ALS) respectively. Through extensive experiments on benchmark datasets and against representative attacks, we empirically demonstrate that BSD surpasses existing defenses by over 20% in average Defense Effectiveness Rating (DER), achieving clean data-free backdoor security. Yangyang Shen, Xiao Tan 0005, Dian Shen, Meng Wang 0009, Beilun Wang |
ICML | 2 |
| 2025 | NoTeNet: Normalized Mutual Information-Driven Tuning-free Dynamic Dependence Network Inference Method for Multimodal DataabstractDynamic Dependence Network (DDN) inference is crucial for understanding evolving relationships in multimodal time series web data, with broad applications in fields like medical and financial network analysis. The inherent dynamic nature, temporal continuity, and heterogeneous data sources in multimodal time series data pose three fundamental challenges: computational efficiency, prediction stability and robustness, and modality quality disparity. Previous methods, generally lacking utilization of multiple modalities, either struggle with computational efficiency due to the time-intensive manual hyperparameter tuning, or compromise prediction stability and robustness by neglecting temporal coherence. To address these challenges, we propose a Normalized mutual information-driven Tuning-free Dynamic Dependence Network inference method for multimodal data, namely NoTeNet. NoTeNet provides a promising paradigm that can integrate two different data modalities to enhance prediction accuracy. It uses normalized mutual information transforms noisy auxiliary data into relationship matrices and employs a kernel function for smooth temporal estimation. Additionally, NoTeNet significantly reduces the need for manual hyperparameter adjustments, offering a tuning-free approach with theoretical guarantees. On various synthetic datasets and real-world data, NoTeNet demonstrates superior prediction accuracy and efficiency without the need for hyperparameter tuning, making it potential for a wide range of web data applications. Xiao Tan 0005, Yangyang Shen, Yan Zhang 0100, Jingwen Shao, Dian Shen, Meng Wang 0009, Beilun Wang |
WWW | 1 |
| 2024 | Factor Model-Based Large Covariance Estimation from Streaming Data Using a Knowledge-Based Sketch MatrixabstractCovariance matrix estimation is an important problem in statistics, with wide applications in finance, neuroscience, meteorology, oceanography, and other fields. However, when the data are high-dimensional and constantly generated and updated in a streaming fashion, the covariance matrix estimation faces huge challenges, including the curse of dimensionality and limited memory space. The existing methods either assume sparsity, ignoring any possible common factor among the variables, or obtain poor performance in recovering the covariance matrix directly from sketched data. To address these issues, we propose a novel method - KEEF: Knowledge-based Time and Memory Efficient Covariance Estimator in Factor Model and its extended variation. Our method leverages historical data to train a knowledge-based sketch matrix, which is used to accelerate the factor analysis of streaming data and directly estimates the covariance matrix from the sketched data. We provide theoretical guarantees, showing the advantages of our method in terms of time and space complexity, as well as accuracy. We conduct extensive experiments on synthetic and real-world data, comparing KEEF with several state-of-the-art methods, demonstrating the superior performance of our method. Xiao Tan 0005, Hao Qian 0003, Jun Zhou 0011, Peibo Duan, Dian Shen, Meng Wang 0009, Beilun Wang |
CIKM | 1 |
| 2024 | Large Covariance Estimation from Streaming Data with Knowledge-Based Sketch Matrix
Xiao Tan 0005, Meng Wang 0009, Dian Shen, Weitong Chen 0001, Beilun Wang |
DASFAA (5) | 1 |
| 2024 | FasMe: Fast and Sample-efficient Meta Estimator for Precision Matrix Learning in Small Sample SettingsabstractPrecision matrix estimation is a ubiquitous task featuring numerous applications such as rare disease diagnosis and neural connectivity exploration. However, this task becomes challenging in small sample settings, where the number of samples is significantly less than the number of dimensions, leading to unreliable estimates. Previous approaches either fail to perform well in small sample settings or suffer from inefficient estimation processes, even when incorporating meta-learning techniques.
To this end, we propose a novel approach FasMe for Fast and Sample-efficient Meta Precision Matrix Learning, which first extracts meta-knowledge through a multi-task learning diagram. Then, meta-knowledge constraints are applied using a maximum determinant matrix completion algorithm for the novel task. As a result, we reduce the sample size requirements to $O(\log p/K)$ per meta-training task and $O(\log\vert \mathcal{G}\vert)$ for the meta-testing task. Moreover, the hereby proposed model only needs $O(p \log\epsilon^{-1})$ time and $O(p)$ memory for converging to an $\epsilon$-accurate solution. On multiple synthetic and biomedical datasets, FasMe is at least ten times faster than the four baselines while promoting prediction accuracy in small sample settings. Xiao Tan 0005, Yangyang Shen, Dian Shen, Meng Wang 0009, Peibo Duan, Beilun Wang |
NeurIPS | 1 |
| 2023 | Graph Inference via the Energy-efficient Dynamic Precision Matrix Estimation with One-bit DataabstractGraph knowledge discovery from graph-structured data is a fascinating data mining topic in various domains, especially in the Internet of Things, where inferring the graph structure from such informative data can benefit many downstream tasks. Deep neural networks are typically used to perform such predictions, but they produce unreliable results without sufficient high-quality data. Therefore, researchers introduce lightweight statistical precision matrix learning to infer the graph structure in many IoT scenarios with limited communication and resolution of sensors. However, these methods still suffer from low-resolution data or the omission of hidden information in time-series data. To address the challenges, we propose a novel approach for Energy-efficient Dynamic Sparse Graph Structure Estimation with one-bit data, EDGE. Our method proposes a novel estimator to estimate the covariance matrix from one-bit data, and then utilize the covariance matrices to capture the dynamic structure. We theoretically demonstrate the effectiveness of the estimators by deriving two non-asymptotic estimation error bounds for the estimated covariance matrix and precision matrix, respectively. The theoretical results show that our method can achieve a consistent result of the precision matrix at the rate O(log p/n). On multiple synthetic and real-world datasets, the experimental results demonstrate that our proposed estimator is able to obtain a relatively high detection rate using one-bit data, which exceeds the baseline by 35%, and identify potentially perturbed nodes in real-time dynamic network inference. Xiao Tan 0005, Yangyang Shen, Meng Wang 0009, Beilun Wang |
CIKM | 1 |