EDBT 2026 Demo / reviewers in the wild / expert
Fan Li 0016
dblp:73/237-16
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2026
0009-0004-2951-2226ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TCGU: Data-Centric Graph Unlearning Based on Transferable CondensationabstractWith growing demands for data privacy and model robustness, graph unlearning (GU), which erases the influence of specific data on trained GNN models, has gained significant attention. However, existing exact unlearning methods suffer from either low efficiency or poor model performance. While more utility-preserving and efficient, current approximate methods require access to the forget set during unlearning, which makes them inapplicable in immediate deletion scenarios, thereby undermining privacy. Additionally, these approximate methods, which attempt to directly perturb model parameters, still raise significant concerns regarding unlearning power in empirical studies. To fill the gap, we propose Transferable Condensation Graph Unlearning (TCGU), a data-centric solution to graph unlearning. Specifically, we first develop a two-level alignment strategy to pre-condense the original graph into a compact yet utility-preserving dataset for subsequent unlearning tasks. Upon receiving an unlearning request, we fine-tune the pre-condensed data with a low-rank plugin, to directly align its distribution with the remaining graph, thus efficiently revoking the information of deleted data without accessing them. A novel similarity distribution matching approach and a discrimination regularizer are proposed to effectively transfer condensed data and preserve its utility in GNN training, respectively. Finally, we retrain the GNN on the transferred condensed data. Extensive experiments on 7 benchmark datasets demonstrate that TCGU can achieve superior performance in terms of model utility, unlearning efficiency, and unlearning efficacy compared to existing GU methods. To the best of our knowledge, this is the first study to explore graph unlearning with immediate data removal using a data-centric approximate method. Fan Li 0016, Xiaoyang Wang 0002, Dawei Cheng, Wenjie Zhang 0001, Chen Chen 0017, Ying Zhang 0001, Xuemin Lin 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2026 | Fairness-Aware Hypergraph Self-Supervised Learning With Sampling-Efficient SignalsabstractSelf-supervised learning (SSL) provides a promising paradigm for hypergraph representation learning without reliance on costly labels. However, existing hypergraph SSL methods predominantly employ contrastive learning with instance level discrimination, encountering two significant challenges: (1) Unreliable negative sampling, where arbitrarily selected negative samples introduce bias by misclassifying similar and dissimilar pairs; and (2) High computational cost, as effective training requires a large number of negative samples. To address these limitations, we propose SE-HSSL, a hypergraph SSL frame work with three sampling-efficient self-supervised objectives. Specifically, two sampling-free objectives based on canonical correlation analysis serve as node- and group-level signals, while a hierarchical membership-level contrastive objective exploits the cascading overlap structure of hypergraphs. Overall, these designs mitigate negative sampling bias and enhance sampling efficiency, leading to better downstream performance and faster training. Beyond these challenges, deep hypergraph models are prone to biased predictions against groups defined by sensitive attributes (e.g., gender and race). To address fairness concerns, we first theoretically show that imbalanced contributions of demographic groups during hypergraph message passing amplify sensitive biases in the training data. Motivated by this, we propose FairHSSL, a fairness-aware variant of SE-HSSL equipped with a two-level debiasing augmentation strategy. To generate a fair hypergraph view, the augmentation integrates two complementary components: feature-level debiasing and structure-level perturbation. Specifically, orthogonal projection is applied at the feature level to decouple node features from sensitive attributes, while rebalance-based perturbation is introduced at the structure level to equalize the contributions of different sensitive groups during message passing. Finally, by aligning the fair augmented and original biased views in SSL, we mitigate the influence of sensitive information. Extensive experiments on 10 real-world hypergraphs demonstrate the superior effectiveness and efficiency of SE-HSSL. Moreover, FairHSSL consistently outperforms state of-the-art (SOTA) baselines in terms of utility-fairness trade-off across all datasets. Fan Li 0016, Xiaoyang Wang 0002, Dawei Cheng, Ying Zhang 0001, Wenjie Zhang 0001, Xuemin Lin 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | Masked Graph Distance Network for Accurate Subgraph Similarity ComputationabstractSubgraph similarity search aims to identify target graphs in the database that approximately contain the query graph, which is a fundamental problem in graph analysis. As a key measure for subgraph similarity computation, Subgraph Edit Distance (SED) has garnered significant research attention. Unfortunately, the exact computation of SED is an NP-hard problem. In recent years, some studies have attempted to leverage Graph Neural Networks (GNNs) to learn SED. However, existing GNN-based methods suffer from two significant limitations: (1) They rely on node-centric message passing, which cannot fully capture the impact of graph topology changes caused by graph edit operations. (2) They struggle to handle the asymmetry of SED, making it challenging to balance the scale differences between the input graphs and their distances in the representation space. To address these issues, this paper proposes a novel Masked Graph Distance Network (MGDN) for accurate SED approximation. First, MGDN utilizes a unified graph encoder to perform message passing based on the original graph structure and its dual hypergraph, effectively capturing the impact of node- and edge-specific edits. Then, we introduce an adaptive graph masking module that flexibly assigns masking scores to nodes and edges in the target graph to address asymmetry. Using multi-head masking, we re-encode the input graphs to focus on substructures relevant to SED computation. Finally, a multi-view predictor is employed at the graph level to approximate the SED, enhancing estimation accuracy by integrating information from multiple perspectives. Extensive experiments on nine benchmark datasets demonstrate that MGDN significantly outperforms state-of-the-art methods. Xijuan Liu, Fan Li 0016, Xiaoyang Wang 0002, Ying Zhang 0001 |
CIKM | 3 |
| 2025 | Efficient Dynamic Attributed Graph GenerationabstractData generation is a fundamental research problem in data management due to its diverse use cases, ranging from testing database engines to data-specific applications. However, real-world entities often involve complex interactions that cannot be effectively modeled by traditional tabular data. Therefore, graph data generation has attracted increasing attention recently. Although various graph generators have been proposed in the literature, there are three limitations: i) They cannot capture the co-evolution pattern of graph structure and node attributes. ii) Few of them consider edge direction, leading to substantial information loss. iii) Current state-of-the-art dynamic graph generators are based on the temporal random walk, making the simulation process time-consuming. To fill the research gap, we introduce VRDAG, a novel variational recurrent framework for efficient dynamic attributed graph generation. Specifically, we design a bidirectional message-passing mechanism to encode both directed structural knowledge and attribute information of a snapshot. Then, the temporal dependency in the graph sequence is captured by a recurrence state updater, generating embeddings that can preserve the evolution pattern of early graphs. Based on the hidden node embeddings, a conditional variational Bayesian method is developed to sample latent random variables at the neighboring timestep for new snapshot generation. The proposed generation paradigm avoids the time-consuming path sampling and merging process in existing random walk-based methods, significantly reducing the synthesis time. Finally, comprehensive experiments on real-world datasets are conducted to demonstrate the effectiveness and efficiency of the proposed model. Fan Li 0016, Xiaoyang Wang 0002, Dawei Cheng, Ying Zhang 0001, Xuemin Lin 0001 |
ICDE | 1 |
| 2025 | PCAN: A Pandemic-Compatible Attentive Neural Network for Retail Sales ForecastingabstractThe outbreak of pandemic has a huge impact on production and consumption in the business world, especially for the retail sector. As a crucial component of decision-support technology in the retail industry, sales forecasting is significant for production planning and optimizing the supply of essential goods during the pandemic. However, due to the irregular fluctuation pattern caused by uncertainty and the complex temporal correlation between multiple covariates and sales, there is still no effective approach for sales forecasting in this extreme event. To fill this gap, we propose a Pandemic-Compatible Attentive Network (PCAN) for retail sales forecasting. Specifically, to capture the irregular fluctuation patterns from the sales series, we design a fluctuation attention mechanism based on association discrepancy in the time series. Then, a parallel attention module is developed to learn the complex relationship between target sales and various dynamic influence factors in a decoupled manner. Finally, we introduce a novel rectification decoding strategy to indicate fluctuation points in prediction. By evaluating PCAN on four real-world retail food datasets from the SF Express international supply chain system, the results show that our method achieves superior performance over the existing state-of-the-art baselines. The model has been deployed in the supply chain system as a fundamental component to serve a world-leading food retailer. Fan Li 0016, Guoxuan Wang, Huiyu Chu, Dawei Cheng, Xiaoyang Wang 0002 |
IJCAI | 1 |
| 2024 | Hypergraph Self-supervised Learning with Sampling-efficient Signals
Fan Li 0016, Xiaoyang Wang 0002, Dawei Cheng, Wenjie Zhang 0001, Ying Zhang 0001, Xuemin Lin 0001 |
IJCAI | 1 |
| 2024 | AdaRisk: Risk-Adaptive Deep Reinforcement Learning for Vulnerable Nodes DetectionabstractVulnerable node detection in uncertain graphs is a typical graph mining problem that seeks to identify nodes at a high risk of breakdown under the joint effect from both the self and contagion risk probability. This is an NP-hard problem that is crucial for risk management in many real-world applications such as networked loans and smart grids. Monte Carlo (MC) simulation and its optimized algorithms are commonly used to approximate the breakdown probability, but these methods require a large number of samples to ensure accuracy, which is computationally expensive for large-scale networks. Although recent studies employ Graph Neural Networks (GNNs) to model the contagion process and accelerate the inference, many of these methods suffer from the over-smoothing problem, leading to suboptimal performance under the long-distance risk contagion process. To this end, we propose a novel risk-adaptive deep reinforcement learning-based framework (AdaRisk) for vulnerable nodes detection in uncertain graphs. In particular, we design the Markov Decision Process (MDP) of the vulnerability estimation process in which our agent would approach the risk adaptively based on contagion probability accumulated in prior iterations. To encode state embeddings that incorporate multi-hop contagion information, the agent utilizes a long-distance adaptable policy network to process the input graph and output actions as the vulnerable probability of nodes. We conducted extensive experiments on four benchmark networks and three real-world financial networks to evaluate our proposed framework's performance. Our results demonstrate that AdaRisk outperforms state-of-the-art baselines in terms of detection performance, and also offers significant running time reductions compared to MC simulation. Fan Li 0016, Dawei Cheng, Xiaoyang Wang 0002 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Fighting against Organized Fraudsters Using Risk Diffusion-based Parallel Graph Neural NetworkabstractMedical insurance plays a vital role in modern society, yet organized healthcare fraud causes billions of dollars in annual losses, severely harming the sustainability of the social welfare system. Existing works mostly focus on detecting individual fraud entities or claims, ignoring hidden conspiracy patterns. Hence, they face severe challenges in tackling organized fraud. In this paper, we proposed RDPGL, a novel Risk Diffusion-based Parallel Graph Learning approach, to fighting against medical insurance criminal gangs. In particular, we first leverage a heterogeneous graph attention network to encode the local context from the beneficiary-provider graph. Then, we devise a community-aware risk diffusion model to infer the global context of organized fraud behaviors with the claim-claim relation graph. The local and global representations are parallel concatenated together and trained simultaneously in an end-to-end manner. Our approach is extensively evaluated on a real-world medical insurance dataset. The experimental results demonstrate the superiority of our proposed approach, which could detect more organized fraud claims with relatively high precision compared with state-of-the-art baselines. Jiacheng Ma 0006, Fan Li 0016, Rui Zhang 0003, Zhikang Xu, Dawei Cheng, Ruihui Zhao, Jianguang Zheng, Yefeng Zheng 0001, Changjun Jiang 0002 |
IJCAI | 2 |