EDBT 2026 Demo / reviewers in the wild / expert
Shaohua Fan
dblp:93/1263
· DBLP profile ↗
22ranked-venue papers
10as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unleashing the Power of Pre-trained Graph Models in Federated Graph Learning
Huabin Sun, Bo Yan 0005, Shaohua Fan, Yang Cao 0011, Chuan Shi 0001 |
DASFAA (2) | 4 |
| 2026 | Protect NTN-IoT Security by Malicious Traffic Detection: A Multidimensional Hypergraph Learning ApproachabstractThe vast number of devices and the complexity of requirements present significant challenges in ensuring the security of Non-Terrestrial Internet of Things (NT-IoT). Although existing studies have proposed methods like to defend against data theft and network interference attacks, there is still a need for more in-depth research on detecting data-level attacks in NTNs. Moreover, the vast and diverse nature of network traffic presents significant challenges in traffic modeling and feature extraction. Hypergraph neural networks have gained considerable attention because of capabilities in data modeling and feature extraction. However, most existing hypergraph neural networks are tailored for specific applications and are not adaptable to the detection of malicious encrypted traffic. To address these challenges, we firstly propose a hypergraph neural network-based malicious encrypted traffic detection framework to enhance the resilience of NT-IoT, enabling attack detection across unmanned aerial vehicles, base stations and satellites. Then, we introduce a Multidimensional Encrypted Traffic HyperGraph Network (METHGN). METHGN models the encrypted traffic from network, connection and time dimensions using hypergraph and uses hypergraph convolution network to extracts and fuse features. We conducted comparative experiments on IoT and The Onion Router Network encrypted traffic datasets for different classification tasks. Extensive experiments demonstrate the effectiveness and superiority of our approach. Xuzeng Li, Tao Zhang 0063, Jian Wang 0015, Zhen Han 0001, Nan Wang 0015, Shaohua Fan, Hongyang Du 0001, Jiawen Kang 0001, Jiqiang Liu, Dusit Niyato |
IEEE Internet Things J. | 6 |
| 2025 | PDMC: Generating Feasible Algorithmic Recourse via Perturbation Data Manifold ConstraintabstractTo provide actionable insights and interpretations for individuals affected by algorithmic decisions, algorithmic recourse-demonstrating how outcomes change with modifications to input features-is introduced to facilitate outcome adjustment.However, existing studies often focus on different notions of feasibility and impose complex optimization constraints, relying on strong assumptions and expert knowledge that may be impractical or not widely applicable.In this paper, we propose leveraging adherence to the perturbation data manifold to model typical feasibility challenges, providing both a theoretical clarification and a practical framework.We design optimization constraints based on this model and introduce our method, the Perturbation Data Manifold Constraint (PDMC), to ensure the feasibility of generated algorithmic recourses.Through extensive experiments on both simulated and real clinical data, we validate the rationale and effectiveness of PDMC. Hao Zou 0001, Han Yu 0009, Shaohua Fan, Haotian Wang 0001, Yue He 0001, Peng Cui 0001 |
KDD (2) | 4 |
| 2025 | AuCoGNN: Enhancing Graph Fairness Learning Under Distribution Shifts With Automated Graph GenerationabstractGraph neural networks (GNNs) have shown strong performance on graph-structured data but may inherit bias from training data, leading to discriminatory predictions based on sensitive attributes like gender and race. Existing fairness methods assume that training and testing data share the same distribution, but how fairness is affected under distribution shifts remains largely unexplored. To address this, we first identify theoretical factors that cause bias in graphs and explore how fairness is influenced by distribution shifts, particularly focusing on representation distances between groups in training and testing graphs. Based on this, we propose FatraGNN, which uses a graph generator to create biased graphs from different distributions and an alignment module to reduce representation distances for specific groups. This improves fairness and classification performance on unseen graphs. However, FatraGNN has limitations in generating realistic graphs and addressing group differentiation. To overcome these, we introduce AuCoGNN, which includes an automated graph generation module and a contrastive alignment mechanism. This ensures better fairness by maximizing the representation distance between the same certain groups while minimizing the representation distance between different groups. Experiments on real-world and semi-synthetic datasets demonstrate the effectiveness of both models in improving fairness and accuracy. Xiao Wang 0017, Yujie Xing, Shaohua Fan, Chuan Shi 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Graph Contrastive Invariant Learning from the Causal PerspectiveabstractGraph contrastive learning (GCL), learning the node representation by contrasting two augmented graphs in a self-supervised way, has attracted considerable attention. GCL is usually believed to learn the invariant representation. However, does this understanding always hold in practice? In this paper, we first study GCL from the perspective of causality. By analyzing GCL with the structural causal model (SCM), we discover that traditional GCL may not well learn the invariant representations due to the non-causal information contained in the graph. How can we fix it and encourage the current GCL to learn better invariant representations? The SCM offers two requirements and motives us to propose a novel GCL method. Particularly, we introduce the spectral graph augmentation to simulate the intervention upon non-causal factors. Then we design the invariance objective and independence objective to better capture the causal factors. Specifically, (i) the invariance objective encourages the encoder to capture the invariant information contained in causal variables, and (ii) the independence objective aims to reduce the influence of confounders on the causal variables. Experimental results demonstrate the effectiveness of our approach on node classification tasks. Yanhu Mo, Xiao Wang 0017, Shaohua Fan, Chuan Shi 0001 |
AAAI | 3 |
| 2024 | Graph Fairness Learning under Distribution ShiftsabstractGraph neural networks (GNNs) have achieved remarkable performance on graph-structured data. However, GNNs may inherit prejudice from the training data and make discriminatory predictions based on sensitive attributes, such as gender and race. Recently, there has been an increasing interest in ensuring fairness on GNNs, but all of them are under the assumption that the training and testing data are under the same distribution, i.e., training data and testing data are from the same graph. Will graph fairness performance decrease under distribution shifts? How does distribution shifts affect graph fairness learning? All these open questions are largely unexplored from a theoretical perspective. To answer these questions, we first theoretically identify the factors that determine bias on a graph. Subsequently, we explore the factors influencing fairness on testing graphs, with a noteworthy factor being the representation distances of certain groups between the training and testing graph. Motivated by our theoretical analysis, we propose our framework FatraGNN. Specifically, to guarantee fairness performance on unknown testing graphs, we propose a graph generator to produce numerous graphs with significant bias and under different distributions. Then we minimize the representation distances for each certain group between the training graph and generated graphs. This empowers our model to achieve high classification and fairness performance even on generated graphs with significant bias, thereby effectively handling unknown testing graphs. Experiments on real-world and semi-synthetic datasets demonstrate the effectiveness of our model in terms of both accuracy and fairness. Xiao Wang 0017, Yujie Xing, Shaohua Fan, Chuan Shi 0001 |
WWW | 4 |
| 2024 | Generalizing Graph Neural Networks on Out-of-Distribution GraphsabstractGraph Neural Networks (GNNs) are proposed without considering the agnostic distribution shifts between training graphs and testing graphs, inducing the degeneration of the generalization ability of GNNs in Out-Of-Distribution (OOD) settings. The fundamental reason for such degeneration is that most GNNs are developed based on the I.I.D hypothesis. In such a setting, GNNs tend to exploit subtle statistical correlations existing in the training set for predictions, even though it is a spurious correlation. This learning mechanism inherits from the common characteristics of machine learning approaches. However, such spurious correlations may change in the wild testing environments, leading to the failure of GNNs. Therefore, eliminating the impact of spurious correlations is crucial for stable GNN models. To this end, in this paper, we argue that the spurious correlation exists among subgraph-level units and analyze the degeneration of GNN in causal view. Based on the causal view analysis, we propose a general causal representation framework for stable GNN, called StableGNN. The main idea of this framework is to extract high-level representations from raw graph data first and resort to the distinguishing ability of causal inference to help the model get rid of spurious correlations. Particularly, to extract meaningful high-level representations, we exploit a differentiable graph pooling layer to extract subgraph-based representations by an end-to-end manner. Furthermore, inspired by the confounder balancing techniques from causal inference, based on the learned high-level representations, we propose a causal variable distinguishing regularizer to correct the biased training distribution by learning a set of sample weights. Hence, GNNs would concentrate more on the true connection between discriminative substructures and labels. Extensive experiments are conducted on both synthetic datasets with various distribution shift degrees and eight real-world OOD graph datasets. The results well verify that the proposed model StableGNN not only outperforms the state-of-the-arts but also provides a flexible framework to enhance existing GNNs. In addition, the interpretability experiments validate that StableGNN could leverage causal structures for predictions. Shaohua Fan, Xiao Wang 0017, Chuan Shi 0001, Peng Cui 0001, Bai Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Debiased Graph Neural Networks With Agnostic Label Selection BiasabstractMost existing graph neural networks (GNNs) are proposed without considering the selection bias in data, i.e., the inconsistent distribution between the training set with the test set. In reality, the test data are not even available during the training process, making selection bias agnostic. Training GNNs with biased selected nodes leads to significant parameter estimation bias and greatly impacts the generalization ability on test nodes. In this article, we first present an experimental investigation, which clearly shows that the selection bias drastically hinders the generalization ability of GNNs, and theoretically proves that the selection bias will cause the biased estimation on GNN parameters. Then to remove the bias in GNN estimation, we propose a novel debiased GNNs (DGNN) with a differentiated decorrelation regularizer. The differentiated decorrelation regularizer estimates a sample weight for each labeled node such that the spurious correlation of learned embeddings could be eliminated. We analyze the regularizer in causal view and it motivates us to differentiate the weights of the variables based on their contribution to the confounding bias. Then, these sample weights are used for reweighting GNNs to eliminate the estimation bias, and thus, help to improve the stability of prediction on unknown test nodes. Comprehensive experiments are conducted on several challenging graph datasets with two kinds of label selection biases. The results well verify that our proposed model outperforms the state-of-the-art methods and DGNN is a flexible framework to enhance existing GNNs. Shaohua Fan, Xiao Wang 0017, Chuan Shi 0001, Kun Kuang 0001, Nian Liu 0001, Bai Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Directed Acyclic Graph Structure Learning from Dynamic GraphsabstractEstimating the structure of directed acyclic graphs (DAGs) of features (variables) plays a vital role in revealing the latent data generation process and providing causal insights in various applications. Although there have been many studies on structure learning with various types of data, the structure learning on the dynamic graph has not been explored yet, and thus we study the learning problem of node feature generation mechanism on such ubiquitous dynamic graph data. In a dynamic graph, we propose to simultaneously estimate contemporaneous relationships and time-lagged interaction relationships between the node features. These two kinds of relationships form a DAG, which could effectively characterize the feature generation process in a concise way. To learn such a DAG, we cast the learning problem as a continuous score-based optimization problem, which consists of a differentiable score function to measure the validity of the learned DAGs and a smooth acyclicity constraint to ensure the acyclicity of the learned DAGs. These two components are translated into an unconstraint augmented Lagrangian objective which could be minimized by mature continuous optimization techniques. The resulting algorithm, named GraphNOTEARS, outperforms baselines on simulated data across a wide range of settings that may encounter in real-world applications. We also apply the proposed approach on two dynamic graphs constructed from the real-world Yelp dataset, demonstrating our method could learn the connections between node features, which conforms with the domain knowledge. Shaohua Fan, Xiao Wang 0017, Chuan Shi 0001 |
AAAI | 1 |
| 2023 | A Survey on Heterogeneous Graph Embedding: Methods, Techniques, Applications and SourcesabstractHeterogeneous graphs (HGs) also known as heterogeneous information networks have become ubiquitous in real-world scenarios; therefore, HG embedding, which aims to learn representations in a lower-dimension space while preserving the heterogeneous structures and semantics for downstream tasks (e.g., node/graph classification, node clustering, link prediction), has drawn considerable attentions in recent years. In this survey, we perform a comprehensive review of the recent development on HG embedding methods and techniques. We first introduce the basic concepts of HG and discuss the unique challenges brought by the heterogeneity for HG embedding in comparison with homogeneous graph representation learning; and then we systemically survey and categorize the state-of-the-art HG embedding methods based on the information they used in the learning process to address the challenges posed by the HG heterogeneity. In particular, for each representative HG embedding method, we provide detailed introduction and further analyze its pros and cons; meanwhile, we also explore the transformativeness and applicability of different types of HG embedding methods in the real-world industrial environments for the first time. In addition, we further present several widely deployed systems that have demonstrated the success of HG embedding techniques in resolving real-world application problems with broader impacts. To facilitate future research and applications in this area, we also summarize the open-source code, existing graph learning platforms and benchmark datasets. Finally, we explore the additional issues and challenges of HG embedding and forecast the future research directions in this field. Xiao Wang 0017, Deyu Bo, Chuan Shi 0001, Shaohua Fan, Yanfang Ye 0001, Philip S. Yu |
IEEE Trans. Big Data | 4 |
| 2022 | Debiasing Graph Neural Networks via Learning Disentangled Causal SubstructureabstractMost Graph Neural Networks (GNNs) predict the labels of unseen graphs by learning the correlation between the input graphs and labels. However, by presenting a graph classification investigation on the training graphs with severe bias, surprisingly, we discover that GNNs always tend to explore the spurious correlations to make decision, even if the causal correlation always exists. This implies that existing GNNs trained on such biased datasets will suffer from poor generalization capability. By analyzing this problem in a causal view, we find that disentangling and decorrelating the causal and bias latent variables from the biased graphs are both crucial for debiasing. Inspired by this, we propose a general disentangled GNN framework to learn the causal substructure and bias substructure, respectively. Particularly, we design a parameterized edge mask generator to explicitly split the input graph into causal and bias subgraphs. Then two GNN modules supervised by causal/bias-aware loss functions respectively are trained to encode causal and bias subgraphs into their corresponding representations. With the disentangled representations, we synthesize the counterfactual unbiased training samples to further decorrelate causal and bias variables. Moreover, to better benchmark the severe bias problem, we construct three new graph datasets, which have controllable bias degrees and are easier to visualize and explain. Experimental results well demonstrate that our approach achieves superior generalization performance over existing baselines. Furthermore, owing to the learned edge mask, the proposed model has appealing interpretability and transferability. Shaohua Fan, Xiao Wang 0017, Yanhu Mo, Chuan Shi 0001, Jian Tang 0005 |
NeurIPS | 1 |
| 2020 | Encrypted Traffic Classification Using Graph Convolutional Networks
Shuang Mo, Ding Xiao, Wenrui Wu, Shaohua Fan, Chuan Shi 0001 |
ADMA | 5 |
| 2020 | Decorrelated Clustering with Data Selection BiasabstractMost of existing clustering algorithms are proposed without considering the selection bias in data. In many real applications, however, one cannot guarantee the data is unbiased. Selection bias might bring the unexpected correlation between features and ignoring those unexpected correlations will hurt the performance of clustering algorithms. Therefore, how to remove those unexpected correlations induced by selection bias is extremely important yet largely unexplored for clustering. In this paper, we propose a novel Decorrelation regularized K-Means algorithm (DCKM) for clustering with data selection bias. Specifically, the decorrelation regularizer aims to learn the global sample weights which are capable of balancing the sample distribution, so as to remove unexpected correlations among features. Meanwhile, the learned weights are combined with k-means, which makes the reweighted k-means cluster on the inherent data distribution without unexpected correlation influence. Moreover, we derive the updating rules to effectively infer the parameters in DCKM. Extensive experiments results on real world datasets well demonstrate that our DCKM algorithm achieves significant performance gains, indicating the necessity of removing unexpected feature correlations induced by selection bias when clustering. Xiao Wang 0017, Shaohua Fan, Kun Kuang 0001, Chuan Shi 0001, Jiawei Liu 0006, Bai Wang 0001 |
IJCAI | 2 |
| 2020 | One2Multi Graph Autoencoder for Multi-view Graph ClusteringabstractMulti-view graph clustering, which seeks a partition of the graph with multiple views that often provide more comprehensive yet complex information, has received considerable attention in recent years. Although some efforts have been made for multi-view graph clustering and achieve decent performances, most of them employ shallow model to deal with the complex relation within multi-view graph, which may seriously restrict the capacity for modeling multi-view graph information. In this paper, we make the first attempt to employ deep learning technique for attributed multi-view graph clustering, and propose a novel task-guided One2Multi graph autoencoder clustering framework. The One2Multi graph autoencoder is able to learn node embeddings by employing one informative graph view and content data to reconstruct multiple graph views. Hence, the shared feature representation of multiple graphs can be well captured. Furthermore, a self-training clustering objective is proposed to iteratively improve the clustering results. By integrating the self-training and autoencoder’s reconstruction into a unified framework, our model can jointly optimize the cluster label assignments and embeddings suitable for graph clustering. Experiments on real-world attributed multi-view graph datasets well validate the effectiveness of our model. Shaohua Fan, Xiao Wang 0017, Chuan Shi 0001, Emiao Lu, Ken Lin, Bai Wang 0001 |
WWW | 1 |
| 2019 | Metapath-guided Heterogeneous Graph Neural Network for Intent RecommendationabstractWith the prevalence of mobile e-commerce nowadays, a new type of recommendation services, called intent recommendation, is widely used in many mobile e-commerce Apps, such as Taobao and Amazon. Different from traditional query recommendation and item recommendation, intent recommendation is to automatically recommend user intent according to user historical behaviors without any input when users open the App. Intent recommendation becomes very popular in the past two years, because of revealing user latent intents and avoiding tedious input in mobile phones. Existing methods used in industry usually need laboring feature engineering. Moreover, they only utilize attribute and statistic information of users and queries, and fail to take full advantage of rich interaction information in intent recommendation, which may result in limited performances. In this paper, we propose to model the complex objects and rich interactions in intent recommendation as a Heterogeneous Information Network. Furthermore, we present a novel M etapath-guided E mbedding method for I ntent Rec ommendation~(called MEIRec). In order to fully utilize rich structural information, we design a metapath-guided heterogeneous Graph Neural Network to learn the embeddings of objects in intent recommendation. In addition, in order to alleviate huge learning parameters in embeddings, we propose a uniform term embedding mechanism, in which embeddings of objects are made up with the same term embedding space. Offline experiments on real large-scale data show the superior performance of the proposed MEIRec, compared to representative methods.Moreover, the results of online experiments on Taobao e-commerce platform show that MEIRec not only gains a performance improvement of 1.54% on CTR metric, but also attracts up to 2.66% of new users to search queries. Shaohua Fan, Junxiong Zhu, Chuan Shi 0001, Linmei Hu, Biyu Ma, Yongliang Li |
KDD | 1 |
| 2018 | Abnormal Event Detection via Heterogeneous Information Network EmbeddingabstractHeteregeneous information networks (HINs) are ubiquitous in the real world, and discovering the abnormal events plays an important role in understanding and analyzing the HIN. The abnormal event usually implies that the number of co-occurrences of entities in a HIN are very rare, so most of the existing works are based on detecting the rare patterns of events. However, we find that the number of co-occurrences of majority entities in events are the same, which brings great challenge to distinguish the normal and abnormal events. Therefore, we argue that considering the heterogeneous information structure only is not sufficient for abnormal event detection and introducing additional valuable information is necessary. In this paper, we propose a novel deep heterogeneous network embedding method which incorporates the entity attributes and second-order structures simultaneously to address this problem. Specifically, we utilize type-aware Multilayer Perceptron (MLP) component to learn the attribute embedding, and adopt the autoencoder framework to learn the second-order aware embedding. Then based on the mixed embeddings, we are able to model the pairwise interactions of different entities, such that the events with small entity compatibilities have large abnormal event score. The experimental results on real world network demonstrate the effectiveness of our proposed method. Shaohua Fan, Chuan Shi 0001, Xiao Wang 0017 |
CIKM | 1 |
| 2015 | Robust and efficient adaptive direct lighting estimation
Yu-Chi Lai, Hsuan-Ting Chou, Kuo-Wei Chen, Shaohua Fan |
Vis. Comput. | 4 |
| 2010 | Animation rendering with Population Monte Carlo image-plane sampler
Yu-Chi Lai, Stephen Chenney, Feng Liu 0015, Yuzhen Niu, Shaohua Fan |
Vis. Comput. | 5 |
| 2007 | Photorealistic Image Rendering with Population Monte Carlo Energy Redistribution
Yu-Chi Lai, Shaohua Fan, Stephen Chenney, Charcle Dyer |
Rendering Techniques | 2 |
| 2006 | Optimizing Control Variate Estimators for RenderingabstractAbstract We present the Optimizing Control Variate (OCV) estimator, a new estimator for Monte Carlo rendering. Based upon a deterministic sampling framework, OCV allows multiple importance sampling functions to be combined in one algorithm. Its optimizing nature addresses a major problem with control variate estimators for rendering: users supply a generic correlated function which is optimized for each estimate, rather than a single highly tuned one that must work well everywhere. We demonstrate OCV with both direct lighting and irradiance‐caching examples, showing improvements in image error of over 35% in some cases, for little extra computation time. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Three‐Dimensional Graphics and Realism Color, shading, shadowing, and texture G.3 [Probability and Statistics]: Probabilistic Algorithms Keywords: direct lighting, deterministic mixture sampling, control variates Shaohua Fan, Stephen Chenney, Kam-Wah Tsui, Yu-Chi Lai |
Comput. Graph. Forum | 1 |
| 2005 | Metropolis Photon Sampling with Optional User Guidance
Shaohua Fan, Stephen Chenney, Yu-Chi Lai |
Rendering Techniques | 1 |
| 2003 | An Automatic System for Classification of Nuclear Sclerosis from Slit-Lamp Photographs
Shaohua Fan, Charles R. Dyer, Larry Hubbard, Barbara Klein |
MICCAI (1) | 1 |