Yangxin Fan

dblp:291/3935 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-1728-9560ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Generating Skyline Datasets for Data Science Models
Mengying Wang 0001, Hanchao Ma, Yiyang Bian, Yangxin Fan, Yinghui Wu 0001
EDBT4
2025 Graph Compression for Interpretable Graph Neural Network Inference At Scale
abstract
We demonstrate ExGIS , a parallel inference query engine to support explainable Graph Neural Network (GNNs) inference analysis in large graphs. (1) For a class of GNNs ℳ L with at most L layers, and a graph G , ExGIS performs an offline, once-for-all compression of G to a small graph G c , such that for any inference query Q that requests the output of any GNN M ∈ ℳ L on any node v in G, G c can be directly queried to yield correct output without decompression. (2) Given a workload W of inference queries that requests the output of GNNs from ℳ over G , ExGIS perform fast online GNN inference and interpretation in parallel. It dynamically partitions W to balance workloads, and (a) executes inference that only consults compressed graph G c without decompression, and (b) directly yields concise, explanatory subgraphs from G c that can clarify the query output with high fidelity, all in parallel. Moreover, ExGIS integrates visual, interactive interfaces for query performance analysis, and a Large Language Models (LLMs)- enabled interpreter to support user-friendly, natural language explanation of query outputs. We demonstrate the compression rate and scalability of ExGIS, and its application in interpretable anomaly detection over bitcoin transaction networks and academic networks.
Yangxin Fan, Haolai Che, Mingjian Lu, Yinghui Wu 0001
Proc. VLDB Endow.1
2025 Inference-friendly Graph Compression for Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have demonstrated promising performance in graph analysis. Nevertheless, the inference process of GNNs remains costly, hindering their applications for large graphs. This paper proposes inference-friendly graph compression (IFGC), a graph compression scheme to accelerate GNNs inference. Given a graph G and a GNN M , an IFGC computes a small compressed graph G c , to best preserve the inference results of M over G , such that the result can be directly inferred by accessing G c with no or little decompression cost. (1) We characterize IFGC with a class of inference equivalence relation. The relation captures the node pairs in G that are not distinguishable for GNN inference. (2) We introduce three practical specifications of IFGC for representative GNNs: structural preserving compression (SPGC), which computes G c that can be directly processed by GNN inference without decompression; ( α,r )-compression, that allows for a configurable trade-off between compression ratio and inference quality, and anchored compression that preserves inference results for specific nodes of interest. For each scheme, we introduce compression and inference algorithms with guarantees of efficiency and quality of the inferred results. We conduct extensive experiments on diverse sets of large-scale graphs, which verifies the effectiveness and efficiency of our graph compression approaches.
Yangxin Fan, Haolai Che, Yinghui Wu 0001
Proc. VLDB Endow.1
2024 Parallel-friendly Spatio-Temporal Graph Learning for Photovoltaic Degradation Analysis at Scale
abstract
Photovoltaic (PV) power stations have become an integral component to the global sustainable energy landscape. Accurately monitoring and estimating the performance of PV systems is critical to their feasibility for power generation and as a financial asset. One of the most challenging problems is to understand and estimate the long-term Performance Loss Rate (PLR) for large fleets of PV inverters. This paper introduces a novel Spatio-Temporal Graph Neural Network empowered, long-term Trend analysis system (ST-GTrend), to estimate PLR of PV systems at fleet-level. ST-GTrend nontrivially integrates spatio-temporal coherence and graph attention to separate PLR as a long-term 'aging' trend from multiple fluctuation terms in the PV input data, with a design that can easily scale to large PV sets with effective, multi-level parallel computation. (1) To cope with diverse degradation patterns in timeseries, ST-GTrend adopts a paralleled graph autoencoder array to extract aging and fluctuation terms simultaneously, and imposes flatness and smoothness regularizations to disentangle between aging and fluctuation. (2) For large PV systems, ST-GTrend enables a multi-level parallelization paradigm to scale the training and inference computation with a provable performance guarantee. ST-GTrend has been deployed in CRADLE, a scientific high performance computing infrastructure. We evaluated ST-GTrend with three real-world large-scale PV datasets, spanning a time period of 10 years. Our results show that ST-GTrend reduces MAPE and Euclidean distance-based errors on average by 34.74% and 33.66% of SOTA methods, and scales well to large PV sets. We also showcase that the advantages of ST-GTrend generalize for the need of long-term trend analysis in financial and economic data.
Yangxin Fan, Raymond Wieser, Laura S. Bruckman, Roger H. French, Yinghui Wu 0001
CIKM1
2023 Spatio-Temporal Denoising Graph Autoencoders with Data Augmentation for Photovoltaic Data Imputation
abstract
The integration of the global Photovoltaic (PV) market with real time data-loggers has enabled large scale PV data analytical pipelines for power forecasting and reliability assessment of PV fleets. Nevertheless, the performance of PV data analysis depends on the quality of PV timeseries data. We propose a novel Spatio-Temporal Denoising Graph Autoencoder STD-GAE framework to impute missing PV Power Data. STD-GAE exploits temporal correlation, spatial coherence, and value dependencies from domain knowledge to recover missing data. It is empowered by two modules. (1) To cope with sparse yet various scenarios of missing data, STD-GAE incorporates a domain-knowledge aware data augmentation module to create plausible variations of missing data patterns. This generalizes STD-GAE to robust imputation over different seasons and environment. (2) STD-GAE nontrivially integrates spatiotemporal graph convolution layers and denoising autoencoder to improve the accuracy of imputation accuracy at PV fleet level. Experimental results on two PV datasets show that STD-GAE can achieve a gain of 43.14% in imputation accuracy and remains less sensitive to missing rate, different seasons, and missing scenarios, compared with state-of-the-art data imputation methods.
Yangxin Fan, Xuanji Yu, Raymond Wieser, David Meakin, Avishai Shaton, Jean-Nicolas Jaubert, Robert Flottemesch, Michael Howell, Jennifer Braid, Laura S. Bruckman, Roger H. French, Yinghui Wu 0001
Proc. ACM Manag. Data1
2023 Understanding Public Opinion Toward the #StopAsianHate Movement and the Relation With Racially Motivated Hate Crimes in the US
abstract
#StopAsianHate and #StopAAPIHate are two of the most commonly used hashtags that represent the current movement to end hate crimes against the Asian American and Pacific Islander communities. We conduct a social media study of public opinion toward the #StopAsianHate and #StopAAPIHate movement based on 46 058 Twitter users across 30 states in USA ranging from March 18, 2021 to April 11, 2021. To facilitate fine-grained analyses, the demographic information of the Twitter users, including age, gender, race/ethnicity, social capital, political affiliation, geolocation, family, income, and religious status, is either retrieved from the user profiles or inferred using classifiers. We find that the movement attracts more participation from women, younger adults, Asian, and Black communities. It is noteworthy that most of them are also active in other online movements related to racial or social issues, such as #BlackLivesMatter or #SayHerName; 51.56% of the Twitter users show direct support, 18.38% are news about anti-Asian hate crimes, and 5.43% show a negative attitude toward the movement. By conducting logistic regression, we find that the public opinion varies across user characteristics. Furthermore, among the states with most racial bias-motivated hate crimes, the negative attitude toward the #StopAsianHate and #StopAAPIHate movement is the weakest. The majority of the top influencers are Asian–American reporters, journalists, or politicians. To the best of our knowledge, this is the first large-scale social media-based study to understand public opinion toward the #StopAsianHate and #StopAAPIHate movement. We hope that our study can provide insights and promote research on anti-Asian hate crimes and ultimately help address such a serious societal issue for the common benefits of all communities.
Hanjia Lyu, Yangxin Fan, Ziyu Xiong, Mayya Komisarchik, Jiebo Luo 0001
IEEE Trans. Comput. Soc. Syst.2