EDBT 2026 Demo / reviewers in the wild / expert
Junwei Su
dblp:226/0880
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Graph learning · 42% Efficient and distributed learning · 19% Generative modeling · 17% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Distributed systems · 28% Parallel and multicore computing · 28% Performance modeling and evaluation · 28% |
Topics — the 25 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
graph neural network |
3.2 | 4 | 2025 | Heta: Distributed Training of Heterogeneous Graph Neural Networks · Proc. VLDB Endow. 2025 On the Interplay between Graph Structure and Learning Algorithms in Graph Neural Networks · ICML 2025 On the Topology Awareness and Generalization Performance of Graph Neural Networks · ECCV (84) 2024 |
Machine learning › Efficient and distributed learning
distributed training |
2.4 | 3 | 2025 | Heta: Distributed Training of Heterogeneous Graph Neural Networks · Proc. VLDB Endow. 2025 MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline · KDD 2024 PRES: Toward Scalable Memory-Based Dynamic Graph Neural Networks · ICLR 2024 |
Machine learning › Graph learning › graph neural network
dynamic graph neural network |
2.4 | 3 | 2025 | Temporal-Aware Evaluation and Learning for Temporal Graph Neural Networks · AAAI 2025 MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline · KDD 2024 PRES: Toward Scalable Memory-Based Dynamic Graph Neural Networks · ICLR 2024 |
Machine learning › Generative modeling
diffusion model |
1.7 | 2 | 2025 | SBGD: Improving Graph Diffusion Generative Model via Stochastic Block Diffusion · ICML 2025 A Non-Asymptotic Convergent Analysis for Scored-Based Graph Generative Model via a System of Stochastic Differential Equations · ICML 2025 |
Machine learning › Graph learning
graph generation |
1.7 | 2 | 2025 | SBGD: Improving Graph Diffusion Generative Model via Stochastic Block Diffusion · ICML 2025 A Non-Asymptotic Convergent Analysis for Scored-Based Graph Generative Model via a System of Stochastic Differential Equations · ICML 2025 |
Parallel and multicore computing › parallel programming models
automatic parallelization |
1.0 | 1 | 2026 | HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026 |
Distributed systems › distributed machine learning
distributed training |
1.0 | 1 | 2026 | HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026 |
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training |
0.9 | 1 | 2025 | Heta: Distributed Training of Heterogeneous Graph Neural Networks · Proc. VLDB Endow. 2025 |
Machine learning › Optimization for machine learning
convergence analysis |
0.9 | 1 | 2025 | A Non-Asymptotic Convergent Analysis for Scored-Based Graph Generative Model via a System of Stochastic Differential Equations · ICML 2025 |
Machine learning › Learning theory › statistical learning theory
excess risk |
0.9 | 1 | 2025 | On the Interplay between Graph Structure and Learning Algorithms in Graph Neural Networks · ICML 2025 |
Machine learning › Learning theory
generalization bounds |
0.9 | 1 | 2025 | On the Interplay between Graph Structure and Learning Algorithms in Graph Neural Networks · ICML 2025 |
Machine learning › Generative modeling › diffusion model
graph diffusion model |
0.9 | 1 | 2025 | SBGD: Improving Graph Diffusion Generative Model via Stochastic Block Diffusion · ICML 2025 |
Machine learning › Graph learning › graph neural network
heterogeneous graph neural network |
0.9 | 1 | 2025 | Heta: Distributed Training of Heterogeneous Graph Neural Networks · Proc. VLDB Endow. 2025 |
Machine learning › Generative modeling › diffusion model
score-based generative model |
0.9 | 1 | 2025 | A Non-Asymptotic Convergent Analysis for Scored-Based Graph Generative Model via a System of Stochastic Differential Equations · ICML 2025 |
Machine learning › Deep learning architectures and training
training objective |
0.9 | 1 | 2025 | Temporal-Aware Evaluation and Learning for Temporal Graph Neural Networks · AAAI 2025 |
Machine learning › Graph learning › dynamic graph learning
memory-based temporal graph neural networks |
0.8 | 1 | 2024 | PRES: Toward Scalable Memory-Based Dynamic Graph Neural Networks · ICLR 2024 |
Performance modeling and evaluation
performance prediction |
0.8 | 1 | 2024 | CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor Programs · EuroSys 2024 |
Machine learning › Learning paradigms
incremental learning |
0.7 | 1 | 2023 | Towards Robust Graph Incremental Learning on Evolving Graphs · ICML 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › rule learning
inductive learning |
0.7 | 1 | 2023 | Towards Robust Graph Incremental Learning on Evolving Graphs · ICML 2023 |
Machine learning › Graph learning › graph neural network
node classification |
0.7 | 1 | 2023 | Towards Robust Graph Incremental Learning on Evolving Graphs · ICML 2023 |
Hardware accelerators and domain-specific architectures › accelerator architecture
heterogeneous accelerator |
0.3 | 1 | 2026 | HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026 |
GPUs and heterogeneous computing › GPU memory management
GPU cache management |
0.3 | 1 | 2025 | Heta: Distributed Training of Heterogeneous Graph Neural Networks · Proc. VLDB Endow. 2025 |
Machine learning › Efficient and distributed learning › distributed training › asynchronous training
staleness-aware training |
0.2 | 1 | 2024 | MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline · KDD 2024 |
Performance modeling and evaluation › processor performance modeling
accelerator performance modeling |
0.2 | 1 | 2024 | CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor Programs · EuroSys 2024 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.2 | 1 | 2023 | Towards Robust Graph Incremental Learning on Evolving Graphs · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
relation-aggregation-first paradigm · 1.7meta-partitioning · 1.7heterogeneity-aware caching · 1.7random forest · 1.0monte carlo tree search · 1.0cost model · 1.0communication overlap · 1.0volatility cluster statistics · 0.9stochastic gradient descent · 0.9stochastic differential equation · 0.9stochastic block diffusion · 0.9spectral graph theory · 0.9ridge regression · 0.9modularization · 0.9domain adaptation · 0.8compact AST representation · 0.8KMeans sampling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed TrainingabstractAs large neural network models (e.g., LLMs) grow in scale, single-cluster resources become insufficient, making cross-cluster distributed training essential. Cross-cluster training is challenging: hardware heterogeneity complicates load balancing and parallelization strategy and introduces hardware compatibility issues in implementation; cross-cluster communication bottlenecks severely impact training throughput. We present HetAuto, an automatic parallelization system for efficient cross-cluster heterogeneous large model training. HetAuto contributes three key innovations: (1) a principle-guided MCTS algorithm with a random forest-enhanced cost model that efficiently searches parallelization strategies and quickly evaluates their performance under heterogeneous configurations; (2) cross-cluster communication optimizations including Virtual-1F1B scheduling that overlaps communication with computation and an optimized resharding strategy for inter-stage communication; and (3) a unified API enabling seamless integration of diverse accelerators. We evaluate HetAuto across 4 different clusters with up to 736 heterogeneous devices. The evaluation results show that HetAuto achieves up to 1.57× training throughput improvement over representative baselines, and strikes an efficient balance between solution quality and search overhead. Guicheng Qi, Junwei Su, Liqi Yang, Tingwen Xie, Yerui Sun, Chuan Wu 0001 |
EuroSys | 2 |
| 2025 | Temporal-Aware Evaluation and Learning for Temporal Graph Neural NetworksabstractTemporal Graph Neural Networks (TGNNs) are a family of graph neural networks designed to model and learn dynamic information from temporal graphs. Given their substantial empirical success, there is an escalating interest in TGNNs within the research community. However, the majority of these efforts have been channelled towards algorithm and system design, with the evaluation metrics receiving comparatively less attention. Effective evaluation metrics are crucial for providing detailed performance insights, particularly in the temporal domain. This paper investigates the commonly used evaluation metrics for TGNNs and illustrates the failure mechanisms of these metrics in capturing essential temporal structures in the predictive behaviour of TGNNs. We provide a mathematical formulation of existing performance metrics and utilize an instance-based study to underscore their inadequacies in identifying volatility clustering (the occurrence of emerging errors within a brief interval). This phenomenon has profound implications for both algorithm and system design in the temporal domain. To address this deficiency, we introduce a new volatility-aware evaluation metric (termed volatility cluster statistics), designed for a more refined analysis of model temporal performance. Additionally, we demonstrate how this metric can serve as a temporal-volatility-aware training objective to alleviate the clustering of temporal errors. Through comprehensive experiments on various TGNN models, we validate our analysis and the proposed approach. The empirical results offer revealing insights: 1) existing TGNNs are prone to making errors with volatility clustering, and 2) TGNNs with different mechanisms to capture temporal information exhibit distinct volatility clustering patterns. Moreover, our empirical findings demonstrate that our proposed training objective effectively reduces volatility clusters in error. Junwei Su |
AAAI | 1 |
| 2025 | On the Interplay between Graph Structure and Learning Algorithms in Graph Neural NetworksabstractThis paper studies the interplay between learning algorithms and graph structure for graph neural networks (GNNs). Existing theoretical studies on the learning dynamics of GNNs primarily focus on the convergence rates of learning algorithms under the interpolation regime (noise-free) and offer only a crude connection between these dynamics and the actual graph structure (e.g., maximum degree). This paper aims to bridge this gap by investigating the excessive risk (generalization performance) of learning algorithms in GNNs within the generalization regime (with noise). Specifically, we extend the conventional settings from the learning theory literature to the context of GNNs and examine how graph structure influences the performance of learning algorithms such as stochastic gradient descent (SGD) and Ridge regression. Our study makes several key contributions toward understanding the interplay between graph structure and learning in GNNs. First, we derive the excess risk profiles of SGD and Ridge regression in GNNs and connect these profiles to the graph structure through spectral graph theory. With this established framework, we further explore how different graph structures (regular vs. power-law) impact the performance of these algorithms through comparative analysis. Additionally, we extend our analysis to multi-layer linear GNNs, revealing an increasing non-isotropic effect on the excess risk profile, thereby offering new insights into the over-smoothing issue in GNNs from the perspective of learning algorithms. Our empirical results align with our theoretical predictions, collectively showcasing a coupling relation among graph structure, GNNs and learning algorithms, and providing insights on GNN algorithm design and selection in practice. Junwei Su, Chuan Wu 0001 |
ICML | 1 |
| 2025 | A Non-Asymptotic Convergent Analysis for Scored-Based Graph Generative Model via a System of Stochastic Differential EquationsabstractThis paper investigates the convergence behavior of score-based graph generative models (SGGMs). Unlike common score-based generative models (SGMs) that are governed by a single stochastic differential equation (SDE), SGGMs utilize a system of dependent SDEs, where the graph structure and node features are modeled separately, while accounting for their inherent dependencies. This distinction makes existing convergence analyses from SGMs inapplicable for SGGMs. In this work, we present the first convergence analysis for SGGMs, focusing on the convergence bound (the risk of generative error) across three key graph generation paradigms: (1) feature generation with a fixed graph structure, (2) graph structure generation with fixed node features, and (3) joint generation of both graph structure and node features. Our analysis reveals several unique factors specific to SGGMs (e.g., the topological properties of the graph structure) which significantly affect the convergence bound. Additionally, we offer theoretical insights into the selection of hyperparameters (e.g., sampling steps and diffusion length) and advocate for techniques like normalization to improve convergence. To validate our theoretical findings, we conduct a controlled empirical study using a synthetic graph model. The results in this paper contribute to a deeper theoretical understanding of SGGMs and offer practical guidance for designing more efficient and effective SGGMs. Junwei Su, Chuan Wu 0001 |
ICML | 1 |
| 2025 | SBGD: Improving Graph Diffusion Generative Model via Stochastic Block DiffusionabstractGraph diffusion generative models (GDGMs) have emerged as powerful tools for generating high-quality graphs. However, their broader adoption faces challenges in scalability and size generalization. GDGMs struggle to scale to large graphs due to their high memory requirements, as they typically operate in the full graph space, requiring the entire graph to be stored in memory during training and inference. This constraint limits their feasibility for large-scale real-world graphs. GDGMs also exhibit poor size generalization, with limited ability to generate graphs of sizes different from those in the training data, restricting their adaptability across diverse applications. To address these challenges, we propose the stochastic block graph diffusion (SBGD) model, which refines graph representations into a block graph space. This space incorporates structural priors based on real-world graph patterns, significantly reducing memory complexity and enabling scalability to large graphs. The block representation also improves size generalization by capturing fundamental graph structures. Empirical results show that SBGD achieves significant memory improvements (up to 6$\times$) while maintaining comparable or even superior graph generation performance relative to state-of-the-art methods. Furthermore, experiments demonstrate that SBGD better generalizes to unseen graph sizes. The significance of SBGD extends beyond being a scalable and effective GDGM; it also exemplifies the principle of modularization in generative modelling, offering a new avenue for exploring generative models by decomposing complex tasks into more manageable components. Junwei Su |
ICML | 1 |
| 2025 | Heta: Distributed Training of Heterogeneous Graph Neural NetworksabstractHeterogeneous Graphs (HetGs) that capture relationships among different types of nodes are ubiquitous in real-world applications such as academic networks and e-commerce. Although Heterogeneous Graph Neural Networks (HGNNs) have demonstrated superior performance in learning from these complex structures, distributed training of HGNNs on large-scale graphs with billions of edges faces substantial communication overhead. This challenge is exacerbated by heterogeneous characteristics such as varying feature dimensions across node types and featureless nodes requiring learnable parameters. Existing systems and communication reduction techniques designed for homogeneous graphs become suboptimal or even inapplicable for HetGs and HGNNs by overlooking both these heterogeneous characteristics and the inherent computational structure of HGNNs. We present Heta , a framework designed to address the communication bottleneck in distributed HGNN training. Heta leverages the key insight that HGNN aggregation is order-invariant and decomposable into relation-specific computations. Built on this insight, we introduce three key innovations: (1) a Relation-Aggregation-First (RAF) paradigm that conducts relation-specific aggregations within partitions and exchanges only partial aggregations across machines, proven to reduce communication complexity; (2) a meta-partitioning strategy that divides a HetG based on its graph schema and HGNN computation dependency while minimizing cross-partition communication and maintaining computation and storage balance; and (3) a heterogeneity-aware GPU cache system that accounts for varying miss-penalty ratios across node types. Through extensive evaluation of billion-edge heterogeneous graphs, we demonstrate that Heta achieves up to 5.3X and 4.4X speedup over state-of-the-art systems DGL and GraphLearn while maintaining model accuracy. Yuchen Zhong, Junwei Su, Chuan Wu 0001 |
Proc. VLDB Endow. | 2 |
| 2024 | On the Topology Awareness and Generalization Performance of Graph Neural Networks
Junwei Su, Chuan Wu 0001 |
ECCV (84) | 1 |
| 2024 | CDMPP: A Device-Model Agnostic Framework for Latency Prediction of Tensor ProgramsabstractDeep Neural Networks (DNNs) have shown excellent performance in a wide range of machine learning applications. Knowing the latency of running a DNN model or tensor program on a specific device is useful in various tasks, such as DNN graph- or tensor-level optimization and device selection. Considering the large space of DNN models and devices that impedes direct profiling of all combinations, recent efforts focus on building a predictor to model the performance of DNN models on different devices. However, none of the existing attempts have achieved a cost model that can accurately predict the performance of various tensor programs while supporting both training and inference accelerators. We propose CDMPP, an efficient tensor program latency prediction framework for both cross-model and cross-device prediction. We design an informative but efficient representation of tensor programs, called compact ASTs, and a pre-order-based positional encoding method, to capture the internal structure of tensor programs. We develop a domain-adaption-inspired method to learn domain-invariant representations and devise a KMeans-based sampling algorithm, for the predictor to learn from different domains (i.e., different DNN operators and devices). Our extensive experiments on a diverse range of DNN models and devices demonstrate that CDMPP significantly outperforms state-of-the-art baselines with 14.03% and 10.85% prediction error for cross-model and cross-device prediction, respectively, and one order of magnitude higher training efficiency. The implementation and the expanded dataset are available at https://github.com/joapolarbear/cdmpp. Hanpeng Hu, Junwei Su, Juntao Zhao 0002, Yanghua Peng, Yibo Zhu 0001, Haibin Lin, Chuan Wu 0001 |
EuroSys | 2 |
| 2024 | MTRGL: Effective Temporal Correlation Discerning Through Multi-Modal Temporal Relational Graph LearningabstractIn this study, we explore the synergy of deep learning and financial market applications, focusing on pair trading. This market-neutral strategy is integral to quantitative finance and is apt for advanced deep-learning techniques. A pivotal challenge in pair trading is discerning temporal correlations among entities, necessitating the integration of diverse data modalities. Addressing this, we introduce a novel framework, Multi-modal Temporal Relation Graph Learning (MTRGL). MTRGL combines time series data and discrete features into a temporal graph and employs a memory-based temporal graph neural network. This approach reframes temporal correlation identification as a temporal graph link prediction task, which has shown empirical success. Our experiments on real-world datasets confirm the superior performance of MTRGL, emphasizing its promise in refining automated pair trading strategies. Junwei Su |
ICASSP | 1 |
| 2024 | PRES: Toward Scalable Memory-Based Dynamic Graph Neural NetworksabstractMemory-based Dynamic Graph Neural Networks (MDGNNs) are a family of dynamic graph neural networks that leverage a memory module to extract, distill, and memorize long-term temporal dependencies, leading to superior performance compared to memory-less counterparts. However, training MDGNNs faces the challenge of handling entangled temporal and structural dependencies, requiring sequential and chronological processing of data sequences to capture accurate temporal patterns. During the batch training, the temporal data points within the same batch will be processed in parallel, while their temporal dependencies are neglected. This issue is referred to as temporal discontinuity and restricts the effective temporal batch size, limiting data parallelism and reducing MDGNNs' flexibility in industrial applications. This paper studies the efficient training of MDGNNs at scale, focusing on the temporal discontinuity in training MDGNNs with large temporal batch sizes. We first conduct a theoretical study on the impact of temporal batch
size on the convergence of MDGNN training. Based on the analysis, we propose PRES, an iterative prediction-correction scheme combined with a memory coherence learning objective to mitigate the effect of temporal discontinuity, enabling MDGNNs to be trained with significantly larger temporal batches without sacrificing generalization performance. Experimental results demonstrate that our approach enables up to a 4 $\times$ larger temporal batch (3.4$\times$ speed-up) during MDGNN training. Junwei Su, Difan Zou, Chuan Wu 0001 |
ICLR | 1 |
| 2024 | MSPipe: Efficient Temporal GNN Training via Staleness-Aware PipelineabstractMemory-based Temporal Graph Neural Networks (MTGNNs) are a class of temporal graph neural networks that utilize a node memory module to capture and retain long-term temporal dependencies, leading to superior performance compared to memory-less counterparts. However, the iterative reading and updating process of the memory module in MTGNNs to obtain up-to-date information needs to follow the temporal dependencies. This introduces significant overhead and limits training throughput. Existing optimizations for static GNNs are not directly applicable to MTGNNs due to differences in training paradigm, model architecture, and the absence of a memory module. Moreover, these optimizations do not effectively address the challenges posed by temporal dependencies, making them ineffective for MTGNN training. In this paper, we propose MSPipe, a general and efficient framework for memory-based TGNNs that maximizes training throughput while maintaining model accuracy. Our design specifically addresses the unique challenges associated with fetching and updating node memory states in MTGNNs by integrating staleness into the memory module. However, simply introducing a predefined staleness bound in the memory module to break temporal dependencies may lead to suboptimal performance and lack of generalizability across different models and datasets. To overcome this, we introduce an online pipeline scheduling algorithm in MSPipe that strategically breaks temporal dependencies with minimal staleness and delays memory fetching to obtain fresher memory states. This is achieved without stalling the MTGNN training stage or causing resource contention. Additionally, we design a staleness mitigation mechanism to enhance training convergence and model accuracy. Furthermore, we provide convergence analysis and demonstrate that MSPipe maintains the same convergence rate as vanilla sampling-based GNN training. Experimental results show that MSPipe achieves up to 2.45× speed-up without sacrificing accuracy, making it a promising solution for efficient MTGNN training. The implementation of our paper can be found at the following link: https://github.com/PeterSH6/MSPipe. Guangming Sheng, Junwei Su, Chao Huang 0001, Chuan Wu 0001 |
KDD | 2 |
| 2023 | Towards Robust Graph Incremental Learning on Evolving GraphsabstractIncremental learning is a machine learning approach that involves training a model on a sequence of tasks, rather than all tasks at once. This ability to learn incrementally from a stream of tasks is crucial for many real-world applications. However, incremental learning is a challenging problem on graph-structured data, as many graph-related problems involve prediction tasks for each individual node, known as Node-wise Graph Incremental Learning (NGIL). This introduces non-independent and non-identically distributed characteristics in the sample data generation process, making it difficult to maintain the performance of the model as new tasks are added. In this paper, we focus on the inductive NGIL problem, which accounts for the evolution of graph structure (structural shift) induced by emerging tasks. We provide a formal formulation and analysis of the problem, and propose a novel regularization-based technique called Structural-Shift-Risk-Mitigation (SSRM) to mitigate the impact of the structural shift on catastrophic forgetting of the inductive NGIL problem. We show that the structural shift can lead to a shift in the input distribution for the existing tasks, and further lead to an increased risk of catastrophic forgetting. Through comprehensive empirical studies with several benchmark datasets, we demonstrate that our proposed method, Structural-Shift-Risk-Mitigation (SSRM), is flexible and easy to adapt to improve the performance of state-of-the-art GNN incremental learning frameworks in the inductive setting. Junwei Su, Difan Zou, Chuan Wu 0001 |
ICML | 1 |