EDBT 2026 Demo / reviewers in the wild / expert
Jie Sun 0017
dblp:54/5330-17
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-7030-0146ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bat: Efficient Generative Recommender Serving with Bipartite AttentionabstractGenerative Recommenders (GRs) have recently emerged as promising alternatives to traditional Deep Learning Recommendation Models (DLRMs). Despite their potential, GRs remain computationally expensive in inference, exhibiting compute-bound characteristics similar to the prefill stage of Large Language Model (LLM) inference. Prefix caching can reduce redundant computation by reusing previously constructed KV caches. However, the unique properties of GRs, i.e., highly personalized user profiles and real-time item retrieval, make cache reuse across queries challenging, resulting in limited computational savings. Jie Sun 0017, Shaohang Wang, Zimo Zhang, Peng Sun 0006, Bo Zhao 0019, Bingsheng He, Fei Wu 0001, Zeke Wang |
ASPLOS (2) | 1 |
| 2025 | CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage AccessabstractWith the wide adoption of GPU and the explosion in data volumes, existing accelerator-centric systems require massive storage access. They adopt high-performance storage devices like NVMe SSDs to scale up single-node systems cost-effectively and leverage the CPU to manage these SSDs. However, they suffer from performance bottlenecks because of the high CPU OS kernel overhead and the CPU memory intermediated data transfer. To address this issue, GPU-initiated and GPU-managed SSD management is proposed to allow the GPU to fully manipulate SSDs: 1) direct data transfer from SSD to GPU memory (data plane) and 2) GPU-managed SSD control (control plane). This can potentially enable these GPU systems to fully leverage the SSD bandwidth. However, we still identify two severe issues. First, the GPU-management SSD control leads to low GPU Streaming Multiprocessor utilization. Second, it leads to the serial execution of SSD accesses with GPU computation, which slows down the overall computing task. To this end, we propose CAM, the first asynchronous GPU-initialized, CPU-managed SSD management for batching storage access. It 1) offloads the SSD control plane from GPU to CPU, thus maximizing GPU streaming multiprocessor utilization, and 2) adopts asynchronous user-friendly APIs that allow programmers to easily overlap GPU computation and SSD I/O operations while keeping a synchronous programming experience. As such, CAM enables us to achieve the best of two worlds: high performance and high programmability. The experimental results show that CAM can perform GNN model training, mergesort, and GEMM up to$\mathbf{1.84}\times, \mathbf{1.5}\times$, and$\mathbf{1.84}\times \mathbf{faster}$compared to the existing state-of-the-art GPU systems, while keeping high programmability. Ziyu Song, Jie Zhang 0081, Jie Sun 0017, Mo Sun 0001, Zihan Yang 0004, Xuzheng Chen, Fei Wu 0001, Huajin Tang, Zeke Wang |
ICDE | 3 |
| 2025 | Hyperion: Co-Optimizing SSD Access and GPU Computation for Cost-Efficient GNN TrainingabstractSSDs are traditionally regarded as a cheap but slow way to scale up GNN training. Several GNN systems explore cheap single-machine single-GPU out-of-core training but fall short in terms of TPC (throughput per monetary cost). The underlying reason is that the existing systems 1) overly focus on minimizing the number of SSD accesses, which results in substantial unnecessary overhead on the CPU side, or 2) exhaust all GPU parallelism to saturate SSD but fail to overlap SSD accesses with GNN computation. In this work, we present Hyperion, a cost-efficient system for terabyte-scale GNN training. We argue that co-optimizing GPU-initiated asynchronous SSD access and GNN computation pipeline enables us to only add cheap NVMe SSDs, rather than expensive GPU servers, to achieve in- memory-like throughput and thus maximal TPC of GNN training. However, this is non-trivial due to imbalanced workloads and interference among IO submission, IO completion, and cache lookup. To tackle the challenges, Hyperion proposes three key designs. First, Hyperion proposes the first GPU-initiated pipeline- friendly asynchronous disk IO stack, which only requires about 1% GPU cores to saturate SSD throughput and wastes no GPU cores between IO submission and completion to fully overlap disk IO and computation. Second, we propose a new GPU-managed, disaggregated, and unified cache that disaggregates cache lookup from disk IO and fully utilizes CPU/GPU memory hierarchy by a unified static cache policy. Third, we propose a GNN-aware general TPC-analytical model that precisely predicts TPC under diverse hardware settings and GNN models and provide a hint to guide users to select hardware, e.g., number of SSDs, under a limited budget to maximize TPC. Experiments demonstrate that Hyperion can improve the TPC by over 3.1x on terabyte-scale graphs compared to SOTA out-of-core baselines and improve 60 x TPC compared to distributed in-memory baselines. Jie Sun 0017, Mo Sun 0001, Zuocheng Shi, Zihan Yang 0004, Jie Zhang 0081, Zeke Wang, Fei Wu 0001 |
ICDE | 1 |
| 2025 | Helios: Efficient Distributed Dynamic Graph Sampling for Online GNN InferenceabstractOnline GNN inference has been widely explored by applications such as online recommendation and financial fraud detection systems, where even minor delays can result in significant financial impact. Real-time dynamic graph sampling enables online GNN inference to reflect the latest graph updates in real-world graphs. However, online GNN inference typically demands millisecond-level latency Service Level Objectives (SLOs) as its performance guarantees, which poses great challenges for existing dynamic graph sampling approaches based on graph databases. The issues mainly arise from two aspects: long tail latency due to imbalanced data-dependent sampling and large communication overhead incurred by distributed sampling. To address these issues, we propose Helios, an efficient distributed dynamic graph sampling service to meet the stringent latency SLOs. The key ideas of Helios are 1) pre-sampling the dynamic graph in an event-driven approach, and 2) maintaining a query-aware sample cache to build the complete K-hop sampling results locally for inference requests. Experiments on multiple datasets show that Helios achieves up to 67× higher serving throughput and up to 32× lower P99 query latency compared to baselines. Jie Sun 0017, Zuocheng Shi, Li Su 0005, Wenting Shen, Zeke Wang, Yong Li 0045, Wenyuan Yu, Wei Lin 0016, Fei Wu 0001, Bingsheng He, Jingren Zhou 0001 |
PPoPP | 1 |
| 2025 | Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN TrainingabstractGraph Neural Networks (GNNs) are widely employed in applications like recommendation systems, social network analysis, and fraud detection, but training large-scale GNNs is challenging due to its memory limitations. Existing systems face a trade-off between throughput and monetary cost: Distributed systems require expensive memory scaling, while single-machine out-of-core systems are limited by GPU/PCIe throughput. To this end, we propose Moment, a physical communication topology and data placement co-optimizer to enable high-throughput and low-cost GNN training in a single multi-GPU machine. Moment addresses communication contention and GPU load imbalance issues by modeling the physical topology as capacity-constrained directed graphs and formulating communication scheduling as a max-flow problem. It also introduces a data-distribution-aware knapsack algorithm for optimized data placement. Experimental results show that Moment outperforms out-of-core systems by up to 6.51 × and distributed systems by up to 3.02 ×, with only 50% monetary cost. Zuocheng Shi, Jie Sun 0017, Ziyu Song, Mo Sun 0001, Fei Wu 0001, Zeke Wang |
SC | 2 |
| 2024 | TorchGT: A Holistic System for Large-Scale Graph Transformer TrainingabstractGraph Transformer is a new architecture that surpasses GNNs in graph learning. While there emerge inspiring algorithm advancements, their practical adoption is still limited, particularly on real-world graphs involving up to millions of nodes. We observe existing graph transformers fail on large-scale graphs mainly due to heavy computation, limited scalability and inferior model quality. Motivated by these observations, we propose TORCHGT, the first efficient, scalable, and accurate graph transformer training system. TORCHGT optimizes training at three different levels. At algorithm level, by harnessing the graph sparsity, TORCHGT introduces a Dual-interleaved Attention which is computation-efficient and accuracy-maintained. At runtime level, TORCHGT scales training across workers with a communicationlight Cluster-aware Graph Parallelism. At kernel level, an Elastic Computation Reformation further optimizes the computation by reducing memory access latency in a dynamic way. Extensive experiments demonstrate that TORCHGT boosts training by up to 62.7× and supports graph sequence lengths of up to 1M. Meng Zhang 0045, Jie Sun 0017, Qinghao Hu 0004, Peng Sun 0006, Zeke Wang, Yonggang Wen 0001, Tianwei Zhang 0004 |
SC | 2 |
| 2024 | SparseACC: A Generalized Linear Model Accelerator for Sparse DatasetsabstractStochastic gradient descent (SGD) is widely used for training generalized linear models (GLMs), such as support vector machine and logistic regression, on large industry datasets. Such a training consumes plenty of computing power and therefore plenty of accelerators are proposed to accelerate the GLM training. However, real-world datasets are always highly sparse. For example, YouTube’s social network connectivity contains only 2.31% nonzero elements (NZs). It is not trivial to design an accelerator that is able to efficiently train on a sparse dataset that is stored in a compressed sparse format (e.g., compressed sparse row (CSR) format). The design of such an accelerator faces three challenges: 1) bank conflicts, which may happen when multiple processing engines in the accelerator access multiple memory banks; 2) complex interconnections, which are necessary to allow all processing engines to access any memory bank; and 3) high-synchronization overhead, since each sample in sparse dataset has a different number of NZs and these elements have different distributions, thus it is hard to overlap gradient computation and model update of neighboring batches. To this end, we propose SparseACC, a sparsity-aware accelerator for training generalized linear models (GLMs). SparseACC is based on two key mechanisms. First, a software/hardware co-design approach solves the first two design challenges by proposing a novel bank-conflict-free (BCF) and bank-balanced CSR format. Second, a weight-aware ping-pong model solves the third challenge, thus maximizing the utilization of the processing engines. SparseACC leverages these two mechanisms to orchestrate training over sparse datasets, such that the training time decreases linearly with the sparsity of the dataset. We prototype SparseACC on a Xilinx Alveo U280 FPGA (Xilinx, 2020). The experimental evaluation shows that SparseACC converges up to$3.5\times $,$18\times $,$38\times $, and$110\times $faster than the state-of-the-art counterparts on a sparse accelerator, a Tesla V100 GPU, an Intel i9-10900k CPU, and a dense accelerator, respectively. Jie Zhang 0081, Hongjing Huang, Jie Sun 0017, Juan Gómez-Luna, Onur Mutlu, Zeke Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | Staleness-Reduction Mini-Batch K-MeansabstractK -means (km) is a clustering algorithm that has been widely adopted due to its simple implementation and high clustering quality. However, the standard km suffers from high computational complexity and is therefore time-consuming. Accordingly, the mini-batch (mbatch) km is proposed to significantly reduce computational costs in a manner that updates centroids after performing distance computations on just a mbatch, rather than a full batch, of samples. Even though the mbatch km converges faster, it leads to a decrease in convergence quality because it introduces staleness during iterations. To this end, in this article, we propose the staleness-reduction mbatch (srmbatch) km, which achieves the best of two worlds: low computational costs like the mbatch km and high clustering quality like the standard km. Moreover, srmbatch still exposes massive parallelism to be efficiently implemented on multicore CPUs and many-core GPUs. The experimental results show that srmbatch can converge up to 40× - 130× faster than mbatch when reaching the same target loss, and srmbatch is able to reach 0.2%-1.7% lower final loss than that of mbatch. Xueying Zhu, Jie Sun 0017, Zhenhao He, Jiantong Jiang, Zeke Wang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN Training
Jie Sun 0017, Li Su 0005, Zuocheng Shi, Wenting Shen, Zeke Wang, Lei Wang 0004, Jie Zhang 0081, Yong Li 0020, Wenyuan Yu, Jingren Zhou 0001, Fei Wu 0001 |
USENIX ATC | 1 |
| 2023 | P4SGD: Programmable Switch Enhanced Model-Parallel Training on Generalized Linear Models on Distributed FPGAsabstractGeneralized linear models (GLMs) are a widely utilized family of machine learning models in real-world applications. As data size increases, it is essential to perform efficient distributed training for these models. However, existing systems for distributed training have a high cost for communication and often use large batch sizes to balance computation and communication, which negatively affects convergence. Therefore, we argue for an efficient distributed GLM training system that strives to achieve linear scalability, while keeping batch size reasonably low. As a start, we propose P4SGD, a distributed heterogeneous training system that efficiently trains GLMs through model parallelism between distributed FPGAs and through forward-communication-backward pipeline parallelism within an FPGA. Moreover, we propose a light-weight, latency-centric in-switch aggregation protocol to minimize the latency of the AllReduce operation between distributed FPGAs, powered by a programmable switch. As such, to our knowledge, P4SGD is the first solution that achieves almost linear scalability between distributed accelerators through model parallelism. We implement P4SGD on eight Xilinx U280 FPGAs and a Tofino P4 switch. Our experiments show P4SGD converges up to 6.5X faster than the state-of-the-art GPU counterpart. Hongjing Huang, Yingtao Li 0001, Jie Sun 0017, Xueying Zhu, Jie Zhang 0081, Jialin Li 0001, Zeke Wang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | PASER: A Pattern-Based Approach to Service Requirements AnalysisabstractInconsistent specification are an inevitable intermediate product of a service requirements engineering process. In order to reduce requirements inconsistencies, we propose PASER, a Pattern-based Approach to Service Requirements analysis. The PASER approach first extracts the process information from service documents via natural language processing (NLP) techniques, then uses a requirements modeling language – Workflow-Patterns-based Process Language (WPPL) — to build the process model. Finally, through matching with workflow patterns, the inconsistencies in service requirements are identified and resolved by checking against a set of checking rules. We have conducted a preliminary experiment to evaluate it. An ATM service case study is presented as a running example to illustrate our approach. Ye Wang 0012, Jie Sun 0017 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2013 | Representing and Elaborating Quality Requirements: The QRA Approach
Jie Sun 0017, Peri Loucopoulos, Liping Zhao 0001 |
ER | 1 |
| 2013 | Deriving problem frames from business process and object analysis modelsabstractAbstract While Problem Frames have become a useful approach for requirements analysis, little research has been made to explore how to derive them from a complex problem context. The purpose of this paper is to propose such an approach. The proposed approach consists of three steps to drive the development of Problem Frames. In the first step, business process models are developed to capture the behavioural view of the problem context. In the second step, object analysis models are used to capture the structural view of the problem context. Together, these two views collectively and adequately capture the early context knowledge. These two types of model will then be used in the third step to construct context diagrams and derive Problem Frames. A complex real‐world problem – equity trading problem – is used to illustrate this approach. Xinyu Wang 0001, Jie Sun 0017, Xiaohu Yang 0001, Ye Wang 0012, Shanping Li, Aleksander J. Kavs |
Expert Syst. J. Knowl. Eng. | 2 |
| 2012 | Satisfying quality requirements in the design of a partition-based, distributed stock trading systemabstractSUMMARY Although quality requirements (QRs) have become a major drive in today's software development, there have been very few real‐world examples in the literature that demonstrate how to meet these requirements. This paper presents such an example. Specifically, the paper describes the design of a partition‐based distributed stock trading service system that satisfies a set of QRs related to resource utilization, performance, scalability and availability. The paper evaluates this design through detailed experiments and discusses some design alternatives and the lessons learned. Central to this design are a static load distribution strategy and a dynamic load balancing strategy. The first strategy is to achieve an initial balanced workload on the system's server cluster during the system initialization time, whereas the second strategy is to maintain this balanced workload throughout the system execution time. Together, these two strategies work in unison to ensure that the server resources are efficiently utilized; the user requests are processed with the required speed; the application is partitioned with sufficient room to scale; and the system is highly available. Copyright © 2011 John Wiley & Sons, Ltd. Xiaohu Yang 0001, Liping Zhao 0001, Xinyu Wang 0001, Ye Wang 0012, Jie Sun 0017, Albert Jerry Cristoforo |
Softw. Pract. Exp. | 5 |