EDBT 2026 Demo / reviewers in the wild / expert
Shuangyan Yang
dblp:276/1037
· DBLP profile ↗
9ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered MemoryabstractDeep learning recommendation models (DLRMs) are widely used in industry, and their memory capacity requirements reach the terabyte scale. Tiered memory architectures provide a cost-effective solution but introduce challenges in embedding-vector placement due to complex embedding-access patterns. We propose RecMG, a machine learning (ML)-guided system for vector caching and prefetching on tiered memory. RecMG accurately predicts accesses to embedding vectors with long reuse distances or few reuses. The design of RecMG focuses on making ML feasible in the context of DLRM inference by addressing unique challenges in data labeling and navigating the search space for embedding-vector placement. By employing separate ML models for caching and prefetching, plus a novel differentiable loss function, RecMG narrows the prefetching search space and minimizes on-demand fetches. Compared to state-of-the-art temporal, spatial, and ML-based prefetchers, RecMG reduces on-demand fetches by $2.2 \times, 2.8 \times$, and $1.5 \times$, respectively. In industrial-scale DLRM inference scenarios, RecMG effectively reduces end-to-end DLRM inference time by up to 43%. Jie Ren 0015, Bin Ma 0025, Shuangyan Yang, Benjamin Francis, Ehsan K. Ardestani, Min Si, Dong Li 0001 |
HPCA | 3 |
| 2025 | Buffalo: Enabling Large-Scale GNN Training via Memory-Efficient BucketizationabstractGraph Neural Networks (GNNs) have demonstrated outstanding results in many graph-based deep-learning tasks. However, training GNNs on a large graph can be difficult due to memory capacity limitations. To address this problem, we can divide the graph into multiple partitions. However, this strategy faces a memory explosion problem. This problem stems from a long tail in the degree distribution of graph nodes. This strategy also suffers from time-consuming graph partitioning, difficulty in estimating the memory consumption of each partition, and time-consuming data preparation (e.g., block generation). To address the above problems, we introduce Buffalo, a GNN training system. Buffalo enables flexible mapping between the nodes and partitions to address the memory explosion problem, and enables fast graph partitioning based on node bucketing. Buffalo also introduces lightweight analytical modeling for memory estimation, and reduces block generation time by leveraging graph sampling. Evaluating large-scale real-world datasets (including billionscale datasets), we show that Buffalo effectively addresses the memory capacity limitation, enabling scalable GNN training and outperforming prior works in the compute-vs-memory efficiency Pareto frontier. With a limited memory budget, Buffalo achieves an end-to-end reduction of training time by 70.9% on average, compared to state-of-the-art (DGL [73], PyG [12], and Betty [93]). Shuangyan Yang, Minjia Zhang, Dong Li 0001 |
HPCA | 1 |
| 2025 | Performance Characterization of CXL Memory and Its Use CasesabstractCompute eXpress Link (CXL) is emerging as a promising memory interface technology. However, its performance characteristics remain largely unclear due to the limited availability of production hardware. Key questions include: What are the use cases for the CXL memory? What are the impacts of the CXL memory on application performance? How to use the CXL memory in combination with existing memory components? In this work, we study the performance of three genuine CXL memory-expansion cards from different vendors. We characterize the basic performance of the CXL memory, study how HPC applications and large language models (LLM) can benefit from the CXL memory, and study the interplay between memory tiering and page interleaving. We also propose a novel data object-level interleaving policy to match the interleaving policy with memory access patterns. Our findings reveal the challenges and opportunities of using the CXL memory. Xi Wang 0027, Jie Liu 0096, Shuangyan Yang, Jie Ren 0015, Bhanu Shankar, Dong Li 0001 |
IPDPS | 4 |
| 2025 | Three-Dimensional Forward Modeling for Grounded-Source Semi-Airborne Transient Electromagnetic Method With IP Effect Directly in Time Domain Based on SOE ApproximationabstractIt is very useful for the exploration of metallic sulfide deposits to simulate the induced polarization (IP) effect in semi-airborne transient electromagnetic (TEM) data and analyze its response characteristics. To improve the computational efficiency and the ability to handle the complex models, we develop a fast 3D forward modeling algorithm for semi-airborne TEM method with IP effect directly in time domain based on the sum-of-exponentials (SOE) approximation and the time-domain unstructured vector finite-element method. By adopting the SOE method to approximate the kernel function of Caputo fractional derivative, the convolution operation in the time direction is transformed into a piecewise analytical solution and a recursive relationship. This approach resolves the problem of huge storage and calculation caused by the traditional L1 approximation relying on all historical information, thereby improving the efficiency of 3D forward modeling. Meanwhile, the flexibility of the unstructured finite-element method provides our algorithm with the ability to deal with the models with undulating terrain and complex-shaped bodies. The numerical accuracy and efficiency of our method are verified by comparing it with the 1D analytical solution on a polarized half space and the L1 approximation respectively. On this basis, we analyze the influence of transmitting-waveform parameters on the IP responses on a polarizable ellipsoid model. Finally, our algorithm is applied to a polarizable model with topography to check its ability to handle the complex model. We further analyze the influence of the observation mode and the direction of the transmitter sources on the IP responses. Shuangyan Yang, Yanfu Qi, Chenchen Shu, Zetong Wang, Peihang Xie |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Enabling Large Dynamic Neural Network Training with Learning-based Memory ManagementabstractDynamic neural network (DyNN) enables high computational efficiency and strong representation capability. However, training DyNN can face a memory capacity problem because of increasing model size or limited GPU memory capacity. Managing tensors to save GPU memory is challenging, because of the dynamic structure of DyNN. We present DyNN-Offload, a memory management system to train DyNN. DyNN-Offload uses a learned approach (using a neural network called the pilot model) to increase predictability of tensor accesses to facilitate memory management. The key of DyNN-Offload is to enable fast inference of the pilot model in order to reduce its performance overhead, while providing high inference (or prediction) accuracy. DyNNOffload reduces input feature space and model complexity of the pilot model based on a new representation of DyNN; DyNNOffload converts the hard problem of making prediction for individual operators into a simpler problem of making prediction for a group of operators in DyNN. DyNN-Offload enables 8 × larger DyNN training on a single GPU compared with using PyTorch alone (unprecedented with any existing solution). Evaluating with AlphaFold (a production-level, large-scale DyNN), we show that DyNN-Offload outperforms unified virtual memory (UVM) and dynamic tensor rematerialization (DTR), the most advanced solutions to save GPU memory for DyNN, by 3 × and 2.1 × respectively in terms of maximum batch size. Jie Ren 0015, Dong Xu 0024, Shuangyan Yang, Christian Navasca, Chenxi Wang 0005, Guoqing Harry Xu, Dong Li 0001 |
HPCA | 3 |
| 2024 | 3-D Forward Modeling of Semi-Airborne TEM Method With IP Effect Based on Goal-Oriented Adaptive Finite Element AlgorithmabstractThe coupling of induced polarization (IP) effects and electromagnetic induction significantly complicates the transient electromagnetic (TEM) diffusion, leading to serious distortion of the observed responses. In this article, we develop a 3-D forward modeling algorithm based on a goal-oriented adaptive finite-element method for the semi-airborne TEM method with IP effects to obtain high-precision simulation results. First, the Caputo fractional derivative is discretized by piecewise linear interpolation, and the Cole–Cole model is wholly introduced into the time-domain finite-element governing equation. Then, we simulate the semi-airborne TEM responses with IP effect directly in time domain by solving this governing equation. Because the Caputo fractional derivative depends on all historical information, it leads to a large storage and calculation. In order to solve this problem, we adopt the piecewise local time step discretization strategy to reduce the storage occupation and improve the computational efficiency. Finally, combined with the goal-oriented adaptive grid refinement technology, the forward modeling mesh is automatically optimized based on the weighted hybrid posterior error estimations. The accuracy of our codes is verified by comparing them with 1-D analytical solutions. We further simulate the electromagnetic responses of the polarizable body model with complex shape and topography to analyze the influence of the posterior error estimation method and the IP effects on the adaptive mesh. Chenchen Shu, Yanfu Qi, Shuangyan Yang, Jianmei Zhou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Betty: Enabling Large-Scale GNN Training with Batch-Level Graph PartitioningabstractThe Graph Neural Network (GNN) is showing outstanding results in improving the performance of graph-based applications. Recent studies demonstrate that GNN performance can be boosted via using more advanced aggregators, deeper aggregation depth, larger sampling rate, etc. While leading to promising results, the improvements come at a cost of significantly increased memory footprint, easily exceeding GPU memory capacity. In this paper, we introduce a method, Betty, to make GNN training more scalable and accessible via batch-level partitioning. Different from DNN training, a mini-batch in GNN has complex dependencies between input features and output labels, making batch-level partitioning difficult. Betty introduces two noveltechniques, redundancy-embedded graph (REG) partitioning and memory-aware partitioning, to effectively mitigate the redundancy and load imbalances issues across the partitions. Our evaluation of large-scale real-world datasets shows that Betty can significantly mitigate the memory bottleneck, enabling scalable GNN training with much deeper aggregation depths, larger sampling rate, larger training batch sizes, together with more advanced aggregators, with a few as a single GPU. Shuangyan Yang, Minjia Zhang, Wenqian Dong, Dong Li 0001 |
ASPLOS (2) | 1 |
| 2021 | ZeRO-Offload: Democratizing Billion-Scale Model Training
Jie Ren 0015, Samyam Rajbhandari, Reza Yazdani, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li 0001, Yuxiong He |
USENIX ATC | 5 |
| 2020 | BionoiNet: ligand-binding site classification with off-the-shelf deep neural networkabstractMOTIVATION: Fast and accurate classification of ligand-binding sites in proteins with respect to the class of binding molecules is invaluable not only to the automatic functional annotation of large datasets of protein structures but also to projects in protein evolution, protein engineering and drug development. Deep learning techniques, which have already been successfully applied to address challenging problems across various fields, are inherently suitable to classify ligand-binding pockets. Our goal is to demonstrate that off-the-shelf deep learning models can be employed with minimum development effort to recognize nucleotide- and heme-binding sites with a comparable accuracy to highly specialized, voxel-based methods. RESULTS: We developed BionoiNet, a new deep learning-based framework implementing a popular ResNet model for image classification. BionoiNet first transforms the molecular structures of ligand-binding sites to 2D Voronoi diagrams, which are then used as the input to a pretrained convolutional neural network classifier. The ResNet model generalizes well to unseen data achieving the accuracy of 85.6% for nucleotide- and 91.3% for heme-binding pockets. BionoiNet also computes significance scores of pocket atoms, called BionoiScores, to provide meaningful insights into their interactions with ligand molecules. BionoiNet is a lightweight alternative to computationally expensive 3D architectures. AVAILABILITY AND IMPLEMENTATION: BionoiNet is implemented in Python with the source code freely available at: https://github.com/CSBG-LSU/BionoiNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jeffrey Mitchell Lemoine, Abd-El-Monsif A. Shawky, Manali Singha, Limeng Pu, Shuangyan Yang, J. Ramanujam, Michal Brylinski |
Bioinform. | 6 |