EDBT 2026 Demo / reviewers in the wild / expert
Siyu Yang 0002
dblp:72/7121-2
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0000-0071-6218ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PctoDL: Adaptive GPU Throughput Optimization for Deep Learning Inference with Power ConstraintsabstractThe proliferation of deep learning inference services in power-constrained environments necessitates GPU management strategies that maximize throughput within strict power envelopes. Existing approaches often treat frequency scaling and resource partitioning as orthogonal problems or rely on static hardware assumptions, leading to suboptimal energy efficiency. This article presents PctoDL , a power-aware scheduling system that maximizes aggregate inference throughput by jointly optimizing spatial resource partitioning, batch size, and SM/memory frequency settings. To address the throughput–power tradeoff in power-constrained multi-tenant inference, PctoDL couples resource partitioning with coordinated frequency control under a fixed power cap. It combines a physics-informed iterative greedy partitioning algorithm, a thermodynamic model-predictive controller for runtime frequency regulation, and an online joint optimization mechanism for adaptive refinement. On the NVIDIA RTX 3080 Ti platform, PctoDL improves average throughput over BatchDVFS by 108.41%, with a peak gain of 262.74%. On the NVIDIA A100 platform, it delivers an average gain of 19.74% and a maximum gain of 57.03%. Compared with Morak’s coarse-grained partitioning approach, PctoDL achieves average/peak gains of 79.05%/137.93% on the RTX 3080 Ti and 26.33%/70.21% on the A100. Meng Hao 0002, Zikun Wu, Xueyang Tian, Siyu Yang 0002, Guotong Guo, Yiming Wang 0010, Farui Wang, Desheng Wang 0002, Weizhe Zhang |
ACM Trans. Archit. Code Optim. | 4 |
| 2026 | GreenDLS: An Energy-Efficient and SLO-Aware Deep Learning Serving SystemabstractThe growing demand for deploying deep learning (DL) models, particularly large language models (LLMs), has made it imperative to optimize GPU energy consumption while meeting service-level objectives (SLOs). Significant energy use and carbon dioxide (CO2) emissions from GPU-based inference tasks contribute substantially to the environmental footprint of the DL deployment. Existing approaches primarily rely on batching and dynamic voltage and frequency scaling (DVFS) to optimize service performance or throughput, but often overlook memory frequency adjustments and holistic energy optimization under dynamic workloads. This study presents GREENDLS, a DL serving system that integrates deep reinforcement learning (DRL) with offline prediction models to optimize energy consumption while adhering to inference latency SLOs, achieving significant energy savings. GREENDLS dynamically adjusts batch size, GPU streaming multiprocessor (SM) frequency, and GPU memory frequency based on inference request rates. It also accounts for GPU energy consumption during idle phases, such as batch filling, enabling multi-parameter and fine-grained energy optimization. Compared to the Clipper system, GREENDLS achieves energy savings of up to 45.57% on the RTX 3080Ti and 39.44% on the Tesla V100S. When compared to the EAIS system, which only combines batching with GPU SM frequency adjustment, GREENDLS achieves energy savings of up to 36.44% on the RTX 3080Ti and 15.44% on the Tesla V100S. Against the method proposed by Yu et al., GREENDLS attains energy savings of up to 46.34% on the RTX 3080Ti and 33.74% on the Tesla V100S. In LLM inference tasks using Qwen, GREENDLS reduces average energy consumption by 40.79% compared to Clipper, 10.42% compared to EAIS, and 42.42% compared to Yu et al. These results clearly demonstrate that GREENDLS more effectively optimizes energy consumption compared to traditional methods that rely primarily on batching or a combination of batching and DVFS, while still ensuring SLO compliance. Meng Hao 0002, Xueyang Tian, Siyu Yang 0002, Yiming Wang 0010, Desheng Wang 0002, Weizhe Zhang |
IEEE Trans. Computers | 3 |
| 2025 | Deep Learning Workload Mapping Optimization on Jetson PlatformsabstractTo improve the performance and energy efficiency of deep learning (DL) applications, recent edge computing platforms have built-in heterogeneous accelerators, such as general-purpose graphics processing units (GPUs) and neural processing units (NPUs). For example, widely used NVIDIA Jetson platforms contain CPU, GPU, and deep learning accelerator (DLA), a type of NPU. It is non-trivial to map DL workloads to suitable accelerators to improve performance, energy efficiency, or even both. This article presents JDIMO, 1 a Jetson-aware deep-learning inference workload mapping optimization framework, to simultaneously improve energy efficiency and performance. JDIMO first measures energy-performance data of the fundamental nodes and the sub-networks with energy-efficiency improvement potential according to the topology structure of a DL network. Then, under the guidance of an analytical energy-performance model, the framework exploits an algorithm based on the variable-length sliding window to find the optimal mapping configuration and the optimal number of CUDA streams. We evaluate JDIMO by applying it to seven DL applications on a Jetson Orin NX (16GB) platform. JDIMO saves 47.5% EDP (energy delay product) and 22.6% energy and improves 138.3% QPS (queries per second) on average compared to the DLA-possible configuration. JDIMO saves 22.5% EDP and 12.6% energy and improves 13.5% QPS on average compared to JEDI, the most similar work to ours. Meanwhile, JDIMO also reduces 93.8% optimization time on average compared to JEDI. Farui Wang, Meng Hao 0002, Siyu Yang 0002, Weizhe Zhang |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | Optimizing depthwise separable convolution on DCUabstractAbstract The integration of Large Language Models (LLMs) with Convolutional Neural Networks (CNNs) is significantly advancing the development of large models. However, the computational cost of large models is high, necessitating optimization for greater efficiency. One effective way to optimize the CNN is the use of depthwise separable convolution (DSC), which decouples spatial and channel convolutions to reduce the number of parameters and enhance efficiency. In this study, we focus on porting and optimizing DSC kernel functions from the GPU to the Deep Computing Unit (DCU), a computing accelerator developed in China. For depthwise convolution, we implement a row data reuse algorithm to minimize redundant data loading and memory access overhead. For pointwise convolution, we extend our dynamic tiling strategy to improve hardware utilization by balancing resource allocation among blocks and threads, and we enhance arithmetic intensity through a channel distribution algorithm. We implement depthwise and pointwise convolution kernel functions and integrate them into PyTorch as extension modules. Experiments demonstrate that our optimized kernel functions outperform the MIOpen library on the DCU, achieving up to a 3.59 $$\times$$ × speedup in depthwise convolution and up to a 3.54 $$\times$$ × speedup in pointwise convolution. These results highlight the effectiveness of our approach in leveraging the DCU’s architecture to accelerate deep learning operations. Meng Hao 0002, Weizhe Zhang, Gangzhao Lu, Xueyang Tian, Siyu Yang 0002, Mingdong Xie, Chenyu Yuan, Desheng Wang 0002 |
CCF Trans. High Perform. Comput. | 6 |