Yiming Wang 0010

dblp:71/3182-10 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-3298-5134ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PctoDL: Adaptive GPU Throughput Optimization for Deep Learning Inference with Power Constraints
abstract
The proliferation of deep learning inference services in power-constrained environments necessitates GPU management strategies that maximize throughput within strict power envelopes. Existing approaches often treat frequency scaling and resource partitioning as orthogonal problems or rely on static hardware assumptions, leading to suboptimal energy efficiency. This article presents PctoDL , a power-aware scheduling system that maximizes aggregate inference throughput by jointly optimizing spatial resource partitioning, batch size, and SM/memory frequency settings. To address the throughput–power tradeoff in power-constrained multi-tenant inference, PctoDL couples resource partitioning with coordinated frequency control under a fixed power cap. It combines a physics-informed iterative greedy partitioning algorithm, a thermodynamic model-predictive controller for runtime frequency regulation, and an online joint optimization mechanism for adaptive refinement. On the NVIDIA RTX 3080 Ti platform, PctoDL improves average throughput over BatchDVFS by 108.41%, with a peak gain of 262.74%. On the NVIDIA A100 platform, it delivers an average gain of 19.74% and a maximum gain of 57.03%. Compared with Morak’s coarse-grained partitioning approach, PctoDL achieves average/peak gains of 79.05%/137.93% on the RTX 3080 Ti and 26.33%/70.21% on the A100.
Meng Hao 0002, Zikun Wu, Xueyang Tian, Siyu Yang 0002, Guotong Guo, Yiming Wang 0010, Farui Wang, Desheng Wang 0002, Weizhe Zhang
ACM Trans. Archit. Code Optim.8
2026 GreenDLS: An Energy-Efficient and SLO-Aware Deep Learning Serving System
abstract
The growing demand for deploying deep learning (DL) models, particularly large language models (LLMs), has made it imperative to optimize GPU energy consumption while meeting service-level objectives (SLOs). Significant energy use and carbon dioxide (CO2) emissions from GPU-based inference tasks contribute substantially to the environmental footprint of the DL deployment. Existing approaches primarily rely on batching and dynamic voltage and frequency scaling (DVFS) to optimize service performance or throughput, but often overlook memory frequency adjustments and holistic energy optimization under dynamic workloads. This study presents GREENDLS, a DL serving system that integrates deep reinforcement learning (DRL) with offline prediction models to optimize energy consumption while adhering to inference latency SLOs, achieving significant energy savings. GREENDLS dynamically adjusts batch size, GPU streaming multiprocessor (SM) frequency, and GPU memory frequency based on inference request rates. It also accounts for GPU energy consumption during idle phases, such as batch filling, enabling multi-parameter and fine-grained energy optimization. Compared to the Clipper system, GREENDLS achieves energy savings of up to 45.57% on the RTX 3080Ti and 39.44% on the Tesla V100S. When compared to the EAIS system, which only combines batching with GPU SM frequency adjustment, GREENDLS achieves energy savings of up to 36.44% on the RTX 3080Ti and 15.44% on the Tesla V100S. Against the method proposed by Yu et al., GREENDLS attains energy savings of up to 46.34% on the RTX 3080Ti and 33.74% on the Tesla V100S. In LLM inference tasks using Qwen, GREENDLS reduces average energy consumption by 40.79% compared to Clipper, 10.42% compared to EAIS, and 42.42% compared to Yu et al. These results clearly demonstrate that GREENDLS more effectively optimizes energy consumption compared to traditional methods that rely primarily on batching or a combination of batching and DVFS, while still ensuring SLO compliance.
Meng Hao 0002, Xueyang Tian, Siyu Yang 0002, Yiming Wang 0010, Desheng Wang 0002, Weizhe Zhang
IEEE Trans. Computers7
2025 STAMImputer: Spatio-Temporal Attention MoE for Traffic Data Imputation
abstract
Traffic data imputation is fundamentally important to support various applications in intelligent transportation systems such as traffic flow prediction. However, existing time-to-space sequential methods often fail to effectively extract features in block-wise missing data scenarios. Meanwhile, the static graph structure for spatial feature propagation significantly constrains the model's flexibility in handling the distribution shift issue for the nonstationary traffic data. To address these issues, this paper proposes a Spatio-Temporal Attention Mixture of experts network named STAMImputer for traffic data imputation. Specifically, we introduce a Mixture of Experts (MoE) framework to capture latent spatio-temporal features and their influence weights, effectively imputing block missing. A novel Low-rank guided Sampling Graph ATtention (LrSGAT) mechanism is designed to dynamically balance the local and global correlations across road networks. The sampled attention vectors are utilized to generate dynamic graphs that capture real-time spatial correlations. Extensive experiments are conducted on four traffic datasets for evaluation. The result shows STAMImputer achieves significantly performance improvement compared with existing SOTA approaches. Our codes are available at https://github.com/RingBDStack/STAMImupter.
Yiming Wang 0010, Hao Peng 0001, Senzhang Wang, Haohua Du, Jia Wu 0001, Guanlin Wu
IJCAI1
2025 Privacy-enhanced data distillation with probability distribution matching
Ke Pan 0001, Yuxin Wen, Yiming Wang 0010, Maoguo Gong, Hui Li 0006, Shanfeng Wang
Neurocomputing3
2025 Dynamic Power Management Through Multi-agent Deep Reinforcement Learning for Heterogeneous Systems
abstract
Power management and optimization play a significant role in modern computer systems, from battery-powered devices to servers running in data centers. Existing approaches for power capping fail to meet the requirements presented by dynamic workloads, and the situation becomes even more severe, given the divergent energy efficiency of workloads on heterogeneous hardware platforms. Adaptively optimizing energy consumption for dynamic workloads presents a great challenge to heterogeneous systems. To tackle this challenge, we present a machine learning based method to improve system-level power efficiency. We employ multi-agent deep reinforcement learning (MADRL) to automatically explore the relationship between long-term performance and the power budget for workloads of different types on classic CPU-GPU heterogeneous platforms. Our framework equips each device with an agent, enabling decentralized control over its power budget while maintaining centralized coordination to maximize the running time of applications within a power cap. We evaluate our approach against state-of-the-art methods on CPU-GPU platforms. Experimental results show that our method improves performance by an average of 8.5%. Additionally, our method is significantly more stable compared to the state-of-the-art heuristic approach.
Yiming Wang 0010, Weizhe Zhang, Meng Hao 0002, Weizhi Kong, Yuan Wen
ACM Trans. Archit. Code Optim.1
2024 DRLCAP: Runtime GPU Frequency Capping With Deep Reinforcement Learning
abstract
Power and energy consumption is the limiting factor of modern computing systems. As the GPU becomes a mainstream computing device, power management for GPUs becomes increasingly important. Current works focus on GPU kernel-level power management, with challenges in portability due to architecture-specific considerations. We presentDRLCap, a general runtime power management framework intended to support power management across various GPU architectures. It periodically monitors system-level information to dynamically detect program phase changes and model the workload and GPU system behavior. This elimination from kernel-specific constraints enhances adaptability and responsiveness. The framework leverages dynamic GPU frequency capping, which is the most widely used power knob, to control the power consumption.DRLCapemploys deep reinforcement learning (DRL) to adapt to the changing of program phases by automatically adjusting its power policy through online learning, aiming to reduce the GPU power consumption without significantly compromising the application performance. We evaluateDRLCapon three NVIDIA and one AMD GPU architectures. Experimental results show thatDRLCapimproves prior GPU power optimization strategies by a large margin. On average, it reduces the GPU energy consumption by 22% with less than 3% performance slowdown on NVIDIA GPUs. This translates to a 20% improvement in the energy efficiency measured by the energy-delay product (EDP) over the NVIDIA default GPU power management strategy. For the AMD GPU architecture,DRLCapsaves energy consumption by 10%, on average, with a 4% percentage loss, and improves energy efficiency by 8%.
Yiming Wang 0010, Meng Hao 0002, Weizhe Zhang, Qiuyuan Tang, Zheng Wang 0001
IEEE Trans. Sustain. Comput.1
2022 HiGIL: Hierarchical Graph Inference Learning for Fact Checking
abstract
Fact-checking is vital for countering fake news. This process requires verifying the truthfulness of a claim by reasoning about multiple pieces of evidence. The current dominant approach depends upon capturing the claim-evidence relations from a claim-evidence interaction graph. Existing solutions utilize phrase-level semantics on a single-granularity but ignore other hierarchical features, such as fact- and sentence-level textual semantics and their logical topology. Since the hierarchical features often provide hints to infer collaborative high-order clues that can be essential for fact-checking, they should not be overlooked. This paper proposes a better method to model the claim-evidence graph in a multi-granularity manner. Doing so allows one to exploit more textual semantics and logical topology between a claim and its evidence. To achieve the target, we first employ a graph inference learning framework to infer graph nodes on different granular semantic units within their hierarchical topology. Then, an inference learning procedure is designed to optimize the global textual similarity and local topological reachability from the claim-evidence graph. We evaluate our approach by applying it to fact-checking on an open dataset, and experimental results show that our technique outperforms existing graph-based techniques by a large margin.
Qianren Mao, Yiming Wang 0010, Linfeng Du, Hao Peng 0001, Jia Wu 0001, Jianxin Li 0002, Zheng Wang 0001
ICDM2
2022 A new two-stage based evolutionary algorithm for solving multi-objective optimization problems
Yiming Wang 0010, Weifeng Gao, Maoguo Gong, Hong Li 0007, Jin Xie 0003
Inf. Sci.1
2022 Online Power Management for Multi-Cores: A Reinforcement Learning Based Approach
abstract
Power and energy is the first-class design constraint for multi-core processors and is a limiting factor for future-generation supercomputers. While modern processor design provides a wide range of mechanisms for power and energy optimization, it remains unclear how software can make the best use of them. This article presents a novel approach for runtime power optimization on modern multi-core systems. Our policy combines power capping and uncore frequency scaling to match the hardware power profile to the dynamically changing program behavior at runtime. We achieve this by employing reinforcement learning (RL) to automatically explore the energy-performance optimization space from training programs, learning the subtle relationships between the hardware power profile, the program characteristics, power consumption and program running times. Our RL framework then uses the learned knowledge to adapt the chip's power budget and uncore frequency to match the changing program phases for any new, previously unseen program. We evaluate our approach on two computing clusters by applying our techniques to 11 parallel programs that were not seen by our RL framework at the training stage. Experimental results show that our approach can reduce the system-level energy consumption by 12 percent, on average, with less than 3 percent of slowdown on the application performance. By lowering the uncore frequency to leave more energy budget to allow the processor cores to run at a higher frequency, our approach can reduce the energy consumption by up to 17 percent while improving the application performance by 5 percent for specific workloads.
Yiming Wang 0010, Weizhe Zhang, Meng Hao 0002, Zheng Wang 0001
IEEE Trans. Parallel Distributed Syst.1
2021 Fine-Grained Powercap Allocation for Power-Constrained Systems Based on Multi-Objective Machine Learning
abstract
Power capping is an important solution to keep the system within a fixed power constraint. However, for the over-provisioned and power-constrained systems, especially the future exascale supercomputers, powercap needs to be reasonably allocated according to the workloads of compute nodes to achieve trade-offs among performance, energy and powercap. Thus it is necessary to model performance and energy and to predict the optimal powercap allocation strategies. Existing power allocation approaches have insufficient granularity within nodes. Modeling approaches usually model performance and energy separately, ignoring the correlation between objectives, and do not expose the Pareto-optimal powercap configurations. Therefore, this article combines the powercap with uncore frequency scaling and proposes an approach to predict the Pareto-optimal powercap configurations on the power-constrained system for input MPI and OpenMP parallel applications. Our approach first uses the elaborately designed micro-benchmarks and a small number of existing benchmarks to build the training set, and then applies a multi-objective machine learning algorithm which combines the stacked single-target method with extreme gradient boosting to build multi-objective models of performance and energy. The models can be used to predict the optimal processor and memory powercap settings, helping compute nodes perform fine-grained powercap allocation. When the optimal powercap configuration is determined, the uncore frequency scaling is used to further optimize the energy consumption. Compared with the reference powercap configuration, the predicted optimal configurations predicted by our method can achieve an average powercap reduction of 31.35 percent, an average energy reduction of 12.32 percent, and average performance degradation of only 2.43 percent.
Meng Hao 0002, Weizhe Zhang, Yiming Wang 0010, Gangzhao Lu, Farui Wang, Athanasios V. Vasilakos
IEEE Trans. Parallel Distributed Syst.3