VLDB 2026 Research / reviewers in the wild / expert
Xiao Li 0038
dblp:66/2069-38
· DBLP profile ↗
8ranked-venue papers
1as first author
7since 2021 · last 2024
0000-0002-9728-3267ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PC-oriented Prediction-based Runtime Power Management for GPGPU using Knowledge TransferabstractAs Moore's law slows down, computing systems must prioritize higher energy efficiency to sustain performance scaling. GPUs have emerged as the primary workhorses of computing resources, making the achievement of high energy efficiency in GPUs a critical concern. However, implementing runtime power management on GPUs poses significant challenges due to the high variations and complexities arising from workloads and hardware configurations, which render offline optimization and reactive-based methods less effective. In this paper, we present a program counter (PC)-oriented prediction-based power management approach for GPGPUs. Our approach leverages the benefits of prediction to address online variations while enhancing prediction capability through knowledge transfer across different levels of resources. Experiments conducted on realistic applications demonstrate that our proposed method achieves the maximum energy savings under a user-defined performance constraint compared to state-of-the-art designs. Lin Chen 0029, Xiao Li 0038, Shixi Chen, Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
SPAA | 2 |
| 2024 | Deep Reinforcement Learning-Based Power Management for Chiplet-Based Multicore SystemsabstractChiplet technology has emerged as a promising solution to address the increasing demand for high-performance computing in light of the slowdown of Moore’s law. While chiplet-based multicore systems offer higher performance through heterogeneous integration, they also pose challenges for power delivery system (PDS) design. The integration of additional vertical and inter-chiplet connections, along with higher power density, impose stringent requirements on power delivery. Moreover, PDS efficiency is affected by workload variations at runtime, necessitating the need to design and manage PDSs and processors as a whole to improve system energy efficiency while balancing performance. In this article, we propose an offline-online co-design optimization methodology that combines offline PDS design optimization with online power management. To address the power consumption and delivery mismatch, we introduce a centralized deep Q-network (DQN)-based online control scheme for power co-management in chiplet-based multicore systems. By carefully designing the state space and reward functions, our approach achieves workload-aware adaptive control to reduce the energy-delay-product (EDP) while maintaining PDS efficiency under a given performance target (PT). We conduct evaluations on realistic applications to validate the effectiveness of our approach. For 64-core systems, our method achieves an average EDP reduction of 67% while meeting a 90% PT, surpassing state-of-the-art modular Q-learning (MQL)-based and heuristic-based approaches by up to 4% and 16%, respectively. Additionally, our approach demonstrates wiser action selection policies, higher control stability, and lower implementation overhead compared to the MQL-based approach. Xiao Li 0038, Lin Chen 0029, Shixi Chen, Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | Smart Knowledge Transfer-based Runtime Power ManagementabstractAs Moore's law slows down, computing systems must pivot towards higher energy efficiency to continue scaling performance. Reinforcement learning (RL) performs more adaptively than conventional methods in runtime power management under varied hardware configurations and varying software workloads. However, prior works on either model-free or model-based RL approaches face a non-negligible challenge: relearning the policies to adapt to the new environment is unacceptably time-consuming, especially when encountering significant variances in workloads or hardware configurations. Moreover, existing research on accelerating learning has focused on the speedup while largely ignoring the efficiency degradation of the results. In this paper, we present a smart transfer-enabled Q-learning (STQL) approach to boost the learning process and guarantee the learning efficiency through a contradiction checking mechanism, which wisely evicts inappropriate transferred knowledge. Experiments on realistic applications show that the proposed method can speed up the learning process to up to 2.3x and achieve a 6.2% energy-delay product (EDP) reduction compared to the state-of-the-art design. Lin Chen 0029, Xiao Li 0038, Fan Jiang 0015, Chengeng Li, Jiang Xu 0001 |
DATE | 2 |
| 2023 | RONet: Scaling GPU System with Silicon Photonic ChipletabstractModern GPU systems integrate hundreds of SMs on a single die, and future scaling envisions even more SMs being incorporated. However, the limited number of transistors per die constrains this growth. While current chiplet technology shows promise, its performance is limited by the bandwidth and energy efficiency of existing chiplet interconnect technologies. In contrast, optical interconnects offer ultra-high bandwidth and energy efficiency, making them ideal for high-performance chiplet-based GPUs. This work proposes a novel region-based optical network, called RONet, that divides a chiplet-based GPU system with a 2D Mesh layout into multiple row and column regions, where each region is connected by a separate optical link. Additionally, RONet employs a tuning-free transmission mechanism to further enhance inter-chiplet bandwidth. Experimental results show that RONet achieves 43% improvement on performance and 25.4% reduction on system energy consumption over the baseline. Chengeng Li, Fan Jiang 0015, Shixi Chen, Yinyi Liu, Lin Chen 0029, Xiao Li 0038, Jiang Xu 0001 |
ICCAD | 7 |
| 2022 | Improve the Stability and Robustness of Power Management through Model-free Deep Reinforcement LearningabstractAchieving high performance with low energy consumption has become a primary design objective in multi-core systems. Recently, power management based on reinforcement learning has shown great potential in adapting to dynamic environments without much prior knowledge. However, conventional Q-learning (QL) algorithms adopted in most existing works encounter serious problems about scalability, instability, and overestimation. In this paper, we present a deep reinforcement learning-based approach to improve the stability and robustness of power management while reducing the energy-delay product (EDP) under user-specified performance requirements. The comprehensive status of the system is monitored periodically, making our controller sensitive to environmental change. To further improve the learning effectiveness, knowledge sharing among multiple devices is implemented in our approach. Experimental results on multiple realistic applications show that the proposed method can reduce the instability up to 68% compared with QL. Through knowledge sharing among multiple devices, our federated approach achieves around 4.8% EDP improvement over QL on average. Lin Chen 0029, Xiao Li 0038, Jiang Xu 0001 |
DATE | 2 |
| 2022 | Fast and Accurate Statistical Simulation of Shared-Memory Applications on Multicore SystemsabstractDetailed cycle-accurate simulation of multicore systems is naturally slow. Statistical simulation is one alternative that permits trading off simulation speed for accuracy. However, there is a lack of effective memory locality models for multicore applications. Hence, existing statistical simulators neglect data-sharing between threads. Additionally, the standard method to speed up statistical simulations is to blindly reduce the trace length to be synthesized. While this gives good control over the speedup, it leaves the simulation error unbounded. In this work, we introduce a novel statistical simulation methodology for exploration of shared-memory multicore systems. It includes a newsharing-localitymodel (Shalom) that captures and reproduces data-sharing in multithread applications. Furthermore, we propose a method to bound the simulation error for a particular metric while maximizing speedup. The technique works by monitoring the convergence of the statistical synthesis. It is referred to asconvergence-deterministicsimulation (Condens). The combination ofShalomandCondensis around 130x faster than cycle-accurate simulations with reasonable accuracy loss. Our approach is also 5x faster than state-of-the-art sampling simulation under the same accuracy level. Compared to previous statistical simulators ignoring sharing, our approach is 2x more accurate for performance metrics and 8x more accurate for cache miss estimations. Fan Jiang 0015, Rafael Kioji Vivas Maeda, Jun Feng 0008, Shixi Chen, Lin Chen 0029, Xiao Li 0038, Jiang Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | Multi-Core Power Management through Deep Reinforcement LearningabstractAchieving high energy efficiency is a primary design objective for multi-core systems. Dynamic voltage and frequency scaling (DVFS) is one of the most widely-adopted low-power techniques. In this paper, we present a reinforcement learning-based DVFS control approach to reduce energy consumption under user-specified performance requirements. The learning agent periodically selects the voltage and frequency level for all cores based on observations of their computation intensiveness, memory behaviors as well as synchronization among cores. Experimental results on multiple real applications show that the proposed method can achieve significant energy reduction. Zhongyuan Tian, Lin Chen 0029, Xiao Li 0038, Jun Feng 0008, Jiang Xu 0001 |
ISCAS | 3 |
| 2020 | Efficient Optical Power Delivery System for Hybrid Electronic-Photonic Manycore ProcessorsabstractA lot of efforts have been devoted to optically enabled high-performance communication infrastructures for future manycore processors. Silicon photonic network promises high bandwidth, high energy efficiency and low latency. However, the ever-increasing network complexity results in high optical power demands, which stress the optical power delivery and affect delivery efficiency. Facing these challenges, we propose Ring-based Optical Active Delivery (ROAD) system, to effectively manage and efficiently deliver high optical power throughout photonic-electronic hybrid systems. Experimental results demonstrate up to 5.49X energy efficiency improvement compared to traditional design without affecting processor performance. Shixi Chen, Jiang Xu 0001, Xuanqi Chen, Jun Feng 0008, Zhongyuan Tian, Xiao Li 0038 |
DATE | 8 |