EDBT 2026 Demo / reviewers in the wild / expert
Chengeng Li
dblp:331/6679
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2024
0000-0003-0740-0974ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SCNoCs: An Adaptive Heterogeneous Multi-NoC with Selective Compression and Power GatingabstractIn-network compression has been proposed recently to support efficient communication. However, we find employing compression blindly cannot always pay off since de/compression leads to extra packet transmission delay. We thereby propose selective compression which compresses data adaptively based on network state and predicted compression ratio. Moreover, we observe that simply applying selective compression in a conventional single network is not energy efficient. Therefore, we propose SCNoCs, a heterogeneous Multi-NoC (Main-Net and HelperNet) architecture with the support of selective compression and power gating. SCNoCs can dynamically adjust the policy of selective compression and the utilization degree of the Helper-Net according to the network state at runtime. Experimental results show that our selective compression outperforms conventional compression by 1.5$ \times $. Besides, our proposed SCNoCs achieves comparable performance while reducing energy consumption by 43.4%, compared with the baseline. Fan Jiang 0015, Chengeng Li, Lin Chen 0029, Wei Zhang 0012, Jiang Xu 0001 |
ASPDAC | 2 |
| 2024 | Collaborative Coalescing of Redundant Memory Access for GPU SystemabstractGPU-based computing serves as the primary solution driving the performance of HPC systems. However, modern GPU systems encounter performance bottlenecks resulting from heavy memory access traffic and insufficient NoC bandwidth. In this work, we propose a collaborative coalescing mechanism aimed at eliminating redundant memory access and boosting GPU system performance. To achieve this, we design a coalescing unit for each memory partition, effectively merging requests from both inter-cluster and intra-cluster SMs. Additionally, we introduce a hierarchical multicast module to replicate and distribute the coalesced reply messages to multiple destination SMs. Experimental results show that our method achieves 20.6% improvement on performance and 27.1% reduction on NoC traffic over the baseline. Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
ASPDAC | 2 |
| 2024 | Towards Scalable GPU System with Silicon Photonic ChipletabstractGPU-based computing has emerged as a predominant solution for high-performance computing and machine learning applications. The continuously escalating computing demand fore-sees a requirement for larger-scale GPU systems in the future. However, this expansion is constrained by the finite number of transistors per die. Although chip let technology shows potential for building large-scale systems, current chiplet interconnection technologies suffer from limitations in both bandwidth and en-ergy efficiency. In contrast, optical interconnect has ultra-high bandwidth and energy efficiency, and thereby is promising for constructing chiplet-based GPU systems. Yet, previously proposed optical networks lack scalability and cannot be directly applied to existing chiplet-based GPU systems. In this work, we address the challenges of designing large-scale G PU systems with silicon photonic chiplets. We propose GROOT, a group-based optical network that divides the entire system into groups and facilitates resource sharing among the chiplets within each group. Additionally, we design dedicated channel mapping and allocation policies tailored for the request network and the reply network, respectively. Experimental results show that GROOT achieves 48% improvement on performance and 24.5% reduction on system energy consumption over the baseline. Chengeng Li, Shixi Chen |
DATE | 1 |
| 2024 | PhotonNTT: Energy-Efficient Parallel Photonic Number Theoretic Transform AcceleratorabstractFully homomorphic encryption (FHE) presents a promising opportunity to remove privacy barriers in various scenarios including cloud computing and secure database search, by enabling computation on encrypted data. However, integrating FHE with real-world applications remains challenging due to its significant computational overhead. In the FHE scheme, Number Theoretic Transform (NTT) consumes the primary computing resources and has great potential for acceleration. For the first time, we present a photonic NTT accelerator, PhotonNTT, with high energy efficiency and parallelism to address the above challenge. Our approach involves formulating the NTT into matrix-vector multiplication (MVM) operations and mapping the data flow into parallel photonic MVM units. A dedicated data mapping scheme is proposed to introduce free spectral range (FSR) and distributed RAM design into the system, which enables a high bit-wise parallelism level. The system's reliability is validated through the Monte-Carlo BER analysis. The experimen-tal evaluation shows that the proposed architecture outperforms SOTA CiM-based NTT accelerators with an improvement of 50x in throughput and 63x improvement in energy efficiency. Yinyi Liu, Chengeng Li, Shixi Chen, Fengshi Tian, Wei Zhang 0012, Jiang Xu 0001 |
DATE | 6 |
| 2024 | NEOCNN: NTT-Enabled Optical Convolution Neural Network AcceleratorabstractIn the realm of neural network computation, optical neural network accelerators (ONNs) have emerged as a promising solution, leveraging the inherent speed and parallelism of optical systems. Despite their potential, current ONN designs often fall short due to inefficient data movement and reliance on traditional electronics-based dataflows. Yinyi Liu, Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
ICS | 4 |
| 2024 | PC-oriented Prediction-based Runtime Power Management for GPGPU using Knowledge TransferabstractAs Moore's law slows down, computing systems must prioritize higher energy efficiency to sustain performance scaling. GPUs have emerged as the primary workhorses of computing resources, making the achievement of high energy efficiency in GPUs a critical concern. However, implementing runtime power management on GPUs poses significant challenges due to the high variations and complexities arising from workloads and hardware configurations, which render offline optimization and reactive-based methods less effective. In this paper, we present a program counter (PC)-oriented prediction-based power management approach for GPGPUs. Our approach leverages the benefits of prediction to address online variations while enhancing prediction capability through knowledge transfer across different levels of resources. Experiments conducted on realistic applications demonstrate that our proposed method achieves the maximum energy savings under a user-defined performance constraint compared to state-of-the-art designs. Lin Chen 0029, Xiao Li 0038, Shixi Chen, Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
SPAA | 5 |
| 2024 | Deep Reinforcement Learning-Based Power Management for Chiplet-Based Multicore SystemsabstractChiplet technology has emerged as a promising solution to address the increasing demand for high-performance computing in light of the slowdown of Moore’s law. While chiplet-based multicore systems offer higher performance through heterogeneous integration, they also pose challenges for power delivery system (PDS) design. The integration of additional vertical and inter-chiplet connections, along with higher power density, impose stringent requirements on power delivery. Moreover, PDS efficiency is affected by workload variations at runtime, necessitating the need to design and manage PDSs and processors as a whole to improve system energy efficiency while balancing performance. In this article, we propose an offline-online co-design optimization methodology that combines offline PDS design optimization with online power management. To address the power consumption and delivery mismatch, we introduce a centralized deep Q-network (DQN)-based online control scheme for power co-management in chiplet-based multicore systems. By carefully designing the state space and reward functions, our approach achieves workload-aware adaptive control to reduce the energy-delay-product (EDP) while maintaining PDS efficiency under a given performance target (PT). We conduct evaluations on realistic applications to validate the effectiveness of our approach. For 64-core systems, our method achieves an average EDP reduction of 67% while meeting a 90% PT, surpassing state-of-the-art modular Q-learning (MQL)-based and heuristic-based approaches by up to 4% and 16%, respectively. Additionally, our approach demonstrates wiser action selection policies, higher control stability, and lower implementation overhead compared to the MQL-based approach. Xiao Li 0038, Lin Chen 0029, Shixi Chen, Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | Smart Knowledge Transfer-based Runtime Power ManagementabstractAs Moore's law slows down, computing systems must pivot towards higher energy efficiency to continue scaling performance. Reinforcement learning (RL) performs more adaptively than conventional methods in runtime power management under varied hardware configurations and varying software workloads. However, prior works on either model-free or model-based RL approaches face a non-negligible challenge: relearning the policies to adapt to the new environment is unacceptably time-consuming, especially when encountering significant variances in workloads or hardware configurations. Moreover, existing research on accelerating learning has focused on the speedup while largely ignoring the efficiency degradation of the results. In this paper, we present a smart transfer-enabled Q-learning (STQL) approach to boost the learning process and guarantee the learning efficiency through a contradiction checking mechanism, which wisely evicts inappropriate transferred knowledge. Experiments on realistic applications show that the proposed method can speed up the learning process to up to 2.3x and achieve a 6.2% energy-delay product (EDP) reduction compared to the state-of-the-art design. Lin Chen 0029, Xiao Li 0038, Fan Jiang 0015, Chengeng Li, Jiang Xu 0001 |
DATE | 4 |
| 2023 | RONet: Scaling GPU System with Silicon Photonic ChipletabstractModern GPU systems integrate hundreds of SMs on a single die, and future scaling envisions even more SMs being incorporated. However, the limited number of transistors per die constrains this growth. While current chiplet technology shows promise, its performance is limited by the bandwidth and energy efficiency of existing chiplet interconnect technologies. In contrast, optical interconnects offer ultra-high bandwidth and energy efficiency, making them ideal for high-performance chiplet-based GPUs. This work proposes a novel region-based optical network, called RONet, that divides a chiplet-based GPU system with a 2D Mesh layout into multiple row and column regions, where each region is connected by a separate optical link. Additionally, RONet employs a tuning-free transmission mechanism to further enhance inter-chiplet bandwidth. Experimental results show that RONet achieves 43% improvement on performance and 25.4% reduction on system energy consumption over the baseline. Chengeng Li, Fan Jiang 0015, Shixi Chen, Yinyi Liu, Lin Chen 0029, Xiao Li 0038, Jiang Xu 0001 |
ICCAD | 1 |
| 2022 | Accelerating Cache Coherence in Manycore Processor through Silicon Photonic ChipletabstractCache coherence overhead in manycore systems is becoming prominent with the increase of system scale. However, traditional electrical networks restrict the efficiency of cache coherence transactions in the system due to the limited bandwidth and long latency. Optical network promises high bandwidth and low latency, and supports both efficient unicast and multicast transmission, which can potentially accelerate cache coherence in manycore systems. This work proposes a novel photonic cache coherence network with a physically centralized logically distributed directory called PCCN for chiplet-based manycore systems. PCCN adopts a channel sharing method with a contention solving mechanism for efficient long-distance coherence-related packet transmission. Experiment results show that compared to state-of-the-art proposals, PCCN can speed up application execution time by 1.32x, reduce memory access latency by 26%, and improve energy efficiency by 1.26x, on average, in a 128-core system. Chengeng Li, Fan Jiang 0015, Shixi Chen, Yinyi Liu, Jiang Xu 0001 |
ICCAD | 1 |