Chen Li 0015

dblp:l/ChenLi15 · also Li Chen 0023 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-0684-1754ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 WSGraph: A Framework for Tackling Redundant and Irregular Data Access in Streaming Graph Processing
abstract
The demand for real-time streaming graph analysis has grown significantly, as hundreds of thousands of updates come every second. Monotonic graph algorithms such as Shortest Path are widely used in real-time analytics, but there are two bottlenecks that limit their performance, specially on planar graphs. One is massive redundant data accesses due to irregular state propagations and the other is high memory latency caused by irregular data accesses. We observe that existing systems mainly focus on general-purpose graph algorithms. If the properties of specific graph algorithms are exploited, the analysis performance can be further improved. Moreover, these systems typically tackle these two bottlenecks separately through either software or hardware mechanisms, but not both. However, both bottlenecks need to be addressed simultaneously in real scenarios such as road navigation. This article proposes WSGraph, a software-hardware co-design framework for high-performance streaming graph processing. WSGraph tackles these two challenges by enforcing regularized processing orders and enabling precise data prefetching. Specifically, at the software level, WSGraph integrates a priority-based work scheduler with sliding-window bucket mapping scheme to regulate state propagations, thereby drastically reducing redundant data accesses. At the hardware level, WSGraph incorporates a lightweight in-core Proactive Data Engine (PDE). By exploiting intra-vertex access regularity, the PDE accurately prefetches relevant graph data to effectively hide the high latency of irregular memory accesses. Experimental results demonstrate that WSGraph achieves significant performance improvements over existing systems. Compared with the state-of-the-art software system KickStarter, WSGraph gains a 2.13× speedup primarily by reducing graph data accesses by an average of 78.6%.
Xuanyi Li, Chen Li 0015, Zhengyi Dai, Jianzhuang Lu, Yang Guo 0003
ACM Trans. Archit. Code Optim.2
2025 LAMP: A Locality-Aware TLB Sharing Framework for Multi-Tenant GPUs via TLB Comprehensive Profiling
abstract
With the growing prevalence of cloud services, GPUs are increasingly shared across multiple tenants, making address translation a performance-critical path. Our study makes three key observations: 1) Shared L2 TLB becomes the bottleneck in multi-tenant systems. Intense competition for entries can lead to high miss rates and thrashing under heavy pressure. 2) Benchmarks vary significantly in TLB sensitivity due to differing access patterns and memory reuse potential. 3) The prevailing management strategies, including free-for-all sharing and static partitioning, fail to account for tenant access behavior. And the state-of-the-art token-based strategy relies solely on miss rate, an inadequate metric that fails to capture per-tenant access patterns. Based on these observations, we propose LAMP, a localityaware TLB sharing framework. LAMP comprehensively profiles the access pattern considering three key aspects: spatial locality, temporal locality, and access density. These aspects allow LAMP to assess each tenant's sensitivity to TLB capacity and their reuse potential. Based on these insights, LAMP optimizes the sharing by dynamically partitioning L2 TLB, granting more entries to reuse-efficient tenants while limiting wasteful occupancy. This reduces thrashing and cross-tenant interference. Experimental results show that LAMP improves system performance by$\mathbf{1 2. 3 \%}$on average across a range of multi-tenant workloads, with negligible hardware overhead.
Chen Li 0015, Xuanyi Li, Jianzhuang Lu, Yang Guo 0003
HPCC2
2025 A Multi-objective Mapping Optimization Strategy for Large-Scale NoC Design
Chen Li 0015, Jianzhuang Lu, Yang Guo 0003
ICA3PP (5)2
2024 LWECC: A Lightweight ECC Technology for HPC Accelerators Supporting Multi-granularity Memory Access
abstract
Heterogeneous multi-core accelerators have exhibited remarkable computational performance, establishing their competitiveness within the realm of high-performance computing. The on-chip memory to support reliable and correct data access is a crucial element of the accelerator. With the continuous scaling down of semiconductor process nodes, memory units become more susceptible to soft errors caused by Single Event Upset (SEU), thus significantly increasing the risk of memory access errors. Additionally, data interaction between cores supports a variety of granularities, making it challenging to achieve error tolerant designs for multi-granular data with low hardware overhead. This paper proposes the Light-Weight ECC (LWECC) integrated in the 64-bit accelerator’s Scalar Memory (SM) to improve the efficiency of data protection for multi-granularity memory access. LWECC adopts a software-hardware co-design approach, effectively balancing area overhead and providing comprehensive and flexible data protection. Compared to the finest-grained ECC solution as a baseline, using LWECC can reduce redundant memory area by 80.90% and ECC circuit area by 16.87%, markedly reducing the whole hardware overhead of the SM. In addition, LWECC efficiently caters to diverse application domains, supporting not only high-performance computing applications but also accommodating fine-grained memory access applications with low operational latency.
Lanting Guo, Chen Li 0015, Sheng Liu 0001
ISCAS3
2024 AFTP: An Adaptive High Cost-Effectiveness Fault-Tolerant NoC Design Based on Prediction
abstract
Network-on-Chip (NoC), known for its high bandwidth and scalability, is extensively utilized in chip multiprocessors. However, as technology advances to the nanometer scale, NoC is becoming increasingly vulnerable to errors caused by crosstalk, radiation, electromagnetic interference, and other factors. In addition to ensuring network reliability, designers must consider the overhead instead of blindly pursuing fault-tolerance capabilities in NoC design. In this paper, we analyze the characteristics of conventional End-to-End (E2E) and Switch-to-Switch (S2S) designs and propose AFTP, an adaptive high cost-effectiveness fault-tolerant NoC design based on prediction. Our design prioritizes stronger protection for head flits and relaxed protection for body and tail flits, and it is capable of dynamically adjusting the decoding times for flits in response to changes in error rate, aiming to achieve a balance between the overhead and the reliability of NoC. Our design demonstrates a significant improvement in cost-effectiveness compared to conventional E2E and S2S designs, achieving a 5.9x and 8.9x improvement, respectively, under common synthetic traffic and PARSEC benchmarks.
Ta Tan, Chen Li 0015, Jianzhuang Lu, Qijin Zhu
ISPA3
2022 Adaptive Low-Cost Loop Expansion for Modulo Scheduling
Hongli Zhong, Zhong Liu 0003, Sheng Liu 0001, Sheng Ma, Chen Li 0015
NPC5
2022 Optimizing convolutional neural networks on multi-core vector accelerator
Zhong Liu 0003, Xin Xiao 0008, Chen Li 0015, Sheng Ma, Rangyu Deng
Parallel Comput.3
2021 Improving Inter-kernel Data Reuse With CTA-Page Coordination in GPGPU
abstract
Although modern GPUs are equipped with expanding memory, accommodating the entire working set of large-scale workloads can still be a challenge. With the support of unified virtual memory and demand paging, programmers can transparently oversubscribe the main memory. However, this transparent management still comes at a severe performance cost, especially for applications with inter-kernel data sharing. While there have been many efforts to reduce additional data migrations caused by the memory oversubscription, few consider the reuse of shared data during the boundary of adjacent kernels. Due to limited memory capacity, we observe that adjacent kernel often demands shared pages that were evicted by the previous kernel, resulting in a significant number of costly data migrations. In this paper, we propose a CTA-Page collaborative framework, called CPC, that transparently reduces the impact of memory oversubscription using CTA dispatch switching and page replacement switching coordinately to reuse inter-kernel shared data. We evaluate CPC with a variety of GPGPU benchmark suites. Experimental results show that the system performance is improved by 65 % compared with the state-of-the-art technique for applications with inter-kernel data sharing.
Xuanyi Li, Chen Li 0015, Yang Guo 0003, Rachata Ausavarungnirun
ICCAD2
2021 Advancing DSP into HPC, AI, and beyond: challenges, mechanisms, and future directions
Chen Li 0015, Chang Liu 0019, Sheng Liu 0001, Yuanwu Lei, Jian Zhang 0022, Yang Guo 0003
CCF Trans. High Perform. Comput.2
2020 Active one-shot learning by a deep Q-network strategy
Chen Li 0015, Honglan Huang, Yang-He Feng, Guangquan Cheng, Jincai Huang 0001, Zhong Liu 0002
Neurocomputing1
2020 A Dynamic and Proactive GPU Preemption Mechanism Using Checkpointing
abstract
The demand for multitasking GPUs increases whenever the GPU may be shared by multiple applications, either spatially or temporally. This requires that GPUs can be preempted and switch context to a new application while already executing one. Unlike CPUs, context switching in GPUs is prohibitively expensive due to the large context states to swap out. There have been a number of efforts on reducing the overhead of preemption, through reducing the context sizes or overlapping context switching with execution. All those techniques are reactive approaches, meaning that context switching occurs when the preemption request arrives. In this paper, we propose a dynamic and proactive mechanism to reduce the latency of preemption. We observe that kernel execution is almost always preceded by known commands in both CUDA and OpenCL implementations. Hence, a preemption can be anticipated before the actual request arrives. We study such lead time and develop a prediction scheme to perform an early state saving. When the actual preemption is invoked, an incremental update relative to the previous saved state is performed, much like the conventional checkpointing mechanism. Our design can also choose to drain or checkpointing dynamically and accurately according to the feature of kernels in the runtime. This design effectively reduces the stall time of the preempting kernel due to context switching by 58.6%. Moreover, through careful handling of the saved state, we can also reduce the overall size of saved state by an average of 23.3%, compared with a full context switching.
Chen Li 0015, Andrew Zigerelli, Jun Yang 0002, Youtao Zhang, Sheng Ma, Yang Guo 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 A Framework for Memory Oversubscription Management in Graphics Processing Units
abstract
Modern discrete GPUs support unified memory and demand paging. Automatic management of data movement between CPU memory and GPU memory dramatically reduces developer effort. However, when application working sets exceed physical memory capacity, the resulting data movement can cause great performance loss.
Chen Li 0015, Rachata Ausavarungnirun, Christopher J. Rossbach, Youtao Zhang, Onur Mutlu, Yang Guo 0003, Jun Yang 0002
ASPLOS1
2018 PEP: proactive checkpointing for efficient preemption on GPUs
abstract
The demand for multitasking GPUs increases whenever the GPU may be shared by multiple applications, either spatially or temporally. This requires that GPUs can be preempted and switch context to a new application while already executing one. Unlike CPUs, context switching in GPUs is prohibitively expensive due to the large context states to swap out. There have been a number of efforts on reducing the overhead of preemption, through reducing the context sizes or overlapping context switching with execution. All those techniques are reactive approaches, meaning that context switching occurs when the preemption request arrives.
Chen Li 0015, Andrew Zigerelli, Jun Yang 0002, Yang Guo 0003
DAC1
2017 Fairness-Oriented and Location-Aware NUCA for Many-Core SoC
abstract
Non-uniform cache architecture (NUCA) is often employed to organize the last level cache (LLC) by Networks-on-Chip (NoC). However, along with the scaling up for network size of Systems-on-Chip (SoC), two trends gradually begin to emerge. First, the network latency is becoming the major source of the cache access latency. Second, the communication distance and latency gap between different cores is increasing. Such gap can seriously cause the network latency imbalance problem, aggravate the degree of non-uniform for cache access latencies, and then worsen the system performance.
Zicong Wang, Chen Li 0015, Yang Guo 0003
NOCS3
2017 A high performance reliable NoC router
Lu Wang 0019, Sheng Ma, Chen Li 0015, Wei Chen 0009, Zhiying Wang 0003
Integr.3
2016 DLL: A dynamic latency-aware load-balancing strategy in 2.5D NoC architecture
abstract
As the 3D stacking technology still faces several challenges, the 2.5D stacking technology gains better application prospects nowadays. With the silicon interposer, the 2.5D stacking can improve the bandwidth and capacity of the memory system. To satisfy the communication requirements of the integrated memory system, the free routing resources in the interposer should be explored to implement an additional network. Yet, the performance is strongly limited by the unbalanced loads between the CPU-layer network and the interposer-layer network. In this paper, to address this issue, we propose a dynamic latency-aware load-balancing (DLL) strategy. Our key innovations are detecting congestion of the network layer via the average latency of recent packets and making the network layer selection at each source node. We leverage the free routing resources in the interposer to implement a latency propagation ring. With the ring, the latency information tracked at destination nodes is propagated back to source nodes. We achieve load-balance by using these information. Experimental results show that compared with the baseline design, a destination-detection strategy and a buffer-aware strategy, our DLL strategy achieves 45%, 14.9% and 6.5% of average throughput improvements with minor overheads.
Chen Li 0015, Sheng Ma, Lu Wang 0019, Zicong Wang, Xia Zhao 0004, Yang Guo 0003
ICCD1
2016 A heterogeneous low-cost and low-latency Ring-Chain network for GPGPUs
abstract
To achieve high throughput, core count in compute accelerators such as General-Purpose Graphics Processing Units (GPGPUs) increases continuously. The communication demand of these cores boosts the demand for a low-latency packet switched network. As packet latency is mainly composed of per-hop latency, contention latency and serialization latency, a favorable Network-on-Chip (NoC) design should efficiently decrease these three latency contributors to meet the communication demand while keeping hardware cost low. In this paper, we first make two observations about the NoC differences between CMPs and GPGPUs, and then design a Heterogeneous Ring-Chain network (HRCnet) for the GPGPU reply network. HRCnet eliminates conflicts in the network by proposing a ring-similar topology, using a novel node placement and introducing unidirectional channels. Eliminating conflicts reduces the per-hop latency and removes the contention latency, and exploiting the ring-similar topology reduces the serialization latency. Experimental results show the benefits of the low-cost low-latency design. With the same bisection bandwidth compared to the baseline mesh, our work yields a 45% performance improvement while reducing the area by 42% and reducing energy consumption by 60%. Compared to two state-of-the-art GPGPU NoCs, BENoC and DA2mesh, HRCnet achieves more than 42% performance gain at reduced hardware cost. Our work also achieves the highest power and area efficiency among the designs.
Xia Zhao 0004, Sheng Ma, Chen Li 0015, Lieven Eeckhout, Zhiying Wang 0003
ICCD3
2015 Adaptive remaining hop count flow control: Consider the interaction between packets
abstract
The interaction between packets affects performance and global fairness of Network-on-Chip. Preferentially transferring packets with small remaining hop counts (PPSR) can reduce the flying packet amount to improve the performance. Yet, the global fairness is negatively affected. In contrast, preferentially transferring packets with large remaining hop counts (PPLR) can achieve better global fairness with a poorer performance. In this paper, we propose adaptive remaining hop count flow control, which dynamically switches between PPSR and PPLR. In this way, we can achieve higher performance and better global fairness.
Peng Wang 0036, Sheng Ma, Hongyi Lu, Zhiying Wang 0003, Chen Li 0015
ASP-DAC5