EDBT 2026 Demo / reviewers in the wild / expert
Le Luo 0002
dblp:35/4441-2
· DBLP profile ↗
8ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-2368-8085ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Structure-aware thread throttling for energy-efficient graph processing on shared-memory systems
Yili Chen, Le Luo 0002 |
Future Gener. Comput. Syst. | 2 |
| 2025 | Improving Energy Efficiency of Graph Processing on Shared-Memory SystemsabstractWith the number of cores increasing in shared memory systems, the energy consumption of parallel computing on them is becoming increasingly prominent. Currently, researchers concern with the performance optimization, while ignoring the energy efficiency of graph processing. Meanwhile, existing works that optimize energy efficiency involve mainly the general benchmarks by using dynamic voltage and frequency scaling and thread throttling methods. However, these methods cannot be directly transplanted to graph processing, because most graph algorithms converge in fewer iterations and traditional energy efficiency optimization methods are not applicable to them and will produce much overhead, resulting in the fact that the loss outweighs the gain. And some energy-saving methods estimate the subsequent CPU frequency based on the run-time system state, which leads to an inaccurate prediction of the optimal energy-saving CPU frequency. In view of the above issues, we propose a pre-allocated thread throttling method and a static frequency scaling method. The former achieves thread throttling by establishing a pre-allocated scheduling method, which calculates the optimal energy-saving number of threads promptly when the graph is loaded; On this basis, in order to reduce the cost of dynamic frequency setting at runtime and improve the energy efficiency further, the latter introduces the static frequency scaling method to reduce the execution speed of some tasks by relaxing thread execution time. The experimental results show that the pre-allocated thread throttling method improves the energy efficiency by about 10% compared to the original framework, and the static frequency scaling method further improves it by about 20% with trivial performance loss. Le Luo 0002, Chao Li 0009, Jinyang Guo 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2025 | Identifying Optimal Workload Offloading Partitions for CPU-PIM Graph Processing AcceleratorsabstractThe integrated architecture that features both in-memory logic and host processors, or so-called “processing-in-memory” (PIM) architecture, is an emerging and promising solution to bridge the performance gap between the memory and host processors. In spite of the considerable potential of PIM, the workload offloading policy, which partitions the program and determines where code snippets are executed, is still a main challenge in PIM. In order to determine the best PIM offloading partitions, existing methods require in-depth program profiling to create the control flow graph (CFG) and then transform it into a graph-cut problem. These CFG-based solutions depend on detailed profiling of a crucial element, the execution time of basic blocks, to accurately assess the benefits of PIM offloading. The issue is that these execution times can change significantly in PIM, leading to inaccurate offloading decisions. To tackle this challenge, we present a novel PIM workload offloading framework called “RDPIM” for CPU-PIM graph processing accelerators, which systematically considers the variations in the execution time of basic blocks. By analyzing the relationship between data dependencies among workloads and the connectivity of input graphs, we identified three key features that can lead to variations in execution time. We developed a novel reuse distance (RD)-based model to predict the exact performance of basic blocks for optimal offloading decisions. We evaluate RDPIM using real-world graphs and compare it with some state-of-the-art PIM offloading approaches. Experiments have demonstrated that our method achieves an average speedup of$2\times $compared to CPU-only executions and up to$1.6\times $compared to state-of-the-art PIM offloading schemes. Le Luo 0002, Wu Zhou 0007, Xiaoming Chen 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | Reconfigurable Fault-Tolerant Link With Bandwidth Expansion for 2.5-D Chiplet-Based SystemsabstractThe 2.5-D chiplet-based systems offer a promising path toward higher performance and integration density, but the reliability of interchiplet links, particularly vertical links (VLs), poses a significant challenge. This article proposes reconfigurable bidirectional link (ReBL), a novel architecture providing robust fault tolerance and enhanced bandwidth for these critical interconnects. ReBL features three key innovations: 1) a robust fault-tolerant mechanism leveraging ReBLs that dynamically adapts to and mitigates permanent link failures; 2) a dynamic bandwidth expansion technique that significantly enhances the system efficiency by utilizing idle links to optimize resource allocation and throughput; and 3) a virtual channel (VC) allocation strategy that guarantees deadlock-free operations through strategic channel partitioning and assignment. Evaluations using synthetic traffic and PARSEC benchmarks demonstrate ReBL’s significant advantages under high fault rates. Compared with the state-of-the-art reliable and deadlock-free routing (ReD) approach, ReBL achieves an average reduction of 21.3% in packet latency and 4.9% in application execution time across the evaluated benchmarks. These benefits are achieved with only a 6.06% area overhead over baseline. Wu Zhou 0007, Le Luo 0002, Fulong Chen 0002, Tianming Ni, Xiaoqing Wen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | DrPIM: An Adaptive and Less-blocking Data Replication Framework for Processing-in-Memory ArchitectureabstractProcessing-in-Memory (PIM) architecture suffers from frequent data synchronization, which raises inevitable coherent stalls and consumes more energy. This paper presents DrPIM, an adaptive and less-blocking data replication framework for Processing-in-Memory architecture to address such issues. The key insights are 1) finer-grained data management for the conflict data region and 2) the automatic generation and invalidations of data replications. The above two schemes significantly decrease the critical path of data synchronization and the access latency for conflict data, thus providing a less-blocking execution. Evaluations show that DrPIM achieves a speedup of 1.5x over the state-of-the-art and reduces the data-moving energy by 49%. Hongyu Xue, Le Luo 0002, Xingqi Zou |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Magas: matrix-based asynchronous graph analytics on shared memory systems
Le Luo 0002, Yi Liu 0013, Hailong Yang 0002, Depei Qian 0001 |
J. Supercomput. | 1 |
| 2020 | Processing graphs with barrierless asynchronous parallel model on shared-memory systems
Le Luo 0002, Yi Liu 0013 |
Future Gener. Comput. Syst. | 1 |
| 2019 | Improving parallel efficiency for asynchronous graph analytics using Gauss-Seidel-based matrix computationabstractSummary Graph analytics is extensively used in big‐data applications such as social networks, web analysis, bio‐informatics, etc. Most graph processing frameworks adopt vertex‐centric model due to its ease of use and programming. However, when dealing with asynchronous graph analytics, frameworks based on vertex programming perform inefficiently. The reason is that first, vertex programming must guarantee the sequential consistency, which means frequent use of locks or atomic operations, and second, the algorithms are parallelized in vertex level and latent parallelism of the algorithms cannot be exploited. To improve parallel efficiency of asynchronous graph processing, the Gauss‐Seidel style algorithms in particular, this paper proposes a scheduling model using Gauss‐Seidel‐based matrix computation, which converts the vertex programming into two main matrix operations and then algorithms are parallelized by row and column vectors. Compared to vertex programming, our model parallelizes algorithms in a finer way to exploit more latent parallelism, while retains the ease‐of‐programming advantage of vertex programming. Instead of using locks to guarantee the sequential consistency, our model uses a hybrid synchronization policy to reduce serializability among threads and overheads of context switching. Furthermore, this model strengthens locality of the program. Experiment results show that our model outperforms vertex‐centric asynchronous frameworks in both performance and scalability. Moreover, it even surpasses the matrix‐based synchronous framework GraphMat with some non‐Gauss‐Seidel style algorithms. Le Luo 0002, Yi Liu 0013 |
Concurr. Comput. Pract. Exp. | 1 |