EDBT 2026 Demo / reviewers in the wild / expert
Tao Lu 0012
dblp:03/5189-12
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0009-0926-0451ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NBCache: An Efficient and Scalable Non-Blocking Cache for Coherent Multi-Chiplet Systems
Zhirong Ye, Tao Lu 0012, Zhaolin Li, Zhiyi Yu, Mingyu Wang 0003 |
ASP-DAC | 4 |
| 2026 | MAX-SM: High-Utilization Dynamic SM Partitioning for Heterogeneous Workloads on Multitasking Chiplet-Based GPUs
Mingyu Wang 0003, Tao Lu 0012, Baiqing Zhong, Zhaolin Li, Zhiyi Yu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | CINOC: Computing in Network-On-Chip With Tiled Many-Core Architectures for Large-Scale General Matrix MultiplicationsabstractLarge-scale general matrix multiplications (LMMs) are the key bottlenecks in various computation domains such as Transformer applications. However, it is a challenge to perform LMMs efficiently on traditional multi/many-core processor systems due to the large amount of memory access and the tight dependence of data transmission. By analyzing the aforementioned problems, we propose a computing in network-on-chip paradigm to perform LMMs by mitigating the performance losses caused by limited on-chip cache resources and memory bandwidth. Specifically, we propose a co-design of computable network-on-chip and the last-level cache method in tiled many-core architectures, which can reconstruct the redundant cache capacity as computable input buffer to balance the demands of computing, storage, and communication for the running LMM applications. Furthermore, a data-aware thread execution mechanism is also proposed to maximize the computational efficiency of thread streams in computable network. At the software level, memory-friendly matrix partitioning strategy, hybrid routing method and programming model are designed to bridge the gap between application demands and mismatched hardware/software interfaces. Experimental evaluations demonstrate that this proposed work achieves a computational latency reduction of 45% compared to the state-of-the-art GPU architecture, and the inference performance is improved by$2\times $of the GPT network. Mingyu Wang 0003, Jiahua Yan, Tao Lu 0012, Zhiyi Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | A Scalable Deadlock-Free Static Routing Algorithm for Chiplet-Based SystemsabstractThe utilization of the Chiplet methodology can accelerate VLSI system development and provide better flexibility. Building interconnection networks across multiple Chiplets and ensuring high-performance deadlock-free routing in systems with diverse irregular topologies is a challenging task.To avoid the reordering introduced by adaptive routing algorithms, a scalable static deadlock-free routing algorithm specifically designed for Chiplets is proposed. This approach capitalizes on static routing, a feature that ensures consistent message order due to fixed paths, fundamentally averting reordering issues. By adaptively configuring the state of the local router, it is possible to proactively initiate turns or exit detour loops, thereby effectively preventing deadlocks. Furthermore, by employing state configuration, the routing remains deadlock-free even when there are variations in the number of Chiplets, scale, or internal topology. This showcases the system’s scalability.Due to limited wiring resources, we chose classical routing algorithms up*/down* to compare, and the results showed a significant latency advantages, with a saturation injection rate approximately 1.5 to 2 times higher. Mingyu Wang 0003, Yicong Zhang, Tao Lu 0012, Zhiyi Yu |
ICPADS | 4 |