Tao Lu 0012

dblp:03/5189-12 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0009-0926-0451ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2026 NBCache: An Efficient and Scalable Non-Blocking Cache for Coherent Multi-Chiplet Systems
Zhirong Ye, Tao Lu 0012, Zhaolin Li, Zhiyi Yu, Mingyu Wang 0003
ASP-DAC4
2026 MAX-SM: High-Utilization Dynamic SM Partitioning for Heterogeneous Workloads on Multitasking Chiplet-Based GPUs
Mingyu Wang 0003, Tao Lu 0012, Baiqing Zhong, Zhaolin Li, Zhiyi Yu
IEEE Trans. Parallel Distributed Syst.2
2025 CINOC: Computing in Network-On-Chip With Tiled Many-Core Architectures for Large-Scale General Matrix Multiplications
abstract
Large-scale general matrix multiplications (LMMs) are the key bottlenecks in various computation domains such as Transformer applications. However, it is a challenge to perform LMMs efficiently on traditional multi/many-core processor systems due to the large amount of memory access and the tight dependence of data transmission. By analyzing the aforementioned problems, we propose a computing in network-on-chip paradigm to perform LMMs by mitigating the performance losses caused by limited on-chip cache resources and memory bandwidth. Specifically, we propose a co-design of computable network-on-chip and the last-level cache method in tiled many-core architectures, which can reconstruct the redundant cache capacity as computable input buffer to balance the demands of computing, storage, and communication for the running LMM applications. Furthermore, a data-aware thread execution mechanism is also proposed to maximize the computational efficiency of thread streams in computable network. At the software level, memory-friendly matrix partitioning strategy, hybrid routing method and programming model are designed to bridge the gap between application demands and mismatched hardware/software interfaces. Experimental evaluations demonstrate that this proposed work achieves a computational latency reduction of 45% compared to the state-of-the-art GPU architecture, and the inference performance is improved by$2\times $of the GPT network.
Mingyu Wang 0003, Jiahua Yan, Tao Lu 0012, Zhiyi Yu
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 A Scalable Deadlock-Free Static Routing Algorithm for Chiplet-Based Systems
abstract
The utilization of the Chiplet methodology can accelerate VLSI system development and provide better flexibility. Building interconnection networks across multiple Chiplets and ensuring high-performance deadlock-free routing in systems with diverse irregular topologies is a challenging task.To avoid the reordering introduced by adaptive routing algorithms, a scalable static deadlock-free routing algorithm specifically designed for Chiplets is proposed. This approach capitalizes on static routing, a feature that ensures consistent message order due to fixed paths, fundamentally averting reordering issues. By adaptively configuring the state of the local router, it is possible to proactively initiate turns or exit detour loops, thereby effectively preventing deadlocks. Furthermore, by employing state configuration, the routing remains deadlock-free even when there are variations in the number of Chiplets, scale, or internal topology. This showcases the system’s scalability.Due to limited wiring resources, we chose classical routing algorithms up*/down* to compare, and the results showed a significant latency advantages, with a saturation injection rate approximately 1.5 to 2 times higher.
Mingyu Wang 0003, Yicong Zhang, Tao Lu 0012, Zhiyi Yu
ICPADS4