EDBT 2026 Demo / reviewers in the wild / expert
Linfeng Tao
dblp:218/1166
· DBLP profile ↗
5ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-7001-8893ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACsabstractGeneral matrix-matrix multiplication (GEMM), serving as a cornerstone of AI computations, has positioned tensor processing engines (TPEs) as increasingly critical components within existing GPUs and domain-specific architectures (DSA). Our analysis identifies that the prevailing architectures primarily focus on dataflow or operand reuse strategies, when considering the combination of matrix multiplication with multiply-accumulator (MAC) itself, it provides greater optimization space for the design of TPEs. This work introduces a novel perspective on matrix multiplication from a hardware standpoint, focusing on the bit-weight dimension of MACs. Through this lens, we propose a finer-grained TPE notation, using matrix triple loops as an example, introducing new methods and ideas for designing and optimizing PE microarchitecture. Based on the new notation and transformations, we propose four optimization techniques that achieve varying degrees of improvement in timing, area, and power consumption. We implement our design in RTL using the SMIC-28nm process. Applying our methods to four classic TPE architectures (include systolic array [20], 3D-Cube [27], multiplier-adder tree [48], and 2D-Matrix [30]), we achieved area efficiency improvements of $1.27 \times, 1.28 \times, 1.56 \times$, and $1.44 \times$, and $1.04 \times, 1.56 \times, 1.49 \times$, and $1.20 \times$ for energy efficiency respectively. When applied to a bit-slice architecture, we achieved a $12.10 \times$ improvement in energy efficiency and $2.85 \times$ in area efficiency compared to Laconic [38]. Our Verilog HDL code, along with timing, area, and power reports for circuit synthesis in URL: https://github.com/wqzustc/High-Performance-Tensor-Processing-Engines. Qizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng 0002, Zerong He, Linfeng Tao, Letian Zhao, Zhaoxi Zeng, Wei Yuan 0006, Xi Jin 0002 |
HPCA | 6 |
| 2025 | SageSC: Accelerating GraphSAGE Minibatch Inference on Memory-Intensive GraphsabstractGraph neural networks demonstrate excellent performance on node classification tasks in graph datasets. For inference tasks on memory-intensive graphs, the storage burden, memory access bottlenecks, and load imbalance issues arise. The minibatch inference proposed in GraphSAGE is an effective method for minimizing these problems. However, minibatch inference introduces new challenges: while it facilitates subsequent computation, the irregular random memory access pressure shifts to the minibatch construction phase, creating performance bottlenecks in the system. In this work, to address the aforementioned challenges, we propose a novel scattered minibatch construction and aggregation (SMCA) algorithm to optimize sampling, batch construction, and aggregation computations for minimizing their latency. This method distributes memoryintensive workloads and exploits the parallelism between memory groups. Evaluation results show that the proposed accelerator SageSC achieves speedups ranging from 3x to 96x compared to CPU/GPU baselines, especially on memory-intensive graphs, while outperforming existing state-of-the-art designs. Yuchen Gui, Wei Yuan 0006, Qizhe Wu, Huawen Liang, Letian Zhao, Linfeng Tao, Zhongguang Xu, Xi Jin 0002 |
ICCD | 6 |
| 2025 | MHE-TPE: Multi-Operand High-Radix Encoder for Mixed-Precision Fixed-Point Tensor Processing Engines
Qizhe Wu, Jinyi Zhou, Zhanhe Hu, Zhichen Zeng 0002, Huawen Liang, Jiuru Zhu, Linfeng Tao, Xin Zhang 0176, Zekang Cheng, Letian Zhao, Wei Yuan 0006, Xi Jin 0002 |
MICRO | 7 |
| 2019 | CINT - An Energy-efficient Mixed-signal In-Memory CNN Accelerator Based on NOR Flash MemoryabstractConvolutional neural network (CNN) is a power-hungry and resource-consuming application, which makes it hard to deploy on end devices. We propose a method to perform convolution operations in NOR flash memory. Experiment results show that our method has great performance and high energy efficiency. Linfeng Tao, Teng Tian, Zikun Xiang, Xi Jin 0002, Zhengda Li, Chenxia Li |
MobiSys | 1 |
| 2018 | An efficient resource-optimized learning prefetcher for solid state drivesabstractIn recent years, solid-state drives (SSDs) have been widely deployed in modern storage systems. To increase the performance of SSDs, prefetchers for SSDs have been designed both at operating system (OS) layer and flash translation layer (FTL). Prefetchers in FTL have many advantages like OS-independence, easy-using, and compatibility. However, due to the limitation of computing capabilities and memory resources, existing prefetchers in FTL merely employ simple sequential prefetching which may incur high penalty cost for I/O access stream with complex patterns. In this paper, an efficient learning prefetcher implemented in FTL is proposed. Considering the resource limitation of SSDs, a learning algorithm based on Markov chains is employed and optimized so that high hit ratio and low penalty cost can be achieved even for complex access patterns. To validate our design, a simulator with the prefetcher is designed and implemented based on Flashsim. The TPC-H benchmark and an application launch trace are tested on the simulator. According to experimental results of the TPC-H benchmark, more than 90% of memory cost can be saved in comparison with a previous design at OS layer. The hit ratio can be increased by 24.1% and the number of times of misprefetching can be reduced by 95.8% in comparison with the simple sequential prefetching strategy. Xi Jin 0002, Linfeng Tao, Shuaizhi Guo, Zikun Xiang, Teng Tian |
DATE | 3 |