EDBT 2026 Demo / reviewers in the wild / expert
Letian Zhao
dblp:268/1874
· DBLP profile ↗
8ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 6 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACsabstractGeneral matrix-matrix multiplication (GEMM), serving as a cornerstone of AI computations, has positioned tensor processing engines (TPEs) as increasingly critical components within existing GPUs and domain-specific architectures (DSA). Our analysis identifies that the prevailing architectures primarily focus on dataflow or operand reuse strategies, when considering the combination of matrix multiplication with multiply-accumulator (MAC) itself, it provides greater optimization space for the design of TPEs. This work introduces a novel perspective on matrix multiplication from a hardware standpoint, focusing on the bit-weight dimension of MACs. Through this lens, we propose a finer-grained TPE notation, using matrix triple loops as an example, introducing new methods and ideas for designing and optimizing PE microarchitecture. Based on the new notation and transformations, we propose four optimization techniques that achieve varying degrees of improvement in timing, area, and power consumption. We implement our design in RTL using the SMIC-28nm process. Applying our methods to four classic TPE architectures (include systolic array [20], 3D-Cube [27], multiplier-adder tree [48], and 2D-Matrix [30]), we achieved area efficiency improvements of $1.27 \times, 1.28 \times, 1.56 \times$, and $1.44 \times$, and $1.04 \times, 1.56 \times, 1.49 \times$, and $1.20 \times$ for energy efficiency respectively. When applied to a bit-slice architecture, we achieved a $12.10 \times$ improvement in energy efficiency and $2.85 \times$ in area efficiency compared to Laconic [38]. Our Verilog HDL code, along with timing, area, and power reports for circuit synthesis in URL: https://github.com/wqzustc/High-Performance-Tensor-Processing-Engines. Qizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng 0002, Zerong He, Linfeng Tao, Letian Zhao, Zhaoxi Zeng, Wei Yuan 0006, Xi Jin 0002 |
HPCA | 8 |
| 2025 | SageSC: Accelerating GraphSAGE Minibatch Inference on Memory-Intensive GraphsabstractGraph neural networks demonstrate excellent performance on node classification tasks in graph datasets. For inference tasks on memory-intensive graphs, the storage burden, memory access bottlenecks, and load imbalance issues arise. The minibatch inference proposed in GraphSAGE is an effective method for minimizing these problems. However, minibatch inference introduces new challenges: while it facilitates subsequent computation, the irregular random memory access pressure shifts to the minibatch construction phase, creating performance bottlenecks in the system. In this work, to address the aforementioned challenges, we propose a novel scattered minibatch construction and aggregation (SMCA) algorithm to optimize sampling, batch construction, and aggregation computations for minimizing their latency. This method distributes memoryintensive workloads and exploits the parallelism between memory groups. Evaluation results show that the proposed accelerator SageSC achieves speedups ranging from 3x to 96x compared to CPU/GPU baselines, especially on memory-intensive graphs, while outperforming existing state-of-the-art designs. Yuchen Gui, Wei Yuan 0006, Qizhe Wu, Huawen Liang, Letian Zhao, Linfeng Tao, Zhongguang Xu, Xi Jin 0002 |
ICCD | 5 |
| 2025 | MHE-TPE: Multi-Operand High-Radix Encoder for Mixed-Precision Fixed-Point Tensor Processing Engines
Qizhe Wu, Jinyi Zhou, Zhanhe Hu, Zhichen Zeng 0002, Huawen Liang, Jiuru Zhu, Linfeng Tao, Xin Zhang 0176, Zekang Cheng, Letian Zhao, Wei Yuan 0006, Xi Jin 0002 |
MICRO | 10 |
| 2024 | Efficient Message Passing Architecture for GCN Training on HBM-based FPGAs with Orthogonal Topology On-Chip NetworksabstractGraph Convolutional Networks (GCNs) are state-of-the-art deep learning models for representation learning on graphs. However, the efficient training of GCNs is hampered by constraints in memory capacity and bandwidth, compounded by the irregular data flow that results in communication bottlenecks. To address these challenges, we propose a message-passing architecture that leverages NUMA-based memory access properties and employs a parallel multicast routing algorithm based on a 4-D hypercube network within the accelerator for efficient message passing in graphs. Additionally, we have re-engineered the backpropagation algorithm specific to GCNs within our proposed accelerator. This redesign strategically mitigates the memory demands prevalent during the training phase and diminishes the computational overhead associated with the transposition of extensive matrices. Compared to the state-of-the-art HP-GNN architecture we achieved a performance improvement of 1.03×~1.81×. Qizhe Wu, Letian Zhao, Yuchen Gui, Huawen Liang, Xi Jin 0002 |
FPGA | 2 |
| 2023 | MCANet: Multiscale Cross-Modality Attention Network for Multispectral Pedestrian Detection
Letian Zhao, Xi Jin 0002 |
MMM (1) | 2 |
| 2022 | FP-GNN: Adaptive FPGA accelerator for Graph Neural Networks
Teng Tian, Letian Zhao, Qizhe Wu, Wei Yuan 0006, Xi Jin 0002 |
Future Gener. Comput. Syst. | 2 |
| 2022 | G-NMP: Accelerating Graph Neural Networks with DIMM-based Near-Memory Processing
Teng Tian, Letian Zhao, Xuecang Zhang, Fangmin Lu, Xi Jin 0002 |
J. Syst. Archit. | 3 |
| 2020 | Exploration of Memory Access Optimization for FPGA-based 3D CNN AcceleratorabstractThree-dimensional convolutional networks (3D CNNs) are used efficiently in various video recognition applications. Compared to traditional 2D CNNs, extra temporal dimension causes 3D CNNs more computationally intensive and to have a larger memory footprint. Therefore, the memory optimization is extremely crucial in this case. This paper presents a design space exploration of memory access optimization for FPGA-based 3D CNN accelerator. We present a non-overlapping data tiling method for contiguous off-chip memory access and explore on-chip data reuse opportunity by leveraging different loop ordering strategies. We propose a hardware architecture design which can flexibly support different loop ordering strategies for each 3D CNN layer. With the help of hardware/software co-design, we can provide the optimal configuration toward an energy-efficient and high-performance accelerator design. According to the experiments on AlexNet, VGG16, and C3D, our optimal model reduces up to 84% DRAM accesses and 55% energy consumption on C3D compared to a baseline model, and demonstrates state-of-the-art performance compared to prior FPGA implementations. Teng Tian, Xi Jin 0002, Letian Zhao |
DATE | 3 |