Hai Li 0008

dblp:30/5330-8 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0001-7668-569XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
abstract
Transformer models have achieved state-of-the-art performance across a wide range of machine learning tasks. There is growing interest in training transformers on resource-constrained edge devices due to considerations such as privacy, domain adaptation, and on-device scientific machine learning. However, the significant computational and memory demands required for transformer training often exceed the capabilities of an edge device. Leveraging low-rank tensor compression, this paper presents the first on-FPGA accelerator for transformer training. On the algorithm side, we present a bi-directional contraction flow for tensorized transformer training, significantly reducing the computational FLOPS and intra-layer memory costs compared to existing tensor operations. On the hardware side, we store all highly compressed model parameters and gradient information on chip, creating an on-chip-memory-only framework for each stage in training. This reduces off-chip communication and minimizes latency and energy costs. Additionally, we implement custom computing kernels for each training stage and employ intra-layer parallelism and pipe-lining to further enhance run-time and memory efficiency. Through experiments on transformer models within 36.7 to 93.5 MB using FP-32 data formats on the ATIS dataset, our tensorized FPGA accelerator could conduct single-batch end-to-end training on the AMD Alevo U50 FPGA, with a memory budget of less than 6-MB BRAM and 22.5-MB URAM. Compared to uncompressed training on the NVIDIA RTX 3090 GPU, our on-FPGA training achieves a memory reduction of 30× to 51×. Our FPGA accelerator also achieves up to 4.0× less energy cost per epoch compared with tensor transformer training on an NVIDIA RTX 3090 GPU. As an initial result, this work highlights the significant potential of large-scale tensor training on edge devices.
Jinming Lu, Hai Li 0008, Cong Hao, Ian A. Young, Zheng Zhang 0005
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 Enhanced Operator Learning for Scalable and Ultra-fast Thermal Simulation in 3D-IC Design
abstract
Thermal simulation plays a critical role in the design of 3D integrated circuits (3D-ICs), where accurate and efficient temperature predictions are essential to ensure component reliability and performance. Recently, deep learning methods have shown great potential in accelerating these simulations. DeepOHeat [1] is one such approach, designed to learn solution operators that map single or multiple configurations of heat equations---such as surface power, volume power, or heat transfer coefficients---directly to the 3D temperature distribution. By utilizing a physics-informed DeepONet framework[2], DeepOHeat effectively captures complex relationships between design parameters and temperature fields, even when no training data is available and only PDE constraints are imposed.
Xinling Yu, Ziyue Liu 0003, Hai Li 0008, Ian A. Young, Zheng Zhang 0005
ASP-DAC3
2025 Poor Man's Training on MCUs: A Memory-Efficient Quantized Back-Propagation-Free Approach
abstract
Back propagation (BP) is the default solution for gradient computation in neural network training. However, implementing BP-based training on various edge devices such as FPGA, microcontrollers (MCUs), and analog computing platforms faces multiple major challenges, such as the lack of hardware resources, long time-to-market, and dramatic errors in a low-precision setting. This article presents a simple BP-free training scheme on an MCU, which makes edge training hardware design as easy as inference hardware design. We adopt a quantized zeroth-order method to estimate the gradients of quantized model parameters, which can overcome the error of a straight-through estimator in a low-precision BP scheme. We further employ a few dimension reduction methods (e.g., node perturbation, sparse training) to improve the convergence of zeroth-order training. Experiment results show that our BP-free training achieves comparable performance as BP-based training on adapting a pre-trained image classifier to various corrupted data on resource-constrained edge devices (e.g., an MCU with 1024-KB SRAM for dense full-model training, or an MCU with 256-KB SRAM for sparse training). This method is most suitable for application scenarios where memory cost and time-to-market are the major concerns, but longer latency can be tolerated.
Yequan Zhao, Hai Li 0008, Ian A. Young, Zheng Zhang 0005
ACM Trans. Design Autom. Electr. Syst.2
2024 Digital CIM with Noisy SRAM Bit: A Compact Clustered Annealer for Large-Scale Combinatorial Optimization
abstract
Combinatorial optimization problems (COP) are NP-hard and intractable to solve using conventional computing. The Ising model-based annealer has gained increasing attention recently due to its efficiency and speed in finding approximate solutions. However, Ising solvers for travelling salesman problems (TSP) usually suffer from a scalability issue due to quadratically increasing number of spins. In this paper, we propose a digital computing-in-memory (CIM) based clustered annealer to solve tens of thousands of city-scale TSP with only a few mega-byte (MB) of static random access memory (SRAM), using hierarchical clustering to solve input sparsity and digital CIM flexibility to solve weight sparsity. The intrinsic process variations between SRAM devices are utilized to generate the noisy bit errors during pseudo-read under reduced supply voltage, realizing the annealing process. The design space of cluster size and programmability is explored to understand the trade-offs of solution quality and hardware cost, for TSP scale ranging from 3080 to 85900 cities. The proposed design speeds up the convergence by >109× with <25% solution quality overhead compared with the CPU baseline. The comparison with state-of-the-art scalable annealers shows a >1013× improvement on functionally normalized area and power.
Anni Lu, Yuan-Chun Luo, Hai Li 0008, Ian A. Young, Shimeng Yu
DAC4