Jinming Lu

dblp:224/1175 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HoloLUT: An Efficient LUT-Based Engine via Holistic Data Processing for Low-bit LLM Inference
abstract
Weight-only quantization enhances the efficiency of large language models (LLMs) by storing weights in low-bit integers while retaining activations at a higher precision. However, the absence of efficient mixed-precision computation support on general-purpose hardware tends to hinder the potential computational gains from low-bit LLMs. While lookup table (LUT)-based methods offer a promising alternative, conventional bit-serial architectures introduce shift-and-accumulate bottlenecks, limiting throughput and energy efficiency. To overcome these limitations, we propose HoloLUT, a novel LUT-based engine that incorporates the Unitary Data Operation paradigm. This paradigm processes weights holistically rather than bit-serially, thereby eliminating shift-and-accumulate operations. Furthermore, a precision-adaptive mapping strategy combined with a unified LUT generator allows HoloLUT to flexibly and efficiently handle various precisions with negligible hardware overhead. Implemented in 28nm CMOS technology, HoloLUT achieves 1.86 × and 2.18 × improvements in area and power efficiency, respectively, compared to state-of-the-art LUT-based accelerators, demonstrating its strong potential for deploying low-bit LLMs in resource-constrained scenarios.
Hui Wang 0083, Weize Ma, Jinming Lu, Jun Lin 0001
ACM Great Lakes Symposium on VLSI3
2026 Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
abstract
Transformer models have achieved state-of-the-art performance across a wide range of machine learning tasks. There is growing interest in training transformers on resource-constrained edge devices due to considerations such as privacy, domain adaptation, and on-device scientific machine learning. However, the significant computational and memory demands required for transformer training often exceed the capabilities of an edge device. Leveraging low-rank tensor compression, this paper presents the first on-FPGA accelerator for transformer training. On the algorithm side, we present a bi-directional contraction flow for tensorized transformer training, significantly reducing the computational FLOPS and intra-layer memory costs compared to existing tensor operations. On the hardware side, we store all highly compressed model parameters and gradient information on chip, creating an on-chip-memory-only framework for each stage in training. This reduces off-chip communication and minimizes latency and energy costs. Additionally, we implement custom computing kernels for each training stage and employ intra-layer parallelism and pipe-lining to further enhance run-time and memory efficiency. Through experiments on transformer models within 36.7 to 93.5 MB using FP-32 data formats on the ATIS dataset, our tensorized FPGA accelerator could conduct single-batch end-to-end training on the AMD Alevo U50 FPGA, with a memory budget of less than 6-MB BRAM and 22.5-MB URAM. Compared to uncompressed training on the NVIDIA RTX 3090 GPU, our on-FPGA training achieves a memory reduction of 30× to 51×. Our FPGA accelerator also achieves up to 4.0× less energy cost per epoch compared with tensor transformer training on an NVIDIA RTX 3090 GPU. As an initial result, this work highlights the significant potential of large-scale tensor training on edge devices.
Jinming Lu, Hai Li 0008, Cong Hao, Ian A. Young, Zheng Zhang 0005
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 WiFlow: A Precision-Scalable DNN Training Accelerator Through Winograd Algorithm and Dataflow Co-Design
abstract
To address performance degradation from the domain shift and to support user-specific services while considering privacy, security, and communication overhead, there is an urgent need for efficient on-device training accelerators for deep neural networks (DNNs). Given limited computing resources and battery capacity constraints, implementing complex DNN training on edge devices is extremely challenging. To address these issues, we introduce a Winograd-Integrated Gradient Optimization Framework (WIGOF) for cross-phase operand sharing in the Winograd domain, which significantly reduces the number of multiplications and additions. Additionally, we develop WiFlow, an efficient, precision-scalable on-device training accelerator, minimizing area and power overheads of the dedicated Winograd transformation unit. The WiFlow supports 16-bit floating point (FP16), 16-bit brain floating point (BF16), and 8 and 4-bit fixed point (INT8 and INT4), demonstrating scalable improvements in both computational throughput (TOPS) and energy efficiency (TOPS/W) at low precision. A novel data rearrangement pattern, named channel augmentation, addresses the imperfect decomposition to enhance the utilization of processing element units. Furthermore, we propose a Winograd interleaved block-execution dataflow (WInBlock), along with Hierarchical Adaptive Reuse Memory Optimization (HARM) to improve data reuse and reduce both the amount of DRAM and SRAM access. The end-to-end training of WiFlow is achieved on Xilinx XCVU440 FPGA. WiFlow is also synthesized with a 28nm CMOS technology, achieving an area efficiency of 624 GOPS/mm2and an energy efficiency of 4.4 TOPS/W at a supply voltage of 0.9V and an operating frequency of 500 MHz. WiFlow accomplishes$7.75\times $higher area efficiency and$2.91\times $higher energy efficiency in actual DNN training compared with the state-of-the-art on-device training accelerators.
Hui Wang 0083, Jinming Lu, Weize Ma, Zhongfeng Wang 0001, Jun Lin 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion Model
abstract
Fine-tuning large diffusion models for custom applications demands substantial power and time, which poses significant challenges for efficient implementation on mobile devices. In this paper, we develop a novel training accelerator specifically for Low-Rank Adaptation (LoRA) of diffusion models, aiming to streamline the process and reduce computational complexity. By leveraging a fully quantized training scheme for LoRA fine-tuning, we achieve substantial reductions in memory usage and power consumption while maintaining high model fidelity. The proposed accelerator features flexible dataflow, enabling high utilization for irregular and variable tensor shapes during the LoRA process. Experimental results show up to 1.81× training speedup and 5.50× energy efficiency improvements compared to the baseline, with minimal impact on image generation quality.
Jinming Lu, Minghao She, Wendong Mao, Zhongfeng Wang 0001
ISCAS1
2025 A DeepSeek-powered locally deployed closed-loop system for enhancing quality control in electronic nursing documentation: development and clinical validation
abstract
OBJECTIVES: To develop a locally deployed DeepSeek-powered closed-loop system for electronic nursing documentation quality control (QC) and evaluate its clinical efficacy through a multidimensional validation framework. MATERIALS AND METHODS: We implemented a three-dimensional (3D) QC framework (real-time, final, and vertical QC). A retrospective analysis of 556 electronic nursing records was conducted to evaluate pre- and postimplementation outcomes, with documentation accuracy and audit efficiency assessed via blinded nurse evaluations. RESULTS: After implementation, omission rates decreased from 7.19% to 1.79%, the prevalence of logical inconsistencies decreased from 9.35% to 0.72%, and the prevalence of timeliness errors decreased from 8.63% to 0%. The QC time per record decreased by 3.2-fold. Nurse satisfaction was evaluated using the Clinical Nursing Information System Effectiveness Evaluation Scale (Zhao Y, Gu Y, Zhang X, et al. Developed the clinical nursing information system effectiveness evaluation scale based on the new D&M model and conducted reliability and validity evaluation. Chin J Prae Nurs. 2020;36:544-550. https://doi.org/10.3760/cma.j.issn.1672-7088.2020.07.013), yielding a total score of 102.73 ± 3.25 out of a maximum 115 points. DISCUSSION: This study demonstrates that the Artificial Intelligence (AI)-powered closed-loop QC system significantly enhances documentation accuracy and workflow efficiency while ensuring data security. The 3D framework (real-time, final, and vertical QC) represents a paradigm shift from reactive to proactive quality governance in nursing practice. High nurse satisfaction (102.73/115) confirms clinical viability, offering a scalable model for intelligent health-care quality ecosystems. Future work should explore federated learning for multicenter deployment and regulatory frameworks for clinical AI. CONCLUSION: DeepSeek demonstrated robust efficacy in enhancing QC accuracy and workflow efficiency, with localized deployment ensuring data security. This system redefines nursing documentation management, heralding an era of "intelligent negative feedback" in health-care quality ecosystems.
Jinhong Lv, Mengzhu Jiang, Yuanhao Lv, Jialu Sun, Jinming Lu, Hongru Wang 0013
J. Am. Medical Informatics Assoc.6
2024 WinTA: An Efficient Reconfigurable CNN Training Accelerator With Decomposition Winograd
abstract
Convolutional neural networks (CNNs) are expected to bridge the domain shift between the training data and real-world tasks. Moreover, the efficient training of CNNs on resource-constrained platforms has become more important because of communication latency and privacy concerns. However, deploying CNN training on edge devices is challenging due to the intensive computation and diverse computational patterns. In this work, we firstly propose a hybrid decomposition Winograd (HDW) method that significantly reduces the number of multiplications and flexibly handles various convolution operations during training. Secondly, we design a reconfigurable CNN training accelerator, named WinTA, utilizing a set of unified transformation units to support various Winograd operations. Thirdly, we implement an efficient and flexible data access scheme using a hierarchical barrel shifter network (HBSN). Experimental results on the Xilinx Alveo U50 FPGA Card demonstrate that WinTA effectively accelerates CNN training. Compared to CPU and GPU implementations, WinTA achieves speedups of 7.1 texttimes and 1.65 texttimes, respectively, while improving energy efficiency by 26.6 texttimes and 10.4 texttimes, respectively. Additionally, our design provides 1.24 texttimes and 2.04 texttimes improvements in terms of throughput and resource efficiency compared to prior-art FPGA-based training accelerator.
Jinming Lu, Hui Wang 0083, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 ETA: An Efficient Training Accelerator for DNNs Based on Hardware-Algorithm Co-Optimization
abstract
Recently, the efficient training of deep neural networks (DNNs) on resource-constrained platforms has attracted increasing attention for protecting user privacy. However, it is still a severe challenge since the DNN training involves intensive computations and a large amount of data access. To deal with these issues, in this work, we implement an efficient training accelerator (ETA) on field-programmable gate array (FPGA) by adopting a hardware-algorithm co-optimization approach. A novel training scheme is proposed to effectively train DNNs using 8-bit precision with arbitrary batch sizes, in which a compact but powerful data format and a hardware-oriented normalization layer are introduced. Thus the computational complexity and memory accesses are significantly reduced. In the ETA, a reconfigurable processing element (PE) is designed to support various computational patterns during training while avoiding redundant calculations from nonunit-stride convolutional layers. With a flexible network-on-chip (NoC) and a hierarchical PE array, computational parallelism and data reuse can be fully exploited, and memory accesses are further reduced. In addition, a unified computing core is developed to execute auxiliary layers such as normalization and weight update (WU), which works in a time-multiplexed manner and consumes only a small amount of hardware resources. The experiments show that our training scheme achieves the state-of-the-art accuracy across multiple models, including CIFAR-VGG16, CIFAR-ResNet20, CIFAR-InceptionV3, ResNet18, and ResNet50. Evaluated on three networks (CIFAR-VGG16, CIFAR-ResNet20, and ResNet18), our ETA on Xilinx VC709 FPGA achieves 610.98, 658.64, and 811.24 GOPS in terms of throughput, respectively. Compared with the prior art, our design demonstrates a speedup of 3.65× and an energy efficiency improvement of 8.54× on CIFAR-ResNet20.
Jinming Lu, Chao Ni 0005, Zhongfeng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 An Efficient Training Accelerator for Transformers With Hardware-Algorithm Co-Optimization
abstract
Transformers have achieved significant success in deep learning, and training Transformers efficiently on resource-constrained platforms has been attracting continuous attention for domain adaptions and privacy concerns. However, deploying Transformers training on these platforms is still challenging due to its dynamic workloads, intensive computations, and massive memory accesses. To address these issues, we propose an Efficient Training Accelerator for TRansformers (TRETA) through a hardware-algorithm co-optimization strategy. First, a hardware-friendly mixed-precision training algorithm is presented based on a compact and efficient data format, which significantly reduces the computation and memory requirements. Second, a flexible and scalable architecture is proposed to achieve high utilization of computing resources when processing arbitrary irregular general matrix multiplication (GEMM) operations during training. These irregular GEMMs lead to severe under-utilization when simply mapped on traditional systolic architectures. Third, we develop training-oriented architectures for the crucial Softmax and layer normalization functions in Transformers, respectively. These area-efficient modules have unified and flexible microarchitectures to meet various computation requirements of different training phases. Finally, TRETA is implemented under Taiwan Semiconductor Manufacturing Company (TSMC) 28-nm technology and evaluated on multiple benchmarks. The experimental results show that our training framework achieves the same accuracy as the full precision baseline. Moreover, TRETA can achieve 14.71 tera operations per second (TOPS) and 3.31 TOPS/W in terms of throughput and energy efficiency, respectively. Compared with prior arts, the proposed design shows 1.4–$24.5\times $speedup and 1.5–$25.4\times $energy efficiency improvement.
Haikuo Shao, Jinming Lu, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2022 An Efficient Hardware Architecture for DNN Training by Exploiting Triple Sparsity
abstract
Recently, on-device DNN training has attracted much attention due to its high performance on edge devices and great ability to protect user privacy. Low-power and high throughput implementations of DNN training are highly desired for resource-limited devices. In this paper, we present an efficient hardware accelerator that exploits triple sparsity to reduce the number of unnecessary operations during DNN training. The gradients pruning algorithm is employed to bring error sparsity. Firstly, sparse data are represented in compressed sparse block format, which is suitable for different memory access patterns in all training phases. Secondly, an efficient sparsity detection logic based on the aforementioned data storage format is proposed, which adopts a 2-level grained mechanism. Coarse-grained mask-matching units are reused to improve the energy efficiency, while fine-grained mask-matching units make PEs work independently to enhance throughput. Thirdly, based on the above sparsity detection logic, we propose an efficient architecture for DNN training. Experimental results show that our design can achieve up to 42.1 TOPS and 174.0 TOPS/W in terms of throughput and energy efficiency, respectively. The energy efficiency of our design is $2.12\times $ higher than the state-of-the-art training processor. For training a ResNet-50 model on the CIFAR10 dataset, the energy efficiency of our design achieves 14.10, 96.57, and 84.43 TOPS/W in the FP, BP, and WG phases, respectively.
Jinming Lu, Zhongfeng Wang 0001
ISCAS2
2022 THETA: A High-Efficiency Training Accelerator for DNNs With Triple-Side Sparsity Exploration
abstract
Training deep neural networks (DNNs) on edge devices has attracted increasing attention in real-world applications for domain adaption and privacy protection. However, deploying DNN training on resource-limited edge devices is challenging as there are massive computations and data transportation in training. To address this issue, we propose an energy-efficient training accelerator in this work by employing a hybrid compression strategy. Here, various data redundancies are fully exploited, and the real triple-side sparsity is achieved. Hence, the computational complexity is drastically reduced with negligible accuracy loss across a range of transfer learning tasks. To facilitate triple-side zero-skipping operations during different training stages, we first present a novel sparse data representation and a triple-sparsity index matching scheme. Second, a sparse tensor processing unit (STPU) arranged in a hierarchical structure is developed, which enables a flexible dataflow to process convolutional (Conv) and fully connected (FC) layers with diverse computational patterns throughout the entire training. Third, an auxiliary processing unit (APU) is designed to execute some postprocessing operations, such as rectified linear unit (ReLU) and on-the-fly pruning. Finally, the training accelerator is implemented under Taiwan Semiconductor Manufacturing Company (TSMC) 28-nm process and evaluated on multiple benchmarks. The experimental results show that THETA achieves 7.28–22.32 tera operations per second (TOPS) and 45.24–133.70 TOPS/W in performance and energy efficiency, reducing 40–$72\times $training time and 19–$63\times $energy consumption over dense training, respectively. Compared with the prior art, our design offers$1.6\times $throughput and$1.9\times $energy efficiency, respectively.
Jinming Lu, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2021 Evaluations on Deep Neural Networks Training Using Posit Number System
abstract
The training of Deep Neural Networks (DNNs) brings enormous memory requirements and computational complexity, which makes it a challenge to train DNN models on resource-constrained devices. Training DNNs with reduced-precision data representation is crucial to mitigate this problem. In this article, we conduct a thorough investigation on training DNNs with low-bit posit numbers, a Type-III universal number (Unum). Through a comprehensive analysis of quantization with various data formats, it is demonstrated that the posit format shows great potential to be employed in the training of DNNs. Moreover, a DNN training framework using 8-bit posit is proposed with a novel tensor-wise scaling scheme. The experiments show the same performance as the state-of-the-art (SOTA) across multiple datasets (MNIST, CIFAR-10, ImageNet, and Penn Treebank) and model architectures (LeNet-5, AlexNet, ResNet, MobileNet-V2, and LSTM). We further design an energy-efficient hardware prototype for our framework. Compared to the standard floating-point counterpart, our design achieves a reduction of 68, 51, and 75 percent in terms of area, power, and memory capacity, respectively.
Jinming Lu, Chao Fang 0005, Mingyang Xu, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Computers1