Jiajun Wu 0006

dblp:285/1532 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-2477-3553ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2025 TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
abstract
Modern transformer-based deep neural networks present unique technical challenges for effective acceleration in real-world applications. Apart from the vast amount of linear operations needed due to their sizes, modern transformer models are increasingly reliance on precise non-linear computations that make traditional low-bitwidth quantization methods and fixed-dataflow matrix accelerators ineffective for end-to-end acceleration. To address this need to accelerate both linear and non-linear operations in a unified and programmable framework, this article introduces TATAA. TATAA employs 8-bit integer ( int8 ) arithmetic for quantized linear layer operations through post-training quantization, while it relies on bfloat16 floating-point arithmetic to approximate non-linear layers of a transformer model. TATAA hardware features a transformable arithmetic architecture that supports both formats during runtime with minimal overhead, enabling it to switch between a systolic array mode for int8 matrix multiplications and a SIMD mode for vectorized bfloat16 operations. An end-to-end compiler is presented to enable flexible mapping from emerging transformer models to the proposed hardware. Experimental results indicate that our mixed-precision design incurs only 0.14% to 1.16% accuracy drop when compared with the pre-trained single-precision transformer models across a range of vision, language, and generative text applications. Our prototype implementation on the Alveo U280 FPGA currently achieves 2,935.2 GOPS throughput on linear layers and a maximum of 189.5 GFLOPS for non-linear operations, outperforming related works by up to \(1.45\times\) in end-to-end throughput and \(2.29\times\) in DSP efficiency, while achieving \(2.19\times\) higher power efficiency than modern NVIDIA RTX4090 GPU.
Jiajun Wu 0006, Mo Song, Jingmin Zhao, Yizhao Gao 0002, Jia Li 0057, Hayden Kwok-Hay So
ACM Trans. Reconfigurable Technol. Syst.1
2024 DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network Inference
abstract
To accelerate the inference of deep neural networks (DNNs), quantization with low-bitwidth numbers is actively researched. A prominent challenge is to quantize the DNN models into low-bitwidth numbers without significant accuracy degradation, especially at very low bitwidths (< 8 bits). This work targets an adaptive data representation with variablelength encoding called DyBit. DyBit can dynamically adjust the precision and range of separate bit-fields to be adapted to the DNN weights/activations distribution. We also propose a hardware-aware quantization framework with a mixed-precision accelerator to trade-off the inference accuracy and speedup. Experimental results demonstrate that the ImageNet inference accuracy via DyBit is 1.97% higher than the state-of-the-art at 4-bit quantization, and the proposed framework can achieve up to 8.1× speedup compared with the original ResNet-50 model.
Jiajun Zhou 0004, Jiajun Wu 0006, Yizhao Gao 0002, Yuhao Ding, Chaofan Tao, Fengbin Tu, Kwang-Ting Cheng, Hayden Kwok-Hay So, Ngai Wong 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 MSD: Mixing Signed Digit Representations for Hardware-efficient DNN Acceleration on FPGA with Heterogeneous Resources
abstract
By quantizing weights with different precision for different parts of a network, mixed-precision quantization promises to reduce the hardware cost and improve the speed of deep neural network (DNN) accelerators that typically operate with a fixed quantization scheme. However, the additional control needed, and the decreased hardware efficiency arising from multi-precision operations have made mixed-precision quantization schemes challenging to deploy in practice. In this paper, a practical mixed-precision quantization framework called MSD that leverages the heterogeneous computing resources on FPGA to perform bit-serial and bit-parallel operations simultaneously is presented. MSD combines the use of a custom restricted signed digit (RSD) representation, which utilizes a limited number of effectual bits, and the conventional 2's complement representation to quantize DNN weights. Depending on the availability of fine-grained and coarse-grained resources, MSD encodes a subset of weights with RSD to allow highly efficient bit-serial multiply-accumulate implementation using LUT resources. Furthermore, the number of effectual bits used in RSD is optimized to match the bit-serial hardware latency to the bit-parallel operation on the coarse-grained resources to ensure the highest run-time utilization of all on-chip resources. Experiments show that MSD achieved a 1.36× speedup on the ResNet-18 model over the state-of-the-art, and a remarkable 4.91% higher accuracy on MobileNet-V2.
Jiajun Wu 0006, Jiajun Zhou 0004, Yizhao Gao 0002, Yuhao Ding, Ngai Wong 0001, Hayden Kwok-Hay So
FCCM1
2023 Model-Platform Optimized Deep Neural Network Accelerator Generation through Mixed-Integer Geometric Programming
abstract
Although there are distinct power-performance advantages in customizing an accelerator for a specific combination of FPGA platform and neural network model, developing such highly customized accelerators is a challenging task due to the massive design space spans from the range of network models to be accelerated, the target platform's compute capability, and its memory capacity and performance characteristics. To address this architectural customization problem, an automatic design space exploration (DSE) framework using a mixed-integer geometric programming (MIGP) approach is presented. Given the set of DNN models to be accelerated and a generic description of the target platform's compute and memory capabilities as input, the proposed framework automatically customizes an architectural template for the platform-model combination and produces the associated I/O schedule to maximize its end-to-end performance. By formulating DNN inference as a multi-level loop tiling problem, the proposed framework first customizes an accelerator template that consists of a parameterizable array architecture with SIMD execution cores and a customizable memory hierarchy using a MIGP to maximize the expected resource utilization. Subsequently, a second MIGP is used to schedule memory and compute operations as tiles to improve on-chip data reuse and memory bandwidth utilization. Experimental results from a wide range of neural network models and FPGA platform combinations show that the proposed scheme is able to produce accelerators with performance comparable to the state-of-the-art. The proposed DSE framework and the resulting hardware/software generator are available as an open-source package called AGNA with the hope that it may facilitate vendor-agnostic DNN accelerator development from the research community in the future.
Yuhao Ding, Jiajun Wu 0006, Yizhao Gao 0002, Maolin Wang 0002, Hayden Kwok-Hay So
FCCM2
2022 Energy-Efficient Intelligent Pulmonary Auscultation for Post COVID-19 Era Wearable Monitoring Enabled by Two-Stage Hybrid Neural Network
abstract
This paper proposes an energy-efficient intelligent pulmonary auscultation system for post COVID-19 era wearable monitoring. This system consists of a tightly coupled two-stage hybrid neural network (TC-TSHNN) model and a corresponding multi-task training paradigm to improve prediction accuracy and generalization ability based on the fact that the number of COVID-19 patients is far less than that of normal people. At the first stage, two-category coarse classification is performed to identify normal and abnormal lung sounds. If the lung sound is abnormal, the second stage would be triggered to perform a four-category fine-grained classification. Besides, discrete wavelet transform is utilized for feature extraction, denoising and data reduction. In addition, advanced lightweight convolutional neural networks are used to reduce the model’s computation and improve the model’s performance. The hybrid network model can achieve 92% computation reduction and energy saving compared with a direct four-category classification when the input lung sound is normal, which is the majority of cases. Experiment results with inter-patient classification on the COVID-19 lung sound dataset from Tongji Hospital in Wuhan City and the ICBHI’17 dataset show that the proposed TC-TSHNN model can significantly reduce power consumption while maintaining competitive performance against the state-of-the-art work.
Bingqiang Liu, Ziyuan Wen, Hongling Zhu, Jinsheng Lai, Jiajun Wu 0006, Heng Ping, Wenqing Liu, Guoyi Yu, Zuozhu Liu, Hesong Zeng, Chao Wang 0096
ISCAS5
2022 In Situ Aging-Aware Error Monitoring Scheme for IMPLY-Based Memristive Computing-in-Memory Systems
abstract
Stateful logic through memristor is a promising technology to build Computing-in-Memory (CIM) systems. However, aging-induced degradation of memristors’ threshold voltage imposes a major challenge to the reliability and guardbands estimation of memristive CIM systems, especially the Material Implication (IMPLY) logic based CIM systems. In this paper, a novel in-situ aging-aware error monitoring scheme for memristor-based IMPLY logic is proposed. The proposed in-situ error monitoring scheme can achieve faster error detection speed and higher detection accuracy than the straightforward program-verify monitoring scheme. Simulation results under Monte-Carlo simulation show that the proposed monitoring scheme can effectively detect the major operation failures existing in IMPLY logic operations with a detection accuracy up to 99.95%. Moreover, a case study of error monitoring design of 4-bit IMPLY-based adder is carried out. The analysis result exhibits that the proposed in-situ monitoring scheme can achieve 75.2% improvement on the detection speed against the program-verify scheme. Further analysis on a convolution filter in VGG-11 based Binarized Neural Network shows that 74% improvement on the detection speed can also be achieved by using the proposed monitoring scheme, which suggests that the proposed in-situ error monitoring scheme is an efficient solution to improve the reliability of IMPLY-based memristive CIM systems.
Jiajun Wu 0006, Xinglong Ji, Guoyi Yu, Chao Wang 0096
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 Efficient Design of Spiking Neural Network With STDP Learning Based on Fast CORDIC
abstract
In emerging Spiking Neural Network (SNN) based neuromorphic hardware design, energy efficiency and on-line learning are attractive advantages mainly contributed by bio-inspired local learning with nonlinear dynamics and at the cost of associated hardware complexity. This paper presents a novel SNN design employing fast COordinate Rotation DIgital Computer (CORDIC) algorithm to achieve fast spike timing–dependent plasticity (STDP) learning with high hardware efficiency. In this study, a system design and evaluation method of CORDIC-based SNN is proposed for finding optimal CORDIC type and precision, from theoretical CORDIC-level error to application-level learning performance. From the proposed design and evaluation method, a reconfigurable SNN design based on fast-convergence CORDIC is designed to achieve high classification accuracy on MNIST, fast on-line learning and good energy efficiency. By utilizing SNN’s fault tolerance and time-division-multiplexing (TDM) strategy, the reconfigurable SNN design employs 8-bit fast-convergence CORDIC and TDM-based hardware accelerator for high efficiency. FPGA implementation results confirm that the proposed fast-convergence CORDIC SNN design outperforms the state-of-the-art CORDIC method by 38.5%−45.3% in terms of learning speed and energy efficiency, with the STDP learning of 30.2 ns/SOP, energy efficiency of 176.6 pJ/SOP, processing speed of 6.1 ms/image, and on-line learning convergence of 21.4 s (time to reach the final accuracy, on average), on MNIST benchmark.
Jiajun Wu 0006, Zixuan Peng, Xinglong Ji, Guoyi Yu, Chao Wang 0096
IEEE Trans. Circuits Syst. I Regul. Pap.1