Jingbo Jiang

dblp:39/10642 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0003-0268-9844ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploiting the Irregular Input Sparsity in Systolic Array-based DNN Accelerators via Local Soft Pooling
abstract
One promising approach to mitigating the computational complexity of deep neural networks is to leverage the sparsity of input activations that results from the application of the ReLU function. However, the irregular distribution of zero-valued inputs poses a challenge for efficient implementation in existing regular architectures, such as systolic arrays. Previous works usually depend on specialized architectures to bypass the redundant computations during runtime. In contrast to these prior strategies, we propose a local soft pooling method to efficiently exploit the irregular input sparsity in systolic array-based architectures. Through local soft pooling, adjacent input rows can be safely merged at runtime, compressing the sparse input matrix into a compact format that is only 1/3 to 1/2 of its original size. The compact matrix can then be directly fed into the systolic array for computation. A computation saving of 67.78% is achieved across various networks on both CIFAR-10 and ImageNet with negligible accuracy loss. As a result, the throughput and energy efficiency are improved by 2.72 and 2.07 times, respectively.
Desheng Fu, Jingbo Jiang, Jingyang Zhu, Xizi Chen, Chi-Ying Tsui
ASP-DAC2
2024 Data-Pattern-Based Predictive On-Chip Power Meter in DNN Accelerator
abstract
Advanced power management techniques, such as voltage drop mitigation and fast power management, can greatly enhance energy efficiency in contemporary hardware design. Nevertheless, the implementation of these innovative techniques necessitates accurate and fine-grained power modeling, as well as timely responses for effective coordination with the power management unit. Additionally, existing performance-counter-based and RTL-based on-chip power meters have difficulty in providing sufficient response time for fast power and voltage management scenarios. In this article, we propose PROPHET, a data-pattern-based power modeling method for multiply-accumulate-based (MACC) deep neural network (DNN) accelerators. Our proposed power model extracts the predefined data patterns during memory access and then a pretrained power model can predict the dynamic power of the DNN accelerators. Thus, PROPHET can predict dynamic power and provide sufficient responding time for power management units. In the experiments, we evaluate our predictive power model in four DNN accelerators with different dataflows and data types. In power model training and verification, our proposed data-patterns-based power model can realize the 2-cycle temporal resolution with$R^{2} \gt 0.9$, normalized mean absolute error <7%, and the area and power overhead lower than 4.5%.
Tingyuan Liang, Jingbo Jiang, Yipu Zhang 0002, Zhe Lin 0007, Zhiyao Xie, Wei Zhang 0012
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Accelerating Large Kernel Convolutions with Nested Winograd Transformation
abstract
Recent literature has shown that convolutional neural networks (CNNs) with large kernels outperform vision transformers (ViTs) and CNNs with stacked small kernels in many computer vision tasks, such as object detection and image restoration. The Winograd transformation helps reduce the number of repetitive multiplications in convolution and is widely supported by many commercial AI processors. Researchers have proposed accelerating large kernel convolutions by linearly decomposing them into many small kernel convolutions and then sequentially accelerating each small kernel convolution with the Winograd algorithm. This work proposes a nested Winograd algorithm that iteratively decomposes a large kernel convolution into small kernel convolutions and proves it to be more effective than the linear decomposition Winograd transformation algorithm. Experiments show that compared to the linear decomposition Winograd algorithm, the proposed algorithm reduces the total number of multiplications by 1.4 to 10.5 times for computing 4×4 to 31×31 convolutions.
Jingbo Jiang, Xizi Chen, Chi-Ying Tsui
VLSI-SoC1
2023 Tight Compression: Compressing CNN Through Fine-Grained Pruning and Weight Permutation for Efficient Implementation
abstract
The unstructured sparsity after pruning poses a challenge to the efficient implementation of deep learning models in existing regular architectures like systolic arrays. On the other hand, coarse-grained structured pruning is suitable for implementation in regular architectures but tends to have higher accuracy loss than unstructured pruning when the pruned models are of the same size. In this work, we propose a model compression method based on a novel weight permutation scheme to fully exploit the fine-grained weight sparsity in the hardware design. Through permutation, the optimal arrangement of the weight matrix is obtained, and the sparse weight matrix is further compressed to a small and dense format to make full use of the hardware resources. Two pruning granularities are explored. In addition to the unstructured weight pruning, we also propose a more fine-grained subword-level pruning to further improve the compression performance. Compared to the state-of-the-art works, the matrix compression rate is significantly improved from$5.88\times $to$14.13\times $. As a result, the throughput and energy efficiency are improved by 2.75 and 1.86 times, respectively.
Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 A 16-bit Encrypted On-chip Embedded System for Implantable Medical Devices
abstract
A 16-bit on-chip embedded encryption system built upon eFUSE, cipher, hash functions, and EDCs for optical nerve stimulation is presented. The foundry-provided eFUSE IP is modified with a one-shot block to support wireless power transfer operation by mitigating the supply voltage drop problem during sensing to avoid subsequent resetting. Novel logic gate-based auxiliary circuit facilitates different sensing and programming modes in eFUSE. A 128-bit cipher is reduced to 16 bits with cascade structure using the proposed divide-and-conquer algorithm, keeping the cipher strength constant. The developed resource sharing technique reduces the area and the power consumption of the ciper circuits by 2.7 times and 5.1 times, respectively. The whole system with the optical nerve simulator is fabricated with 0.18$\mu$m BCDlite process, and measurement results show the correct encryption operation when powered by wirelessly transferred power.
Sayan Sarkar, Jingbo Jiang, Wing-Hung Ki, Chi-Ying Tsui
ISCAS2
2021 Shell thickening for extrusion-based ceramics printing
Haisen Zhao, Jingbo Jiang, Lin Lu 0001
Comput. Graph.6
2020 A Comparison of the Taguchi Method and Evolutionary optimization in Multivariate Testing
abstract
Multivariate testing has recently emerged as a promising technique in web interface design. In contrast to the standard A/B testing, multivariate approach aims at evaluating a large number of values in a few key variables systematically. The Taguchi method is a practical implementation of this idea, focusing on orthogonal combinations of values. It is the current state of the art in applications such as Adobe Target. This paper evaluates an alternative method: population-based search, i.e. evolutionary optimization. Its performance is compared to that of the Taguchi method in several simulated conditions, including an orthogonal one designed to favor the Taguchi method, and two realistic conditions with dependences between variables. Evolutionary optimization is found to perform significantly better especially in the realistic conditions, suggesting that it forms a good approach for web interface design and other related applications in the future.
Jingbo Jiang, Diego Legrand, Robert Severn, Risto Miikkulainen
CEC1
2020 Tight Compression: Compressing CNN Model Tightly Through Unstructured Pruning and Simulated Annealing Based Permutation
abstract
The unstructured sparsity after pruning poses a challenge to the efficient implementation of deep learning models in existing regular architectures like systolic arrays. The coarse-grained structured pruning, on the other hand, tends to have higher accuracy loss than unstructured pruning when the pruned models are of the same size. In this work, we propose a compression method based on the unstructured pruning and a novel weight permutation scheme. Through permutation, the sparse weight matrix is further compressed to a small and dense format to make full use of the hardware resources. Compared to the state-of-the-art works, the matrix compression rate is effectively improved from 5.88x to 10.28x. As a result, the throughput and energy efficiency are improved by 2.12 and 1.57 times, respectively.
Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui
DAC3
2019 CompRRAE: RRAM-based convolutional neural network accelerator with reduced computations through a runtime activation estimation
abstract
Recently Resistive-RAM (RRAM) crossbar has been used in the design of the accelerator of convolutional neural networks (CNNs) to solve the memory wall issue. However, the intensive multiply-accumulate computations (MACs) executed at the crossbars during the inference phase are still the bottleneck for the further improvement of energy efficiency and throughput. In this work, we explore several methods to reduce the computations for the RRAM-based CNN accelerators. First, the output sparsity resulting from the widely employed Rectified Linear Unit is exploited, and a significant portion of computations are bypassed through an early detection of the negative output activations. Second, an adaptive approximation is proposed to terminate the MAC early when the sum of the partial results of the remaining computations is considered to be within a certain range of the intermediate accumulated result and thus has an insignificant contribution to the inference. In order to determine these redundant computations, a novel runtime estimation on the maximum and minimum values of each output activation is developed and used during the MAC operation. Experimental results show that around 70% of the computations can be reduced during the inference with a negligible accuracy loss smaller than 0.2%. As a result, the energy efficiency and the throughput are improved by over 2.9 and 2.8 times, respectively, compared with the state-of-the-art RRAM-based accelerators.
Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui
ASP-DAC3
2019 SubMac: Exploiting the subword-based computation in RRAM-based CNN accelerator for energy saving and speedup
Xizi Chen, Jingbo Jiang, Jingyang Zhu, Chi-Ying Tsui
Integr.2
2018 A high-throughput and energy-efficient RRAM-based convolutional neural network using data encoding and dynamic quantization
abstract
To solve the scaling, memory wall and high power density issues, recently RRAM-based accelerators, which show a better energy and area efficiency compared with the CMOS-based counterparts, have been proposed for convolutional neural networks. However, the RRAM-based architectures still face several design challenges, including the high energy and timing overhead at the analog/digital (A/D) conversion and interfacing circuits. To address these issues, we propose several novel optimization schemes in this work. First an encoding scheme for the synaptic weights and the input feature maps is proposed to reduce the energy of the in-situ computation and the bit-resolution of the A/D conversion. Then the resolution of the A/D conversion is further optimized for a lower energy consumption. Moreover, a dynamic quantization scheme for the multiply-accumulate operations (MACs) is proposed to improve the throughput and the energy efficiency by reducing the number of partial products. Experimental results show that the throughput, the energy efficiency and the area efficiency are improved by 2 to 4 times when compared with the state-of-the-art RRAM-based accelerators.
Xizi Chen, Jingbo Jiang, Jingyang Zhu, Chi-Ying Tsui
ASP-DAC2
2018 SparseNN: An energy-efficient neural network accelerator exploiting input and output sparsity
abstract
Contemporary Deep Neural Network (DNN) contains millions of synaptic connections with tens to hundreds of layers. The large computational complexity poses a challenge to the hardware design. In this work, we leverage the intrinsic activation sparsity of DNN to substantially reduce the execution cycles and the energy consumption. An end-to-end training algorithm is proposed to develop a lightweight (less than 5% overhead) run-time predictor for the output activation sparsity on the fly. Furthermore, an energy-efficient hardware architecture, SparseNN, is proposed to exploit both the input and output sparsity. SparseNN is a scalable architecture with distributed memories and processing elements connected through a dedicated on-chip network. Compared with the state-of-the-art accelerators which only exploit the input sparsity, SparseNN can achieve a 10%-70% improvement in throughput and a power reduction of around 50%.
Jingyang Zhu, Jingbo Jiang, Xizi Chen, Chi-Ying Tsui
DATE2