EDBT 2026 Demo / reviewers in the wild / expert
Jun-Shen Wu
dblp:253/7491
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-3816-3350ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reliability and Optimization for Neural Network Accelerators using Value-Aware Error-Marking Pattern with Sequential Access Error Correction
Jun-Shen Wu, Ren-Shuo Liu |
ISCAS | 1 |
| 2024 | ISSA: Architecting CNN Accelerators Using Input-Skippable, Set-Associative Computing-in-MemoryabstractAmong several emerging architectures, computing in memory (CIM), which featuresin-situ analog computation, is a potential solution to the data movement bottleneck of the Von Neumann architecture for artificial intelligence (AI). Interestingly, more strengths of CIM significantly different from in-situ analog computation are not widely known yet. In this work, we point out thatmutually stationary vectors (MSVs), which can be maximized by introducingassociativityto CIM, are another inherent power unique to CIM. By MSVs, CIM exhibits significant freedom to dynamically vectorize the stored data (e.g., weights) to perform agile computation using the dynamically formed vectors. We have designed and realized an SA-CIM silicon prototype and corresponding architecture and acceleration schemes in the TSMC 28 nm process. More specifically, the contributions of this paper are fivefold: 1) We identify MSVs as new features that can be exploited to improve the current performance and energy challenges of the CIM-based hardware. 2) We propose SA-CIM to enhance MSVs (input-reordering flexibility) for skipping the zeros, small values, and sparse vectors. 3) We propose channel swapping to enhance the zero-skipping technique. 4) We propose a transposed systolic dataflow to efficiently conduct conv3×3 while being capable of exploiting input-skipping schemes. 5) We propose a design flow to search for optimal aggressive skipping scheme setups while satisfying the accuracy loss constraint. The proposed ISSA architecture improves the throughput by 1.91× to 2.97× speedup and the energy efficiency by 2.5× to 4.2×. Yun-Chen Lo, Jun-Shen Wu, Chia-Chun Wang, Yu-Chih Tsai, Chih-Chen Yeh, Wen-Chien Ting, Ren-Shuo Liu |
IEEE Trans. Computers | 2 |
| 2023 | Exploiting and Enhancing Computation Latency Variability for High-Performance Time-Domain Computing-in-Memory Neural Network AcceleratorsabstractTo address the inefficiency resulting from data movement in Von Neumann architecture, computing-in-memory (CIM) is a promising solution due to its in-situ analog computation. Among the various types of CIMs, time-domain CIM stands out as a promising solution for achieving high energy efficiency and high readout resolution by employing time-to-digital converters (TDC) instead of analog-to-digital converters (ADC) to convert time-domain delays into digital values. However, the performance of the accelerator may be constrained by the maximum operating frequency of time-domain CIM, which is significantly lower than that of digital circuits.This paper proposes an architecture for a time-domain CIM-based neural network accelerator that leverages the varying output time of the TDC. The key contributions of this work are as follows: 1) We introduce an early-termination scheme for time-domain CIM, which dynamically determines the length of the CIM clock period by deriving the maximum possible multiply-accumulate (MAC) value based on the current input. This approach reduces computation time for low-MAC results. 2) We propose an input-inversion scheme to decrease the computation time for high-MAC results. By employing linear combination, we perform bit-inversion on large inputs and compensate for the results using a low-cost digital circuit. 3) We propose a hardware optimization on the compensation circuit by combining it with shift-adders in traditional neural network accelerators.Experiments show that our schemes could gain 2× ∼ 2.9× speedup under different clock period specifications with 5.82% area overhead compared to the CIM macro. Chia-Chun Wang, Yun-Chen Lo, Jun-Shen Wu, Yu-Chih Tsai, Chia-Cheng Chang, Tsen-Wei Hsu, Min-Wei Chu, Chuan-Yao Lai, Ren-Shuo Liu |
ICCD | 3 |
| 2023 | FM-P2L: An Algorithm Hardware Co-design of Fixed-Point MSBs with Power-of-2 LSBs in CNN AcceleratorsabstractConvolutional neural networks (CNNs) are a focal point for advancing the field of artificial intelligence, enabling significant advances in image and speech recognition, natural language processing, and other complex tasks. However, due to the high requirements of memory storage and computational resources, implementing CNNs can be challenging.To alleviate these deficiencies, this work presents a novel algorithm hardware co-design centered on a new number format, fixed-point MSBs with power-of-2 LSBs (FM-P2L), which can significantly reduce the requirement of computational resources of the CNN accelerators by trading negligible accuracy loss. First, we propose the novel FM-P2L number format that uses fixed-point to represent MSBs and power-of-2 to represent LSBs, which can reduce the computation complexity of multiplications. The optimal bitwidths of MSBs and LSBs are determined from our proposed algorithm with the capability to preserve accuracy. Second, we propose a novel multiplier that best matches FM-P2L with CNN accelerators, which can significantly reduce the area and power of the computational units. Finally, to evaluate the benefits of FM-P2L, we compared FM-P2L with fixed-point and low-bitwidith floating-point on a weight-stationary based systolic array accelerator with vector-vector multiplication PEs, which is adopted by Google TPU and NVDLA, in TSMC 40nm technology. Our evaluation results demonstrate that FM-P2L can achieve up to 44% computing power and 50% area reduction compared with fixed-point and up to 55% computing power and 65% area reduction compared with floating-point, while maintaining negligible inference accuracy loss on the state-of-the-art CNN models. Jun-Shen Wu, Ren-Shuo Liu |
ICCD | 1 |
| 2023 | SG-Float: Achieving Memory Access and Computing Power Reduction Using Self-Gating Float in CNNsabstractConvolutional neural networks (CNNs) are essential for advancing the field of artificial intelligence. However, since these networks are highly demanding in terms of memory and computation, implementing CNNs can be challenging. To make CNNs more accessible to energy-constrained devices, researchers are exploring new algorithmic techniques and hardware designs that can reduce memory and computation requirements. In this work, we present self-gating float (SG-Float), algorithm hardware co-design of a novel binary number format, which can significantly reduce memory access and computing power requirements in CNNs. SG-Float is a self-gating format that uses the exponent to self-gate the mantissa to zero, exploiting the characteristic of floating-point that the exponent determines the magnitude of a floating-point value and the error tolerance property of CNNs. SG-Float represents relatively small values using only the exponent, which increases the proportion of ineffective mantissas, corresponding to reducing mantissa multiplications of floating-point numbers. To minimize the accuracy loss caused by the approximation error introduced by SG-Float, we propose a fine-tuning process to determine the exponent thresholds of SG-Float and reclaim the accuracy loss. We also develop a hardware optimization technique, called the SG-Float buffering strategy, to best match SG-Float with CNN accelerators and further reduce memory access. We apply the SG-Float buffering strategy to vector-vector multiplication processing elements (PEs), which NVDLA adopts, in TSMC 40nm technology. Our evaluation results demonstrate that SG-Float can achieve up to 35% reduction in memory access power and up to 54% reduction in computing power compared with AdaptivFloat, a state-of-the-art format, with negligible power and area overhead. Additionally, we show that SG-Float can be combined with neural network pruning methods to further reduce memory access and mantissa multiplications in pruned CNN models. Overall, our work shows that SG-Float is a promising solution to the problem of CNN memory access and computing power. Jun-Shen Wu, Tsen-Wei Hsu, Ren-Shuo Liu |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2022 | ISSA: Input-Skippable, Set-Associative Computing-in-Memory (SA-CIM) Architecture for Neural Network AcceleratorsabstractAmong several emerging architectures, computing in memory (CIM), which features in-situ analog computation, is a potential solution to the data movement bottleneck of the Von Neumann architecture for artificial intelligence (AI). Interestingly, more strengths of CIM significantly different from in-situ analog computation are not widely known yet. In this work, we point out that mutually stationary vectors (MSVs), which can be maximized by introducing associativity to CIM, are another inherent power unique to CIM. By MSVs, CIM exhibits significant freedom to dynamically vectorize the stored data (e.g., weights) to perform agile computation using the dynamically formed vectors. Yun-Chen Lo, Chih-Chen Yeh, Jun-Shen Wu, Chia-Chun Wang, Yu-Chih Tsai, Wen-Chien Ting, Ren-Shuo Liu |
ICCAD | 3 |
| 2021 | Value-Aware Error Detection and Correction for SRAM Buffers in Low-Bitwidth, Floating-Point CNN AcceleratorsabstractLow-power CNN accelerators are a key technique to enable the future artificial intelligence world. Dynamic voltage scaling is an essential low-power strategy, but it is bottlenecked by on-chip SRAM. More specifically, SRAM can exhibit stuck-at (SA) faults at a rate as high as 0.1% when the supply voltage is lowered to, e.g., 0.5 V. Although this issue has been studied in CPU cache design, since their solutions are tailored for CPUs instead of CNN accelerators, they inevitably incur unnecessary design complexity and SRAM capacity overhead. Jun-Shen Wu, Chi-En Wang, Ren-Shuo Liu |
ASP-DAC | 1 |