EDBT 2026 Demo / reviewers in the wild / expert
Yu-Chih Tsai
dblp:258/6658
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-5142-081XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Input Reuse, Weight-Stationary Dataflow and Mapping Strategy for Depthwise Convolution in Computing-in-Memory Neural Network AcceleratorsabstractComputing-in-Memory (CIM) is a promising solution to address the bottleneck of data movement in traditional Von Neumann architecture by performing in-situ computation in the memory. However, naively mapping depthwise convolution onto general CIM results in high computation latency or storage overhead due to poor input reuse. Unlike previous works focusing on designing CIM to support depthwise convolution effectively, we propose an Input Reuse, Weight-Stationary dataflow and mapping strategy to support depthwise convolution without modifying the macro, achieving a good balance between computation latency and storage size. The key aspect of our dataflow and mapping is to maximize the reuse of input features by allowing CIM to produce partial sums of output features in the same cycle. Additionally, we propose a detailed hardware design of a partial sum processing unit to handle the reconstruction of output features from the partial sums. The experimental results show that our approach achieves up to $2.87 \times$ model-wise energy efficiency compared to the baseline design with only 6.3% area overhead of the CIM macro. Chia-Chun Wang, Yu-Chih Tsai, Ren-Shuo Liu |
ASP-DAC | 2 |
| 2026 | Clipping Error Compensation for Accuracy Recovery and Throughput Improvement in Computing-in-Memory
Yu-Chih Tsai, Hsuan-Hung Shen, Ren-Shuo Liu |
ISCAS | 1 |
| 2025 | DMP-BFP: Dynamic Mixed-Precision Block Floating-Point and Exponent-Guided Precision AdjustmentabstractBlock Floating-Point (BFP), an emerging datatype, has demonstrated significant potential in model accuracy and hardware efficiency. This paper presents a dynamic mixedprecision BFP processing engine (PE), an accompanying framework, and optimization techniques to improve hardware efficiency. First, we propose a strategy for identifying accuracysensitive inner products within BFP models by comparing exponent values against predefined thresholds. This enables precision adjustments according to the sensitivity of the calculations at runtime. Second, we observe that only a small subset of inner products require full-precision (i.e., accuracy-sensitive inner products). Furthermore, a full-precision multiplication can be decomposed into four low-precision multiplications. Based on this, we propose the design of a low-precision PE capable of supporting full-precision mode, thereby reducing area overhead. Third, we optimize the BFP quantization scheme and datatype representation within the PE, significantly mitigating quantization errors in low-precision mode and reducing power consumption during datatype conversion. Finally, experimental results demonstrate that our dynamic mixed-precision BFP approach maintains accuracy while employing over 80% lowprecision operations and increases this ratio to 95% through retraining. Compared to state-of-the-art BFP architectures, our design improves inference speed, area efficiency, and energy efficiency by up to$1.64 \times, 1.42 \times$, and$1.47 \times$, respectively. Yu-Chih Tsai, Chia-Cheng Chang, Ren-Shuo Liu |
ICCD | 1 |
| 2024 | ISSA: Architecting CNN Accelerators Using Input-Skippable, Set-Associative Computing-in-MemoryabstractAmong several emerging architectures, computing in memory (CIM), which featuresin-situ analog computation, is a potential solution to the data movement bottleneck of the Von Neumann architecture for artificial intelligence (AI). Interestingly, more strengths of CIM significantly different from in-situ analog computation are not widely known yet. In this work, we point out thatmutually stationary vectors (MSVs), which can be maximized by introducingassociativityto CIM, are another inherent power unique to CIM. By MSVs, CIM exhibits significant freedom to dynamically vectorize the stored data (e.g., weights) to perform agile computation using the dynamically formed vectors. We have designed and realized an SA-CIM silicon prototype and corresponding architecture and acceleration schemes in the TSMC 28 nm process. More specifically, the contributions of this paper are fivefold: 1) We identify MSVs as new features that can be exploited to improve the current performance and energy challenges of the CIM-based hardware. 2) We propose SA-CIM to enhance MSVs (input-reordering flexibility) for skipping the zeros, small values, and sparse vectors. 3) We propose channel swapping to enhance the zero-skipping technique. 4) We propose a transposed systolic dataflow to efficiently conduct conv3×3 while being capable of exploiting input-skipping schemes. 5) We propose a design flow to search for optimal aggressive skipping scheme setups while satisfying the accuracy loss constraint. The proposed ISSA architecture improves the throughput by 1.91× to 2.97× speedup and the energy efficiency by 2.5× to 4.2×. Yun-Chen Lo, Jun-Shen Wu, Chia-Chun Wang, Yu-Chih Tsai, Chih-Chen Yeh, Wen-Chien Ting, Ren-Shuo Liu |
IEEE Trans. Computers | 4 |
| 2023 | Built-in Self-Test and Built-in Self-Repair Strategies Without Golden Signature for Computing in MemoryabstractThis paper proposes built-in self-test (BIST) and built-in self-repair (BISR) strategies for computing in memory (CIM), including a novel test method and two repair schemes. They all focus on mitigating the impacts of inherent and in-evitable CIM inaccuracy on convolution neural networks (CNNs). Regarding the proposed BIST strategy, it exploits the distributive law to achieve at-speed CIM tests without storing testing vectors or golden results. Besides, it can assess the severity of the inherent inaccuracies among CIM bitlines instead of only offering a pass/fail outcome. In addition to BIST, we propose two BISR strategies. First, we propose to slightly offset the dynamic range of CIM outputs toward the negative side to create a margin for negative noises. By not cutting CIM outputs off at zero, negative noises are preserved to cancel out positive noises statistically, and accuracy impacts are mitigated. Second, we propose to remap the bitlines of CIM according to our BIST outcomes. Briefly speaking, we propose to map the least noisy bitlines to be the MSBs. This remapping can be done in the digital domain without touching the CIM internals. Experiments show that our proposed BIST and BISR strategies can restore CIM to less than 1% Top-1 accuracy loss with slight hardware overhead. Yu-Chih Tsai, Wen-Chien Ting, Chia-Chun Wang, Chia-Cheng Chang, Ren-Shuo Liu |
DATE | 1 |
| 2023 | CNN Inference Accelerators with Adjustable Feature Map Compression RatiosabstractRecently, an increasing interest has been in developing a convolution neural network (CNN) with adjustable configurations, enabling instant adaption to different resource constraints during inference. The trained CNN in run-time can switch to different modes to achieve a certain accuracy-energy trade-off point, similar to DVFS (dynamic voltage and frequency scaling) and turbo boost, which are widely adopted in CPUs. In this paper, we propose strategies to enable CNN inference accelerators to have an adjustable feature map compression ratio, making them tunable regarding their external memory access amount. We resort to the mature JPEG technique to compress those intermediate feature maps. The critical challenge is to support such adjustable compression ratios using one single CNN instead of multiple CNNs corresponding to multiple ratios. In response, we propose compression-aware joint-training and switchable batch normalization.We use ResNet18, ResNet50, and MobileNetV2 on ImageNet to demonstrate our design, achieve inference-time compression ratio adjustability, and reduce external memory access bandwidth requirements. The result shows that our proposed strategies can maintain the Top-1 accuracy and reduce external memory access by at most 22.7× ∼ 28.3× only using a single CNN model with sets of BN parameters corresponding to multiple compression ratios. Yu-Chih Tsai, Chung-Yueh Liu, Chia-Chun Wang, Tsen-Wei Hsu, Ren-Shuo Liu |
ICCD | 1 |
| 2023 | Exploiting and Enhancing Computation Latency Variability for High-Performance Time-Domain Computing-in-Memory Neural Network AcceleratorsabstractTo address the inefficiency resulting from data movement in Von Neumann architecture, computing-in-memory (CIM) is a promising solution due to its in-situ analog computation. Among the various types of CIMs, time-domain CIM stands out as a promising solution for achieving high energy efficiency and high readout resolution by employing time-to-digital converters (TDC) instead of analog-to-digital converters (ADC) to convert time-domain delays into digital values. However, the performance of the accelerator may be constrained by the maximum operating frequency of time-domain CIM, which is significantly lower than that of digital circuits.This paper proposes an architecture for a time-domain CIM-based neural network accelerator that leverages the varying output time of the TDC. The key contributions of this work are as follows: 1) We introduce an early-termination scheme for time-domain CIM, which dynamically determines the length of the CIM clock period by deriving the maximum possible multiply-accumulate (MAC) value based on the current input. This approach reduces computation time for low-MAC results. 2) We propose an input-inversion scheme to decrease the computation time for high-MAC results. By employing linear combination, we perform bit-inversion on large inputs and compensate for the results using a low-cost digital circuit. 3) We propose a hardware optimization on the compensation circuit by combining it with shift-adders in traditional neural network accelerators.Experiments show that our schemes could gain 2× ∼ 2.9× speedup under different clock period specifications with 5.82% area overhead compared to the CIM macro. Chia-Chun Wang, Yun-Chen Lo, Jun-Shen Wu, Yu-Chih Tsai, Chia-Cheng Chang, Tsen-Wei Hsu, Min-Wei Chu, Chuan-Yao Lai, Ren-Shuo Liu |
ICCD | 4 |
| 2022 | ISSA: Input-Skippable, Set-Associative Computing-in-Memory (SA-CIM) Architecture for Neural Network AcceleratorsabstractAmong several emerging architectures, computing in memory (CIM), which features in-situ analog computation, is a potential solution to the data movement bottleneck of the Von Neumann architecture for artificial intelligence (AI). Interestingly, more strengths of CIM significantly different from in-situ analog computation are not widely known yet. In this work, we point out that mutually stationary vectors (MSVs), which can be maximized by introducing associativity to CIM, are another inherent power unique to CIM. By MSVs, CIM exhibits significant freedom to dynamically vectorize the stored data (e.g., weights) to perform agile computation using the dynamically formed vectors. Yun-Chen Lo, Chih-Chen Yeh, Jun-Shen Wu, Chia-Chun Wang, Yu-Chih Tsai, Wen-Chien Ting, Ren-Shuo Liu |
ICCAD | 5 |