Chia-Chun Wang

dblp:261/6280 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Input Reuse, Weight-Stationary Dataflow and Mapping Strategy for Depthwise Convolution in Computing-in-Memory Neural Network Accelerators
abstract
Computing-in-Memory (CIM) is a promising solution to address the bottleneck of data movement in traditional Von Neumann architecture by performing in-situ computation in the memory. However, naively mapping depthwise convolution onto general CIM results in high computation latency or storage overhead due to poor input reuse. Unlike previous works focusing on designing CIM to support depthwise convolution effectively, we propose an Input Reuse, Weight-Stationary dataflow and mapping strategy to support depthwise convolution without modifying the macro, achieving a good balance between computation latency and storage size. The key aspect of our dataflow and mapping is to maximize the reuse of input features by allowing CIM to produce partial sums of output features in the same cycle. Additionally, we propose a detailed hardware design of a partial sum processing unit to handle the reconstruction of output features from the partial sums. The experimental results show that our approach achieves up to $2.87 \times$ model-wise energy efficiency compared to the baseline design with only 6.3% area overhead of the CIM macro.
Chia-Chun Wang, Yu-Chih Tsai, Ren-Shuo Liu
ASP-DAC1
2026 Adaptive Row-wise Attention Score Pruning and Compensation for Attention Acceleration
Chia-Chun Wang, Yu-Lin Lin, Ren-Shuo Liu
ISCAS1
2026 Tris-GCN: A 3D NAND Flash-based In-Storage Processing Architecture for GCN Acceleration
abstract
Graph convolutional networks (GCNs) excel in many applications, but scaling to large graphs is bottlenecked by heavy data movement. Existing in-storage processing (ISP) solutions offload I/O-intensive operations to the SSD controller to reduce PCIe traffic, but limited parallelism and flash bandwidth still constrain energy efficiency, even with accuracy-degrading neighbor sampling. We propose Tris-GCN, a 3D NAND-based ISP design that executes core GCN computations in situ by leveraging inherent flash computing capabilities, reducing channel traffic without accuracy loss. It incorporates mapping and scheduling optimizations for energy efficiency, alongside wear-leveling with selective recomputation for reliability. Results show that Tris-GCN achieves average 11.7× speedup (up to 41.0×) and 99.7% energy savings over CPU baselines, and 2.45× speedup and 63.8% energy savings over the SOTA ISP on the Amazon dataset.
Yi-Wa Wu, Jia-You Li, Chi-Jung Chen, Ching (Ryan) Cheng, Chia-Chun Wang, Chin-Fu Nien, Hsiang-Yun Cheng
ISLPED5
2025 Access Frequency-Aware Storage Reduction for Deep Learning Recommendation Model
abstract
Personalized recommendation is one of the flourishing AI applications in recent years. However, its powerful functionality comes with the need of a significant amount of embeddings, which hinders the model performance by causing large memory footprints. In addition, each inference request requires multiple memory access to these embeddings, which puts even more stress on the already-suffering memory bandwidth. In this paper, we aim to reduce the storage size of the embeddings required by personalized recommendation while inducing minimal impact on the accuracy. To achieve such goals, we take advantage of the fact that not all data contribute equally towards the final result. Small values and less frequently accessed embeddings have little impact on the accuracy. In addition, we opt to avoid any extra steps of training to restore the accuracy since training recommendation models can take hours, consuming a huge amount of computing power, which is undesirable. Instead, to restore accuracy, we further identify the insight that some data, although less frequently accessed, do not offer good storage reduction-accuracy trade-offs. Following these key guidelines, we judiciously choose certain embeddings to apply element-wise pruning, leaving the rest untouched. We are able to prune the embedding data by more than 99% while inducing less than 1% accuracy drop.
Chia-Chun Wang, Chuan-Yao Lai, Ren-Shuo Liu
ICCD1
2024 ISSA: Architecting CNN Accelerators Using Input-Skippable, Set-Associative Computing-in-Memory
abstract
Among several emerging architectures, computing in memory (CIM), which featuresin-situ analog computation, is a potential solution to the data movement bottleneck of the Von Neumann architecture for artificial intelligence (AI). Interestingly, more strengths of CIM significantly different from in-situ analog computation are not widely known yet. In this work, we point out thatmutually stationary vectors (MSVs), which can be maximized by introducingassociativityto CIM, are another inherent power unique to CIM. By MSVs, CIM exhibits significant freedom to dynamically vectorize the stored data (e.g., weights) to perform agile computation using the dynamically formed vectors. We have designed and realized an SA-CIM silicon prototype and corresponding architecture and acceleration schemes in the TSMC 28 nm process. More specifically, the contributions of this paper are fivefold: 1) We identify MSVs as new features that can be exploited to improve the current performance and energy challenges of the CIM-based hardware. 2) We propose SA-CIM to enhance MSVs (input-reordering flexibility) for skipping the zeros, small values, and sparse vectors. 3) We propose channel swapping to enhance the zero-skipping technique. 4) We propose a transposed systolic dataflow to efficiently conduct conv3×3 while being capable of exploiting input-skipping schemes. 5) We propose a design flow to search for optimal aggressive skipping scheme setups while satisfying the accuracy loss constraint. The proposed ISSA architecture improves the throughput by 1.91× to 2.97× speedup and the energy efficiency by 2.5× to 4.2×.
Yun-Chen Lo, Jun-Shen Wu, Chia-Chun Wang, Yu-Chih Tsai, Chih-Chen Yeh, Wen-Chien Ting, Ren-Shuo Liu
IEEE Trans. Computers3
2023 Built-in Self-Test and Built-in Self-Repair Strategies Without Golden Signature for Computing in Memory
abstract
This paper proposes built-in self-test (BIST) and built-in self-repair (BISR) strategies for computing in memory (CIM), including a novel test method and two repair schemes. They all focus on mitigating the impacts of inherent and in-evitable CIM inaccuracy on convolution neural networks (CNNs). Regarding the proposed BIST strategy, it exploits the distributive law to achieve at-speed CIM tests without storing testing vectors or golden results. Besides, it can assess the severity of the inherent inaccuracies among CIM bitlines instead of only offering a pass/fail outcome. In addition to BIST, we propose two BISR strategies. First, we propose to slightly offset the dynamic range of CIM outputs toward the negative side to create a margin for negative noises. By not cutting CIM outputs off at zero, negative noises are preserved to cancel out positive noises statistically, and accuracy impacts are mitigated. Second, we propose to remap the bitlines of CIM according to our BIST outcomes. Briefly speaking, we propose to map the least noisy bitlines to be the MSBs. This remapping can be done in the digital domain without touching the CIM internals. Experiments show that our proposed BIST and BISR strategies can restore CIM to less than 1% Top-1 accuracy loss with slight hardware overhead.
Yu-Chih Tsai, Wen-Chien Ting, Chia-Chun Wang, Chia-Cheng Chang, Ren-Shuo Liu
DATE3
2023 BICEP: Exploiting Bitline Inversion for Efficient Operation-Unit-Based Compute-in-Memory Architecture: No Retraining Needed!
abstract
Compute-in-memory (CIM) architecture is promising for its in-situ analog computing ability. However, one practical constraint for CIM architectures is the limited number of activated rows in an operation Unit (OU). OU-based CIM architecture only activates a subgroup of memory cells to ensure a large signal margin and enough consideration of non-ideal device/circuit effect, which pays the cost of lowered computing throughput. In short, the OU-based CIM architectures suffer from array underutilization to ensure high accuracy.This work proposes a novel architecture, BICEP, which exploits bitline inversion technique to enlarge the OU size without the need to prune, approximate, and retrain. More specifically, the key contributions of this work are threefold: 1) We propose a bitline inversion scheme, which guarantees more than 2× larger OU size without affecting the numerical results and the ADC resolution. The key insight is to selectively apply code inversion on heavy bitlines to constrain their MAC outputs and compensate using low-cost compensation units. We mathematically prove that the proposal can be applied to both single- and multi-level cells (SLC and MLC). 2) We propose an inversion-aware weight swapping scheme, which swaps the weight order to maximize the OU size exploiting bitline inversion. 3) We propose weight order propagation to enable inversion-aware weight swapping without storage overheads. The extensive experiments on ImageNet classification tasks demonstrate that this work outperforms state-of-the-art OU-based CIM architecture (DL-RSIM) by up to 2.06× speedup and 1.97× energy efficiency.
Yun-Chen Lo, Chia-Chun Wang, Ren-Shuo Liu
ICCD2
2023 CNN Inference Accelerators with Adjustable Feature Map Compression Ratios
abstract
Recently, an increasing interest has been in developing a convolution neural network (CNN) with adjustable configurations, enabling instant adaption to different resource constraints during inference. The trained CNN in run-time can switch to different modes to achieve a certain accuracy-energy trade-off point, similar to DVFS (dynamic voltage and frequency scaling) and turbo boost, which are widely adopted in CPUs. In this paper, we propose strategies to enable CNN inference accelerators to have an adjustable feature map compression ratio, making them tunable regarding their external memory access amount. We resort to the mature JPEG technique to compress those intermediate feature maps. The critical challenge is to support such adjustable compression ratios using one single CNN instead of multiple CNNs corresponding to multiple ratios. In response, we propose compression-aware joint-training and switchable batch normalization.We use ResNet18, ResNet50, and MobileNetV2 on ImageNet to demonstrate our design, achieve inference-time compression ratio adjustability, and reduce external memory access bandwidth requirements. The result shows that our proposed strategies can maintain the Top-1 accuracy and reduce external memory access by at most 22.7× ∼ 28.3× only using a single CNN model with sets of BN parameters corresponding to multiple compression ratios.
Yu-Chih Tsai, Chung-Yueh Liu, Chia-Chun Wang, Tsen-Wei Hsu, Ren-Shuo Liu
ICCD3
2023 Exploiting and Enhancing Computation Latency Variability for High-Performance Time-Domain Computing-in-Memory Neural Network Accelerators
abstract
To address the inefficiency resulting from data movement in Von Neumann architecture, computing-in-memory (CIM) is a promising solution due to its in-situ analog computation. Among the various types of CIMs, time-domain CIM stands out as a promising solution for achieving high energy efficiency and high readout resolution by employing time-to-digital converters (TDC) instead of analog-to-digital converters (ADC) to convert time-domain delays into digital values. However, the performance of the accelerator may be constrained by the maximum operating frequency of time-domain CIM, which is significantly lower than that of digital circuits.This paper proposes an architecture for a time-domain CIM-based neural network accelerator that leverages the varying output time of the TDC. The key contributions of this work are as follows: 1) We introduce an early-termination scheme for time-domain CIM, which dynamically determines the length of the CIM clock period by deriving the maximum possible multiply-accumulate (MAC) value based on the current input. This approach reduces computation time for low-MAC results. 2) We propose an input-inversion scheme to decrease the computation time for high-MAC results. By employing linear combination, we perform bit-inversion on large inputs and compensate for the results using a low-cost digital circuit. 3) We propose a hardware optimization on the compensation circuit by combining it with shift-adders in traditional neural network accelerators.Experiments show that our schemes could gain 2× ∼ 2.9× speedup under different clock period specifications with 5.82% area overhead compared to the CIM macro.
Chia-Chun Wang, Yun-Chen Lo, Jun-Shen Wu, Yu-Chih Tsai, Chia-Cheng Chang, Tsen-Wei Hsu, Min-Wei Chu, Chuan-Yao Lai, Ren-Shuo Liu
ICCD1
2022 ISSA: Input-Skippable, Set-Associative Computing-in-Memory (SA-CIM) Architecture for Neural Network Accelerators
abstract
Among several emerging architectures, computing in memory (CIM), which features in-situ analog computation, is a potential solution to the data movement bottleneck of the Von Neumann architecture for artificial intelligence (AI). Interestingly, more strengths of CIM significantly different from in-situ analog computation are not widely known yet. In this work, we point out that mutually stationary vectors (MSVs), which can be maximized by introducing associativity to CIM, are another inherent power unique to CIM. By MSVs, CIM exhibits significant freedom to dynamically vectorize the stored data (e.g., weights) to perform agile computation using the dynamically formed vectors.
Yun-Chen Lo, Chih-Chen Yeh, Jun-Shen Wu, Chia-Chun Wang, Yu-Chih Tsai, Wen-Chien Ting, Ren-Shuo Liu
ICCAD4