EDBT 2026 Demo / reviewers in the wild / expert
Yung-Chin Chen
dblp:32/1031
· DBLP profile ↗
7ranked-venue papers
5as first author
3since 2021 · last 2025
0009-0006-7064-1905ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ASiM: Modeling and Analyzing Inference Accuracy of SRAM-Based Analog CiM CircuitsabstractStatic random-access memory (SRAM)-based analog compute-in-memory (ACiM) demonstrates promising energy efficiency for deep neural network (DNN) processing. Nevertheless, efforts to optimize efficiency frequently compromise accuracy, and this trade-off remains insufficiently studied due to the difficulty of performing full-system validation. Specifically, existing simulation tools rarely target SRAM-based ACiM and exhibit inconsistent accuracy predictions, highlighting the need for a standardized, SRAM compute-in-memory (CiM) circuit-aware evaluation methodology. This article presents ASiM, a simulation framework for evaluating inference accuracy in SRAM-based ACiM systems. ASiM captures critical effects in SRAM-based analog compute in memory systems, such as analog-to-digital converter (ADC) quantization, bit-parallel encoding, and analog noise, which must be modeled with high fidelity due to their distinct behavior in charge-domain architectures compared to other memory technologies. ASiM supports a wide range of modern DNN workloads, including CNN and Transformer-based models such as ViT, and scales to large-scale tasks like ImageNet classification. Our results indicate that bit-parallel encoding can improve energy efficiency with only modest accuracy degradation; however, even 1 LSB of analog noise can significantly impair inference performance, particularly in complex tasks such as ImageNet. To address this, we explore hybrid analog-digital execution and majority voting schemes, both of which enhance robustness without negating energy savings. ASiM bridges the gap between hardware design and inference performance, offering actionable insights for energy-efficient, high-accuracy ACiM deployment. The code is available athttps://github.com/Keio-CSG/ASiM Wenlun Zhang, Shimpei Ando, Yung-Chin Chen, Kentaro Yoshioka |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | OSA-HCIM: On-The-Fly Saliency-Aware Hybrid SRAM CIM with Dynamic Precision ConfigurationabstractComputing-in-Memory (CIM) has shown great potential for enhancing efficiency and performance for deep neural networks (DNNs). However, the lack of flexibility in CIM leads to an unnecessary expenditure of computational resources on less critical operations, and a diminished Signal-to-Noise Ratio (SNR) when handling more complex tasks, significantly hindering the overall performance. Hence, we focus on the integration of CIM with Saliency-Aware Computing—a paradigm that dynamically tailors computing precision based on the importance of each input. We propose On-the-fly Saliency-Aware Hybrid CIM (OSA-HCIM) offering three primary contributions: (1) On-the-fly Saliency-Aware (OSA) precision configuration scheme, which dynamically sets the precision of each multiply-and-accumulate (MAC) operation based on its saliency, (2) Hybrid CIM Array (HCIMA), which enables simultaneous operation of digital-domain CIM (DCIM) and analog-domain CIM (ACIM) via split-port 6T SRAM, and (3) an integrated framework combining OSA and HCIMA to fulfill diverse accuracy and power demands.Implemented on a 65nm CMOS process, OSA-HCIM demon-strates an exceptional balance between accuracy and resource utilization. Notably, it is the first CIM design to incorporate a dynamic digital-to-analog boundary, providing unprecedented flexibility for saliency-aware computing. OSA-HCIM achieves a 1. 95x enhancement in energy efficiency, while maintaining minimal accuracy loss compared to DCIM when tested on CIFAR100 dataset. Yung-Chin Chen, Shimpei Ando, Daichi Fujiki, Shinya Takamaeda-Yamazaki, Kentaro Yoshioka |
ASPDAC | 1 |
| 2024 | PACiM: A Sparsity-Centric Hybrid Compute-in-Memory Architecture via Probabilistic ApproximationabstractApproximate computing emerges as a promising approach to enhance the efficiency of compute-in-memory (CiM) systems in deep neural network processing. However, traditional approximate techniques often significantly trade off accuracy for power efficiency, and fail to reduce data transfer between main memory and CiM banks, which dominates power consumption. This paper introduces a novel probabilistic approximate computation (PAC) method that leverages statistical techniques to approximate multiply-and-accumulation (MAC) operations, reducing approximation error by 4× compared to existing approaches. PAC enables efficient sparsity-based computation in CiM systems by simplifying complex MAC vector computations into scalar calculations. Moreover, PAC enables sparsity encoding and eliminates the LSB activations transmission, significantly reducing data reads and writes. This sets PAC apart from traditional approximate computing techniques, minimizing not only computation power but also memory accesses by 50%, thereby boosting system-level efficiency. We developed PACiM, a sparsity-centric architecture that fully exploits sparsity to reduce bit-serial cycles by 81% and achieves a peak 8b/8b efficiency of 14.63 TOPS/W in 65 nm CMOS while maintaining high accuracy of 93.85/72.36/66.02% on CIFAR-10/CIFAR-100/ImageNet benchmarks using a ResNet-18 model, demonstrating the effectiveness of our PAC methodology. Software simulation framework is available at GitHub. Wenlun Zhang, Shimpei Ando, Yung-Chin Chen, Satomi Miyagi, Shinya Takamaeda-Yamazaki, Kentaro Yoshioka |
ICCAD | 3 |
| 1993 | Performance Evaluation of Memory Caches in MultiprocessorsabstractLarge-scale MIN-based shared-memory multiprocessor systems have long shared memory latency. Private caches can improve memory access latency but they may suf fer from the cache coherence problem and potentially lower data locality due to data sharing and multiproces sor scheduling. These two problems also increase shared memory load and may result in frequent memory stalls. In this paper, we evaluate the performance of memory caches, a cache memory placed in front of shared memory, in a large-scale multiprocessor system in the presence of pro cessor caches. The memory cache is shown to have good performance and scalability. Yung-Chin Chen, Alexander V. Veidenbaum |
ICPP (1) | 1 |
| 1992 | An Effective Write Policy for Software Coherence SchemesabstractThe authors study the write behavior and evaluate the performance of various write strategies and buffering techniques for a MIN-based multiprocessor system using the simple software coherence scheme. Hit ratios, memory latencies, total execution time, and total write traffic are used as the performance indices. The write-through write-allocate no-fetch cache using a write-back write buffer is shown to have a better performance than both write-through and write-back caches. This type of write buffer is effective in reducing the volume as well as bursts of write traffic. On average, the use of a write-back cache reduces by 60% the total write traffic generated by a write-through cache.> Yung-Chin Chen, Alexander V. Veidenbaum |
SC | 1 |
| 1991 | A software coherence scheme with the assistance of directoriesabstractArticle A software coherence scheme with the assistance of directories Share on Authors: Yung-Chin Chen Center for Supercomputing Research and Development, University of Illinois at Urbana-Champaign, Urbana, Illinois Center for Supercomputing Research and Development, University of Illinois at Urbana-Champaign, Urbana, IllinoisView Profile , Alexander V. Veidenbaum Center for Supercomputing Research and Development, University of Illinois at Urbana-Champaign, Urbana, Illinois Center for Supercomputing Research and Development, University of Illinois at Urbana-Champaign, Urbana, IllinoisView Profile Authors Info & Claims ICS '91: Proceedings of the 5th international conference on SupercomputingJune 1991 Pages 284–294https://doi.org/10.1145/109025.109095Online:01 June 1991Publication History 6citation215DownloadsMetricsTotal Citations6Total Downloads215Last 12 Months5Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Yung-Chin Chen, Alexander V. Veidenbaum |
ICS | 1 |
| 1991 | Comparison and analysis of software and directory coherence schemesabstractDirectory schemes and software schemes have been proposed to solve the cache coherence problem for the MIN-based large-scale multiprocessor system.We compare the performance of the two schemes using irate-driven simulation including the effect of fake sharing caused by a nontrivial cache line size.It shows that the simplest software scheme can have a hit ratio and shared memory trafic comparable to those of the directory scheme.The invalidations and the sharing behavior of the directory scheme are classified and analyzed. Yung-Chin Chen, Alexander V. Veidenbaum |
SC | 1 |