EDBT 2026 Demo / reviewers in the wild / expert
Yu-Chuan Chuang
dblp:227/2828
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0001-7940-9033ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AIRCHITECT v2: Learning the Hardware Accelerator Design Space Through Unified RepresentationsabstractDesign space exploration (DSE) plays a crucial role in enabling custom hardware architectures, particularly for emerging applications like AI, where optimized and specialized designs are essential. With the growing complexity of deep neural networks (DNNs) and the introduction of advanced large language models (LLMs), the design space for DNN accelerators is expanding at an exponential rate. Additionally, this space is highly non-uniform and non-convex, making it increasingly difficult to navigate and optimize. Traditional DSE techniques rely on search-based methods, which involve iterative sampling of the design space to find the optimal solution. However, this process is both time-consuming and often fails to converge to the global optima for such design spaces. Recently, AIRCHITECT vl, the first attempt to address the limitations of search-based techniques, transformed DSE into a constant-time classification problem using recommendation networks. However, AIRCHITECT v1 lacked generalizability and had poor performance in complex design spaces. In this work, we propose AIRCHITECT v2, a more accurate and generalizable learning-based DSE technique applicable to large-scale design spaces that overcomes the shortcomings of earlier approaches. Specifically, we devise an encoder-decoder transformer model that ($a$) encodes the complex design space into a uniform intermediate representation using contrastive learning and (b) leverages a novel unified representation blending the advantages of classification and regression to effectively explore the large DSE space without sacrificing accuracy. Experimental results evaluated on 105real DNN workloads demonstrate that, on average, AIRCHITECT v2 outperforms existing techniques by 15% in identifying optimal design points. Furthermore, to demonstrate the generalizability of our method, we evaluate performance on unseen model workloads and attain a 1.7 x improvement in inference latency on the identified hardware architecture. Code and dataset are available at: https://github.com/maestro-project/AIrchitect-v2. Jamin Seo, Akshat Ramachandran, Yu-Chuan Chuang, Anirudh Itagi, Tushar Krishna |
DATE | 3 |
| 2024 | BFP-CIM: Data-Free Quantization with Dynamic Block-Floating-Point Arithmetic for Energy-Efficient Computing-In-Memory-based AcceleratorabstractConvolutional neural networks (CNNs) are known for their exceptional performance in various applications; however, their energy consumption during inference can be substantial. Analog Computing-In-Memory (CIM) has shown promise in enhancing the energy efficiency of CNNs, but the use of analog-to-digital converters (ADCs) remains a challenge. ADCs convert analog partial sums from CIM crossbar arrays to digital values, with high-precision ADCs accounting for over 60% of the system’s energy. Researchers have explored quantizing CNNs to use low-precision ADCs to tackle this issue, trading off accuracy for efficiency. However, these methods necessitate data-dependent adjustments to minimize accuracy loss. Instead, we observe that the first most significant toggled bit indicates the optimal quantization range for each input value. Accordingly, we propose a range-aware rounding (RAR) for runtime bit-width adjustment, eliminating the need for pre-deployment efforts. RAR can be easily integrated into a CIM accelerator using dynamic block-floating-point arithmetic. Experimental results show that our methods maintain accuracy while achieving up to 1.81 × and 2.08 × energy efficiency improvements on CIFAR-10 and ImageNet datasets, respectively, compared with state-of-the-art techniques. Cheng-Yang Chang, Chi-Tse Huang, Yu-Chuan Chuang, Kuang-Chao Chou, An-Yeu Wu |
ASPDAC | 3 |
| 2024 | A 40nm 24.6TOPS/W Scalable EfficientDet Processor for Object DetectionabstractObject detection is a crucial technology used to identify and locate objects in a wide range of applications. Google’s EfficientDet, a scalable solution, employs a compound scaling method to systematically adjust the network’s depth, width, and input resolution, meeting different resource constraints on edge devices. This paper presents the first dedicated processor for EfficientDet, featuring three key elements: 1) an adaptive channel/input-wise (CIW) mapper to improve hardware utilization by applying distinct mapping strategies for layers with varying data shapes, 2) a tri-mode activation compression (AC) engine to reduce external memory access (EMA) by leveraging the sparsity level of activations, and 3) a unified aggregation core (AggrCore) to flexibly handle different computations. The chip is fabricated using TSMC 40nm CMOS technology and achieves a maximum energy efficiency of 24.6TOPS/W. Compared to the state-of-the-art object detection processor, our chip demonstrates 3.7× and 2.1× improvements in energy and area efficiencies, respectively. Yu-Chuan Chuang, Ming-Guang Lin, Chi-Tse Huang, Chieh-Fang Teng, Cheng-Yang Chang, Yi-Ta Chen, An-Yeu Wu |
ISCAS | 1 |
| 2024 | BFP-CIM: Runtime Energy-Accuracy Scalable Computing-in-Memory-Based DNN Accelerator Using Dynamic Block-Floating-Point ArithmeticabstractConvolutional neural networks (CNNs) are known for their exceptional performance in various applications; however, their energy consumption during inference can be substantial. Analog Computing-In-Memory (CIM) has shown promise in enhancing the energy efficiency of CNNs, but the use of analog-to-digital converters (ADCs) remains a challenge. In analog CIM-based accelerators, ADCs convert analog partial sums from CIM crossbar arrays to digital values, with high-precision ADCs accounting for over 60% of the system’s energy consumption. To prevent ADCs from damaging the energy efficiency benefits of CIM, researchers have explored quantizing CNNs to use low-precision ADCs, trading off accuracy for energy efficiency. However, these approaches often necessitate data-dependent adjustments to minimize accuracy loss. Instead, we observe that the first most significant toggled bit indicates the optimal quantization range for each input value. Accordingly, we propose a range-aware rounding (RAR) method for runtime bit-width adjustment, eliminating the need for pre-deployment efforts. RAR can be easily integrated into a CIM accelerator using dynamic block-floating-point arithmetic. We also seamlessly incorporate a bit-level zero-skipping mechanism by dynamically forming input blocks. Experimental results demonstrate that our methods maintain accuracy while achieving up to 1.81$\bm{\times }$and 2.08$\bm{\times }$energy efficiency improvements on the CIFAR-10 and ImageNet datasets, respectively, compared with state-of-the-art techniques. Cheng-Yang Chang, Chi-Tse Huang, Yu-Chuan Chuang, Kuang-Chao Chou, An-Yeu Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | BWA-NIMC: Budget-based Workload Allocation for Hybrid Near/In-Memory-ComputingabstractTo enable efficient computation for convolutional neural networks, in-memory-computing (IMC) is proposed to perform computation within memory. However, the non-ideality significantly degrades the accuracy of IMC. In this work, we leverage a hybrid near/in-memory-computing architecture (NIMC) that allocates sensitive weights to error-free NMC and computes remained weights with high-efficient IMC. We further propose a Budget-based Workload Allocation for NIMC (BWA-NIMC). Specifically, we consider the resource difference between NMC and IMC to effectively allocate workloads under a targeted resource budget. Simulation results show that BWA-NIMC improves the accuracy by 18.38-48.54% under limited budgets (e.g., energy and latency) compared with prior works. Chi-Tse Huang, Cheng-Yang Chang, Yu-Chuan Chuang, An-Yeu Wu |
DAC | 3 |
| 2023 | S-QRD-ELM: Scalable QR-Decomposition-Based Extreme Learning Machine Engine Supporting Online Class-Incremental Learning for ECG-Based User IdentificationabstractUser identification enables secure access to data and machines in smart factories. Compared with other modalities, ECG-based user identification is rising due to its intrinsic liveness proof and invulnerability to spoofing without contact. On the other hand, as new employees are registered at the factory, the ECG-based user identification system needs to be updated based on the new coming data. This scenario can be defined as an online class-incremental learning (O-CIL) problem. By exploiting hardware-software co- design, this work presents a Scalable QR-decomposition-based extreme learning machine (S-QRD-ELM) engine that can effectively and efficiently support O-CIL for ECG-based user identification. At the software level, we apply the concept of “the others” class and inversion-free QR-decomposition (QRD) recursive least squares to the S-QRD-ELM. This makes S-QRD-ELM achieve 79.7% higher accuracy in the O-CIL scenario compared with the neural network trained with back-propagation (BP-NN). At the hardware level, a one-dimensional diagonally-mapped linear array (1D-DMLA) is proposed to efficiently compute the QRD and back-substitution (BS) operations inside the S-QRD-ELM, reducing 98.5% of the silicon area. Moreover, the integrated processing element (PE) design with the unified COordinate Rotation DIgital Computer (u-CORDIC) further reduces 15.3% of the area and 22.4% of the power consumption. This engine is fabricated in 40nm CMOS technology with a$1.33\times 1.33$mm2 die area. The chip achieves$0.02\mu \text{J}$/sample and$2.47\mu \text{J}$/sample inferencing and learning energy efficiency, respectively, which is$6.4\times $and$28.5\times $than the state-of-the-art. To the best of our knowledge, the proposed highly energy-efficient S-QRD-ELM engine is the first chip to meet the requirements of O-CIL for ECG-based user identification. Yi-Ta Chen, Yu-Chuan Chuang, Li-Sheng Chang, An-Yeu Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Automated Quantization Range Mapping for DAC/ADC Non-linearity in Computing-In-MemoryabstractComputing-in-memory (CIM) has demonstrated the great potential of analog computing in improving the energy efficiency of matrix-vector multiplications for deep learning applications. Albeit low-power feature of CIM, the non-linearity of digital-to-analog converters (DACs)/analog-to-digital converters (ADCs) causes deviation between the computed outputs and desired values, thus degrading classification accuracy. This paper proposes Automated Quantization Range Mapping (A-QRM) mechanism to mitigate the negative effect of non-linearity on model accuracy. Instead of fixing the quantization range for quantized deep learning models, the proposed A-QRM automatically finds a better quantization range that balances the model capability and quantization errors caused by the non-linearity. Experimental results show that our proposed A-QRM achieves 89.02% and 86.93% of top-1 accuracy in ResNet20 and VGG8 on Cifar-10, respectively, under the non-linearity of DACs/ADCs. Chi-Tse Huang, Yu-Chuan Chuang, Ming-Guang Lin, An-Yeu Wu |
ISCAS | 2 |
| 2021 | MulTa-HDC: A Multi-Task Learning Framework For Hyperdimensional ComputingabstractBrain-inspired Hyperdimensional computing (HDC) has shown its effectiveness in low-power/energy designs for edge computing in the Internet of Things (IoT). Due to limited resources available on edge devices, multi-task learning (MTL), which accommodates multiple cognitive tasks in one model, is considered a more efficient deployment of HDC. However, as the number of tasks increases, MTL-based HDC (MTL-HDC) suffers from the huge overhead of associative memory (AM) and performance degradation. This hinders MTL-HDC from the practical realization on edge devices. This article aims to establish an MTL framework for HDC to achieve a flexible and efficient trade-off between memory overhead and performance degradation. For the shared-AM approach, we propose Dimension Ranking for Effective AM Sharing (DREAMS) to effectively merge multiple AMs while preserving as much information of each task as possible. For the independent-AM approach, we propose Dimension Ranking for Independent MEmory Retrieval (DRIMER) to extract and concatenate informative components of AMs while mitigating interferences among tasks. By leveraging both mechanisms, we propose a hybrid framework of Multi-Tasking HDC, called MulTa-HDC. To adapt an MTL-HDC system to an edge device given a memory resource budget, MulTa-HDC utilizes three parameters to flexibly adjust the proportion of the shared AM and independent AMs. The proposed MulTa-HDC is widely evaluated across three common benchmarks under two standard task protocols. The simulation results of ISOLET, UCIHAR, and MNIST datasets demonstrate that the proposed MulTa-HDC outperforms other state-of-the-art compressed HD models, including SparseHD and CompHD, by up to 8.23% in terms of classification accuracy. Cheng-Yang Chang, Yu-Chuan Chuang, En-Jui Chang, An-Yeu Wu |
IEEE Trans. Computers | 2 |