Chi-Tse Huang

dblp:295/5360 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-0393-3976ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 5 first-author · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Energy-Efficient In-Memory Vector Similarity Search With Fidelity-Aware Range Encoding
abstract
Vector similarity search (VSS) is a crucial operation in applications of machine learning, but it often incurs high energy consumption due to frequent memory accesses. Previous works have adopted ternary content-addressable memories (TCAMs) to perform parallel VSS within memory. Among these approaches, Exact-Match TCAM (EX-TCAM) combined with range encoding has shown promise for executing in-memory VSS under theL∞norm distance. Existing EX-TCAM-based approaches initiate the search from the positions of query vectors and iteratively expand the search ranges to identify the closest stored vectors. However, this initialization strategy leads to excessive search iterations and long codewords. To address these challenges, we present FORE, an EX-TCAM-based framework that significantly improves latency, energy efficiency, and accuracy for in-memory VSS. To reduce redundant search iterations, we first initialize the search from ranges based on theL∞norm distance. To shorten the codeword length, we then develop a range encoding scheme that supports range-to-range matching. In addition, we introduce a novel metric, “range fidelity,” to evaluate the quality of range encoding. Building on the insight that a certain degree of loss in range fidelity is tolerable in EX-TCAM-based VSS, we further propose a lossy range encoding scheme that yields more compact codewords without significantly compromising accuracy. Finally, FORE incorporates aL∞norm distance-based training mechanism that further reduces search iterations and enhances classification accuracy in EX-TCAM-based VSS. Experimental results demonstrate that FORE improves energy efficiency by 45.25× to 59.72× and reduces latency by 4.47× to 5.66× compared to previous EX-TCAM-based methods within 1% accuracy loss. Moreover, FORE outperforms Best-Match TCAM (Best-TCAM)-based approaches by 6.9× in energy efficiency under non-ideal conditions.
Chi-Tse Huang, Hsiang-Yun Cheng, An-Yeu Wu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 Efficient and Reliable Vector Similarity Search Using Asymmetric Encoding with NAND-Flash for Many-Class Few-Shot Learning
abstract
While memory-augmented neural networks (MANNs) offer an effective solution for few-shot learning (FSL) by integrating deep neural networks with external memory, the capacity requirements and energy overhead of data movement become enormous due to the large number of support vectors in many-class FSL scenarios. Various in-memory search solutions have emerged to improve the energy efficiency of MANNs. NAND-based multi-bit content addressable memory (MCAM) is a promising option due to its high density and large capacity. Despite its potential, MCAM faces limitations such as a restricted number of word lines, limited quantization levels, and non-ideal effects like varying string currents and bottleneck effects, which lead to significant accuracy drops. To address these issues, we propose several innovative methods. First, the Multi-bit Thermometer Code (MTMC) leverages the extensive capacity of MCAM to enhance vector precision using cumulative encoding rules, thereby mitigating the bottleneck effect. Second, the Asymmetric Vector Similarity Search (AVSS) reduces the precision of the query vector while maintaining that of the support vectors, thereby minimizing the search iterations and improving efficiency in many-class scenarios. Finally, the Hardware-Aware Training (HAT) method optimizes controller training by modeling the hardware characteristics of MCAM, thus enhancing the reliability of the system. Our integrated framework reduces search iterations by up to 32×, and increases overall accuracy by 1.58% to 6.94%.
Hao-Wei Chiang, Chi-Tse Huang, Hsiang-Yun Cheng, Po-Hao Tseng, Ming-Hsiu Lee, An-Yeu Wu
ASP-DAC2
2025 Energy-Efficient Large-Scale Vector Similarity Search in NAND-Flash via Hybrid Matching
abstract
Vector similarity search (VSS) is crucial in many AI applications, such as few-shot learning (FSL) and approximate nearest-neighbor search (ANNS), but it demands significant memory capacity and incurs substantial energy costs for data transfers during large-scale comparisons. Various in-memory search technologies have been developed to improve energy efficiency, with NAND-based multi-bit content-addressable memory (MCAM) standing out as a promising solution for its high density and large capacity. MCAM can operate in exact-search (ES) mode, supporting only perfect matches with low energy cost, or in approximate-search (AS) mode, enabling flexible VSS. However, AS mode incurs significant energy waste when comparing queries with non-target stored vectors. To address this issue, we propose Hybrid-M, a 3D NAND-based in-memory VSS architecture that integrates both modes into a single hybrid matching process, using ES mode as a filter to reduce redundant searches for AS mode. We apply three techniques to optimize this integration: range encoding for multi-level cells (MLC) to enhance filtering, search voltage shifts to mitigate the impact on AS accuracy and reduce matching currents, and a filtering-aware training method to further improve reliability and energy efficiency. Results show that Hybrid-M achieves comparable accuracy while reducing energy consumption by 67% to 83% compared to MACM-based VSS using only AS mode, across various many-class FSL and ANNS workloads.
Chih-Yu Hu, Chi-Tse Huang, Hao-Wei Chiang, Hsiang-Yun Cheng, Po-Hao Tseng, Ming-Hsiu Lee, An-Yeu Wu
DAC2
2025 Segmented Angular Pre-Processing for Accurate and Efficient In-Memory Vector Similarity Search
abstract
Vector similarity search (VSS) is a fundamental operation in modern AI applications, including few-shot learning (FSL) and approximate nearest neighbor search (ANNS). Cosine similarity is widely regarded as the optimal metric for VSS. However, VSS incurs substantial energy and computational overhead, primarily due to frequent vector transfers and the complexity of cosine similarity calculations in high-dimensional spaces. Prior research has explored the use of ternary content addressable memories (TCAMs) for parallel in-memory VSS to reduce vector movement. Exact-Match TCAM (EX-TCAM) enables exact bitmatching, and Best-Match TCAM (Best-TCAM) supports Hamming distance calculations, both of which are spatial metrics and computationally efficient. As a result, existing TCAM-based VSS approaches have focused on developing frameworks to efficiently support more complex spatial metrics such as the $L_{\infty}$ and $L_{1}$ norms. However, these spatial metrics exhibit notable discrepancies compared to angular metrics like cosine similarity. To overcome this limitation, we propose Seg-Cos, a TCAM-based framework that directly approximates cosine similarity within TCAM for angular VSS. Seg-Cos introduces a dedicated preprocessing technique and encoding scheme that segments vectors and encodes them as circular ranges based on their angles and magnitudes. Seg-Cos is the first angular VSS framework compatible with both EX-TCAM and Best-TCAM, enabling accurate and energy-efficient VSS in the angular domain. Simulation results demonstrate that Seg-Cos improves energy efficiency by $1.41 \times$ and achieves up to $2.2 \%$ higher accuracy over prior EX-TCAMbased methods in FSL. In ANNS, Seg-Cos enhances recall rate by $10 \%$ to $52 \%$ and improves energy efficiency by $2 \times$ compared to previous Best-TCAM approaches with $L_{1}$ norm.
Chi-Tse Huang, Jen-Chieh Wang, Hsiang-Yun Cheng, An-Yeu Wu
DAC1
2024 BFP-CIM: Data-Free Quantization with Dynamic Block-Floating-Point Arithmetic for Energy-Efficient Computing-In-Memory-based Accelerator
abstract
Convolutional neural networks (CNNs) are known for their exceptional performance in various applications; however, their energy consumption during inference can be substantial. Analog Computing-In-Memory (CIM) has shown promise in enhancing the energy efficiency of CNNs, but the use of analog-to-digital converters (ADCs) remains a challenge. ADCs convert analog partial sums from CIM crossbar arrays to digital values, with high-precision ADCs accounting for over 60% of the system’s energy. Researchers have explored quantizing CNNs to use low-precision ADCs to tackle this issue, trading off accuracy for efficiency. However, these methods necessitate data-dependent adjustments to minimize accuracy loss. Instead, we observe that the first most significant toggled bit indicates the optimal quantization range for each input value. Accordingly, we propose a range-aware rounding (RAR) for runtime bit-width adjustment, eliminating the need for pre-deployment efforts. RAR can be easily integrated into a CIM accelerator using dynamic block-floating-point arithmetic. Experimental results show that our methods maintain accuracy while achieving up to 1.81 × and 2.08 × energy efficiency improvements on CIFAR-10 and ImageNet datasets, respectively, compared with state-of-the-art techniques.
Cheng-Yang Chang, Chi-Tse Huang, Yu-Chuan Chuang, Kuang-Chao Chou, An-Yeu Wu
ASPDAC2
2024 BORE: Energy-Efficient Banded Vector Similarity Search with Optimized Range Encoding for Memory-Augmented Neural Network
abstract
Memory-augmented neural networks (MANNs) in-corporate external memories to address the significant issue of catastrophic forgetting in few-shot learning applications. MANNs rely on vector similarity search (VSS), which incurs substantial energy and computational overhead due to frequent data transfers and complex cosine similarity calculations. To tackle these challenges, prior research has proposed adopting ternary content addressable memories (TCAMs) for parallel VSS within memory. One promising approach is to use Exact-Match TCAM (EX-TCAM) with range encoding to find the vector with the minimum$L_{\infty}$distance, avoiding the need for sensing circuit modifications as required by Best-Match TCAM (Best-TCAM), However, this method demands multiple search iterations and longer code words, limiting its practicality. In this paper, we propose an energy-efficient EX-TCAM-based design called BORE. BORE skips redundant search iterations and reduces code word length through performing Banded$L_{\infty}$distance search with Optimized Range Encoding. Additionally, we consider the characteristics of the similarity metric and develop a distance-based training mechanism aimed at improving classification accuracy. Simulation results demonstrate that BORE enhances energy efficiency by$9.35\times$to$11.84\times$and accuracy by 2.95 % to 4.69 % compared to previous EX-TCAM-based approaches. Furthermore, BORE improves energy efficiency by$1.04\times$to$1.63\times$over prior works of Best-TCAM-based VSS.
Chi-Tse Huang, Cheng-Yang Chang, Hsiang-Yun Cheng, An-Yeu Wu
DATE1
2024 A 40nm 24.6TOPS/W Scalable EfficientDet Processor for Object Detection
abstract
Object detection is a crucial technology used to identify and locate objects in a wide range of applications. Google’s EfficientDet, a scalable solution, employs a compound scaling method to systematically adjust the network’s depth, width, and input resolution, meeting different resource constraints on edge devices. This paper presents the first dedicated processor for EfficientDet, featuring three key elements: 1) an adaptive channel/input-wise (CIW) mapper to improve hardware utilization by applying distinct mapping strategies for layers with varying data shapes, 2) a tri-mode activation compression (AC) engine to reduce external memory access (EMA) by leveraging the sparsity level of activations, and 3) a unified aggregation core (AggrCore) to flexibly handle different computations. The chip is fabricated using TSMC 40nm CMOS technology and achieves a maximum energy efficiency of 24.6TOPS/W. Compared to the state-of-the-art object detection processor, our chip demonstrates 3.7× and 2.1× improvements in energy and area efficiencies, respectively.
Yu-Chuan Chuang, Ming-Guang Lin, Chi-Tse Huang, Chieh-Fang Teng, Cheng-Yang Chang, Yi-Ta Chen, An-Yeu Wu
ISCAS3
2024 BFP-CIM: Runtime Energy-Accuracy Scalable Computing-in-Memory-Based DNN Accelerator Using Dynamic Block-Floating-Point Arithmetic
abstract
Convolutional neural networks (CNNs) are known for their exceptional performance in various applications; however, their energy consumption during inference can be substantial. Analog Computing-In-Memory (CIM) has shown promise in enhancing the energy efficiency of CNNs, but the use of analog-to-digital converters (ADCs) remains a challenge. In analog CIM-based accelerators, ADCs convert analog partial sums from CIM crossbar arrays to digital values, with high-precision ADCs accounting for over 60% of the system’s energy consumption. To prevent ADCs from damaging the energy efficiency benefits of CIM, researchers have explored quantizing CNNs to use low-precision ADCs, trading off accuracy for energy efficiency. However, these approaches often necessitate data-dependent adjustments to minimize accuracy loss. Instead, we observe that the first most significant toggled bit indicates the optimal quantization range for each input value. Accordingly, we propose a range-aware rounding (RAR) method for runtime bit-width adjustment, eliminating the need for pre-deployment efforts. RAR can be easily integrated into a CIM accelerator using dynamic block-floating-point arithmetic. We also seamlessly incorporate a bit-level zero-skipping mechanism by dynamically forming input blocks. Experimental results demonstrate that our methods maintain accuracy while achieving up to 1.81$\bm{\times }$and 2.08$\bm{\times }$energy efficiency improvements on the CIFAR-10 and ImageNet datasets, respectively, compared with state-of-the-art techniques.
Cheng-Yang Chang, Chi-Tse Huang, Yu-Chuan Chuang, Kuang-Chao Chou, An-Yeu Wu
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 BWA-NIMC: Budget-based Workload Allocation for Hybrid Near/In-Memory-Computing
abstract
To enable efficient computation for convolutional neural networks, in-memory-computing (IMC) is proposed to perform computation within memory. However, the non-ideality significantly degrades the accuracy of IMC. In this work, we leverage a hybrid near/in-memory-computing architecture (NIMC) that allocates sensitive weights to error-free NMC and computes remained weights with high-efficient IMC. We further propose a Budget-based Workload Allocation for NIMC (BWA-NIMC). Specifically, we consider the resource difference between NMC and IMC to effectively allocate workloads under a targeted resource budget. Simulation results show that BWA-NIMC improves the accuracy by 18.38-48.54% under limited budgets (e.g., energy and latency) compared with prior works.
Chi-Tse Huang, Cheng-Yang Chang, Yu-Chuan Chuang, An-Yeu Wu
DAC1
2022 Automated Quantization Range Mapping for DAC/ADC Non-linearity in Computing-In-Memory
abstract
Computing-in-memory (CIM) has demonstrated the great potential of analog computing in improving the energy efficiency of matrix-vector multiplications for deep learning applications. Albeit low-power feature of CIM, the non-linearity of digital-to-analog converters (DACs)/analog-to-digital converters (ADCs) causes deviation between the computed outputs and desired values, thus degrading classification accuracy. This paper proposes Automated Quantization Range Mapping (A-QRM) mechanism to mitigate the negative effect of non-linearity on model accuracy. Instead of fixing the quantization range for quantized deep learning models, the proposed A-QRM automatically finds a better quantization range that balances the model capability and quantization errors caused by the non-linearity. Experimental results show that our proposed A-QRM achieves 89.02% and 86.93% of top-1 accuracy in ResNet20 and VGG8 on Cifar-10, respectively, under the non-linearity of DACs/ADCs.
Chi-Tse Huang, Yu-Chuan Chuang, Ming-Guang Lin, An-Yeu Wu
ISCAS1