Yung-Chun Lee

dblp:46/3378 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2026 In-3-D nand Flash Computing for Vector Similarity Search Acceleration on Edge Devices
abstract
Vector Similarity Search (VSS) on edge devices is increasingly essential for data privacy but faces substantial latency and energy overhead due to frequent data transfers between storage and DRAM. To address these challenges, we propose the Intelligent Cognition Engine (ICE), a fully digital non-volatile in-memory computing (nvIMC) framework designed for integration with commercial 3DnandFlash. ICE avoids the use of analog-digital conversion (ADC/DAC) and reduces data movement by performing vector similarity computation directly withinnandstorage. Key architectural components include a digital page multiplier, a two’s-complement accumulator that supports signed computation with minimal circuit modification, and a hierarchical Top-N search strategy to reduce unnecessary data accesses. The proposed framework is evaluated through a combination of post-layout circuit simulations, representative silicon measurements, and system-level modeling on edge platforms. Experimental results indicate that ICE can achieve 17.8–$122.6\times $speedup and 11.5–$162\times $improvement in energy efficiency compared to conventional von Neumann-based approaches, demonstrating the feasibility and scalability of fully digital 3Dnand-based nvIMC for edge AI workloads.
Han-Wen Hu, Yuan-Hao Chang 0001, Bo-Rong Lin, Huai-Mu Wang, Yung-Chun Lee, Hsiang-Pang Li, Chung Kuang Chen, Tei-Wei Kuo, Meng-Fan Chang
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 Accelerating Genome Alignment Pipeline with In-NAND Search Technology and Group Testing Techniques
abstract
Genomic sequence analysis deciphers and interprets an organism’s DNA, offering crucial insights into personalized medicine, disease diagnosis, evolutionary biology, and agricultural biotechnology. While Next-Generation Sequencing (NGS) has revolutionized genomics by providing a fast and cost-effective method for generating genomic sequences, the computational complexity of aligning short reads back to a reference genome remains a significant bottleneck. The exact-match-based preseeding filter has emerged as an effective and general methodology to address this issue, capable of removing 70% to 80% of exact-matched genomic reads at the source and applicable to a wide range of alignment tools. However, the state-of-the-art exact-match filter architecture, GenStore, encounters performance limitations due to the need to load reference sequences from NAND flash memory to the controller page by page.In this work, we propose a novel Solid-State Drive (SSD) architecture that leverages computing-in-NAND-flash techniques to perform match detection directly within memory. By harnessing the two-dimensional input capability of 3D NAND flash memory and integrating group testing methods, our design enables comparisons across hundreds of pages in a single read cycle and supports simultaneous multi-query searches. Combined with a Bloom filter for in-NAND search, our architecture significantly reduces data movement by 48% to 96%, achieves a speedup of 1.60× to 4.99× over GenStore, and delivers 30% higher energy efficiency with only a 4.5% circuit overhead.
Ming-Hsiang Tsai, Ming-Liang Wei, Chia-Chun Chien, Po-Hao Tseng, Yung-Chun Lee, Hsiang-Pang Li, Chia-Lin Yang
ICCAD5
2023 A digital 3D TCAM accelerator for the inference phase of Random Forest
abstract
Random forest is a popular ensemble machine-learning algorithm for classification and regression tasks. However, the irregular tree shapes and non-deterministic memory access patterns make it hard for the current von Neumann architecture to handle random forest efficiently. This paper proposes a digital 3D TCAM-based accelerator for the random forest, adopting the idea of processing-in-memory (PIM) to reduce data movement. By utilizing this accelerator, we propose a TCAM-based approach to provide real-time inference with low energy consumption, making it suitable for edge or embedded environments. In the experiments, the proposed approach achieves an average of 3.13 times higher throughput with 22 times more energy saving than the GPU approach.
Chieh-Lin Tsai, Chun-Feng Wu, Yuan-Hao Chang 0001, Han-Wen Hu, Yung-Chun Lee, Hsiang-Pang Li, Tei-Wei Kuo
DAC5
2022 ICE: An Intelligent Cognition Engine with 3D NAND-based In-Memory Computing for Vector Similarity Search Acceleration
abstract
Vector similarity search (VSS) for unstructured vectors generated via machine learning methods is a promising solution for many applications, such as face search. With increasing awareness and concern about data security requirements, there is a compelling need to store data and process VSS applications locally on edge devices rather than send data to servers for computation. However, the explosive amount of data movement from NAND storage to DRAM across memory hierarchy and data processing of the entire dataset consume enormous energy and require long latency for VSS applications. Specifically, edge devices with insufficient DRAM capacity will trigger data swap and deteriorate the execution performance. To overcome this crucial hurdle, we propose an intelligent cognition engine (ICE) with cognitive 3D NAND, featuring non-volatile in-memory computing (nvIMC) to accelerate the processing, suppress the data movement, and reduce data swap between the processor and storage. This cognitive 3D NAND features digital nvIMC techniques (i. e., ADClDAC-free approach), high-density 3D NAND, and compatibility with standard 3D NAND products with minor modifications. To facilitate parallel INT8/INT4 vector-vector multiplication (VVM) and mitigate the reliability issue of 3D NAND, we develop a bit-error-tolerance data encoding and a two’s complement-based digital accumulator. VVM can support similarity computations (e.g., cosine similarity and Euclidean distance), which are required to search “the most similar data” right where they are stored. In addition, the proposed solution can be realized on edge storage products, e.g., embedded Multi-Media Card (eMMC). The measured and simulated results on real 3D NAND chips show that ICE enhances the system execution time by $17\times to 95\times$ and energy efficiency by $11\times to 140\times$, compared to traditional von Neumann approaches using state-of-the-art edge systems with MobileFaceNet on CASIA-WebFace dataset. To the best of our knowledge, this work demonstrates the first 3D NAND-based digital nvIMC technique with measured silicon data.
Han-Wen Hu, Wei-Chen Wang 0002, Yuan-Hao Chang 0001, Yung-Chun Lee, Bo-Rong Lin, Huai-Mu Wang, Yen-Po Lin, Chong-Ying Lee, Tzu-Hsiang Su, Chih-Chang Hsieh, Chia-Ming Hu, Yi-Ting Lai, Chung Kuang Chen, Han-Sung Chen, Hsiang-Pang Li, Tei-Wei Kuo, Meng-Fan Chang, Keh-Chung Wang, Chun-Hsiung Hung, Chih-Yuan Lu
MICRO4