Xuchu Huang

dblp:366/0548 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 VQT-CiM: Accelerating Vector Quantization Enhanced Transformer with Ferroelectric Compute-in-Memory
abstract
Transformer models have achieved state-of-the-art performance in various natural language processing (NLP) and computer vision (CV) tasks. To meet their substantial computational demands, the compute-in-memory (CiM) architectures, which alleviate the memory wall problem and enable efficient vector-matrix multiplication (VMM), have been adopted for transformer accelerators. However, the dynamic VMM involved in the attention mechanism, which necessitates runtime write operations, presents significant challenges for non-volatile memory (NVM)-based CiM designs. High write overhead, complex compute-write-compute (CWC) dependencies, and limited endurance reduce their effectiveness. In this paper, we propose VQT-CiM, a ferroelectric FET (FeFET)-based CiM design that accelerates vector quantization (VQ) enhanced transformers by eliminating the runtime write operations. VQT-CiM quantizes keys and values in self-attention to convert dynamic VMMs in inner-product and weighted-sum into static VMMs with the codebooks, enabling efficient calculations with CiM crossbars. However, directly applying VQ hinders the accuracy of transformer model due to its limited representation capability. To address this, we introduce a vector quantization scheme that integrates residual VQ (RVQ) and product VQ (PVQ) for enhanced representation space. We present an efficient hardware implementation for the proposed VQT-CiM with optimized dataflow in RVQ, which incorporates the FeFET-based CiM crossbars and peripheral digital circuits. Evaluation results suggest that VQT-CiM achieves the $3.54 \times$ and $4.53 \times$ improvements in energy efficiency and throughput, respectively, compared to state-of-the-art NVM-based CiM transformer designs.
Xuchu Huang, Haonan Du, Zheyu Yan, Cheng Zhuo, Xunzhao Yin
DAC1
2025 FactorHD: A Hyperdimensional Computing Model for Multi-Object Multi-Class Representation and Factorization
abstract
Neuro-symbolic artificial intelligence (neurosymbolic AI) excels in logical analysis and reasoning. Hyperdimensional Computing (HDC), a promising braininspired computational model, is integral to neuro-symbolic AI. Various HDC models have been proposed to represent class-instance and class-class relations, but when representing the more complex class-subclass relation, where multiple objects associate different levels of classes and subclasses, they face challenges for factorization, a crucial task for neuro-symbolic AI systems. In this article, we propose FactorHD, a novel HDC model capable of representing and factorizing the complex class-subclass relation efficiently. FactorHD features a symbolic encoding method that embeds an extra memorization clause, preserving more information for multiple objects. In addition, it employs an efficient factorization algorithm that selectively eliminates redundant classes by identifying the memorization clause of the target class. Such model significantly enhances computing efficiency and accuracy in representing and factorizing multiple objects with class-subclass relation, overcoming limitations of existing HDC models such as “superposition catastrophe” and “the problem of 2 “. Evaluations show that FactorHD achieves approximately $5667 \times$ speedup at a representation size of $10^{9}$ compared to existing HDC models. When integrated with the ResNet-18 neural network, FactorHD achieves $92.48 \%$ factorization accuracy on the Cifar-10 dataset.
Xuchu Huang, Chenyu Ni, Zheyu Yan, Xunzhao Yin, Cheng Zhuo
DAC2
2025 High-Performance In-Memory Bayesian Inference With Multi-Bit Ferroelectric FET
abstract
Conventional neural network-based machine learning algorithms often encounter difficulties in data-limited scenarios or where interpretability is critical. Conversely, Bayesian inference-based models excel with reliable uncertainty estimates and explainable predictions. Recently, many in-memory computing (IMC) architectures achieve exceptional computing capacity and efficiency for neural network tasks leveraging emerging nonvolatile memory (NVM) technologies. However, their application in Bayesian inference remains limited because the operations in Bayesian inference differ substantially from those in neural networks. In this article, we introduce a compact in-memory Bayesian inference engine with high efficiency and performance utilizing a multi-bit ferroelectric field-effect transistor (FeFET). This design encodes a Bayesian model within a compact FeFETbased crossbar by mapping quantized probabilities to discrete FeFET states. Consequently, the crossbar’s outputs naturally represent the output posteriors of the Bayesian model. Our design facilitates efficient Bayesian inference, accommodating various input types and probability precisions, without additional calculation circuitry. As the first FeFET-based in-memory Bayesian inference engine, our design demonstrates a notable storage density of 26.32 Mb/mm2and a computing efficiency of 581.40 TOPS/W in a representative Bayesian classification task, indicating a 10.7×/43.4× compactness/efficiency improvement compared to the state-of-the-art alternative. Utilizing the proposed Bayesian inference engine, we develop a feature selection system that efficiently addresses a representative NP-hard optimization problem, showcasing our design’s capability and potential to enhance various Bayesian inference-based applications. Test results suggest that our design identifies the essential features, enhancing the model’s performance while reducing its complexity, surpassing the latest implementation in operation speed and algorithm efficiency by 2.9×/2.0×, respectively.
Chao Li 0065, Xuchu Huang, Ruibin Mao, Thomas Kämpfe, Kai Ni 0004, Can Li 0024, Xunzhao Yin, Cheng Zhuo
IEEE Trans. Computers2
2025 SenHDC: A 3-D NAND Flash-Based Processing-in-Sensor Hyperdimensional Computing Architecture
abstract
The rapid growth of the Internet of Things (IoT) promotes vision application deployments in embedded devices. To alleviate the data conversion and transmission overheads in CMOS image sensor (CIS), processing-in-sensor (PIS) has been proposed, enabling computations to occur directly within sensory systems. Meanwhile, brain-inspired hyperdimensional computing (HDC) emerges as a promising computing paradigm well-suited for PIS due to its high accuracy and efficiency in various cognitive tasks. However, HDC requires transforming the sensed signals into long hypervectors (HVs), which imposes significant overhead for PIS systems. Thus, in this article, we propose SenHDC, a hardware-software co-design framework for HDC-based PIS. We introduce a novel hardware-friendly encoding paradigm that eliminates complex HV transformations, coupled with a noise-resilient position HV generation method that mitigates the impact of capacitance load imbalance. We further present an efficient 3-D NAND flash-based compute-in-memory (CIM) hardware design, comprising an encoding module that performs encoding in charge domain and an associative search module that checks the similarity between the input and prestored classes for classification. Experimental results show that SenHDC improves energy efficiency by$40.6\times $and$1.6\times $in the encoding module and associative search module, respectively, compared to the state-of-the-art HDC CIM implementations.
Xuchu Huang, Qingrong Huang, Zheyu Yan, Cheng Zhuo, Xunzhao Yin
IEEE Trans. Very Large Scale Integr. Syst.1
2024 Low Power and Temperature- Resilient Compute-In-Memory Based on Subthreshold-FeFET
abstract
Compute-in-memory (CiM) is a promising solution for addressing the challenges of artificial intelligence (AI) and the Internet of Things (IoT) hardware such as “memory wall” issue. Specifically, CiM employing nonvolatile memory (NVM) devices in a crossbar structure can efficiently accelerate multiply-accumulation (MAC) computation, a crucial operator in neural networks among various AI models. Low power CiM designs are thus highly desired for further energy efficiency optimization on AI models. Ferroelectric FET (FeFET), an emerging device, is attractive for building ultra-low power CiM array due to CMOS compatibility, high ION /$I$O F F ratio, etc. Recent studies have explored FeFET based CiM designs that achieve low power consumption. Nevertheless, subthreshold-operated FeFETs, where the operating voltages are scaled down to the subthreshold region to reduce array power consumption, are particularly vulnerable to temperature drift, leading to accuracy degradation. To address this challenge, we propose a temperature-resilient 2T-1FeFET CiM design that performs MAC operations reliably at subthreahold region from 0°C to 85°C, while consuming ultra-low power. Benchmarked against the VGG neural network architecture running the CIFAR-10 dataset, the proposed 2T1FeFET CiM design achieves 89.45% CIFAR-10 test accuracy. Compared to previous FeFET based CiM designs, it exhibits immunity to temperature drift at an 8-bit wordlength scale, and achieves better energy efficiency with 2866 TOPS/W.
Xuchu Huang, Jianyi Yang 0003, Kai Ni 0004, Hussam Amrouch, Cheng Zhuo, Xunzhao Yin
DATE2