Zihao Xuan

dblp:307/9783 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-4573-4556ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SeDA: Secure and Efficient DNN Accelerators with Hardware/Software Synergy
abstract
Ensuring the confidentiality and integrity of DNN accelerators is paramount across various scenarios spanning autonomous driving, healthcare, and finance. However, current security approaches typically require extensive hardware resources, and incur significant off-chip memory access overheads. This paper introduces SeDA, which utilizes 1) a bandwidth-aware encryption mechanism to improve hardware resource efficiency, 2) optimal block granularity through intra-layer and inter-layer tiling patterns, and 3) a multi-level integrity verification mechanism that minimizes, or even eliminates, memory access overheads. Experimental results show that SeDA decreases performance overhead by over 12% for both server and edge neural processing units (NPUs), while ensuring robust scalability.11SeDA source code:https://github.com/wayne4s/seda.git
Lang Feng 0001, Ning Lin, Zihao Xuan, Rongliang Fu, Tsung-Yi Ho, Yuzhong Jiao, Luhong Liang
DAC5
2025 YOCO: A Hybrid In-Memory Computing Architecture with 8-bit Sub-PetaOps/W In-Situ Multiply Arithmetic for Large-Scale AI
abstract
In this paper, we further explore the potential of analog in-memory computing (AiMC) and introduce an innovative artificial intelligence (AI) accelerator architecture named YOCO, featuring three key proposals: (1) YOCO proposes a novel 8-bit in-situ multiply arithmetic (IMA) achieving 123.8 TOPS/W energy-efficiency and 34.9 TOPS throughput through efficient charge-domain computation and time-domain accumulation mechanism. (2) YOCO employs a hybrid ReRAM-SRAM memory structure to balance computational efficiency and storage density. (3) YOCO tailors an IMC-friendly attention computing flow with an efficient pipeline to accelerate the inference of transformer-based AI models. Compared to three SOTA baselines, YOCO on average improves energy efficiency by up to $3.9 \times \sim 19.9 \times$ and throughput by up to $6.8 \times \sim 33.6 \times$ across $10 \mathrm{CNN} /$ transformer models.
Zihao Xuan, Yuxuan Yang 0009, Zijia Su, Song Chen 0001, Yi Kang
DAC1
2025 A dynamic decoder with speculative termination for low latency inference in spiking neural networks
Zihao Xuan, Yi Kang
Neurocomputing2
2025 A neuromorphic hardware architecture based on TTFS coding with temporal quantization for spiking neural networks
Yuxuan Yang 0009, Qihu Xie, Zihao Xuan, Song Chen 0001, Yi Kang
Integr.3
2025 HNM-CIM: An Algorithm-Hardware Co-designed SRAM-based CIM for Transformer Acceleration Exploiting Hybrid N:M Sparsity
abstract
SRAM-based computing-in-memory (CIM) is an efficient technology for computing neural networks where matrix operations are dominated. However, leveraging sparsity in CIM presents challenges due to the crossbar architecture, which complicates the avoidance of zero element calculations. Previous CIM designs have demonstrated that sparsity can improve energy efficiency, but these approaches often lead to non-negligible accuracy loss or substantial hardware overhead. To address this challenge, we propose a hybrid N:M CIM (HNM-CIM), an algorithm-architecture co-design framework for accelerating Transformers. At the algorithm level, we propose a hybrid N:M pruning (HNMP), a method that combines structured and unstructured sparsity. This approach maintains regularity while preserving the random distributions of sparsity, thereby enhancing model sparsity with negligible accuracy loss and ensuring CIM compatibility. At the hardware level, we introduce a hybrid N:M sparse digital CIM (HNM-CIM) to support HNMP, which can accelerate Transformers with hybrid N:M sparsity patterns. Experimental results show that HNMP can reduce Transformer models by about 3.1× on model size with negligible accuracy loss. Compared with state-of-the-art references, HNM-CIM yields about 2.46× speed up and 1.43× area savings.
Yuang Ma, Yulong Meng, Zihao Xuan, Song Chen 0001, Yi Kang
ACM Trans. Design Autom. Electr. Syst.3
2024 TQ-TTFS: High-Accuracy and Energy-Efficient Spiking Neural Networks Using Temporal Quantization Time-to-First-Spike Neuron
abstract
In recent years, spiking neural networks (SNNs) have gained attention for their biological realistic and event-driven characteristics, which align well with neuromorphic hardware. Time-to-First-Spike (TTFS) coding is an coding scheme for SNNs, where neurons are fired only once throughout the inference process, reducing the number of spikes and improving energy efficiency. However, the SNNs with TTFS coding face the issue of low classification accuracy. This paper introduces TQ-TTFS, a temporal quantization on TTFS neuron model to address this issue. In addition, the temporal quantization neurons can apply lower clock frequency without increasing inference latency, which can lead to higher energy efficiency. The experimental results show the effectiveness of the proposed temporal quantization neuron model in improving both classification accuracy and energy efficiency. In our simulations TQ-TTFS achieves classification accuracy of 98.6% on MNIST dataset and 90.2% on FashionMNIST dataset which are among SOTA of temporal coding SNNs. An analysis is also given to show that TQ-TTFS on an example SNN can have $2.94 \times$ energy efficiency improvement compared with tranditional TTFS coding.
Zihao Xuan, Yi Kang
ASPDAC2
2022 HPSW-CIM: A Novel ReRAM-Based Computing-in-Memory Architecture with Constant-Term Circuit for Full Parallel Hybrid-Precision-Signed-Weight MAC Operation
abstract
Non-volatile memory (NVM) based computing-- (CIM) systems can provide low-latency and high-efficiency parallel multiply-accumulate (MAC) operations, which shows great potential in accelerating edge AI computing. In response to the limitations of miniaturization, parallelism, and low-power consumption of edge AI devices, this work proposes: 1. A hybrid-precision-signed-weight ITIR (HPSWITIR) sub-array realizes the signed-weight by constant-term resistor circuit without additional ReRAM array overhead. 2. A computing array reduces the IR-drop and transistor errors of the array by assembling a series of HPSW-1T1R sub-arrays and achieve high-density array integration. 3. Global-share ADCs (GS-ADCs) optimize the analog signals conversion process to improve computing parallelism. 4. Odd-Channel-Input-WeightInverse Coding (ORIWI) can further reduce IR drop by decreasing the accumulative SL current (ISL) inside the HPSWITIR sub-array. In these manners, this work constructed a high computing-density HPSW-CIM calculation core for fully parallel MAC operation. The circuit-level evaluation shows that the energy cost of data conversion reduced by 2×-2.7×, reduces the peak column average current by 3.1×-6.2×, and improves array computing density by at least twice. The system-level evaluation shows that, in an HPSW-CIM core with 256KB memory in the 28nm process, the peak energy efficiency reaches 95.8TOPS/W@8bIN-4bW-8bO.
Zihao Xuan, Yi Kang
ISCAS1
2022 High-Efficiency Data Conversion Interface for Reconfigurable Function-in-Memory Computing
abstract
Recently, analog in-memory computing (IMC) systems exhibit the considerable potential to break through the inherent high computational latency and energy cost of Von Neumann’s computer architecture. However, inefficient data convertor will inhibit the performance improvement of this system. The tradeoff between different data conversion circuit technologies has turned into one of the major driving forces for the analog IMC system-level improvements. The primary contribution is in two aspects. First, this article shows a digital-to-time-to-analog converter (DTAC) with the tradeoff of latency, area, and power consumption compared to a digital-to-time converter (DTC) and digital-to-analog converter (DAC). Second, we develop an innovative reconfigurable joint-quantization nonlinear analog-to-digital convertor (JQNL-ADC) architecture with lower quantization error by merging the two paradigms of uniform input quantization and uniform output quantization. Compared to conventional DAC, DTAC can reduce power and area by$50\times $and$3\times $, respectively. Compared to SAR-ADC, our JQNL-ADC can reduce area and power by$1.6\times $and$2\times $, respectively. In an example of ReRAM-based reconfigurable function-IMC (RFIMC) macro with 256-kb memory, our design can reach 112.9 TOPS/W@8bIN-8bW-8bO under the 28-nm process conditions.
Zihao Xuan, Yi Kang
IEEE Trans. Very Large Scale Integr. Syst.1