EDBT 2026 Demo / reviewers in the wild / expert
Rui Xiao 0003
dblp:94/1463-3
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-3078-0178ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A FeFET-Based Compute-in-Memory Architecture on FPGA for Neural Network InferenceabstractImplementing compute-in-memory (CIM) architectures on FPGA offers an effective solution to the von Neumann bottleneck by enabling fast configuration and computation directly within memory. Traditional custom solutions rely on the modification of block RAM (BRAM) to implement memory computing. However, single-word-line activation of BRAM results in low parallelism, and the need for additional adder trees to accumulate partial sums further limits efficiency. To overcome these limitations, we propose a CIM core based on a 2T1C structure as a replacement for BRAM units. This core utilizes a charge redistribution mechanism and reuse of ADC capacitors, achieving high parallelism, low power consumption, and a compact area. By incorporating computational capabilities within a single cell, our design enables dual parallelism, further enhancing performance and efficiency. In addition, we present an automated deployment and mapping tool for deep neural networks (DNNs) on FPGA, allowing users to rapidly develop FPGA-based solutions for different network architectures. Compared to state-of-the-art solutions, our design achieves a peak throughput improvement of 4.5× and a reduction in area by 53%. Minghan Jiang, Yonggen Li, Rui Xiao 0003, Haibin Shen, Kejie Huang |
FCCM | 3 |
| 2025 | A 1FeFET-1T-1C based Compute-in-Memory Macro with Capacitor Reused Pipeline SAR ADCabstractComputing-in-memory (CIM) significantly reduces latency and power consumption by combining computation and memory, typically utilizing non-volatile memories (NVM). However, device manufacturing non-uniformity on NVMs can cause output deviations. Additionally, the necessity for bit-shifting circuits and Analog-to-Digital Converters (ADC) increases the area and power overhead. To tackle these challenges, we propose a high-density 1FeFET-1T-1C based CIM macro, integrated with a pipeline Successive-Approximation-Register (SAR) ADC. The design introduces a capacitor structure that counters the non-uniformity issues inherent in FeFET devices. Also, the capacitor array is reused as charge-redistribution and ADCs, substantially minimizing the area and power overhead. Moreover, the pipeline architecture accelerates the conversion process, achieving high speed and high precision. The design is implemented using SMIC 55nm PDK. The energy efficiency (EF) and area efficiency (AF) of the proposed macro are 80.9 TOPS/W and 1.161 TOPS/mm2, respectively. The inference accuracy reaches 91.2% on the CIFAR-10 dataset. Minghan Jiang, Rui Xiao 0003, Shuaiting Li, Yishu Zhang, Haibin Shen, Kejie Huang |
ISCAS | 2 |
| 2025 | A Robust Computing-in-Memory Macro With 2T1R1C Cells and Reused Capacitors for Successive-Approximation ADCabstractComputing-in-memory (CIM) has emerged as a practical paradigm to bypass the von Neumann bottleneck. However, traditional CIM schemes face challenges due to the nonideal characteristics of nonvolatile memory (NVM). To address this issue, this work provides a resistive random access memory (RRAM)-based CIM macro employing two-transistor-one-RRAM–one-capacitor (2T1R1C) cells, with capacitors reused for the successive-approximation analog-to-digital converter (SAR ADC). Single-level RRAM is utilized to mitigate resistance variation. The multiply-accumulate (MAC) operation is performed via the charge and discharge of capacitors, enhancing robustness across different process, voltage, and temperature (PVT) corners. The capacitors in 2T1R1C cells are repurposed as sampling capacitors to integrate the ADC with the array. A precision-adjustable SAR (PA-SAR) logic is proposed to generate partial sums at varying precision levels aligned with different input bits, optimizing energy efficiency while maintaining reliability. Our proposed 2T1R1C array features an average area of$3.403~\mu $m2 for each cell, which accounts for 87.46% of the total macro area. The total macro area is 1.020 mm2 with a capacity of 256 Kb, achieving an energy density of 0.201 TOPS/mm2. The PA-SAR logic boosts energy efficiency to 44.71 TOPS/W, marking a 38.55% improvement over conventional full-precision schemes. Rui Xiao 0003, Minghan Jiang, Haibin Shen, Kejie Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | A Low-Power In-Memory Multiplication and Accumulation Array With Modified Radix-4 Input and Canonical Signed Digit WeightsabstractData transfer between the processing and storage units has become a significant bottleneck in modern von Neumann computing systems for artificial intelligence (AI) tasks. Computing in memory (CIM) has emerged as a promising candidate for lowering latency and power consumption. However, the conventional analog CIM schemes are suffering from reliability issues, which may significantly degenerate the accuracy of the computation. Recently, digitized input data and weights have been utilized for high-reliable in-memory computing. However, the properties of the digital memory and input data are not fully utilized. This article presents a novel low-power CIM scheme to further reduce the power consumption by using a modified radix-4 (M-RD4) booth algorithm at the input and a modified canonical signed digit (M-CSD) for the network weights. The simulation results show that M-RD4 and M-CSD reduce the number of nonzero activation bits by 24.2% and the number of nonzero weight bits by 36.0% in AlexNet, respectively. The power consumption can be reduced by 41.6% on average. The computing-power ratio at the fixed-point 8 bit is 60.7 tera operations per second per watt (TOPS/W), and the density is 0.177 TOPS/mm2. Rui Xiao 0003, Yewei Zhang, Bo Wang 0020, Yanfeng Xu, Jicong Fan 0002, Haibin Shen, Kejie Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | An 8-Bit in Resistive Memory Computing Core With Regulated Passive Neuron and Bitline Weight MappingabstractThe rapid development of artificial intelligence (AI) and Internet of Things (IoT) increase the requirement for edge computing with low power and relatively high processing speed devices. The computing-in-memory (CIM) schemes based on emerging resistive nonvolatile memory (NVM) show great potential in reducing the power consumption for AI computing. However, the inconsistency of the NVM may significantly degenerate the performance of the neural network. In this article, we propose a low power resistive RAM (RRAM)-based CIM core to not only achieve high computing efficiency but also greatly enhance the robustness by bit line (BL) regulator and BL weight mapping algorithm. The simulation results show that the power consumption of our proposed 8-bit CIM core is only 12.6 mW ($256\times 256$at 8b). The spurious-free dynamic range (SFDR) and signal to noise and distortion ratio (SNDR) of the CIM core achieve 62.64 and 45.92 dB, respectively. The proposed BL weight mapping scheme improves the top-1 accuracy by 2.46% and 3.47% for AlexNet and VGG16 on ImageNet Large Scale Visual Recognition Competition 2012 (ILSVRC 2012) in 8-bit mode, respectively. Yewei Zhang, Kejie Huang, Rui Xiao 0003, Bo Wang 0020, Yanfeng Xu, Jicong Fan 0002, Haibin Shen |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |