VLDB 2026 Research / reviewers in the wild / expert
Xiangshui Miao
dblp:71/11430
· DBLP profile ↗
32ranked-venue papers
0as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 11 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MemSearch: An Efficient Memristive In-memory Search Engine with Configurable Similarity Measures
Yingjie Yu, Houji Zhou, Jiancong Li, Tong Hu, Jia Chen 0032, Yi Li 0049, Xiangshui Miao |
ASP-DAC | 7 |
| 2026 | OAH-CIM: Outlier-Aware Hybrid RRAM-SRAM CIM Accelerator with Variation-Robust Sparsity
Tong Hu, Han Bao 0013, Houji Zhou, Yuyang Fu, Jiancong Li, Jia Chen 0032, Yi Li 0049, Xiangshui Miao |
ASP-DAC | 9 |
| 2026 | Delta-STDP: Enabling Hardware-Friendly Supervised Learning in Spiking Neural NetworksabstractThe deployment of supervised learning in Spiking Neural Networks (SNNs) on energy-efficient hardware is hindered by the computational complexity of gradient-based algorithms like SpikeProp, involving intricate temporal derivative calculations and non-parallelizable operations. This paper introduces Delta-STDP, a novel learning rule through hardware-algorithm co-design. We simplify the postsynaptic potential (PSP) waveform from a complex alpha function to a ReLU shape, transforming its temporal derivative into a binary causal mask. Concurrently, we binarize the error signal to its sign, resulting in a weight update rule that factorizes into a global 1-bit error direction and a local spike-timing-dependent amplitude. Information-theoretic analysis shows ReLU optimally shifts information from error to PSP, minimizing loss from error sign binarization, while MNIST results confirm that under sign-based updates, ReLU outperforms alpha and exponential-rise PSPs, reversing their order with full-precision updates. Furthermore, we outline a direct mapping to hardware, where forward propagation simplifies to integrate-and-fire with square-pulse inputs, backpropagation becomes a causally gated memristor crossbar operation, and weight update translates to sign-modulated spike-timing-dependent plasticity (STDP). This work bridges the gap between algorithmic effectiveness and hardware efficiency, providing a pathway towards scalable neuromorphic learning accelerators. Dayou Zhang, Xiangshui Miao, Yuhui He |
ACM Great Lakes Symposium on VLSI | 4 |
| 2026 | A Scaling Annealing Method for Combinatorial Optimization in Asynchronous Memristive Hopfield Network
Han Bao 0013, Kehong Xu, Yibai Xue, Yuyang Fu, Jiancong Li, Jia Chen 0032, Yi Li 0049, Xiangshui Miao |
ISCAS | 9 |
| 2026 | SPEAR: Spike-Aware Point Pruning and Fused FPS-KNN Sampling for Spiking Point Transformer Acceleration Co-Design
Yilong Fang, Pinfeng Jiang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ISCAS | 6 |
| 2026 | Energy-Efficient Acceleration of Fourier-based Transformers on RRAM-CIM via Mixed-Precision and DFT Symmetry
Yuyang Fu, Jiancong Li, Yi Li 0049, Xiangshui Miao |
ISCAS | 5 |
| 2026 | Compute-in-memory compatible ANN-SNN conversion via ternary nonpolar neuron
Dayou Zhang, Xinyu Wen, Xiangshui Miao, Yuhui He |
Neurocomputing | 5 |
| 2026 | Memristor-Based Circuit Implementation and Circuitry Optimized Algorithm for Mamba Language NetworkabstractLanguage networks are crucial in artificial intelligence, with the novel Mamba architecture significantly reducing computations and consumption compared to the traditional transformer network. However, a full-circuit implementation of the Mamba network has not been proposed due to the complexity of computations and data storage. Additionally, optimized hardware-aware parallel algorithms for Mamba inference in circuits remain undeveloped. This work addresses these challenges by presenting a memristor-based full-circuit implementation of the Mamba network and introducing a computing-in-memory parallel-aware algorithm tailored for circuit-level inference. The implementation includes: 1) Standard 1T1M memristor crossbar and depthwise separable convolution memristor crossbar for different convolutions. 2) Computing-in-memory implicit latent state circuits for the computation and transition of latent states. 3) Functional circuits for SiLU activation, RMS normalization, and multi-layer multiply-accumulate operations. 4) Optimized algorithm and circuit implementation for hardware-aware inference, achieving parallel scanning and hardware awareness in circuits. The proposed circuit enables analog signal computations and eliminates redundant analog-to-digital conversions and intermediate storage. A basic single-sentence generation task was simulated in PSPICE, validating the circuit’s correctness. Analyses of analog computation accuracy, circuit stability, and power consumption demonstrate the proposed circuit’s advantages, highlighting its potential as a fundamental module for large-scale circuit integration and complex text generation tasks. Zheyuan Sheng, Huajun Sun, Chuanbo Zhu 0002, Xiangshui Miao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | ReSMiPS: A ReRAM-based Sparse Mixed-precision Solver with Fast Matrix Reordering AlgorithmabstractThe solution of sparse matrix equations is essential in scientific computing. However, traditional solvers on digital computing platforms are limited by memory bottlenecks in largescale sparse matrix storage and computation. Resistive Random Access Memory (ReRAM)-based computing-in-memory (CIM) offers a promising solution to this challenge but faces constraints in achieving high solution precision and energy efficiency simultaneously in sparse matrix computations. In this work, we propose ReSMiPS, a ReRAM-accelerated sparse mixed-precision solver. ReSMiPS incorporates a novel Fast Sparse Matrix Reordering (FSMR) algorithm and introduces an In-memory float64 (IF64) data format, enabling efficient floating-point sparse matrix computation directly within the analog ReRAM array. By combining our floating-point CIM macro design with a hybrid-domain solution framework, ReSMiPS achieves precision comparable to CPU and GPU-based BiCGSTAB solvers (with errors below $10^{-15}$) on real-world workloads, while demonstrating over two orders of magnitude improvement in both computational speed and energy efficiency. Yuyang Fu, Jiancong Li, Jia Chen 0032, Houji Zhou, Wenlong Peng, Yi Li 0049, Xiangshui Miao |
DAC | 8 |
| 2025 | SPARTA: Spike-Aware Token Skipping Co-Optimization with Heterogeneous ReRAM-CIM Architecture for Spiking Transformer AccelerationabstractSpiking neural networks (SNNs) have emerged as a promising paradigm for energy-efficient neural computation, offering advantages over artificial neural networks (ANNs) by using sparse spike activations. Among SNN models, spiking transformers have shown great potential in achieving high accuracy while maintaining low energy consumption, making them ideal for resource-constrained applications. However, a significant challenge in accelerating Spiking Transformers lies in managing the inherent unstructured sparsity in spike activations. This sparsity introduces substantial hardware overhead, limiting the efficiency of existing accelerators.To address these challenges, we propose SPARTA, an algorithm-hardware co-optimization framework designed specifically for spiking transformers. SPARTA utilizes a spike-aware dynamic token skipping algorithm, which applies reinforcement learning to selectively skip less informative tokens, achieving structured token-level sparsity in both spatial and temporal domains. Additionally, we introduce the spike-aware token prediction algorithm to predict and eliminate inactive tokens, further improving efficiency. Meanwhile, we present a dedicated heterogeneous ReRAM-based Compute-in-Memory (CIM) hardware architecture tailored to support token-level sparsity, which integrates a ReRAM analog CIM engine for linear layer and a token-spike fusion engine for optimized token routing and spike attention. Experimental results show that SPARTA achieves up to 543.1× and 10.2× speedup with 308.0× and 5.2× energy efficiency improvement compared to GPU and the state-of-the-art SNN accelerator "COMPASS", while preserving high model accuracy. Pinfeng Jiang, Yilong Fang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ICCAD | 6 |
| 2025 | ASMA: An Anisotropy Scaling Memristor-Based Accelerator for LLM InferenceabstractLarge Language Models (LLMs) present substantial computational and memory challenges, particularly within their Feed-Forward Network (FFN) layers. While Computing-in-Memory (CIM) using Memristor offers a path to mitigate data movement bottlenecks, naively scaling existing CIM architectures for LLMs leads to new, severe communication overheads due to their typically isotropic (uniform) design. This paper introduces ASMA, an Anisotropy Scaling Memristor-based Accelerator, specifically co-designed for efficient LLM inference. ASMA identifies and exploits the inherent anisotropic scaling characteristics of LLM FFN layers—categorized into additive, concatenative, and sequential dimensions. Key innovations include a novel Subtile hierarchy for optimized pipelining, a hierarchical anisotropic Network-on-Chip (NoC) featuring distinct channels for broadcast and localized communication, and a co-designed compiler that maps LLM computations by leveraging these architectural specializations. Evaluations using a transaction-level simulator demonstrate that ASMA achieves latency improvement up to 81.7% and energy improvement up to 91.5% compared to conventional hierarchical CIM designs. ASMA offers a new paradigm for designing LLM-specific CIM accelerators by embracing workload anisotropy. Zijian Xiong, Yaoyu Tao, Xiangshui Miao, Yuhui He |
ICCD | 7 |
| 2025 | A Novel Memristor-Based Majority Logic and Efficient Approximate Full Adder for Image ProcessingabstractThe memristor, an emerging non-volatile memory device, is well-suited for in-memory computing (IMC) due to its capability to simultaneously store data and perform logic operations. Majority (MAJ) logic is a type of expressive Boolean logic that performs well in logic synthesis. This paper proposes a novel MAJ logic implementation scheme to improve the latency of n-bit full adder (FA) from 6n+1 to 5n+1. Additionally, an approximate computing approach is introduced, yielding an n-bit approximate adder with a significantly reduced delay of 2n+1, at the cost of precision. Both exact and approximate adders were verified experimentally, and the error quality metrics of the approximate design were thoroughly evaluated. Furthermore, hybrid configurations combining exact and approximate adders were applied to image processing tasks, where the peak signal-to-noise ratio (PSNR) for most hybrid designs remained within an acceptable range (>30dB). Zhouchao Gan, Yifeng Xiong, Fan Yang 0148, Xiangshui Miao, Xingsheng Wang |
ISCAS | 5 |
| 2025 | ReBA: A Hybrid Sparse Reconfigurable Butterfly Accelerator for Solving Partial Differential Equations via Hardware and Algorithm Co-DesignabstractPartial differential equations (PDEs) are widely used in many scientific and engineering fields. Traditional numerical methods for solving PDEs often struggle with complex equations and require extensive computation. Recent advances in neural operators, particularly the Galerkin Transformer (GT-FNO), offer faster solutions but present unique challenges for hardware due to complex neural network operations such as attention, Fourier transforms, and convolution. To address these challenges, we propose ReBA, a dedicated hardware and algorithm co-design framework specified for GT-FNO. Our design implements hybrid sparsity schemes for efficient workload balance, significantly reducing hardware overhead while maintaining high accuracy. ReBA introduces a reconfigurable dual-engine architecture, featuring an adaptable butterfly processing element (BPE), which efficiently supports FFT and other DNN operations through flexible BPE allocation, offering high parallelism and low-latency computation. Experiments validate that ReBA delivers significant speedups of 34.57×, 2.70×, and 1.82 51.26× over conventional CPUs, GPUs, and prior state-of-the-art accelerators in solving PDEs. Pinfeng Jiang, Yilong Fang, Ming-De Zhu, Xiangshui Miao, Xingsheng Wang |
ISCAS | 6 |
| 2025 | Optimizing hardware-software co-design based on non-ideality in memristor crossbars for in-memory computing
Pinfeng Jiang, Danzhe Song, Menghua Huang, Fan Yang 0148, Xiangshui Miao, Xingsheng Wang |
Sci. China Inf. Sci. | 7 |
| 2025 | 3D self-rectifying memristive ternary content addressable memory for massive and exact in-memory search
Yingjie Yu, Shengguang Ren, Yi Li 0049, Xiangshui Miao |
Sci. China Inf. Sci. | 5 |
| 2025 | ArPCIM: An Arbitrary-Precision Analog Computing-in-Memory Accelerator With Unified INT/FP ArithmeticabstractAnalog Computing-in-memory (ACIM) breaks the von Neumann bottleneck and significantly improves energy efficiency by enabling parallel matrix-vector-multiplication (MVM) operations. However, most existing ACIM accelerators are highly customized for specific precision formats, lacking the generality to support efficiently arbitrary precision in both integer (INT) and floating-point (FP) formats. In this work, we present an arbitrary precision analog computing-in-memory (ArPCIM) accelerator with unified INT/FP arithmetic to address this limitation. We introduce a CIM-friendly INT/FP arithmetic to convert FP numbers into INT numbers for efficient execution on CIM, minimizing precision loss through local pre-alignment and dynamic bit-weight slicing methods. In addition, we implement multi-level reconfigurable precision circuits, featuring both intra- and inter-processing element (PE) reconfigurability, which supports precision ranging from 1-bit to 47-bit. Experimental results show that our ArPCIM accelerator achieves up to$4.09\times $and$5.37\times $improvement in energy efficiency and area efficiency, respectively, compared with state-of-the-art arbitrary precision digital CIM. Our ArPCIM accelerator offers the flexibility to meet diverse computational needs while maintaining high energy efficiency and accuracy, paving the way for versatile CIM acceleration across various fields. Jiancong Li, Han Jia, Houji Zhou, Han Bao 0013, Yuyang Fu, Yi Li 0049, Xiangshui Miao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2025 | ISARA: An Island-Style Systolic Array Reconfigurable Accelerator Based on Memristors for Deep Neural NetworksabstractThe demand for edge artificial intelligence (AI) is significant, particularly in revolutionary technological areas such as the Internet of Things, autonomous driving, and industrial control. However, reliable and high-performance edge AI is still constrained by computing hardware, and improving the performance and reliability of edge AI accelerators remains a key focus for researchers. This work proposes a memristor/resistive random access memory (RRAM)-based island-style systolic array reconfigurable accelerator (ISARA) that meets the reliability and performance requirements of edge AI. Inspired by the island-style architecture of FPGAs, this work proposes a flexible-tile architecture based on RRAM processing element (PE) islands, optimizing the data flow within the systolic array. The design of network-on-chip reduces data processing latency. In addition, to enhance computational efficiency, this work incorporates a bit-fusion scheme within the flexible tile, which reduces analog-to-digital converter (ADC) power consumption and addresses the conductance variation of RRAM. To date, only a few works have completed the entire process from simulation, design, and fabrication to hardware testing. This work fully realizes the design and validation of a new accelerator based on RRAM chips, demonstrating the reliability of RRAM-based systolic array accelerators for the first time. After deploying algorithms, the hardware accelerator achieved recognition rates comparable to software. Compared to similar works, ISARA’s computational efficiency exceeds theirs and has flexible reconfigurability. The same deep neural network (DNN) models are adopted for evaluation and compared to other accelerators, and ISARA’s processing latency is reduced by 200 times. Fan Yang 0148, Pinfeng Jiang, Xiangshui Miao, Xingsheng Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | Multi-order Differential Neural Network for TCAD Simulation of the Semiconductor DevicesabstractTechnology Computer Aided Design (TCAD) is a crucial step in the design and manufacturing of semiconductor devices. It involves solving physical equations that describe the behavior of semiconductor devices to predict various device parameters. Traditional TCAD methods, such as finite volume and finite element methods, discretize relevant physical equations to achieve numerical simulations of devices, significantly burdening the computation resources. For the first time, this paper proposes a novel method for TCAD simulation based on Physics-Informed Neural Networks (PINNs). We proposed multi-order differential neural network (MDNN), an improved Radial Basis Function Neural Network (RBFNN) model. By training MDNN, it achieves the couple solution of the Poisson equation and drift-diffusion equation under steady-state conditions, without the need for a pre-existing dataset. To the best of our knowledge, this marks the first instance of an ML-TCAD simulation that does not require any pre-existing data. For an example of PN junction diode, this method effectively simulates the basic physical characteristics of the device, with a self-consistent solution error of less than 1×10−5. Zifei Cai, Anaoxue Huang, Yifeng Xiong, Dejiang Mu, Xiangshui Miao, Xingsheng Wang |
DAC | 5 |
| 2024 | Efficient implementation of majority-inverter graph logic and arithmetic functions with memristor arrays
Zhouchao Gan, Yinghao Ma, Xiangshui Miao, Xingsheng Wang |
Sci. China Inf. Sci. | 5 |
| 2024 | Mitigating set-stuck failure in 3D phase change memory: substituting square pulses with surge pulses
Ninghua Li, Wang Cai, Weiming Cheng, Xiangshui Miao |
Sci. China Inf. Sci. | 6 |
| 2024 | A self-selecting memory element based on a method of interconnected ovonic threshold switching device
Jinyu Wen, Jiangxi Chen, Xiangshui Miao |
Sci. China Inf. Sci. | 5 |
| 2024 | Energy Efficient Memristive Transiently Chaotic Neural Network for Combinatorial OptimizationabstractThe utilization of memristive analog-digital mixed in-memory computing has significantly tackled the issues of massive computing resources and time delays in solving combinatorial optimization problems. However, further improvements in computing energy efficiency are still desirable for resource-constrained conditions and practical applications. Therefore, in this work, a memristive analog transiently chaotic neural network (TCNN) system is proposed to solve the traveling salesman problems (TSPs), which is consisted of 1) a memristor array to perform matrix-vector multiplication for the network iteration; 2) neuronal modules and nonlinear activation function modules to emulate the basic functions of the network; 3) chaotic simulated annealing (CSA) modules to improve the solution performance at extremely low hardware overhead. Based on these, after mapping the TSPs onto the memristor array, the proposed TCNN can self-iterate to convergence states to solve the problems with high performance. Compared to prior analog-digital mixed ones, the analog system can eliminate the extra control and data conversions during the network solving process, and achieve an$8\times $reduction in the total energy consumption in consecutive solving tasks. Han Bao 0013, Kehong Xu, Houji Zhou, Jiancong Li, Yi Li 0049, Xiangshui Miao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2024 | In-Memory Wallace Tree Multipliers Based on Majority Gates Within Voltage-Gated SOT-MRAM Crossbar ArraysabstractIn-memory computing represents an efficient paradigm for high-performance computing using crossbar arrays of emerging nonvolatile devices. While various techniques have emerged to implement Boolean logic in memory, the latency of arithmetic circuits, particularly multipliers, significantly increases with bit-width. In this work, we introduce an in-memory Wallace tree multiplier based on majority gates within voltage-gated spin-orbit torque (SOT) magnetoresistive random access memory (MRAM) crossbar arrays. By utilizing a resistance sum, the majority gate is implemented during READ operations in voltage-gated SOT-MRAM crossbar arrays, resulting in reduced read currents and improved energy efficiency. We employ a series of READ and WRITE operations to perform multiplier calculations, leveraging the fast READ and WRITE speeds of voltage-gated SOT-MRAM devices. Furthermore, the use of five-input majority gates simplifies multiplication by employing uniform logic gates and reducing logic depth, thereby lowering the operation’s complexity and the total number of occupied cells. Our experimental results demonstrate that the proposed in-memory Wallace tree multipliers consume three times less energy for in-memory operations than previously reported$4\times 4$multipliers. Moreover, the proposed method reduces the delay overhead from O ($n^{2}$) to O ($\log _{2}{n}$), where$\mathit {n}$represents the number of bits. Yajuan Hui, Qingzhen Li, Leimin Wang, Cheng Liu 0008, Deming Zhang, Xiangshui Miao |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2023 | A memristive neural network based matrix equation solver with high versatility and high energy efficiency
Jiancong Li, Houji Zhou, Yi Li 0049, Xiangshui Miao |
Sci. China Inf. Sci. | 4 |
| 2023 | A comprehensive study of device variability of sub-5 nm nanosheet transistors and interplay with quantum confinement variation
Haowen Luo, Xiangshui Miao, Xingsheng Wang |
Sci. China Inf. Sci. | 3 |
| 2023 | Two-dimensional materials-based integrated hardware
Zhuiri Peng, Runfeng Lin, Langlang Xu, Xiangxiang Yu, Xiaohan Meng, Xiangshui Miao |
Sci. China Inf. Sci. | 11 |
| 2023 | All-van der Waals stacking ferroelectric field-effect transistor based on In2Se3 for high-density memory
Zeyang Feng, Jingwei Cai, Xiangshui Miao |
Sci. China Inf. Sci. | 5 |
| 2023 | Modeling and physical mechanism analysis of the effect of a polycrystalline-ferroelectric gate on FE-FinFETs
Chengxu Wang, Zichong Zhang, Xiangshui Miao, Xingsheng Wang |
Sci. China Inf. Sci. | 5 |
| 2022 | Low-time-complexity document clustering using memristive dot product engine
Houji Zhou, Yi Li 0049, Xiangshui Miao |
Sci. China Inf. Sci. | 3 |
| 2022 | Complementary Memtransistor-Based Multilayer Neural Networks for Online Supervised Learning Through (Anti-)Spike-Timing-Dependent PlasticityabstractWe propose a complete hardware-based architecture of multilayer neural networks (MNNs), including electronic synapses, neurons, and periphery circuitry to implement supervised learning (SL) algorithm of extended remote supervised method (ReSuMe). In this system, complementary (a pair of n- and p-type) memtransistors (C-MTs) are used as an electrical synapse. By applying the learning rule of spike-timing-dependent plasticity (STDP) to the memtransistor connecting presynaptic neuron to the output one whereas the contrary anti-STDP rule to the other memtransistor connecting presynaptic neuron to the teacher one, extended ReSuMe with multiple layers is realized without the usage of those complicated supervising modules in previous approaches. In this way, both the C-MT-based chip area and power consumption of the learning circuit for weight updating operation are drastically decreased comparing with the conventional single memtransistor (S-MT)-based designs. Two typical benchmarks, the linearly nonseparable benchmark XOR problem and Mixed National Institute of Standards and Technology database (MNIST) recognition have been successfully tackled using the proposed MNN system while impact of the nonideal factors of realistic devices has been evaluated. Nuo Xu 0002, Bin Gao 0006, Fuwei Zhuge, Zijian Tang, Xinchen Deng, Yi Li 0049, Yuhui He, Xiangshui Miao |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2021 | Global Exponential Stability of Memristive Neural Networks With Mixed Time-Varying DelaysabstractThis article investigates the Lagrange exponential stability and the Lyapunov exponential stability of memristive neural networks with discrete and distributed time-varying delays (DMNNs). By means of inequality techniques, theories of the M-matrix, and the comparison strategy, the Lagrange exponential stability of the underlying DMNNs is considered in the sense of Filippov, and the globally exponentially attractive set is estimated through employing the M-matrix and external input. Especially, when the external input is not concerned, the Lyapunov exponential stability of the corresponding DMNNs is developed immediately in the form of an M-matrix, which contains some published outcomes as special cases. Furthermore, by constructing an M-matrix-based differential system, the Lyapunov exponential stability of the DMNNs is studied, which is less conservative than some existing ones. Finally, three simulation examples are carried out to examine the validness of the theories. Yin Sheng, Tingwen Huang, Zhigang Zeng, Xiangshui Miao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Memcomputing: fusion of memory and computing
Yi Li 0049, Yaxiong Zhou, Zhuorui Wang, Xiangshui Miao |
Sci. China Inf. Sci. | 4 |